80+ free AI services for chat, image, video, voice & APIs (may sometimes include access to lead gen ai models for free)
-
Updated
Aug 18, 2026
80+ free AI services for chat, image, video, voice & APIs (may sometimes include access to lead gen ai models for free)
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
Very fast, accurate speaker diarization
Stable Audio LoRA Trainer of salty goodness
Implementation of the model "AudioFlamingo" from the paper: "Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities"
text-to-audio-latent-diffusion
Code to train a custom time-domain autoencoder to dereverb audio
Memory-optimized SongGeneration (v2 Large) for 16GB VRAM GPUs. Features 8-bit µ-law KV-caching, fused layers, and SDPA/Triton integration.
ACE Step 1.5 XL music generation with dynamics-preserving mastering. Part of the AEON Media Production family.
Safe, production-ready starter for voice cloning via SV2TTS (RTVC wrapper). CLI, tests, Docker, CI, pre-commit. No model weights included.
Guide to deploying neural networks in VST plugins, with a specific focus on embedded devices using the Elk Audio OS
AI 环绕声智能上混系统 — 基于 Demucs 深度学习模型,将普通立体声音频分离为人声/贝斯/鼓/乐器 4 音轨,智能路由到 5.1/7.1 多声道布局。支持 NVIDIA CUDA / AMD ROCm / Intel Arc 显卡加速,Gradio WebUI 一键操作,输出 WAV/FLAC/AAC 环绕声文件。
AI-powered voice chatbot that coaches salespeople using behavioral psychology principles from Cialdini, Voss, and Kahneman. Built with LFM2.5-Audio, Modal, Pinecone, and Streamlit.
🗣️ Audio AI: Your Audio & Video Transcription Powerhouse!
Swift client for the ElevenLabs API, generated from the official OpenAPI spec using swift-openapi-generator
Melo: The AI-Native Voice Agent. Your Intimate Voice Buddy!
⚡ Accelerate speaker diarization with Senko, processing 1 hour of audio in just 5 seconds on powerful hardware—boost your audio analysis efficiency.
Python scripts to handle a two way voice conversation with Anthropic Claude, using ElevenLabs, Faster-Whisper, and Pygame.
SoundX - AI-Native DAW (Digital Audio Workstation).
To associate your repository with the audio-ai topic, visit your repo's landing page and select "manage topics."