Free, open-source curriculum for making money with generative AI image, video, and audio — for creators and agencies.
-
Updated
Aug 21, 2026
Free, open-source curriculum for making money with generative AI image, video, and audio — for creators and agencies.
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.
Full-featured DAW, DJ, and VJ app interoperable w/Ableton, Reaper, Resolume & more. Stable Audio 3, Magenta RT2, Suno API, Chimera track fusion, Demucs stems, MIDI generate/notate, img > spectrogram > music, drawing > music, VST3 & .gan plugins, automix & key-lock, GLSL shaders, volumetric video, Quest 3 XR interface, MIDI auto-map, RAG assistant
Draft to Take beta: local-first AI audio production studio powered by IndexTTS2, Docker, Qwen, OmniVoice, SFX, ambience, and music sidecars.
Free real-time AI Noise Gate VST3/AU plugin. Removes coughs, sneezes, and other artifacts from your live streams, podcasts, and videos.
ARCHIVED USE theDAW at https://github.com/gantasmo/theDAW - Browser-based AI audio DAW for Stable Audio 3 with text-to-audio, inpainting, LoRA training, FFmpeg effects, waveform editing, sequencer, piano roll, and persistent library.
Real-Time Deepfake Pipeline
Local-first CLI that turns Markdown scripts into multi-speaker podcast-style audio using Coqui XTTS v2.
AI-powered audio manipulation studio — sample pack creation, algorithmic composition, text-to-audio generation, and ChatGPT on one screen
Community list of AI tools for audio and music
Music Generation Using Deep Learning🎶🎵
AI Audio Content Creation Platform for Podcasts, Narration, Voice Generation and Audio Production. Create Professional Audio Content from Text with Modern Audio Workflows.
Windows GUI for building better Kokoro TTS voice outputs. Combines Kokoro voice random-walk search, target-audio scoring, RVC model training, and automatic post-generation voice conversion into a single queue-based workflow. Bootstraps its own Python environment from one executable.
Aurora Audio Studio 1.4.1:面向 Windows 的本地 AI 音频工作台,统一音乐生成、声音克隆、歌声转换、智能分轨、MIDI 扒谱与视频字幕。
AI Voice Agents: Exploring the Next Generation of Human-Machine Interaction! 🎙️🤖🎧
High-performance KittenTTS API server with a built-in web UI, OpenAI-compatible routes, long-form text support, and optional CUDA acceleration.
A local-first EPUB reader with high-fidelity neural text-to-speech, word-level synchronization, and Next.js/FastAPI/ONNX stack.
AudioInsight is a web application that processes audio, generates transcriptions, and allows users to ask questions about the related audio.
To associate your repository with the ai-audio topic, visit your repo's landing page and select "manage topics."