Voice pipeline for Cloudflare Agents -- continuous STT, TTS, streaming, and real-time audio over WebSocket.
The published package includes the complete Voice guide at docs/index.md.
Experimental. This API is under active development and will break between releases. Pin your version and expect to rewrite when upgrading.
npm install @cloudflare/voice| Export path | What it provides |
|---|---|
@cloudflare/voice |
Server-side mixins (withVoice, withVoiceInput), provider types, Workers AI providers, SFU utilities |
@cloudflare/voice/react |
React hooks (useVoiceAgent, useVoiceInput) |
@cloudflare/voice/client |
Framework-agnostic VoiceClient class |
Adds the complete voice pipeline: continuous STT, LLM turn handling, streaming TTS, interruption, and conversation persistence. When the transcriber reports speech start, the pipeline aborts active LLM/TTS work and tells the client to stop any queued playback so users can barge in before a final transcript is available.
import { Agent } from "agents";
import {
withVoice,
WorkersAIFluxSTT,
WorkersAITTS,
type VoiceTurnContext
} from "@cloudflare/voice";
const VoiceAgent = withVoice(Agent);
export class MyAgent extends VoiceAgent<Env> {
transcriber = new WorkersAIFluxSTT(this.env.AI);
tts = new WorkersAITTS(this.env.AI);
async onTurn(transcript: string, context: VoiceTurnContext) {
return "Hello! I heard you say: " + transcript;
}
}onTurn() can also return streaming text, including AI SDK stream values:
import { streamText } from "ai";
async onTurn(transcript: string, context: VoiceTurnContext) {
const result = streamText({
model: myModel,
instructions: "You are a helpful voice assistant. Keep replies short.",
messages: [
...context.messages,
{ role: "user", content: transcript }
]
});
return result.stream;
}context.messages contains completed conversation history before the current transcript. Append transcript exactly once when constructing the LLM request. The pipeline persists the transcript before onTurn() runs, so calling getConversationHistory() directly inside the hook returns stored history that includes the current transcript.
| Property | Type | Required | Description |
|---|---|---|---|
transcriber |
Transcriber |
Yes | Continuous per-call STT provider |
tts |
TTSProvider |
Yes | Text-to-speech provider |
| Method | Description |
|---|---|
onTurn(transcript, context) |
Required. Handle a user utterance. Return string, AI SDK stream, or AsyncIterable<string>. |
createTranscriber(connection) |
Override to create a transcriber dynamically per connection. |
onCallStart(connection) |
Called when a voice call begins. |
onCallEnd(connection) |
Called when a voice call ends. |
onInterrupt(connection) |
Called when user interrupts playback, either from client audio-level detection or model-detected speech start. |
beforeCallStart(connection) |
Return false to reject a call. |
onMessage(connection, message) |
Handle non-voice WebSocket messages (voice protocol is intercepted automatically). |
| Method | Description |
|---|---|
afterTranscribe(transcript, connection) |
Process transcript after STT. Return null to skip. |
beforeSynthesize(text, connection) |
Process text before TTS. Return null to skip. |
afterSynthesize(audio, text, connection) |
Process audio after TTS. Return null to skip. |
speak(connection, text)-- synthesize and send audio to one connectionspeakAll(text)-- synthesize and send audio to all connectionsforceEndCall(connection)-- programmatically end a callsaveMessage(role, content)-- persist a message to conversation historygetConversationHistory()-- retrieve conversation history from SQLite
STT-only mixin -- no TTS, no LLM. Use when you only need speech-to-text (e.g., dictation, transcription).
import { Agent } from "agents";
import { withVoiceInput, WorkersAINova3STT } from "@cloudflare/voice";
const InputAgent = withVoiceInput(Agent);
export class DictationAgent extends InputAgent<Env> {
transcriber = new WorkersAINova3STT(this.env.AI);
onTranscript(text: string, connection: Connection) {
console.log("User said:", text);
}
}import { useVoiceAgent } from "@cloudflare/voice/react";
function App() {
const selectedSpeakerId = "default";
const {
status, // "idle" | "listening" | "thinking" | "speaking"
transcript, // TranscriptMessage[]
interimTranscript, // string | null (real-time partial transcript)
turnMetrics, // VoiceTurnMetrics | null (latest stable terminal summary)
audioLevel, // number (0-1)
isMuted, // boolean
connected, // boolean
error, // string | null
outputDeviceError, // string | null
startCall, // () => Promise<void>
endCall, // () => void
toggleMute, // () => void
sendText, // (text: string) => void
sendJSON // (data: Record<string, unknown>) => void
} = useVoiceAgent({
agent: "my-agent",
// Route assistant playback to a selected audiooutput device when supported.
outputDeviceId: selectedSpeakerId,
// Set false to delay connecting until async prerequisites are ready.
enabled: true
});
return <div>Status: {status}</div>;
}When enabled is false, the hook does not create or connect a VoiceClient, returns the idle/disconnected state, and action callbacks such as startCall(), sendText(), and sendJSON() are safe no-ops. The first change from disabled to enabled connects with the current options without firing onReconnect; later connection identity changes while enabled do fire onReconnect.
outputDeviceId accepts a MediaDeviceInfo.deviceId from an audiooutput device. Browsers without HTMLMediaElement.setSinkId() support continue playing through the default output and set outputDeviceError for non-default devices. Use "default" or undefined to return to the system default output. Device labels may be blank until the user grants microphone permission.
For voice input only:
import { useVoiceInput } from "@cloudflare/voice/react";
const {
transcript,
interimTranscript,
turnMetrics,
isListening,
start,
stop,
clear
} = useVoiceInput({ agent: "DictationAgent" });import { VoiceClient } from "@cloudflare/voice/client";
const client = new VoiceClient({ agent: "my-agent" });
const selectedSpeakerId = "default";
client.addEventListener("statuschange", () => console.log(client.status));
client.connect();
await client.startCall();
// Switch assistant playback without reconnecting the call.
await client.setOutputDevice(selectedSpeakerId);All default providers use Workers AI bindings -- no API keys required:
| Class | Type | Workers AI model | Recommended for |
|---|---|---|---|
WorkersAIFluxSTT |
Continuous STT | @cf/deepgram/flux |
withVoice |
WorkersAINova3STT |
Continuous STT | @cf/deepgram/nova-3 |
withVoiceInput |
WorkersAITTS |
TTS | @cf/deepgram/aura-1 |
Both |
WorkersAIFluxSTT uses Flux StartOfTurn events for low-latency barge-in and EndOfTurn events for final utterances. Custom transcribers can provide the same behavior by calling onSpeechStart from TranscriberSessionOptions when user speech begins, then onUtterance when the turn is complete.
| Package | What it provides |
|---|---|
@cloudflare/voice-assemblyai |
Continuous STT (AssemblyAI Universal 3.5 Pro Realtime) |
@cloudflare/voice-deepgram |
Continuous STT (Deepgram Nova) |
@cloudflare/voice-elevenlabs |
Continuous STT and TTS (ElevenLabs) |
@cloudflare/voice-telnyx |
Continuous STT, TTS, and phone transport (Telnyx) |
@cloudflare/voice-twilio |
Telephony adapter (Twilio Media Streams) |
examples/voice-agent-- full voice agent example with provider togglesexamples/voice-input-- voice input (dictation) example