What features would you like to see added?
Audio should play progressively while chunks are being received
More details
Problem
When using TTS (text-to-speech) with streaming-enabled backends (like edge-tts), audio data is transferred in chunks (chunked transfer encoding), but LibreChat waits for the entire transfer to complete before starting playback.
Expected Behavior
Audio should start playing as soon as the first chunk is received, not after all chunks are transferred. This is especially important for longer texts where the delay is noticeable.
Current Implementation
In client/src/hooks/Input/useTextToSpeechExternal.ts:103-106:
onSuccess: async (data: ArrayBuffer, variables) => {
const audioBlob = new Blob([data], { type: 'audio/mpeg' });
autoPlayAudio(blobUrl);
The ArrayBuffer received is the complete audio even if the backend streams it. The Blob is only created after the entire response is received.
Suggested Solution
Replace the useTextToSpeechMutation approach with a streaming-capable fetch that feeds audio chunks to HTMLAudioElement as they arrive. For example:
- Use fetch() with response.body.getReader() to read chunks progressively
- Create a MediaSource or use a progressive HTMLAudioElement with a readable stream
- Start playback after the first chunk is available
Environment
- LibreChat version: latest
- TTS Provider: openai-edge-tts (or any streaming-capable TTS backend)
- Backend streaming is confirmed working (chunked transfer encoding verified)
Which components are impacted by your request?
No response
Pictures
No response
Code of Conduct
What features would you like to see added?
Audio should play progressively while chunks are being received
More details
Problem
When using TTS (text-to-speech) with streaming-enabled backends (like edge-tts), audio data is transferred in chunks (chunked transfer encoding), but LibreChat waits for the entire transfer to complete before starting playback.
Expected Behavior
Audio should start playing as soon as the first chunk is received, not after all chunks are transferred. This is especially important for longer texts where the delay is noticeable.
Current Implementation
In
client/src/hooks/Input/useTextToSpeechExternal.ts:103-106:The ArrayBuffer received is the complete audio even if the backend streams it. The Blob is only created after the entire response is received.
Suggested Solution
Replace the useTextToSpeechMutation approach with a streaming-capable fetch that feeds audio chunks to HTMLAudioElement as they arrive. For example:
Environment
Which components are impacted by your request?
No response
Pictures
No response
Code of Conduct