Proposal: Media Studio — image and video generation on LibreChat's providers, billing and storage #16140
usnavy13
started this conversation in
Feature Requests & Suggestions
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
media-studio-demo.mp4
Image and video generation in LibreChat today happens through per-provider agent tools. Each tool owns its own credentials, request shape and output handling, results are not billed through the balance, video jobs cannot outlive the tool call, and native image output from Gemini (#6065) has no home in the message model. This proposal adds one shared media pipeline, exposed three ways: a Media Studio workspace, two agent tools, and native Gemini images in chat. The existing image tools are untouched.
The work is complete on a branch and running in a test deployment.
@librechat/agentspull request: feat: add native media port for Google image output and host model-invocation tracing agents#553What it adds
Media Studio. A workspace next to chat for creating and iterating on images and video. Each creation is a thread: the original stays immutable, edits and retries are separate turns, and results can be compared across models side by side. Presets save provider, model, parameters and uploaded reference images. The gallery is searchable, supports multi-select delete, and Data Controls gains "Delete all Studio creations". Temporary creations stay out of the library and expire on their own deadline. A result can be sent to a chat with "Use in chat", and any chat image has an "Open in Media Studio" action.
Agent tools.
media_generateandmedia_statusregister in the existing tool registry. Image requests wait a bounded time (default 120 s) and otherwise return a durable receipt; video requests always return a receipt the agent can check later. Every configured provider is reachable from agents through the same service Studio uses.Native Gemini images in chat. Gemini image models return ordered text and image parts on an ordinary assistant message. Images are stored as normal
image_generationfiles, and the private continuation signatures Gemini needs for follow-up turns are kept on the message and replayed on the next request. Chat keeps ownership of the model call and its billing.Providers. Adapters for OpenAI Images and video (gpt-image and DALL-E profiles), Azure OpenAI v1, OpenRouter image and video with capability discovery, Gemini, Vertex Veo, xAI, Black Forest Labs, Recraft, Krea, Runway, MiniMax, Alibaba, HeyGen, Sourceful, Atlas, Seed and Microsoft. An integration references an existing endpoint (
builtin,custom,vertex) to reuse its credentials, headers and user-provided key slot, or uses adirectconnection for providers without a chat endpoint. Model controls (size, quality, format, background, duration and so on) come from per-model capability profiles, so the form only shows what the selected model accepts.Billing. Media uses the existing balance and transactions settings. In balance mode each integration declares
maxCostUSD; credits are reserved before dispatch and settled through the Transaction ledger from provider-reported cost, frozen token rates or a completed-result estimate. With balance off, transactions record usage without a debit. Uncertain paid outcomes (a cancel that may have completed) stay as liabilities until reconciled instead of being retried. Native chat is billed once by chat and never enters the Studio ledger.Storage and retention. Originals are ordinary File records on the deployment's configured file strategy (local, S3, CloudFront, Azure Blob, Firebase) and follow the chat retention policy. Thumbnails and video posters are best-effort derivatives through ffmpeg, which the Dockerfiles now install. Deleting an original removes its derivatives; the shared expired-file sweep owns expiry.
Operator and admin. Everything sits behind
media.enabledplusinterface.mediarole permissions (use,create; USER off by default, ADMIN on). Surfaces (studio,chat,tools), queue and concurrency limits, worker leases, transfer caps, polling and title generation are all inlibrechat.yaml.Realtime. Studio and the sidebar receive snapshot-invalidation events through the existing event-transport interface, with polling as the recovery path for reconnects and deployments without cross-replica delivery.
One library across providers, with image and video results side by side.
How it fits
The pipeline reuses what LibreChat already has rather than adding parallel copies: provider keys and user-provided credentials, file strategies, balances and transactions, rate limiters, moderation, SSRF guards, Prometheus metrics, Langfuse tracing,
SystemCapabilities, role permissions and the event transport. New backend code is TypeScript inpackages/api/src/media, database contracts live inpackages/data-schemas, and shared types inpackages/data-provider. Withmedia.enabled: false(the default) no worker runs and no generation is accepted; a deployment that never enabled it sees no change.Related requests
direct/openai/v1connection; the receipt model avoids the blocking poll loop the community Sora PRs usegpt-image-1support #6592 maintainer position that image generation is delivered through agent tools — honoured: the tools stay, andmedia_generateis itself a registered toolScope limits
Vertex serves Studio and the tools but is not on the native chat path; native chat is Gemini Developer API only. Azure is supported through
directconnections only. Studio threads have no archive, pin, tag or public share. Not in this contribution: Bedrock Canvas and Reel, Imagen, agent-wide default reference sets (discussion #13724), zip download (#14486), watermarking (#10589), a timeline editor and standalone audio generation. These are named as follow-ups rather than half-built.Dependencies
Native chat output needs two additions to
@librechat/agents: an injectedNativeMediaPortthat the Google adapter calls to persist image parts before any output is emitted, and atraceModelInvocationAPI so hosts can trace model calls made outside a graph run with the SDK's own Langfuse lifecycle. Those are in the companion pull request above. The LibreChat branch consumes a build of that SDK branch until a release exists; the pull request will move to a released version.Trying it
Check out the branch, build, and enable one integration:
librechat.example.yamldocuments every field. The mock e2e suite has Studio, tools, video and native chat journeys that run without provider keys.Question for maintainers
Thumbnails and video posters are produced with ffmpeg, so the branch adds it to the official Docker images. Is that acceptable, or would you rather derivatives stay disabled unless the operator installs ffmpeg and points
media.assets.derivatives.ffmpegPathat it?All reactions