Repository navigation
mediagen : LTX-2 image, video and audio generation - #28540
ericcurtin wants to merge 6 commits into
Conversation
llmman-manatee.mp4puppy_cuda.mp4 |
|
Hi @ericcurtin, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Opened #28541 as the RFC for this change, including a proposed split into smaller PRs. Happy to restructure based on the discussion there. |
|
/bot review |
Automated code reviewReview complete. Here is my review of PR #28540. Static review:
|
|
Build is failing on a flake |
e723ca2 to
7358429
Compare
There was a problem hiding this comment.
🟡 Changes recommended
Eight unresolved moderate issues affect correctness, resource safety, privacy, request limits, and routing.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Adds native LTX-2.3 image, video, and audio generation across libmediagen, CLI tooling, model downloads, and server APIs.
Changes:
- Introduces LTX diffusion, VAE, vocoder, prompt-enhancement, and encoding pipelines.
- Adds Hugging Face sidecar and text-encoder resolution plus mediagen CLI options.
- Exposes OpenAI-compatible image, video, and speech endpoints.
File summaries
| File | Description | Review |
|---|---|---|
tools/server/server.cpp |
Integrates mediagen routing and lifecycle. | Moderate: Router-mode video GET URLs lack the required model association. |
tools/server/server-mediagen.h |
Declares media endpoint state and handlers. | No issues identified. |
tools/server/server-mediagen.cpp |
Implements image, video, and speech endpoints. | Moderate: Missing aggregate generation limits permit excessive allocations; accepted long speech requests are silently truncated. |
tools/server/README.md |
Documents media APIs. | No issues identified. |
tools/server/CMakeLists.txt |
Links the server with libmediagen. | No issues identified. |
tools/mediagen/README.md |
Documents models, usage, and validation. | No issues identified. |
tools/mediagen/mediagen.h |
Defines the public mediagen API. | No issues identified. |
tools/mediagen/mediagen.cpp |
Implements loading, generation, and encoding. | Moderate: FPS is not validated; enhanced prompts are logged at INFO level. |
tools/mediagen/mediagen-ltx.h |
Defines LTX structures and interfaces. | No issues identified. |
tools/mediagen/mediagen-ltx-vae.cpp |
Implements video VAE decoding. | No issues identified. |
tools/mediagen/mediagen-ltx-text.cpp |
Implements text conditioning. | No issues identified. |
tools/mediagen/mediagen-ltx-graph.h |
Provides shared graph operations. | No issues identified. |
tools/mediagen/mediagen-ltx-enhance.cpp |
Implements prompt enhancement. | No issues identified. |
tools/mediagen/mediagen-ltx-dit.cpp |
Implements diffusion transformer inference. | No issues identified. |
tools/mediagen/mediagen-ltx-audio.cpp |
Implements audio decoding and bandwidth extension. | No issues identified. |
tools/mediagen/mediagen-impl.h |
Declares internal loading and backend utilities. | Moderate: Backend and scheduler resources leak when initialization fails. |
tools/mediagen/mediagen-common.cpp |
Implements weight loading and graph execution. | No issues identified. |
tools/mediagen/mediagen-cli.cpp |
Adds the llama-mediagen CLI. |
No issues identified. |
tools/mediagen/CMakeLists.txt |
Builds and installs mediagen targets. | No issues identified. |
tools/CMakeLists.txt |
Enables the mediagen subdirectory. | No issues identified. |
scripts/sync_vendor.py |
Adds stb image-write synchronization. | No issues identified. |
common/download.h |
Extends download plans for diffusion assets. | No issues identified. |
common/download.cpp |
Resolves diffusion models and sidecars. | Moderate: Distilled and fallback selection can choose a non-first split-model shard. |
common/common.h |
Adds mediagen configuration parameters. | No issues identified. |
common/arg.h |
Tracks the text-encoder download plan. | No issues identified. |
common/arg.cpp |
Adds mediagen arguments and download handling. | No issues identified. |
Review details
- Files reviewed: 26/27 changed files
- Comments generated: 8
- Review effort level: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
9317741 to
ac5464a
Compare
Repos with VAE / text projection sidecars are diffusion models: prefer the distilled transformer, pick the matching sidecars and fetch the text encoder of the model family. Adds common_params_mediagen and its arguments. Assisted-by: Claude
ac5464a to
cc782bd
Compare
New tools/mediagen library and llama-mediagen CLI for the LTX-2.3 audio-video diffusion transformer, with a Gemma 3 text encoder via libllama. Assisted-by: Claude
Serve /v1/images/generations, /v1/videos and /v1/audio/speech with libmediagen when the loaded model is a diffusion model. Assisted-by: Claude
cc782bd to
bc66033
Compare
|
The Copilot and bot review points are addressed in earlier pushes. @ggerganov @ngxson PTAL when you get a chance. Thank you! |
Overview
RFC: #28541
Native image, video and audio generation with the LTX-2.3 audio-video diffusion transformer, so that
works out of the box and serves OpenAI compatible
POST /v1/images/generations,POST /v1/videos(+GET /v1/videos/{id}/content) andPOST /v1/audio/speech.tools/mediagen: the DiT from GGUF, text conditioning from all hidden states of a Gemma 3 text encoder loaded with libllama, video VAE, audio VAE + vocoder + bandwidth extension (48 kHz), prompt enhancement with the LTX-2 system prompt, PNG/WAV/mp4 (ffmpeg) output,llama-mediagenCLI. Runs on the ggml backends, tested on Metal, CUDA and Vulkan.common:-hfrecognises diffusion repos (VAE / text projection sidecars), prefers the distilled transformer, downloads the matching sidecars and the text encoder of the model family. New arguments--vae,--audio-vae,--text-proj,--text-encoder[-hf],--no-diffusion-auto,--enhance-prompt.server: diffusion models are served by libmediagen; the text endpoints answer with a not-supported error,/v1/modelsand/propsreport the capabilities.Decoders and connectors validated against the reference PyTorch implementation (video VAE 1e-3 max abs diff, connectors 0.2%, mel 7e-5, 48 kHz audio 1.4% mean rel diff). See
tools/mediagen/README.md.Additional information
Two ggml issues worked around here and worth separate fixes:
ggml_conv_1dlays out its output incorrectly for a batch larger than one with more than one output channel, and the CUDA pad kernel mapsne1togridDim.y(limited to 65535).Requirements