Skip to content

mediagen : LTX-2 image, video and audio generation - #28540

Open
ericcurtin wants to merge 6 commits into
ggml-org:masterfrom
ericcurtin:mediagen-ltx
Open

ericcurtin wants to merge 6 commits into
ggml-org:masterfrom
ericcurtin:mediagen-ltx

Conversation

@ericcurtin

@ericcurtin ericcurtin commented Sep 7, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

RFC: #28541

Native image, video and audio generation with the LTX-2.3 audio-video diffusion transformer, so that

llama-server -hf unsloth/LTX-2.3-GGUF

works out of the box and serves OpenAI compatible POST /v1/images/generations, POST /v1/videos (+ GET /v1/videos/{id}/content) and POST /v1/audio/speech.

  • tools/mediagen: the DiT from GGUF, text conditioning from all hidden states of a Gemma 3 text encoder loaded with libllama, video VAE, audio VAE + vocoder + bandwidth extension (48 kHz), prompt enhancement with the LTX-2 system prompt, PNG/WAV/mp4 (ffmpeg) output, llama-mediagen CLI. Runs on the ggml backends, tested on Metal, CUDA and Vulkan.
  • common: -hf recognises diffusion repos (VAE / text projection sidecars), prefers the distilled transformer, downloads the matching sidecars and the text encoder of the model family. New arguments --vae, --audio-vae, --text-proj, --text-encoder[-hf], --no-diffusion-auto, --enhance-prompt.
  • server: diffusion models are served by libmediagen; the text endpoints answer with a not-supported error, /v1/models and /props report the capabilities.

Decoders and connectors validated against the reference PyTorch implementation (video VAE 1e-3 max abs diff, connectors 0.2%, mel 7e-5, 48 kHz audio 1.4% mean rel diff). See tools/mediagen/README.md.

Additional information

Two ggml issues worked around here and worth separate fixes: ggml_conv_1d lays out its output incorrectly for a batch larger than one with more than one output channel, and the CUDA pad kernel maps ne1 to gridDim.y (limited to 65535).

Requirements

@ericcurtin
ericcurtin requested review from a team and ggerganov as code owners September 7, 2026 10:49
@ericcurtin

Copy link
Copy Markdown
Collaborator Author
llmman-manatee.mp4
puppy_cuda.mp4

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown

Hi @ericcurtin, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Large PR: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ericcurtin

Copy link
Copy Markdown
Collaborator Author

Opened #28541 as the RFC for this change, including a proposed split into smaller PRs. Happy to restructure based on the discussion there.

@github-actions github-actions Bot added documentation Improvements or additions to documentation server vendor labels Sep 7, 2026
@pwilkin

pwilkin commented Sep 7, 2026

Copy link
Copy Markdown
Member

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 7, 2026 •

Copy link
Copy Markdown
Automated code review

Review complete. Here is my review of PR #28540.


Static review: mediagen - LTX-2 image, video and audio generation

Scope: new tools/mediagen subsystem + common download/args + server integration + vendored stb_image_write.h. Checklists run: scope gate, security, approach/design, server, general. The RFC (#28541) is linked and AI disclosure is filled in - good. The change is large but coherent as one feature.

Blocking / correctness and security

(point 1) Unvalidated tensor dims in the GGUF/safetensors loaders can abort the process on a crafted file (tools/mediagen/mediagen-common.cpp). mg_gguf_read_header reads each ne as a raw uint64 cast to int64_t with no positivity check, and only validates d.ne[0] % blck == 0. A negative or huge dim flows into ggml_new_tensor() (GGML asserts on n_dims, not on ne), then into ggml_nbytes/alloc, where crafted values hit GGML_ASSERT/GGML_ABORT (e.g. the hash-set assert in ggml_backend_sched_alloc_graph) - a hard abort instead of the graceful MG_ERR path the rest of the loader uses. Same for the safetensors reader: shape entries are used unchecked; the mg_nelements product can overflow int64_t (UB) before the d.nbytes != expected check fires. Since the -hf sidecar resolution downloads these files and they are otherwise attacker-controlled, validate before creating tensors: every ne[i] > 0, total element count under a sane cap, and the nbytes product checked for overflow before the arithmetic.

(point 2) Router mode breaks the video content URL (tools/server/server.cpp). In router mode mediagen.post_videos is replaced by models_routes->proxy_post, but mediagen.get_video and mediagen.get_video_content keep their local handlers (empty videos list). A POST /v1/videos proxied to a child returns a job with content_url: /v1/videos/{id}/content; the client then GETs that on the router, which always 404s. Either proxy the two GET handlers too, or make the child-returned content_url point at the child.

(point 3) Text servers now answer the new endpoints with 500 (tools/server/server.cpp). The mediagen.* handlers are registered unconditionally, and in a plain text-model server mediagen.post_images_generations throws "media generation model not loaded", which ex_wrapper maps to a 500. Before this PR the path did not exist (404). Consider returning ERROR_TYPE_NOT_SUPPORTED (501/400) when ctx == nullptr, or registering the routes only in mediagen mode.

(point 4) Repo classification is heuristic and can hijack a regular server. is_diffusion_sidecar flags any repo containing a file whose name contains video_vae, _vae. or embeddings_connectors. For llama-server -hf <repo> this decides the entire server mode before the model is downloaded; a regular text/multimodal repo that happens to ship a matching filename will have its model selection switched to find_best_diffusion_model and then fail the post-download check "the downloaded model is not a supported diffusion model" and exit - a regression for that repo. At minimum log prominently which file triggered the diffusion detection; consider a stricter match (exact sidecar directory patterns) or a confirmation before switching modes.

Will slow the review

(point 5) std::vector<float> mean, std; in ltx_vae_decode (tools/mediagen/mediagen-ltx-vae.cpp) declares a local named std. It compiles today because nothing after it uses std::, but it shadows the namespace and is a trap for the next editor. The audio file already uses the better name stdv - rename here too.

(point 6) Graph capacity vs scheduler capacity mismatch. mg_backend::init(..., 24 * 1024) sizes the backend sched's max_nodes, but ltx_audio_decode builds mg_graph g(64 * 1024). If the audio graph ever exceeds 24k nodes, ggml_backend_sched_alloc_graph hits GGML_ASSERT(hash_set.size >= n_nodes + n_leafs) - abort, not a graceful failure. Derive one constant from the other (e.g. pass the sched capacity when constructing the graphs).

(point 7) Hardcoded VAE temporal scale in post_videos (tools/server/server-mediagen.cpp): gp.n_frames = (n_frames - 1) / 8 * 8 + 1; assumes vae_scale_t == 8. mediagen_generate re-fixes n_frames from the real hparam, so the behavior is benign, but the server-side rounding is silently wrong for any other model - drop it and let the library normalize.

(point 8) User tag interpolated into a regex unescaped (common/download.cpp, find_best_diffusion_model): std::regex pattern(t + "[.-]", ...) misbehaves for tags containing ., +, ( etc. Use a plain substring search (like find_best_model does) or escape the tag.

(point 9) Uncaught exceptions from the CLI. mg_weights::get throws std::runtime_error for missing tensors; in mediagen-cli.cpp nothing catches it (the server's ex_wrapper does). A mismatched sidecar file gives the user an abort instead of the intended error message. Catch in mediagen_init/mediagen_generate (or in main) and return the documented failure.

(point 10) Distilled schedule silently ignores --steps > 8 (ltx_get_sigmas): n_steps >= n returns the full 8-entry schedule, so --steps 20 quietly runs 8 steps. Warn or document.

(point 11) Resource bounds on the video path. MG_MAX_DIM 4096 x MG_MAX_FRAMES 1025 allows a client to request a ~51 GB raw RGB frame file (written to disk by the ffmpeg muxer) plus an enormous compute graph. Generation will fail, but only after a long stall on the single generation mutex. Consider a combined widthheightframes cap or a default much lower than the maximums.

(point 12) ltx_text_encoder depends on internal llama.cpp graph tensor names ("inp_scaled", "l_out-N") via cb_eval. This is an implicit contract with the build graph that future refactors of llama-build will break. It fails loudly (hidden state %d was not captured), which is good - please add a comment in mediagen-ltx-text.cpp stating the dependency so it is obvious where to look when it fires.

Nits

(point 13) Double blank line after llama_numa_init in mediagen-cli.cpp (leftover from editing).

(point 14) common_params_mediagen::empty() appears unused - remove or use it.

(point 15) res->data = v.mp4 in get_video_content copies the full mp4 on every request while it is already stored in the job; and the 8-entry videos ring has no total-size bound - fine for now, worth a comment.

(point 16) mediagen_encode_wav truncates data_size to uint32_t - harmless within the 60 s cap, but it is a public API; add a guard or a comment.

Notes

  • The two worked-around ggml bugs (ggml_conv_1d batch>1 output layout, CUDA pad kernel gridDim.y limit) are flagged in the description - please do split those fixes into their own PRs, since both workarounds (mel_stft per-channel loop, zero_stuff concat trick) cost clarity here.
  • Validation against the PyTorch reference is documented in the README with concrete tolerances - good. There is no automated coverage at all of the pure functions (ltx_rope_tables, ltx_get_sigmas, the safetensors/gguf readers); a small test for the loaders (valid + crafted-negative-dim headers) would directly guard point (point 1). Keep it minimal per the project's test guidance.
  • Style, naming, ASCII and comment conventions in the new code match the project rules; vendoring of stb_image_write.h is wired through scripts/sync_vendor.py correctly.

This review was generated automatically by pi coding agent using zai-org/GLM-5.3. It may contain mistakes. Maintainers make the final call.

@ericcurtin

Copy link
Copy Markdown
Collaborator Author

Build is failing on a flake

@ericcurtin
ericcurtin force-pushed the mediagen-ltx branch 3 times, most recently from e723ca2 to 7358429 Compare September 8, 2026 02:19
@ericcurtin
ericcurtin requested a balanced review from Copilot September 8, 2026 03:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Eight unresolved moderate issues affect correctness, resource safety, privacy, request limits, and routing.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Adds native LTX-2.3 image, video, and audio generation across libmediagen, CLI tooling, model downloads, and server APIs.

Changes:

  • Introduces LTX diffusion, VAE, vocoder, prompt-enhancement, and encoding pipelines.
  • Adds Hugging Face sidecar and text-encoder resolution plus mediagen CLI options.
  • Exposes OpenAI-compatible image, video, and speech endpoints.
File summaries
File Description Review
tools/server/server.cpp Integrates mediagen routing and lifecycle. Moderate: Router-mode video GET URLs lack the required model association.
tools/server/server-mediagen.h Declares media endpoint state and handlers. No issues identified.
tools/server/server-mediagen.cpp Implements image, video, and speech endpoints. Moderate: Missing aggregate generation limits permit excessive allocations; accepted long speech requests are silently truncated.
tools/server/README.md Documents media APIs. No issues identified.
tools/server/CMakeLists.txt Links the server with libmediagen. No issues identified.
tools/mediagen/README.md Documents models, usage, and validation. No issues identified.
tools/mediagen/mediagen.h Defines the public mediagen API. No issues identified.
tools/mediagen/mediagen.cpp Implements loading, generation, and encoding. Moderate: FPS is not validated; enhanced prompts are logged at INFO level.
tools/mediagen/mediagen-ltx.h Defines LTX structures and interfaces. No issues identified.
tools/mediagen/mediagen-ltx-vae.cpp Implements video VAE decoding. No issues identified.
tools/mediagen/mediagen-ltx-text.cpp Implements text conditioning. No issues identified.
tools/mediagen/mediagen-ltx-graph.h Provides shared graph operations. No issues identified.
tools/mediagen/mediagen-ltx-enhance.cpp Implements prompt enhancement. No issues identified.
tools/mediagen/mediagen-ltx-dit.cpp Implements diffusion transformer inference. No issues identified.
tools/mediagen/mediagen-ltx-audio.cpp Implements audio decoding and bandwidth extension. No issues identified.
tools/mediagen/mediagen-impl.h Declares internal loading and backend utilities. Moderate: Backend and scheduler resources leak when initialization fails.
tools/mediagen/mediagen-common.cpp Implements weight loading and graph execution. No issues identified.
tools/mediagen/mediagen-cli.cpp Adds the llama-mediagen CLI. No issues identified.
tools/mediagen/CMakeLists.txt Builds and installs mediagen targets. No issues identified.
tools/CMakeLists.txt Enables the mediagen subdirectory. No issues identified.
scripts/sync_vendor.py Adds stb image-write synchronization. No issues identified.
common/download.h Extends download plans for diffusion assets. No issues identified.
common/download.cpp Resolves diffusion models and sidecars. Moderate: Distilled and fallback selection can choose a non-first split-model shard.
common/common.h Adds mediagen configuration parameters. No issues identified.
common/arg.h Tracks the text-encoder download plan. No issues identified.
common/arg.cpp Adds mediagen arguments and download handling. No issues identified.
Review details
  • Files reviewed: 26/27 changed files
  • Comments generated: 8
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread common/download.cpp
Comment thread common/download.cpp
Comment thread tools/mediagen/mediagen-impl.h
Comment thread tools/mediagen/mediagen.cpp
Comment thread tools/mediagen/mediagen.cpp Outdated
Comment thread tools/server/server-mediagen.cpp Outdated
Comment thread tools/server/server-mediagen.cpp Outdated
Comment thread tools/server/server.cpp
@ericcurtin
ericcurtin force-pushed the mediagen-ltx branch 2 times, most recently from 9317741 to ac5464a Compare September 8, 2026 03:54
Repos with VAE / text projection sidecars are diffusion models: prefer the
distilled transformer, pick the matching sidecars and fetch the text encoder
of the model family. Adds common_params_mediagen and its arguments.

Assisted-by: Claude
New tools/mediagen library and llama-mediagen CLI for the LTX-2.3
audio-video diffusion transformer, with a Gemma 3 text encoder via libllama.

Assisted-by: Claude
Serve /v1/images/generations, /v1/videos and /v1/audio/speech with
libmediagen when the loaded model is a diffusion model.

Assisted-by: Claude
@ericcurtin

Copy link
Copy Markdown
Collaborator Author

The Copilot and bot review points are addressed in earlier pushes. @ggerganov @ngxson PTAL when you get a chance. Thank you!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation server vendor

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants