llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) - #29818
Conversation
| "LLaDAMoEModelLM": "llada", | ||
| "LLaDAModelLM": "llada", | ||
| "LLaMAForCausalLM": "llama", | ||
| "KevModel": "lev", |
There was a problem hiding this comment.
Is this correct "lev" instead of "kev" here?
There was a problem hiding this comment.
it's expected, they are quite similar so I grouped them into one file
ggerganov
left a comment
There was a problem hiding this comment.
I did some testing with OpenJEV and Laya and seems to work correct. Nice work!
| if not (dir_model / "config.json").is_file(): | ||
| config = ModelBase.load_hparams_guess(dir_model) | ||
| if config is not None: | ||
| return config |
There was a problem hiding this comment.
Are we sure this does not negatively affect any models?
There was a problem hiding this comment.
it should only affect models with load_hparams_guess, i.e. models that had config.json at all
now the load_hparams_guess expanded to models that already had config.json, but we want to use a derived class for conversion
btw, now having this, we should probably remove --mistral-format or at least somehow make these mistral-related flags more agnostic
There was a problem hiding this comment.
The models missing config.json was what concerned me, but I'm not sure which those are.
There was a problem hiding this comment.
only pockettts.py currently using it, which should be fine
|
Hi @ngxson, thanks for bringing TypeSafe API parity upstream. The parity table methodology is quite helpful. I've been testing Kev conversions independently (see fork in case it helps) and have quantized 0.8B/4B/9B GGUFs with embedded decision heads at espetro/kev-4b-demo-gguf (also 0.8b, 9b), produced with an end-to-end probe harness that measures argmax flips and probability drift vs the Python F16 reference. Couple of things that could help out:
One more thing: KevModel subclasses Qwen3_5TextModel, and the date_facts preprocessing is marked TODO - so Kev 1.0's Qwen3.8-based 27B model needs both a new arch and date_facts support. Is that on the roadmap? Happy to test. |
|
@espetro I don't quite get what you are talking about: merging LoRA to base is done at conversion time, on pytorch, there is no quantization there. if you are referring to your fork, that's completely unrelated to this PR date_facts is intentionally left out for simplicity. I don't plan to support it, it's just quite fragile: the date matching using regex, that won't work well on multi-language setup |
Clean upstream sync past 254b177. Relevant to the fork's deployment: - 4e2713c qwen4exp : optimize mask constructions (ggml-org#29824) - 631109b ggml : add alloc_buffer_n to the buffer type interface (ggml-org#23671) plus sycl / vulkan / opencl kernel work, llama warning and abort cleanups, and a new server /v1/systemone API (ggml-org#29818). Why it matters here: adopting upstream's Qwen4Exp MTP (ggml-org#29761) costs roughly +2.6 GB of compute buffers per device on the Qwen3.8-Flash-Next route, because the fixed attention path (ggml-org#29751) builds a k-pool input per QSA layer and makes the indexer cache track the attention cache cell for cell. ggml-org#29824 reduces the mask construction cost of that path. Fork-own work (RAM prompt-cache retention, selective CUDA P2P transport, server prefill/decode phase isolation, DFlash M-RoPE inference, docs, tests) is unchanged.
|
Will this support https://huggingface.co/Cloudflare/clef? Seems to be the best model of this class released so far. |
|
How to use decision models, aka system one endpoint, with the llama.cpp embedded model router: Adding just the info for anyone wondering how to use it with the model router (as I wondered 10 min ago): just add the "model": "< model name >" in the json, same level as "state". |
Overview
Support 5 decision models: laya, julia-1, lev, openjev (+vision), kev
Pre-converted GGUF:
Additional information
"Decision models" are just fancy wrappers around traditional embedding models (BERT/Qwen/etc), so changes to support them are kept to be as small and as self-contained as possible.
Since the underlay infra is just existing embedding models, changes to libllama are minimal. Main part is
server_decision_contextthat switches the input/output handling based on{arch}.decision.typegguf metadata.Recap changes per-component:
/v1/systemoneAPI, add the mainserver_decision_contexthandling logic, add tests (using a slice of laya model)TODO:
maybe support https://huggingface.co/Cloudflare/clef , just out 2 hours ago at the time writing thisplanned as a follow-upTesting
Table represent result in format: reference → llama.cpp
input json
{ "state": { "message": "Hi, I was charged twice for my order #4471 and I want a refund.", "plan": "pro", "order": { "id": 4471, "items": ["phone case", "charger"] } }, "questions": { "intent": { "type": "choice", "instructions": "What does the customer want?", "criteria": { "refund": "wants money back", "cancel": "wants to cancel an order", "track": "wants to know where an order is", "other": "anything else" } }, "urgent": { "type": "noul", "instructions": "Does this need a human within the hour?" }, "frustration": { "type": "score", "instructions": "How frustrated is the customer?", "criteria": ["calm", "mildly annoyed", "annoyed", "angry"] }, "refund": { "type": "noul", "instructions": "Is a refund requested?", "criteria": { "true": "money back is asked", "false": "no money back is asked" } }, "team": { "type": "choice", "instructions": "Which team?", "criteria": { "billing": null, "shipping": null, "technical": null, "sales": null, "legal": null, "returns": null, "fraud": null, "accounts": null, "retention": null, "other": null } } } }Requirements