Skip to content

llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) - #29818

Merged
ngxson merged 17 commits into
ggml-org:masterfrom
ngxson:xsn/decision
Oct 2, 2026
Merged

ngxson merged 17 commits into
ggml-org:masterfrom
ngxson:xsn/decision

Conversation

@ngxson

@ngxson ngxson commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

Support 5 decision models: laya, julia-1, lev, openjev (+vision), kev

Pre-converted GGUF:

llama-server -hf ggml-org/OpenJev-GGUF
llama-server -hf ggml-org/lev-GGUF
llama-server -hf ggml-org/Laya-GGUF
llama-server -hf ggml-org/Julia-1-GGUF
llama-server -hf ggml-org/Kev-4B-GGUF

Additional information

"Decision models" are just fancy wrappers around traditional embedding models (BERT/Qwen/etc), so changes to support them are kept to be as small and as self-contained as possible.

Since the underlay infra is just existing embedding models, changes to libllama are minimal. Main part is server_decision_context that switches the input/output handling based on {arch}.decision.type gguf metadata.

Recap changes per-component:

  • Conversion: handle conversion + new tensors/metadata + chat template injection
  • libllama: add decision head on graph, add some getters API
  • Server: add /v1/systemone API, add the main server_decision_context handling logic, add tests (using a slice of laya model)

TODO:

  • maybe support https://huggingface.co/Cloudflare/clef , just out 2 hours ago at the time writing this planned as a follow-up
  • add docs (+ dev docs for those who will add future model support)
  • support parallel shared prompt prefix (only for causal models)
  • support vision input

Testing

Table represent result in format: reference → llama.cpp

input json
{
  "state": {
    "message": "Hi, I was charged twice for my order #4471 and I want a refund.",
    "plan": "pro",
    "order": {
      "id": 4471,
      "items": ["phone case", "charger"]
    }
  },
  "questions": {
    "intent": {
      "type": "choice",
      "instructions": "What does the customer want?",
      "criteria": {
        "refund": "wants money back",
        "cancel": "wants to cancel an order",
        "track": "wants to know where an order is",
        "other": "anything else"
      }
    },
    "urgent": {
      "type": "noul",
      "instructions": "Does this need a human within the hour?"
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["calm", "mildly annoyed", "annoyed", "angry"]
    },
    "refund": {
      "type": "noul",
      "instructions": "Is a refund requested?",
      "criteria": {
        "true": "money back is asked",
        "false": "no money back is asked"
      }
    },
    "team": {
      "type": "choice",
      "instructions": "Which team?",
      "criteria": {
        "billing": null,
        "shipping": null,
        "technical": null,
        "sales": null,
        "legal": null,
        "returns": null,
        "fraud": null,
        "accounts": null,
        "retention": null,
        "other": null
      }
    }
  }
}
question laya julia-1 lev openjev kev-4b
intent (top option, refund) 0.9871 → 0.9870 0.9899 → 0.9910 0.8159 → 0.8209 0.9994 → 0.9994 0.6537 → 0.6550
urgent (noul) 0.1013 → 0.1016 0.6447 → 0.6232 0.4680 → 0.4681 0.2505 → 0.2507 0.7051 → 0.7045
frustration (top level) 0.3492 → 0.3492 0.6749 → 0.6895 0.2627 → 0.2633 0.5295 → 0.5295 0.4406 → 0.4407
refund (noul) 0.8916 → 0.8916 0.5608 → 0.5636 0.9503 → 0.9505 0.9901 → 0.9901 0.9472 → 0.9474
team (top option) 0.9666 → 0.9664 not run 0.3815 → 0.3787 0.9874 → 0.9874 0.4598 → 0.4599
worst diff, any probability 8.4e-4 2.2e-2 5.2e-3 1.8e-4 1.3e-3
input tokens 414 = 414 not reported by ref 576 vs 1194 619 = 619 406 = 406

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: most of the code is AI-written, I own the design

@ngxson ngxson changed the title llama, server: support decision models llama, server: support decision models (laya, julia-1, lev, openjev, kev-4b) Oct 1, 2026
@ngxson ngxson changed the title llama, server: support decision models (laya, julia-1, lev, openjev, kev-4b) llama, server: support decision models (laya, julia-1, lev, openjev, kev) Oct 1, 2026
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Oct 1, 2026
@ngxson ngxson changed the title llama, server: support decision models (laya, julia-1, lev, openjev, kev) llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) Oct 1, 2026
@ngxson
ngxson marked this pull request as ready for review October 1, 2026 19:01
@ngxson
ngxson requested review from a team, CISC and ggerganov as code owners October 1, 2026 19:01
@ggerganov ggerganov self-assigned this Oct 1, 2026
Comment thread conversion/__init__.py
"LLaDAMoEModelLM": "llada",
"LLaDAModelLM": "llada",
"LLaMAForCausalLM": "llama",
"KevModel": "lev",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this correct "lev" instead of "kev" here?

@ngxson ngxson Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's expected, they are quite similar so I grouped them into one file

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some testing with OpenJEV and Laya and seems to work correct. Nice work!

@ggerganov ggerganov added the highlight Changes that will be highlighted in the next release notes label Oct 2, 2026
Comment thread conversion/base.py
Comment on lines -1283 to -1286
if not (dir_model / "config.json").is_file():
config = ModelBase.load_hparams_guess(dir_model)
if config is not None:
return config

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we sure this does not negatively affect any models?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it should only affect models with load_hparams_guess, i.e. models that had config.json at all

now the load_hparams_guess expanded to models that already had config.json, but we want to use a derived class for conversion

btw, now having this, we should probably remove --mistral-format or at least somehow make these mistral-related flags more agnostic

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The models missing config.json was what concerned me, but I'm not sure which those are.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

only pockettts.py currently using it, which should be fine

@espetro

espetro commented Oct 2, 2026 •

Copy link
Copy Markdown

Hi @ngxson, thanks for bringing TypeSafe API parity upstream. The parity table methodology is quite helpful.

I've been testing Kev conversions independently (see fork in case it helps) and have quantized 0.8B/4B/9B GGUFs with embedded decision heads at espetro/kev-4b-demo-gguf (also 0.8b, 9b), produced with an end-to-end probe harness that measures argmax flips and probability drift vs the Python F16 reference. Couple of things that could help out:

  1. Footgun worth a line in the conversion docs: merging a decision head into an already-quantized base GGUF (tested with an unrelated Unsloth quant) flipped argmax on 10/17 questions, max |dp| 0.875. Heads need to be merged pre-quantization.
  2. If you want extra parity targets: across a 17-question probe suite, our imatrix quants show 0-1 argmax flips, max |dp| 0.051-0.140 (we focused on quantized-vs-F16-reference; different baseline than your F16 parity table, so not directly comparable).

One more thing: KevModel subclasses Qwen3_5TextModel, and the date_facts preprocessing is marked TODO - so Kev 1.0's Qwen3.8-based 27B model needs both a new arch and date_facts support. Is that on the roadmap? Happy to test.

@ngxson

ngxson commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

@espetro I don't quite get what you are talking about: merging LoRA to base is done at conversion time, on pytorch, there is no quantization there. if you are referring to your fork, that's completely unrelated to this PR

date_facts is intentionally left out for simplicity. I don't plan to support it, it's just quite fragile: the date matching using regex, that won't work well on multi-language setup

@ngxson
ngxson merged commit a4cb4c6 into ggml-org:master Oct 2, 2026
34 of 37 checks passed
novkien added a commit to novkien/llama.cpp-fork that referenced this pull request Oct 2, 2026
Clean upstream sync past 254b177. Relevant to the fork's deployment:

- 4e2713c qwen4exp : optimize mask constructions (ggml-org#29824)
- 631109b ggml : add alloc_buffer_n to the buffer type interface (ggml-org#23671)

plus sycl / vulkan / opencl kernel work, llama warning and abort cleanups,
and a new server /v1/systemone API (ggml-org#29818).

Why it matters here: adopting upstream's Qwen4Exp MTP (ggml-org#29761) costs roughly
+2.6 GB of compute buffers per device on the Qwen3.8-Flash-Next route, because
the fixed attention path (ggml-org#29751) builds a k-pool input per QSA layer and makes
the indexer cache track the attention cache cell for cell. ggml-org#29824 reduces the
mask construction cost of that path.

Fork-own work (RAM prompt-cache retention, selective CUDA P2P transport, server
prefill/decode phase isolation, DFlash M-RoPE inference, docs, tests) is
unchanged.
@kroaton

kroaton commented Oct 2, 2026

Copy link
Copy Markdown

Will this support https://huggingface.co/Cloudflare/clef? Seems to be the best model of this class released so far.

@paschembri

paschembri commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

How to use decision models, aka system one endpoint, with the llama.cpp embedded model router:

Adding just the info for anyone wondering how to use it with the model router (as I wondered 10 min ago):

just add the "model": "< model name >" in the json, same level as "state".

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion documentation Improvements or additions to documentation highlight Changes that will be highlighted in the next release notes model Model specific server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants