Skip to content

Add prefill method support without generation - #183

Open
craftingmod wants to merge 13 commits into
JamePeng:mainfrom
craftingmod:prefill-feature
Open

craftingmod wants to merge 13 commits into
JamePeng:mainfrom
craftingmod:prefill-feature

Conversation

@craftingmod

@craftingmod craftingmod commented Sep 26, 2026 •

Copy link
Copy Markdown

Summary

  • Add create_chat_prefill method which could be used to prompt evaluation.
    • Expose a prefill-only chat API that evaluates and returns the final next-token without generation.
  • Add chat_template_kwargs support in create_chat_completion for overriding chat_template_kwargs each completion. : reverted

Maybe useful when implementing probability classification.

AI Disclose

I've used GPT 6 Luna to implementing features, but I manually reviewed and understood how to works.

@craftingmod

craftingmod commented Sep 26, 2026 •

Copy link
Copy Markdown
Author
  • Why i made this PR: I want trying something like system-one in comfyui 😕
  • Code quality might be bad so feel free to edit or close PR 😉

@JamePeng

JamePeng commented Oct 1, 2026

Copy link
Copy Markdown
Owner

I believe a PR shouldn't include changes unrelated to the submission's purpose; additionally, using grammar constraints might be a simpler way to control the model's selective output.

@craftingmod
craftingmod marked this pull request as draft October 1, 2026 14:56
@craftingmod

craftingmod commented Oct 1, 2026 •

Copy link
Copy Markdown
Author

I believe a PR shouldn't include changes unrelated to the submission's purpose; additionally, using grammar constraints might be a simpler way to control the model's selective output.

Additional modification is pushed.

  • Removed unrelated changes. (Added because prefill requires non-thinking and should be a way to disable thinking without new llama context. Correctly removed.)
  • Added Grammer parameter to create_chat_prefill, which could be used like LlamaGrammar.from_string('root ::= "Y" | "N"').

@craftingmod
craftingmod marked this pull request as ready for review October 1, 2026 17:03
@JamePeng

JamePeng commented Oct 2, 2026

Copy link
Copy Markdown
Owner

I can go ahead and merge this, but there are still several areas requiring fixes—such as the passing and handling of add_generation_prompt. Also, the handler already returns a PrefillResult containing an independent copy of the logits, yet create_chat_prefill() creates another PrefillResult after calculating the probabilities.

@craftingmod
craftingmod marked this pull request as draft October 2, 2026 07:29
@JamePeng
JamePeng marked this pull request as ready for review October 2, 2026 15:37
@JamePeng

JamePeng commented Oct 2, 2026

Copy link
Copy Markdown
Owner

https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp

ggml-org/llama.cpp#29831

It appears that work is underway on the underlying llama.cpp to adapt the execution logic for models similar to Jev.

@craftingmod

Copy link
Copy Markdown
Author

https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp

ggml-org/llama.cpp#29831

It appears that work is underway on the underlying llama.cpp to adapt the execution logic for models similar to Jev.

YAY

Seems should be wait until official API is introduced.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants