AI systems to production: inference optimization, GPU kernel and quantization work, and benchmark engineering, measured on our own RTX PRO 6000 Blackwell (SM120) and DGX Spark (GB10, SM121) hardware. Engagements are written-only, fixed price, 24-48h turnaround.
This is the engineering and benchmarking account of Conatus AI, founded by @Dev-Jahn.
-
vLLM #55571, Xid 13 crashes on an RTX PRO 5000 (SM120) serving an FP8 model on vLLM 0.28.0, where the reporter had already tied the symptom to a FlashInfer shared-workspace bug. From source, narrowed it to one condition: that FlashInfer 0.6.16.post3's autotuner had picked the cuDNN
bmm_fp8runner for the reporter's prefill shapes, the only runner that resizes that workspace, and gave a cache check for it. The reporter's autotune cache then showed cuDNN winning all four entries at the M=512 bucket those prefills map to. A follow-up corrected the first comment's claim that upgrading FlashInfer alone would not help: the 0.6.18 SM12x gate (flashinfer#4165) keeps the cuDNN runner out of that path on cuDNN below 9.23.1, and the workspace fix (flashinfer#4666) is merged to main and in no release; a later note added that the three FlashInfer packages have to move together. After the swap the cache names no cuDNN runner and a third clean 981-question arm ran without a fault, against three faults in the first ~1,500 questions with the FlashInfer FP8 kernel enabled; the reporter is running the upgraded stack. -
llama.cpp #28196, an RTX 5090 decode-speed report, over three rounds on an RTX PRO 6000. A dense-model control on the same card narrowed the cost specific to the qwen35 architecture against the residual a conventional dense model also pays, and the reporter combined that with his own WSL2 run to withdraw the title claim and rescope the issue to speculative decoding. Reading the loader then turned up a regression that would have wasted his next benchmark: between llama.cpp #28159 and #28173 the MTP layer loads with a KV head count of zero, so
--spec-type draft-mtpcannot start on a file carrying an MTP head. That was hit on Apple Metal with the 9B file and not run on CUDA; the fix first ships in b10749. His own Ollama blob then went from 67.7 tok/s undrafted to 119.8 at draft 4 (1.77x, acceptance 0.470), and an nsys pass showed the large FFN matrices running at 66 to 71% of the bandwidth bound at the five-column verify width against 84 to 87% single-column, while the lm_head matvec sits at 93 to 95% at either width. Raw logs, per-request JSON and the kernel table. -
vLLM #54906, a thinking budget ignored under MTP speculative decoding on Thor (SM110), checked on SM120 at the reporter's exact commit, where it is honoured both under MTP4 and in the no-speculation control. That narrows the field without clearing shared code, since our checkpoint loads as
modelopt_mixedrather than his top-level NVFP4. He asked whether two flags were worth instrumenting; the follow-up endorsed one, showed the other is set True by his owntop_kandtop_pregardless of the budget, and located the V2 gate on the first line ofThinkingBudgetState.apply. The thread continued through the reporter's allocator-history trace; on that evidence the latest round drops two writer candidates raised in earlier rounds and proposes one allocator-isolation start to separate the pointer and overrun cases, with the target's GDN warmup as the first place to look if the answer is overrun. -
vLLM #54945, FlashInfer CUTLASS NVFP4 fused-finalize nondeterminism reported on GB10, independently reproduced on SM120 on a different stock checkpoint,
nvidia/Qwen3-30B-A3B-NVFP4. Six identical temperature-0 requests return six distinct top-20 logprob signatures and three distinct texts; with the reporter's ownuse_fused_finalize=Falsethey come back bit-identical. The patched arm also pinned--moe-backend, so that pair is not a single-variable A/B; what the run adds is that the fault survives a change of checkpoint. -
vLLM #53960 / PR #53899: the Flash-Next PLE-offload TP=1 hang: triple-diagnosed the missing-worker state on GB10 (
ps, IPC socket, py-spy frames identical to both reporters), connected the issue thread to the five-line uniproc-executor fix the morning it landed upstream, and verified that fix on the same hardware with an overlay A/B where only the five lines differed. -
SGLang #36596 / #36599: GLM-5.3-Flash NVFP4 bring-up triage: surfaced the checkpoint-side ignore-list mitigation (verified by config diff across revisions) with the old-revision caveat, and pinned the NextN draft-quantization override as still present at the support PR's head, with the fix shape the existing resolver allows.
-
SGLang PR #36556: SM120 A/B of the QSA sparse-attention resolver-gate fix (base fails in the pip FA4 CuTe path, patched boots on identical flags), which also surfaced that #36545 is incidentally fixed by the same PR; the author subsequently expanded the PR scope to SM120/SM121.
-
FlashInfer #3170: SM121 / DGX Spark audit datapoint measured on a physical GB10: the bare-metal uv install path with timings, sm_121 backend resolution on the vLLM release, a 16/16 long-range recall matrix, and the release-vs-nightly XQA decode gate with exact code cites. Bare-metal guide.
-
SGLang #34895: A/B verification of the FP8 lm_head
weight_scalebug on an NVFP4 mixed-precision checkpoint, one commit before and after the fix on identical hardware and flags (single SM120, TP=1): pre-fix reproduces degenerate repetition with empty content, current main answers correctly, and no release up to v0.5.18 contains the fix. Error-index entry. -
vLLM #53748: cross-architecture confirmation that the MLA-decode Triton shared-memory overflow reported on GB10 also reproduces on workstation SM120 (same 101,376-byte per-block opt-in limit), with device budget tables and a
num_stagessweep narrowing the failing configuration. -
vLLM #49476: quantified the FlashInfer
flashinfer_b12xSM120 workspace on a 96 GB Blackwell against a marlin baseline on identical flags (11.47 GiB of extra reservation before the KV cache), which explains why the same config OOMs 16-32 GB cards duringprofile_run. Full repro and logs. -
FlashInfer #4549: independent 96 GB reproduction of the SM120 CUTLASS FP4 GEMM per-call workspace allocation, with the per-M table that bounds it to the decode and small-prefill regime (M 1-256, gone at M 384+).
-
vLLM #53787: independent reproduction attempt of a reported GDN prefill regression on SM120, with pinned image digests; the non-reproduction on large-VRAM SM120 was subsequently confirmed by the reporter on his own hardware.
-
vLLM #53788: SM120 data for the Helion default-on RFC: out-of-box kernel availability, autotuned configs with measured speedups vs the CUDA baseline (1.50x and 3.29x geomean), and a root-cause trace for one kernel's autotune failure.
-
vLLM #53775 / #53777: sm_120 shared-memory measurements and commit-graph verification of a reporter's build state.
-
Writing: dev.to/conatusai and conatusai.hashnode.dev
- blackwell-doctor: a zero-dependency probe that reports your NVIDIA Blackwell (sm_120 / sm_121) GPU, serving stack, and a stable matrix key for the exact cell you are running. On GB10 it reports unified memory correctly as of v0.1.1. Run it with
uvx --from git+https://github.com/jahnclawdmonet/blackwell-doctor blackwell-doctor(not on PyPI yet). - blackwell-serving-matrix: measured serving results on Blackwell hardware (which model/runtime/quantization/backend cells start, OOM, output garbage, and how fast), one JSON line per cell, including before/after rows for verified upstream fixes.
- Single-cell serving verification, 59 USD: one public model, one runtime, one quantization and one topology measured on Blackwell, handed back with the exact command and raw logs. conatus.jahn.ai/serving-check
- GPU inference stack benchmark and tuning, 349 USD: conatus.jahn.ai/ai-engineering
- Custom CUDA and Triton kernels, from 1,000 USD, quoted after a benchmark isolates the bottleneck
- Agent and LLM integration repair sprint, 99-199 USD by scope
Sample deliverables: benchmark report, SM120 workspace repro, and redacted repair handoff.
Contact: jahn.clawd.monet@gmail.com