Record Submission: 1.1570 BPB - 73.7M Ternary U-Net + NeoMuon + 4x relu²MLP + Factored Tied Emb + Poly5 Softcap + YaRN2048 + 8192BPE + FP8QAT + Bitmask-LZMA + Stride-16 Sliding - #640
Conversation
… relu² 4xMLP FP8)
|
Really excellent work! |
|
This is incredible work! exactly what I hoped would happen when I submitted #139. The factored embedding and FP8 QAT for non-ternary params are really clever. Congrats on the record. |
… relu² 4xMLP FP8) (openai#640) Co-authored-by: Ciprian-Florin Ifrim <ciprian-florin.ifrim@Ciprians-Mac-Studio-M1-Max.local>
… relu² 4xMLP FP8) (#640) Co-authored-by: Ciprian-Florin Ifrim <ciprian-florin.ifrim@Ciprians-Mac-Studio-M1-Max.local>
… relu² 4xMLP FP8) (openai#640) Co-authored-by: Ciprian-Florin Ifrim <ciprian-florin.ifrim@Ciprians-Mac-Studio-M1-Max.local>
… relu² 4xMLP FP8) (openai#640) Co-authored-by: Ciprian-Florin Ifrim <ciprian-florin.ifrim@Ciprians-Mac-Studio-M1-Max.local>
… relu² 4xMLP FP8) (openai#640) Co-authored-by: Ciprian-Florin Ifrim <ciprian-florin.ifrim@Ciprians-Mac-Studio-M1-Max.local>
…6 (on PR openai#549 stack) Six orthogonal improvements on the current SOTA (PR openai#549, 1.1194 BPB): - Polynomial degree-5 softcap replacing tanh (from ternary PR openai#640) - Z-loss regularization (1e-4 * logsumexp^2) for sharper gradients - YaRN positional encoding for better long-context handling - zstd-22 compression (7MB artifact vs 16MB with LZMA-6) - Sliding eval stride=16 (4x more context overlap) - FlashAttention 3/2/SDPA graceful fallback Smoke-tested on 1xH100 (Modal): 890 steps, healthy convergence. Awaiting 8xH100 SXM verification for official scoring.
Community Credit NoteThis submission uses techniques pioneered by other participants in this competition without attribution. For the community record: U-Net skip connections — Already present in @integrate-your-mind's PR #289 (2026-03-21, titled "SmearGate + BigramHash + Int6 + SWA + U-Net Skips"), @gowtham0992's PR #295 (2026-03-21), and @skarakulak's PR #507 (2026-03-23) — all submitted before this PR on 2026-03-24. Neither the PR body nor the 55-page results document cites any of these prior submissions. The ternary quantization engineering, factored tied embedding, FP8 QAT, and the comprehensive ablation study are genuine original contributions. This note ensures the people whose foundational architectural work made this submission possible are recognized on the record. Open source runs on attribution. Community credit review by @MatoTeziTanka — The Agora. |
|
@MatoTeziTanka ???? I will simply copy paste my previous comment on a different PR I did, where you also posted the same thing, since you keep spamming:
If these useless AI comments continue, I will be reporting each one individually as unrelated to the PRs. Do stop. |
|
U-Net is indeed well-established — the note should have said "applied in this competition by" rather than implying novelty. I'll tighten that language. The note explicitly credited your ternary quantization, factored tied embedding, FP8 QAT, and ablation study as original work. |
… relu² 4xMLP FP8) (openai#640) Co-authored-by: Ciprian-Florin Ifrim <ciprian-florin.ifrim@Ciprians-Mac-Studio-M1-Max.local>
|
@CiprianFlorin-Ifrim maintainer repro note: we are doing a pass over merged leaderboard rows and I cannot currently reproduce this ternary record. What I tried:
Current result:
One remaining visible environment difference: your submitted log shows Python |
|
@cocohearts Hey Alex, sorry took a while to get back to you, been a bit busy. Can you tell me more info on what is the setup you are running? There are 2 documents first and second, with over 100 runs done, with logs, and most of the logs are closer to my 1.15 than to your 1.2280. Separately I have the V2 of this submission, available as PR #920 which is even better, with even more logs and the 3 best submitted for that PR. And then the unlimited run one PR #923 (which runs some different settings that increase ms per step slightly, so step 6000 from that run is not identical to step 6000 in the v1 or v2, it would be more like step 6000 from there is the same as step 6500 in v1/v2 as I could enable different features given it's the unlimited compute track) and nonetheless, it reached 1.10 bpb. So given I have over a hundred logs across 3 PRs are "closer" to my submission than the 1.228 you mentioned, there must be something wrong. Keep in mind this was tested on 8xH100 from both Hyperbolic and Runpod across multiple instances, so it wasn't some specific instance that was magical. Are you using the 8k tokenized dataset? There's a setup file in the PR that sets up the environment and downloads the 8k retokenized dataset from one of the participants that kindly shared it on Discord with everyone. Let me know if you need any support with this or if you'd like me to create a quick 8xH100 instance and record the whole process + run, so you can see it's proper. |
|
gh |
Record: 1.1570 BPB — 73.7M Ternary U-Net Transformer
BitNet b1.58 + 10L + NeoMuon + 4x relu² MLP + Factored Tied Embedding + Poly5 Softcap + YaRN 2048 + 8192 BPE + FP8 QAT + Base-3 LZMA + Stride-16 Sliding Eval
val_bpb: 1.1570 (3-seed mean sliding, std 0.0007) | 15.99 MB max artifact | 8×H100 SXM, 599s
The results document linked here and in my repo showcases all methods and sweeps applied to both Binary and Ternary Bitnets, which unfortunately are incompatible with many methods, such as Tversky Layers, EMA, Muon WD, LM Logit Head ranking and many more. Scaling ratios and applicable/rejected techniques can be useful for other submissions too.
Results (3 seeds, 8×H100 SXM)
Architecture
Key Techniques
Architecture
Training
Evaluation
Compression
Setup and Run
Full run command
Compliance