diff --git a/research/arxiv_tnf/SUBMISSION.md b/research/arxiv_tnf/SUBMISSION.md index 12189e0eb..ed1d5e62d 100644 --- a/research/arxiv_tnf/SUBMISSION.md +++ b/research/arxiv_tnf/SUBMISSION.md @@ -81,3 +81,21 @@ Kept in the paper where they were made: 2. *One trit per tripling is a new regime class* — it is takum's, off by an additive 1 3. *The gap is empty because an intermediate regime is too expensive* — it costs 18% more than posit's, which ships 4. *TNF64 will cost 4,762 LUTs* — it costs 7,479; the single power law was low by 36% + +## Journal submission record (added 2026-09-05) + +The section above predates the journal submission and its "nothing has been +submitted" no longer holds for the journal route. + +| event | date | manuscript | tree | +|---|---|---|---| +| submitted to *Microprocessors and Microsystems* | 2026-09-03 00:25 (+07) | MICPRO-D-26-00839 | last commit before submission `a0fb006`; `tnf_paper.tex` last changed in `64400a1` (2026-08-26) | +| revision R1 prepared, not yet submitted | 2026-09-05 | same number | this commit | + +R1 changes, all in `tnf_paper.tex`: a paragraph naming the repository and the +root each cited path is relative to (the `measurements/…` paths are relative to +this directory, not to the repository root); the Lloyd–Max paragraph states the +scale rule it used and no longer calls its row a ceiling; the "2026 literature +varies the transform" sentence now names the work that varies the scale field +(HBQ, M2XFP, VS-Quant, shared microexponents) and cites it; four `\bibitem`s +added. No number in any table changed. diff --git a/research/arxiv_tnf/tnf_paper.tex b/research/arxiv_tnf/tnf_paper.tex index e1bf93c21..ddb7f4754 100644 --- a/research/arxiv_tnf/tnf_paper.tex +++ b/research/arxiv_tnf/tnf_paper.tex @@ -869,6 +869,8 @@ \subsection{Four staircase forms, and only four}\label{sec:taxonomy} four class counts below sum to $53$: the total closes on $83$ and is asserted in the harness rather than trusted to the reader's addition. +\paragraph{Where the record files are.} Every path of the form \texttt{measurements/\ldots}, \texttt{research/\ldots} or \texttt{tools/\ldots} cited in this paper is a file in the public repository \texttt{github.com/gHashTag/trinity-fpga}. Paths beginning \texttt{measurements/} are relative to the paper's own directory, \texttt{research/arxiv\_tnf/}; paths beginning \texttt{research/} or \texttt{tools/} are relative to the repository root. The journal submission of 2026-09-03 was made from the tree at commit \texttt{a0fb006}, in which \texttt{tnf\_paper.tex} last changed at \texttt{64400a1}; the commit of each later revision is recorded in \texttt{research/arxiv\_tnf/SUBMISSION.md}. + \begin{table}[t]\centering \caption{The catalogue by staircase form. The parameter is the taper constant in the units natural to each form. Measured against each format's own usable range. @@ -3896,11 +3898,9 @@ \section{The block axis, and the half of it that is ours} Everything above concerns per-channel scaling. The other deployed axis shares one scale across a short block, and there our formats lose. It is worth stating -how thoroughly, because the answer is not a ranking but a ceiling. +how thoroughly. What follows is a ranking under one distortion measure, not a ceiling on perplexity: Lloyd--Max minimises squared error, and squared error does not order perplexity across codebook families on this table, so the Lloyd--Max perplexity is one more row rather than a bound. -Lloyd--Max on the within-block distribution --- $18{,}850{,}950$ values, block of -$32$ along the contraction axis, E8M0 shared scale as the MX specification -mandates --- gives the eight-level codebook minimising squared error. It is +Lloyd--Max on the within-block distribution gives the eight-level codebook minimising squared error. The distribution is $18{,}850{,}950$ values: every linear weight tensor of SmolLM2-135M except the output head, cut into blocks of $32$ along the contraction axis ($3{,}317{,}760$ blocks in full), each block divided by its own maximum, then a strided subsample of at most about $80{,}000$ values per tensor --- a sample of values, not of whole blocks (\texttt{research/block/block\_bound.py}). The shared scale is E8M0 as the MX specification mandates, assigned as $s=2^{\lceil\log_2(a_{\max}/\mathrm{top})\rceil}$ with $\mathrm{top}$ the codebook's largest magnitude ($6.0$ for E2M1), so that no element saturates. The specification's own alignment, $s=2^{\lfloor\log_2 a_{\max}\rfloor-2}$, permits saturation and gives $23.5380$ for MXFP4 on the same run; the convention is stated because every perplexity below moves with it (\texttt{research/block/MXFP4\_SCALE\_CONVENTION\_2026-08-11.md}). It is neither multiply-free nor a ladder; it is the best any eight magnitudes can do. The territory between a learned codebook and a fixed field layout is being occupied from the other side as well: parametric non-uniform codebooks for one- to @@ -5498,11 +5498,9 @@ \subsection{Where the block literature is not looking}\label{sec:blockrelated} current accelerators, which is why the stop condition for this work is stated against them and not against posit or takum. -What the 2026 literature then varies is almost entirely the \emph{transform} -applied before quantisation, not the format underneath it: learnable block-wise -optimisation for outlier resilience, and rounding fitted to the E2M1 -level set, augmented residual channels, and activation-sparsity coupling. Each -reports gains, and each holds the element format at E2M1. +Much of the 2026 literature varies the \emph{transform} applied before quantisation while holding the element format at E2M1: learnable block-wise +optimisation for outlier resilience, rounding fitted to the E2M1 +level set, augmented residual channels, and activation-sparsity coupling. Not all of it does. HBQ~\cite{hbq2026} refines the scale grid with a shift-and-add second level, M2XFP~\cite{m2xfp2026} ships a sub-binade scale set $\{1.0,1.25,1.5,1.75\}\cdot2^{E}$ per subgroup, and VS-Quant~\cite{vsquant2021} ablated 3-, 4- and 6-bit per-vector scales in 2021; the shared-microexponent framework~\cite{bdr2023} makes the level-1 scale width a design variable before pinning it. The scale field, then, is not a closed axis, and the claim below is about the element field only. That is exactly the variable Theorem~\ref{thm:optimal} addresses. E2M1 spends two of its three non-sign positions on an exponent, \emph{inside a block that already @@ -5834,12 +5832,10 @@ \subsection{Figures whose generating data is not in the repository} rows have records and the rows they are ranked above do not is not evidence of the ordering, and is presented here as illustrative. -Worse, and decisive for the release question: \textbf{no place-and-route log is in -the repository tree at all}. \texttt{git ls-files} finds zero \texttt{*.log} files; -the only one in the working tree is the \LaTeX{} log of this document. A -withdrawal note in \texttt{research/frontier/} states that the ladder frequencies +Worse, and decisive for the release question: \textbf{no place-and-route log behind +any frequency quoted in this paper is in the repository tree}. At commit \texttt{a0fb006}, \texttt{git ls-files} finds five \texttt{*.log} files, all under \texttt{measurements/pnr\_logs/} and all five seeds of one design (\texttt{e2m11\_add}) that no frequency here is sourced to; beyond those, the only log in the working tree is the \LaTeX{} log of this document. The note \texttt{research/frontier/WITHDRAWAL\_FMAX\_UNSOURCED\_2026-08-10.md}, which retracts itself in its own heading, states that the ladder frequencies rest on $298$ \texttt{nextpnr-xilinx} logs under \texttt{fpga/phiscale/}; that -directory holds $145$ files and none of them is a log. So even the twenty-four +directory holds $163$ files at the same commit and none of them is a log. So even the twenty-four sourced frequencies are sourced to a \emph{document that reports a measurement}, not to the tool output that produced it. Under the standard this paper applies elsewhere --- a claim is closed by a command, its inputs, its log and a hash --- the @@ -7907,6 +7903,14 @@ \section*{Disclosure} \bibitem{earplynch2026powers} B.~Earp-Lynch, S.~Earp-Lynch, O.~Kihel and P.~Tiebekabe, ``Powers as Fibonacci Sums,'' arXiv:2608.04445, 5~August 2026. \url{https://arxiv.org/abs/2608.04445} +\bibitem{hbq2026} Chen et al., ``HBQ: Hierarchical Block Quantization,'' MICRO 2026, arXiv:2609.00450, 2026. + +\bibitem{m2xfp2026} Hu et al., ``M2XFP,'' arXiv:2601.19213, 2026. + +\bibitem{vsquant2021} S.~Dai, R.~Venkatesan, H.~Ren, B.~Zimmer, W.~J.~Dally, and B.~Khailany, ``VS-Quant: Per-vector scaled quantization for accurate low-precision neural network inference,'' MLSys 2021, arXiv:2102.04503. + +\bibitem{bdr2023} B.~Darvish Rouhani et al., ``With shared microexponents, a little shifting goes a long way,'' ISCA 2023, arXiv:2302.08007. + \end{thebibliography} \end{document}