Repository navigation
feat(fpga): pipeline the phi scale and take the cycles back - #592
Merged
Merged
Conversation
FMAX.md measured the iterative phi scale honestly and the answer was bad: smaller and faster on the clock, and 2.15x SLOWER per element at the |k| = 8 a deployed layer needs. That was the implementation, not the lattice. The closed form was rejected on purpose: phi^k = F(k-1) + F(k)*phi makes the scale two multiplications by Fibonacci constants, which puts back the operator this line exists to remove. scale_phi_pipe.v unrolls the STEP instead -- K_MAX stages, one adder and one mux each, latency K_MAX, one element per cycle, no multiplier. Verified before measured: 40 elements back to back, 0 errors, negative control catches 38 of 40 when the adder is broken by one. Measured, nextpnr-ice40 hx8k, N=8 ACC=16: phi pipelined 660 LC 204.08 MHz 204.08 M elem/s multiplier 1098 LC 69.21 MHz 69.21 M elem/s 1.66x smaller and 2.95x faster on clock AND throughput. The throughput ratio is 2.95x and not 5.90x: the multiplier arm has two cycles of LATENCY and still accepts one element per cycle. Counting latency as cycles-per-element would have doubled the claim. The first version of the pipeline synthesised to 2 logic cells for an eight-stage 16-bit design -- two always blocks drove the same array and yosys pruned nearly all of it. A fold reports a number far too small rather than an error.
gHashTag
added a commit
to gHashTag/t27
that referenced
this pull request
Aug 18, 2026
The combined form [[.check_runs[]|select(.status!="completed")]|length, [...]|length]|@TSV failed with 'expected an object but got: array' the first time --wait ran, on the very PR it was added to gate. A wait loop that errors out is worse than no wait loop: it turns a gate into an interruption. Two plain queries instead. Verified against gHashTag/trinity-fpga#592.
This was referenced Aug 18, 2026
gHashTag
added a commit
to gHashTag/t27
that referenced
this pull request
Aug 19, 2026
…ea (#2223) * Make tri build from a clean checkout, add rtl check, build it in CI (Refs #2176) cargo build -p tri failed on main: dlc10 embeds a bitstream with include_bytes! that was committed to a feature branch and never to main. Nothing caught it because none of this repository's 35 workflows built the crate, while an autonomous loop kept committing into cli/tri/src/. tri rtl check runs the five structural checks t27.ai offers, locally. It names the yosys version beside the cell count, because that number is version-dependent: 0.33 reports 45 cells where 0.65 reports 49 on the same design, while wires and flip-flops agree exactly. Closes #2222 * Add tri gates: find workflows that have never once succeeded (Refs #2176) Eighteen active workflows across three repositories have never produced a green run, consuming 8182 runs between them. The largest are 1705, 1541 and 1224 runs with zero successes. This was a hand-run loop of gh api calls three times before it became a command. It reports and disables nothing: choosing between fix, workflow_dispatch-only and delete belongs to whoever owns the workflow. Verified against t27: 4 workflows, 1565 runs, matching the manual count. Closes #2222 * Add tri red: what is failing on the default branch right now (Refs #2176) Written because of a specific failure of mine, not a hypothetical. The publisher for t27.ai failed six consecutive times, the site served hours-old content, and site-live-gate.yml caught it correctly and went red five times on my own commits. I read it zero times and found the outage by accident when a page I had just published returned 404. Detection was never the problem. Reading it was. Measured on trinity: 27 workflows red on main, several since February. A saturated streak prints as '30+' rather than '30', because the count is bounded by the page size — the same silent truncation this command exists to surface, which appeared in the command itself first. Closes #2222 * feat(tri): add `tri mutate` — find constants a checker never checks Two vacuous checks turned up within an hour of each other. A workflow step read an 8-digit date out of `yosys -V`, which contains no date, so the guard skipped and the step passed without comparing anything. A verifier had a lookup table on both sides of the identity it tested, so flipping an entry cancelled with itself and left every claim green. Neither was found by reading. Both were found by changing a constant and noticing nothing went red. This automates that: perturb one numeric literal at a time, re-run the checker, report the ones it did not notice. Refuses to run on a file with uncommitted changes -- it rewrites the file and restores it, and an interrupted run must be recoverable with git. Dogfooded on the verifier that prompted it: 28 survivors, all but one in prose. Masking comments and strings cut that to 17, and the one real finding underneath was a quantiser compared only against itself, which would have made the whole inertness claim pass vacuously. Columns are reported because two identical literals on one line are otherwise indistinguishable; I read such a report, hand-checked the wrong one, and wrongly concluded the tool was lying. Closes #2222 * fix(tri): verify the restore instead of assuming it A measurement taken against a file this command left perturbed is not a measurement. That is not hypothetical: a perturbed constant survived a hand-run mutation on this machine, was read back as if it were the real value, and produced a written-up finding that did not exist -- an 'unreachable top binade' that the clean file does not have. Each mutant now re-reads the file after restoring it and fails loudly with the git command to recover, rather than continuing quietly. Closes #2222 * fix(tri): back the file up instead of demanding a clean tree The clean-tree requirement was safe and was the wrong trade. Every mutation run on work-in-progress needed a throwaway commit first, and a throwaway commit is how a 'wip-for-mutation' subject reached a repository whose format gate rejects it -- twice in two iterations, the second time an hour after I wrote the lesson down. A sibling backup gives the same recovery guarantee without asking the caller to commit anything. Git cleanliness is still reported, because git checkout is the nicer recovery path when it is available, and the backup is removed on success. Verified on a deliberately dirty file: the run completes, the file comes back byte-for-byte, and no backup is left behind. Closes #2222 * fix(tri): clear derived bytecode caches, not just the source Verifying the source came back byte-for-byte was not enough. Most mutations here preserve the file's LENGTH -- 5 becomes 6, 16 becomes 17 -- and Python decides a .pyc is current by comparing the source's (mtime, size). Restore inside the same filesystem second and both match, so the interpreter serves bytecode compiled from the mutant. Measured, not theorised: a benchmark's format table read back as e5m11 and then e6m10 on consecutive runs while the file on disk said e5m10 both times. Clearing __pycache__ made every assertion pass. Second contamination of this kind in two iterations -- first the file itself, now its cache. The restore has to reach derived artefacts or the next measurement in that session is against a mutant nobody can see. Closes #2222 * feat(tri): add `tri pr ready` — is this PR actually safe to merge? Written after merging a pull request whose language audit was red. The failure was in the list and I read a summary line I had written myself instead of the list. The gate was correct; I was not. The judgement that matters is not 'is anything failing' -- in these repositories something always is -- but 'is anything failing HERE that is not already failing everywhere else'. So it classifies every failure against the default branch and the last few merged PRs, and ends with one unambiguous verdict line. Pending checks produce WAIT rather than a verdict: a verdict computed from a partial list reads exactly like a complete one. Checked against the two real cases: #588 (safe) and #832, the one I merged red -- it names `checks` as the single failure unique to that PR and says DO NOT MERGE. Closes #2222 * feat(tri): add `tri synth area` — area with the instrument named Written after getting the same task wrong twice in one session, both times producing a plausible number instead of an error. First: the stat parser matched `Number of cells:`, which yosys 0.33 prints and 0.65 does not. It found nothing and reported zero, and zero looks exactly like a small design. Second: the yosys output was captured into a shell variable and printed with `echo`. zsh's echo interprets backslash escapes and yosys writes identifiers with a leading backslash, so `\tern_node` became a tab and a `\c` truncated the stream. Every count came back zero again, from data that was correct on disk. Now the counts are parsed by cell NAME (which did not change between versions), both streams are read, the yosys banner is printed beside the numbers, and a run that never reaches a stat block is an error carrying yosys's own message rather than a row of zeros. Reproduces the hand-measured layer numbers exactly: 4299/0/254/220 and 3906/3/123/204. It also states what it does not measure. `ltp` is not a frequency substitute: it counts topological hops in netlists whose structure differs between designs, so it does not compare across them -- measured today, where it reported 213 hops for a one-adder scale path and 20 for a 32x16 multiplier. Closes #2222 * fix(tri): let `pr ready` wait, because the verdict raced the checks I merged a pull request while this command's own output said WAIT -- ten checks still running -- because the merge was in the same batch as the readiness check, so the verdict gated nothing. The polling loop that fed it had its own bug: it counted rows matching 'pending' and exited on zero. Zero rows also means the checks have not STARTED, which is exactly the state it hit. So: --wait blocks until the list is both non-empty and quiet, and an empty list is treated as 'not started' for four rounds before being believed. The verdict can no longer be computed against a list that has not appeared yet. Nothing broke on main this time. That was luck, not process. Closes #2222 * fix(tri): split the wait-loop jq, which errored on first use The combined form [[.check_runs[]|select(.status!="completed")]|length, [...]|length]|@TSV failed with 'expected an object but got: array' the first time --wait ran, on the very PR it was added to gate. A wait loop that errors out is worse than no wait loop: it turns a gate into an interruption. Two plain queries instead. Verified against gHashTag/trinity-fpga#592. Closes #2222 * feat(tri): add `tri sweep area` — catch a synthesis fold by its shape Written after two folds in one afternoon, both of which reported a number rather than an error. A layer whose operand memory had no write port: yosys propagated the never-written memory as constant and pruned the design. 12 logic cells for a 64-tap multiplier layer, and area flat across fan-in. The same layer with one memory location per lane: reading N per cycle is N read ports, so muxes rather than block RAM. 23052 cells on a 7680-cell part, zero RAM inferred. Neither is an error and both were caught only because a human found the number implausible. This checks the shape: sweep the parameter and say plainly when the area does not respond to it, or responds far too steeply. Both real cases are unit tests, with their real numbers. It stays quiet on proportional growth, because a detector that fires on the ordinary case stops being read. Closes #2222 * fix(tri): let the wait loop survive a transient API failure The first time --wait met a TLS handshake timeout it propagated the error, the caller's merge ran anyway, and the gate protected nothing -- the third time in this project a verdict has failed to gate. Now a failed poll is retried up to five times before giving up, and a failed FINAL read makes the verdict WAIT rather than an invented clean list. Unknown is not zero. Closes #2222 * fix(tri): give each command back its own description Inserting a new variant used `Red {` as the anchor, which put the new command BETWEEN Red's doc comment and Red itself. Clap then read that comment as the new command's description and left Red with none: mutate What is failing on the default branch right now, and since when. Find the constants in a checker that nothing checks red (blank) Four commands were added against that same anchor this session, so the help text has been wrong since the first of them. Nothing behaved differently -- it just described itself incorrectly, which for a set of commands whose whole point is being readable is not a small thing. Closes #2222 * feat(tri): let `pr ready --merge` perform the merge itself The verdict cannot gate anything if the caller puts `gh pr merge` in the same batch as this command. It prints WAIT, the merge runs anyway, and nobody reads the line. That happened four times in one session -- including on the pull request that added the waiting mode, and again on the one that made the wait survive network failures. Each time the fix addressed the tool and the next failure came from the same place: the human batching past it. So the action moves inside the verdict. With --merge, the command merges on 'safe to merge' and refuses on WAIT or DO NOT MERGE, telling the caller to re-run with --wait. There is no longer a gap to batch through. Closes #2222 * docs(now): the CLI wave's NOW entry (Closes #2222) * ci(tri): build --all-targets -- the plain build hid two master test-build breaks (Closes #2222)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#591 measured the iterative φ scale honestly and the honest answer was bad: smaller and faster on the clock, and 2.15× slower per output element at the
|k| = 8a deployed layer needs.That was a property of the implementation, not of the lattice. This removes it.
Why not the closed form
φ^k = F(k−1) + F(k)·φ, so scaling(a,b)byφ^kisa' = a·F(k-1) + b·F(k),b' = a·F(k) + b·F(k+1)— two multiplications by Fibonacci constants. A barrel built that way puts back the operator this whole line exists to remove, so it was not built.What was built
scale_phi_pipe.vunrolls the step: K_MAX stages, stage j applying(a,b) → (b, a+b)whenj < kand passing through otherwise. One adder and one mux per stage, latency K_MAX, one element per cycle, no multiplier.Verified before measured — 40 elements fed back to back against an independent golden model:
Negative control: breaking the adder by one is caught on 38 of 40.
Measured — nextpnr-ice40, hx8k ct256, N=8, ACC=16
1.66× smaller and 2.95× faster, on clock and throughput.
The ratio is 2.95× and not 5.90×. The multiplier arm has two cycles of latency and still accepts one element per cycle —
prodis combinational from a registered accumulator and nothing gates the input. Counting its latency as cycles-per-element would have doubled the claim, and I checked the RTL before quoting it.Cost against the iterative version
3.6× the scale-block area for k× the throughput. At the deployed
|k| = 8that is a good trade; below k = 4 the iterative version delivers the same rate for less.A fold, caught by its own implausibility
The first version synthesised to 2 logic cells for an eight-stage 16-bit pipeline. Two always blocks drove the same array — a reset block writing every element and the generate stages writing their own — so yosys pruned nearly all of it. A fold never announces itself: it reports a number far too small rather than an error, and 2 is what made it visible.
What this does not establish
iCE40, not Artix-7, and iCE40 has no DSP blocks so the multiplier arm is maximally penalised. N=8 because at 16 the pair output needs 217 pins on a 206-pin package. No board, one output element, no weight or activation memory.