Repository navigation
feat(fpga): measure a real Fmax, and find the clock is not the throughput - #591
Merged
Merged
Conversation
…hput Every area number here carried "no Fmax, no place-and-route installed". nextpnr-xilinx still is not and neither is Vivado, but nextpnr-ice40 is installable and the RTL is generic, so: iCE40 HX8K, fan-in 8, ACC=16. phi, Fibonacci step 425 LC 144.80 MHz multiplier 1098 LC 69.21 MHz 2.58x smaller and 2.09x faster on the clock -- the phi scale path is one adder where the multiplier's is a multiply. And then the number that matters. The phi scale takes k cycles where the multiplier takes one, so per output element the break-even is k = 3.18. README.md states the deployed alpha = mean|W| ~ 0.02, and log_phi 0.02 = -8.13, so a real layer needs |k| = 8 -- past break-even. At |k| = 8 the MULTIPLIER arm is 2.15x faster per element despite half the clock. Fabric caveat stated first in the file: iCE40 has no DSP blocks at all, so the multiplier arm is maximally penalised. On xc7 with DSPs it is smaller in LUTs than the phi arm. Also records why ltp was not used as a timing substitute: it reported 213 topological hops for a one-adder path and 20 for a 32x16 multiplier, because ABC restructures each design differently. Neither was published. And why fan-in 8: at 16 the phi arm needs 217 pins on a 206-pin package and P&R stops. The pair representation doubles the output width, and on a pin-limited part that binds before the logic does -- it fits at 10%.
This was referenced Aug 18, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every area number in
fpga/phiscale/has carried the same disclaimer: no Fmax, because no place-and-route is installed.nextpnr-xilinxstill is not, and neither is Vivado — butnextpnr-ice40is installable and the RTL is generic.Fabric caveat first, because it bounds everything: iCE40 HX8K has no DSP blocks at all, so the multiplier arm is maximally penalised here. On xc7 with DSPs available the multiplier arm is smaller in LUTs (3906 vs 4299) and pays three DSP48.
Measured — nextpnr-ice40, hx8k ct256, fan-in 8, ACC=16
2.58× smaller and 2.09× faster on the clock. The structure showing through: the φ scale's register-to-register path is one adder, the multiplier's is a multiply.
And then the number that actually matters
The φ scale takes k cycles where the multiplier takes one, so the clock is not the throughput:
Break-even is k = 3.18.
README.mdin that directory states the deployed scale isα = mean|W| ≈ 0.02, andlog_φ 0.02 = −8.13, so a real layer needs |k| = 8 — past break-even. At |k| = 8 the multiplier arm is 2.15× faster per output element, despite running at half the clock.So the honest summary of the φ scale path on a DSP-less fabric: 2.6× the area efficiency, 2.1× the clock, and 2.2× less throughput at the exponent a real layer uses. An unrolled or barrel variant trades that back at more area; it is not built.
Two things that did not become claims
ltpis not a timing substitute. Before P&R worked it looked like one. It reported 213 topological hops for a one-adder scale path and 20 for a 32×16 multiplier — it counts hops in netlists ABC restructures differently per design. Neither number was published as depth.Fan-in 8, not 16. At N=16 the φ arm needs 217 pins on a 206-pin package and P&R stops with
Unable to find a placement location. The pair representation carries two 24-bit components where the multiplier carries one; on a pin-limited part that binds before the logic does — the logic fits at 10% utilisation.