Skip to content

feat(fpga): measure a real Fmax, and find the clock is not the throughput - #591

Merged
gHashTag merged 1 commit into
mainfrom
measure/depth-and-fmax
Aug 18, 2026
Merged

gHashTag merged 1 commit into
mainfrom
measure/depth-and-fmax

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

Every area number in fpga/phiscale/ has carried the same disclaimer: no Fmax, because no place-and-route is installed. nextpnr-xilinx still is not, and neither is Vivado — but nextpnr-ice40 is installable and the RTL is generic.

Fabric caveat first, because it bounds everything: iCE40 HX8K has no DSP blocks at all, so the multiplier arm is maximally penalised here. On xc7 with DSPs available the multiplier arm is smaller in LUTs (3906 vs 4299) and pays three DSP48.

Measured — nextpnr-ice40, hx8k ct256, fan-in 8, ACC=16

arm ICESTORM_LC Fmax
φ, Fibonacci step 425 144.80 MHz
multiplier 1098 69.21 MHz

2.58× smaller and 2.09× faster on the clock. The structure showing through: the φ scale's register-to-register path is one adder, the multiplier's is a multiply.

And then the number that actually matters

The φ scale takes k cycles where the multiplier takes one, so the clock is not the throughput:

k φ cycles φ ns multiplier ns winner
1–3 2–4 13.8–27.6 28.9 φ
4 5 34.5 28.9 multiplier
8 9 62.2 28.9 multiplier

Break-even is k = 3.18. README.md in that directory states the deployed scale is α = mean|W| ≈ 0.02, and log_φ 0.02 = −8.13, so a real layer needs |k| = 8 — past break-even. At |k| = 8 the multiplier arm is 2.15× faster per output element, despite running at half the clock.

So the honest summary of the φ scale path on a DSP-less fabric: 2.6× the area efficiency, 2.1× the clock, and 2.2× less throughput at the exponent a real layer uses. An unrolled or barrel variant trades that back at more area; it is not built.

Two things that did not become claims

ltp is not a timing substitute. Before P&R worked it looked like one. It reported 213 topological hops for a one-adder scale path and 20 for a 32×16 multiplier — it counts hops in netlists ABC restructures differently per design. Neither number was published as depth.

Fan-in 8, not 16. At N=16 the φ arm needs 217 pins on a 206-pin package and P&R stops with Unable to find a placement location. The pair representation carries two 24-bit components where the multiplier carries one; on a pin-limited part that binds before the logic does — the logic fits at 10% utilisation.

…hput

Every area number here carried "no Fmax, no place-and-route installed".
nextpnr-xilinx still is not and neither is Vivado, but nextpnr-ice40 is
installable and the RTL is generic, so: iCE40 HX8K, fan-in 8, ACC=16.

  phi, Fibonacci step   425 LC   144.80 MHz
  multiplier           1098 LC    69.21 MHz

2.58x smaller and 2.09x faster on the clock -- the phi scale path is one
adder where the multiplier's is a multiply.

And then the number that matters. The phi scale takes k cycles where the
multiplier takes one, so per output element the break-even is k = 3.18.
README.md states the deployed alpha = mean|W| ~ 0.02, and log_phi 0.02 =
-8.13, so a real layer needs |k| = 8 -- past break-even. At |k| = 8 the
MULTIPLIER arm is 2.15x faster per element despite half the clock.

Fabric caveat stated first in the file: iCE40 has no DSP blocks at all,
so the multiplier arm is maximally penalised. On xc7 with DSPs it is
smaller in LUTs than the phi arm.

Also records why ltp was not used as a timing substitute: it reported
213 topological hops for a one-adder path and 20 for a 32x16 multiplier,
because ABC restructures each design differently. Neither was published.

And why fan-in 8: at 16 the phi arm needs 217 pins on a 206-pin package
and P&R stops. The pair representation doubles the output width, and on
a pin-limited part that binds before the logic does -- it fits at 10%.
@gHashTag
gHashTag merged commit 8aad57f into main Aug 18, 2026
32 of 43 checks passed
@gHashTag
gHashTag deleted the measure/depth-and-fmax branch August 18, 2026 08:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant