Block FIR — a stateful accelerator, in fixed point, built two ways
This example builds on interleaver. That one added a real compute stage to a
data mover; every firing of it was still independent — a gather reads its inputs, writes its outputs,
and remembers nothing. fir_block is the first design in the tree where that stops being true.
A block FIR filters a signal y[i] = Σₖ h[k]·x[i−k] one block at a time. Two things therefore have
to survive from one firing to the next:
- the coefficients
h, loaded once by aLOAD_TAPScommand and read by everyFILTERafter it; - the tail of the previous block — the last
T−1samples — becausey[0]of a block needs samples that arrived in the previous one.
Neither is a buffer passed between components, and neither is memory on the far side of a bus. Both are
storage the module owns, declared with add_state. That is the
headline of this example, and it is why the design is a filter rather than something smaller: it needs
two flavours of state with different lifetimes in one module, so no single-flavour toy can stand in
for it.
Two other things come with the territory. A filter is arithmetic, so this is the example where the
fixed-point story is told end to end — one format for samples,
coefficients and output, and an accumulator that is derived rather than hand-sized. And a filter is
a loop, so it is the natural place to show that the same specification can be realized more than one
way: fir_block ships two hand-written kernels computing bit-identical results, selected by one
parameter, and measures what each costs.
On the word “firing”
This example needs two words that are easy to conflate, so they are used strictly throughout:
- A firing is one execution of a task body — one job, one command. State persists across firings.
- An iteration is one trip of a loop inside a body. The two kernels here differ in what happens per iteration of the filter loop.
They are independent axes: unroll_lane changes the iteration structure and does not change what a
firing is, while add_state changes what survives a firing and does not change any loop. Every other
page in the guide uses the words this way — see
free-running codegen, where a task body is defined as one
firing rather than a loop.
Learning Objectives
In going through this example, you will learn to:
- Give a hardware module memory between firings with
add_state— both load-once, held state and per-firing carry state, and add a declared reset path - Perform a common DSP calculation in fixed point with one declared format, leveraging the format algebra to derive the full-precision accumulators, adding lossy steps where necessary.
- Hand-write kernel task bodies with different unrolling structures — one output per iteration versus a whole lane per iteration — and create a compile-time parameter to select between the options
- Verify stateful hardware, using a golden that is deliberately stateless so that agreeing with it is the proof the state is right, plus falsification tests that break each flavour of state in turn
- Identify the firing patterns that can wedge a free-running pipeline — a stage that consumes without emitting, or a zero-length transfer handled as if it were non-zero — and avoid them by keeping every stage’s token count uniform across opcodes, so that even a no-output command issues a (zero-length) transfer and lands its completion
- Sweep the bitwidth parameters and build a resource model from what comes back — encoding the DSP48’s geometry as a prior that needs no fitting at all, fitting only the counters no closed form reaches, and composing the parts into a whole-design estimate that reports its own confidence
In this example
The pages build the design up from Python, parallel to interleaver:
- Module overview — the filter, the four stages, one leaf dispatching two opcodes, and why the tap load deliberately does not overlap the compute.
- Cross-firing state — the two flavours,
add_state, where thestaticlands in a free-running task, the declared reset path, and the evidence that it survives in real RTL. - Fixed point — one format for everything, the derived accumulator, why the window
reduction is
fixed_sumand never a loop ofadd, and the width ceiling that lives in Python. - Python — building the design: the command/descriptor split, samples versus words, and the composite.
- Testbench (Python) — the stateless golden, and the falsification tests that prove the gate can fail.
- The two kernels — the serial and unrolled task bodies, the shared delay line, and the seeding rule that has bitten this kernel twice.
- DUT codegen — the composite top, how
unroll_laneselects a body, and the storage declarations generated straight fromadd_state. - RTL simulation — the generated BFM harness, the run, and bit-exactness for both realizations against one golden.
- Resource models — what the design asserts about its own area: a zero-parameter DSP prior from the DSP48E1’s geometry, a BRAM prior that asserts zero, a LUT/FF fit over structural features, and the one method per class that installs them.
- The sweep and its results — what 24 syntheses in 20 minutes actually showed, and a composed whole-design estimate checked against totals that trained none of it.
Table of contents
- Module Overview
- Cross-firing state
- Fixed point
- Python
- Testbench (Python)
- The two kernels
- DUT codegen
- RTL simulation
- Resource models - What this design declares about its own area, and how five modules compose into one estimate. FirCompute states its structure — how many multipliers of what width, that nothing goes in block RAM, and the terms LUT and FF may grow in — and a stock VitisResourceModel prices it. FirBlock declares only the interface term. Three modules declare nothing at all and keep the inherited lookup, which is the ratio to expect: the authoring effort concentrates where area actually moves.
- The sweep and its results - What 24 syntheses in 20 minutes actually showed. The DSP prior lands exactly at every point with zero fitted parameters; the LUT/FF fit is 9.8%/7.1% mean under leave-one-out; the composed whole-design estimate is validated against totals that trained none of it, with the warning that the 3.3% design-level figure flatters the model because most of the design is known rather than predicted. Includes the two results that only appear because the design has coupled step functions in it -- the DSP packing win at 8 bits, and the unrolled plateau that is two effects cancelling -- and a design finding: the right realization inverts with sample width.