MAC at a glance
A 16-bit BFloat16 multiply–accumulate unit with an FP32 accumulator. Current AI needs optimized workload hardware; MAC is one of those blocks. It is a controller-style datapath built for low-precision ML on a legacy open PDK: BF16 in, FP32 partial sums out, hardened for SkyWater 130 nm and targeted at TinyTapeout 07.
About this manual
This is the detailed technical reference for the project.
It is the companion to the concise, GitHub-facing README.md:
-
README.md— project identity, status, and quick start. -
docs/(this manual) — architecture, interfaces, verification, and implementation detail in engineering-manual style.
Conventions used throughout:
-
monospacemarks commands, file paths, identifiers, and literals. -
BF16 means 1 sign + 8 exponent + 7 mantissa bits; FP32 means IEEE-754 single precision. Bit fields are written
a[15],a[14:7],a[6:0]. -
Unmeasured values are written as
TBD (unmeasured)— never silently estimated. -
Normative pin and format detail lives alongside this manual as Markdown (
docs/*.md,info.yaml) and is linked, not duplicated.
| Start with Introduction for scope and reading order, then Architecture for the system overview. |
1. Introduction
This chapter defines what the system is, what this manual covers, and how to navigate the rest of the document.
1.1. Purpose
MAC is a 16-bit BFloat16 multiply–accumulate unit with an FP32 accumulator, built for low-precision ML on a legacy open PDK.
It computes acc += A × B where A and B are BF16 and acc
is IEEE-754 single precision, so long dot products keep FP32
dynamic range while the multiplier stays 8×8 instead of 11×11.
This manual is the engineering reference for that system. It records design intent, interfaces, verification strategy, hardening flow, and measured behaviour in one numbered document.
The fastest way to see what the system is: the hardened-layout preview and links at the top of this manual.
1.2. Scope
The manual covers:
-
system architecture and pipeline (see Architecture)
-
arithmetic data structures and special cases (see Design)
-
streaming bus and pin interfaces (see Interfaces)
-
the drive model (see Programming Model)
-
correctness strategy (see Verification)
-
hardening layout and signoff (see Implementation)
-
measured performance (see Performance)
-
honest limits (see Known Limitations)
It does not replace:
-
README.md— quick start and repository map -
docs/architecture.md,docs/specs.md,docs/verification.md,docs/librelane.md,docs/info.md— Markdown authorities beside this manual -
log/— build journal and motivation -
module-level comments in
src/mac_core.svandsrc/tt_um_tensor_mac.v
1.3. How to read this manual
Read chapters in order the first time. Chapters 1–2 give context; chapters 3–5 give the working model; chapters 6–8 give the evidence that the system behaves as claimed.
Each chapter is self-contained enough to be entered directly from the sidebar once the overview in Architecture is familiar.
1.4. Conventions
-
Internal references look like this: Verification.
-
External links look like this: TinyTapeout, the shuttle target for this project.
-
Inline code uses monospace, e.g.
mac_core,tt_um_tensor_mac,ui_in. -
File paths are repository-relative, e.g.
src/mac_core.sv.
-
Unordered lists state properties or inventories.
-
Ordered lists state procedures where sequence matters:
-
Assert
rst_nlow, then release. -
Stream four bytes per MAC.
-
Reconstruct the 32-bit result from the two 16-bit phases.
-
Honest status is a project rule.
Roadmap items stay labelled roadmap; unverified claims stay out of
this manual. Where a number has not been measured it is written
TBD (unmeasured).
|
| The sidebar on the left is collapsible. Select a chapter title to navigate; select its arrow control to expand or collapse its subsections without leaving the page. |
2. Architecture
System overview: what the pieces are, how operands flow in, how results flow out, and where each piece lives in the tree.
2.1. System overview
The two paths are deliberately asymmetric: input is pin-bound and streaming; the core itself is a fast 2-stage pipeline.
ui_in (8-bit, 4 cycles)
-> input buffers (operand_a, operand_b)
-> mac_core S1: BF16 multiply -> product_fp32
-> mac_core S2: FP32 accumulate -> acc_reg
-> result_buffer -> uo_out / uio_out (2 phases)
A 4-cycle counter (state_cnt) owns the asymmetry:
four input bytes per MAC at the pins, two output phases for
the 32-bit result. The core latency is 2 cycles; the pin
throughput is 1 MAC / 4 cycles.
2.2. Pipeline stages
Stage 1 (Mul): a[15:0], b[15:0] (BF16)
-> sign XOR, exp add - 127, 8x8 mantissa
-> product_fp32 (FP32) + valid_s1
Stage 2 (Acc): product_fp32 + acc_reg (FP32)
-> align, add/sub, normalize (highest_bit_pos)
-> next_acc -> acc_reg + valid
Splitting multiply and accumulate buys a clean 20 ns budget per phase at the 50 MHz target, instead of one giant combinational MAC through the 32-bit adder and normalization path.
2.3. Module map
| Module | Owns | Stage |
|---|---|---|
|
BF16 multiply, FP32 adder, |
datapath |
|
4-cycle streaming controller, input buffers, |
interface |
| Every module named above runs in this tree and is covered by the verification gates in Verification — see Known Limitations for honest boundaries. |
2.4. Further reading
-
docs/ARCHITECTURE.md— pipeline detail and I/O tables -
docs/SPECS.md— formats, pins, PPA targets -
src/mac_core.sv— Stage 1 / Stage 2 implementation -
src/tt_um_tensor_mac.v— streaming controller -
Design — arithmetic behind each box above
3. Design
Arithmetic data structures and transformations. This is the conceptual companion to the pin-level detail in Interfaces and the byte-level signoff in Implementation.
3.1. Number formats
BF16 keeps FP32’s 8-bit exponent — same dynamic range, less silicon. The accumulator stays FP32 so long dot products do not collapse on precision.
| Bits | Field | Notes |
|---|---|---|
15 |
sign |
|
14:7 |
exponent |
8-bit, bias 127 |
6:0 |
mantissa |
7-bit explicit, 1 implied ( |
| Bits | Field | Notes |
|---|---|---|
31 |
sign |
|
30:23 |
exponent |
8-bit, bias 127 |
22:0 |
mantissa |
23-bit explicit, 1 implied |
3.2. Stage 1: BF16 multiply
new_sign = a[15] ^ b[15]
exp_sum = a[14:7] + b[14:7]
ma, mb = {1'b1, a[6:0]}, {1'b1, b[6:0]}
m_prod = ma * mb // 8x8 -> 16 bits
final_exp = exp_sum - 127 (+1 if m_prod[15])
product_fp32 = {new_sign, final_exp, {m_prod narrowed to 23 bits}}
Special cases, in priority order:
-
NaN in either input → FP32 QNaN
0x7FC00000. -
Inf × zero → QNaN; Inf otherwise → signed Inf.
-
Zero in either input →
0. -
Exponent overflow (
>= 255) → signed Inf +overflow_s1. -
Exponent underflow (
== 0or carry out) → flush to zero.
Denormals are flushed to zero on input; the product is promoted directly to FP32 format for the accumulator.
Query the overflow flag; it is sticky (overflow | overflow_s1 |
overflow_add). It clears only on rst_n or clr_acc.
|
3.3. Stage 2: FP32 accumulate
sa, ea, ma = unpack(acc_reg) // flush denormals to zero
sb, eb, mb = unpack(product_fp32)
align smaller mantissa by |ea - eb| (with 3 guard bits)
m_sum = same-sign ? add : subtract larger-minus-smaller
if m_sum[27]: renormalize right by 1, exp + 1
else: shift left by highest_bit_pos(m_sum[26:0]), exp - shift
next_acc = {s_res, e_res, m_sum narrowed}
NaN propagates; Inf + -Inf (opposite signs) yields QNaN;
like-signed Inf wins directly. A zero m_sum returns 0.
Overflow on the carry path returns signed Inf + overflow_add.
Think of Stage 1 as the tensor product, Stage 2 as the partial-sum
accumulator, and acc_reg as the running dot-product state.
result is a continuous assignment of acc_reg.
|
4. Interfaces
How operators and testbenches touch the system: the streaming bus, the result phases, and the control signals.
4.1. Input protocol (ui_in)
Two 16-bit operands do not fit in 8 pins, so the wrapper streams
A[7:0] → A[15:8] → B[7:0] → B[15:8] over four clocks and triggers
compute on the last byte.
Cycle (state_cnt) |
ui_in |
Meaning |
|---|---|---|
0 |
|
Operand A low byte |
1 |
|
Operand A high byte |
2 |
|
Operand B low byte |
3 |
|
Operand B high byte (triggers MAC via |
The counter is free-running while ena is high; compute_trigger
is high for one cycle at state_cnt == 3.
4.2. Output protocol (uo_out, uio_out)
The 32-bit accumulator is time-multiplexed over two phases from
result_buffer, which captures result whenever valid is high
so reads stay stable during the next 4-cycle load.
state_cnt[1] |
uo_out |
uio_out |
Meaning |
|---|---|---|---|
0 |
|
|
Lower 16 bits of accumulator |
1 |
|
|
Upper 16 bits of accumulator |
Reconstruct with {high_phase, low_phase}. The MAC pipeline means
the result for operation N is ready at phase 1 of operation N+1 —
see Programming Model.
4.3. Control and pin map
| Port | Width | Dir | Notes |
|---|---|---|---|
|
1 |
in |
System clock, 50 MHz target |
|
1 |
in |
Active-low async reset; clears buffers, counter, accumulator |
|
1 |
in |
Chip enable; counter and pipeline advance only when high |
|
8 |
in |
Streaming data bus (see table above) |
|
8 |
out |
Result low/high byte (see phases above) |
|
8 |
out |
Result high byte of the current 16-bit word |
|
8 |
in |
Unused |
|
8 |
out |
Always |
Core-level extras on mac_core (not pinned out): clr_acc
(clear accumulator; tied to 1’b0 in the wrapper — accumulation
only, reset via rst_n), valid (result strobe), overflow
(sticky, see Design).
uio_oe is hard-wired to output. Do not drive uio_in
expecting a response; it is ignored.
|
5. Programming Model
How to drive the system: the per-MAC sequence, output reconstruction, and the accumulation pattern used for dot products.
5.1. Drive sequence
-
Assert
rst_nlow for several cycles, then release. -
Hold
enahigh during normal operation. -
Stream four bytes per MAC: A low, A high, B low, B high.
-
Read
uo_out/uio_outwhile the counter runs; reconstruct the 32-bit FP32 result from the low and high 16-bit phases.
# Pseudocode for one MAC (A, B are BF16 halfwords)
drive(ui_in = A[7:0]) # state_cnt 0
drive(ui_in = A[15:8]) # state_cnt 1
drive(ui_in = B[7:0]) # state_cnt 2
drive(ui_in = B[15:8]) # state_cnt 3 -> compute_trigger
lo = {uio_out, uo_out} # while state_cnt[1] == 0
hi = {uio_out, uo_out} # while state_cnt[1] == 1
result_fp32 = {hi, lo}
Set ena high throughout; the free-running counter advances only
under ena. Assert rst_n low to clear the accumulator between
independent dot products.
5.2. Accumulation pattern
The wrapper ties clr_acc low, so every triggered MAC adds into
acc_reg. A dot product is a stream of MACs without reset:
rst_n pulse -> acc = 0
MAC(1.0, 2.0) -> acc = 2.0
MAC(1.5, 2.0) -> acc = 5.0
read {hi, lo} -> 0x40A00000 (5.0 in FP32)
This 1.0×2.0 + 1.5×2.0 chain is the first non-reset check in the
testbench (see Verification).
When a result surprises you, check phase alignment first.
Reading the high word during state_cnt[1] == 0 (or vice versa)
silently swaps halves. Most surprises are bus phasing, not arithmetic.
|
5.3. Timing model
-
Core latency: 2 cycles (Mul → Acc).
-
Pin throughput: 1 MAC / 4 cycles (streaming bound, not logic bound).
-
Result availability: operation N commits to
result_bufferwhenvalidpulses, readable during operation N+1’s window.
Never assume the result is ready in the same 4-cycle window that triggered it; pipeline it like any 2-stage datapath.
6. Verification
How the project proves the MAC is correct — and how to re-run that proof locally.
6.1. Golden model
An intentionally straightforward Python model computes the expected answer for every stimulus: BF16 truncation on inputs, FP32 accumulation reference, relaxed tolerance for precision loss. Every RTL result must agree with it:
-
reset clears all registers and output buffers
-
zero inputs:
0 × N = 0 -
identity:
1 × N = N, sign-flip handling -
chained accumulation:
1.0×2.0 + 1.5×2.0 -
1,000 random BF16 vectors uniform in
[-2.0, 2.0]
| No RTL change lands without golden-model agreement. Simulation answers "does the logic work?"; the P&R flow answers "does it close on silicon?" (see Implementation). |
6.2. Evidence
6.3. Running the gates
cd test && make # Linux / macOS
cd test; .\run_test.ps1 # Windows
Expected output:
test_mac.test_reset PASS
test_mac.test_bf16_mac_simple PASS
test_mac.test_random_1000_bf16 PASS
This is the same gate CI runs. Unit stimulus lives in test/test_mac.py
and test/test_mac_core.py; benches in test/tb.v and
test/tb_mac_core.v. Details, screenshots, and methodology:
test/README.md.
| Tool | Purpose |
|---|---|
Icarus Verilog |
Event-driven Verilog simulator |
cocotb |
Python testbench framework driving the pseudo-SPI stream |
pytest |
Test runner and reporting |
NumPy |
Numerical reference generation |
| When adding arithmetic, add the golden-model comparison first. A failing reference test is the specification; a passing one is the proof. The suite targets full functional coverage of the pipeline control plus meaningful arithmetic coverage over standard ML ranges. |
7. Implementation
Hardening layout and signoff rules.
docs/SPECS.md is normative for targets; docs/LIBRELANE.md is
normative for the local flow; this chapter summarizes both and states
the invariants the flow enforces.
7.1. Targets
Process: SkyWater 130 nm (sky130_fd_sc_hd, core 1.8 V)
Clock: 50 MHz (period 20 ns)
Die: 400 x 400 um first pass (mac_core.json); shrink after area known
Tile: 1x1 TinyTapeout (increase only on area failure)
info.yaml pins the shuttle contract: top_module tt_um_tensor_mac,
clock_hz 50000000, tiles 1x1, sources tt_um_tensor_mac.v
mac_core.sv, and the ui/uo/uio pinout tables duplicated in
Interfaces.
7.2. Hardening flow
LibreLane is a staged ASIC flow; TinyTapeout tooling is not used locally:
Verilog RTL
-> Synthesis (Yosys, SV-aware for mac_core.sv)
-> Floorplan + PDN
-> Placement
-> CTS
-> Routing
-> Signoff (DRC / LVS / STA)
-> GDSII
Two configs are provided:
| Config | Top module | Use |
|---|---|---|
|
|
Start here — smaller, pure datapath |
|
|
Full TinyTapeout wrapper + streaming bus (later) |
7.3. Running the flow
make librelane-check # one-time Docker + PDK smoke test
make librelane # full GDS for mac_core -> runs/mac/
make librelane-wrapper # full wrapper (when the core closes)
make view-gds # KLayout app if installed, else PNG preview
Outputs land in runs/mac/:
| Path | Contents |
|---|---|
|
Layout for KLayout |
|
Gate-level netlist (feeds |
|
Area, cells, timing summary |
|
Per-step logs (read these when something fails) |
7.4. Config knobs (mac_core.json)
| Variable | Value | Meaning |
|---|---|---|
|
|
50 MHz target (ns); raise to |
|
|
Core utilization % — lower = easier routing |
|
|
Placer density — raise toward |
|
|
Generous first pass; shrink after area known |
Never hand-edit generated GDS or the gate-level netlist.
Tune the JSON knobs and re-run; the flow is the only writer that
preserves DRC / LVS / STA invariants. mac_core.sv uses SystemVerilog
(logic, always_ff) — on SV synthesis errors check
runs/mac/logs/synthesis/yosys.log first.
|
8. Performance
Measured behaviour only. Unmeasured values are TBD (unmeasured).
Every hardening change follows synthesize → place/route → measure →
document.
8.1. Throughput and latency
| Metric | Value | Notes |
|---|---|---|
Throughput |
1 MAC / 4 cycles |
Pin-bound by the 8-bit streaming bus, not the core |
Latency |
2 cycles |
Core pipeline (Mul → Acc) |
Frequency |
50 MHz (20 ns) |
Target; closes only when STA passes in |
Throughput scales with pins, not logic: widening ui_in would raise
MACs/cycle, but the TinyTapeout template fixes 8 input pins.
8.2. Area and power
| Metric | Value | Notes |
|---|---|---|
Area |
~0.10–0.12 mm² ( |
Estimate for 1x1 tile; real number comes from |
Dynamic power |
~10–20 mW @ 50 MHz ( |
Simulation estimate only |
Static power |
< 1 mW ( |
Sky130 HD leakage estimate |
Report area, cell count, and worst slack together from
runs/mac/final/metrics.csv — size alone never justifies a change.
Until the first make librelane closes, all silicon numbers above
stay labelled estimates.
8.3. Benchmarks
# RTL correctness (must stay green)
cd test && make
# Gate-level regression (needs PDK netlist from a harden run)
make gl-test # == cd test && make clean GATES=yes
# Hardening + inspection
make librelane
make view-gds-preview
Power or area numbers copied from another project do not
belong here. If it was not measured in this tree, it is
TBD (unmeasured).
|
9. Known Limitations
Honest boundaries: what the system is, and what it is not. This chapter overrides any optimistic reading of earlier chapters.
9.1. What is real today
-
RTL for
mac_core(BF16 multiply → FP32 accumulate, NaN/Inf/zero handling, sticky overflow) andtt_um_tensor_mac(4-cycle stream, 2-phase result mux). -
cocotb proof: reset,
1.0×2.0 + 1.5×2.0chain, and 1,000 random BF16 vectors against a Python golden model. -
LibreLane configs (
mac_core.json,tt_um_tensor_mac.json) and a documented local hardening path (make librelane-check,make librelane). -
TinyTapeout contract in
info.yaml(50 MHz, 1x1 tile, pinout).
9.2. What is explicitly not here
-
Closed GDS: no
runs/directory exists yet — timing, DRC, and LVS closure is the next milestone, not a completed result. -
Measured silicon PPA: area and power are estimates until
runs/mac/final/metrics.csvexists (TBD (unmeasured)). -
Gate-level regression:
make gl-testrequires a PDK netlist from a harden run that has not happened yet. -
External hardware: none required; the streaming interface uses only dedicated TinyTapeout I/O pins.
9.3. Limits to design around
-
I/O is the bottleneck by design: 1 MAC / 4 cycles at the pins while the core could retire faster. That is a shuttle constraint, not an architecture mistake.
-
Precision is BF16 by design: FP32-range exponents, ~3-decimal-digit mantissas, accumulator-only FP32. Workloads needing FP32 inputs do not belong here.
-
clr_accis tied low in the wrapper: the only way to start a fresh dot product isrst_n. Streaming a new vector without reset folds it into the previous sum. -
Built alongside Tomato, which taught transistor-to-ISA thinking; MAC applies that instinct to an inference block — see
log/2026-08-03 - Reassessing Mac for optimization.md.