MAC at a glance

A 16-bit BFloat16 multiply–accumulate unit with an FP32 accumulator. Current AI needs optimized workload hardware; MAC is one of those blocks. It is a controller-style datapath built for low-precision ML on a legacy open PDK: BF16 in, FP32 partial sums out, hardened for SkyWater 130 nm and targeted at TinyTapeout 07.

MAC core hardened layout on Sky130
Figure 1. Demo — the hardened core on Sky130 (LibreLane GDS preview)
Author Profile GitHub Repository

About this manual

This is the detailed technical reference for the project. It is the companion to the concise, GitHub-facing README.md:

  • README.md — project identity, status, and quick start.

  • docs/ (this manual) — architecture, interfaces, verification, and implementation detail in engineering-manual style.

Conventions used throughout:

  • monospace marks commands, file paths, identifiers, and literals.

  • BF16 means 1 sign + 8 exponent + 7 mantissa bits; FP32 means IEEE-754 single precision. Bit fields are written a[15], a[14:7], a[6:0].

  • Unmeasured values are written as TBD (unmeasured) — never silently estimated.

  • Normative pin and format detail lives alongside this manual as Markdown (docs/*.md, info.yaml) and is linked, not duplicated.

Start with Introduction for scope and reading order, then Architecture for the system overview.

1. Introduction

This chapter defines what the system is, what this manual covers, and how to navigate the rest of the document.

1.1. Purpose

MAC is a 16-bit BFloat16 multiply–accumulate unit with an FP32 accumulator, built for low-precision ML on a legacy open PDK.

It computes acc += A × B where A and B are BF16 and acc is IEEE-754 single precision, so long dot products keep FP32 dynamic range while the multiplier stays 8×8 instead of 11×11.

This manual is the engineering reference for that system. It records design intent, interfaces, verification strategy, hardening flow, and measured behaviour in one numbered document.

The fastest way to see what the system is: the hardened-layout preview and links at the top of this manual.

1.2. Scope

The manual covers:

It does not replace:

  • README.md — quick start and repository map

  • docs/architecture.md, docs/specs.md, docs/verification.md, docs/librelane.md, docs/info.md — Markdown authorities beside this manual

  • log/ — build journal and motivation

  • module-level comments in src/mac_core.sv and src/tt_um_tensor_mac.v

1.3. How to read this manual

Read chapters in order the first time. Chapters 1–2 give context; chapters 3–5 give the working model; chapters 6–8 give the evidence that the system behaves as claimed.

Each chapter is self-contained enough to be entered directly from the sidebar once the overview in Architecture is familiar.

1.4. Conventions

Normal cross-reference and link usage
  • Internal references look like this: Verification.

  • External links look like this: TinyTapeout, the shuttle target for this project.

  • Inline code uses monospace, e.g. mac_core, tt_um_tensor_mac, ui_in.

  • File paths are repository-relative, e.g. src/mac_core.sv.

List conventions
  • Unordered lists state properties or inventories.

  • Ordered lists state procedures where sequence matters:

    1. Assert rst_n low, then release.

    2. Stream four bytes per MAC.

    3. Reconstruct the 32-bit result from the two 16-bit phases.

Honest status is a project rule. Roadmap items stay labelled roadmap; unverified claims stay out of this manual. Where a number has not been measured it is written TBD (unmeasured).
The sidebar on the left is collapsible. Select a chapter title to navigate; select its arrow control to expand or collapse its subsections without leaving the page.

2. Architecture

System overview: what the pieces are, how operands flow in, how results flow out, and where each piece lives in the tree.

2.1. System overview

The two paths are deliberately asymmetric: input is pin-bound and streaming; the core itself is a fast 2-stage pipeline.

ui_in (8-bit, 4 cycles)
  -> input buffers (operand_a, operand_b)
  -> mac_core S1: BF16 multiply -> product_fp32
  -> mac_core S2: FP32 accumulate -> acc_reg
  -> result_buffer -> uo_out / uio_out (2 phases)

A 4-cycle counter (state_cnt) owns the asymmetry: four input bytes per MAC at the pins, two output phases for the 32-bit result. The core latency is 2 cycles; the pin throughput is 1 MAC / 4 cycles.

2.2. Pipeline stages

Stage 1 (Mul):  a[15:0], b[15:0] (BF16)
                  -> sign XOR, exp add - 127, 8x8 mantissa
                  -> product_fp32 (FP32) + valid_s1

Stage 2 (Acc):  product_fp32 + acc_reg (FP32)
                  -> align, add/sub, normalize (highest_bit_pos)
                  -> next_acc -> acc_reg + valid

Splitting multiply and accumulate buys a clean 20 ns budget per phase at the 50 MHz target, instead of one giant combinational MAC through the 32-bit adder and normalization path.

2.3. Module map

Table 1. Modules and their responsibilities
Module Owns Stage

mac_core

BF16 multiply, FP32 adder, acc_reg, valid / overflow

datapath

tt_um_tensor_mac

4-cycle streaming controller, input buffers, result_buffer, output mux, uio_oe = 0xFF

interface

Every module named above runs in this tree and is covered by the verification gates in Verification — see Known Limitations for honest boundaries.

2.4. Further reading

  • docs/ARCHITECTURE.md — pipeline detail and I/O tables

  • docs/SPECS.md — formats, pins, PPA targets

  • src/mac_core.sv — Stage 1 / Stage 2 implementation

  • src/tt_um_tensor_mac.v — streaming controller

  • Design — arithmetic behind each box above

3. Design

Arithmetic data structures and transformations. This is the conceptual companion to the pin-level detail in Interfaces and the byte-level signoff in Implementation.

3.1. Number formats

BF16 keeps FP32’s 8-bit exponent — same dynamic range, less silicon. The accumulator stays FP32 so long dot products do not collapse on precision.

Table 2. BF16 input (16 bits)
Bits Field Notes

15

sign

0 = positive, 1 = negative

14:7

exponent

8-bit, bias 127

6:0

mantissa

7-bit explicit, 1 implied ({1’b1, a[6:0]})

Table 3. FP32 accumulator (32 bits)
Bits Field Notes

31

sign

0 = positive, 1 = negative

30:23

exponent

8-bit, bias 127

22:0

mantissa

23-bit explicit, 1 implied

3.2. Stage 1: BF16 multiply

new_sign = a[15] ^ b[15]
exp_sum  = a[14:7] + b[14:7]
ma, mb   = {1'b1, a[6:0]}, {1'b1, b[6:0]}
m_prod   = ma * mb            // 8x8 -> 16 bits
final_exp = exp_sum - 127 (+1 if m_prod[15])
product_fp32 = {new_sign, final_exp, {m_prod narrowed to 23 bits}}

Special cases, in priority order:

  • NaN in either input → FP32 QNaN 0x7FC00000.

  • Inf × zero → QNaN; Inf otherwise → signed Inf.

  • Zero in either input → 0.

  • Exponent overflow (>= 255) → signed Inf + overflow_s1.

  • Exponent underflow (== 0 or carry out) → flush to zero.

Denormals are flushed to zero on input; the product is promoted directly to FP32 format for the accumulator.

Query the overflow flag; it is sticky (overflow | overflow_s1 | overflow_add). It clears only on rst_n or clr_acc.

3.3. Stage 2: FP32 accumulate

sa, ea, ma = unpack(acc_reg)      // flush denormals to zero
sb, eb, mb = unpack(product_fp32)
align smaller mantissa by |ea - eb| (with 3 guard bits)
m_sum = same-sign ? add : subtract larger-minus-smaller
if m_sum[27]: renormalize right by 1, exp + 1
else: shift left by highest_bit_pos(m_sum[26:0]), exp - shift
next_acc = {s_res, e_res, m_sum narrowed}

NaN propagates; Inf + -Inf (opposite signs) yields QNaN; like-signed Inf wins directly. A zero m_sum returns 0. Overflow on the carry path returns signed Inf + overflow_add.

Think of Stage 1 as the tensor product, Stage 2 as the partial-sum accumulator, and acc_reg as the running dot-product state. result is a continuous assignment of acc_reg.

4. Interfaces

How operators and testbenches touch the system: the streaming bus, the result phases, and the control signals.

4.1. Input protocol (ui_in)

Two 16-bit operands do not fit in 8 pins, so the wrapper streams A[7:0] → A[15:8] → B[7:0] → B[15:8] over four clocks and triggers compute on the last byte.

Table 4. Input bytes per MAC
Cycle (state_cnt) ui_in Meaning

0

A[7:0]

Operand A low byte

1

A[15:8]

Operand A high byte

2

B[7:0]

Operand B low byte

3

B[15:8]

Operand B high byte (triggers MAC via compute_trigger)

The counter is free-running while ena is high; compute_trigger is high for one cycle at state_cnt == 3.

4.2. Output protocol (uo_out, uio_out)

The 32-bit accumulator is time-multiplexed over two phases from result_buffer, which captures result whenever valid is high so reads stay stable during the next 4-cycle load.

Table 5. Result phases
state_cnt[1] uo_out uio_out Meaning

0

result[7:0]

result[15:8]

Lower 16 bits of accumulator

1

result[23:16]

result[31:24]

Upper 16 bits of accumulator

Reconstruct with {high_phase, low_phase}. The MAC pipeline means the result for operation N is ready at phase 1 of operation N+1 — see Programming Model.

4.3. Control and pin map

Table 6. Ports (info.yaml, src/tt_um_tensor_mac.v)
Port Width Dir Notes

clk

1

in

System clock, 50 MHz target

rst_n

1

in

Active-low async reset; clears buffers, counter, accumulator

ena

1

in

Chip enable; counter and pipeline advance only when high

ui_in

8

in

Streaming data bus (see table above)

uo_out

8

out

Result low/high byte (see phases above)

uio_out

8

out

Result high byte of the current 16-bit word

uio_in

8

in

Unused

uio_oe

8

out

Always 0xFF — uio is output-only in this design

Core-level extras on mac_core (not pinned out): clr_acc (clear accumulator; tied to 1’b0 in the wrapper — accumulation only, reset via rst_n), valid (result strobe), overflow (sticky, see Design).

uio_oe is hard-wired to output. Do not drive uio_in expecting a response; it is ignored.

5. Programming Model

How to drive the system: the per-MAC sequence, output reconstruction, and the accumulation pattern used for dot products.

5.1. Drive sequence

  1. Assert rst_n low for several cycles, then release.

  2. Hold ena high during normal operation.

  3. Stream four bytes per MAC: A low, A high, B low, B high.

  4. Read uo_out / uio_out while the counter runs; reconstruct the 32-bit FP32 result from the low and high 16-bit phases.

# Pseudocode for one MAC (A, B are BF16 halfwords)
drive(ui_in = A[7:0])    # state_cnt 0
drive(ui_in = A[15:8])   # state_cnt 1
drive(ui_in = B[7:0])    # state_cnt 2
drive(ui_in = B[15:8])   # state_cnt 3 -> compute_trigger
lo = {uio_out, uo_out}   # while state_cnt[1] == 0
hi = {uio_out, uo_out}   # while state_cnt[1] == 1
result_fp32 = {hi, lo}

Set ena high throughout; the free-running counter advances only under ena. Assert rst_n low to clear the accumulator between independent dot products.

5.2. Accumulation pattern

The wrapper ties clr_acc low, so every triggered MAC adds into acc_reg. A dot product is a stream of MACs without reset:

rst_n pulse -> acc = 0
MAC(1.0, 2.0) -> acc = 2.0
MAC(1.5, 2.0) -> acc = 5.0
read {hi, lo} -> 0x40A00000 (5.0 in FP32)

This 1.0×2.0 + 1.5×2.0 chain is the first non-reset check in the testbench (see Verification).

When a result surprises you, check phase alignment first. Reading the high word during state_cnt[1] == 0 (or vice versa) silently swaps halves. Most surprises are bus phasing, not arithmetic.

5.3. Timing model

  • Core latency: 2 cycles (Mul → Acc).

  • Pin throughput: 1 MAC / 4 cycles (streaming bound, not logic bound).

  • Result availability: operation N commits to result_buffer when valid pulses, readable during operation N+1’s window.

Never assume the result is ready in the same 4-cycle window that triggered it; pipeline it like any 2-stage datapath.

6. Verification

How the project proves the MAC is correct — and how to re-run that proof locally.

6.1. Golden model

An intentionally straightforward Python model computes the expected answer for every stimulus: BF16 truncation on inputs, FP32 accumulation reference, relaxed tolerance for precision loss. Every RTL result must agree with it:

  • reset clears all registers and output buffers

  • zero inputs: 0 × N = 0

  • identity: 1 × N = N, sign-flip handling

  • chained accumulation: 1.0×2.0 + 1.5×2.0

  • 1,000 random BF16 vectors uniform in [-2.0, 2.0]

No RTL change lands without golden-model agreement. Simulation answers "does the logic work?"; the P&R flow answers "does it close on silicon?" (see Implementation).

6.2. Evidence

Cocotb test run
Figure 2. Test run — cocotb driving the streaming bus
Passing MAC test suite
Figure 3. Passing suite — reset, chain, and 1,000 random vectors

6.3. Running the gates

cd test && make          # Linux / macOS
cd test; .\run_test.ps1  # Windows

Expected output:

test_mac.test_reset                 PASS
test_mac.test_bf16_mac_simple       PASS
test_mac.test_random_1000_bf16      PASS

This is the same gate CI runs. Unit stimulus lives in test/test_mac.py and test/test_mac_core.py; benches in test/tb.v and test/tb_mac_core.v. Details, screenshots, and methodology: test/README.md.

Table 7. Tools
Tool Purpose

Icarus Verilog

Event-driven Verilog simulator

cocotb

Python testbench framework driving the pseudo-SPI stream

pytest

Test runner and reporting

NumPy

Numerical reference generation

When adding arithmetic, add the golden-model comparison first. A failing reference test is the specification; a passing one is the proof. The suite targets full functional coverage of the pipeline control plus meaningful arithmetic coverage over standard ML ranges.

7. Implementation

Hardening layout and signoff rules. docs/SPECS.md is normative for targets; docs/LIBRELANE.md is normative for the local flow; this chapter summarizes both and states the invariants the flow enforces.

7.1. Targets

Process:  SkyWater 130 nm (sky130_fd_sc_hd, core 1.8 V)
Clock:    50 MHz (period 20 ns)
Die:      400 x 400 um first pass (mac_core.json); shrink after area known
Tile:     1x1 TinyTapeout (increase only on area failure)

info.yaml pins the shuttle contract: top_module tt_um_tensor_mac, clock_hz 50000000, tiles 1x1, sources tt_um_tensor_mac.v
mac_core.sv, and the ui/uo/uio pinout tables duplicated in Interfaces.

7.2. Hardening flow

LibreLane is a staged ASIC flow; TinyTapeout tooling is not used locally:

Verilog RTL
  -> Synthesis (Yosys, SV-aware for mac_core.sv)
  -> Floorplan + PDN
  -> Placement
  -> CTS
  -> Routing
  -> Signoff (DRC / LVS / STA)
  -> GDSII

Two configs are provided:

Table 8. Configs (librelane/)
Config Top module Use

mac_core.json

mac_core

Start here — smaller, pure datapath

tt_um_tensor_mac.json

tt_um_tensor_mac

Full TinyTapeout wrapper + streaming bus (later)

7.3. Running the flow

make librelane-check   # one-time Docker + PDK smoke test
make librelane         # full GDS for mac_core -> runs/mac/
make librelane-wrapper # full wrapper (when the core closes)
make view-gds          # KLayout app if installed, else PNG preview

Outputs land in runs/mac/:

Path Contents

runs/mac/final/gds/mac_core.gds

Layout for KLayout

runs/mac/final/verilog/mac_core.v

Gate-level netlist (feeds make gl-test / GATES=yes regression)

runs/mac/final/metrics.csv

Area, cells, timing summary

runs/mac/logs/

Per-step logs (read these when something fails)

7.4. Config knobs (mac_core.json)

Variable Value Meaning

CLOCK_PERIOD

20

50 MHz target (ns); raise to 25 if timing fails

FP_CORE_UTIL

40

Core utilization % — lower = easier routing

PL_TARGET_DENSITY

0.45

Placer density — raise toward 0.6-0.7 on GPL-0302

DIE_AREA

400x400 um

Generous first pass; shrink after area known

Never hand-edit generated GDS or the gate-level netlist. Tune the JSON knobs and re-run; the flow is the only writer that preserves DRC / LVS / STA invariants. mac_core.sv uses SystemVerilog (logic, always_ff) — on SV synthesis errors check runs/mac/logs/synthesis/yosys.log first.

8. Performance

Measured behaviour only. Unmeasured values are TBD (unmeasured). Every hardening change follows synthesize → place/route → measure → document.

8.1. Throughput and latency

Metric Value Notes

Throughput

1 MAC / 4 cycles

Pin-bound by the 8-bit streaming bus, not the core

Latency

2 cycles

Core pipeline (Mul → Acc)

Frequency

50 MHz (20 ns)

Target; closes only when STA passes in runs/mac/logs/

Throughput scales with pins, not logic: widening ui_in would raise MACs/cycle, but the TinyTapeout template fixes 8 input pins.

8.2. Area and power

Metric Value Notes

Area

~0.10–0.12 mm² (TBD (unmeasured))

Estimate for 1x1 tile; real number comes from metrics.csv

Dynamic power

~10–20 mW @ 50 MHz (TBD (unmeasured))

Simulation estimate only

Static power

< 1 mW (TBD (unmeasured))

Sky130 HD leakage estimate

Report area, cell count, and worst slack together from runs/mac/final/metrics.csv — size alone never justifies a change. Until the first make librelane closes, all silicon numbers above stay labelled estimates.

8.3. Benchmarks

# RTL correctness (must stay green)
cd test && make

# Gate-level regression (needs PDK netlist from a harden run)
make gl-test   # == cd test && make clean GATES=yes

# Hardening + inspection
make librelane
make view-gds-preview
Power or area numbers copied from another project do not belong here. If it was not measured in this tree, it is TBD (unmeasured).

9. Known Limitations

Honest boundaries: what the system is, and what it is not. This chapter overrides any optimistic reading of earlier chapters.

9.1. What is real today

  • RTL for mac_core (BF16 multiply → FP32 accumulate, NaN/Inf/zero handling, sticky overflow) and tt_um_tensor_mac (4-cycle stream, 2-phase result mux).

  • cocotb proof: reset, 1.0×2.0 + 1.5×2.0 chain, and 1,000 random BF16 vectors against a Python golden model.

  • LibreLane configs (mac_core.json, tt_um_tensor_mac.json) and a documented local hardening path (make librelane-check, make librelane).

  • TinyTapeout contract in info.yaml (50 MHz, 1x1 tile, pinout).

9.2. What is explicitly not here

  • Closed GDS: no runs/ directory exists yet — timing, DRC, and LVS closure is the next milestone, not a completed result.

  • Measured silicon PPA: area and power are estimates until runs/mac/final/metrics.csv exists (TBD (unmeasured)).

  • Gate-level regression: make gl-test requires a PDK netlist from a harden run that has not happened yet.

  • External hardware: none required; the streaming interface uses only dedicated TinyTapeout I/O pins.

9.3. Limits to design around

  • I/O is the bottleneck by design: 1 MAC / 4 cycles at the pins while the core could retire faster. That is a shuttle constraint, not an architecture mistake.

  • Precision is BF16 by design: FP32-range exponents, ~3-decimal-digit mantissas, accumulator-only FP32. Workloads needing FP32 inputs do not belong here.

  • clr_acc is tied low in the wrapper: the only way to start a fresh dot product is rst_n. Streaming a new vector without reset folds it into the previous sum.

  • Built alongside Tomato, which taught transistor-to-ISA thinking; MAC applies that instinct to an inference block — see log/2026-08-03 - Reassessing Mac for optimization.md.