This manual is the detailed technical reference for this project.
The repository root README.md remains the concise project overview;
everything normative about architecture, interfaces, programming,
verification, and implementation lives here.
Sections are assembled from sections/ with include:: directives.
To add, delete, rename, or reorder a chapter, edit the include list
below and the files under sections/ — no theme or build changes needed.
1. Introduction
This chapter introduces the project, its goals, and how this manual is organized.
It distills the repository README.md, the frozen P1 brief
(docs/p1-specification.md), docs/milestones.md, and docs/provenance.md.
1.1. Purpose
Pineapple GPU P1 is a standalone programmable 3D graphics processor for the
Nexys A7-100T (xc7a100tcsg324-1).
It renders indexed triangle meshes — textured, lit, depth-tested — with a
programmable vertex stage and a programmable fragment stage.
The textured pineapple is the first showcase workload, not the architecture:
the same GPU RTL renders the cube asset with no GPU changes
(rtl/cube_top.v is parameters only).
There is no CPU, no DDR, no flash image; make program loads volatile SRAM only.
1.2. Scope of this manual
-
Architecture — system map: every diagram box to file, clock, and width.
-
Design — pipeline stages and RTL organization.
-
Interfaces — PineBus register contract.
-
Programming model — shader ISA, number formats, runtime upload.
-
Verification — testbenches, golden model, silicon evidence.
-
Implementation — FPGA flow, clocks, pins, resources.
-
Performance — measured cycles and frame rates.
-
Known limitations — honest gaps and next steps.
Legacy Markdown sources under docs/*.md are preserved alongside this manual;
this manual is the central, cross-linked reference.
1.3. Conventions
Inline code denotes RTL identifiers, register names, and file paths.
Source listings state their language explicitly:
make test # fast unit suites (Icarus)
make verify # pixel-exact RTL-vs-reference frames (minutes)
make fpga
make program
| Timing and frame-rate numbers in this manual are measured, never promised. Each cites the build, date, and conditions under which it was taken. |
| Start with Architecture for the one-page contract, then Programming model if you write shaders. |
2. Architecture
System-level contract for the Nexys A7-100T.
Source of truth: rtl/, constr/nexys.xdc, docs/architecture.md.
2.1. System overview
PineBus command submission → indexed vertex fetch → programmable vertex shader → primitive assembly, cull, and setup → top-left edge rasterizer → perspective interpolation → early depth test → programmable fragment shader and texture unit → RGB444 back buffer → vblank presentation → board scanout.
rtl/gpu/ holds the graphics processor and never knows board pins or asset
identity.
rtl/board/ owns clocks, buttons, timing, and DVI pins.
rtl/host/ (demo FSM plus camera) is the first PineBus host; a future host
replaces it without touching rtl/gpu/.
2.2. Clock domains
All GPU logic runs on gpu_clk (50 MHz).
All scanout logic runs on pix_clk (25 MHz).
The only crossing is the framebuffer swap plus scanout read.
NEXYS A7-100T (xc7a100tcsg324-1), 100 MHz E3 -+- div/2 -> gpu_clk 50 MHz (GPU, host, keypad, UART)
+- div/4 -> pix_clk 25 MHz (scanout, DVI, vblank)
2.3. System map
Every diagram node maps to exact files, clocks, and widths.
| Node | RTL | Domain | Notes |
|---|---|---|---|
Five buttons |
|
|
2-flop sync, ~1.31 ms tick, ~16 ms debounce. Outputs |
Demo / host controller |
|
|
Per-frame PineBus program: |
PineBus commands |
|
|
|
Command processor |
|
|
Folded into the GPU FSM. |
Vertex / index RAM |
|
|
|
Programmable vertex core |
|
|
Shared Pine core, ~5 cycles per instruction. See Programming model. |
Primitive, clip, cull |
|
|
|
Triangle setup |
|
|
32-cycle restoring divider, per-vertex |
Rasterizer |
|
|
Pixel-center edge walker, top-left fill rule, bbox clamped to 320x180. |
Interpolator |
|
|
7 attributes, one component per 3-cycle pass (~21 cycles per fragment). |
Texture RAM |
|
|
256x256 RGB444, 4 texels per word, 2-cycle read. |
Programmable fragment core |
|
|
Same core reused per fragment. Runtime upload over USB-UART. |
Depth test (early Z) |
|
|
Nearer wins, clear is far ( |
Double framebuffer |
|
|
Two 57,600x12 banks. |
Scanout |
|
|
800x525 timing, pixel-doubled 640x360 centered in 640x480. Output 59.524 Hz. |
DVI output |
|
|
12-bit payload registered, clock forwarded through |
Second-asset proof: rtl/cube_top.v is parameters only (ASSET, INDEX_COUNT).
rtl/gpu/ is untouched between pineapple and cube builds.
rtl/board/bringup_top.v.| The brief sketches 13 discrete stage modules. The as-built RTL folds 7 of them into `pineapple_gpu.v’s sequential FSM. Function is complete; modularity is the documented P2 step. See Known limitations. |
3. Design / microarchitecture
How the pipeline is actually built. Files named here are the specification; comments in RTL resolve any ambiguity.
3.1. Pipeline dataflow
CLEAR / DRAW_INDEXED / PRESENT (PineBus)
-> index fetch (states 2/3) -> vertex shade -> clip/cull/viewport
-> setup (reciprocal) -> rasterize -> interpolate -> early-Z
-> fragment shade (+TEX2D) -> color+Z write -> PRESENT (vblank swap)
Only one vertex or fragment is in flight at a time, except at the
rasterizer boundary, which uses ready/valid backpressure.
Throughput claims about “one fragment per cycle” refer to the rasterizer’s
generation rate, not the pipeline’s consumption rate.
3.2. Key modules
-
rtl/gpu/pineapple_gpu.v— command processor, fetch, clip/cull/setup, interpolator, depth, and PRESENT handshake (one large FSM). -
rtl/gpu/shader_core.v— shared programmable core used by both stages, selected byprogram_base. -
rtl/gpu/rasterizer.v— standalone edge walker withbusy/donebracketing. -
rtl/gpu/reciprocal.v— 32-cycle restoring divider for1/wand1/area. -
rtl/gpu/shader_loader.v,rtl/board/uart_rx.v— staged, checksummed, atomic-between-frames shader upload. -
rtl/memory/render_target.v— dual-bank framebuffer with vblank-only swap. -
rtl/memory/texture_mem.v— 256x256 RGB444 texture ROM. -
rtl/host/demo_host.v— per-frame PineBus program (uniforms, shaders, clear, draw, present);holdinput freezes it during upload commits. -
rtl/host/camera_controller.v— D-pad orbit, pitch, zoom chord, shader-mode cycle. -
rtl/board/—clock_reset.v,keypad.v,gpu_scanout.v,dvi_out.v, plus preservedbringup_top.vdiagnostics.
3.3. Fixed-point and raster math
Full formats live in Programming model. The pipeline-relevant rules:
-
Clip:
w < 512/4096kills (behind/near guard). -
Screen:
x = 160+((px*160)>>24),y = 90-((py*90)>>24); reject outside ±1024. -
Depth:
(pz*q)>>8clamped to0…65535; clear0xFFFFis far; nearer wins, equal fails. -
Edge functions evaluated at pixel centers with the top-left rule, so shared edges fill exactly once.
3.4. Memory organization
vertex_mem[2048]x144 ({uv, normal, pos}), index_mem[16384]x16,
shader_mem[512] (8 slots x 64), zmem[57600]x16, two color banks of
57,600x12, texture 16,384x48.
(* rom_style="block" *) with $readmemh at build; texture is read-only
during draw and packs cleanly.
vertex_mem, index_mem, shader_mem, and the render target live
inside the GPU file rather than as the separate rtl/memory/ files the brief
sketches. This is intentional in v0.3.0 and does not change behavior.
|
4. Interfaces
PineBus is the whole host contract: one 32-bit register per idle-window write,
plus combinatorial status and counter reads.
Source of truth: rtl/gpu/pineapple_gpu.v (reads) and rtl/host/demo_host.v
(the only writer).
4.1. Handshake
write + valid → ready; writes commit only while the GPU is idle
(ready = (state==0) || read).
Reads are combinatorial and always ready; there is no read handshake.
demo_host freezes its stream on hold during shader-upload commits so no
command double-issues.
4.2. Control and geometry registers
| Addr | Name | Bits | Meaning |
|---|---|---|---|
|
STATUS (read) |
|
|
|
CLEAR_COLOR |
|
Written every frame by the demo host. Showcase uses grape-dark blue |
|
VERTEX_BASE |
|
Word base into |
|
INDEX_BASE |
|
Word base into |
|
INDEX_COUNT |
|
3216 (pineapple), 36 (cube). Must be a multiple of 3 with |
|
VS_PROGRAM |
|
Vertex program slot base (words, 0..511). Currently 0. |
|
FS_PROGRAM |
|
Fragment slot: 64/128/…/384 for modes 0..5. Driven by D-pad mode. |
|
CULL |
|
Write-only backface-culling enable. Not readable. |
|
COMMAND |
|
|
4.3. Uniform window (write-only)
Any address with addr[15:8]==1 writes one S18 lane:
u{addr[3:2]}[addr[7:4]] <= wdata[17:0]; // 16 vec4 uniforms u0..u3[0..15]
The demo host fills the 4 MVP rows as 0x100+component*4 (16 words per frame
from assets/camera.mem) plus light and ambient constants at
0x140/0x144/0x148/0x14C.
Shader LDU reads the addressed vec4.
4.4. Counters (read-only, reset-cleared)
| Addr | Name | Counts |
|---|---|---|
|
FRAMES |
Presented frames |
|
FRAME_CYCLES |
|
|
VERTICES / TRIANGLES / CULLED / RASTERIZED |
Geometry flow |
|
FRAGMENTS / KILLED / SHADED |
Early-Z kills versus shaded |
|
TEXTURE_REQUESTS / INSTRUCTIONS |
|
|
PROG_LOADS |
Committed UART shader uploads |
Unmapped reads return 0. There is no interrupt, no DMA, no burst — by design.
FRAME_CYCLES plus the geometry and shading counters is how
Performance attributes bottlenecks instead of guessing.
|
5. Programming model
One vector ISA, two stages (vertex and fragment share the core RTL, selected
by program_base).
Source of truth: rtl/gpu/shader_core.v, tools/pineasm.py,
tools/reference_renderer.py:shader().
5.1. Machine model
-
16 vector registers
r0…r15, each 4x S18Q12 lanes ({x,y,z,w}). -
Straight-line only: no branches, no predication.
-
Max 64 instructions per program (
steps==63faults withoutEND); 512 program words total (8 slots x 64). -
32-bit encoding:
op[31:26] dst[25:22] a[21:18] b[17:14] c[13:10] mask[9:6] uaddr[5:0].LDIrepurposes the low 18 bits as{mask[21:18], imm[17:0]}.OUTrepurposes as{a[21:18], …, slot[1:0]}. -
Per-instruction latency in RTL is ~5 cycles (fetch, operand, multiply, execute, writeback);
TEX2Dpays texture RAM latency on top. -
The integer golden model is bit-exact: wrap-18 arithmetic,
(a*b)>>12truncation, same masks, same texture quantize.
Number format (S18Q12): signed 18-bit, 12 fractional bits, 1.0 = 4096:
\(wrap18(v) = ((v + 131072)\ mod\ 262144) - 131072\)
Multiply: \(rd = ((ra \cdot rb) >> 12)\) truncated, then wrapped.
5.2. Opcodes
| # | Mnemonic | Form | Semantics (per masked lane) |
|---|---|---|---|
0 |
|
|
Halt, success. Missing |
1 |
|
|
|
2 |
|
|
|
3 |
|
|
|
4 |
|
|
|
5 |
|
|
|
6 |
|
|
|
7 |
|
|
|
8 |
|
|
3-lane dot, broadcast to masked lanes |
9 |
|
|
4-lane dot, broadcast |
10 |
|
|
|
11 |
|
|
|
12 |
|
|
clamp to |
13 |
|
|
|
14 |
|
|
|
15 |
|
|
Sample 256x256 RGB444 at |
16 |
|
|
|
Write mask .xyzw defaults to all lanes.
5.3. Stage ABIs (fixed by pineapple_gpu.v, not by shaders)
-
Vertex in:
r0={1.0,x,y,z} r1={0,nx,ny,nz} r2={0,0,u,v}(+ uniforms). Vertex out:OUT 0= clipp,OUT 1= normal,OUT 2= uv. -
Fragment in:
r0={1,0,v,u} r1={0,nz,ny,nx} r2={1,z,z,z}. Fragment out:OUT 0= linear RGB (quantized to RGB444 at writeback). -
Six showcase fragment slots (
fs_base = (mode+1)*64): 0 diffuse, 1 normal, 2 toon, 3 unlit, 4 uv, 5 depth.
Example — canonical vertex shader (shaders/basic.vert):
LDU r3, u0 ; MVP row 0
DP4 r4.x, r3, r0
LDU r3, u1
DP4 r4.y, r3, r0
LDU r3, u2
DP4 r4.z, r3, r0
LDU r3, u3
DP4 r4.w, r3, r0
OUT 0, r4
OUT 1, r1
OUT 2, r2
END
5.4. Runtime shader upload
New shaders run on the FPGA with no resynthesis.
tools/pineload.py assembles a source file and sends it over USB-UART
(115200 8N1, RsRx pin C4); rtl/gpu/shader_loader.v stages the payload and
commits it to instruction RAM atomically between frames.
Wire frame (PINE magic, LE16 base and count, payload, 8-bit checksum):
| Field | Bytes | Notes |
|---|---|---|
magic |
|
Exact match, driven from true-idle. |
base |
LE16 |
Program-RAM word address, 0..511. |
count |
LE16 |
Payload words, 1..512, |
payload |
4 x count |
Words little-endian. |
checksum |
1 |
Sum of every post-magic byte mod 256. |
Slots in assets/program.mem (seven baked programs, eighth slot free):
| Base | Content | Selected by |
|---|---|---|
0 |
|
fixed |
64 |
|
D-pad mode 0 |
128 |
|
mode 1 |
192 |
|
mode 2 |
256 |
|
mode 3 |
320 |
|
mode 4 |
384 |
|
mode 5 |
448 |
free |
|
Usage:
pip install pyserial # once
python3 tools/pineload.py shaders/mine.frag --mode 2
Select D-pad mode 2 first; the new fragment program takes effect on the next frame.
The serial port is write-only: the screen is the acknowledgement.
A bad frame leaves program RAM untouched and sticks error until the next
good commit or reset. Uploads survive mode switches only in their own slot;
reprogramming the bitstream restores the baked image.
|
6. Verification
What has actually been checked, with what, and what remains open.
Evidence lives in tb/, tools/, and docs/validation/.
6.1. Test layers
-
make test— fast unit suites (Icarus): board frames, keypad, camera control, framebuffer CDC and vblank swap, reciprocal divider, DRAW-error path, 96 randomized shader programs versus the integer reference, 24 rasterizer triangles under randomized backpressure, UART bytes/framing/ staging and a wired upload. -
make verify— pixel-exact full-pipeline check (minutes): every RTL frame must match the integer golden model with zero mismatched pixels out of 57,600: cube shader modes 0–5 (yaw 5) plus the 1072-triangle pineapple showcase (reference CRC = RTL CRC =63134D29). -
make fpga/make program— synthesis, place-and-route, timing, and on-silicon HDMI confirmation (see Implementation).
make test # fast unit suites (Icarus)
make verify # pixel-exact RTL-vs-reference frames (minutes)
6.2. Golden model
tools/reference_renderer.py is integer-only and bit-exact with RTL —
[s18, mul, edge, project] must be read as the specification, not as an
approximation.
tools/check_frame.py compares RTL frame dumps against it (CRC plus mismatch
count); tools/generate_tests.py produces the randomized shader and raster
vectors.
6.3. Silicon evidence (v0.2.0 and v0.3.0)
-
v0.2.0:
make test8/8 PASS;make verifyPASS (cube modes 0–5 and pineapple, zero mismatches);make fpgaPASS at 50 MHz post-route (gpu_clk50.36 MHz);make programPASS, DONE set. -
v0.3.0:
rtl/gpu/untouched; cube build PASS at 50 MHz post-route (gpu_clk60.88 MHz);make program-cubePASS with matching mode-0 yaw-0 frame (CRC641A94BF, 0 mismatches); pineapple rebuilt and re-confirmed on screen. -
2026-10-01 re-verification (Icarus 13.0, Apple Silicon):
make testPASS (all 9 binaries),make verifyPASS (showcase CRC63134D29, 1,998,853 cycles; cube modes 0–5 CRC-matched).
Logs: docs/validation/icarus.log, docs/validation/fpga-summary.log.
Photos: docs/validation/hardware-pineapple.jpg,
docs/validation/hardware-cube.jpg.
6.4. Bench sessions A1/A2 — pending
Builds: make program (pineapple), make program-cube (cube).
| # | Build | Input | Expected | Photo | Result |
|---|---|---|---|---|---|
1 |
pineapple |
power-on boot |
lit textured pineapple on |
||
2 |
pineapple |
RIGHT x3 |
yaw steps, matches sim poses |
||
3 |
pineapple |
UP x2 |
pitch steps, matches sim poses |
||
4 |
pineapple |
CENTER+UP / CENTER+DOWN |
zoom in / out |
||
5 |
pineapple |
CENTER tap modes 0–5 |
0 diffuse, 1 normal, 2 toon, 3 unlit, 4 uv, 5 depth |
||
6 |
cube |
boot + RIGHT x2 + modes 0,1 |
same behavior, cube asset |
Live shader upload:
| # | Step | Expected | Result |
|---|---|---|---|
1 |
mode 2 selected, |
flat-shaded frame, no resynthesis or reboot |
|
2 |
CENTER-tap away and back to mode 2 |
upload persists in its slot |
|
3 |
|
+1, |
| Any photo diverging from the sim-predicted frame is a P1 defect: record the mode, yaw, and reference CRC alongside it. |
7. Implementation
How the design is turned into a bitstream and what it costs.
Target is fixed: xc7a100tcsg324-1 on the Nexys A7-100T.
7.1. Toolchain
Yosys synthesis → nextpnr-xilinx place-and-route → Project X-Ray bitstream,
reusing the installed Tomato hardware/fpga/common.mk flow in place
(TOMATO_FPGA override; see root README.md).
No Vivado, no CPU firmware, no copied toolchain, no destructive setup targets.
make fpga # synthesize + route TOP=pineapple_top
make program # load volatile FPGA SRAM (no flash image)
make fpga-cube # same flow with TOP=cube_top
make program-cube
Synthesis uses -nodsp: the installed nextpnr-xilinx cannot reliably route
the mapped DSP carry ports, so 18x18 multiplies stay in fabric by choice, not
by accident. Widths are unchanged, so timing — not function — is affected.
7.2. Clocks and reset
100 MHz oscillator on E3 → rtl/board/clock_reset.v → divided gpu_clk
50 MHz (GPU, host, keypad, UART) and pix_clk 25 MHz (scanout, DVI),
distributed with BUFG.
Reset is asserted asynchronously and released synchronously with the pixel
clock running (CPU_RESETN on C12).
The fabric-divided clocks are proven but fragile (~0.7% margin on the pineapple build below). Moving to Vivado/MMCM is the first timing-robustness upgrade, not a functional one.
7.3. Pins
Clock, reset, five D-pad buttons, USB-UART RsRx on C4, and 12-bit TFP410
DVI PMOD across JC/JD only (constr/nexys.xdc).
Unused peripheral constraints from the Tomato source are excluded; see
docs/provenance.md.
7.4. Resources (measured post-route, never bit-counted)
Payload alone is 3,090,432 bits (2x 57,600x12 color + 57,600x16 depth
65,536x12 texture ≈ 84 RAMB36 at ideal packing).
Port width and depth rounding plus index, vertex, shader, uniform, and camera
tables cost the rest.
Measured v0.2.0: 22% LUT, 42% RAMB18, 41% RAMB36 of the 100T’s 135 RAMB36. Any capacity claim must cite a post-route report, never a bit total.
7.5. Timing (measured post-route)
-
Pineapple
make fpga:gpu_clk50.36 MHz vs 50 MHz target;pix_clk151.10 MHz. -
Cube
make fpga-cube:gpu_clk60.88 MHz,pix_clk120.32 MHz. -
Pineapple rebuild after parameter plumbing:
gpu_clk55.43 MHz.
| Only the post-route number gates. Pre-route estimates (e.g. 40.03 MHz seen once) are informational. |
8. Performance
Measured frame loop at 50 MHz (frame_cycles, host overhead included).
Pose- and mode-dependent — fps is a measurement, not a spec.
8.1. Frame table (Icarus 13.0, 2026-10-01)
| Workload | Cycles | fps | Notes |
|---|---|---|---|
cube yaw 0, mode 0 (diffuse) |
1,359,877 |
36.8 |
face-on: 9,120 of 18,240 fragments early-Z killed |
cube yaw 5, mode 0 (diffuse) |
1,835,013 |
27.2 |
14,052 shaded x ~11 instr |
cube yaw 5, mode 1 (normal) |
1,409,029 |
35.5 |
14,052 shaded x ~5 instr |
cube yaw 5, mode 2 (toon) |
1,835,013 |
27.2 |
same cost class as diffuse |
cube yaw 5, modes 3/4/5 |
1.20–1.28 M |
39–42 |
cheapest shaders |
pineapple showcase |
1,998,853 |
25.0 |
1,072 tris, 12,780 shaded x ~14 instr |
Hardware PRESENT can add up to one 16.7 ms vblank wait on top.
30 fps is a measured target, never a guarantee.
8.2. Bottleneck attribution (pineapple frame)
From counters plus FSM latencies:
-
Shader execution ~45% (179,172 instr x ~5 cycles — irreducible, it is the workload).
-
Divider-blocked reciprocal waits ~27% (16k fragment + 3.2k vertex + 1k setup waits x 32 cycles;
CLEARis a fixed 57,600). -
Serial interpolation ~17% (7 attributes x 3 cycles x 16,165 fragments).
-
Raster stepping and host overhead the rest.
-
Early-Z kills 3,385 fragments before shading — without it the frame would cost ~25% more.
One structural win clears 30 fps: a pipelined reciprocal (2-cycle
throughput instead of 32-cycle blocking) removes ~0.5 M cycles, giving
pineapple ≈1.5 M ≈ 33 fps.
It is provably bit-exact because every numerator is a site constant
(2^24 for q, 2^30 for 1/area) — only denominators vary over
enumerable ranges, so exhaustive equivalence simulation is feasible.
Interpolation widening (~−11%) is the second lever, not the first.
|
9. Known limitations
What P1 does not do, what is fragile, and what changes first in P2. Read this before proposing optimizations.
9.1. Functional gaps
-
Near-plane clipping is not implemented — the demo camera is constrained to never cross it.
-
UART upload is write-only; the screen is the only acknowledgement.
-
Uploads persist per slot across mode switches but not across bitstream reprogramming (baked
assets/program.memreturns). -
CULLis write-only; several geometry registers are effectively fixed at 0 in the demo host configuration. -
DRAW validation uses a serial mod-3 check (state 42, 14 cycles per DRAW) after
%was found to infer a ~32 ns divider.
9.2. Timing and flow fragility
-
Fabric-divided clocks with ~0.7% margin on the pineapple build; Vivado/MMCM is the first robustness upgrade.
-
-nodspsynthesis works around a nextpnr DSP-carry routing limitation on this install; multiplies stay in fabric. -
nextpnr ignores
[current_design]configuration properties in the reused XDC; Project X-Ray falls back to its slower Python FASM parser. Neither prevents bitstream generation.
9.3. Verification debt
-
Physical D-pad confirmation (orbit, pitch, zoom, shader-mode cycling) and bench sessions A1/A2 remain open — see Verification.
-
Second-asset silicon proof is closed (cube,
rtl/gpu/untouched); the remaining gate is the physical pass, not simulation.
9.4. Correct P2 order
-
Split
pineapple_gpu.vFSM into the sketched stage modules with pixel-exact parity (make verify: zero mismatches, CRC63134D29). -
Pipeline the reciprocal (throughput, not latency).
-
Widen interpolation only against counter data (
fragments/killed/shaded/instructions/texture_requests).
Each stage should be a separately reviewable change with passing Icarus tests and synthesis for hardware changes.