This manual is the detailed technical reference for this project. The repository root README.md remains the concise project overview; everything normative about architecture, interfaces, programming, verification, and implementation lives here.

Sections are assembled from sections/ with include:: directives. To add, delete, rename, or reorder a chapter, edit the include list below and the files under sections/ — no theme or build changes needed.

1. Introduction

This chapter introduces the project, its goals, and how this manual is organized. It distills the repository README.md, the frozen P1 brief (docs/p1-specification.md), docs/milestones.md, and docs/provenance.md.

1.1. Purpose

Pineapple GPU P1 is a standalone programmable 3D graphics processor for the Nexys A7-100T (xc7a100tcsg324-1). It renders indexed triangle meshes — textured, lit, depth-tested — with a programmable vertex stage and a programmable fragment stage.

The textured pineapple is the first showcase workload, not the architecture: the same GPU RTL renders the cube asset with no GPU changes (rtl/cube_top.v is parameters only). There is no CPU, no DDR, no flash image; make program loads volatile SRAM only.

1.2. Scope of this manual

Legacy Markdown sources under docs/*.md are preserved alongside this manual; this manual is the central, cross-linked reference.

1.3. Conventions

Inline code denotes RTL identifiers, register names, and file paths. Source listings state their language explicitly:

make test     # fast unit suites (Icarus)
make verify   # pixel-exact RTL-vs-reference frames (minutes)
make fpga
make program
Timing and frame-rate numbers in this manual are measured, never promised. Each cites the build, date, and conditions under which it was taken.
Start with Architecture for the one-page contract, then Programming model if you write shaders.

2. Architecture

System-level contract for the Nexys A7-100T. Source of truth: rtl/, constr/nexys.xdc, docs/architecture.md.

2.1. System overview

PineBus command submission → indexed vertex fetch → programmable vertex shader → primitive assembly, cull, and setup → top-left edge rasterizer → perspective interpolation → early depth test → programmable fragment shader and texture unit → RGB444 back buffer → vblank presentation → board scanout.

rtl/gpu/ holds the graphics processor and never knows board pins or asset identity. rtl/board/ owns clocks, buttons, timing, and DVI pins. rtl/host/ (demo FSM plus camera) is the first PineBus host; a future host replaces it without touching rtl/gpu/.

2.2. Clock domains

All GPU logic runs on gpu_clk (50 MHz). All scanout logic runs on pix_clk (25 MHz). The only crossing is the framebuffer swap plus scanout read.

NEXYS A7-100T (xc7a100tcsg324-1), 100 MHz E3 -+- div/2 -> gpu_clk 50 MHz (GPU, host, keypad, UART)
                                               +- div/4 -> pix_clk 25 MHz (scanout, DVI, vblank)

2.3. System map

Every diagram node maps to exact files, clocks, and widths.

Node RTL Domain Notes

Five buttons

rtl/board/keypad.v, rtl/host/camera_controller.v

gpu_clk

2-flop sync, ~1.31 ms tick, ~16 ms debounce. Outputs yaw/pitch/zoom/mode. Pins N17 C, M18 U, P18 D, P17 L, M17 R.

Demo / host controller

rtl/host/demo_host.v

gpu_clk

Per-frame PineBus program: INDEX_COUNT, clear color, uniforms, CLEAR → DRAW_INDEXED → PRESENT. hold freezes it during shader upload.

PineBus commands

rtl/gpu/pineapple_gpu.v

gpu_clk

addr/wdata/write/valid → ready. Writes commit only while idle. See Interfaces.

Command processor

pineapple_gpu.v states 0/1/40/41/42

gpu_clk

Folded into the GPU FSM. CLEAR fills color and depth; DRAW_INDEXED validates count; PRESENT handshakes the swap.

Vertex / index RAM

pineapple_gpu.v

gpu_clk

index_mem[16384], vertex_mem[2048]. 2-cycle fetch, bounds-checked; violations raise error.

Programmable vertex core

rtl/gpu/shader_core.v

gpu_clk

Shared Pine core, ~5 cycles per instruction. See Programming model.

Primitive, clip, cull

pineapple_gpu.v states 5/8/9/13

gpu_clk

w clip kill, viewport reject, zero-area kill, cull-gated backface kill, CCW repair.

Triangle setup

pineapple_gpu.v + rtl/gpu/reciprocal.v

gpu_clk

32-cycle restoring divider, per-vertex q, viewport map, 16-bit depth.

Rasterizer

rtl/gpu/rasterizer.v

gpu_clk

Pixel-center edge walker, top-left fill rule, bbox clamped to 320x180.

Interpolator

pineapple_gpu.v states 21/22/36/23

gpu_clk

7 attributes, one component per 3-cycle pass (~21 cycles per fragment).

Texture RAM

rtl/memory/texture_mem.v

gpu_clk

256x256 RGB444, 4 texels per word, 2-cycle read.

Programmable fragment core

shader_core.v + shader_loader.v + uart_rx.v

gpu_clk

Same core reused per fragment. Runtime upload over USB-UART.

Depth test (early Z)

pineapple_gpu.v states 24/31

gpu_clk

Nearer wins, clear is far (0xFFFF). Single fragment in flight, no Z hazard.

Double framebuffer

rtl/memory/render_target.v

gpu_clk + pix_clk

Two 57,600x12 banks. PRESENT swaps at vblank_start only via 2-flop CDC.

Scanout

rtl/board/gpu_scanout.v

pix_clk

800x525 timing, pixel-doubled 640x360 centered in 640x480. Output 59.524 Hz.

DVI output

rtl/board/dvi_out.v

pix_clk

12-bit payload registered, clock forwarded through ODDR. JC/JD PMOD.

Second-asset proof: rtl/cube_top.v is parameters only (ASSET, INDEX_COUNT). rtl/gpu/ is untouched between pineapple and cube builds.

Bring-up test pattern
Figure 1. Figure: bring-up test pattern produced by rtl/board/bringup_top.v.
The brief sketches 13 discrete stage modules. The as-built RTL folds 7 of them into `pineapple_gpu.v’s sequential FSM. Function is complete; modularity is the documented P2 step. See Known limitations.

3. Design / microarchitecture

How the pipeline is actually built. Files named here are the specification; comments in RTL resolve any ambiguity.

3.1. Pipeline dataflow

CLEAR / DRAW_INDEXED / PRESENT (PineBus)
  -> index fetch (states 2/3) -> vertex shade -> clip/cull/viewport
  -> setup (reciprocal) -> rasterize -> interpolate -> early-Z
  -> fragment shade (+TEX2D) -> color+Z write -> PRESENT (vblank swap)

Only one vertex or fragment is in flight at a time, except at the rasterizer boundary, which uses ready/valid backpressure. Throughput claims about “one fragment per cycle” refer to the rasterizer’s generation rate, not the pipeline’s consumption rate.

3.2. Key modules

  • rtl/gpu/pineapple_gpu.v — command processor, fetch, clip/cull/setup, interpolator, depth, and PRESENT handshake (one large FSM).

  • rtl/gpu/shader_core.v — shared programmable core used by both stages, selected by program_base.

  • rtl/gpu/rasterizer.v — standalone edge walker with busy/done bracketing.

  • rtl/gpu/reciprocal.v — 32-cycle restoring divider for 1/w and 1/area.

  • rtl/gpu/shader_loader.v, rtl/board/uart_rx.v — staged, checksummed, atomic-between-frames shader upload.

  • rtl/memory/render_target.v — dual-bank framebuffer with vblank-only swap.

  • rtl/memory/texture_mem.v — 256x256 RGB444 texture ROM.

  • rtl/host/demo_host.v — per-frame PineBus program (uniforms, shaders, clear, draw, present); hold input freezes it during upload commits.

  • rtl/host/camera_controller.v — D-pad orbit, pitch, zoom chord, shader-mode cycle.

  • rtl/board/ — clock_reset.v, keypad.v, gpu_scanout.v, dvi_out.v, plus preserved bringup_top.v diagnostics.

3.3. Fixed-point and raster math

Full formats live in Programming model. The pipeline-relevant rules:

  • Clip: w < 512/4096 kills (behind/near guard).

  • Screen: x = 160+((px*160)>>24), y = 90-((py*90)>>24); reject outside ±1024.

  • Depth: (pz*q)>>8 clamped to 0…65535; clear 0xFFFF is far; nearer wins, equal fails.

  • Edge functions evaluated at pixel centers with the top-left rule, so shared edges fill exactly once.

3.4. Memory organization

vertex_mem[2048]x144 ({uv, normal, pos}), index_mem[16384]x16, shader_mem[512] (8 slots x 64), zmem[57600]x16, two color banks of 57,600x12, texture 16,384x48. (* rom_style="block" *) with $readmemh at build; texture is read-only during draw and packs cleanly.

vertex_mem, index_mem, shader_mem, and the render target live inside the GPU file rather than as the separate rtl/memory/ files the brief sketches. This is intentional in v0.3.0 and does not change behavior.

4. Interfaces

PineBus is the whole host contract: one 32-bit register per idle-window write, plus combinatorial status and counter reads. Source of truth: rtl/gpu/pineapple_gpu.v (reads) and rtl/host/demo_host.v (the only writer).

4.1. Handshake

write + valid → ready; writes commit only while the GPU is idle (ready = (state==0) || read). Reads are combinatorial and always ready; there is no read handshake. demo_host freezes its stream on hold during shader-upload commits so no command double-issues.

4.2. Control and geometry registers

Addr Name Bits Meaning

0x00

STATUS (read)

{30’b0, error, busy}

busy = state!=0; error sticks until reset. Writes have no effect.

0x04

CLEAR_COLOR

[11:0] RGB444

Written every frame by the demo host. Showcase uses grape-dark blue 0x013.

0x08

VERTEX_BASE

[10:0]

Word base into vertex_mem[2048]. Currently 0.

0x0C

INDEX_BASE

[13:0]

Word base into index_mem[16384]. Currently 0.

0x10

INDEX_COUNT

[13:0]

3216 (pineapple), 36 (cube). Must be a multiple of 3 with base+count ⇐ 16384, else error and DRAW aborts.

0x14

VS_PROGRAM

[8:0]

Vertex program slot base (words, 0..511). Currently 0.

0x18

FS_PROGRAM

[8:0]

Fragment slot: 64/128/…/384 for modes 0..5. Driven by D-pad mode.

0x1C

CULL

[0]

Write-only backface-culling enable. Not readable.

0x50

COMMAND

wdata

1=CLEAR (fill color plus 0xFFFF depth), 2=DRAW_INDEXED, 3=PRESENT (vblank swap). Anything else raises error.

4.3. Uniform window (write-only)

Any address with addr[15:8]==1 writes one S18 lane:

u{addr[3:2]}[addr[7:4]] <= wdata[17:0];  // 16 vec4 uniforms u0..u3[0..15]

The demo host fills the 4 MVP rows as 0x100+component*4 (16 words per frame from assets/camera.mem) plus light and ambient constants at 0x140/0x144/0x148/0x14C. Shader LDU reads the addressed vec4.

4.4. Counters (read-only, reset-cleared)

Addr Name Counts

0x20

FRAMES

Presented frames

0x24

FRAME_CYCLES

gpu_clk cycles in the last frame (host overhead included)

0x28 / 0x2C / 0x30 / 0x34

VERTICES / TRIANGLES / CULLED / RASTERIZED

Geometry flow

0x38 / 0x3C / 0x40

FRAGMENTS / KILLED / SHADED

Early-Z kills versus shaded

0x44 / 0x48

TEXTURE_REQUESTS / INSTRUCTIONS

TEX2D executions / shader instructions retired

0x4C

PROG_LOADS

Committed UART shader uploads

Unmapped reads return 0. There is no interrupt, no DMA, no burst — by design.

FRAME_CYCLES plus the geometry and shading counters is how Performance attributes bottlenecks instead of guessing.

5. Programming model

One vector ISA, two stages (vertex and fragment share the core RTL, selected by program_base). Source of truth: rtl/gpu/shader_core.v, tools/pineasm.py, tools/reference_renderer.py:shader().

5.1. Machine model

  • 16 vector registers r0…r15, each 4x S18Q12 lanes ({x,y,z,w}).

  • Straight-line only: no branches, no predication.

  • Max 64 instructions per program (steps==63 faults without END); 512 program words total (8 slots x 64).

  • 32-bit encoding: op[31:26] dst[25:22] a[21:18] b[17:14] c[13:10] mask[9:6] uaddr[5:0]. LDI repurposes the low 18 bits as {mask[21:18], imm[17:0]}. OUT repurposes as {a[21:18], …, slot[1:0]}.

  • Per-instruction latency in RTL is ~5 cycles (fetch, operand, multiply, execute, writeback); TEX2D pays texture RAM latency on top.

  • The integer golden model is bit-exact: wrap-18 arithmetic, (a*b)>>12 truncation, same masks, same texture quantize.

Number format (S18Q12): signed 18-bit, 12 fractional bits, 1.0 = 4096:

\(wrap18(v) = ((v + 131072)\ mod\ 262144) - 131072\)

Multiply: \(rd = ((ra \cdot rb) >> 12)\) truncated, then wrapped.

5.2. Opcodes

# Mnemonic Form Semantics (per masked lane)

0

END

END

Halt, success. Missing END, unknown op>16, or steps==63 faults; DRAW aborts with error.

1

MOV

MOV rD[.m], rA

rD = rA

2

LDI

LDI rD[.m], imm

rD = s18(imm); assembler takes float x4096, range -131072…131071

3

LDU

LDU rD[.m], uN

rD = uniform[N] (whole vec4 addressed by uaddr)

4

ADD

ADD rD[.m], rA, rB

rD = rA+rB (wrapped)

5

SUB

SUB rD[.m], rA, rB

rD = rA-rB

6

MUL

MUL rD[.m], rA, rB

rD = (rA·rB)>>12

7

MAD

MAD rD[.m], rA, rB, rC

rD = ((rA·rB)>>12)+rC

8

DP3

DP3 rD[.m], rA, rB

3-lane dot, broadcast to masked lanes

9

DP4

DP4 rD[.m], rA, rB

4-lane dot, broadcast

10

MIN

MIN rD[.m], rA, rB

min

11

MAX

MAX rD[.m], rA, rB

max

12

SAT

SAT rD[.m], rA

clamp to 0…4096 (0.0…1.0)

13

CMP

CMP rD[.m], rA, rB

rA<rB ? 4096 : 0

14

SEL

SEL rD[.m], rA, rB, rC

rA!=0 ? rB : rC

15

TEX2D

TEX2D rD, rA

Sample 256x256 RGB444 at (u,v)=rA.xy (nearest, clamp); rD = {1.0, b', g', r'}. Counts texture_requests.

16

OUT

OUT slot, rA

slot 0→out0, 1→out1, 2→out2; any other slot faults.

Write mask .xyzw defaults to all lanes.

5.3. Stage ABIs (fixed by pineapple_gpu.v, not by shaders)

  • Vertex in: r0={1.0,x,y,z} r1={0,nx,ny,nz} r2={0,0,u,v} (+ uniforms). Vertex out: OUT 0 = clip p, OUT 1 = normal, OUT 2 = uv.

  • Fragment in: r0={1,0,v,u} r1={0,nz,ny,nx} r2={1,z,z,z}. Fragment out: OUT 0 = linear RGB (quantized to RGB444 at writeback).

  • Six showcase fragment slots (fs_base = (mode+1)*64): 0 diffuse, 1 normal, 2 toon, 3 unlit, 4 uv, 5 depth.

Example — canonical vertex shader (shaders/basic.vert):

LDU r3, u0      ; MVP row 0
DP4 r4.x, r3, r0
LDU r3, u1
DP4 r4.y, r3, r0
LDU r3, u2
DP4 r4.z, r3, r0
LDU r3, u3
DP4 r4.w, r3, r0
OUT 0, r4
OUT 1, r1
OUT 2, r2
END

5.4. Runtime shader upload

New shaders run on the FPGA with no resynthesis. tools/pineload.py assembles a source file and sends it over USB-UART (115200 8N1, RsRx pin C4); rtl/gpu/shader_loader.v stages the payload and commits it to instruction RAM atomically between frames.

Wire frame (PINE magic, LE16 base and count, payload, 8-bit checksum):

Field Bytes Notes

magic

50 49 4E 45 (PINE)

Exact match, driven from true-idle.

base

LE16

Program-RAM word address, 0..511.

count

LE16

Payload words, 1..512, base+count ⇐ 512.

payload

4 x count

Words little-endian.

checksum

1

Sum of every post-magic byte mod 256.

Slots in assets/program.mem (seven baked programs, eighth slot free):

Base Content Selected by

0

basic.vert (vertex)

fixed vs_base

64

diffuse.frag

D-pad mode 0

128

normal.frag

mode 1

192

toon.frag

mode 2

256

unlit.frag

mode 3

320

uv.frag

mode 4

384

depth.frag

mode 5

448

free

--base 448

Usage:

pip install pyserial                       # once
python3 tools/pineload.py shaders/mine.frag --mode 2

Select D-pad mode 2 first; the new fragment program takes effect on the next frame.

The serial port is write-only: the screen is the acknowledgement. A bad frame leaves program RAM untouched and sticks error until the next good commit or reset. Uploads survive mode switches only in their own slot; reprogramming the bitstream restores the baked image.

6. Verification

What has actually been checked, with what, and what remains open. Evidence lives in tb/, tools/, and docs/validation/.

6.1. Test layers

  • make test — fast unit suites (Icarus): board frames, keypad, camera control, framebuffer CDC and vblank swap, reciprocal divider, DRAW-error path, 96 randomized shader programs versus the integer reference, 24 rasterizer triangles under randomized backpressure, UART bytes/framing/ staging and a wired upload.

  • make verify — pixel-exact full-pipeline check (minutes): every RTL frame must match the integer golden model with zero mismatched pixels out of 57,600: cube shader modes 0–5 (yaw 5) plus the 1072-triangle pineapple showcase (reference CRC = RTL CRC = 63134D29).

  • make fpga / make program — synthesis, place-and-route, timing, and on-silicon HDMI confirmation (see Implementation).

make test     # fast unit suites (Icarus)
make verify   # pixel-exact RTL-vs-reference frames (minutes)

6.2. Golden model

tools/reference_renderer.py is integer-only and bit-exact with RTL — [s18, mul, edge, project] must be read as the specification, not as an approximation. tools/check_frame.py compares RTL frame dumps against it (CRC plus mismatch count); tools/generate_tests.py produces the randomized shader and raster vectors.

6.3. Silicon evidence (v0.2.0 and v0.3.0)

  • v0.2.0: make test 8/8 PASS; make verify PASS (cube modes 0–5 and pineapple, zero mismatches); make fpga PASS at 50 MHz post-route (gpu_clk 50.36 MHz); make program PASS, DONE set.

  • v0.3.0: rtl/gpu/ untouched; cube build PASS at 50 MHz post-route (gpu_clk 60.88 MHz); make program-cube PASS with matching mode-0 yaw-0 frame (CRC 641A94BF, 0 mismatches); pineapple rebuilt and re-confirmed on screen.

  • 2026-10-01 re-verification (Icarus 13.0, Apple Silicon): make test PASS (all 9 binaries), make verify PASS (showcase CRC 63134D29, 1,998,853 cycles; cube modes 0–5 CRC-matched).

Lit pineapple on silicon
Figure 2. Figure: lit textured pineapple on silicon (HDMI capture).

Logs: docs/validation/icarus.log, docs/validation/fpga-summary.log. Photos: docs/validation/hardware-pineapple.jpg, docs/validation/hardware-cube.jpg.

6.4. Bench sessions A1/A2 — pending

Builds: make program (pineapple), make program-cube (cube).

# Build Input Expected Photo Result

1

pineapple

power-on boot

lit textured pineapple on 0x013

2

pineapple

RIGHT x3

yaw steps, matches sim poses

3

pineapple

UP x2

pitch steps, matches sim poses

4

pineapple

CENTER+UP / CENTER+DOWN

zoom in / out

5

pineapple

CENTER tap modes 0–5

0 diffuse, 1 normal, 2 toon, 3 unlit, 4 uv, 5 depth

6

cube

boot + RIGHT x2 + modes 0,1

same behavior, cube asset

Live shader upload:

# Step Expected Result

1

mode 2 selected, python3 tools/pineload.py shaders/flat.frag --mode 2

flat-shaded frame, no resynthesis or reboot

2

CENTER-tap away and back to mode 2

upload persists in its slot

3

0x4C PROG_LOADS before/after

+1, error flag clear

Any photo diverging from the sim-predicted frame is a P1 defect: record the mode, yaw, and reference CRC alongside it.

7. Implementation

How the design is turned into a bitstream and what it costs. Target is fixed: xc7a100tcsg324-1 on the Nexys A7-100T.

7.1. Toolchain

Yosys synthesis → nextpnr-xilinx place-and-route → Project X-Ray bitstream, reusing the installed Tomato hardware/fpga/common.mk flow in place (TOMATO_FPGA override; see root README.md). No Vivado, no CPU firmware, no copied toolchain, no destructive setup targets.

make fpga          # synthesize + route TOP=pineapple_top
make program       # load volatile FPGA SRAM (no flash image)
make fpga-cube     # same flow with TOP=cube_top
make program-cube

Synthesis uses -nodsp: the installed nextpnr-xilinx cannot reliably route the mapped DSP carry ports, so 18x18 multiplies stay in fabric by choice, not by accident. Widths are unchanged, so timing — not function — is affected.

7.2. Clocks and reset

100 MHz oscillator on E3 → rtl/board/clock_reset.v → divided gpu_clk 50 MHz (GPU, host, keypad, UART) and pix_clk 25 MHz (scanout, DVI), distributed with BUFG. Reset is asserted asynchronously and released synchronously with the pixel clock running (CPU_RESETN on C12).

The fabric-divided clocks are proven but fragile (~0.7% margin on the pineapple build below). Moving to Vivado/MMCM is the first timing-robustness upgrade, not a functional one.

7.3. Pins

Clock, reset, five D-pad buttons, USB-UART RsRx on C4, and 12-bit TFP410 DVI PMOD across JC/JD only (constr/nexys.xdc). Unused peripheral constraints from the Tomato source are excluded; see docs/provenance.md.

7.4. Resources (measured post-route, never bit-counted)

Payload alone is 3,090,432 bits (2x 57,600x12 color + 57,600x16 depth
65,536x12 texture ≈ 84 RAMB36 at ideal packing). Port width and depth rounding plus index, vertex, shader, uniform, and camera tables cost the rest.

Measured v0.2.0: 22% LUT, 42% RAMB18, 41% RAMB36 of the 100T’s 135 RAMB36. Any capacity claim must cite a post-route report, never a bit total.

7.5. Timing (measured post-route)

  • Pineapple make fpga: gpu_clk 50.36 MHz vs 50 MHz target; pix_clk 151.10 MHz.

  • Cube make fpga-cube: gpu_clk 60.88 MHz, pix_clk 120.32 MHz.

  • Pineapple rebuild after parameter plumbing: gpu_clk 55.43 MHz.

Only the post-route number gates. Pre-route estimates (e.g. 40.03 MHz seen once) are informational.

8. Performance

Measured frame loop at 50 MHz (frame_cycles, host overhead included). Pose- and mode-dependent — fps is a measurement, not a spec.

8.1. Frame table (Icarus 13.0, 2026-10-01)

Workload Cycles fps Notes

cube yaw 0, mode 0 (diffuse)

1,359,877

36.8

face-on: 9,120 of 18,240 fragments early-Z killed

cube yaw 5, mode 0 (diffuse)

1,835,013

27.2

14,052 shaded x ~11 instr

cube yaw 5, mode 1 (normal)

1,409,029

35.5

14,052 shaded x ~5 instr

cube yaw 5, mode 2 (toon)

1,835,013

27.2

same cost class as diffuse

cube yaw 5, modes 3/4/5

1.20–1.28 M

39–42

cheapest shaders

pineapple showcase

1,998,853

25.0

1,072 tris, 12,780 shaded x ~14 instr

Hardware PRESENT can add up to one 16.7 ms vblank wait on top. 30 fps is a measured target, never a guarantee.

8.2. Bottleneck attribution (pineapple frame)

From counters plus FSM latencies:

  • Shader execution ~45% (179,172 instr x ~5 cycles — irreducible, it is the workload).

  • Divider-blocked reciprocal waits ~27% (16k fragment + 3.2k vertex + 1k setup waits x 32 cycles; CLEAR is a fixed 57,600).

  • Serial interpolation ~17% (7 attributes x 3 cycles x 16,165 fragments).

  • Raster stepping and host overhead the rest.

  • Early-Z kills 3,385 fragments before shading — without it the frame would cost ~25% more.

One structural win clears 30 fps: a pipelined reciprocal (2-cycle throughput instead of 32-cycle blocking) removes ~0.5 M cycles, giving pineapple ≈1.5 M ≈ 33 fps. It is provably bit-exact because every numerator is a site constant (2^24 for q, 2^30 for 1/area) — only denominators vary over enumerable ranges, so exhaustive equivalence simulation is feasible. Interpolation widening (~−11%) is the second lever, not the first.

9. Known limitations

What P1 does not do, what is fragile, and what changes first in P2. Read this before proposing optimizations.

9.1. Functional gaps

  • Near-plane clipping is not implemented — the demo camera is constrained to never cross it.

  • UART upload is write-only; the screen is the only acknowledgement.

  • Uploads persist per slot across mode switches but not across bitstream reprogramming (baked assets/program.mem returns).

  • CULL is write-only; several geometry registers are effectively fixed at 0 in the demo host configuration.

  • DRAW validation uses a serial mod-3 check (state 42, 14 cycles per DRAW) after % was found to infer a ~32 ns divider.

9.2. Timing and flow fragility

  • Fabric-divided clocks with ~0.7% margin on the pineapple build; Vivado/MMCM is the first robustness upgrade.

  • -nodsp synthesis works around a nextpnr DSP-carry routing limitation on this install; multiplies stay in fabric.

  • nextpnr ignores [current_design] configuration properties in the reused XDC; Project X-Ray falls back to its slower Python FASM parser. Neither prevents bitstream generation.

9.3. Verification debt

  • Physical D-pad confirmation (orbit, pitch, zoom, shader-mode cycling) and bench sessions A1/A2 remain open — see Verification.

  • Second-asset silicon proof is closed (cube, rtl/gpu/ untouched); the remaining gate is the physical pass, not simulation.

9.4. Correct P2 order

  1. Split pineapple_gpu.v FSM into the sketched stage modules with pixel-exact parity (make verify: zero mismatches, CRC 63134D29).

  2. Pipeline the reciprocal (throughput, not latency).

  3. Widen interpolation only against counter data (fragments/killed/shaded/instructions/texture_requests).

Each stage should be a separately reviewable change with passing Icarus tests and synthesis for hardware changes.