A 64-bit out-of-order superscalar RISC-V core (RV64IMAC target), from RTL to silicon feasibility.
Authority: rtl/ SystemVerilog wins over this prose.
Current facts: Introduction and Known Limitations.
Machine-readable ISA contract: docs/isa/rv64xo3.csv.
History in docs/log/ is never authoritative for current status.
1. Introduction
riscv64xO3 is a 64-bit superscalar out-of-order RISC-V project and core. Target ISA: RV64IMAC + Zicsr/Zifencei. The core fetches, decodes, issues, and commits two instructions per cycle, using Tomasulo reservation stations with in-order ROB commit. Native bus: AXI4-Lite. Privilege: M-mode with precise traps.
1.1. What this manual is
This manual is the detailed technical reference.
The root README.md remains the concise GitHub-facing project overview;
it does not duplicate this content.
1.2. Repository map
| Path | Purpose |
|---|---|
|
SystemVerilog source (synthesizable). Executable authority. |
|
Unit, cocotb, compliance, and formal testbenches |
|
This manual (source) plus ISA contract and history |
|
CRT, linker scripts, bare-metal examples |
|
Yosys synthesis and OpenLane feasibility flow |
|
Documentation guardrails, ISA generators |
|
Build and reporting scripts, including |
|
Development and CI container definitions |
1.3. Canonical vocabulary
-
riscv64xO3 is the 64-bit superscalar out-of-order RISC-V project and core.
-
RTL simulation means Verilator-driven simulation of
rtl/. Label its results as simulation, never silicon or board execution. -
Compliance means the
riscv-testsISA suites run against the RTL model. -
FPGA prototype means a future
hw/fpgamapping; it does not exist yet. -
ASIC feasibility means Yosys/OpenLane smoke reports in
asic/; it is not a tapeout.
1.4. Documentation authority
When sources disagree, this order wins:
-
rtl/SystemVerilog (the executable authority). -
Known Limitations and the facts stated in this introduction.
-
Machine-readable contracts (
docs/isa/rv64xo3.csv,tools/generators). -
Other prose in this manual.
-
docs/log/history (never authoritative for current status).
Status words are used exactly: present / in progress / planned / verified. Simulation, FPGA, and silicon claims are kept distinct.
Dated engineering history lives in docs/log/.
Do not cite a log entry for present status unless a current chapter confirms it.
|
1.5. Conventions used in this manual
mono marks file paths, signal names, and commands.
Cross references look like Architecture and resolve in both the
source view and the built site.
Admonitions mark non-obvious obligations:
| Start with this chapter and Known Limitations before trusting any number elsewhere. |
2. Architecture
Target: RV64IMAC + Zicsr/Zifencei, 2-wide superscalar out-of-order, Tomasulo with in-order commit. AXI4-Lite native. M-mode precise traps.
2.1. Pipeline overview
Five programmer-visible stages: fetch, decode/rename/dispatch, issue/execute, memory, commit. Mispredicts and traps roll back to the ROB head and redirect fetch.
The commit stage retires up to two instructions per cycle and is the only point where architectural state becomes visible. See Design for stage internals and Interfaces for bus behavior.
2.2. Top-level parameters
| Parameter | Value | Notes |
|---|---|---|
|
64 |
RV64 datapath |
|
2 |
Superscalar width |
|
64 |
In-order retire buffer |
|
96 |
Physical registers (32 arch {rightarrow} 96 phys) |
|
8 / 4 / 4 |
Reservation stations |
|
16 |
Load-store queue |
|
|
Core default ( |
2.3. Branch prediction
-
Bimodal predictor: 256-entry branch history table of 2-bit saturating counters, indexed by PC bits, initialized weakly not taken.
-
JALis always predicted taken with a computed target; conditional branches use the counter’s taken bit. -
Updates come from the execute stage; resolution happens in execute.
-
Mispredict flushes instructions younger than the ROB branch entry and restores the rename checkpoint. Source:
rtl/rv64xo3_bp.sv.
2.4. Memory and commit model
-
16-entry LSQ (scaffolding): ready/forward ports exist, but address matching, store-set prediction, and AXI issue are TODO.
-
AXI4-Lite instruction and data masters; data misalignment is split in hardware (see Design), misaligned fetch targets trap. See Interfaces.
-
CSR,
mret,ecall, andebreakretire at the ROB head only, giving precise M-mode traps (mepc/mcause/mtvalset, redirect tomtvec).
3. Design
This chapter describes the microarchitecture behind the pipeline in Architecture. All module paths are relative to the repository root.
3.1. Fetch (IF)
Sources: rtl/rv64xo3_if.sv, rtl/rv64xo3_bp.sv.
-
2-wide fetch on a 4-byte-aligned PC with an instruction buffer for downstream stalls. The fetch path is 32-bit today; RVC expansion is planned.
-
Bimodal direction prediction as described in Architecture.
-
Instruction memory reached through the AXI4-Lite master;
fetch_stallasserts on miss.
3.2. Decode, rename, dispatch (ID)
Sources: rtl/rv64xo3_id.sv, rtl/issue/rv64xo3_rename.sv, rtl/rv64xo3_decoder.sv.
-
2-wide decode driven by
docs/isa/rv64xo3.csv. Illegal encodings poison the ROB entry and raise a precise trap. -
Rename maps 32 architectural registers onto 96 physical registers using a map table plus freelist, with a checkpoint per branch.
-
Dispatch targets ALU (x8), multiply (x4), and LSU (x4) reservation stations and allocates ROB entries in order. Dispatch stalls when RS, ROB, or LSQ is full.
| Structure | Entries | Function |
|---|---|---|
ROB |
64 |
In-order commit, precise traps |
PRF |
96 x 64 bit |
Renamed data storage |
RS ALU / MUL / LSU |
8 / 4 / 4 |
Tomasulo wakeup and select |
LSQ |
16 |
Forwarding, ordering, commit drain |
3.3. Issue and execute (IS)
Sources: rtl/issue/rv64xo3_rs.sv, rtl/rv64xo3_ex.sv, rtl/rv64xo3_alu.sv, rtl/rv64xo3_muldiv.sv.
-
One single-cycle combinational ALU plus a shared multi-cycle multiply/divide unit (compute states with an iterative divider).
-
Reservation stations expose CDB tag/valid ports; tag RAM, ready bits, oldest-ready select, and broadcast are TODO scaffolding.
-
Branch resolution in execute triggers flush and map-table restore.
The ALU operation set (from rtl/include/rv64xo3_pkg.sv):
typedef enum logic [4:0] {
ALU_ADD, ALU_SUB, ALU_SLL, ALU_SLT, ALU_SLTU,
ALU_XOR, ALU_SRL, ALU_SRA, ALU_OR, ALU_AND,
ALU_ADDW, ALU_SUBW, ALU_SLLW, ALU_SRLW, ALU_SRAW
} alu_op_e;
3.4. Memory (MEM)
Sources: rtl/memory/rv64xo3_lsq.sv, rtl/rv64xo3_mem.sv.
-
The bus moves one 4-byte-aligned word per transaction. Accesses up to 8 bytes that cross the word boundary are split in hardware into at most two transactions, so data misalignment never traps.
-
No MEM exception source remains: bus faults are not modeled and the exception ports are tied off for future PMP/bus-error use.
-
The LSQ (
rtl/memory/rv64xo3_lsq.sv) is scaffolding: ready/forward ports exist, address CAM, store-set prediction, and AXI issue are TODO.
| Data misalignment is split in hardware. Only misaligned instruction-fetch targets trap. |
4. Interfaces
The core is AXI4-Lite native on both instruction and data ports.
There is no Wishbone in the core path; the legacy wb_to_axi4lite
bridge (rtl/wb_to_axi4lite.sv) is retained only for UART and debug
bring-up.
4.1. Bus architecture
-
Separate AXI4-Lite IMEM and DMEM masters from
rtl/rv64xo3_top.sv. -
Single clock domain with explicit reset strategy.
-
External memory interface only; no internal caches or TCM in this revision.
| Interface | Protocol | Notes |
|---|---|---|
IMEM master |
AXI4-Lite |
Instruction fetch, |
DMEM master |
AXI4-Lite |
LSQ spill/fill, store drain at commit |
UART bridge |
AXI4-Lite via |
Bring-up and |
4.2. Clocks, reset, and memory map
| Item | Value |
|---|---|
Clocking |
Single clock domain |
Reset |
Synchronous; core PC loads |
Privilege |
M-mode trap handling; supervisor CSR shadows and |
Misaligned data |
Split in hardware across bus transactions, never traps |
Misaligned fetch target |
Precise trap |
Bare-metal programs link against sw/link64.ld with the reset
vector at 0x8000_0000 and enter through sw/crt0_64.S.
See Implementation for the software flow.
|
4.3. Debug and UART
-
tb/tb_uart_axi.svandrtl/uart_axi.svcover the UART path. -
The
hello.cend-to-end demo exercises UART output through the bridge. -
Handshake, mutual-exclusion, and trap-entry assertions are documented in Verification.
5. Programming Model
5.1. ISA support
Machine-readable contract: docs/isa/rv64xo3.csv
(columns: opcode,funct3,funct7/instr,name,class,description).
| Extension | Status | Notes |
|---|---|---|
RV64I (incl. W-suffixed word ops) |
present |
|
M (incl. W ops) |
present |
|
Zicsr / Zifencei |
present |
CSR access, fences |
A |
in progress |
Bring-up; see Known Limitations |
C |
in progress |
16-bit parcels via RVC expander (bring-up) |
The full opcode map covers integer computation (immediate and
register-register), control transfer, loads/stores, AUIPC/LUI,
system, and fence encodings. Illegal encodings trap precisely.
FENCE and FENCE.I decode as NOPs (single in-order pipe, no caches).
5.2. Control and status registers
| CSR | Address | Purpose |
|---|---|---|
|
|
Machine status |
|
|
ISA identity |
|
|
Interrupt enable / pending |
|
|
Trap vector |
|
|
Machine scratch |
|
|
Trap state |
|
|
Counters |
|
|
Identity |
CSR instructions CSRRW/CSRRS/CSRRC and immediates are supported,
including masked writes. mret restores the pre-trap PC; sret is
supported alongside the supervisor CSR shadows. misa reports MXL=2
with I+M only.
5.3. Traps and exceptions
| Cause | Meaning |
|---|---|
0 |
Instruction address misaligned |
1 |
Instruction access fault |
2 |
Illegal instruction |
3 |
Breakpoint |
4 |
Load address misaligned |
5 |
Load access fault |
6 |
Store address misaligned |
7 |
Store access fault |
11 |
Environment call from M-mode |
ecall, ebreak, CSR operations, and mret retire at the ROB head,
so trap entry updates the PC and mret restores it atomically.
6. Verification
Multi-layered strategy: block-level unit tests, cocotb integration,
and riscv-tests compliance against the RTL model, all cycle-accurate
under Verilator.
6.1. Strategy
| Layer | Tools | Scope |
|---|---|---|
Unit (pytest) |
|
ALU, decoder, CSR, MulDiv, memory in isolation |
Integration (cocotb) |
|
Pipeline interaction, hazards, bus models |
Compliance |
|
Official |
Formal / SVA |
|
14 assert properties: x0, PC/mutual-exclusion, handshake liveness ( |
Directed tests cover corner cases (divide-by-zero, mixed signs, misaligned traps, masked CSR writes); randomized tests cover arithmetic and constrained-random instruction streams. Bus-functional models drive the AXI4-Lite interfaces.
6.2. Current results
Phase 0 baseline (docs/log/2026-10-01-phase0-foundation.md,
riscv-tests submodule @ 7af6beb):
-
make compliance(RV64 default): RV64UI 54/54, RV64UM 13/13, RV64MI 15/15 (includesma_addr,ma_fetch, misaligned family,sbreak). -
Legacy baselines: RV32UI 37/37 + RV32UM 8/8 (
make compliance-rv32), legacy asm 8/8, C++ harness suites 15/15. -
make unit: 58 passed.make cocotb: 7 passed (ALU, decoder, CSR, memory, MulDiv, AXI IMEM+DMEM, top). -
Line/toggle coverage is collected with Verilator and tracked toward PRD targets (>=95% line, >=80% toggle); regenerate with
scripts/generate_coverage_report.pyrather than quoting old figures.
| Numbers here describe RTL simulation only. See Known Limitations for what simulation evidence does and does not prove. |
6.3. Running the suites
make lint # Verible + Verilator lint
make unit # pytest module tests
make cocotb-smoke # fast integration subset
make cocotb # full integration
make compliance # RV64 riscv-tests via Verilator sim
make regress # lint + unit + cocotb + compliance
Waveforms land in waves/ as .fst; view them with:
make waves # requires GTKWave on the host
7. Implementation
7.1. Getting started
Two flows are supported. Docker is recommended for exact tool versions.
Option A — Docker (recommended):
make docker-build
make docker-shell
# inside the container:
make regress
Option B — local install (Linux/macOS):
brew install verilator # macOS
# or: sudo apt install verilator # Ubuntu
pip install -r requirements.txt
# riscv64-unknown-elf-gcc must be on PATH
Focused checks after setup:
python3 tools/check_docs.py
make lint
make unit cocotb-smoke
7.2. Build and test commands
| Command | Effect |
|---|---|
|
Verible + Verilator lint |
|
Build the Verilator simulator binary |
|
Unit / integration tests |
|
RV64 compliance (RV32 kept as |
|
Full regression: lint + all test layers |
|
Yosys synthesis smoke |
|
OpenLane feasibility (needs PDK image) |
|
Build this manual into |
Generated artifacts go to build/ (simulation, logs) and waves/
(waveforms). Both are gitignored and never committed.
7.3. Synthesis and ASIC feasibility
Synthesis status (Yosys, rv64xo3_top).
| Aspect | Status | Notes |
|---|---|---|
RTL synthesizability |
verified |
Yosys completes; no latches or combinational loops |
FPGA target |
Artix-7 family |
Vendor run pending; LUT/FF counts estimated only |
ASIC feasibility |
Sky130 / OpenLane2 |
Flow configured ( |
Clocking |
single domain |
Explicit synchronous reset |
make synth # Yosys smoke (see scripts/synth/)
make asic-feas # OpenLane feasibility (see asic/)
Estimates only, pending real runs: core area on the order of 0.1-0.3 mm^2 in Sky130, 50-150 MHz post-layout, 5-8k LUTs on Artix-7. Do not quote these as measured results.
8. Performance
All figures below are measured on Verilator RTL simulation unless noted. They describe the model, not silicon or a board.
8.1. Dhrystone 2.1
Config: 100 runs, GCC 15.1.0 -O2, RV64IM_Zicsr.
| Metric | Value |
|---|---|
Cycles per run |
1174 |
DMIPS/MHz |
0.48 |
IPC |
TBD (counter issue in the measured run) |
The score follows the standard definition:
Current performance is bounded by the absence of caches: instruction fetch and data access over AXI4-Lite cost multiple cycles per access. A dedicated instruction cache or tightly coupled memory would move IPC toward the theoretical 2-wide ceiling.
8.2. Other benchmarks
-
CoreMark: pending porting.
-
FPGA synthesis on Artix-7: pending vendor run.
-
Post-layout frequency and power: TBD after a full OpenLane run.
8.3. Resource and timing estimates
| Item | Estimate basis |
|---|---|
LUTs / FFs |
~5-8k / ~2-3k from RTL complexity (5-stage + OoO scaffolding) |
Block RAM / DSP |
Minimal / 0-2 depending on multiplier mapping |
Target clock |
50-100 MHz conservative (Artix-7); 50-150 MHz post-layout (Sky130) |
Critical path |
Likely ALU or hazard detection |
9. Known Limitations
This chapter bounds what the repository proves. It outranks any optimistic reading of earlier chapters.
9.1. Status boundaries
-
Source-present: in-order 5-stage pipeline renamed to
rv64xo3_*; out-of-order issue/commit scaffolding present; full Tomasulo in progress. -
Testable without hardware:
make lint unit cocotb-smoke compliance. -
Hardware-dependent: FPGA mapping, board boot, and silicon require live evidence. Do not infer them from source or simulation.
9.2. Open items
-
A and C extensions are in bring-up, not verified.
-
No PMP unit (
pmpaddrunimplemented). -
No instruction/data caches; AXI4-Lite latency bounds IPC.
-
Misaligned data accesses are handled in hardware (split bus transactions, no trap); misaligned instruction-fetch targets still trap.
-
FPGA prototype mapping does not exist yet.
-
ASIC work is feasibility smoke (Yosys + OpenLane2), not a tapeout.
-
CoreMark and FPGA synthesis numbers are pending.
| Never relabel simulation, feasibility, or estimates as silicon, board, or tapeout results. |
9.3. History and policy
Dated decisions live in docs/log/ (YYYY-MM-DD-topic.md with what was
run and observed). Log entries are historical and never override this
chapter. The standing authority order is rtl/, then this manual’s
current facts, then machine-readable contracts, then other prose, then
history.