A 64-bit out-of-order superscalar RISC-V core (RV64IMAC target), from RTL to silicon feasibility.

Authority: rtl/ SystemVerilog wins over this prose. Current facts: Introduction and Known Limitations. Machine-readable ISA contract: docs/isa/rv64xo3.csv. History in docs/log/ is never authoritative for current status.

1. Introduction

riscv64xO3 is a 64-bit superscalar out-of-order RISC-V project and core. Target ISA: RV64IMAC + Zicsr/Zifencei. The core fetches, decodes, issues, and commits two instructions per cycle, using Tomasulo reservation stations with in-order ROB commit. Native bus: AXI4-Lite. Privilege: M-mode with precise traps.

1.1. What this manual is

This manual is the detailed technical reference. The root README.md remains the concise GitHub-facing project overview; it does not duplicate this content.

1.2. Repository map

Table 1. Repository layout. Use it to locate evidence for every claim.
Path Purpose

rtl/

SystemVerilog source (synthesizable). Executable authority.

tb/

Unit, cocotb, compliance, and formal testbenches

docs/

This manual (source) plus ISA contract and history

sw/

CRT, linker scripts, bare-metal examples

asic/

Yosys synthesis and OpenLane feasibility flow

tools/

Documentation guardrails, ISA generators

scripts/

Build and reporting scripts, including build-docs.sh

docker/

Development and CI container definitions

1.3. Canonical vocabulary

  • riscv64xO3 is the 64-bit superscalar out-of-order RISC-V project and core.

  • RTL simulation means Verilator-driven simulation of rtl/. Label its results as simulation, never silicon or board execution.

  • Compliance means the riscv-tests ISA suites run against the RTL model.

  • FPGA prototype means a future hw/fpga mapping; it does not exist yet.

  • ASIC feasibility means Yosys/OpenLane smoke reports in asic/; it is not a tapeout.

1.4. Documentation authority

When sources disagree, this order wins:

  1. rtl/ SystemVerilog (the executable authority).

  2. Known Limitations and the facts stated in this introduction.

  3. Machine-readable contracts (docs/isa/rv64xo3.csv, tools/ generators).

  4. Other prose in this manual.

  5. docs/log/ history (never authoritative for current status).

Status words are used exactly: present / in progress / planned / verified. Simulation, FPGA, and silicon claims are kept distinct.

Dated engineering history lives in docs/log/. Do not cite a log entry for present status unless a current chapter confirms it.

1.5. Conventions used in this manual

mono marks file paths, signal names, and commands. Cross references look like Architecture and resolve in both the source view and the built site. Admonitions mark non-obvious obligations:

Start with this chapter and Known Limitations before trusting any number elsewhere.

2. Architecture

Target: RV64IMAC + Zicsr/Zifencei, 2-wide superscalar out-of-order, Tomasulo with in-order commit. AXI4-Lite native. M-mode precise traps.

2.1. Pipeline overview

Five programmer-visible stages: fetch, decode/rename/dispatch, issue/execute, memory, commit. Mispredicts and traps roll back to the ROB head and redirect fetch.

Pipeline stages
Figure 1. Pipeline stages and rollback path.

The commit stage retires up to two instructions per cycle and is the only point where architectural state becomes visible. See Design for stage internals and Interfaces for bus behavior.

2.2. Top-level parameters

Table 2. Parameters from rtl/include/rv64xo3_pkg.sv.
Parameter Value Notes

XLEN

64

RV64 datapath

FETCH_WIDTH / COMMIT_WIDTH

2

Superscalar width

ROB_SZ

64

In-order retire buffer

PRF_SZ

96

Physical registers (32 arch {rightarrow} 96 phys)

RS_ALU_SZ / RS_MUL_SZ / RS_LSU_SZ

8 / 4 / 4

Reservation stations

LSQ_SZ

16

Load-store queue

RESET_PC

0x8000_0000

Core default (rv64xo3_top/rv64xo3_if parameter; the AXI system wrapper re-parameterizes it)

2.3. Branch prediction

  • Bimodal predictor: 256-entry branch history table of 2-bit saturating counters, indexed by PC bits, initialized weakly not taken.

  • JAL is always predicted taken with a computed target; conditional branches use the counter’s taken bit.

  • Updates come from the execute stage; resolution happens in execute.

  • Mispredict flushes instructions younger than the ROB branch entry and restores the rename checkpoint. Source: rtl/rv64xo3_bp.sv.

2.4. Memory and commit model

  • 16-entry LSQ (scaffolding): ready/forward ports exist, but address matching, store-set prediction, and AXI issue are TODO.

  • AXI4-Lite instruction and data masters; data misalignment is split in hardware (see Design), misaligned fetch targets trap. See Interfaces.

  • CSR, mret, ecall, and ebreak retire at the ROB head only, giving precise M-mode traps (mepc/mcause/mtval set, redirect to mtvec).

3. Design

This chapter describes the microarchitecture behind the pipeline in Architecture. All module paths are relative to the repository root.

3.1. Fetch (IF)

Sources: rtl/rv64xo3_if.sv, rtl/rv64xo3_bp.sv.

  • 2-wide fetch on a 4-byte-aligned PC with an instruction buffer for downstream stalls. The fetch path is 32-bit today; RVC expansion is planned.

  • Bimodal direction prediction as described in Architecture.

  • Instruction memory reached through the AXI4-Lite master; fetch_stall asserts on miss.

3.2. Decode, rename, dispatch (ID)

Sources: rtl/rv64xo3_id.sv, rtl/issue/rv64xo3_rename.sv, rtl/rv64xo3_decoder.sv.

  • 2-wide decode driven by docs/isa/rv64xo3.csv. Illegal encodings poison the ROB entry and raise a precise trap.

  • Rename maps 32 architectural registers onto 96 physical registers using a map table plus freelist, with a checkpoint per branch.

  • Dispatch targets ALU (x8), multiply (x4), and LSU (x4) reservation stations and allocates ROB entries in order. Dispatch stalls when RS, ROB, or LSQ is full.

Table 3. Register and queue sizing.
Structure Entries Function

ROB

64

In-order commit, precise traps

PRF

96 x 64 bit

Renamed data storage

RS ALU / MUL / LSU

8 / 4 / 4

Tomasulo wakeup and select

LSQ

16

Forwarding, ordering, commit drain

3.3. Issue and execute (IS)

Sources: rtl/issue/rv64xo3_rs.sv, rtl/rv64xo3_ex.sv, rtl/rv64xo3_alu.sv, rtl/rv64xo3_muldiv.sv.

  • One single-cycle combinational ALU plus a shared multi-cycle multiply/divide unit (compute states with an iterative divider).

  • Reservation stations expose CDB tag/valid ports; tag RAM, ready bits, oldest-ready select, and broadcast are TODO scaffolding.

  • Branch resolution in execute triggers flush and map-table restore.

The ALU operation set (from rtl/include/rv64xo3_pkg.sv):

typedef enum logic [4:0] {
  ALU_ADD, ALU_SUB, ALU_SLL, ALU_SLT, ALU_SLTU,
  ALU_XOR, ALU_SRL, ALU_SRA, ALU_OR, ALU_AND,
  ALU_ADDW, ALU_SUBW, ALU_SLLW, ALU_SRLW, ALU_SRAW
} alu_op_e;

3.4. Memory (MEM)

Sources: rtl/memory/rv64xo3_lsq.sv, rtl/rv64xo3_mem.sv.

  • The bus moves one 4-byte-aligned word per transaction. Accesses up to 8 bytes that cross the word boundary are split in hardware into at most two transactions, so data misalignment never traps.

  • No MEM exception source remains: bus faults are not modeled and the exception ports are tied off for future PMP/bus-error use.

  • The LSQ (rtl/memory/rv64xo3_lsq.sv) is scaffolding: ready/forward ports exist, address CAM, store-set prediction, and AXI issue are TODO.

Data misalignment is split in hardware. Only misaligned instruction-fetch targets trap.

3.5. Commit (CM)

Source: rtl/commit/rv64xo3_rob.sv.

  • In-order retirement of up to two instructions per cycle.

  • Precise traps flush younger entries and redirect fetch to mtvec.

4. Interfaces

The core is AXI4-Lite native on both instruction and data ports. There is no Wishbone in the core path; the legacy wb_to_axi4lite bridge (rtl/wb_to_axi4lite.sv) is retained only for UART and debug bring-up.

4.1. Bus architecture

  • Separate AXI4-Lite IMEM and DMEM masters from rtl/rv64xo3_top.sv.

  • Single clock domain with explicit reset strategy.

  • External memory interface only; no internal caches or TCM in this revision.

Table 4. Core memory interfaces.
Interface Protocol Notes

IMEM master

AXI4-Lite

Instruction fetch, fetch_stall on miss

DMEM master

AXI4-Lite

LSQ spill/fill, store drain at commit

UART bridge

AXI4-Lite via wb_to_axi4lite

Bring-up and hello.c demo only

4.2. Clocks, reset, and memory map

Table 5. Reset and address behavior.
Item Value

Clocking

Single clock domain

Reset

Synchronous; core PC loads RESET_PC (default 0x8000_0000, re-parameterizable per wrapper)

Privilege

M-mode trap handling; supervisor CSR shadows and sret present

Misaligned data

Split in hardware across bus transactions, never traps

Misaligned fetch target

Precise trap

Bare-metal programs link against sw/link64.ld with the reset vector at 0x8000_0000 and enter through sw/crt0_64.S. See Implementation for the software flow.

4.3. Debug and UART

  • tb/tb_uart_axi.sv and rtl/uart_axi.sv cover the UART path.

  • The hello.c end-to-end demo exercises UART output through the bridge.

  • Handshake, mutual-exclusion, and trap-entry assertions are documented in Verification.

5. Programming Model

5.1. ISA support

Machine-readable contract: docs/isa/rv64xo3.csv (columns: opcode,funct3,funct7/instr,name,class,description).

Table 6. ISA status. RV64IM is implemented; A/C are in bring-up.
Extension Status Notes

RV64I (incl. W-suffixed word ops)

present

ADDW/SUBW/SLLW/SRLW/SRAW, LD/SD, LWU

M (incl. W ops)

present

MUL[H[SU][U]]/DIV[U]/REM[U], MULW/DIVW[D]/REMW[D]

Zicsr / Zifencei

present

CSR access, fences

A

in progress

Bring-up; see Known Limitations

C

in progress

16-bit parcels via RVC expander (bring-up)

The full opcode map covers integer computation (immediate and register-register), control transfer, loads/stores, AUIPC/LUI, system, and fence encodings. Illegal encodings trap precisely. FENCE and FENCE.I decode as NOPs (single in-order pipe, no caches).

5.2. Control and status registers

Table 7. Minimum CSR set (rtl/include/rv64xo3_pkg.sv).
CSR Address Purpose

mstatus

0x300

Machine status

misa

0x301

ISA identity

mie / mip

0x304 / 0x344

Interrupt enable / pending

mtvec

0x305

Trap vector

mscratch

0x340

Machine scratch

mepc / mcause / mtval

0x341 / 0x342 / 0x343

Trap state

cycle / instret

0xC00 / 0xC02

Counters

mvendorid / marchid / mimpid / mhartid

0xF11-0xF14

Identity

CSR instructions CSRRW/CSRRS/CSRRC and immediates are supported, including masked writes. mret restores the pre-trap PC; sret is supported alongside the supervisor CSR shadows. misa reports MXL=2 with I+M only.

5.3. Traps and exceptions

Table 8. Precise M-mode trap causes (mcause values).
Cause Meaning

0

Instruction address misaligned

1

Instruction access fault

2

Illegal instruction

3

Breakpoint

4

Load address misaligned

5

Load access fault

6

Store address misaligned

7

Store access fault

11

Environment call from M-mode

ecall, ebreak, CSR operations, and mret retire at the ROB head, so trap entry updates the PC and mret restores it atomically.

6. Verification

Multi-layered strategy: block-level unit tests, cocotb integration, and riscv-tests compliance against the RTL model, all cycle-accurate under Verilator.

6.1. Strategy

Table 9. Verification layers and entry points.
Layer Tools Scope

Unit (pytest)

tb/unit/

ALU, decoder, CSR, MulDiv, memory in isolation

Integration (cocotb)

tb/cocotb/

Pipeline interaction, hazards, bus models

Compliance

tb/compliance/Makefile.rv64

Official riscv-tests RV64 suites

Formal / SVA

rtl/rv64xo3_top.sv, rtl/rv64xo3_hazard_sva.sv

14 assert properties: x0, PC/mutual-exclusion, handshake liveness (tb/formal/ reserved, empty)

Directed tests cover corner cases (divide-by-zero, mixed signs, misaligned traps, masked CSR writes); randomized tests cover arithmetic and constrained-random instruction streams. Bus-functional models drive the AXI4-Lite interfaces.

6.2. Current results

Phase 0 baseline (docs/log/2026-10-01-phase0-foundation.md, riscv-tests submodule @ 7af6beb):

  • make compliance (RV64 default): RV64UI 54/54, RV64UM 13/13, RV64MI 15/15 (includes ma_addr, ma_fetch, misaligned family, sbreak).

  • Legacy baselines: RV32UI 37/37 + RV32UM 8/8 (make compliance-rv32), legacy asm 8/8, C++ harness suites 15/15.

  • make unit: 58 passed. make cocotb: 7 passed (ALU, decoder, CSR, memory, MulDiv, AXI IMEM+DMEM, top).

  • Line/toggle coverage is collected with Verilator and tracked toward PRD targets (>=95% line, >=80% toggle); regenerate with scripts/generate_coverage_report.py rather than quoting old figures.

Numbers here describe RTL simulation only. See Known Limitations for what simulation evidence does and does not prove.

6.3. Running the suites

make lint            # Verible + Verilator lint
make unit            # pytest module tests
make cocotb-smoke    # fast integration subset
make cocotb          # full integration
make compliance      # RV64 riscv-tests via Verilator sim
make regress         # lint + unit + cocotb + compliance

Waveforms land in waves/ as .fst; view them with:

make waves   # requires GTKWave on the host

6.4. Quality gates

  • Verible and Verilator lint clean: no inferred latches, no combinational loops, all registers explicitly clocked.

  • Pre-commit hooks run the documentation guardrails (python3 tools/check_docs.py) and YAML parsing.

7. Implementation

7.1. Getting started

Two flows are supported. Docker is recommended for exact tool versions.

Option A — Docker (recommended):

make docker-build
make docker-shell
# inside the container:
make regress

Option B — local install (Linux/macOS):

brew install verilator            # macOS
# or: sudo apt install verilator  # Ubuntu
pip install -r requirements.txt
# riscv64-unknown-elf-gcc must be on PATH

Focused checks after setup:

python3 tools/check_docs.py
make lint
make unit cocotb-smoke

7.2. Build and test commands

Table 10. Canonical make targets.
Command Effect

make lint

Verible + Verilator lint

make sim

Build the Verilator simulator binary

make unit / make cocotb[-smoke]

Unit / integration tests

make compliance

RV64 compliance (RV32 kept as compliance-rv32 baseline)

make regress

Full regression: lint + all test layers

make synth

Yosys synthesis smoke

make asic-feas

OpenLane feasibility (needs PDK image)

make docs

Build this manual into build/docs/

Generated artifacts go to build/ (simulation, logs) and waves/ (waveforms). Both are gitignored and never committed.

7.3. Synthesis and ASIC feasibility

Synthesis status (Yosys, rv64xo3_top).

Aspect Status Notes

RTL synthesizability

verified

Yosys completes; no latches or combinational loops

FPGA target

Artix-7 family

Vendor run pending; LUT/FF counts estimated only

ASIC feasibility

Sky130 / OpenLane2

Flow configured (asic/); not a tapeout

Clocking

single domain

Explicit synchronous reset

make synth        # Yosys smoke (see scripts/synth/)
make asic-feas    # OpenLane feasibility (see asic/)

Estimates only, pending real runs: core area on the order of 0.1-0.3 mm^2 in Sky130, 50-150 MHz post-layout, 5-8k LUTs on Artix-7. Do not quote these as measured results.

7.4. Software flow

  • sw/crt0_64.S startup, sw/link64.ld link script, sw/uart.h driver.

  • Examples: sw/hello.c (UART end-to-end), sw/test_misalign.c.

  • 32-bit crt0.S/link.ld retained for the legacy RV32 baseline.

Rebuild software images with make -C sw; never commit generated .elf/.hex/.mem files.

8. Performance

All figures below are measured on Verilator RTL simulation unless noted. They describe the model, not silicon or a board.

8.1. Dhrystone 2.1

Config: 100 runs, GCC 15.1.0 -O2, RV64IM_Zicsr.

Table 11. Measured results.
Metric Value

Cycles per run

1174

DMIPS/MHz

0.48

IPC

TBD (counter issue in the measured run)

The score follows the standard definition:

\[\mathrm{DMIPS/MHz} = \frac{10^6 / \mathrm{cycles\_per\_run}}{1757}\]

Current performance is bounded by the absence of caches: instruction fetch and data access over AXI4-Lite cost multiple cycles per access. A dedicated instruction cache or tightly coupled memory would move IPC toward the theoretical 2-wide ceiling.

8.2. Other benchmarks

  • CoreMark: pending porting.

  • FPGA synthesis on Artix-7: pending vendor run.

  • Post-layout frequency and power: TBD after a full OpenLane run.

8.3. Resource and timing estimates

Table 12. Pending measurement; do not quote as results.
Item Estimate basis

LUTs / FFs

~5-8k / ~2-3k from RTL complexity (5-stage + OoO scaffolding)

Block RAM / DSP

Minimal / 0-2 depending on multiplier mapping

Target clock

50-100 MHz conservative (Artix-7); 50-150 MHz post-layout (Sky130)

Critical path

Likely ALU or hazard detection

9. Known Limitations

This chapter bounds what the repository proves. It outranks any optimistic reading of earlier chapters.

9.1. Status boundaries

  • Source-present: in-order 5-stage pipeline renamed to rv64xo3_*; out-of-order issue/commit scaffolding present; full Tomasulo in progress.

  • Testable without hardware: make lint unit cocotb-smoke compliance.

  • Hardware-dependent: FPGA mapping, board boot, and silicon require live evidence. Do not infer them from source or simulation.

9.2. Open items

  • A and C extensions are in bring-up, not verified.

  • No PMP unit (pmpaddr unimplemented).

  • No instruction/data caches; AXI4-Lite latency bounds IPC.

  • Misaligned data accesses are handled in hardware (split bus transactions, no trap); misaligned instruction-fetch targets still trap.

  • FPGA prototype mapping does not exist yet.

  • ASIC work is feasibility smoke (Yosys + OpenLane2), not a tapeout.

  • CoreMark and FPGA synthesis numbers are pending.

Never relabel simulation, feasibility, or estimates as silicon, board, or tapeout results.

9.3. History and policy

Dated decisions live in docs/log/ (YYYY-MM-DD-topic.md with what was run and observed). Log entries are historical and never override this chapter. The standing authority order is rtl/, then this manual’s current facts, then machine-readable contracts, then other prose, then history.