FPU
Tuesday 28 May 2024, by // CPU 63F09
The floating-point unit 63F39
The 63F09’s FPU is split cleanly in two: a 16-level register stack hosted in the CPU core itself, and a separate, stateless execution engine (referred to as 63F39) that the core drives through a small request/response interface. The core owns all FPU state; the engine only computes.
A 16-level, RPN-style operand stack, not a flat register file
REG_FP is sixteen 64-bit registers, REG_FP(0) to REG_FP(15) — but software addresses them as stack levels, not fixed register numbers: REG_FP(0) is always the top of the stack. Every FPU instruction ultimately drives one value into FP_STACK_CTRL (type FP_STACK_TYPE), decoded in a single process on the falling edge of the internal clock, and the whole array shifts as one unit:
- Push — every register moves up one slot (
REG_FP(i+1)<=REG_FP(i)forifrom 14 downto 0) and the new value lands inREG_FP(0). Used to load an operand from memory (LOAD_FROM_MEM_DATA) or from the integer side of the core (LOAD_FROM_O, moving the 64-bit pseudo registerO = Q:Vacross). - Pop — every register moves down one slot, the vacated top entries are cleared to zero (
DROP_FP, and its N-level generalizationDROPN_FP). - Stack juggling —
DUP_FP/DUP2_FP/SWAP_FP/OVER_FP/ROT_FPcover the fixed, one- or two-level Forth-style primitives;DUPN_FP/DROPN_FP/ROLL_FP/ROLLD_FP/PICK_FP/EXG_FPgeneralize duplication, dropping, rotation and exchange to an arbitrary level, andCLEAR_FPempties the whole stack.
Feeding the execution engine
An operation reads its operands directly off the top of this stack. LOAD_FPU_ARG1/LOAD_FPU_ARG2/LOAD_FPU_ARG3 copy the top one, two or three levels into the engine’s own input latches, and START_FPU raises FPU_IN.FP_EXE_I.ENABLE to actually launch the computation. Once the engine reports back, LOAD_FROM_FPU takes the 64-bit RESULT and the 5-bit IEEE exception FLAGS, and LOAD_CCFPU updates just the flags without touching the stack at all — the path a comparison uses (see below).
Whether consuming a value counts as a real pop is a single bit, FP_STACK_CONF (FP_STACK_CONF_KEEP_ARGS vs FP_STACK_CONF_DISCARD_ARGS, itself selected by one bit of the opcode): with discard, LOAD_FPU_ARGn shifts the consumed levels away and LOAD_FROM_FPU pushes the result on top — a normal, stack-popping operation, net stack effect one level shallower per result produced. With keep, neither shift happens at all: LOAD_FPU_ARGn reads the operands in place and LOAD_FROM_FPU simply overwrites REG_FP(0) with the result — the stack depth is unchanged, and any operand below the top (REG_FP(1) and beyond) survives untouched for a following operation to reuse. This is what lets a value outlive the operation that reads it, without a manualDUP beforehand.
The execution engine itself
fp_exe_in_type/fp_exe_out_type (fp_wire.vhd) define the engine’s own interface: three 64-bit operands, a one-hot fp_operation_type selector (fmadd/fmsub/fnmadd/fnmsub, fadd/fsub/fmul/fdiv/fsqrt, fsgnj, fcmp, fmax, fclass, fmv_i2f/fmv_f2i, fcvt_f2f/fcvt_i2f/fcvt_f2i plus a 2-bit fcvt_op sub-selector), a 2-bit format field and a 3-bit rounding-mode field — the same shape as a RISC-V “F/D” extension unit. The engine itself carries no stack, no state beyond one computation: the CPU-side stack machine described above is entirely what turns this generic, three-operand functional unit into the 63F09’s own Forth-style FPU instruction set.
Private or shared: the FPU_BY_CORE choice
Everything above describes the interface between a core and an FPU engine — but on a multi-core build, that engine can be instantiated two different ways, selected by a single generic, FPU_BY_CORE, threaded down from the SoC top level through 63F81.vhd/63F29.vhd/ 63F09.vhd:
- Per-core (
FPU_BY_CORE=true) — each core gets its own privateFP_UNITinstance, declared right inside the CPU core’s own architecture (63F09.vhd,FPU_INTERFACE : if FPU_BY_CORE generate). No arbitration is needed at all: a core’sREG_FPstack talks directly to an engine that belongs to it alone. - Shared (
FPU_BY_CORE=false) — a singleFP_UNITinstance lives at the top level instead (63F81.vhd, generate blockSHARED_FPU), clocked by the chip-levelE— the current bus master’s ownE, from the same arbiter described in the companion document. Every core still drives its ownFPU_IN(i)entry (a per-core vector,fp_unit_in_vector_type) exactly as if it owned the engine outright, but only one entry at a time is actually multiplexed into the real unit’s input (FPU_MUX_IN <= FPU_IN (CURRENT_CPU)), and the engine’s single output,FPU_OUT, is wired back identically to every core’sFPU_63F39_OUTport — the same signal, broadcast. A core only sees a meaningful result once the arbiter has actually picked its own request.
The arbiter that picks CURRENT_CPU mirrors the bus arbiter almost line for line: on each falling edge of E, a round robin starting from the currently-selected core (FPU_ACTIVE_CPU) looks for the first core whose FP_EXE_I.ENABLE is asserted, and sticks with whichever core it finds for as long as that core keeps its own ENABLE raised — the same “sticky, self-first” pattern as SIGNAL_LAST_ACTIVE_CPU for the bus, just applied to the FPU request instead of the bus request. A core that loses the race simply waits, its own ENABLE still high, until the current holder releases the engine and the round robin reaches it.
Comparisons: result goes to the CPU’s own flags, not the stack
fcmp is the one operation that does not produce a floating-point value: it feeds LOAD_CCFPU instead of LOAD_FROM_FPU, so the comparison’s outcome only ever updates REG_CCFPU’s exception flags — the actual true/false predicate is delivered to the CPU’s own condition codes (Z), the same register a plain integer CMPx would set. A comparing instruction therefore never touches the operand stack itself; a non-popping variant leaves both compared values exactly where they were.
Status and rounding mode: REG_CCFPU
REG_CCFPU is a single 8-bit register, split in two: bits 7-5 hold the current IEEE-754 rounding mode, set by six dedicated stack-control values (SET_RNE_FP/SET_RTZ_FP/SET_RDN_FP/SET_RUP_FP/ SET_RMM_FP/SET_DYN_FP — nearest-even, toward-zero, downward, upward, nearest-max-magnitude, or the dynamic mode driven by whatever the engine itself decides); bits 4-0 are the sticky IEEE exception flags (invalid, divide-by-zero, overflow, underflow, inexact) from the last operation that actually ran, refreshed by LOAD_FROM_FPU or LOAD_CCFPU.
Saving and restoring the stack
The full 16-level stack can be pushed to or pulled from memory one byte at a time — PULL_FP steers each incoming byte to the right slice of REG_FP(0) under a one-hot byte selector (REG_FP_BYTE), and PULL_DUP_FP additionally pushes a fresh level before landing the first byte, so a full 8-byte transfer naturally grows the stack by one level per value restored. A separate control, REG_FP_LIMIT, lets this transfer stop after fewer than sixteen levels — the mechanism behind the partial-stack save/restore instructions, which only move as many levels as a function actually uses instead of always paying for the full sixteen.