FPU

Tuesday 28 May 2024, by 63F09 // CPU 63F09

The floating-point unit 63F39

The 63F09’s FPU is split cleanly in two: a 16-level register stack hosted in the CPU core itself, and a separate, stateless execution engine (referred to as 63F39) that the core drives through a small request/response interface. The core owns all FPU state; the engine only computes.

A 16-level, RPN-style operand stack, not a flat register file

REG_FP is sixteen 64-bit registers, REG_FP(0) to REG_FP(15) — but software addresses them as stack levels, not fixed register numbers: REG_FP(0) is always the top of the stack. Every FPU instruction ultimately drives one value into FP_STACK_CTRL (type FP_STACK_TYPE), decoded in a single process on the falling edge of the internal clock, and the whole array shifts as one unit:

  • Push — every register moves up one slot (REG_FP(i+1)<=REG_FP(i) for i from 14 downto 0) and the new value lands in REG_FP(0). Used to load an operand from memory (LOAD_FROM_MEM_DATA) or from the integer side of the core (LOAD_FROM_O, moving the 64-bit pseudo register O = Q:V across).
  • Pop — every register moves down one slot, the vacated top entries are cleared to zero (DROP_FP, and its N-level generalization DROPN_FP).
  • Stack juggling — DUP_FP/DUP2_FP/SWAP_FP/OVER_FP/ ROT_FP cover the fixed, one- or two-level Forth-style primitives; DUPN_FP/DROPN_FP/ROLL_FP/ROLLD_FP/PICK_FP/EXG_FP generalize duplication, dropping, rotation and exchange to an arbitrary level, and CLEAR_FP empties the whole stack.

Feeding the execution engine

An operation reads its operands directly off the top of this stack. LOAD_FPU_ARG1/LOAD_FPU_ARG2/LOAD_FPU_ARG3 copy the top one, two or three levels into the engine’s own input latches, and START_FPU raises FPU_IN.FP_EXE_I.ENABLE to actually launch the computation. Once the engine reports back, LOAD_FROM_FPU takes the 64-bit RESULT and the 5-bit IEEE exception FLAGS, and LOAD_CCFPU updates just the flags without touching the stack at all — the path a comparison uses (see below).

Whether consuming a value counts as a real pop is a single bit, FP_STACK_CONF (FP_STACK_CONF_KEEP_ARGS vs FP_STACK_CONF_DISCARD_ARGS, itself selected by one bit of the opcode): with discard, LOAD_FPU_ARGn shifts the consumed levels away and LOAD_FROM_FPU pushes the result on top — a normal, stack-popping operation, net stack effect one level shallower per result produced. With keep, neither shift happens at all: LOAD_FPU_ARGn reads the operands in place and LOAD_FROM_FPU simply overwrites REG_FP(0) with the result — the stack depth is unchanged, and any operand below the top (REG_FP(1) and beyond) survives untouched for a following operation to reuse. This is what lets a value outlive the operation that reads it, without a manualDUP beforehand.

The execution engine itself

fp_exe_in_type/fp_exe_out_type (fp_wire.vhd) define the engine’s own interface: three 64-bit operands, a one-hot fp_operation_type selector (fmadd/fmsub/fnmadd/fnmsub, fadd/fsub/fmul/fdiv/fsqrt, fsgnj, fcmp, fmax, fclass, fmv_i2f/fmv_f2i, fcvt_f2f/fcvt_i2f/fcvt_f2i plus a 2-bit fcvt_op sub-selector), a 2-bit format field and a 3-bit rounding-mode field — the same shape as a RISC-V “F/D” extension unit. The engine itself carries no stack, no state beyond one computation: the CPU-side stack machine described above is entirely what turns this generic, three-operand functional unit into the 63F09’s own Forth-style FPU instruction set.

Private or shared: the FPU_BY_CORE choice

Everything above describes the interface between a core and an FPU engine — but on a multi-core build, that engine can be instantiated two different ways, selected by a single generic, FPU_BY_CORE, threaded down from the SoC top level through 63F81.vhd/63F29.vhd/ 63F09.vhd:

  • Per-core (FPU_BY_CORE=true) — each core gets its own private FP_UNIT instance, declared right inside the CPU core’s own architecture (63F09.vhd, FPU_INTERFACE : if FPU_BY_CORE generate). No arbitration is needed at all: a core’s REG_FP stack talks directly to an engine that belongs to it alone.
  • Shared (FPU_BY_CORE=false) — a single FP_UNIT instance lives at the top level instead (63F81.vhd, generate block SHARED_FPU), clocked by the chip-level E — the current bus master’s own E, from the same arbiter described in the companion document. Every core still drives its own FPU_IN(i) entry (a per-core vector, fp_unit_in_vector_type) exactly as if it owned the engine outright, but only one entry at a time is actually multiplexed into the real unit’s input (FPU_MUX_IN <= FPU_IN (CURRENT_CPU)), and the engine’s single output, FPU_OUT, is wired back identically to every core’s FPU_63F39_OUT port — the same signal, broadcast. A core only sees a meaningful result once the arbiter has actually picked its own request.

The arbiter that picks CURRENT_CPU mirrors the bus arbiter almost line for line: on each falling edge of E, a round robin starting from the currently-selected core (FPU_ACTIVE_CPU) looks for the first core whose FP_EXE_I.ENABLE is asserted, and sticks with whichever core it finds for as long as that core keeps its own ENABLE raised — the same “sticky, self-first” pattern as SIGNAL_LAST_ACTIVE_CPU for the bus, just applied to the FPU request instead of the bus request. A core that loses the race simply waits, its own ENABLE still high, until the current holder releases the engine and the round robin reaches it.

Comparisons: result goes to the CPU’s own flags, not the stack

fcmp is the one operation that does not produce a floating-point value: it feeds LOAD_CCFPU instead of LOAD_FROM_FPU, so the comparison’s outcome only ever updates REG_CCFPU’s exception flags — the actual true/false predicate is delivered to the CPU’s own condition codes (Z), the same register a plain integer CMPx would set. A comparing instruction therefore never touches the operand stack itself; a non-popping variant leaves both compared values exactly where they were.

Status and rounding mode: REG_CCFPU

REG_CCFPU is a single 8-bit register, split in two: bits 7-5 hold the current IEEE-754 rounding mode, set by six dedicated stack-control values (SET_RNE_FP/SET_RTZ_FP/SET_RDN_FP/SET_RUP_FP/ SET_RMM_FP/SET_DYN_FP — nearest-even, toward-zero, downward, upward, nearest-max-magnitude, or the dynamic mode driven by whatever the engine itself decides); bits 4-0 are the sticky IEEE exception flags (invalid, divide-by-zero, overflow, underflow, inexact) from the last operation that actually ran, refreshed by LOAD_FROM_FPU or LOAD_CCFPU.

Saving and restoring the stack

The full 16-level stack can be pushed to or pulled from memory one byte at a time — PULL_FP steers each incoming byte to the right slice of REG_FP(0) under a one-hot byte selector (REG_FP_BYTE), and PULL_DUP_FP additionally pushes a fresh level before landing the first byte, so a full 8-byte transfer naturally grows the stack by one level per value restored. A separate control, REG_FP_LIMIT, lets this transfer stop after fewer than sixteen levels — the mechanism behind the partial-stack save/restore instructions, which only move as many levels as a function actually uses instead of always paying for the full sixteen.