DMC 63F41
Friday 23 August 2024, by // Peripherals
The 63F09 L1 data cache (DMC 63F41)
Each 63F09 core embeds its own private L1 data cache. It is a direct-mapped, write-through cache: every store is propagated to external memory immediately, and no dirty write-back ever happens.
A write-back policy would require flushing on every context switch and would delay writes aimed at memory-mapped peripherals — both unacceptable for this design. Write-through sidesteps both problems at the cost of writing external memory on every store, hit or miss.
Address layout and storage
A physical address is 36 bits, split as:
- TAG — 20 bits, upper address bits
- BLOCK — 16 bits, lower address bits, used directly as the row index into the cache
Because BLOCK covers the entire low 16 bits of the address with no further line/offset structure, the cache holds one row per byte over a 64 KB window — there is no cache-line granularity beyond a single byte. Two block RAMs are addressed in parallel by the same BLOCK index:
VALID_64k(componentRAM_61F64x1) — one bit of VALID, the 20-bit TAG, and one DIRTY bit per other core in the system, for every block indexCACHE_64K(componentRAM_61F512) — the cached data byte itself, for every block index
Both are true block RAMs, clocked by the same signal, so a tag lookup and a data lookup for the same block always complete together.
CURRENT_TAG/CURRENT_BLOCK are pure combinational functions of the CPU’s own address bus CPU_A — they track the CPU’s requested address instantly, with no latency of their own.
State machine and the fast-path budget
The controller is a state machine with a strict latency budget for the common case: a cache hit costs exactly one CPU cycle, no more (IDLE or IDLE_2 → COMPARE_TAG → READ_CACHE →READ_CACHE_2), four ticks, atching one full rotation of the four phase-shifted clocks that also drive the CPU’s own E clock. IDLE/IDLE_2 are two functionally identical states that simply alternate to balance clock-tree fan-out while the CPU address is unchanged or CPU_VMA is deasserted; as soon as a new address is presented, the machine moves to COMPARE_TAG.
Read hit
In COMPARE_TAG, for a CPU read, the controller compares the requested tag against the stored one.
DIRTY_FLAG folds in the cross-core coherency state: even a matching, valid tag is treated as a miss if another core has more recently written that same block. On a genuine hit, BUS_REQUEST is deasserted, the external address bus is disabled, and the machine proceeds straight to READ_CACHE/READ_CACHE_2, which simply present CACHE_D_OUT (the data BRAM’s output) to the CPU as CPU_D_IN and resynchronize the machine on the CPU’s own E clock before returning to IDLE. A cache hit therefore costs exactly one CPU cycle, matching what a direct, uncached read would already cost — the cache adds no overhead of its own on the hit path.
Read miss
If the tag does not match, is not valid, or is marked dirty by another core, the controller requests the external bus and, once granted, moves to ALLOCATE. ALLOCATE drives the real address out and waits for the external memory’s own ready signal (MRDY, self-looping on ALLOCATE while it is low); once ready, ALLOCATE_2 both hands the fetched data to the CPU and writes it into the cache in the same cycle before returning to IDLE. A block that is not cacheable (peripheral space, for instance) is simply never marked valid, so it is always re-fetched. Populating a new cache entry this way costs 1.5 CPU cycles in the fast case — bus already free, external memory ready without extra wait — on top of the read itself; a busy bus or a slower external memory extends this by however long BUS_GRANTED/MRDY actually take to arrive.
Write (always write-through)
CPU writes never distinguish hit from miss: a store always goes both to the cache row and out to external memory, in the same transaction. From COMPARE_TAG, the controller requests the bus and, once granted, enters WRITE_THROUGH, which drives the address and data out and updates the cache, in the very same cycle it exits toward WRITE_THROUGH_2. WRITE_THROUGH_2 (and, on multi-core configurations, an extra WRITE_THROUGH_3 resync tick — see below) waits for the external write to actually complete before returning to IDLE.
Because the write is unconditional (never gated on a hit), external memory is guaranteed to always hold the up-to-date value for everyaddress ever written — the precondition that lets the cache skip any write-back logic entirely.
Cross-core invalidation
On a multi-core build, each core’s cache carries one DIRTY bit per other core for every block, plus a matching Cross Cache Signaling (CCS) bus that broadcasts every write to all other cores:
CCS_WRITE_PULSE— a single-tick pulse, raised exactly on the
WRITE_THROUGH→WRITE_THROUGH_2transition (never earlier), because that is the one point where the bus is already known to be granted and the write is certain to actually happen — broadcasting any earlier would risk announcing a write that never occurs.CCS_ADDRESS_OUT/CCS_ADDRESS_RW_n_OUT— the written address and direction, broadcast alongside the pulse, received by every other core asCCS_ADDRESS_IN(i)/CCS_ADDRESS_RW_n_IN(i).
Each core runs a separate EXTERNAL_DIRTY_SUPERVISION process per remote core, watching this broadcast. It reads its own tag/valid BRAM
through a second port at the broadcast block index; if the tag matches and the entry is valid, it sets its own DIRTY_OUT(i) bit for that block. Back in COMPARE_TAG, any core’s own read hit test folds in every other core’s dirty bit for the block — a block flagged dirty by a peer is treated as a miss and re-fetched from external memory, guaranteeing the freshest value is picked up rather than a stale local copy.
This snoop path has its own read-latency compensation (COMPARE_TAG_2, an explicit extra settling tick before trusting TAG_DIRTY_OUT/CCS_VALID), deliberately added because this particular path carries no “one CPU cycle” budget constraint — unlike the CPU-facing hit/miss path described above, which cannot afford it at all.
LOCAL_CACHE_MRDY further gates the whole state-commit process on multi-core builds: a core only advances its own state machine while every peer reports CCS_MRDY_IN(i)='1', so no core commits mid-way through a cross-core invalidation window. On single-core builds (NCPUS=1), all of this collapses away — CACHE_MRDY/LOCAL_CACHE_MRDY are simply tied to '1', and DIRTY_FLAG is always false.