BAC 63F81
Friday 23 August 2024, by // Peripherals
The shared bus arbiter 63F81
On a multi-core 63F09 build, all cores share a single external bus interface — one set of address, data, and control pins. 63F81.vhd instantiates one DMC_63F41 (with its own L1 cache, see the companion document) per core, and arbitrates which core’s signals actually drive that shared interface at any given moment.
One master at a time, chosen by a sticky round robin
The arbiter tracks a single integer, SIGNAL_LAST_ACTIVE_CPU — the index of the core currently granted the bus. It is recomputed by a purely combinational process, sensitive only to the per-core BUS_REQUEST/LOCK_REQUEST vectors:
for i in 0 to NCPUS loop
CURRENT_CPU := (SIGNAL_LAST_ACTIVE_CPU + i) mod NCPUS;
if ((BUS_REQUEST(CURRENT_CPU)='1') or (LOCK_REQUEST(CURRENT_CPU)='1'))
and (REG_CONFIGURATION(CURRENT_CPU)='1') then
exit;
end if;
end loop;
SIGNAL_LAST_ACTIVE_CPU <= CURRENT_CPU;The search always starts from the current master (i=0 maps back to SIGNAL_LAST_ACTIVE_CPU itself). As long as that core keeps asserting BUS_REQUEST or LOCK_REQUEST and remains enabled in REG_CONFIGURATION, the loop exits immediately on its very first iteration and the same core keeps the bus — there is no forced, time-sliced rotation while a master keeps asking. Only once the current master’s own request drops (or it is disabled) does the search continue forward, round-robin, to the next requesting, enabled core. If nothing at all is requesting, the loop runs to completion (i=NCPUS wraps back to the current master) and the arbiter simply parks on the same core by default.
Grants are a strict one-hot
BUS_GRANTED, LOCK_GRANTED and REG_CURRENT_CPU are all derived from SIGNAL_LAST_ACTIVE_CPU alone, in a second, equally combinational process: every entry is cleared, then exactly the one at index SIGNAL_LAST_ACTIVE_CPU is set. Exactly one core ever holds the bus (and the lock) at a time; every other core’s BUS_GRANTED(i) reads '0' and its request simply waits for its turn.
Only the master sees the real external signals
The same process default-initializes every core’s CPU_MRDY, CPU_DMA_BREQ_n, CPU_D_IN and CPU_CACHEABLE to an inert value (CPU_MRDY(i)<='1', meaning “never wait”, CPU_D_IN(i)<=all-ones, and so on) for every core, then overwrites only the current master’s own entry with the real, top-level signal — CPU_MRDY (SIGNAL_LAST_ACTIVE_CPU) <= MRDY, and likewise for the other three. A core that does not currently hold the bus is therefore never stalled by a memory transaction it is not even part of; it keeps running against its own private L1 cache exactly as if it were the only core in the system. The chip-level MMU_MRDY output follows the identical idiom, one level up — MMU_MRDY <= CORE_MMU_MRDY(SIGNAL_LAST_ACTIVE_CPU) — deliberately not an OR across all cores, since that would only stay correct for as long as every non-master core’s own CORE_MMU_MRDY output happens to be an unconditional '1' (a property of 63F29.vhd, not guaranteed by the arbiter itself).
The chip-level bus signals proper — E/E_n/Q/RW_n/A/ D_OUT/BA/BS/FIC/VMA/MEMCLK/CLKVECT/BUS_MASTER — are all simply the current master’s own per-core signals, muxed by SIGNAL_LAST_ACTIVE_CPU: E <= CPU_E(SIGNAL_LAST_ACTIVE_CPU), and so on for every one of them.
Protecting E across a bus hand-off
Each core keeps generating its own local E_LOCAL/E_n_LOCAL continuously, bus or no bus — a core not holding the bus still runs against its own cache. But the instant a core is granted the bus, its own local E phase may not line up cleanly with the moment ownership actually changed hands. A small per-core shift register guards against this:
if BUS_GRANTED(i) = '1' then
if REGISTER_BUS_GRANTED(1) = '0' then
CPU_E(i) <= '1'; CPU_E_n(i) <= '0'; -- held, right after the grant
else
CPU_E(i) <= E_LOCAL; CPU_E_n(i) <= E_n_LOCAL; -- settled, pass through
end if;
else
CPU_E(i) <= E_LOCAL; CPU_E_n(i) <= E_n_LOCAL; -- not the master: always local
end if;REGISTER_BUS_GRANTED is a 2-bit shift register, sampling BUS_GRANTED(i) on the falling edge of each of the four CLKVECT phases in turn. Right after a grant, CPU_E(i) is held high for as long as BUS_GRANTED(i) has not yet been seen stably asserted across two such samples; once it has, the core’s real E_LOCAL/E_n_LOCAL is allowed through as the externally-visible CPU_E(i). The register is reset the instant BUS_GRANTED(i) falls back to '0', so the very next hand-off to this core goes through the same settling window again.
The lock mechanism
LOCK_REQUEST/LOCK_GRANTED give a core a way to hold the bus for an indivisible sequence (for instance an atomic read-modify-write, or the dedicated LOCK/UNLOCK instructions). The arbiter treats a lock request exactly like a bus request for the purpose of keeping a core selected — either one is enough to make the round-robin search stop on that core — and grants it in the very same one-hot process as BUS_GRANTED, at the same index. As long as a core keeps LOCK_REQUEST asserted, the sticky round robin above guarantees no other core can become SIGNAL_LAST_ACTIVE_CPU in between.