CPU
Ziren V2.0 has no CPU chip. The work a CPU table used to do (fetch the instruction, read and write registers, advance the clock and the program counter) is done by an instruction frame: a fixed group of columns and constraints that every instruction chip embeds in its rows. An executed instruction therefore costs one row in the chip of its opcode family and nothing else. The frame is defined in crates/core/machine/src/frame/mod.rs.
Frame columns
InstructionFrameCols has:
shard: the shard number, range-checked to 16 bits.clk_16bit_limb,clk_high_limb: the per-shard clock \( clk = clk_{16} + 2^{16} \cdot clk_{high} \). The low limb is range-checked to 16 bits and the high limb to 10 bits, so timestamps are 26-bit values.instruction: the decoded instruction (opcode,op_a,op_b,op_c, the immediate flagsimm_b,imm_c, andop_a_0, which marks a write to$zero).op_a_access,op_b_access,op_c_access: the three register accesses.op_ais a read-write access (previous and new value),op_bandop_care reads.
Jump, MovCond and MiscInstrs use this full frame. Narrower variants save columns where the instruction format allows it:
ITypeFrameCols:op_aandop_bregisters and an immediateop_cword. Used byAddSubImm,BitwiseImm,LtImm,CloClz,Branchand the memory-instruction chips.RTypeFrameCols: three registers, with the opcode and operand indices stored as single columns. Used byAddSub,Bitwise,Lt,Mul,DivRem,ShiftLeft,ShiftRightandSyscallInstrs.ShamtFrameCols:op_aandop_bregisters and a 5-bit shift amount held in one column. Used byShiftLeftImmandShiftRightImm.
Each register access holds the value, the clock of the previous access to that register, and one 16-bit limb of the clock difference: six columns. The previous access is always in the same shard because the MemoryBump chip inserts a read of every touched register at (shard, 0) (see Memory).
Frame constraints
eval_instruction_frame constrains, for a real row:
- Program fetch. The row sends
(pc, instruction)on theProgrambus, so the instruction must be the one the preprocessed program table stores atpc. The instruction's opcode must equal the chip's opcode, which the chip computes from its own selector columns. - Clock. The shard and the two clock limbs are range-checked.
- Immediates. If
imm_b(orimm_c) is set, the operand value equals the immediate in the instruction and no register is read. - Registers. The accesses happen at fixed sub-cycle offsets:
op_cat \( clk + 1 \),op_bat \( clk + 2 \),op_aat \( clk + 3 \). A memory access, where there is one, uses \( clk + 0 \), and theHIregister written byMul,DivRemand the multiply-accumulate instructions ofMiscInstrsuses \( clk + 4 \). Ifop_a_0is set, the written value is zero. The written value's four bytes are range-checked. - State hand-off. The row receives
(shard, clk, pc, next_pc)on theStatebus and sends(shard, clk + 5 + extra, next_pc, next_next_pc), whereextrais the number of extra cycles a syscall takes (zero for every other chip).
The chip supplies pc, next_pc and next_next_pc as expressions. pc and next_pc are whatever the row receives from its predecessor; in particular a delay-slot instruction receives a branch target as its next_pc. A sequential chip only fixes \( next\_next\_pc = next\_pc + 4 \). The branch and jump chips carry next_pc and next_next_pc as range-checked words and compute the target. The halt row of the syscall chip receives its predecessor's pc + 4 as next_pc and sends next_pc = 0, the exit marker.
Because every instruction advances clk by at least 5 and consumes one State tuple while producing one, balancing the State bus forces the rows of all instruction chips to form one chain from the shard's initial state to its final state (see State Machine).
Program counters and the delay slot
MIPS executes the instruction after a branch or jump (the delay slot) before the target. The frame carries two program counters:
next_pcis the address of the next instruction to execute; for a branch or jump this is the delay slot.next_next_pcis the address after that. For sequential instructions it isnext_pc + 4; a branch or jump sets it to the taken target or the fall-through address.
The delay-slot instruction receives (next_pc, next_next_pc) and passes the target on as its own next_pc.
The shard's public values carry both pairs, (start_pc, start_next_pc) and (next_pc, next_next_pc), but a pending branch target never crosses a shard boundary. The executor never closes a shard in front of a delay slot, and the recursion leaf enforces it: for a shard that executes instructions it asserts start_next_pc = start_pc + 4 and next_next_pc = next_pc + 4, so both boundaries are sequential. The recursion then chains only pc: each shard's start_pc must equal the previous shard's next_pc.