Repository navigation
MX changes the on-chip MX format and layout; MX has no production users yet, so they stay on 0.7.x.
MX software written for 0.7.3 needs updating (see the MX entries below).
Fixed
- inst64 writes with a non-zero user (multicast, reduction) always go out on AXI, so they reach
the cluster router instead of the local TCDM over OBI
#264. - Remove the structural combinational loops between the backend and an
axi_to_obislave
(w_dp_busyno longer depends on the write port ready)
#265, and between the read handshake and the
compute busy flag; declare the compute helper types at module scope for VCS
#271. - Transpose and MX compute over OBI: transpose edge strobes reach an OBI destination
#235. - Hold the ND midend's zero-transfer rejection until it is accepted. It was lost when the response
side was not ready, and a rejection taken while the backend stalled left the next transfer
starting from the rejected transfer's address
#267.
Changed
- Public CI takes Verilator 5.052 from its official Docker image #263.
- MX format: the MX block scale byte is now OCP E8M0 (unsigned, bias 127,
0xFF= NaN)
instead of two's complement; the shared exponent clamps to [-127, 127] and a0xFFscale
dequantizes to NaN. MX blocks written by earlier versions decode differently
#259. - MX format: MX quant poisons a block holding an Inf or NaN (scale
0xFF, every element
0x7D); the per-transfermx_options_t.poison_disbit (mx_cfgregister, DMOPCrs1[18])
restores the finite-lane behaviour. - MX dequant to FP32 is exact: FP32 subnormal results instead of a flush to zero, Inf instead of
saturation. - The MX quantizer is a four-stage pipeline on whole beats with an output queue of destination
beats. Its reg2reg paths close at 0.45 ns in GF12 (1.17 ns before). - MX dequant is a beat-granular pipeline (input buffer, extract, expand lanes, output queue)
that retires one output beat per cycle. - MX ops left the compute config interlock: every MX beat carries its config in a beat tag
(r_dp_req_t.mx), so transfers with different MX ops and copies stream back to back without
draining. MX beats go from the read side to their engine past the dataflow element: the quant
engine has an input queue ofBufferDepthbeats (one more withdfe_ready_cut), the dequant
input buffer holds three more, and the dataflow element carries copy and transpose beats only,
so a copy or a waiting dequant input does not block the next transfer's input. Both engines take
whole beats on a registered credit and have no transfer-boundary clear; the write side
selects the burst's output queue after the write shifter and pops it on the W handshake. The
interlock remains for transpose.tb_idma_mxclearchecks the empty-queue pop guard of both
engines. - MX format: MX transfers use a data plane (32 B of elements per block) and a separate
scale plane (one E8M0 byte per block) in one pass; the inline 33 B block layout is removed.
mx_options_t(compute_params_t.mx, registermx_cfg, DMOPCrs1fields generated
from the same database) carriespoison_dis,rceil,elem_fmtandgroup(G = 64 or 32
blocks per scale chunk); the scale plane address isidma_req_t.scale_addr. The legalizer
interleaves data bursts and per-group scale bursts on the write side (quant) and scale-first
reads on the read side (dequant). The dequantlengthis the data-plane length.
idma_pkg::MxBlockBytesis gone, and theidma_mx_golden.hquant/dequant helpers take a data
and a scale buffer. - MX format: MX compute needs
StrbWidth <= 64; wider MX builds stop at elaboration. - Interface:
IDMA_TYPEDEF_REQ_Taddsscale_addr, the MX scale plane address, 64 B
aligned (ComputeMxScaleAligned), andIDMA_TYPEDEF_D_REQ_Taddsscale_strides; code that
builds these structs field by field must drive them (zero outside MX). The ND midend steps
scale_addrper row withscale_strides. The register frontend hasscale_addrand
per-dimensionmx_dim[d].scale_strideregisters aftermx_cfg, inst64 sets them with the
DMOPC setters0x28and0x29.
Added
- The register frontend refuses an MX launch with a reserved element format, a scale plane that
is not 64 B aligned, a used dimension's scale stride that is not a multiple of 64, a data plane
or a used dimension's data stride off a beat (new parameterDataWidth, default 512) or a
length that is not a whole number of blocks:next_idreads 0 and no transfer starts. inst64
refuses such aDMCPYwith id 0 and the error bit. - Both frontends also refuse a compute op the backend does not elaborate (register frontend:
new parameterComputeOps) or a reserved op code, and an MX dequant whose written length does
not fit the length field; inst64 refuses an MXDMCPYwith a plane in the TCDM window. The
backend runs an op that is not elaborated as a plain copy instead of hanging. - Backend parameter
TimingCuts(idma_pkg::timing_cuts_t, default'0= unchanged):
dfe_ready_cut(dataflow element without same-cycle refill on write-side pops, one more entry),
dfe_reg_flags(registered dataflow flags, pointers without load enables),wdp_head_spill
(spill register on the write datapath request head) andoutst_cnt_reg(outstanding-transfer
counter offreq_ready_o).idma_inst64_toppasses it to its backends.mx_in_reg(MX beats
enter their engine through a one-beat register stage; one cycle more MX latency) and
dfe_ready_ahead(dataflow element ready from flops; one more entry per lane, taken only when
the next write beat pops the lane; new write port outputbuffer_out_offer_o) and
mx_beat_push(MX engine push on the AXI read beat without the byte-lane masks;idma_axi_read
outputbuffer_in_beat_o) are added as leading fields, so the values of the earlier fields are
unchanged. - The register-map tables of the register frontend page and the DMOPC opcode table of the
Snitch page are generated from the RDL andidma_dmopc.yml(make idma_doc_tables);
make idma_verify_codegenfails when a page differs. - MX quant and dequant support E4M3 elements next to E5M2, selected per transfer by
mx_options_t.elem_fmt(E4M3: max normal 448 =0x7E, NaN0x7F, no Inf); a poisoned E4M3
block has0x7Felements. - Per-transfer RCEIL scale rounding (
mx_options_t.rceil): the shared exponent is
ceil(log2(amax / max_normal))instead offloor(log2(amax)) - emax, so the block max never
saturates. FLOOR (OCP MX v1.0) stays the default.
Fixed
- MX quant of FP32 subnormal inputs dropped the leading zeros (implicit 1), e.g. 32 x 2^-127
quantized to 1.5 instead of 1.0. - The legalizer took a request only when both its read and write sides had emitted the previous
one, so a stream of one-block MX quant transfers (two write bursts per read burst) left the
read channel no lead and lost cycles after every read stall. The write side now queues up to
three requests (MxWqDepth) behind an MX transfer while the read side runs ahead; the read
and write sides keep separate MX plane state, so a dequant behind a short quant reads at once. make idma_verify_codegentakes the compute guards from the generated legalizers, and every
one of them has a negative test: mxneg case 11 now runs (ComputeMxdequantLengthFits), and
ComputeDstTilelinkis only generated for variants with a TileLink write port.