Skip to content

CP regfile Fmax, early-Z fix + default on, asic_gate 4x4 top/gfx, CI shards + ccache - #425

Merged
tinebp merged 8 commits into
masterfrom
ci-earlyz-asic-gate
Oct 6, 2026
Merged

tinebp merged 8 commits into
masterfrom
ci-earlyz-asic-gate

Conversation

@tinebp

@tinebp tinebp commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • CP register file:
    • The AXI-Lite register file now decodes addresses flat instead of through a priority chain.
    • A Yosys equiv_induct proof shows it is equivalent to the old RTL. The proof was checked by confirming it catches a deliberately planted bug.
    • ASIC top Fmax rises from 341 to 620 MHz, so it now meets the 400 MHz target.
  • Early-Z:
    • Fixes a raster bug: the depth plane was not delayed through the quad-evaluator stage, so early-Z tested against the wrong primitive's plane.
    • Early-Z is now on by default wherever RASTER and OM are both enabled.
    • -DVX_CFG_RASTER_EARLYZ_DISABLE turns it off. The gfx_earlyz-* cells become gfx_no_earlyz-*.
    • Perf baselines are re-recorded. draw3d cycles change by −1.61% / +0.03% / +0.02%.
  • ASIC gate:
    • top and gfx now synthesize at NT=NW=4. At 8×8 they peak at 18.6 and 19.5 GB of Yosys memory, which exceeds the 16 GB hosted runners. At 4×4 they peak at 11.2 and 14.6 GB.
    • tensor and vm stay at 8×8.
    • Fmax is now checked against the target clock only. ABC maps to the target period and spends any extra slack on area, so Fmax above the target is noise. Cell area is still gated against the baseline.
    • The rtu baseline is re-recorded from CI run 37263315588. Its area growth is the cost of the Vulkan ray-tracing feature (Vulkan rt #424).
  • FPGA gate: the top and gfx baselines are re-recorded with Vivado 2024.2.
  • CI:
    • Cells that ran longer than 2h30 are split into shards. Tests are assigned to shards by their runtimes recorded in GitHub run 37241939097 (ci/runtimes.json, ci/record_runtimes.py, pytest --shard K/N).
    • Job timeout is 180 min.
    • Host cells now exclude tests that belong to a specific driver.
    • Tests that need the same simulator build run next to each other, so each build happens once per job.
    • ccache is back on, because hosted runners start every job with an empty cache.

Test plan

  • Early-Z:
    • SimX: graphics, graphics_parity and Vulkan pass.
    • rtlsim: graphics 43/43 and graphics_parity 6/6 pass.
    • xrtsim: Vulkan passes.
  • Perf, ASIC (Yosys/ASAP7) and FPGA (Vivado) baselines recorded on the final RTL.
  • The register-file equivalence proof passes.
  • Shards, timeout and ccache: only checked locally with --collect-only. This PR's CI run is the first full validation.
  • ASIC nightly with top and gfx at 4×4: gfx peaks at 14.6 GB on a 16 GB runner, which is tight.

🤖 Generated with Claude Code

tinebp and others added 8 commits October 5, 2026 11:13
read_reg() and is_decoded() tested every register with its own early
return, which synthesizes as a priority chain of ~30 address compares.
On ASIC (Yosys + ABC, ASAP7) that chain was the CP's critical path, about
80 logic levels into rresp, and it held the whole `top` DUT to 341 MHz
against its 400 MHz target.

Both windows decode every word (the 16-byte block padding fills the gaps),
so is_decoded() reduces to word alignment plus window membership, and
read_reg() becomes one case per window. Yosys proves the new module
sequentially equivalent to the old one (equiv_induct, 2445/2445 $equiv
cells; the same check flags a one-bit change to the DECERR sentinel).

  VX_cp_axil_regfile alone   362 -> 1286 MHz, cells 5732 -> 3779
  VX_cp_core alone           363 -> 1169 MHz
  top (8x8, asic_gate)       341 ->  620 MHz

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The block evaluator fed its output FIFO zplane_r from the stage before
VX_raster_qe, while the stamps come from VX_raster_qe's registered output.
Each wave therefore left with the depth plane of the block behind it, and
whenever consecutive blocks belonged to different primitives, early-Z
evaluated a quad against the wrong triangle's plane. With the wrong
candidate depth it culled a fragment that late-Z keeps, so the image
differed: gfx_draw3d-6 (evilskull 128) failed with early-Z on rtlsim, 2
pixels at 2 slices and 1 pixel at 1 slice, while SimX passed.

The plane now goes through a register with the evaluator's enable, so each
wave carries its own primitive's plane. Early-Z on, rtlsim: gfx_draw3d-6
passes at 1 and 2 slices, matching SimX.

The early-Z cases in graphics.yaml never caught this: they passed
-DVX_CFG_RASTER_EARLYZ, which is not a config knob, so they built the
knob-off core.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
VX_CFG_RASTER_EARLYZ_ENABLE now defaults to (EXT_RASTER && EXT_OM) instead
of false. Early-Z reads committed depth through the OM's ocache, so it is
legal only with both units (VX_define.vh errors otherwise); the expression
keeps every other build unchanged, and -DVX_CFG_RASTER_EARLYZ_DISABLE opts a
build out.

The early-Z cases in graphics.yaml passed -DVX_CFG_RASTER_EARLYZ, which
resolves to nothing, so CI never ran the early-Z path. Every RASTER+OM case
now covers it; those five cases become gfx_no_earlyz-* and pass
-DVX_CFG_RASTER_EARLYZ_DISABLE to keep the knob-off build covered.

Validated with early-Z on (after the be.sv plane fix): graphics simx 35/35,
graphics_parity simx 6/6, vulkan simx 13/13, graphics rtlsim 43/43,
graphics_parity rtlsim 6/6, vulkan xrtsim 1/1.

The generated config header changed, so every perf baseline's config_hash
is re-recorded. Cycles are identical outside graphics; draw3d moves by
-1.61% (nt16, +1041 instrs), +0.03% (nt4) and +0.02% (mc).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
top, gfx, tensor and vm were added to asic_gate in 4966e76 at NT=NW=8
without baselines, so the nightly reported them STALE -- and lost top and
gfx outright: at 8x8 their Yosys peaks (18.6 and 19.5 GB) exceed a hosted
runner's 16 GB, and run 37263315588 killed both mid-synthesis (exit 143).
top and gfx are gated at NT=NW=4, where they fit; tensor and vm stay at 8x8.

Recorded with the CI toolchain (Yosys 0.69, OpenSTA 3.1.0, sv2v 0.0.13-13),
400 MHz target, on this branch's RTL (top includes the CP regfile fix, gfx
the early-Z plane fix):

  top     4x4   620.9 MHz  1.83M cells  11.2 GB peak RSS
  gfx     4x4   850.0 MHz  2.31M cells  14.6 GB peak RSS
  tensor  8x8   976.9 MHz  1.01M cells   8.0 GB peak RSS
  vm      8x8  1028.5 MHz  1.37M cells   5.3 GB peak RSS

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Timeout and sharding. The tests job times out at 180 minutes (was 300).
A cell whose recorded runtime exceeds SHARD_BUDGET_S (150 min) is planned
as several shards, split longest-first onto the least-loaded shard by the
per-case wall times in ci/runtimes.json; `pytest --shard=K/N` derives the
same split from the same ids. ci/record_runtimes.py refreshes the table
from a CI run's junit artifacts (gh api). Recorded from run 37241939097:
model_parity rtlsim 32 (180 min), perf_gate rtlsim 32 (199), tensor_wg
rtlsim 32 (185) and 64 (195) each split into two 90-100 min shards.

Host cells. The planner now owns each cell's -m expression
(`testcase.py expr`). A host cell selected `<category> and not <checks>`
with no driver term, so the regression host cell re-ran all 19 of the
category's simx/rtlsim/xrtsim cases that their driver cells already run
(~100 min per xlen). It now holds only driverless cases.

Build reuse. Cases are ordered so those with one sim build signature run
back to back: a blackbox case rebuilds the sim whenever its flags differ
from the previous case's, so A, B, A compiled A twice (perf_gate rtlsim:
49 builds -> 32). ccache is re-enabled: it was disabled against the
fmt::v8 stale-object class, which needs a cache populated against other
headers -- impossible on a hosted runner, whose cache starts empty each job.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Yosys gate now gates only cell area against the baseline; Fmax is held
to the build's clock target by the existing BELOW TARGET check. ABC maps
against that target (-D) and spends the slack beyond it recovering area, so
above the target Fmax moves with whatever else the design contains. After
#424 the RTU's worst path is an FMA stage (tri_pe fsub_r, align -> add ->
LZC) in VX_fma_unit_rtl, which #424 did not touch, yet it moved 957 -> 903
MHz against a 400 MHz target and failed the baseline check. The fpga_gate
is unchanged: Vivado reports what the routed design achieves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#424 grew the RTU by 8.0% cell area (91600 -> 98965 um^2) and SRAM area by
10%: the per-context state its Vulkan ray-tracing features add (candidate
records, object-ray staging, the any-hit/intersection resume state) and the
watertight triangle PE. Measured by the nightly on 3c3a84b (run
37263315588, Yosys 0.69 / OpenSTA 3.1.0, config_hash unchanged) and
recorded from that run's report; the RTU RTL is the same on this branch.
Fmax 903 MHz against the 400 MHz target.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both DUTs contain RTL this branch changes: top the CP AXI-Lite regfile
(flat decode), gfx the raster block evaluator (early-Z plane register).
Recorded with Vivado 2024.2, config_hash unchanged; the previous baselines
also predate Vulkan rt (#424), which reshaped the RTU's memories.

  top  250 MHz target: 252.6 -> 250.9 MHz, LUT 136066 -> 136762 (+0.5%)
  gfx  250 MHz target: 250.1 -> 250.1 MHz, LUT 202801 -> 204871 (+1.0%),
       FF +3%, LUTRAM 10287 -> 16721, BRAM 312 -> 165 (RTU memories now
       map to LUTRAM)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@tinebp
tinebp merged commit 9643f3e into master Oct 6, 2026
100 of 103 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant