Repository navigation
CP regfile Fmax, early-Z fix + default on, asic_gate 4x4 top/gfx, CI shards + ccache - #425
Merged
Merged
Conversation
read_reg() and is_decoded() tested every register with its own early return, which synthesizes as a priority chain of ~30 address compares. On ASIC (Yosys + ABC, ASAP7) that chain was the CP's critical path, about 80 logic levels into rresp, and it held the whole `top` DUT to 341 MHz against its 400 MHz target. Both windows decode every word (the 16-byte block padding fills the gaps), so is_decoded() reduces to word alignment plus window membership, and read_reg() becomes one case per window. Yosys proves the new module sequentially equivalent to the old one (equiv_induct, 2445/2445 $equiv cells; the same check flags a one-bit change to the DECERR sentinel). VX_cp_axil_regfile alone 362 -> 1286 MHz, cells 5732 -> 3779 VX_cp_core alone 363 -> 1169 MHz top (8x8, asic_gate) 341 -> 620 MHz Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The block evaluator fed its output FIFO zplane_r from the stage before VX_raster_qe, while the stamps come from VX_raster_qe's registered output. Each wave therefore left with the depth plane of the block behind it, and whenever consecutive blocks belonged to different primitives, early-Z evaluated a quad against the wrong triangle's plane. With the wrong candidate depth it culled a fragment that late-Z keeps, so the image differed: gfx_draw3d-6 (evilskull 128) failed with early-Z on rtlsim, 2 pixels at 2 slices and 1 pixel at 1 slice, while SimX passed. The plane now goes through a register with the evaluator's enable, so each wave carries its own primitive's plane. Early-Z on, rtlsim: gfx_draw3d-6 passes at 1 and 2 slices, matching SimX. The early-Z cases in graphics.yaml never caught this: they passed -DVX_CFG_RASTER_EARLYZ, which is not a config knob, so they built the knob-off core. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
VX_CFG_RASTER_EARLYZ_ENABLE now defaults to (EXT_RASTER && EXT_OM) instead of false. Early-Z reads committed depth through the OM's ocache, so it is legal only with both units (VX_define.vh errors otherwise); the expression keeps every other build unchanged, and -DVX_CFG_RASTER_EARLYZ_DISABLE opts a build out. The early-Z cases in graphics.yaml passed -DVX_CFG_RASTER_EARLYZ, which resolves to nothing, so CI never ran the early-Z path. Every RASTER+OM case now covers it; those five cases become gfx_no_earlyz-* and pass -DVX_CFG_RASTER_EARLYZ_DISABLE to keep the knob-off build covered. Validated with early-Z on (after the be.sv plane fix): graphics simx 35/35, graphics_parity simx 6/6, vulkan simx 13/13, graphics rtlsim 43/43, graphics_parity rtlsim 6/6, vulkan xrtsim 1/1. The generated config header changed, so every perf baseline's config_hash is re-recorded. Cycles are identical outside graphics; draw3d moves by -1.61% (nt16, +1041 instrs), +0.03% (nt4) and +0.02% (mc). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
top, gfx, tensor and vm were added to asic_gate in 4966e76 at NT=NW=8 without baselines, so the nightly reported them STALE -- and lost top and gfx outright: at 8x8 their Yosys peaks (18.6 and 19.5 GB) exceed a hosted runner's 16 GB, and run 37263315588 killed both mid-synthesis (exit 143). top and gfx are gated at NT=NW=4, where they fit; tensor and vm stay at 8x8. Recorded with the CI toolchain (Yosys 0.69, OpenSTA 3.1.0, sv2v 0.0.13-13), 400 MHz target, on this branch's RTL (top includes the CP regfile fix, gfx the early-Z plane fix): top 4x4 620.9 MHz 1.83M cells 11.2 GB peak RSS gfx 4x4 850.0 MHz 2.31M cells 14.6 GB peak RSS tensor 8x8 976.9 MHz 1.01M cells 8.0 GB peak RSS vm 8x8 1028.5 MHz 1.37M cells 5.3 GB peak RSS Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Timeout and sharding. The tests job times out at 180 minutes (was 300). A cell whose recorded runtime exceeds SHARD_BUDGET_S (150 min) is planned as several shards, split longest-first onto the least-loaded shard by the per-case wall times in ci/runtimes.json; `pytest --shard=K/N` derives the same split from the same ids. ci/record_runtimes.py refreshes the table from a CI run's junit artifacts (gh api). Recorded from run 37241939097: model_parity rtlsim 32 (180 min), perf_gate rtlsim 32 (199), tensor_wg rtlsim 32 (185) and 64 (195) each split into two 90-100 min shards. Host cells. The planner now owns each cell's -m expression (`testcase.py expr`). A host cell selected `<category> and not <checks>` with no driver term, so the regression host cell re-ran all 19 of the category's simx/rtlsim/xrtsim cases that their driver cells already run (~100 min per xlen). It now holds only driverless cases. Build reuse. Cases are ordered so those with one sim build signature run back to back: a blackbox case rebuilds the sim whenever its flags differ from the previous case's, so A, B, A compiled A twice (perf_gate rtlsim: 49 builds -> 32). ccache is re-enabled: it was disabled against the fmt::v8 stale-object class, which needs a cache populated against other headers -- impossible on a hosted runner, whose cache starts empty each job. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Yosys gate now gates only cell area against the baseline; Fmax is held to the build's clock target by the existing BELOW TARGET check. ABC maps against that target (-D) and spends the slack beyond it recovering area, so above the target Fmax moves with whatever else the design contains. After #424 the RTU's worst path is an FMA stage (tri_pe fsub_r, align -> add -> LZC) in VX_fma_unit_rtl, which #424 did not touch, yet it moved 957 -> 903 MHz against a 400 MHz target and failed the baseline check. The fpga_gate is unchanged: Vivado reports what the routed design achieves. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#424 grew the RTU by 8.0% cell area (91600 -> 98965 um^2) and SRAM area by 10%: the per-context state its Vulkan ray-tracing features add (candidate records, object-ray staging, the any-hit/intersection resume state) and the watertight triangle PE. Measured by the nightly on 3c3a84b (run 37263315588, Yosys 0.69 / OpenSTA 3.1.0, config_hash unchanged) and recorded from that run's report; the RTU RTL is the same on this branch. Fmax 903 MHz against the 400 MHz target. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both DUTs contain RTL this branch changes: top the CP AXI-Lite regfile (flat decode), gfx the raster block evaluator (early-Z plane register). Recorded with Vivado 2024.2, config_hash unchanged; the previous baselines also predate Vulkan rt (#424), which reshaped the RTU's memories. top 250 MHz target: 252.6 -> 250.9 MHz, LUT 136066 -> 136762 (+0.5%) gfx 250 MHz target: 250.1 -> 250.1 MHz, LUT 202801 -> 204871 (+1.0%), FF +3%, LUTRAM 10287 -> 16721, BRAM 312 -> 165 (RTU memories now map to LUTRAM) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
equiv_inductproof shows it is equivalent to the old RTL. The proof was checked by confirming it catches a deliberately planted bug.-DVX_CFG_RASTER_EARLYZ_DISABLEturns it off. Thegfx_earlyz-*cells becomegfx_no_earlyz-*.ci/runtimes.json,ci/record_runtimes.py,pytest --shard K/N).Test plan
--collect-only. This PR's CI run is the first full validation.🤖 Generated with Claude Code