Conversation
The floating >=2.13.6 lower bound only happened to resolve to 3.1.0 because that's PyPI's current latest. pyproject.toml already pins pybind11==3.1.0 and says to keep it in sync with torch-nexus's pin; tighten requirements.txt to match so a future pybind11 release can't silently desync the two repos' ABI-compatible pybind11 versions.
…ources Schema/API changes: - CoreSubsystem.SubUnits (array of objects with a Type field) becomes UnitTypes, an object keyed by type name. - Per-entry SubunitType (single string) becomes Subunits, an array of contained unit-type names, so a unit can reference several children. - Add CoreSubsystem.ChipType and CoreType, naming the UnitTypes entries for the whole chip and for a single core. - UnitTypes.Size is now optional with a documented default of 1, applied at read time in InfoImpl::getProp (scoped to UnitTypes entries so MemorySubsystem sizes are unaffected). - Property enums and docs/JSON_API.md updated to match. Device data: Every device file was fact-checked against vendor primary sources and rewritten where wrong. Most files contained fabricated rather than merely stale values -- amd-gpu-gfx950 was entirely placeholder text, gfx942 was labelled CDNA2 for a CDNA3 part and carried LLVM features that do not exist on the target, apple-gpu-m4 claimed 16 TB of unified memory, and the Blackhole INT8 rate was actually a BLOCKFP8 figure. Values that cannot be sourced are now omitted or carry their provenance inline. New: aws-npu-trn1, aws-npu-trn2, nvidia-gpu-sm_100, nvidia-gpu-sm_120. Removed: apple-gpu-m4p (duplicate of apple-gpu-applegpu_g16s, which is the same M4 Pro silicon). All 15 files / 20 device entries validate against the schema with no dangling unit references. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The device library knew its speeds and feeds only as English inside Description
strings. This gives the schema a machine-readable performance vocabulary, fills
it in across the corpus, and reshapes the memory subsystem to match the compute
hierarchy it sits beside.
- schema: a shared Performance definition reachable from every
CoreSubsystem.UnitTypes entry and every MemorySubsystem.MemoryTypes entry --
ClockRate, LaneWidth, TransferRate, Latency, a Source provenance enum, and a
Throughput array of {Rate, Unit, Source, Precision?, Mode?}. Unit and
Precision are closed enums; Mode is free text, because fidelity phases and
Dense/Sparse are vendor-specific and an enum would need editing per vendor.
- device_lib: 82 Performance blocks across 20 files. Every figure is either
stated in that same file's prose or derived from that same file's fields and
marked so; 145 are Published, 38 Derived, 8 Measured. Figures per one unit,
never a chip total multiplied up.
- MemorySubsystem.MemoryTypes is now a name-keyed map like
CoreSubsystem.UnitTypes, with the redundant Type field removed and keys
shortened from captions to identifiers -- "SRAM (L1 per Tensix Core)" is now
"Tensix L1". 106 memories, 38 distinct names, so the same concept carries the
same name across devices.
- Repaired 22 UnitTypes.*.Memory references in the Trainium files that resolved
to nothing: they named "SBUF" against a type declared "SBUF (State Buffer,
per NeuronCore)". Nothing had been checking them.
- KernelModel.SubUnits is now KernelModel.Toolchains and sits inside properties.
It had been a sibling of properties, so JSON Schema ignored it entirely, and
it holds toolchain versions (CUDA, HIP, Metal, TT-Metalium), never subunits.
Its property row also declared _prop_str_vec for an array of objects.
- apple-gpu-.json held six device objects in one array, and device_db.cpp keys
a device by filename stem -- so all six registered under "apple-gpu-" and none
was reachable by architecture. Split into six files.
- New tests: device_lib validation against the schema, which nothing did before;
a provenance audit that reproduces every Derived figure from same-file fields
or traces it to a quoted sentence; and property-type coverage proving a
whole-number literal in a float-typed field still reads back as a double.
Adds jsonschema to requirements.txt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Schema/API changes:
- CoreSubsystem.SubUnits (array of objects with a Type field) becomes UnitTypes, an object keyed by type name.
- Per-entry SubunitType (single string) becomes Subunits, an array of contained unit-type names, so a unit can reference several children.
- Add CoreSubsystem.ChipType and CoreType, naming the UnitTypes entries for the whole chip and for a single core.
- UnitTypes.Size is now optional with a documented default of 1, applied at read time in InfoImpl::getProp (scoped to UnitTypes entries so MemorySubsystem sizes are unaffected).
- Property enums and docs/JSON_API.md updated to match.
Also fixed pybind11 pin.