Clifford t compilers - #5
Merged
Merged
Conversation
Move src/compile_circuit.py to data_processing/bqskit_compile_cliffordt.py, and track data_processing/qiskit_compile_cliffordt.py, which was previously untracked. Both now sit next to each other with matching names. bqskit compiler: - Substitute a gauge-fixed ZXZXZDecomposition. bqskit's version splits a diagonal rotation evenly between the two outer rz gates, leaving both generic, so GridSynthPass is called twice where once would do. That split is a free gauge whenever the middle block is Clifford. Takes hubbard_18 from 220400 to 151890 T at equal accuracy, seed 0, and slightly improves fidelity. - Default --seed to 0. Unseeded, T counts varied ~40% between runs, which made gate counts unreproducible and unsafe to compare. - Report per-gate counts, T count and multi-qubit count, and total elapsed time. - Add --synthesis-epsilon, and --no-rz-gauge-fix to reproduce stock output. qiskit compiler: - Guard both exact-synthesis paths by measuring the word they produce against the target. The Clifford lookup key rounds to 7 decimals, so it matched anything within ~5e-8 of a Clifford and discarded small rotations for free at hundreds of times the requested epsilon. Raises QFT(40) from 192147 to 224547 T and restores 184 wrongly cancelled cx gates. - Default --tol to --epsilon, so an exact rewrite is never looser than the approximate path it replaces. - Accumulate every rewrite's error into a reported bound. - Verify at any qubit count: basis check and error bound always, random statevector fidelity to 24 qubits, dense unitary comparison to 10. - Count control-flow block contents in the reported statistics. Also complete the README project structure, which listed 16 of 17 src/ files and 2 of 14 in data_processing/, and omitted tests/ and data/ entirely. Every tracked file is now named, except the ~190 benchmark circuits under data/all_compiled/ and data/transpiled/, which are summarised by directory. Drop the empty, unused .gitmodules.
The data/ benchmark circuits are LFS-tracked, but postCreateCommand never installed git-lfs, so a devcontainer rebuild silently loses the ability to add or pull LFS content: git add on any data/**/*.qasm or *.trans file fails since filter.lfs.required=true refuses to fall back to a plain add. Note in the README that git-lfs is needed for data/ but not for building or testing, since tests/fixtures/ is plain git content.
bqskit_compile_cliffordt.py: drop --optimization_level/-l and everything that existed only to support levels 3/4 (whole-circuit unitary synthesis size limits, held-measurement extraction). Always use level 1, bqskit's own default and the only level this script ever ran at in practice. qiskit_compile_cliffordt.py: drop --opt-level. Pin transpile()'s optimization_level to 1 rather than qiskit's own default (None, which resolves to 2): level 2 restructures which single-qubit gates sit adjacent to which cx gates, and on an 18-qubit/600-gate slice of a Hubbard benchmark that raises the T count from 1858 to 7085 for no change in cx count, so the default was never worth exposing as a choice. The --no-optimize flag is unrelated and untouched: it toggles this script's own post-synthesis cleanup, not qiskit's optimization level.
Replace qiskit_compile_cliffordt.py with compile_cliffordt.py, adding a
--backend {qiskit,bqskit} flag so both paths share the {u,cx} preopt step,
verification, report printing, and QASM writing.
bqskit_compile_cliffordt.py is kept on disk for reference but is superseded
by --backend bqskit.
- Both backends now resynthesise from the same
transpile(basis_gates=["u","cx"]) preopt step; previously only the qiskit
compiler had one. Reduces the bqskit backend's T count on the reference
slice even before any other change here.
- Drop the sk (Solovay-Kitaev) rotation-synthesis option: the existing
pygridsynth fallback already covers gridsynth failing on part of a
circuit, so sk was never doing useful work as a fallback, only serving
as a strictly worse standalone mode.
- Add FastGridSynthPass, swapping bqskit's pygridsynth-based GridSynthPass
for qiskit's Rust gridsynth_rz (~8x faster per rotation, and empirically
shorter sequences at the same epsilon). Runs with num_workers=1 rather
than bqskit's default parallel worker pool: running it in parallel was
found to make the same RZ angle synthesize into different, equally
valid sequences across otherwise-identical runs at the same seed --
a concurrency issue in that Rust extension, not in this code, confirmed
by isolating the extracted angles (bit-identical across runs) and by
comparing a direct in-process call (deterministic) to the same call
routed through bqskit's worker pool (not deterministic). Serial
execution was also measured faster in practice at this pass's rotation
counts, not just safer.
- Give the bqskit backend an error bound via bqskit's own
calculate_error_bound mechanism (exact per-block unitary distance,
composed via PassData.update_error_mul), comparable in spirit though
narrower in scope than qiskit's existing per-rewrite error bound.
- --epsilon's default is now backend-specific (1e-10 qiskit, 1e-8 bqskit)
rather than a single shared value: tightening bqskit's default to match
qiskit's measured ~2.7x the T count for no measurable accuracy gain,
since bqskit already saturated this script's fidelity check at its old
default.
- Extend the recognized Clifford+T basis to include sx/sxdg, which bqskit
emits natively as Clifford generators and the transpile binary already
understands.
- Add per-stage timing (load/preopt/compile/verify/write/total) to both
the printed report and --stats JSON.
Update README.md and data/hubbard_18_slice600.qasm's reference invocations
and numbers to match.
…ed fallback
Basis check and error bound are cheap (no simulation) and now run
unconditionally, regardless of --verify -- previously a broken non-Clifford+T
output only failed the run if --verify happened to be passed.
--verify now gates only the numeric fidelity check, and collapses from four
threshold flags (--verify-max-qubits, --verify-max-gates,
--verify-statevector-qubits, --verify-statevector-ops) down to just --verify
itself; their former defaults become fixed internal constants.
Replace the old "no numeric check" dead end (hit by, e.g., a 64-qubit circuit
needing ~825 exabytes for a direct statevector comparison) with an automatic
random-window sampling fallback: sample a handful of random contiguous
windows of the source circuit's instructions, greedily grown to stay under a
qubit/cost budget derived from the existing statevector check's own tuned
constants, independently recompile and fidelity-check each. This is always
tractable, at a bounded total cost, regardless of the real circuit's size --
window count and size are chosen automatically, not exposed as flags. Circuits
with classical control flow, which the direct checks always skip, now fall
into this same path too, since a window is control-flow-free by construction.
Caught during testing: the window budget check only accounted for the source
window's gate count, but Clifford+T compilation can blow gate count up by two
orders of magnitude (a 59-gate window compiled to 6944 gates), which would
have made its statevector check take on the order of 20 minutes. Added a
second budget check after compiling, using the real compiled gate count, so
an unexpectedly expensive window is skipped in favor of another rather than
run anyway.
Also extract compile_dispatch() from main()'s inline backend branch, since
window verification needs to recompile arbitrary small circuits the same way,
and add a fidelity_method field ("dense"/"statevector"/"windowed") to --stats
JSON so consumers can tell an exact check apart from a window-sampled one.
FastGridSynthPass (qiskit's Rust gridsynth_rz swapped into the bqskit backend's rotation-synthesis stage) gave no measurable end-to-end speedup on real benchmarks -- the initial 2-qubit block instantiation in bqskit's own compile() dominates wall time regardless of how fast the final rotation synthesis is. Worse, calling it repeatedly from inside bqskit's worker-process model proved unreliable on a real 36-qubit circuit (qv_N036) that the stock path compiles successfully: an unguarded pyo3 panic on some angles at coarse epsilon (bqskit's task-stepping loop only catches Exception, not the BaseException a panic surfaces as, so it took the whole worker process down rather than just failing one call), and, once that panic was caught and handled with a pygridsynth fallback, a separate `TypeError: cannot pickle '_MPMathModule' object` and worker hangs on the same circuit. Both point at instability in calling that Rust extension from bqskit's forked workers, not a one-line bug -- not worth the reliability cost for a change that wasn't measurably faster in practice. decompose_rz_gauge_fixed now inlines bqskit's own stock GridSynthPass (pygridsynth-based, no Rust extension involved) in place of FastGridSynthPass, keeping calculate_error_bound=True on that pass so the bqskit backend's error bound still covers rotation synthesis, not just the exact gauge-fix rewrite. Compiler() goes back to bqskit's default parallel workers, since the num_workers=1 restriction existed only to work around FastGridSynthPass's determinism bug and no longer applies. Confirmed: qv_N036 now compiles successfully in ~496s (vs. ~691s for the pre-merge standalone script, and with substantially fewer T gates), and determinism is unaffected. Update the reference numbers in data/hubbard_18_slice600.qasm's header and the module docstring accordingly.
decompose_rz_gauge_fixed computed GridSynthPass's digit-count precision as int(log10(1/epsilon)) + 2, copied from bqskit's own build_cliffordt_workflow. That padding made GridSynthPass synthesize to 1e-10 whenever --epsilon was left at the bqskit backend's default of 1e-8, costing ~25% extra T gates for no accuracy gain -- confirmed by comparing against the previously reverted FastGridSynthPass, which took epsilon as a literal float with no such padding, and by holding epsilon fixed while varying only this precision value across several benchmarks and seeds. Switch to ceil(log10(1/epsilon)): the smallest digit count that still meets epsilon (unlike plain int(), which can undershoot for non-power-of-10 epsilons), with no gratuitous overshoot. Update the module docstring's epsilon-tightening comparison and hubbard_18_slice600.qasm's reference numbers accordingly.
…tion GaugeFixedZXZXZDecomposition duplicated a fix now merged into bqskit itself (this repo's bqskit/ clone, branch ZXZXZ-fix, submitted upstream as a PR). decompose_rz_gauge_fixed now uses stock ZXZXZDecomposition directly instead of patching around the bug locally -- pip install -e ./bqskit picks up the fix ahead of an upstream release; without it, stock bqskit's known doubling bug applies as before. Removed the now-unused class and its helpers/imports, and updated docstrings/help text that referenced it. Also trimmed comments across the file down to explanations of current implementation choices, dropping narration of past changes (FastGridSynthPass, the precision-padding bug) whose code no longer exists to refer to. Updated hubbard_18_slice600.qasm's reference numbers with the fix active -- notably --no-rz-gauge-fix improved too, since that path runs bqskit's own internal ZXZXZDecomposition and was never reachable by the old local patch.
The flag never really controlled "is the ZXZXZ gauge bug patched" -- that now depends solely on which bqskit is installed, unrelated to this script. What it actually toggles is whether compile_bqskit routes rotation resynthesis through this script's own post-processing pass (which tracks an error bound) or bqskit's inline decompose_rz=True workflow (which doesn't). Renamed the flag and its matching internal names (gauge_fix parameter -> use_custom_rz_decomposition, decompose_rz_gauge_fixed -> decompose_rz_tracked) to describe that instead. Also added targeted pyright ignores on the bqskit.ft imports: bqskit.ft is a separate distribution that only resolves via a pkgutil.extend_path() call in bqskit/__init__.py (a runtime sys.path merge static analysis can't evaluate), so pyright otherwise reports them as unresolved regardless of environment.
…qskit
build_multi_qudit_retarget_workflow (core bqskit) is gated on a predicate
that fires for any circuit with 2+ qubits, not just ones needing
retargeting, so it was numerically re-synthesising every already-native
<=3-qubit block and discarding exact Clifford+T structure -- 12x more T
gates measured on a CSWAP-heavy benchmark. build_bqskit_workflow replaces
bqskit's own build_circuit_workflow with the same pass list minus that
stage, since unroll_to_u_cx already guarantees bqskit only ever sees native
{u, cx} gates.
bqskit now beats qiskit's T count on every benchmark measured, so it
becomes the default --backend.
The bqskit-wins conclusion from the previous commit compared each backend at its own (different) epsilon default, which isn't a fair comparison. At matched epsilon, qiskit produces fewer T gates than bqskit on every benchmark measured, so it goes back to being the default backend. Measured via exact dense-unitary fidelity (not this script's coarser --verify checks) that qiskit's own former default of 1e-10 bought no real accuracy over 1e-8 -- infidelity plateaus by ~1e-8 for both backends' resynthesis, so tightening further only cost T gates for no gain. Both backends now share one EPSILON_DEFAULT (1e-8) instead of two different per-backend constants. Also documents that the reported error bound (either backend) is a real but very loose upper bound -- typically several orders of magnitude worse than the actual measured infidelity -- not an estimate of real accuracy.
…entinel The sentinel-and-resolve-after-parsing pattern was needed when the default depended on which backend was chosen. Now that both backends share one EPSILON_DEFAULT, it's unnecessary and was causing --help to show "(default: None)" -- argparse's ArgumentDefaultsHelpFormatter reports the literal default= passed to add_argument, not the value substituted in afterward.
It was previously an untracked plain directory, but it's already its own clone with a remote and full history -- tracking it as a submodule pins an exact upstream commit reproducibly instead of hitting git's "embedded repository" warning.
cyclosynth (git submodule, previous commit) synthesizes a whole single-qubit block's ZYZ Euler angles into a near-T-optimal Clifford+T word in one call, via a diamond-distance lattice search, rather than gridsynth's per-elementary-rotation Ross-Selinger algorithm. It reuses the qiskit backend's rewrite_single_qubit_runs pipeline entirely (compile_qiskit renamed to compile_via_resynthesis, now shared by both), differing only in CyclosynthSynthesizer's resynthesis of a non-Clifford block. Measured at EPSILON_DEFAULT: fewer T gates than the qiskit backend on every benchmark tried (e.g. dnn_n8: 13316 vs 23592 T), at essentially the same real (exact dense-unitary) infidelity, confirming the comparison is fair despite the two backends' epsilon meaning slightly different things. cyclosynth's own lattice search is parallelised via rayon with no per-call seed, so results vary slightly run to run unless pinned to one thread (~15x slower at EPSILON_DEFAULT, measured). --cyclosynth-threads exposes this as an explicit tradeoff, defaulting to rayon's own (fast, non-reproducible) thread count. Also promotes CliffordTSynthesizer._from_word to a module-level circuit_from_word, reused by both synthesizer classes.
Both backends funnel through compile_via_resynthesis's shared rewrite_single_qubit_runs pipeline, so a single addition covers both: count_resynthesis_blocks does a cheap dry pass (matrix merges only, no gridsynth/cyclosynth calls) to get a total up front, and _with_progress wraps the real resynthesize callback to report percentage through it. Progress overwrites a single line in place (\r, no newline until done) rather than printing one line per update, throttled to once per PROGRESS_INTERVAL_SECONDS plus a final line on completion -- a fast compile only shows its completion state instead of spamming near one line per block. Respects -q like all other logging, since it's threaded through the same log() callable. bqskit is deliberately not covered: it exposes no public per-block callback, and the only signal (DEBUG-level runtime log lines) would need splitting decompose_rz_tracked's atomic compile() call in two -- with the two halves' error bounds recombined by hand -- plus unverified logging overhead inside bqskit's own runtime pipeline. Not worth the risk for a backend that's kept for comparison, not because it's competitive.
QFT-family circuits reuse the same handful of rotation angles across O(n^2) blocks, so with plain block counting the progress indicator raced through the first few percent (the actual gridsynth/cyclosynth searches) then jumped straight to 100% on the remaining cache hits. Progress now tracks growth of the synthesizer's own cache dict, so it advances roughly in step with real elapsed time instead of traversal order.
…lify progress display cyclosynth's Clifford+T search can hang or fail to terminate on single-qubit rotations very close to (but not exactly on) a Clifford element -- e.g. a QFT's deep phase gates. Root-caused via an instrumented local build (see cyclosynth-bug-report.md, to be filed upstream): the cost is in cyclosynth's own per-prefix lattice setup, not gridsynth's algorithm, which already handles these angles fine. Detect such blocks up front via distance to the nearest of the 24 single-qubit Cliffords and route them to CliffordTSynthesizer's gridsynth path instead of ever calling cyclosynth, with a catch-None fallback as defense in depth. Required splitting CliffordTSynthesizer.synthesize into a pure compute step and an accounting step so CyclosynthSynthesizer can reuse the former for its own bookkeeping. Also simplify the resynthesis progress line to show only a percentage -- the underlying counts (distinct new cache misses) differ enough in meaning between backends that showing them invited comparing them directly, which isn't meaningful.
Mirrors puremagic.rs's build-info banner (Git branch/Commit/Built), read at run time instead of baked in at compile time since a Python script has no build step -- "Run:" uses the current time instead. Printed unconditionally, even under -q, since it's a one-time identification line rather than per-file progress noise.
…wline The progress counter only incremented by 1 whenever the synthesizer's cache grew "at all", but a single CliffordTSynthesizer block can need gridsynth for more than one Euler angle (Rz and Rx) in one call, growing the cache by 2 or 3 at once -- confirmed via dnn_n8.qasm, where the cache grew by 36 entries across only 17 growing calls, permanently stalling the reported percentage at 47%. Count the actual per-call delta instead of a fixed +1. Also stop inferring completion from count reaching the (only ever estimated) total: compile_via_resynthesis now prints the final 100% line unconditionally once the real pass returns, so a residual mismatch between the estimate and actual cache growth (e.g. from the cyclosynth gridsynth fallback) can no longer stall the line the same way again.
…g it rsgridsynth's known, already-handled panic (falls back to pygridsynth) was still dumping its own message and a ~25-line Rust backtrace directly to fd 2 before the catch ran, regardless of RUST_BACKTRACE (confirmed empirically -- that env var has no effect on this specific output, so an earlier attempt to set it as a default is now dead weight, kept only for the harmless case some other panic in the chain does honor it). Redirect the OS-level stderr fd to a temp file around the gridsynth_rz call and inspect what it captured before deciding what to do with it: the one known panic is now fully silent (already measured and accounted for in error_bound, nothing the user needs to see), while anything else is forwarded verbatim plus a note, since it hasn't been verified safe to hide. Always restores the real fd via try/finally so a bug in the capture logic itself can't leave stderr silently broken for the rest of the process.
compile_via_resynthesis's post-resynthesis cleanup loop (up to 5 rounds of cancel_inverses + shorten_run) was entirely silent, even though on a large enough circuit it costs as much as the main resynthesis pass -- measured on a 36-qubit, ~1.2M-gate QV circuit: ~28s resynthesizing, then ~16s per cleanup round for 5 rounds, invisible before this. Add count_resynthesis_blocks/_with_block_progress: shorten_run never calls gridsynth/cyclosynth, so unlike the main pass there's no cache-hit/cache-miss split to weight by -- plain per-block counting is the right denominator. Blend each round's own percentage into a single continuous "cleanup: X%" line spanning all rounds (this round contributes 1/max_rounds of the overall range) rather than resetting to 0% and starting a new line every round. Since the loop can converge and break before max_rounds is reached, the blended percentage can plateau below 100% -- the final 100% is printed unconditionally once the loop exits, same reasoning as the main pass's.
… resynthesis
Adds merge_phase_polynomial, an exact preopt step that cancels/merges
rz-family rotations sharing the same phase-polynomial "parity" (the t-par
technique): two such rotations anywhere in a {cx, diagonal-gate} region
commute and add exactly, regardless of physical qubit or what runs between
them. QFT-family circuits are full of this redundancy from their
CX-ladder-decomposed controlled-phase gates -- qft_N032's needed rotation
count drops from 1350 to 500, cutting T roughly 2.7x for the qiskit backend
(bqskit benefits too, sharing the same preopt).
The merge isn't always a net win: moving a rotation to satisfy a global
parity match can break a neighbouring gate's local gauge-cancellation,
costing more real rotations than it saves on circuits without much genuine
redundancy (measured 10% more T on a random Quantum Volume benchmark).
CliffordTSynthesizer.count_real_rotations gives a cheap, gridsynth-free way
to compare the merged and unmerged candidates, so unroll_to_u_cx always
keeps whichever is actually cheaper rather than assuming the merge helps.
Also fixes a pre-existing crash: compile_via_resynthesis and
compile_dispatch's default log callbacks didn't accept the `end=` kwarg
the progress wrappers always pass, so any --verify run falling back to
windowed sampling on a large circuit (>24 qubits) with the qiskit or
cyclosynth backend crashed with a TypeError instead of reporting fidelity.
TRbO (arxiv.org/abs/2603.25101) numerically re-optimizes a partitioned block's Rz angles jointly rather than rounding each independently like RoundToDiscreteZPass, catching gauge-freedom reductions that phase- polynomial merging can't (verified: 0% change on QFT-family/Hubbard circuits, a real few-percent T-count reduction on Haar-random circuits like QV). Off by default via --bqskit-trbo since it's a real compile-time cost, not a free rewrite, and its own multi-start search isn't reproducible at a fixed --seed the way the rest of this script is. Also suppresses two RuntimeWarnings from TRbO's own gradient computation (a benign 0**negative singularity at exact convergence) via PYTHONWARNINGS, since bqskit's worker processes don't inherit this process's warnings filters.
…oducible - Enable --gpus all in the devcontainer now that host support is in place. - Pin requirements.txt's bqskit to an editable install of the local ZXZXZ-fix fork instead of stock PyPI bqskit, which was silently reinstalling on every rebuild and losing the T-count fix with no error. - Automate cyclosynth's maturin build in postCreateCommand instead of leaving it as a manual step, including the system libs and PYO3_USE_ABI3_FORWARD_COMPATIBILITY it needs. - Add pyright, used throughout development to verify compile_cliffordt.py.
A devcontainer rebuild wiped the session-state directory, losing every saved transcript, because it lived in the container filesystem with no mount behind it. - Mount a named volume over the state directory so transcripts, resume history, and saved notes survive a rebuild. Docker creates the volume root-owned, so postCreateCommand chowns it to vscode on the rebuild that first creates it; a no-op thereafter. - Drop the ANTHROPIC_AUTH_TOKEN/ANTHROPIC_BASE_URL and ANTHROPIC_*_MODEL remoteEnv passthroughs. Sign-in goes through the browser login, which stores OAuth tokens in the state directory, so these were resolving to empty strings and doing nothing. Keeping AUTH_TOKEN around is a mild hazard: set on the host it would take precedence over the OAuth credentials and silently switch auth paths. - Gitignore the pre-rebuild snapshot directory, since a tarball of the state directory contains live credentials.
Pinning bqskit to an editable install of the local ZXZXZ-fix clone (87ccc86) left `from bqskit import Circuit` failing outright with "unknown location", so compile_cliffordt.py could not start at all under the bqskit backend. bqskit-ft installs its files into site-packages/bqskit/ft/, which relies on bqskit itself being a normal site-packages install so the two share one directory. With bqskit installed editable there is no site-packages/bqskit/__init__.py, leaving that directory a bare namespace portion. setuptools appends its editable finder to sys.meta_path *after* PathFinder, so PathFinder resolves the namespace first and wins; the fork's own __init__.py never executes, which is also why the extend_path call added in the fork (bqskit e94d841) to merge bqskit.ft back in never fired. Naming the fork's repo root in PYTHONPATH makes PathFinder find a real package there and prefer it over the namespace portions, after which extend_path merges site-packages/bqskit/ft in as intended. Verified end to end: qv_N008_12345 compiles to 27084 T, matching the baseline recorded in 625a074 for that circuit, so the ZXZXZ fix is active.
Assistant-generated commits were adding attribution trailers that do not belong in this project's history. Put the rule somewhere that is loaded automatically at the start of every session and is version-controlled, so it survives devcontainer rebuilds and fresh clones rather than living only in per-machine configuration.
CLAUDE.md states the rule but nothing checked it, and the trailers it forbids are exactly the kind an automated committer adds without being asked. The hook rejects Co-Authored-By trailers naming an assistant, and the assistant no-reply address when it appears in a trailer. Both checks ignore commented-out lines, and the address check is anchored to trailer lines rather than matched anywhere, so a message can still discuss the address in prose -- this commit message needs to. Descriptive prose is untouched: documenting a change to assistant configuration still has to name the paths involved. Kept in a tracked .githooks/ with core.hooksPath rather than .git/hooks/, so it is version-controlled and reaches fresh clones. postCreateCommand sets core.hooksPath, since .git/config survives a rebuild but a clone starts without it. Bypass a single commit with --no-verify.
Redirecting a run to a file (2>&1 | tee out, or just > out) recorded every throttled progress update as a literal ^M-separated run on one line: \r only means "move to column 0, about to overwrite" to a terminal, and is just another byte otherwise. _progress_line now checks sys.stdout.isatty() and picks the right rendering: the existing \r-overwrite on a real terminal, unchanged, or one plain line per update when redirected, since overwriting has no meaning there and the old behaviour just produced unreadable log files.
…t.isatty() The previous fix (22dbd42) checked sys.stdout.isatty(), which is false the moment stdout is piped through `tee` even though a real terminal is watching on the other end -- tee duplicates whatever reaches stdout byte-for-byte, so there was no way for the terminal to see an overwritten line while the file tee writes stays clean; the check picked one or the other, and coudn't do both for the same underlying request (the actual ask here). _progress_line now writes the \r-overwrite straight to /dev/tty, cached and opened once, bypassing log()/stdout entirely whenever a controlling terminal exists. tee/redirection never sees these bytes, so the terminal still gets the live single-line ticker and any captured output is clean regardless of how it's captured. Falls back to one plain line per update via log() only when there's no controlling terminal at all (headless: CI, a detached cron job) -- confirmed open("/dev/tty") genuinely raises ENXIO there. -q's suppression, previously implicit in log() being a no-op, is now explicit (_no_log, checked by identity) since the /dev/tty path bypasses log() entirely and would otherwise ignore -q.
Flushing the Clifford queue emits every queued gate, but most of what
that costs buys nothing. Only two-qubit gates spread a T gate's Pauli
across qubits; a single-qubit Clifford merely rotates it, leaving the
weight alone. Yet 1Q gates dominate the bill - S/SX span 3 logical
cycles each against CX's 2, and they are 61-97% of a flush's layer
cost across the benchmarks. So the flush criterion was arguing over a
cost that is largely avoidable.
Each qubit's trailing 1Q run - the gates after that qubit's last 2Q
gate - acts last on its own wire, so the queue splits as U = R*E with
R a tensor product of single-qubit Cliffords. Holding R back rather
than emitting it keeps every later product weight 1 (merely rotated,
since a local Clifford maps Z_q to a single-qubit Pauli on q), and R
then merges with the next segment's leading 1Q gates where
optimize_single_qubit_sequence can collapse the combined run. The
split needs no commutation and no resynthesis: it is exactly the
residue optimize_clifford_sequence already flushes in its final
per-qubit loop.
- split_trailing_1q returns the gates to emit and that trailing layer;
optimize_clifford_sequence is now a thin wrapper over it, so the
fixed-weight path is byte-for-byte unchanged.
- flush_clifford_queue takes defer_trailing, keeping the residual
queued and rebuilding the tableau from it. Products emitted after a
partial flush are conjugated through that residual, not identity.
- predicted_depth scores a transpilation by the layering
Circuit::compute_layers derives downstream, charging each product
its logical-cycle span. transpile_best transpiles at each candidate
weight and keeps the shallowest.
- New --auto (select weight by predicted depth, writing .transauto)
and --defer_trailing, which is orthogonal to the weight threshold
and can be applied to a fixed weight too.
Candidate weights are just {0, 1}. Deferral makes flushes cheap enough
that the optimum collapses to either "never flush" or "keep every
product weight 1": those two reproduce the full 0..=10 sweep's best
depth exactly on 8 of 11 circuits and within 1.5% on the rest. This
also shifts where the old sweep landed - fermi_hubbard 3->1, qft 2->1,
qv 2->1 - so results/circuits/best-weights.txt is stale.
Scheduled with puremagic -u against each circuit's hand-tuned weight,
logical cycles (volume tracks them exactly):
qugan_n111 13371 vs 16739 -20.1%
square_heisenberg_N225 18589 vs 21553 -13.8%
qft_n160 51406 vs 56239 -8.6%
fermi_hubbard_1d_128q 542284 vs 571475 -5.1%
qaoa_..._N149_3reps 15766 vs 16105 -2.1%
ising_n98 2060 vs 2084 -1.2%
qv_N064_12345 154643 vs 155739 -0.7%
dnn_n51 7915 vs 7781 +1.7%
cdkm, grover and knn select weight 0, where nothing is flushed and the
output is byte-identical to .trans0 - already their hand-tuned choice.
Suite total 806034 vs 847715 lcycles, 4.9% fewer. dnn regresses because
its hand-tuned weight of 4 sits outside the candidate set; widening the
set recovers it at the cost of more passes.
Deferral also lowers Y-basis exposure rather than raising it, which was
the worry given a Y operator occupies two data nodes: on qv, Y ops per
product fall 0.514 -> 0.344 and mean node weight 2.08 -> 1.36.
It has no production callers anywhere in the repo -- transpile.rs is the only consumer of Tableau and never calls .len() -- so its only purpose was satisfying its own dedicated test. Remove both.
…cuit_stats puremagic.rs and circuit_stats.rs each reimplemented the same "minimum layers / minimum volume" estimate independently, and had drifted: puremagic used plain subtraction (min_layers - n_t_layers, underflow risk) and a truncating cast for the magic-state throughput bound (n_t_gates / max_t_parallelism, no zero-guard, floors instead of rounding up a required minimum), while circuit_stats already used the safer saturating_sub and a zero-guarded, ceiling-rounded division. Move the shared logic into Circuit::estimate_layer_volume in circuit.rs, using circuit_stats' safer arithmetic as the one canonical version, and have both binaries call it. Flooring a required-capacity lower bound was simply wrong, so puremagic's printed "Max parallelism estimate" and "Volume estimate" shift by a small amount on circuits where the division isn't exact -- the real scheduling output (the .schedule file, actual scheduled cycle counts) is unaffected, since this estimate is a separate diagnostic computed after scheduling already completed. Added direct unit tests for the consolidated function covering the zero-magic-qubit and zero-T-gate edge cases that motivated pulling this out of two copy-pasted inline blocks in the first place.
magic_state_lambda, ancilla_rows, rseed, circuit_fname, and no_t_failures were each declared independently in both binaries' Args structs, and had already drifted once: rseed was u32 in puremagic but u64 in circuit_stats (StdRng::seed_from_u64's native type), for the same "--rseed 29" flag. Pull these into a shared CommonArgs in utils.rs, flattened into each binary's Args via #[command(flatten)], so the flag names, defaults, and types can't diverge between the two independently again. Standardize on u64 for rseed; puremagic casts down to u32 at the two internal call sites (Scheduler::new, TopoGraph::set_topo) that still take u32, which is a no-op for every seed value that was ever representable through the old u32 CLI arg.
puremagic and compile_cliffordt each printed their own git/build banner independently, and had drifted: puremagic sliced the commit SHA with a fixed &sha[0..8], which panics if the SHA is ever shorter than 8 characters, while compile_cliffordt already guarded this with sha[..sha.len().min(8)]. transpile, circuit_stats, and gen_circuit printed no banner at all. Move this into utils::print_banner(name), using the panic-safe slicing, and call it from all 5 mains for consistent UX. puremagic still needs the returned string to fold into its .schedule file's header, so the function returns what it prints rather than being print-only.
The scheduler trace feature stays exactly as it behaves today: a useful diagnostic in debug builds, a no-op in release (debug_sched!/ info_sched! compile out entirely there, per utils.rs). What was missing was documentation of that split, plus a stale doc-comment on puremagic's --log-scheduler field claiming it writes a ".sched" file when the real output is "<name>.sched_trace". Fix the doc-comment, note debug-build-only in the --help text, and update the README's CLI options and Output Files tables to match. No behavioral change.
spread_probability/decay_factor were manually range-checked in main() with eprintln!+std::process::exit(1), despite main() already returning a Result -- inconsistent with itself, and with the value_parser-based validation puremagic.rs already uses for its own custom-constrained args (log_scheduler, plot). Move the [0.0, 1.0] range check into each field's value_parser, matching that existing convention. Exit code stays non-zero on invalid input either way (clap's own parse-error exit code), so tests/integration_test.rs's gen_circuit_rejects_invalid_* tests (which only assert non-zero exit) are unaffected.
set_from_str indexed into the input string by character position from .chars().enumerate(), then sliced the string at that same position (&s[i..]) to pull out the gate tag. For the ASCII-only .trans format this is always in bounds, but it's a latent panic (slicing at a non-char-boundary byte offset) waiting for any malformed input with a multi-byte character earlier in the line. char_indices() gives the byte index directly, which is what the slice actually needs, and is numerically identical to the old char position for every real (ASCII) input -- no behavior change on valid circuits, just no longer relies on ASCII-ness to avoid a panic. Caught by clippy's char_indices_as_byte_indices lint during a repo-wide sweep. A repo-wide cargo clippy --all-targets --all-features run (per-binary, since the 5 binaries are separate crates over the shared modules) otherwise found no further dead code beyond what earlier commits in this series already removed.
convert_lss_to_qasm.py and convert_lss_to_trans.py each reimplemented the same regex-based Rotate/Measure line parser independently, and had drifted: qasm.py raised on any unparseable/unknown-sign line, while trans.py only raised for unknown-sign lines but warned-and-skipped on lines that failed the regex outright (and even that only landed on one of two overlapping "+1"/"-1" sign-acceptance rules between the two files). Neither script is called by anything else in the repo. Move the shared parsing into lss_common.py, harmonized to the stricter behavior (raise on any unparseable/unknown input) in both -- a deliberate, documented change to trans.py's tolerance for malformed input, not just a refactor. Also fixes stale documentation: both scripts' epilogs described an output format neither one actually produces. trans.py's own epilog example even contradicted its own module docstring for the identical input line (<pi/8> vs the correct <pi/4>). qasm.py's docstring described the old `<pi/8>`-tagged Pauli-string format, when it's actually emitted OpenQASM 2.0 gate lines (t/tdg/cx/h/s/sdg/measure) for a long time. Rewrote both to describe their actual current output.
Both scripts independently implemented the same "ensure output dir exists, dump circuit as OpenQASM 2.0, print a confirmation" sequence. Move it into gen_common.py's write_benchmark_qasm, used by both; their actual circuit-construction logic (QFT/transpile vs build_oracle/build_diffusion/build_grover) is unrelated and untouched. Verified byte-identical generated .qasm files before/after for both scripts across a sample of qubit counts.
plot_cultivation_dist.py copy-pasted the identical 8-colour palette, with a comment admitting it was "consistent with plot_puremagic.py" rather than importing it. plot_volume_vs_layers.py already imports from plot_puremagic for its own parsing needs, so this brings plot_cultivation_dist.py in line with that established pattern.
plot_puremagic.py, scheduling_table.py, and circuit_table.py each reimplement their own state machine over puremagic's stdout (their "when does a run start/end" logic differs deliberately per script, so that part stays put), but several of the underlying regex patterns were typed out identically (or near-identically, e.g. greedy vs non-greedy) in two or three of them: the ANSI-escape stripper, magic_state_lambda, Layers, Loaded-circuit, and Scheduled-products-written-to-<name>.schedule. Move just those into puremagic_log.py and have all three import them, so a future fix to one of these patterns doesn't need to be applied three times identically. Left the patterns that only look similar but actually capture different things for different callers (e.g. the lcycles-only vs volume-only halves of the "Scheduled N in L logical cycles, volume V" line) alone, rather than forcing a merge that would have needed each caller's capture-group indexing to change. Verified byte-identical parsed output before/after across 4 real captured puremagic stdout logs, for all 3 scripts.
latex_escape() was byte-for-byte identical in circuit_table.py and scheduling_table.py; generate_pdf() was functionally identical (same pdflatex/xelatex/lualatex-then-pandoc fallback chain, same standalone document template) with only comment wording differing -- scheduling_table.py's own comment even said "same as circuit_table.py" next to it. Move both into table_common.py, imported by both scripts. Verified latex_escape produces identical output from both call sites, and that generate_pdf is now literally the same function object.
circuit_table.py and scheduling_table.py each had their own pretty_name (rule-for-rule identical to each other, scheduling_table.py's own comment said "identical rules to circuit_table.py"), which disagreed with plot_puremagic.py's prettify_circuit_name on 7 of the 11 real benchmark circuits -- e.g. fermi_hubbard_1d_128q -> "Fermi(128)" vs "HUBBARD(128)", square_heisenberg_N225 -> "Heis.(225)" vs "heis(225)", qv_N064_12345 -> "QV(64)" vs "QV(064)". Confirmed by listing both functions' output for all 11 real names side-by-side before touching anything. User chose prettify_circuit_name (plot_puremagic.py's curated version, e.g. ADD/HUBBARD/GROVER for those specific benchmarks) as canonical for tables too. Removed pretty_name and _UPPERCASE_NAMES from both table scripts; both now import prettify_circuit_name from plot_puremagic.py, matching the pattern plot_volume_vs_layers.py already uses. Note: prettify_circuit_name takes only a name, not a qubit-count fallback -- pretty_name always appended "(num_qubits)" even for a name it didn't otherwise recognize, prettify_circuit_name doesn't. All 11 real circuit names already get a "(N)" suffix from prettify_circuit_name's own rules, so this doesn't affect the current benchmark set, but an unrecognized future circuit name spelled without an "_n<digits>"-style suffix would now display without one where it previously always had one.
Only src/transpile.rs had drifted from .rustfmt.toml (3 spots, all in test code) -- pure whitespace/wrapping changes, no program text affected. Verified zero harness diff before/after.
Used --line-length 100 rather than black's default 88 to match this codebase's existing style (and .rustfmt.toml's max_width=100 on the Rust side) instead of reformatting every line in every file. 7 of 20 files needed changes, all pure whitespace/wrapping (comment alignment, blank-line counts, line-wrapping of long calls) -- nothing inside string literals changed. Verified: both LSS converters produce byte-identical output on a sample file, all 3 puremagic-log parsers produce identical parsed output on a real captured log, and both benchmark generators produce byte-identical .qasm files, before and after. wisq/, bqskit/, and bqskit-ft/ are gitignored vendored packages, not touched.
The --flasq -x circuit chart previously showed raw scheduled volume as bars against absolute FLASQ conservative/optimistic lines, making it hard to read the overhead relative to the lower bound at a glance. Bars are now volume/FLASQ-conservative (with a matching --minvol overlay), so the conservative line collapses to a flat 1.0 and the optimistic line is opt/cons. Also thickened both FLASQ lines, drew the optimistic ratio as a solid step plot spanning the full plot width instead of a jagged diagonal line, and lightened the bar fill so the lines stay legible on top.
…c.py pyright without pandas-stubs falls back to pandas' own inline stubs, which are imprecise enough to misreport dozens of ordinary DataFrame/Series calls (dropna, apply, to_numpy, fillna, ...) as attribute-access errors throughout data_processing/. Pulling in pandas-stubs (matching the installed pandas 3.0.5) clears all of that noise. Also initialize _flasq_entries before the `if args.flasq_file` guard: it's read again later in a separate `if args.flasq_file` block for the data_qubits FLASQ overlay, which pyright can't correlate with the first guard, so it flagged the read as possibly unbound.
… date Group this repo's own circuit generators under data/circuit-generators/ (previously loose in data/) and drop data/all_compiled/, a pre-compiled circuit cache superseded by data/circuits/ and its family subdirectories. The READMEs had drifted from the code and from each other over the Clifford+T compiler work, so bring all four back in sync with this move: - Fix two circuit README links left pointing at gen_qft.py/gen_grover.py's old data/ location. - Fix a --sides_only flag documented as --sides-only, a --skip-cyclosynth description with cyclosynth's default backwards (it's now the default, not opt-in), and an incorrect "debug builds only" claim on <name>.topo.txt (it's gated by --plot topo in any build). - Document the previously-unlisted --auto/--defer_trailing transpile flags and the heisenberg/qft ablation subdirectories under data/circuits/. - Rewrite data/README.md, stale since before data/circuits/ and the gen_*.py scripts existed. - Add data_processing/README.md, covering its 18 tracked scripts; there wasn't one. - Fix two of transpile's/puremagic's own --help strings: a stale --dynamic reference (renamed to --auto) and a --plot cstats output path that says .png but is actually .svg.
cyclosynth is a git submodule, but actions/checkout@v4 does not initialize submodules by default, so cargo test failed on CI with a missing cyclosynth/Cargo.toml.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a pure-Rust Clifford+T compilation pipeline (
compile_cliffordt), replacing thePython qiskit/bqskit/cyclosynth-based backend as the primary compiler, and carries
along a batch of scheduler/transpiler correctness and efficiency work developed
alongside it.
Clifford+T compiler (
src/cliffordt/, newcompile_cliffordtbinary)Clifford/non-Clifford partitioning, rotation synthesis (via the
cyclosynthsubmodule), NLS instantiation (
levenberg-marquardt, replacing a hand-rolledsolver), Clifford simplification, and post-synthesis scan-removal cleanup.
cliffordt::qasm) supporting custom gate macros andMQT Bench's OQASM3 export format.
nalgebra,num-complex,levenberg-marquardt,rayon(enabled),cyclosynth(added as a git submodule).data_processing/compile_cliffordt.pykept as a thin qiskit/bqskit baseline forcomparison rather than removed outright.
Scheduler / transpiler fixes and optimizations
apply_t_conjugationalways swapped X↔Y as if a pendingcorrection were Z-basis; it should use the failing T-gate's own axis
(
correction_basis). Real circuits target all three axes, so this was silentlymisrouting later T operands. Fix shifts scheduling outcomes on
dnn_n51.trans0(seed 42): 9903→9875 logical cycles, 7346→7031 A* calls.
always materializing a standalone correction gate first (37/534 vs. 50/534
materialized corrections on
dnn_n51.trans6).new
--autoflag that picks the flush-weight threshold by predicted depth.4.9% fewer logical cycles across the benchmark suite (up to 20% on some
circuits).
Shared-code cleanup
puremagicandcircuit_stats, and unified their CLI args and startup banner (also applied togen_circuit,transpile,compile_cliffordt).Tableau::len().cargo fmtpass acrosssrc/.Data processing / plotting (
data_processing/)compile_cliffordt.pybaseline (qiskit/bqskit),mcmc_cultivation.py,plot_volume_vs_layers.py,remove_y_gates.py.table_common.py(LaTeX tables),lss_common.py(LSS parsing),puremagic_log.py(shared log regexes).prettify_circuit_name.plot_puremagic.pyreworked for the FLASQ circuit-volume conservative-boundnormalization and a layer x-axis option.
Infra / tooling
.githooks/commit-msg: blocks AI-assistant attribution trailers in commits.cyclosynthadded as a git submodule; devcontainer updated for GPU passthroughand reproducible builds of the bqskit fork and cyclosynth.
git-lfsdocumented as a prerequisite;data/circuits/adds a curatedbenchmark circuit set with its own README.
data/refreshed (stale/duplicate result files removed, new runsadded).
Test plan
cargo testpasses, including new Clifford axis-conjugation regressiontests in
scheduler.rscargo fmt --check/cargo clippycleancompile_cliffordtruns end-to-end ondata/circuits/benchmarks andT-count/depth results match the numbers cited in the corresponding commits
data_processing/scripts still generate expected plots/tablesagainst the refreshed
data/