Fix Burgers docs tutorials: Hugging Face dataset mirror + Reactant-safe GridEmbedding - #174
Merged
ChrisRackauckas merged 3 commits intoOct 5, 2026
Conversation
…e Drive
The Burgers tutorials downloaded burgers_data_R10.mat from Google Drive
via gdown. Drive rate-limits shared files ("Too many users have viewed
or downloaded this file recently"), which failed CI docs builds with
FileURLRetrievalError whenever the quota was exhausted.
Fetch the same dataset from the kks32/sciml-dataset Hugging Face mirror
over plain HTTPS with a pinned sha256 checksum instead. This also drops
the PythonCall/CondaPkg/gdown stack from the docs environment. The
DataDep is renamed to BurgersR10 so machines with a stale Drive-zip
cache refetch instead of reusing the old layout.
Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Agent-Harness: OpenCode
Agent-Model: muse-spark-1.3-contributor-free
Agent-Session: local session at /home/crackauc/sandbox/tmp_20260917_040834_91074 (no public conversation URL)
…acement Placing the CPU-built grid via Lux.get_device(x) throws inside Reactant.@compile (get_device errors on TracedRArray), which broke the Burgers FNO docs tutorial training step. Allocate like x and broadcast-assign instead: no runtime device query, and the grid still lands on x's device, preserving the CUDA scalar-indexing fix from SciML#125. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Agent-Harness: OpenCode Agent-Model: muse-spark-1.3-contributor-free Agent-Session: local session at /home/crackauc/sandbox/tmp_20260917_040834_91074 (no public conversation URL)
The mirror URL used `resolve/main`, so any upstream change to the mirror's main branch would force a full re-download before the pinned sha256 could fail the build. Pin to the revision (982685ff) whose content the checksum was verified against, so docs builds are reproducible. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Agent-Harness: OpenCode Agent-Model: muse-spark-1.3-contributor-free Agent-Session: local session at /home/crackauc/sandbox/tmp_20260917_040834_91074 (no public conversation URL)
ChrisRackauckas
marked this pull request as ready for review
October 5, 2026 01:20
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Please ignore until reviewed by @ChrisRackauckas.
What changed and why
The docs build on
mainis red (https://github.com/SciML/NeuralOperators.jl/actions/runs/35161521723). Two independent causes:burgers_data_R10.matfrom Google Drive viagdown. Drive rate-limits shared files, so the build died withPython: FileURLRetrievalError: ... Too many users have viewed or downloaded this file recently, and every later@exampleblock on those pages cascaded intoUndefVarError: x_data/x_data_dev/y_data_dev. The tutorials now fetch the same file from thekks32/sciml-datasetHugging Face mirror over plain HTTPS with the pinned sha256d1a04567..., pinned to an immutable mirror revision (982685ff) rather thanmain. This also removes thePythonCall/CondaPkg/gdownstack anddocs/CondaPkg.tomlfrom the docs environment. The DataDep is renamedBurgers->BurgersR10so machines with a stale Drive-zip cache refetch instead of reusing the old unzipped layout.GridEmbeddingvs Reactant (src/layers.jl). fix: transfer GridEmbedding grid to input device (fixes CUDA scalar indexing #125) #132 placed the CPU-built grid viaLux.get_device(x)(grid), butget_devicethrows onTracedRArrayinsideReactant.@compile, breaking the FNO tutorial training step. Allocate likexand broadcast-assign instead: no runtime device query, grid still lands onx's device, preserving the FNO invokes scalar indexing #125 CUDA fix.Verification (local, Julia 1.12.4)
Full docs build,
julia --project=docs docs/make.jlfrom the repo root (afterjulia --project=docs -e 'using Pkg; Pkg.develop(PackageSpec(path=pwd())); Pkg.instantiate()'):Zero failed
@exampleblocks and nomakedocs encountered an error [:example_block]; both tutorial pages rendered (docs/build/tutorials/burgers_deeponet-*.png258827 B,burgers_fno-*.png247647 B). The output is byte-identical to the pre-pinning build, so the pinned-revision commit changes no behavior.Pinned mirror fetched through DataDeps' own fetch+checksum pipeline into a fresh DataDep (
BurgersR10Pinned, no cache):PATH=.../datadeps/BurgersR10Pinned/burgers_data_R10.mat,SIZE=644427710, sha256d1a0456776255a4bd841dbc18951d3f468266d945d96e24ae531a12f18bb5a1a(matches the pin).Fail-before evidence on the same machine/paths: full build failed at
burgers_deeponet.md:5-50withFileURLRetrievalError, cascading toUndefVarError: x_data / x_data_devon the following blocks; and before theGridEmbeddingcommit the FNO training step failed insideReactant.@compilewithget_device"isn't meant to be called insideReactant.@compilecontext".Not verified
similar+broadcast grid placement (no GPU locally). The repo GPU CI lane passed on this branch and should confirm FNO invokes scalar indexing #125 stays fixed.tests / Core (julia 1, ubuntu-latest, x86)is red on this PR. It fails in a dependency__init__(TypeError: in keyword argument priority, expected Int64, got Int32) and this PR does not touch the test environment; I did not trace it to amainrun from here.@examplerun reused the on-diskBurgersR10dataset cache; the fresh-fetch path was verified separately (above) and CI's Documentation run on this branch (https://github.com/SciML/NeuralOperators.jl/actions/runs/35211262876) performed the fresh download.🤖 Generated with OpenCode (model: muse-spark-1.3-contributor-free); no public session URL (local session at /home/crackauc/sandbox/tmp_20260917_040834_91074).
CI triage (2026-09-25)
Head commit
a534b3c; basemain @ b42958f. The branch is up to date with the base (git merge-base HEAD upstream/main== base HEAD), so no merge commit was needed; GitHub reports MERGEABLE.gh pr checks 174shows exactly one failing check on the head.tests / Core (julia 1, ubuntu-latest, x86)b42958f) run: https://github.com/SciML/NeuralOperators.jl/actions/runs/35161522453/job/105022031954 — Both fail inReactant.Accelerators.CPU.__init__(CPU.jl:50) withTypeError: in keyword argument priority, expected Int64, got a value of type Int32during package__init__/precompile, before any NeuralOperators test code runs. 32-bit-Julia dependency issue; this PR touches onlydocs/**(Burgers tutorials) andsrc/layers.jl(GridEmbedding) and does not touch the test environment.gh pr checks 174— all greenNotes:
typoson changed files clean;Runic --check src/layers.jlclean (matches the green CI lanes). The Core suite was not re-run locally: the only red lane is 32-bit-only and fails identically on base, and all x86_64 Core lanes already pass in CI on this exact head commit.🤖 Posted by an AI agent — harness: opencode 1.18.31 · model: opencode/muse-spark-1.3-contributor-free
Conversation: /home/crackauc/sandbox/fleet-master-jobs/nw-neuraloperators-174/log.txt
Risk assessment
GridEmbeddingand FNO models using positional grids, particularly CUDA/custom arrays; Burgers tutorial downloads. No new dependency internals, weakened tests, or apparent SemVer-breaking API changes.a534b3chas 19 passing checks and one failure:tests / Core (julia 1, ubuntu-latest, x86) / Tests - Core (Julia 1). Head and mainb42958f, plus recentmaincommitfed7f48, fail identically inReactant.Accelerators.CPU.__init__:priorityexpectsInt64, receivesInt32. Pre-existing; no observed CI regression. Documentation, GPU, QA, remaining Core, formatting, spelling, downgrade, and benchmarks pass. No review comments exist.GridEmbeddingcompatibility, an undisclosed coverage gap).🤖 Risk assessment posted by an AI agent (fleet master) — harness: Codex CLI 0.156.1 (Mac) / gpt-6-astra; dispatched by Claude Code head, model claude-opus-5-5[1m]
Conversation: local Claude Code session 3cd6500a-1f81-46b5-ac0b-c466e15b6a53 on Chris's Mac (session ID, no URL)