Skip to content

Fix Burgers docs tutorials: Hugging Face dataset mirror + Reactant-safe GridEmbedding - #174

Merged
ChrisRackauckas merged 3 commits into
SciML:mainfrom
ChrisRackauckas-Claude:docs/burgers-hf-mirror
Oct 5, 2026
Merged

ChrisRackauckas merged 3 commits into
SciML:mainfrom
ChrisRackauckas-Claude:docs/burgers-hf-mirror

Conversation

@ChrisRackauckas-Claude

@ChrisRackauckas-Claude ChrisRackauckas-Claude commented Sep 17, 2026 •

Copy link
Copy Markdown
Member

Please ignore until reviewed by @ChrisRackauckas.

What changed and why

The docs build on main is red (https://github.com/SciML/NeuralOperators.jl/actions/runs/35161521723). Two independent causes:

  1. Burgers dataset download (docs tutorials). Both Burgers tutorials fetched burgers_data_R10.mat from Google Drive via gdown. Drive rate-limits shared files, so the build died with Python: FileURLRetrievalError: ... Too many users have viewed or downloaded this file recently, and every later @example block on those pages cascaded into UndefVarError: x_data / x_data_dev / y_data_dev. The tutorials now fetch the same file from the kks32/sciml-dataset Hugging Face mirror over plain HTTPS with the pinned sha256 d1a04567..., pinned to an immutable mirror revision (982685ff) rather than main. This also removes the PythonCall/CondaPkg/gdown stack and docs/CondaPkg.toml from the docs environment. The DataDep is renamed Burgers -> BurgersR10 so machines with a stale Drive-zip cache refetch instead of reusing the old unzipped layout.
  2. GridEmbedding vs Reactant (src/layers.jl). fix: transfer GridEmbedding grid to input device (fixes CUDA scalar indexing #125) #132 placed the CPU-built grid via Lux.get_device(x)(grid), but get_device throws on TracedRArray inside Reactant.@compile, breaking the FNO tutorial training step. Allocate like x and broadcast-assign instead: no runtime device query, grid still lands on x's device, preserving the FNO invokes scalar indexing #125 CUDA fix.

Verification (local, Julia 1.12.4)

Full docs build, julia --project=docs docs/make.jl from the repo root (after julia --project=docs -e 'using Pkg; Pkg.develop(PackageSpec(path=pwd())); Pkg.instantiate()'):

[ Info: SetupBuildDirectory: setting up build directory.
[ Info: Doctest: running doctests.
[ Info: ExpandTemplates: expanding markdown templates.
[ Info: CrossReferences: building cross-references.
[ Info: CheckDocument: running document checks.
[ Info: Populate: populating indices.
[ Info: RenderDocument: rendering document.
[ Info: HTMLWriter: rendering HTML pages.
...
[ Info: Automatic `version="0.7.3"` for inventory from ../Project.toml
┌ Warning: Documenter could not auto-detect the building environment. Skipping deployment.

Zero failed @example blocks and no makedocs encountered an error [:example_block]; both tutorial pages rendered (docs/build/tutorials/burgers_deeponet-*.png 258827 B, burgers_fno-*.png 247647 B). The output is byte-identical to the pre-pinning build, so the pinned-revision commit changes no behavior.

Pinned mirror fetched through DataDeps' own fetch+checksum pipeline into a fresh DataDep (BurgersR10Pinned, no cache): PATH=.../datadeps/BurgersR10Pinned/burgers_data_R10.mat, SIZE=644427710, sha256 d1a0456776255a4bd841dbc18951d3f468266d945d96e24ae531a12f18bb5a1a (matches the pin).

Fail-before evidence on the same machine/paths: full build failed at burgers_deeponet.md:5-50 with FileURLRetrievalError, cascading to UndefVarError: x_data / x_data_dev on the following blocks; and before the GridEmbedding commit the FNO training step failed inside Reactant.@compile with get_device "isn't meant to be called inside Reactant.@compile context".

Not verified

  • GPU/CUDA path for the new similar+broadcast grid placement (no GPU locally). The repo GPU CI lane passed on this branch and should confirm FNO invokes scalar indexing #125 stays fixed.
  • tests / Core (julia 1, ubuntu-latest, x86) is red on this PR. It fails in a dependency __init__ (TypeError: in keyword argument priority, expected Int64, got Int32) and this PR does not touch the test environment; I did not trace it to a main run from here.
  • The full local @example run reused the on-disk BurgersR10 dataset cache; the fresh-fetch path was verified separately (above) and CI's Documentation run on this branch (https://github.com/SciML/NeuralOperators.jl/actions/runs/35211262876) performed the fresh download.

🤖 Generated with OpenCode (model: muse-spark-1.3-contributor-free); no public session URL (local session at /home/crackauc/sandbox/tmp_20260917_040834_91074).

CI triage (2026-09-25)

Head commit a534b3c; base main @ b42958f. The branch is up to date with the base (git merge-base HEAD upstream/main == base HEAD), so no merge commit was needed; GitHub reports MERGEABLE. gh pr checks 174 shows exactly one failing check on the head.

check classification evidence
tests / Core (julia 1, ubuntu-latest, x86) master-red — identical failure on base PR head run: https://github.com/SciML/NeuralOperators.jl/actions/runs/35291773241/job/105437248640 — Base (b42958f) run: https://github.com/SciML/NeuralOperators.jl/actions/runs/35161522453/job/105022031954 — Both fail in Reactant.Accelerators.CPU.__init__ (CPU.jl:50) with TypeError: in keyword argument priority, expected Int64, got a value of type Int32 during package __init__/precompile, before any NeuralOperators test code runs. 32-bit-Julia dependency issue; this PR touches only docs/** (Burgers tutorials) and src/layers.jl (GridEmbedding) and does not touch the test environment.
all other checks (Documentation, Downgrade, Runic, Typos, benchmark, remaining Core/QA/GPU lanes) pass on head gh pr checks 174 — all green

Notes:


🤖 Posted by an AI agent — harness: opencode 1.18.31 · model: opencode/muse-spark-1.3-contributor-free
Conversation: /home/crackauc/sandbox/fleet-master-jobs/nw-neuraloperators-174/log.txt

Risk assessment

  • Risk: medium
  • Blast radius: GridEmbedding and FNO models using positional grids, particularly CUDA/custom arrays; Burgers tutorial downloads. No new dependency internals, weakened tests, or apparent SemVer-breaking API changes.
  • Evidence: Head a534b3c has 19 passing checks and one failure: tests / Core (julia 1, ubuntu-latest, x86) / Tests - Core (Julia 1). Head and main b42958f, plus recent main commit fed7f48, fail identically in Reactant.Accelerators.CPU.__init__: priority expects Int64, receives Int32. Pre-existing; no observed CI regression. Documentation, GPU, QA, remaining Core, formatting, spelling, downgrade, and benchmarks pass. No review comments exist.
  • Independent review: Codex CLI 0.156.1 (Mac) / gpt-6-astra rated it medium, high confidence: the CI failure is pre-existing, but the description overstates GPU validation because GPU FNO tests omit positional grids and existing grid checks use CPU arrays.
  • Merge: needs human review (no focused regression test was added for the allocation/broadcast change; passing GPU CI does not establish CUDA GridEmbedding compatibility, an undisclosed coverage gap).

🤖 Risk assessment posted by an AI agent (fleet master) — harness: Codex CLI 0.156.1 (Mac) / gpt-6-astra; dispatched by Claude Code head, model claude-opus-5-5[1m]
Conversation: local Claude Code session 3cd6500a-1f81-46b5-ac0b-c466e15b6a53 on Chris's Mac (session ID, no URL)

…e Drive

The Burgers tutorials downloaded burgers_data_R10.mat from Google Drive
via gdown. Drive rate-limits shared files ("Too many users have viewed
or downloaded this file recently"), which failed CI docs builds with
FileURLRetrievalError whenever the quota was exhausted.

Fetch the same dataset from the kks32/sciml-dataset Hugging Face mirror
over plain HTTPS with a pinned sha256 checksum instead. This also drops
the PythonCall/CondaPkg/gdown stack from the docs environment. The
DataDep is renamed to BurgersR10 so machines with a stale Drive-zip
cache refetch instead of reusing the old layout.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Agent-Harness: OpenCode
Agent-Model: muse-spark-1.3-contributor-free
Agent-Session: local session at /home/crackauc/sandbox/tmp_20260917_040834_91074 (no public conversation URL)
…acement

Placing the CPU-built grid via Lux.get_device(x) throws inside
Reactant.@compile (get_device errors on TracedRArray), which broke the
Burgers FNO docs tutorial training step. Allocate like x and
broadcast-assign instead: no runtime device query, and the grid still
lands on x's device, preserving the CUDA scalar-indexing fix from SciML#125.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Agent-Harness: OpenCode
Agent-Model: muse-spark-1.3-contributor-free
Agent-Session: local session at /home/crackauc/sandbox/tmp_20260917_040834_91074 (no public conversation URL)
The mirror URL used `resolve/main`, so any upstream change to the mirror's
main branch would force a full re-download before the pinned sha256 could
fail the build. Pin to the revision (982685ff) whose content the checksum
was verified against, so docs builds are reproducible.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Agent-Harness: OpenCode
Agent-Model: muse-spark-1.3-contributor-free
Agent-Session: local session at /home/crackauc/sandbox/tmp_20260917_040834_91074 (no public conversation URL)
@ChrisRackauckas
ChrisRackauckas marked this pull request as ready for review October 5, 2026 01:20
@ChrisRackauckas
ChrisRackauckas merged commit 8ea3f6b into SciML:main Oct 5, 2026
19 of 20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants