Release/0.4.0 - #2
Merged
Merged
Conversation
…ontract
The package half of the 0.4.0 roadmap (milestones M1-M3).
Performance. Grouping no longer enumerates sample pairs: composite, pairwise
and component derivation run igraph's connected components over the bipartite
sample-target graph. Direct assignments, sample maps, rule-based composites,
validation issues and the JSON writers are vectorised. With the default `via`,
composite derivation of 500 samples went from ~70 s (quadratic) to well under
a second, and 20 000 samples now take ~0.2 s. Every step is linear in nodes
plus edges. `inst/bench/pipeline.R` reproduces the numbers and
`test-performance.R` fails if any step regresses.
Breaking. Removes the 0.1-era aliases deprecated in 0.2.0 (`new_depgraph*`,
`build_depgraph`, `validate_depgraph`) and `validate_graph(checks=)`; removes
the unused `dependency_constraint()` class; composite and pairwise constraints
no longer carry `metadata$projection_edges` (the quadratic artefact), which
remains available from `detect_dependency_components()`.
Conditions. Every error now inherits from `splitgraph_error` and carries a
machine-readable `code`, with schema / reference / ambiguity / validation / io
subclasses. Invalid enumerated arguments are classed too, with partial
matching preserved. See `?splitgraph_conditions`.
Contract (schema 0.3.0, shipped under `inst/schema/0.3.0/`). Adds the
`stratum` annotation and `stratum_var`; makes the pairwise relations
composable inside `mode = "composite"`; records thresholds and column
provenance in `metadata$edge_sources` and in the spec; adds the
`subject_cross_site_overlap` leakage rule; and adds opt-in validation on read.
Empty JSON objects are written as `{}` rather than `[]`, and the validators
now check that, so nothing the package writes fails its own validator.
New API. `subset_graph()`, `combine_graphs()`, `add_edges()`, `export_graph()`
(GraphML / GML / CSV), `plot(focus=)`, a `SummarizedExperiment` method for
`graph_from_metadata()`, and a kinship-matrix input for
`relatedness_edges_from_kinship()`.
Derivation, validation and JSON output are unchanged from 0.3.0 on regression
cohorts apart from these intended additions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…sumer seam Three new vignettes: a ten-minute quick start with a decision table for choosing a constraint mode; FAQ and design notes (why not call the downstream planner directly, when composite grouping over-merges, how thresholds interact with transitive closure, the schema versioning policy); and a case study on a real public cohort, GEO series GSE60424, whose sample metadata is cached under `inst/extdata` with its derivation recorded. The adapter cookbook's rsample adapters now execute when rsample is installed rather than being shown as dead code. Corrects what the documentation claims about the downstream consumer. Verified by installing bioLeak 0.3.8 and running every mode against it: it accepts subject, batch, study and time. It errors on site, region, platform, assay, relatedness and spatial, which are absent from its mode map, and also on composite, which is mapped but reaches a planner mode whose required arguments the adapter never supplies. Earlier text claimed composite was the workaround; it is not. The working interim path is to join `group_id` onto the observation frame and call the planner directly, which the README, `?as_split_spec` and the contract tests now document and pin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tooling CI never set NOT_CRAN, so the Python cross-language conformance test had never actually run in either workflow; both now set it and install Python. A Linux-only perf-budget workflow runs the wall-clock guard, and a pkgdown workflow builds and deploys the site (required before submission, since DESCRIPTION advertises the site URL). Adds `dev/coverage.R` (covr with a 90 % gate; currently 91 %) and `dev/release.md`, a release checklist that records the traps found while running it: which check output is environmental, that the site URL 404s until the workflow runs once, and that a tarball built on a Windows checkout carries CRLF unless it is built in CI. The committed `.lintr` is calibrated so `lint_package()` reports nothing: the eight lines over 160 characters were rewrapped and compound statements split, rather than leaving a linter that emits a hundred findings and is ignored. Removes the stale `MD5` file and the `.DS_Store` files from the tree, and ignores them along with `Rplots.pdf` and the generated site. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Bumps the version, dates the NEWS section, and declares rsample and SummarizedExperiment in Suggests so no test is skipped for a missing optional dependency. Local release checks, with every Suggests installed: 211 tests, 0 failures, 0 skips; `R CMD check --as-cran` with vignettes rebuilt reports 1 warning and 3 notes, all environmental (qpdf absent, clock unverifiable, pandoc not on the check subprocess PATH) except the DESCRIPTION site URL, which 404s until the pkgdown workflow runs once; `lint_package()` 0 findings; coverage 91 %; `urlchecker` clean apart from that same URL; and bioLeak 0.3.8's own 281 tests pass against this tree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Checking a tarball built from a clean clone reported `checking examples ...
NOTE`: `graph_from_metadata` took 5.03 s against CRAN's 5 s budget, because
its example attached SummarizedExperiment to demonstrate the method added in
this release. Attaching Bioconductor dominated the timing; the example itself
is 0.14 s without it.
The demonstration moves to `vignette("faq-design-notes")`, where the limit
does not apply, and the reference page points at it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The three new vignettes covered the new material, but the older four still described the 0.2.0 world. Audited every 0.4.0 change against every vignette and closed the gaps. Corrected: the Python reader's `schema_version` was shown as 0.2.0; the shipped schemas were cited at `inst/schema/` rather than the versioned path they moved to; the JSON round-trip did not mention that the reader can validate. Added where each belongs. The cross-language vignette now carries an outcome through to Python and shows `StratifiedGroupKFold` driven from the spec's stratum annotation, which is the reader's new capability. The workflow vignette shows the stratum column and the declared roles an adapter keys on, the two new plot views, reshaping a graph with `subset_graph()` / `combine_graphs()` / `add_edges()`, and `export_graph()`. The structure vignette gains the `subject_cross_site_overlap` rule and a worked example of combining a pairwise relation with a direct one in a single composite, which is the feature its own subject matter most needs. The README gains an index of all seven vignettes. Every claim in the new passages was checked by running it, which caught one overreach: a composite spec does not carry a scalar `threshold`, because several relations with different cuts may have contributed. The graph's `edge_sources` is the complete record there, and the text now says so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Auditing the quick start against the package turned up four problems. Two were in the text. The canonical-column table omitted `featureset_id`, one of the nine relations `graph_from_metadata()` auto-detects, and it implied every listed column drives a constraint mode, which `featureset_id` does not. The validation paragraph called all leakage findings advisories; a subject spanning two studies or two sites is a warning. One was pedagogical. The lead example grouped by subject *and* batch on a cohort where that collapses eight samples into two groups, so the vignette's first result was a degenerate split it never explained. It now derives by subject first, then shows the composite as a contrast that yields three groups rather than four, and says why: one batch spans two subjects. The cohort was adjusted so each subject sits at a single site, which removes an unexplained cross-site warning from the output. One was a correctness bug in the text: the bioLeak snippet was shown against a composite spec, which the released bioLeak rejects. The spec it now hands over is subject-mode, which bioLeak accepts, and the limitation and its workaround are stated. Verifying the "only `sample_id` is required" claim then found a real bug: such a table failed in the edge binder with `objects must be a non-empty list`, although the documentation promises it and the equivalent explicit construction works. `graph_from_metadata()` now emits an edgeless graph, which validates, derives, serialises and round-trips; a regression test covers it. The vignette also gains the leakage-risk summary, a validating read-back, and pointers to all six other vignettes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Audited the vignette against the installed 0.4.0 package by knitting it and checking every prose claim against the output it actually produces. One claim was wrong: the summary section said the strict composite constraint leaves `missing_time_ordering` unsevered. That rule never fires on this graph (it has timepoints); the two FALSE rows are `per_dataset_featureset` and `shared_featureset_provenance`. Corrected, and the reason those two cannot be severed by any grouping is now stated, along with what the NA rows mean. Twelve exports were never mentioned. Added where each belongs rather than in a list: `as_igraph()` where the text already claims an igraph representation exists; `query_node_type()` / `query_edge_type()` as inventory queries, and `query_paths()` showing the three separate routes -- batch, assay, feature set -- by which S1 and S6 are entangled, which is the thing a single shared column cannot show; `write_dependency_graph()` / `read_dependency_graph()` / `validate_graph_json()` and the two `migrate_*_json()` functions in the persistence section, which previously covered only the spec despite the prose promising the graph too. `combine_graphs()` and `add_edges()` were prose-only in the reshaping section and are now worked examples, the second attaching thresholded kinship to a graph that was already built. The leakage validation rules are now listed in full with their default severities. They are documented nowhere else -- `?validate_graph` describes the arguments and `?depgraph_validation_report` is the constructor -- so the vignette was already the only place to learn them, without saying so. The site section gained the `subject_cross_site_overlap` finding, which the fixture it already builds happens to trigger. Checking the graph round-trip turned up one asymmetry worth stating rather than papering over: a node attribute that is `NA` in R comes back as `NULL`, because that is what JSON null means. Nothing derived from the graph changes, and the vignette now says exactly that instead of asserting object identity. Also corrected `inst/schema/` to the versioned `inst/schema/<schema_version>/` in `?validate_graph_json` and the README, matching where the schemas moved. Verified: 892 tests, 0 failures, 0 skips; lintr 0; all seven vignettes build; `R CMD check --as-cran` on the built tarball from an ASCII path is 1 WARNING (no qpdf) and 4 NOTEs (the site URL that 404s until pkgdown deploys, plus clock, HTML tidy and MiKTeX noise) -- no package findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Audited by knitting the vignette and checking every claim against the output it
produces, then re-running each new passage on its own before writing about it.
The composite example was the real problem. `via = c("Subject", "Site",
"Platform")` collapsed all six samples into a single group -- there was nothing
left to hold out -- and the text presented that as a plain feature demo. Worse,
the page had already shown the reason two chunks earlier without connecting
them: `subject_cross_site_overlap` fires because subject P3 has one sample at
each site, and in a closure that subject is the bridge that fuses NYC and BOS.
So the section now runs a composite that works (`Site` + `Platform`, two
groups), then adds subject to make it collapse, names the validation warning
that predicted it, and says plainly that splitGraph does not flag a one-group
result -- confirmed by checking the constraint warnings, validate_split_spec()
and summarize_leakage_risks(), none of which complain; the risk summary even
reports `split_spec_ready`. Until something does, `metadata$n_groups` is the
check to run, and the two escapes are both already on the page.
The opposite failure had the same silence in the vignette although the package
does detect it: too strict a kinship cut leaves samples alone in their own
groups, and the pairwise deriver records that in `metadata$warnings`. Shown now,
and paired with the collapse as two ends of one trade-off. (The threshold is on
the fraction of *samples* sitting alone, not of groups -- checked against a
3-chain-plus-two-singletons case that separates the two readings.)
Filled in what 0.4.0 added or never showed. `relatedness_edges_from_kinship()`
accepts a square GRM, which is what PLINK `--make-rel square` writes, and takes
`id1` / `id2` / `kinship` for KING's and GCTA's own headers -- both were
pair-table-only here. The metric that passed the threshold is kept on each edge,
so the surviving pairs and their kinship values are now readable rather than
implicit. `region` was omitted "to keep the output short"; all four cluster
modes now appear side by side in one table, which is shorter than the three
separate calls it replaces.
Two traps a real table walks into. `spatial_edges_from_coords()` treats every
numeric column as a coordinate unless `coord_cols` says otherwise, so a QC
column in the same frame drops the edge count from five to zero -- shown, since
it fails silently. And `via` means different things in different places: the
derivation accepts pairwise modes, while `detect_dependency_components()` and
`plot(focus = "sample_projection")` take node types and reject "relatedness".
Both are documented in the man pages but the contrast is easy to trip over right
after reading that pairwise relations are not a separate world.
The mixed-relation graph is now built with `graph_from_metadata()` plus
`add_edges()` rather than hand-assembled node and edge sets, which is the 0.4.0
route and eleven lines shorter; the threshold provenance still survives it.
Verified: 892 tests, 0 failures, 0 skips; lintr 0; all seven vignettes build;
`R CMD check --as-cran` on the tarball from an ASCII path is 1 WARNING (no qpdf)
and 4 NOTEs (the site URL that 404s until pkgdown deploys, plus clock, HTML tidy
and MiKTeX noise) -- no package findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every Python block in this vignette is displayed, never executed, so its
asserted outputs had never been checked by anything. I ran them: loaded the spec
this vignette writes with the shipped reader, asserted each documented value
(`schema_version`, `constraint_mode`, `recommended_resampling`, `stratum_var`,
`strata()`, the `grouping()` dict), and executed all four snippets verbatim
against scikit-learn 1.8. Those were all correct, including that the explicit
StratifiedGroupKFold form and the `stratified_group_kfold()` one-liner yield
identical folds.
One prose claim was not. `strata()` was documented as raising when a sample
lacks a stratum; it does not -- it returns `None` per sample, and it is
`stratified_group_kfold()` that raises `ValueError` rather than hand
scikit-learn a `y` full of `None`. Corrected, and the distinction is now the
point of the paragraph.
The bigger problem was the verification section itself. Its guard was
`nzchar(Sys.which("python3"))`, and on Windows that finds the Store launcher
stub -- a real path to an executable that prints "Python not found" and exits
9009. So the chunk ran, `system2` failed, the `if (status == 0)` body was
skipped, and the vignette rendered a code block with no output at all while the
introduction promised it "does run the shipped Python reader". The package's own
`test-python-conformance.R` had already solved this, probing `python3` then
`python` by executing each; the vignette now uses the same rule, reports the
skip in words when nothing usable is found, and checks `order_rank` and
`stratum` alongside the grouping instead of only claiming they are checked. On
this machine it now prints three TRUEs where it previously printed nothing.
Filled in what an interchange-format vignette should not have been missing. The
file itself is now shown -- top-level keys and one sample row -- since that
artifact, not any object, is the contract; the point that inapplicable columns
are `null` rather than absent, so every row has the same shape, is visible
there. The reader's full surface is tabulated rather than the five attributes
that happened to appear in one snippet, and `group_kfold()` joins the stratified
one-liner it was inconsistently omitted next to. Schema versioning gets its own
section: the reader accepts any major 0 and refuses a future major with a
`ValueError` (checked by feeding it a 1.0.0 file), older minors load with the
new fields simply absent (checked with a 0.2.0 file), and
`migrate_split_spec_json()` is the other direction -- confirmed to bump 0.2.0 to
0.3.0, fill `stratum` with null, and still validate. Concrete `pip install`
lines replace "install/copy the package".
Also aligned `splitspec.__version__`, which said 0.2.0 while `pyproject.toml`
already said 0.3.0.
Verified: 892 tests, 0 failures, 0 skips -- the Python conformance test runs
here rather than skipping; lintr 0; all seven vignettes build; `R CMD check
--as-cran` on the tarball from an ASCII path is 1 WARNING (no qpdf) and 4 NOTEs
(the site URL that 404s until pkgdown deploys, plus clock, HTML tidy and MiKTeX
noise) -- no package findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Checked by knitting the vignette and then running every claim it makes about
bioLeak, rsample and the spec's own fields against the installed packages.
The bioLeak sentence was stale. It promised `as_leaksplits(spec, data, outcome)`
without qualification, but 0.3.8 errors on composite ("'primary_axis' must be a
list with 'type' and 'col' elements") and on all four cluster modes ("subscript
out of bounds") -- I ran each to confirm, and confirmed that
`make_split_plan(group = "group_id")` still produces the split. README,
quick-start, `?as_split_spec`, NEWS and paper.md already carried this caveat;
this vignette was the one place that did not. Now it does.
The stratum annotation was absent from a document whose whole subject is what a
consumer must honor -- it appeared nowhere in 348 lines, although `stratum_var`
is part of the 0.4.0 spec. It now has a section, and the section says the thing
that actually bites: the annotation is per sample while the split unit is the
group, and the two need not agree. In this cohort no subject has a single
stratum, so `rsample::group_vfold_cv(strata = "stratum")` refuses the spec
outright -- shown live, with rsample's own message, rather than described.
Two corrections to Adapter 3. The comment claimed each analysis set "precedes"
its assessment set, which is stronger than the `<=` beside it proves and is
false here: `order_rank` ranks timepoints, not rows, so six samples carry two
distinct ranks and slices 2 and 3 hold a rank-2 sample on both sides. That is
worth a reader's attention, not a quiet overstatement, so the tie is now
demonstrated and the remedy named. Separately, `rolling_origin()` is
[Superseded] in rsample; it does not warn, and this vignette suppresses warnings
anyway, so nothing would have surfaced it. The superseded status is now stated
and the `sliding_window()` equivalent given -- verified to produce identical
assessment sets.
The dispatcher section claimed only that the adapters "cover different shapes".
They cover every shape: `recommended_resampling` takes exactly five values, and
all eleven modes are now shown mapping onto them, so a five-branch dispatcher
cannot be surprised by a spec. Added the mapping table and the check that
produces it.
Also seeded the toy observation frame, which was calling `rnorm()` and
`rbinom()` unseeded in a cookbook.
Verified: 892 tests, 0 failures, 0 skips (the bioLeak contract test and the
Python conformance test both run here); lintr 0; all seven vignettes build;
`R CMD check --as-cran` on the tarball from an ASCII path is 1 WARNING (no qpdf)
and 4 NOTEs (the site URL that 404s until pkgdown deploys, plus clock, HTML tidy
and MiKTeX noise) -- no package findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A case study on real data earns its keep only if the numbers are right, so I recomputed every factual claim against the shipped CSV. Three were wrong. "Every donor contributed several samples (one per cell population)" -- the range printed right below it is 6 to 7. NK cells were sorted for only 14 of the 20 donors, so six donors have six samples, not seven. The comment now says so and the distribution is printed alongside the range. "No cross-study or cross-site overlap, because each donor belongs to exactly one batch and one condition" -- the reason is unrelated to the result. Those two rules cannot fire because the graph has no `Study` or `Site` nodes at all; the CSV carries neither column. The passage now separates that from the reason `heavy_batch_reuse` stays quiet (no batch holds half of 134; each date covers one donor), and adds the question a reader of this cohort will actually have: why a donor spanning six or seven cell populations produces no finding. It is deliberate -- there is no cross-region rule, because one person contributing several sorted populations is the experiment's design, not an anomaly. "Every donor spans every cell population, so a strict composite ... chains the whole cohort" -- the premise is false, and the conclusion is true for a slightly different reason: six of the seven populations were sorted for all twenty donors, which is enough to chain everything. Stated correctly now. That composite prints `n_groups` 1 and `warnings` `character(0)`, and the empty vector went by without comment -- inviting the reading that nothing was wrong. As in the structure vignette, the text now says plainly that splitGraph does not flag a one-group result and that `metadata$n_groups` is the check to run. (This is the open finding recorded against the package, seen here on real data.) Two additions. `by_region` was derived and counted but never interpreted, leaving the reader with 7 groups and no sense of what they mean: grouping by cell population cuts across donors rather than between them, putting all 20 donors in more than one group -- precisely the leak the study is exposed to, and the concrete reason region rides on the spec as a blocking annotation instead. And the graph is never shown; at 186 nodes it cannot be drawn whole, which makes it the right place to demonstrate `focus = "ego"` on one donor. Also `comps$table$component_size` reached past the accessor into the object's internals; `as.data.frame()` is the documented one and is what the other vignettes use. And `validate_split_spec_json(out)` printed a tempfile path into the rendered vignette; `$valid` is what the other vignettes print. Verified: every number above recomputed from `inst/extdata/GSE60424_samples.csv` (134 samples, 20 donors, 7 cell types, 5 conditions, 6 of 7 populations complete, batch sizes 6-7 against a `heavy_batch_reuse` threshold of 67); bioLeak 0.3.8 accepts this subject-mode spec, so that pointer stands. 892 tests, 0 failures, 0 skips; lintr 0; all seven vignettes build; `R CMD check --as-cran` from an ASCII path is 1 WARNING (no qpdf) and 4 NOTEs (the site URL that 404s until pkgdown deploys, plus clock, HTML tidy and MiKTeX noise) -- no package findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`__pycache__` was shipping to CRAN. The rule meant to stop it was `^inst/python/__pycache__$`, but the cache directory sits one level deeper, at `inst/python/splitspec/__pycache__`, so the anchored pattern never matched and two `.pyc` files rode along in every tarball. Matching `__pycache__` at any depth (plus `\.pyc$`) removes them; the tarball goes from 128 entries to 125 and every legitimate Python file still ships. `.Rbuildignore` also had `^\.Rproj\.user$` alongside the correct `^\.Rproj\.user$` -- the doubled backslashes make it a pattern for a path with a literal backslash in it, so it matched nothing. Dropped the dead line and grouped the rest under headings. Added defensively, none of it currently present but all of it produced by the release checklist in dev/release.md: `revdep/`, `cran-comments.md`, `*.Rcheck`, `.Renviron`, `.Rhistory` at any depth, editor directories, and the `vignettes/figure/`, `*_cache/`, `*_files/`, `*.html` leftovers from knitting a vignette in place rather than through `R CMD build`. On the git side the gap was the opposite one: `.Rbuildignore` knew about the built tarball and `splitGraph.Rcheck/`, but `.gitignore` did not, so a careless `git add -A` after a release build would have committed them. Both are ignored now, along with the same knitr and Python leftovers. Also stopped tracking `.claude/settings.local.json`. It is Claude Code's per-machine permission allowlist -- 8.7 KB of paths and commands specific to this checkout -- and was committed. The file stays on disk; only the index entry goes. A shared `.claude/settings.json` would still be committable. Verified: rebuilt the tarball and confirmed no `__pycache__` or `.pyc` entries remain while all seven vignettes, the schema files, `inst/extdata` and the Python reader are still present; `R CMD check --as-cran` on it from an ASCII path reports the same 1 WARNING (no qpdf) and 4 NOTEs as before, with tests, vignette re-building and the top-level/hidden-file checks all OK. Also confirmed by `git check-ignore` that a tarball, an `.Rcheck` directory, a stray `.pyc`, `vignettes/figure/` and the local settings file are now all ignored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The existing paper predated JOSS's current requirements and had four of the six mandatory sections missing. JOSS now requires Summary, Statement of need, State of the field, Software design, Research impact statement and AI usage disclosure; the paper had only the first two. Rewritten to that structure, 1466 words against the 750-1750 limit. State of the field previously named only rsample and scikit-learn, which understated the competition: mlr3spatiotempcv, blockCV and CAST all handle dependence-aware resampling and are, within spatial and temporal autocorrelation, more capable than anything here. The section now says so and argues the distinction that actually holds -- those packages produce resamplings tied to one dependence family and one framework, not a validated portable description of the structure itself -- and gives the build-vs-contribute justification JOSS asks for. Software design and Research impact are new. The impact section rests on what can be evidenced rather than asserted: CRAN availability since July 2026, the bioLeak seam pinned by a contract test, the Python consumer pinned by a conformance test, and the GSE60424 case study. The AI usage disclosure is new and material. Generative AI assisted substantially in this package's development and in drafting this manuscript, and the section says so plainly rather than minimising it, alongside the verification regime that backs the result. Every factual claim was measured rather than recalled: 11 node types and 17 relations read from the schema; the 20,000-sample timing re-run (5.93 s for build, validate, composite derivation and write -- "about six seconds"); 892 test expectations; 91.4% coverage re-measured with covr; seven vignettes; the CRAN publication date and the five-configuration CI matrix read from the workflow. Two drafts of my own overclaimed and were corrected before commit: lintr does not run in CI (it is local), and the equivalence check against the previous release was a pre-release verification, not a shipped test. Every reference was verified against its publisher of record, including the GSE60424 source publication, whose citation I had initially guessed. The bibliography gains Kapoor & Narayanan, Whalen et al., mlr3spatiotempcv, blockCV, CAST, igraph, the GEO source paper and the package's own CRAN DOI, which resolves. The paper compiles under pandoc with citeproc: 12 references resolved, zero unresolved citations. The author's ORCID was checked against the ORCID public API. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Compared the paper against joss.11126 (segregation, published 16 September
2026), which is recent enough to reflect the criteria as they are actually
applied. That paper carries ten sections, not the six that are mandatory: it
adds Core functionality and Example workflow, and it has a figure and code
blocks. Ours had eight sections, no figure and no code.
Added Core functionality (validation layers, the eleven constraint modes split
into direct / pairwise / composite, and what the handoff attaches) and Example
workflow (the five-call R pipeline, then the Python side reading the same JSON).
Both code blocks were executed before being written down: the R block runs
verbatim against a real metadata frame and writes the spec; the Python block was
run against that file with scikit-learn 1.8 and yields five folds.
Added a data-flow figure, generated by dev/paper-figure.R with base graphics
only so it is reproducible from a bare R install. I rendered and looked at it
three times: the first version clipped the function labels behind the boxes, the
second ran off the right edge, and both put the "splitGraph stops here" boundary
before the JSON artifact, which wrongly implied the JSON belonged to the
consumer rather than to us. The committed version fixes all three.
Trimmed Statement of need and Research impact to absorb the new sections: 1677
words against the 1750 ceiling, with the exemplar at 1622 for calibration.
The figure lives in paper-figures/ and is build-ignored alongside paper.md and
paper.bib, so it does not add weight to the CRAN tarball -- confirmed by
building and grepping the tarball. The paper compiles under pandoc with
citeproc: 12 references resolved, figure embedded, no warnings.
One drafting slip caught before commit: a shell heredoc ate the backslash in
\autoref, leaving "utoref{fig:pipeline}" in the text. Fixed, and the label now
sits inside the caption as JOSS documents.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Last of the seven vignettes. Verified every claim it makes, including the ones about other packages, by running them rather than reading them. The performance answer overstated the guard rail. It said `test-performance.R` "fails if any step regresses to superlinear behaviour". It does not: that file is a wall-clock budget at 5,000 samples whose own header calls the limits "deliberately generous", and only the composite derivation gets a scaling check -- 1,000 against 4,000 samples, failing above an 8x ratio, and even that is skipped when the small run is too fast to time. Any other step could degrade without tripping it. Said plainly now. "In seconds on a laptop" was also doing a lot of work. Re-ran the benchmark twice at 20,000 samples: 12.37 s and 12.13 s for the six steps the sentence enumerates. That is "about twelve seconds", and more than half of it is writing the *graph* JSON (~6 s) -- which a reader who only wants the handoff artifact does not need to pay. Both facts are now in the text. The schema-versioning section calls itself the policy "in one place" but described only R's behaviour. Checked what actually happens: a differing major produces a warning naming both versions and suggesting `migrate_split_spec_json()`, then loads; an older minor loads with the missing column present and NA; an unknown column is ignored. All three as documented. The section now demonstrates the first of those with a runnable chunk showing the real warning text, because a permissive boundary is exactly the kind of claim worth showing. It also now records the asymmetry it was silent about: the shipped Python reader is stricter and raises `ValueError` on a differing major rather than warning. Confirmed against both implementations, including that each still accepts every schema sharing major 0. Claims verified and left unchanged: bioLeak's `combined` mode does pick a primary axis and drop training rows sharing a secondary level with the test set (confirmed in `fit_resample.R` and in bioLeak's own `secondary_axis` docs); `stratify` is a real `make_split_plan()` argument; the over-merge example collapses to one component; rule-based grouping yields the three subject groups; the kinship thresholds chain as annotated; `query_neighbors()` on an unknown id raises `splitgraph_reference_error` with code `unknown_node_ids`; the SummarizedExperiment path matches the data-frame path exactly. Verified: 892 tests, 0 failures, 0 skips; lintr 0; all seven vignettes rebuild under `R CMD check` (30 s); check from an ASCII path is unchanged at 1 WARNING (no qpdf) and 4 NOTEs, the only package-attributable one still being the pkgdown URL that 404s until the site is deployed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`https://selcukorkmaz.github.io/splitGraph/` 404s because the site has never been deployed, and it was the one package-attributable finding in `R CMD check --as-cran`: CRAN's incoming-feasibility check reports invalid URLs and would bounce the submission. DESCRIPTION's `URL` now lists only the GitHub repository, and `man/splitGraph-package.Rd` was regenerated to match -- the URL appeared in both, and roxygen derives the Rd from DESCRIPTION. `_pkgdown.yml` keeps its own `url:`, so the site still builds correctly with the right canonical links; nothing about the pkgdown workflow changes. `dev/release.md` described this as a pre-submission choice between deploying the site and dropping the URL. Updated to record which was chosen, and to say how to reverse it: once Pages is live, put the URL back in DESCRIPTION, re-roxygenise, and confirm it resolves. Verified: `urlchecker::url_check(".")` now reports ALL URLs OK across all nine it finds, where it previously failed on this one in two places. This commit also carries the uncommitted `Config/Needs/website: selcukorkmaz/leakdown` line that was already in the working tree from outside this session; git stages DESCRIPTION whole. It is a `Config/*` field, which CRAN ignores. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reports what was actually run rather than the usual boilerplate. Only one environment has been checked: local Windows 11, R 4.5.1. The win-builder, macOS-builder and R-hub lines are TODO markers, because those runs have not happened -- and neither has CI: the `release/0.4.0` branch is pushed, but `R-CMD-check.yaml` triggers only on `main`/`master`, so GitHub Actions has never seen this code. The file says so instead of implying multi-platform coverage. The local result is given as it is, 0 errors / 1 warning / 3 notes, with each of the four quoted verbatim from the check log and attributed to missing local tooling (qpdf, a reachable time server, HTML Tidy, MiKTeX's leftover file). Also recorded: incoming feasibility now passes and urlchecker finds all nine URLs valid, but no spell-checker was installed here, so the aspell check on DESCRIPTION is genuinely unverified. The reverse-dependency section matters because this release is breaking. CRAN lists exactly one revdep, bioLeak, and only under Suggests -- confirmed on CRAN's own page, not just locally. The three specific reasons it is unaffected were each checked against bioLeak 0.3.8's source. Two things flagged for the maintainers that a reviewer would otherwise have to work out: SummarizedExperiment is a Bioconductor Suggests used by one guarded S3 method, with no example touching it; and inst/python ships as data, with its only test skipped on CRAN. cran-comments.md is already build-ignored, confirmed by building the tarball and grepping it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The file said the aspell check had not been exercised locally. It has now: installed `spelling` and ran `spell_check_package()`. Exactly one word in DESCRIPTION is flagged, `inspectable`, which is correctly spelled and already sits in the Description field of 0.3.0 on CRAN — so it is a NOTE a reviewer has accepted before, and cran-comments now says so instead of leaving it open. (The other 88 words the checker flags are outside DESCRIPTION — technical names like igraph, GraphML, PLINK, JOSS, and British spellings in the vignettes — none of which CRAN's incoming check looks at.) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
JOSS's AI usage policy names three required elements. The disclosure had one and a half of them: it said Claude was used and that the author reviewed the output, but gave no model version, did not separate where the assistance applied, and did not assert who made the core design decisions -- which the policy requires explicitly. Restructured into the four labelled parts the policy asks for: tools and versions, where used, nature and scope, and confirmation of review including the design decisions that were the author's. Also added a Conflicts of interest and funding section, which JOSS's policies require and the paper lacked entirely. It discloses the relationship a reviewer would otherwise have to infer: the author maintains bioLeak, the package this paper presents as the reference consumer of split_spec. The funding line states no external support. Trimmed elsewhere to stay inside the limit, since the new material is mandatory: 1749 words against the 1750 ceiling. Compiles clean under pandoc with citeproc, 12 references resolved, eleven sections. Note for the author: the AI disclosure now names Claude Opus 5 "and earlier Claude models during 2026" and asserts specific design decisions as the author's. Both statements need the author's confirmation before submission -- only he knows the full tool history, and the policy makes an incomplete disclosure an ethical matter rather than a formatting one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Checked the paper against JOSS's official example. The structure, figure syntax (`\label` in the caption, `\autoref` in the text) and citation syntax already matched it; the example's Mathematics / Citations / Figures sections are syntax demonstrations rather than real sections, so they are correctly absent here. Two optional metadata fields the example shows were missing. Added `ror: 00xa0xn82` for Trakya University, looked up in the ROR registry and confirmed against two sources rather than guessed, and `corresponding: true` for the sole author. The affiliation fields now follow the example's order. Also verified the point the example makes about references: every venue in paper.bib is spelled out in full -- ACM Transactions on Knowledge Discovery from Data, Journal of Statistical Software, Methods in Ecology and Evolution and the rest -- with no discipline-specific abbreviations. Still 1749 words. Pandoc renders the title, author, all 12 references and the embedded figure, with no unresolved citations. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ADME links CRAN's incoming pre-test rejected 0.4.0 with 1 ERROR and 1 NOTE on both r-devel-linux-x86_64-debian-gcc and r-devel-windows-x86_64. The ERROR was in `export_graph(format = "gml")`. It called `igraph::write_graph()` without an `id` argument; on the igraph present on those machines that NULL default reaches the C layer as a zero-length vector and is rejected with "Size of id vector must match vertex count" (io/gml.c:1057). The call now supplies `seq_len(vcount(g))`, which is what the writer would have generated anyway and which satisfies the C-level length check in every version. Worth recording why this was not caught here: it does not reproduce on igraph 2.2.1, where the same call succeeds. The traceback CRAN returned is the tell -- it shows `write_graph_gml_impl(..., options = "default", ...)`, a signature that does not exist in 2.2.1, so the check machines are on a newer igraph whose generated binding no longer handles the NULL. Confirmed against igraph's current source, where `write.graph.gml()` passes `id` straight through untouched. The test now also asserts one `id` line per node in the written file, so the ids going missing again cannot pass silently. The NOTE was two dangling file URIs: README.md linked to `.github/CONTRIBUTING.md` and `CODE_OF_CONDUCT.md` by relative path, but both are build-ignored and so are absent from the installed package. Both are now absolute GitHub URLs. Checked the rest: the only relative link left in any shipped file is `man/figures/README-plot-1.png`, which does ship. 893 tests, 0 failures, 0 skips; lintr 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.