Curate CALLA2026_OSHV: a screened study backlog and the first raw-reanalysis registration - #4
Merged
Merged
Conversation
Ten candidate studies, verified against NCBI BioProject and the ENA read-run API rather than recalled: accessions, run counts, and group structures come from deposited metadata. Corrects the selection criterion I had been using. "Does the source publish full statistics" is the wrong primary filter — a significance-filtered table carrying only adjusted p-values is the normal format in this literature, not bad luck, so screening on it mostly reproduces the HESSER2024_VCOR dead end. The reliable route to poolable evidence is raw reads in SRA, because AREE's own DESeq2 run then produces lfcSE by construction regardless of what the authors published. Records one rejected candidate in detail. PRJNA623063 looked like the ideal second study from its title — C. gigas larvae, Vibrio challenge, same phenotype as the study already registered — but all 12 runs share one sample_alias across 12 timepoints. It is an unreplicated time course with no valid DE contrast, and that is invisible from the abstract. Also notes that hypoxia has no usable public dataset despite being in the phenotype ontology, and that GEO holds only 10 oyster expression series so a GEO-only search will miss the corpus. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ed metadata Second real study in AREE and the first in raw_reanalysis mode: PRJNA1329250, 42 paired-end NovaSeq libraries of Pacific oyster spat from two hatchery populations challenged with three OsHV-1 microvariant isolates. Published as Calla, Thompson & Burge 2026, Fish & Shellfish Immunology 171:111154. Registered, NOT harmonized. The design, sample sheet, FASTQ manifest, run config, and reference context are complete and verified; the pipeline run that would produce result tables has not been executed (226 GB of reads). The study contributes zero evidence records, and analysis_status says not_started. Why raw data. HESSER2024_VCOR showed that harmonizing a published table leaves a study unpoolable when the source reports no standard error, which is the normal format for supplementary DE tables. Reanalyzing raw reads makes AREE's own DESeq2 emit lfcSE by construction, so poolability stops depending on what the authors chose to publish. Adds `aree fetch-samplesheet`, which reads ENA's run report and BioSample attribute blocks and emits a design sheet, a FASTQ manifest carrying ENA's own MD5 per file, and checksummed provenance. The design here was derived that way rather than transcribed from the paper, which is paywalled. Reading the archive instead of the methods section is what catches an unreplicated design before anyone downloads hundreds of gigabytes — the failure mode already documented for PRJNA623063. The command warns when the smallest group has n<3, and `validate-study` now cross-checks a study's declared BioProject against the one its sample sheet came from. Curation judgements, all of which are tested: * Not registered as resilience evidence. The BioProject is titled "Evaluating Pacific oyster lineages for tolerance to Ostreid herpesvirus", but no survival, mortality, or viral-load measurement exists for these animals in the deposited metadata or the abstract, and the paper frames its results as groundwork for future tolerance breeding. Registered as disease_susceptibility, flagged ambiguous_phenotype_definition. This also corrects docs/candidate_studies.md, which had promised a resilience phenotype on the strength of that title; the correction is recorded on the page rather than quietly edited out. * Exposure dose, route, duration, and timing are absent from the deposited metadata and the paper is not open access, so they are null, not guessed, and a test asserts they stay null. Effect sizes are not dose-comparable until someone reads the methods. * The three viral isolates stay separate comparisons. The authors report no marked difference between the USA and Australian responses; merging arms on that basis is a meta-analysis decision with heterogeneity reported, not a curation shortcut. * Reanalysis targets GCF_963853765.1 / RS_2024_06, the annotation the crosswalk is built from, so this study carries no assembly crossing. One accession in an early draft of the study YAML (sra: SRP617521) was wrong — written from assumption rather than looked up. The verified value is SRP620802, from the ENA run report. Tests: 132 -> 146. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two commits: a screened backlog of real public datasets, and the registration of the first one.
1. A screened curation backlog
docs/candidate_studies.md— ten candidate M. gigas datasets, verified against NCBI BioProject and the ENA read-run API. Accessions, run counts, and group structures come from deposited metadata, not from recall or paper text.It corrects the selection criterion I had been working from. "Does the source publish full statistics" is the wrong primary filter: a significance-filtered table carrying only adjusted p-values is the normal format in this literature, not bad luck, so screening on it mostly reproduces the
HESSER2024_VCORdead end. The reliable route to poolable evidence is raw reads in SRA, because AREE's own DESeq2 run then emitslfcSEby construction.The page also records a rejected candidate in detail.
PRJNA623063looked ideal from its title — C. gigas larvae, Vibrio challenge, nearly the same phenotype as the study already registered — but all 12 runs share onesample_aliasacross 12 timepoints. Unreplicated time course, no valid DE contrast, and invisible from the abstract.Two findings worth knowing: hypoxia has no usable public dataset despite being in the phenotype ontology, and GEO holds only 10 oyster expression series, so a GEO-only search misses the corpus entirely.
2. CALLA2026_OSHV — registered, not harmonized
PRJNA1329250: 42 paired-end NovaSeq libraries, oyster spat from two hatchery populations challenged with three OsHV-1 microvariant isolates. Calla, Thompson & Burge 2026, Fish & Shellfish Immunology (10.1016/j.fsi.2026.111154).Nothing has been harmonized. The design, sample sheet, FASTQ manifest, run config, and reference context are complete and verified; the pipeline run has not been executed (226 GB of reads). The study contributes zero evidence records and
analysis_statusisnot_started.New:
aree fetch-samplesheetReads ENA's run report and BioSample attribute blocks; emits a design sheet, a FASTQ manifest carrying ENA's own MD5 per file, and checksummed provenance. The design above was derived this way rather than transcribed from the paper — which matters twice here, because the paper is paywalled, and because reading the archive is what catches an unreplicated design before someone downloads hundreds of gigabytes.
It warns when the smallest group has n<3, and
validate-studynow cross-checks a study's declared BioProject against the sheet it was generated from.Curation judgements, each with a test
This is not registered as resilience evidence, and that contradicts what I said when proposing it. I described this study as one whose "contrast is a resilience phenotype, not an exposure," on the strength of the BioProject title ("Evaluating Pacific oyster lineages for tolerance to Ostreid herpesvirus"). Curating it showed otherwise: there is no survival, mortality, or viral-load measurement for these animals in the deposited metadata or the abstract, and the paper frames its results as groundwork for future tolerance breeding. Registered as
disease_susceptibility, flaggedambiguous_phenotype_definition.The correction is written onto the backlog page rather than quietly edited out, because the lesson generalizes: judge on whether a phenotype was measured, not on how the title is worded.
Also:
GCF_963853765.1/RS_2024_06, the annotation the crosswalk is built from, so unlikeHESSER2024_VCORthis study carries no assembly crossing.Myagi; the paper spells itMiyagi. Deposited spelling preserved verbatim, discrepancy recorded.An error caught before commit
An early draft of the study YAML carried
sra: SRP617521. I wrote that from assumption rather than looking it up, and it is wrong. The verified value from the ENA run report isSRP620802. Flagging it here because a well-formed but fabricated accession is the kind of thing that survives review.Verification
Lint clean · 146 tests (was 132) · all studies validate · demo pipeline green · real-study path green. Confirmed that
CALLA2026_OSHVproduces 0 evidence records, as it should until the reanalysis runs.Honest state after this
The RNA-seq Nextflow workflow is still a scaffold that has never run against real FASTQ. This PR does not change that — it queues a real, fully specified job, and the first attempt should be expected to surface bugs in the workflow. Random-effects pooling on real data remains unexercised. All of that is stated in
docs/first_raw_reanalysis.mdand the status table.🤖 Generated with Claude Code