Skip to content
GenomarkerPUBLIC BENCHMARK · v3.6 · IN PROGRESS
Illustrative figures — audit in progress

The Public-Corpus Reproducibility Audit.

What happens when the Genomarker assurance model is applied to published scientific work it had no role in preparing? We re-executed published computational biology analyses — data, methods, and environments as published — and let the findings speak. The close-out ships as a preprint and this interactive public benchmark: named, reproducible material failures, replay rates, and preservation grades.

7

Published analyses re-executed end to end

of 9 papers audited · 1 refused with stated reason · 1 unauditable by construction

57

Structured assurance findings emitted

1 hard refuse · 11 block with override · 39 warn · 6 ok

12

Material — capable of changing interpretation

block-with-override + hard-refuse findings

Illustrative figures — audit in progress

71%

Detectable pre-execution, from the spec alone

Illustrative figures — audit in progress

Failure classes observed

Illustrative figures — audit in progress
  • Undeclared batch confoundingTreatment inseparable from batch; not reported in methods.

    14Material class
  • Statistical method mismatchTest applied outside its assumptions, or on pseudo-replicated units.

    11Material class
  • Environment reconstruction failureSoftware versions unrecoverable from the publication and supplements.

    12Provenance class
  • Unrecorded parametersThresholds and covariates absent from methods; recovered only by inference.

    10Provenance class
  • Undeclared nondeterminismStochastic steps without seeds; results vary across identical reruns.

    8Provenance class
  • Post-hoc deviationReported analysis differs from the stated analysis plan.

    6Material class

Replay outcomes

Illustrative figures — audit in progress
  • REPRODUCED_EXACT38%
  • REPRODUCED_EQUIVALENT33%
  • DIVERGED17%
  • NOT_REPLAYABLE12%
  • GM-DE-CONFOUND-004

    corpus paper 07 · pre-executionMaterialSpecimen

    Reported treatment effect is statistically inseparable from sequencing batch (Cramér's V 0.91). Re-analysis with batch adjustment removes 74% of reported hits.

    Detectable before execution — from the published design alone.

  • GM-ENV-RECON-002

    corpus paper 12 · replayProvenanceSpecimen

    Published methods name the package but not the version; three plausible versions produce three different DE gene lists.

    NOT_REPLAYABLE as published — replayable under any one pinned environment.

Every failure the audit surfaces becomes a versioned assurance rule, a numerical fixture, or a preservation policy — the corpus is both external proof and the beginning of a proprietary failure library.

Methodology & corpus selection