Skip to content
GenomarkerTHE PUBLIC-CORPUS REPRODUCIBILITY AUDIT · ROUND 1 · 2026

We audited published science we had no part in.

Nine published gene-expression studies, 2004–2022, re-run end to end on our production system with data and methods as published. Here is what we found — in the studies, and in our own platform.

9

Published studies audited

re-run on our production system

7

Re-executed end to end

1 refused with a stated reason · 1 not auditable from its deposit

57

Structured findings

1 hard refusal · 11 blocking · 39 warnings · 6 confirmations

1

Hidden confound caught — analysis refused

recorded in no metadata field; read from the raw instrument files

The catch

THE CATCH

Hidden in the raw files.

In a large public patient study, none of the 48 days on which samples were scanned held both patients and controls. No metadata field recorded it. Genomarker read the scan dates from the raw instrument files and refused the comparison: the data cannot separate the biology from the scan day.

A statement about what the public deposit supports — not about the published conclusions.

What the audit found in the studies

  1. Batch confounded with the comparison

    In one study the scan day perfectly separated patients from controls; in others, batch was nested in the very variable under test.

  2. Repeated measures treated as independent

    Several deposits carry the same subjects more than once, which inflates significance when the samples are analyzed as independent.

  3. Sample sheets that don't match the data

    Mismatched annotations, a typo that silently drops a tumour sample from the comparison, and a 60-versus-57 sample gap between a deposit and its paper.

  4. The wrong data for the method

    Differential expression run on TPM values, which the method's statistics are not built for.

  5. Batches too small to correct

    Batch levels holding a single sample, which no correction can estimate.

  6. Predictors that cannot be rebuilt

    A published predictor that could not be reconstructed from its own deposit, and a model retrained on that deposit that showed no signal on four independent measures.

Every finding is a statement about what a public deposit supports — not about a paper's conclusions.

What the audit found in Genomarker

24 findings were about our own platform.

An audit that only finds problems in other people's work isn't credible. Ours also found 24 in Genomarker — from a privacy screen that blocked a public, de-identified deposit, to a covariate feature that was advertised but unreachable from a researcher's own data. Every one was filed as a defect and routed to an owner.

How the studies were chosen

Round 1 deliberately mixes calibration cases — three papers later retracted, a signature that failed in a phase III trial, a canonical confounding case — with highly cited studies and a clean counterexample. The test: does Genomarker find the problems we know are there, and pass the study that should pass?

Every failure the audit surfaced is filed to become a versioned rule, a test fixture or a fixed defect — the start of a proprietary library of how real analyses fail.

Methodology & study selection