We audited published science we had no part in.
Nine published gene-expression studies, 2004–2022, re-run end to end on our production system with data and methods as published. Here is what we found — in the studies, and in our own platform.
9
Published studies audited
re-run on our production system
7
Re-executed end to end
1 refused with a stated reason · 1 not auditable from its deposit
57
Structured findings
1 hard refusal · 11 blocking · 39 warnings · 6 confirmations
1
Hidden confound caught — analysis refused
recorded in no metadata field; read from the raw instrument files
The catch
THE CATCH
Hidden in the raw files.
In a large public patient study, none of the 48 days on which samples were scanned held both patients and controls. No metadata field recorded it. Genomarker read the scan dates from the raw instrument files and refused the comparison: the data cannot separate the biology from the scan day.
A statement about what the public deposit supports — not about the published conclusions.
What the audit found in the studies
Batch confounded with the comparison
In one study the scan day perfectly separated patients from controls; in others, batch was nested in the very variable under test.
Repeated measures treated as independent
Several deposits carry the same subjects more than once, which inflates significance when the samples are analyzed as independent.
Sample sheets that don't match the data
Mismatched annotations, a typo that silently drops a tumour sample from the comparison, and a 60-versus-57 sample gap between a deposit and its paper.
The wrong data for the method
Differential expression run on TPM values, which the method's statistics are not built for.
Batches too small to correct
Batch levels holding a single sample, which no correction can estimate.
Predictors that cannot be rebuilt
A published predictor that could not be reconstructed from its own deposit, and a model retrained on that deposit that showed no signal on four independent measures.
Every finding is a statement about what a public deposit supports — not about a paper's conclusions.
What the audit found in Genomarker
24 findings were about our own platform.
An audit that only finds problems in other people's work isn't credible. Ours also found 24 in Genomarker — from a privacy screen that blocked a public, de-identified deposit, to a covariate feature that was advertised but unreachable from a researcher's own data. Every one was filed as a defect and routed to an owner.
How the studies were chosen
Round 1 deliberately mixes calibration cases — three papers later retracted, a signature that failed in a phase III trial, a canonical confounding case — with highly cited studies and a clean counterexample. The test: does Genomarker find the problems we know are there, and pass the study that should pass?
Every failure the audit surfaced is filed to become a versioned rule, a test fixture or a fixed defect — the start of a proprietary library of how real analyses fail.
Methodology & study selection