BLOG

Disease models drift, and a bulk average hides it.

Disease Modeling

July 14, 2026by Jacqueline Marin10 min read

Disease models drift, and a bulk average hides it.

Short tandem repeat (STR) profiling confirms a model's identity. It cannot confirm the model still represents the disease, and that gap decides whether you can trust the program you build on it.

Key takeaways

  • A bulk average can hold steady while a model's clonal composition drifts, so a model can pass STR and still stop representing the disease.
  • Identity is not fidelity. STR confirms the line is what it claims to be; it cannot confirm the model still holds the disease's clonal architecture.
  • Single-cell DNA reads each cell directly, exposing both the drift and the resistant clone an average conceals.

Disease models drift as they passage, engraft, and are engineered: the genetically distinct clones that make up a model shift over time. A bulk average can stay almost unchanged while that mix moves underneath it, so a model can pass STR identity testing and still stop representing the disease. Single-cell DNA sequencing reads each cell directly, revealing both the drift and the resistant clones an average buries.

Every preclinical program rests on a disease model, and every disease model is a population of genetically distinct clones, not a single uniform cell type. That population is not fixed. Xenografts, organoids, cell lines, and engineered lines all change as they grow, because culture and engraftment select for whichever clones grow best, which are rarely the ones that match the patient's disease.

Sanity Image
Infographic with three statistics on PDX model genomic drift: 1,110 patient-derived xenograft samples surveyed for drift (Ben-David et al., 2017); 24 cancer types across which models acquired recurrent genomic changes through clonal selection; and 3→30%, illustrating how far a resistant clone can expand while barely moving the average.

The usual way to confirm a model is STR profiling, which matches genetic markers against a reference and settles one question: is this the line it claims to be. Bulk DNA sequencing adds the average genotype across all the cells at once. Both share a blind spot. They report the model as a single summary, so a shift in which clones dominate barely moves the number. This is disease model drift, and the measurement most teams rely on to detect it, the population average, is built to hide it.

Disease model drift

The change in a model's clonal composition as it passages, engrafts, or is engineered, until the model no longer represents the disease it came from. Documented in xenografts, cell lines, and organoids, and able to alter how a model responds to a drug.

The model keeps its identity while its biology moves. A population average is built to hide exactly the drift you need to detect.

Sanity Image
Disease models drift over time — bulk sequencing sees the average, single-cell analysis sees the subclones. Download the eBook.

What does disease model drift hide from a population average?

A bulk variant allele frequency (VAF) is a single number per locus, pooled across every cell, so two very different models can produce the same value. A resistant clone that grows from 3% to 30% of cells can move the average only slightly, because the other cells dilute it. Three things disappear into that average.

Co-occurrence and zygosity

A bulk VAF cannot tell you whether two mutations sit in the same cell, which makes them one clone, or in separate cells, which makes them two. The therapeutic consequence is opposite in each case, and the average reads them identically.

Rare and emerging subclones

Standard bulk targeted sequencing resolves variants down to a few percent VAF, set by background error, so a resistant subclone below that floor is unreadable until it expands. Single-cell DNA resolves the clonal composition a bulk average pools, including co-occurring resistance alterations in individual surviving cells (Chen et al., Ann Oncol 2022, osimertinib-resistant lung cancer).

Change under selection

When the readout is how composition shifts after a treatment or across passages, a before-and-after average can move very little while the clones underneath it turn over completely.

Sanity Image
What a population average hides in a disease model. The same sample reads as one value on a bulk assay, while a single-cell census shows most clones cleared and one resistant clone surviving and expanding. Concept schematic, not platform data.

What today's confirmation methods can and cannot show

Four methods carry most of the load today, each the right tool for the question it was built for. The difference that matters for drift is whether the method reads each cell's DNA, or reports the population as a summary.

MethodWhat it confirmsReads each cell's DNA?
STR profiling
Identity: the line is what it claims to be, blind to the disease-relevant genome
No
Bulk targeted DNA / WES
Averaged genotype; UMIs push the floor toward 1% VAF, but no link between variants in one cell
No
Single-cell RNA
Cell states after treatment; genotype inferred from expression, not confirmed
No
Flow cytometry
Surface phenotype, fast, across a population; survivors disconnected from their mutations
No
Single-cell DNA
Clonal architecture and mutation co-occurrence, cell by cell
Yes

Each reports one thing: identity, or average genotype, or transcriptional state, or surface phenotype. None reads each cell's DNA across thousands of cells, the measurement that confirms a model still represents the disease and shows which clones a treatment leaves behind.

Identity is not fidelity. A model can be authentically itself and no longer be the disease.

Single-cell DNA sequencing reads what an average masks

Drift spans every model system, which is what makes a per-cell measurement worth the cost. It has been measured in each class a program uses: minor clones expand to dominate xenografts (Eirew et al., Nature 2015), cell-line strains diverge until they respond differently to the same compound (Ben-David et al., Nature 2018), organoids accumulate chromosomal changes in culture (Bolhaqueiro et al., Nat Genet 2019), and pluripotent and engineered lines acquire new mutations as they grow (Merkle et al., Nature 2017; Kuijk et al., Nat Commun 2020).

What the evidence shows. None of this work used a single-cell DNA platform; together it establishes drift as a property of the models themselves. That is the case for reading a model directly, instead of trusting it to be its average.

The methods-level response is to read DNA one cell at a time. Droplet microfluidics partitions thousands of cells, barcodes each, and amplifies a targeted panel, so every variant call is anchored to a single cell. The result is a census rather than an average: which clones are present, which mutations co-occur, and how the composition changes after a treatment. Single-cell mutation analysis on this principle resolves the clonal architecture of myeloid malignancies that a bulk average masks (Miles et al., Nature 2020). Reading protein or RNA from the same cells links a clone's genotype to what it expresses, so a survivor is characterized directly.

Encapsulate

Droplet microfluidics partitions thousands of single cells, one per droplet.

Lyse and barcode

Each cell is opened and tagged, so every variant call is anchored to its cell of origin.

Targeted multiplex PCR

A designed panel amplifies the loci that matter, at clonal depth.

Call cell by cell

The result is a census, not an average: which clones are present, which mutations co-occur, and how the mix changes after a treatment.

Where the Tapestri Platform fits, and where it does not

The Tapestri Platform applies this to disease-model characterization. A two-step droplet workflow encapsulates single cells, lyses them, and runs a targeted multiplex PCR, so single-nucleotide variants, insertions and deletions, copy-number changes, and known translocations are called cell by cell across thousands of cells from one input. The same run can read more than 40 surface proteins, or targeted RNA, from those cells, so a clone's genotype and phenotype come from one measurement.

For a disease model, that answers the two questions a population average cannot.

Fidelity

Is it still the disease?

Does the model still hold the clonal architecture of the disease, or has it drifted since derivation?

Response

Which clones survive?

When a lead candidate is applied, which clones does it clear and which survive? The surviving clone, small before treatment and able to drive relapse, is exactly the population a bulk average masks.

Sanity Image
The Tapestri Platform. A two-step droplet workflow reads each cell's DNA, with protein or RNA, and resolves the clonal architecture a bulk average pools into one number.

The limits matter as much as the capability. The platform characterizes which clones a treatment selects, and the mechanism behind it. It complements the dose-response assay; it does not replace that assay, measure efficacy, IC50, or potency, or screen a compound library. It confirms a lead candidate on a model you trust. Trial design and patient stratification are downstream, done by other teams who receive the clonal profile this work generates.

It is also honest about scope: as a targeted, panel-based assay it reads designed loci at clonal depth, so de novo genome-wide discovery still belongs to bulk sequencing. Teams without the instrument can run it through Mission Bio's Pharma Assay Development (PAD) service.

Before a program rides on a model, confirm:

  • The model still holds the disease's clonal architecture, not just its identity.
  • You can see the rare clones, not only the dominant average.
  • When a candidate is applied, you can tell which clones it clears and which survive.

Frequently asked questions

Q.What is disease model drift?

Disease model drift is the change in a model's clonal composition as it passages, engrafts, or is engineered, so the model stops representing the disease it came from. Drift has been documented in xenografts, cell lines, and organoids, and it can alter how a model responds to a drug (Ben-David et al., Nat Genet 2017).

Q.Does STR profiling confirm a disease model still represents the disease?

STR profiling confirms a model's identity, that it is the cell line it claims to be, by matching short-tandem-repeat markers to a reference. STR does not read the disease-relevant genome or clonal architecture, so it cannot detect drift. Identity and fidelity are different properties, and a model can pass STR while no longer representing the disease.

Q.Why does bulk sequencing miss a resistant subclone?

Bulk sequencing reports one variant frequency per locus, averaged across all cells, so a subclone present in a few percent of cells can sit beneath the detection floor. Single-cell DNA sequencing reads each cell separately, so a rare resistant clone appears as a discrete population; single-cell sequencing has resolved the clonal evolution of resistance in osimertinib-resistant lung cancer this way (Chen et al., Ann Oncol 2022).

Q.How is single-cell DNA sequencing different from single-cell RNA sequencing for disease models?

Single-cell DNA sequencing reads each cell's genotype directly, so it calls the mutations and clonal architecture of a model. Single-cell RNA sequencing reads expression and infers genotype indirectly, so it maps cell states but cannot confirm a resistance genotype at clonal depth. The two are complementary readouts.

Trust the model before the program rides on it

A disease model is the foundation of a preclinical program, and the tools used to trust it report it as a single summary. STR confirms identity. Bulk sequencing confirms the average genotype. Neither confirms the model still represents the disease, or shows which clones decide how it responds. That gap is where drift hides, and where a resistant survivor waits.

Reading the model one cell at a time closes that gap. Mission Bio's single-cell genotype-plus-targeted-gene-expression capability profiles DNA and RNA from the same cells, so the clone that survives a candidate is read directly, cell by cell. The question shifts from how much the model responded to which clones responded, and whether the model was sound to begin with.

Before a program rides on a model, confirm it is still the disease, then see which clones decide the result. The average will not tell you. A census will.

References

  1. Loewa A, Feng JJ, Hedtrich S. Human disease models in drug development. Nat Rev Bioeng. 2023;1:545-559. doi:10.1038/s44222-023-00063-3
  2. Eirew P, et al. Dynamics of genomic clones in breast cancer patient xenografts at single-cell resolution. Nature. 2015;518:422-426. PMID 25470049. doi:10.1038/nature13952
  3. Ben-David U, et al. Patient-derived xenografts undergo mouse-specific tumor evolution. Nat Genet. 2017;49:1567-1575. PMID 28991255. doi:10.1038/ng.3967
  4. Ben-David U, et al. Genetic and transcriptional evolution alters cancer cell line drug response. Nature. 2018;560:325-330. PMID 30089904. doi:10.1038/s41586-018-0409-3
  5. Bolhaqueiro ACF, et al. Ongoing chromosomal instability and karyotype evolution in human colorectal cancer organoids. Nat Genet. 2019;51:824-834. PMID 31036964. doi:10.1038/s41588-019-0399-6
  6. Merkle FT, et al. Human pluripotent stem cells recurrently acquire and expand dominant negative P53 mutations. Nature. 2017;545:229-233. PMID 28445466. doi:10.1038/nature22312
  7. Kuijk E, et al. The mutational impact of culturing human pluripotent and adult stem cells. Nat Commun. 2020;11:2493. PMID 32427826. doi:10.1038/s41467-020-16323-4
  8. Miles LA, et al. Single-cell mutation analysis of clonal evolution in myeloid malignancies. Nature. 2020;587:477-482. PMID 33116311. doi:10.1038/s41586-020-2864-x
  9. Chen J, et al. Single-cell DNA-seq depicts clonal evolution of multiple driver alterations in osimertinib-resistant patients. Ann Oncol. 2022;33(4):434-444. PMID 35066105. doi:10.1016/j.annonc.2022.01.004

For Research Use Only. Not for use in diagnostic procedures.

SHARE THIS PAGE

Request quote