10 minutes ago7 min read
Technical vs. Biological Replicates: The Pseudo-replication Trap That Sinks Papers
- 10 minutes ago
- 7 min read

When researchers audited 200 published animal studies, only 22% had replicated the right thing. Nearly half, 46%, had counted repeated measurements of the same biological unit as independent samples, and another 32% didn't report enough detail to tell. That error has a name, pseudo-replication, and it does something specific and damaging: it inflates your n, shrinks your error bars, and manufactures significance that isn't there.
The uncomfortable part is that this isn't a beginner's mistake. It survives peer review constantly, it's been warned about since 1929, and the reason it persists is that the line between a technical and a biological replicate is genuinely blurry in cell culture, where most of us work.
Here's the distinction that decides your n, why getting it wrong inflates false positives, and how to figure out what actually counts as a replicate in a cell-based experiment.
Technical vs biological replicates: what each one actually measures
The two answer different questions, and only one of them contributes to your sample size.
Biological replicates are independent biological units subjected to the same treatment: five different mice, cells cultured and treated independently, separate patient samples. They capture the natural, random variation between biological entities. That variation is the thing your statistics need, because your hypothesis is a claim about the population, and the spread between independent units is what tells you whether a treatment effect is bigger than the noise biology already contains. Biological replicates set your n.
Technical replicates are repeated measurements of the same biological unit: three wells loaded from the same lysate, the same sample run in triplicate on a plate reader, multiple images of one dish. They capture measurement error from the instrument, the operator, and the protocol. They make your estimate of that one sample more precise. Technical replicates do not increase your n.
The clean way to hold it: technical replicates tell you how well you measured; biological replicates tell you how much biology varies. A t-test is asking whether a difference exceeds biological variation, so feeding it measurement variation instead answers a question nobody asked.
The example that makes pseudo-replication obvious
Suppose you hypothesize that male mice have heavier brains than female mice. Two designs:
Design A: take one male and one female, weigh each brain five times.
Design B: take five males and five females, weigh each brain once.
Both hand you ten numbers. Both will produce a p-value. But the hypothesis is about sexes, and Design A contains exactly one member of each sex, so its p-value is meaningless no matter how small it is. Weighing the same brain five times tells you how reproducible your scale is. It tells you nothing about whether male brains are heavier in general, because you have no idea how much brains vary between males.
Note: Design A's p-value will often be smaller than Design B's, because repeated weighings of one brain agree closely. That's precisely the trap: the wrong design looks more convincing.
Why pseudo-replication inflates your false positives
The damage is mechanical, and it's worth understanding rather than memorizing.
Your standard error shrinks as n grows (SE = SD/√n). Counting technical replicates as independent samples inflates n artificially, so the SE comes out too small, the confidence interval comes out too narrow, and the degrees of freedom used in the test are wrong. The result is a p-value that overstates the evidence, which means an inflated Type I error rate: you reject a true null more often than your stated 5%.
Put in plain terms, pseudo-replication doesn't just make your analysis technically incorrect. It systematically manufactures the appearance of an effect, which is why it's a documented contributor to irreproducible results. Someone repeats your experiment with genuine biological replicates, encounters the real between-unit variation, and the effect evaporates.
Caution: the wrong choice of replicate cannot be repaired after the fact. No clever hierarchical model rescues an experiment that only ever contained one biological unit. This has to be settled at the design stage, before you plate anything.
The hard part: what counts as a biological replicate in cell culture?
With mice it's obvious. With cells it genuinely isn't, and anyone who tells you there's one clean rule hasn't looked closely. Researchers argue about this constantly, and the honest answer is that it depends on what you want to generalize to.
Work down this ladder, from clearly technical to clearly biological:
Multiple wells from the same plate, seeded from the same suspension: technical. They share everything upstream. Splitting one tube of cells across six wells gives you six measurements of one biological unit.
The same cell line, same passage, treated in parallel on the same day: usually technical. Independence is thin; everything traces to one culture at one moment.
Separate cultures maintained independently, treated on different days, at different passages: generally accepted as biological. This is the practical working standard for cell lines. Independent maintenance means independent handling, media lots, confluence histories, and passage-related drift, which is real between-unit variation.
Different donors, different animals, different patient-derived lines: unambiguously biological. Here the biological entity itself differs.
Tip: the useful test is what am I claiming? If the claim is "this drug slows migration in HeLa cells," your replicates need to capture variation between independent HeLa cultures, so independent cultures on separate days are the unit. If the claim is "this drug slows migration in human epithelial cells generally," a single immortalized line cannot support it at any n, because a monoclonal cell line doesn't carry the natural variation that claim is about. That's a limitation to state, not to solve with more wells.
One consequence worth internalizing: with a patient-derived line, you can run a hundred technical repeats and the original patient is still a single biological entity. Sample size is about independent units, never about how many measurements you took.
How to handle technical vs biological replicates correctly
The practical workflow is simpler than the conceptual debate.
Decide your experimental unit before you start, based on the claim you intend to make. This is the design decision everything else follows from.
Run technical replicates anyway. They're useful. They reduce measurement noise and flag pipetting or instrument problems.
Average your technical replicates first, collapsing them into one value per biological unit. Then run your statistics across biological units. This is the standard, defensible approach for most bench experiments.
Or model the structure explicitly with a nested or mixed-effects model, which keeps both levels of variation. Useful when the technical variation itself matters, but note it doesn't create biological units that were never there.
Report both numbers explicitly. State the number of biological replicates and the number of technical replicates per unit, and say which one your n refers to. Reviewers increasingly ask, and ambiguity here is one of the most common reasons a methods section gets flagged. Describing your design in plain language and having it checked before you commit to plating is a cheap sanity check (Sophie's assay design engine is one way to pressure-test whether your proposed replicates are genuinely independent).
Note: "n = 3" is not self-explanatory. Three wells, three flasks, three separate experiments, and three donors are four different experiments, and only your methods section can tell them apart.
The takeaway
The technical vs biological replicates distinction is not bookkeeping. It determines whether your statistics answer your hypothesis at all. Technical replicates measure your instrument and belong averaged into a single value per unit; biological replicates measure your biology and are the only thing that sets your n. Counting the former as the latter inflates sample size, narrows intervals, and produces significance that won't survive replication, which is exactly what nearly half of audited studies did. Decide your experimental unit from the claim you intend to make, fix it before the experiment rather than after, and state both numbers plainly in your methods. The trap only closes on people who never asked which of their replicates were actually independent.
FAQ
What is the difference between a technical and a biological replicate? A biological replicate is an independent biological unit given the same treatment (a different animal, an independently maintained culture, a separate donor) and it captures natural biological variation. A technical replicate is a repeated measurement of the same biological unit (multiple wells from one lysate, triplicate reads of one sample) and it captures measurement error. Only biological replicates contribute to your sample size.
What is pseudo-replication? Pseudoreplication is treating technical replicates, or repeated measurements of the same unit, as if they were independent biological samples. It artificially inflates n, shrinks the standard error and confidence interval, and inflates the Type I error rate, producing significant results that don't replicate.
Do technical replicates count toward n? No. They improve the precision of the measurement for a single biological unit, but they add no information about how much biology varies between units. Average them into one value per biological unit, then run your statistics across those units.
Are three wells on the same plate three biological replicates? Generally no. Wells seeded from the same cell suspension on the same day share their entire upstream history, so they are technical replicates. Independently maintained cultures treated on separate days, at different passages, are the usual working standard for biological replicates with cell lines.
Can I fix pseudo-replication with a mixed-effects model? Not retroactively. Nested and mixed-effects models correctly represent hierarchical data and are a good choice when you have multiple levels, but no model can create independent biological units that the experiment never contained. The choice has to be made at the design stage.
How should I report replicates in a methods section? State the number of biological replicates, the number of technical replicates per biological unit, and explicitly which one your reported n refers to. "n = 3 independent cultures, each measured in triplicate" is unambiguous; "n = 3" alone is not.
References
Lazic, S.E., Clarke-Williams, C.J. & Munafò, M.R. (2018). What exactly is 'N' in cell culture and animal experiments? PLOS Biology 16(4):e2005282 (only 22% replicated correctly; 46% pseudoreplication).
Pseudoreplication in physiology: more means less — Journal of General Physiology, PMC7814346.
Chan, S.W. et al. (2020). Replicates in stem cell models: how complicated! STEM CELLS.
What does "N" mean when using cell lines? — Bitesize Bio (monoclonal lines and natural variation).





