top of page
photo_2026-01-04_19-44-31_edited.jpg

Got Questions?

Two-Way ANOVA Interaction Effects: Step-by-Step SOP for Testing Your Treatment Between Groups

6 days ago
7 min read
Two-Way ANOVA Interaction Effects: A Step-by-Step SOP for Testing Whether Your Treatment Works Differently Between Groups

Your drug shrinks tumors in wild-type mice (p = 0.02) and does nothing significant in the knockout (p = 0.21), so the effect depends on the gene. That conclusion does not follow. Only a two-way ANOVA interaction test can tell you whether an effect differs between groups.

It is one of the most common errors in published biology. Of 157 papers in Science, Nature, Neuron and similar journals that needed an interaction test, 79 compared two p-values instead, and the error looked even more common in cellular and molecular neuroscience. The difference between "significant" and "not significant" is not itself statistically significant.

This SOP runs the correct test in GraphPad Prism, from data entry to a reportable result, with a worked example where the naive reading and the right answer disagree.


What this SOP covers

Ordinary two-way ANOVA: two categorical factors (such as genotype and treatment), one continuous outcome, and independent observations. If you measure the same animal or culture more than once, you need a repeated-measures model instead. If one variable is continuous, use ANCOVA.


Why the two-way ANOVA interaction is the only test of "different effect"

A two-way ANOVA splits your data's variation three ways: a main effect of each factor, and an interaction. In a genotype × treatment design, the interaction is literally a difference between two differences: the drug effect in wild-type minus the drug effect in the knockout.

Separate t-tests never test that quantity. They test each drug effect against zero, so a smaller or noisier effect can miss p < 0.05 while being statistically indistinguishable from the one that made it.

Being a difference of differences, the interaction also carries a larger standard error: exactly twice that of a main effect in a balanced 2 × 2 design. Estimating it as precisely as a main effect of the same size takes four times as many animals, so most interaction tests are underpowered, which is why you must read the test rather than infer it.


Before you start: what the design needs

  • Two categorical factors and one continuous outcome.

  • At least two biological replicates in every cell. With one value per cell there are no degrees of freedom left for the interaction, and Prism cannot test it. Replicates must be independent animals or cultures, never repeat wells from one technical sample.

  • Balanced cells where possible. Prism handles unequal n with Type III sums of squares, but equal n keeps every test at full power.

  • GraphPad Prism 8 or later, which adds a table of cell means to the output.


Entering the data in Prism

Step 1. On the Welcome dialog, choose the Grouped tab.

Step 2. Put the factor whose effect you are testing in the data-set columns (Vehicle, Drug) and the factor you suspect changes that effect in the rows (WT, KO).

Note: The layout decides which comparisons Prism offers in Step 7. With treatment in columns, "compare within each row" means "the drug effect within each genotype," which is usually the question.

Step 3. On the same dialog, give the Grouped table enough replicate subcolumns for your largest group, then enter one biological replicate per subcolumn. Leave missing values blank.

Caution: Enter raw values, not means. Prism can run the ANOVA from mean, SD and n, but you lose the residuals you need in Step 10.

Running the analysis

Step 4. Click Analyze and choose Two-way ANOVA from the grouped analyses.

Step 5. On the RM Design tab, leave both boxes unchecked.

Step 6. On the Factor Names tab, name the row and column factors ("Genotype", "Treatment") so the output is readable.

Step 7. On the Multiple Comparisons tab, choose Compare each cell mean with the other cell mean in that row. With three or more treatment columns, choose Simple effects. Within each row, compare columns instead.

Step 8. On the Options tab, choose Šídák for these planned within-row comparisons, tick Report multiplicity adjusted P value for each comparison where offered, and keep alpha at 0.05.

Note: Use Tukey only to compare every cell with every other, and Dunnett when everything is compared with one control. The Tukey vs Bonferroni guide covers the trade-offs.

Step 9. On the Residuals tab, tick the residual plot and the QQ plot, plus the tests for normality and equal variance. Click OK.

Tip: Prefer plain English to dialogs? Soφ runs the same two-way model from a description like "two genotypes, vehicle or drug, six mice each; does the drug work differently in the knockout?"

Checking the assumptions

Step 10. Open the QQ plot. Residuals should fall close to the diagonal line, and the normality test on the residuals should not be significant.

Caution: Test the residuals, not each group's raw values. The model assumes normal residuals, and six values per group are too few to judge anything separately.

Step 11. Open the residual plot. The spread should look similar across predicted values. If it fans out as the mean grows, which is common for tumor volumes, concentrations and cell counts, log-transform the outcome and rerun from Step 4. The non-normal data guide covers what to do when a transform is not enough.


Reading the two-way ANOVA interaction first

Step 12. In the ANOVA table, read the Interaction row before either main effect.

Step 13. Take the four cell means from the means table and plot them as two lines, one per genotype, with treatment on the x-axis. A quick sketch on paper is enough. Parallel lines mean no interaction. Lines that converge or diverge mean the effect changes in size. Lines that cross mean the effect reverses direction.

Step 14. Decide which effects you may interpret:

  • Interaction p ≥ 0.05: read the main effects. The treatment row is the drug's average effect across genotypes, with no evidence it differs between them.

  • Interaction p < 0.05, lines do not cross: the main effect is real but incomplete. Lead with the within-row comparisons.

  • Interaction p < 0.05, lines cross: ignore the main effects, which average opposite effects into a number that describes no real group. Report the within-row comparisons only.

Step 15. Read the within-row comparisons: each difference, its adjusted confidence interval and its adjusted p-value.

Caution: One comparison significant and the other not is still not evidence that they differ. Only the interaction row in Step 14 can say that.

Worked example: a drug in wild-type and knockout mice

Illustrative data. Tumor volume at day 14, six mice per group:

Genotype

Vehicle (mean ± SD, mm³)

Drug (mean ± SD, mm³)

Drug effect

Wild-type

520 ± 78

400 ± 72

−120

Knockout

500 ± 76

440 ± 81

−60

The naive reading. Separate t-tests give p = 0.020 in wild-type and p = 0.21 in the knockout: the drug seems to need the gene.


The two-way ANOVA:

Source

F (DFn, DFd)

p

% of variation

Interaction

F(1, 20) = 0.92

0.35

3.1%

Treatment

F(1, 20) = 8.26

0.009

28.2%

Genotype

F(1, 20) = 0.10

0.75

0.3%

Residuals pass both checks (Shapiro–Wilk p = 0.31; Brown–Forsythe p = 0.99).


The interaction is not significant. The two drug effects differ by 60 mm³, but the 95% CI on that difference (not part of Prism's default output) runs from −191 to +71 mm³: it includes no difference at all, and a knockout effect larger than the wild-type one. Per Step 14, read the treatment main effect instead: the drug cut tumor volume by 90 mm³ across genotypes (95% CI 25 to 155 mm³, p = 0.009).

The within-row comparisons repeat the naive pattern (wild-type −120 mm³, Šídák-adjusted p = 0.027; knockout −60 mm³, p = 0.35), which is why Step 15 carries its warning. The defensible conclusion: the drug works, and this experiment cannot tell whether the knockout blunts it.


Reporting the result

Report the interaction test first, with both degrees of freedom: "the genotype × treatment interaction was not significant, F(1, 20) = 0.92, p = 0.35." Then report the effect that answers your question, as a difference with a confidence interval rather than a p-value alone. The effect size guide covers standardized versions for readers who need them.


Troubleshooting

Symptom

Cause

Fix

No interaction row in the output

One value per cell

Add biological replicates; at least two per cell

Significant in one group, not the other, interaction not significant

The comparison trap

Report the main effect; do not claim the effects differ

Significant interaction, no main effects

Crossover interaction

Report within-row comparisons only

QQ plot bends away from the line, or residuals fan out

Skewed outcome, variance rising with the mean

Log-transform and rerun

Unequal n across cells

Missing animals or failed wells

Keep all data; Prism uses Type III sums of squares

Same animal appears in more than one cell

Repeated measures

Use a repeated-measures or mixed-effects model


The two-way ANOVA interaction answers the question you were actually asking

"Does my treatment work differently in these groups?" is a question about a difference between effects, and the interaction row is the only line in the output that tests it. Read it first, plot the means, and let it decide which other numbers you are entitled to report.

The data behind a two-way ANOVA usually start at the bench.



Frequently asked questions

What does a significant interaction mean in a two-way ANOVA? The effect of one factor depends on the level of the other: here, the drug's effect differs between genotypes by more than chance would explain.

Can I interpret main effects when the interaction is significant? Only if the interaction lines do not cross. Crossing lines mean the main effect averages opposite effects and describes no real group.

Why is my interaction not significant when the drug only worked in one group? Because each group's p-value compares its effect with zero, not with the other group's effect. Two effects on either side of p = 0.05 can be statistically indistinguishable.

How many replicates do I need for a two-way ANOVA interaction? At least two biological replicates per cell to test it at all, and about four times the sample of a main effect to estimate it as precisely. Power the study on the interaction itself.



  1. Nieuwenhuis S, Forstmann BU, Wagenmakers EJ. Erroneous analyses of interactions in neuroscience: a problem of significance. Nature Neuroscience, 2011;14(9):1105–1107.

  2. Gelman A, Stern H. The Difference Between "Significant" and "Not Significant" is not Itself Statistically Significant. The American Statistician, 2006;60(4):328–331. DOI

  3. NIST/SEMATECH e-Handbook of Statistical Methods. 7.4.3.7 The two-way ANOVA and 7.4.3.8 Models and calculations for the two-way ANOVA.

  4. GraphPad Prism 11 Statistics Guide. Interpreting results: Two-way ANOVA; Entering data for two-way ANOVA; Multiple comparisons tab; Options tab; Residuals tab.

  5. Ford C. Understanding 2-way Interactions. University of Virginia Library.

  6. Kiernan D. Main Effects and Interaction Effect. Natural Resources Biometrics, SUNY College of Environmental Science and Forestry (OpenSUNY).

  7. numiqo. Two-factorial ANOVA without repeated measures.


bottom of page