
FANR 6750
Fall 2026
Following a significant F-test, the next step is to determine which means differ
If all group means are to be compared, then we should correct for multiple testing
That is, conducting many tests increases the probability of finding one that is significant, even if it is not

Aphids are a major crop pest and farmers are interested which brands of pesticide have the biggest influence on crop yield
The data:
Call:
lm(formula = Yield ~ Brand, data = pesticidedata)
Residuals:
Min 1Q Median 3Q Max
-6.932 -1.499 0.335 1.406 5.459
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 36.357 0.494 73.66 < 2e-16 ***
BrandA 3.142 0.698 4.50 1.9e-05 ***
BrandB 3.940 0.698 5.65 1.7e-07 ***
BrandC 5.616 0.698 8.05 2.3e-12 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 2.47 on 96 degrees of freedom
Multiple R-squared: 0.416, Adjusted R-squared: 0.397
F-statistic: 22.7 on 3 and 96 DF, p-value: 3.29e-11
One can always fall back on pairwise, two-sample t-tests (or linear models)
One can always fall back on pairwise, two-sample t-tests (or linear models)…BUT
Running multiple tests increases the probability of finding one that is significant, even if it is not (remember the jellybeans!)

One common way to deal with multiple comparisons is to do pairwise t-tests but adjust the p-values of each test
One of the most common adjustments is the Bonferroni correction
Multiply each p-value by the number of tests
If this results in p > 1, set p to 1
At first, it can be confusing when a p-value is significant in one case (e.g., unadjusted) but not significant in another (e.g., Bonferroni)1
If you find yourself in this situation, take a step back and think about a few things:
What does the p-value measure? What is the risk when rejecting the null hypothesis?
When is falsely rejecting the null hypothesis (Type I error) most likely?
What is so special about \(\alpha = 0.05\)?

The combination of some data and an aching desire for an answer does not ensure that a reasonable answer can be extracted from a given body of data
According to Tukey’s Honestly Significant Difference test, two means (\(\bar{y}_i\) and \(\bar{y}_j\)) are different if:
\[\large |\bar{y}_i - \bar{y}_j | \geq q_{1- \alpha,a,a(n-1)}\sqrt{\frac{MSW}{n}}\]
where \(q\) comes from the “Studentized Range Distribution”(see qtukey in R). MSW = mean squared error from the ANOVA table
Tukey multiple comparisons of means
95% family-wise confidence level
Fit: aov(formula = Yield ~ Brand, data = pesticidedata)
$Brand
diff lwr upr p adj
A-Control 3.1424 1.3174 4.967 0.0001
B-Control 3.9403 2.1153 5.765 0.0000
C-Control 5.6159 3.7909 7.441 0.0000
B-A 0.7979 -1.0271 2.623 0.6639
C-A 2.4735 0.6485 4.299 0.0034
C-B 1.6756 -0.1495 3.501 0.0838
Bonferonni and Tukey’s HSD are the most commonly used types of a posteriori comparisons - comparing all possible pairwise combinations only after a significant F-test
There are many other types of a posteriori comparison, e.g. Fisher’s LSD (see ?p.adjust for even more)
A posteriori comparisons decrease the risk of Type I error (but increase the risk of Type II error)2
By forcing the researcher to correct for all possible combinations, a posteriori comparisons often overly conservative
A better approach in many situations is to use a priori contrasts
Watershed
|
|||
|---|---|---|---|
| Twelvemile | Chattooga | Keowee | Coneross |
| 16 | 18 | 28 | 14 |
| 8 | 25 | 22 | 20 |
| 12 | 22 | 24 | 11 |
| 17 | 16 | 20 | 17 |
| 13 | 26 | 27 | 15 |
Chattooga and Keowee are more forested
Coneross and Twelvemile are more agricultural
Call:
lm(formula = Length ~ Watershed, data = musseldata)
Residuals:
Min 1Q Median 3Q Max
-5.4 -2.5 -0.2 3.0 4.6
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 21.40 1.64 13.02 6.2e-10 ***
WatershedConeross -6.00 2.32 -2.58 0.0201 *
WatershedKeowee 2.80 2.32 1.20 0.2458
WatershedTwelvemile -8.20 2.32 -3.53 0.0028 **
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 3.67 on 16 degrees of freedom
Multiple R-squared: 0.645, Adjusted R-squared: 0.579
F-statistic: 9.7 on 3 and 16 DF, p-value: 0.000692
Pairwise comparisons using t tests with pooled SD
data: musseldata$Length and musseldata$Watershed
Chattooga Coneross Keowee
Coneross 0.120 - -
Keowee 1.000 0.010 -
Twelvemile 0.017 1.000 0.001
P value adjustment method: bonferroni
Instead of comparing all possible combinations, what if we are only interested in some specific comparisons?
Do the agricultural watersheds differ from one another?
Do the forested watersheds differ from one another?
Also, what if we are interested in combinations of treatments?
What are the advantages of focusing only on these specific questions?
Hypotheses
| Comparisons | H0 to test | |
|---|---|---|
| 1 | Forested vs agricultural | \(\frac{\mu_T + \mu_{Co}}{2} - \frac{\mu_{Ch} + \mu_K}{2}=0\) |
| 2 | Twelvemile vs Coneross | \(\mu_{T} - \mu_{Ch} = 0\) |
| 3 | Chattooga vs Keowee | \(\mu_{Ch} - \mu_{K} = 0\) |
Linear combinations
| 1 | \((1/2)\mu_T - (1/2)\mu_{Ch} - (1/2)\mu_{K} + (1/2)\mu_{Co}\) |
| 2 | \((1)\mu_T + (0)\mu_{Ch} + (0)\mu_{K} - (1)\mu_{Co}\) |
| 3 | \((0)\mu_T + (1)\mu_{Ch} - (1)\mu_K + (0)\mu_{Co}\) |
A set of linear combinations is called a set of orthogonal contrasts3 if the following conditions hold for all pairs of linear combinations:
Given
\[L_1 = a_1\mu_1 + a_2\mu_2 + ... + a_a\mu_a\]
and
\[L_2 = b_1\mu_1 + b_2\mu_2 +...+ b_a\mu_a\]
then \(L_1\) and \(L_2\) are orthogonal if:
\[\sum_i a_i=0;\;\; \sum_i b_i=0; \;\;and \; \sum_i a_ib_i=0\]
Returning to the question: “Does mussel size differ among forested and agricultural watersheds?”:
\[H_{0_1} = \frac{\mu_T + \mu_{Co}}{2} - \frac{\mu_{Ch} + \mu_{K}}{2} = 0\]
Multiplying through by 2 gives:
\[H_{0_1} = (\mu_{T} + \mu_{Co}) - (\mu_{Ch} + \mu_K) = 0\]
Which can be written as4: \[L1 = (1)\mu_T + (-1)\mu_{Ch} + (-1)\mu_K + (1)\mu_{Co}\] where the coefficients are \(a_1 = 1\), \(a_2 = -1\), \(a_3 = -1\), and \(a_4 = 1\)
Does each set of coefficients sum to 05?
\(L_1\): \(\sum_i a = 1 - 1 - 1 + 1 = 0\)
\(L_2\): \(\sum_i a = 1 + 0 + 0 - 1 = 0\)
\(L_3\): \(\sum_i a = 0 + 1 - 1 + 0 = 0\)
Do the products of pairs of coefficients sum to 0?
For \(L_1\) and \(L_2\): \((1)(1) + (-1)(0) + (-1)(0) + (1)(-1) = 0\)
For \(L_1\) and \(L_3\): \((1)(0) + (-1)(1) + (-1)(-1) + (1)(0) = 0\)
For \(L_2\) and \(L_3\): \((1)(0) + (0)(1) + (0)(-1) + (-1)(0) = 0\)
To obtain the sums of squares for each contrast, we use the general formula:
\[SS_L =\frac{(\sum_i a_i T_i)^2}{n \sum_i a_i^2}\]
where \(T_i\) is the sum of observations in group \(i\), and \(a_i\) is the corresponding coefficient for group \(i\)
For thefirst hypothesis we have:
\[ SS_{L_1} = \frac{(66-107-121+77)^2}{5(1 + 1 + 1 + 1)} = 361.2\]
For the second hypothesis we have: \(L_2 = \mu_T + 0 + 0 - \mu_{Co}\) with coefficients \(a_i = (1, 0, 0,-1)\)
\[SS_{L_2} = \frac{(66 +0 +0 - 77)^2}{5(1 + 0 + 0 + 1)} = 12.1\]
For the third hypothesis we have: \(L_3 = 0 + \mu_B - \mu_C + 0\) with coefficients \(a_i = (0, 1,-1, 0)\)
\[SS_{L_3} = \frac{(0 + 107 - 121 + 0)^2}{5(0 + 1 + 1 + 0)} = 19.6\]
For each contrast, each SS is divided by 1 d.f., and then divided by MSW
| Source | df | SS | MS | F | p |
|---|---|---|---|---|---|
| Among watersheds | 3 | 393 | 131.0 | 9.70 | 0.007 |
| F vs Ag | 1 | 361 | 361.0 | 26.76 | 9e-5 |
| T vs Co | 1 | 12 | 12.0 | 0.90 | 0.36 |
| Ch vs K | 1 | 20 | 20.0 | 1.45 | 0.25 |
| Within | 16 | 216 | 13.5 |
Notice that the Sums of Squares are partitioned according to hypotheses we are interested in, unlike in multiple comparison procedures.
Orthogonal contrasts are more powerful than multiple comparison procedures
They also require more thought and preparation (good things!)
There can be only \(a - 1\) comparisons
If the contrasts are not orthogonal, they can’t be used to fully partition the Sums of Squares among groups
If the comparisons really represent more than 1 treatment variable, it will be better to use a factorial design. More on that later
Next time: Statistical power
Reading: Quinn chp. 7.3
Especially if you’re a grad student working on your thesis and you really want significant results
We’ll discuss the trade off with Type II error in the next lecture
Statistically, “orthogonal” is a type of independence. With regard to contrasts, orthogonal means that we can partition variability in the response variable between the different a priori hypotheses
Note that it’s easier to work with coefficients that are integers rather than fractions
In practice, finding contrasts that are orthogonal often takes some trial and error