Fisher's Exact Test (Expanded)
Menu location: Analysis_Exact_Expanded Fisher.
This allows you to see a conventional Fisher's exact (Fisher-Irwin) test in more detail, including mid-P values. The complete conditional distribution for the observed marginal totals is displayed (Bailey, 1977); a probability below 0.0000001 is shown as a power of ten, for example 2.6773527659E-015.
Like the chi-square test for fourfold (2 by 2) tables, Fisher's exact test examines the relationship between the two dimensions of the table (classification into rows vs. classification into columns). The null hypothesis is that these two classifications are not different.
The P values in this test are computed by considering all possible tables that could give the row and column totals observed. A mathematical short cut relates these permutations to factorials; a form shown in many textbooks. StatsDirect uses the hypergeometric distribution for the calculation (Conover 1999). The test statistic that is hypergeometrically distributed is the first count A; its expectation is shown.
This exact treatment of the fourfold table should be used instead of the chi-square test when any expected frequency is less than 1 or 20% of expected frequencies are less than or equal to 5. With StatsDirect, it is reasonable to use Fisher's exact test by default because the computational method used can cope with large numbers.
StatsDirect uses the definition of a 2 sided P value described by Bailey (1977) (P values for all possible tables with P less than or equal to that for the observed table are summed). Many authors prefer to simply double the 1 sided P value (Armitage and Berry, 1994; Bland, 2000).
Consider using mid-P values and intervals when you have several similar studies within an overall investigation (Armitage and Berry, 1994; Barnard, 1989).
Assumptions:
· each observation is classified into exactly one cell
· the row and column totals are fixed, not random
The assumption of fixed marginal (row/column) totals is controversial and causes disagreements such as the best approach to two sided inference from this test.
DATA INPUT:
Observed frequencies should be entered as a standard fourfold table. The frequencies should be whole numbers; one that is not is rounded to the nearest whole number. A table with an empty row or an empty column cannot be analysed.
| feature present | feature absent | |
| outcome positive: | a | b |
| outcome negative: | c | d |
Example
From Armitage and Berry (1994, p. 138).
The following data compare malocclusion of teeth with type of feeding received by infants.
| Normal teeth | Malocclusion | |
| Breast fed: | 4 | 16 |
| Bottle fed: | 1 | 21 |
To analyse these data in StatsDirect you must select the expanded Fisher's exact test function from the exact tests section of the analysis menu. Enter the frequencies into the contingency table on screen as shown above.
For this example:
Arranged table and totals:
| 4 | 1 | 5 |
| 16 | 21 | 37 |
| 20 | 22 | 42 |
Expectation of A = 2.380952
| A | Lower Tail | Individual P | Upper Tail |
| 0 | 0.03095684803002 | 0.03095684803002 | 1.00000000000000 |
| 1 | 0.20293933708568 | 0.17198248905566 | 0.96904315196998 |
| 2 | 0.54690431519700 | 0.34396497811132 | 0.79706066291432 |
| 3 | 0.85647279549719 | 0.30956848030019 | 0.45309568480300 |
| 4 | 0.98177432323774 | 0.12530152774055 | 0.14352720450281 |
| 5 | 1.00000000000000 | 0.01822567676226 | 0.01822567676226 |
One sided (upper tail) P = 0.1435 (doubled: P = 0.2871)
Two sided (by summation) P = 0.1745
One sided mid-P = 0.0809
Two sided mid-P = 0.1618
Here we cannot reject the null hypothesis that there is no association between these two classifications, i.e. between feeding mode and malocclusion.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Fisher's exact test (expanded): the StatsDirect help example (Armitage and Berry 1994,
# p. 138, malocclusion of the teeth by the type of feeding received by infants) in R
teeth <- matrix(c(4, 16, 1, 21), 2, byrow = TRUE,
dimnames = list(fed = c("breast", "bottle"),
teeth = c("normal", "malocclusion")))
print(addmargins(teeth))
# R's standard test. Its two sided P sums the probabilities of every table with the
# observed margins that is no more probable than the observed table (0.1745), the
# definition of Bailey (1977) that the report uses. With alternative = "greater" the P
# value is the probability of the observed count of breast-fed infants with normal
# teeth (the top left cell, A) or of a larger one: the upper tail of A (0.1435).
# fisher.test also gives a conditional odds ratio and its confidence interval, which
# the expanded report leaves out.
print(fisher.test(teeth))
print(fisher.test(teeth, alternative = "greater"))
six <- function(x) formatC(x, digits = 6, format = "f", drop0trailing = TRUE)
pv <- function(p) {
if (p < 0.0001) "P < 0.0001" else
paste("P =", formatC(p, digits = 4, format = "f", drop0trailing = TRUE))
}
# Given the margins, A is hypergeometric: the k = 20 breast-fed infants are a sample of
# the 42, of whom m = 5 have normal teeth and n = 37 malocclusion, and A counts the
# normal teeth in the sample
a <- teeth[1, 1]
m <- sum(teeth[, 1])
n <- sum(teeth[, 2])
k <- sum(teeth[1, ])
cat("Expectation of A =", six(k * m / (m + n)), "\n")
# The whole conditional distribution to 14 places, as the report tabulates it: for each
# possible A the probability of that value (individual P, dhyper), of that value or a
# smaller one (lower tail, phyper) and of that value or a larger one (upper tail: the
# probability above A - 1, which phyper gives with lower.tail = FALSE). Each individual
# probability is a ratio of counts of tables, choose(m, A) * choose(n, k - A) divided
# by choose(m + n, k).
A <- max(0, k - n):min(k, m)
prob <- dhyper(A, m, n, k)
lower <- phyper(A, m, n, k)
upper <- phyper(A - 1, m, n, k, lower.tail = FALSE)
fourteen <- function(x) formatC(x, digits = 14, format = "f")
cat("A Lower Tail Individual P Upper Tail\n")
for (i in seq_along(A)) {
cat(A[i], fourteen(lower[i]), fourteen(prob[i]), fourteen(upper[i]), "\n")
}
# One sided P: the tail that holds the observed A, the upper tail here as A = 4 is
# above its expectation; doubling it is the other common two sided P. The two sided P
# by summation adds up every value of A that is no more probable than the observed one
# (a tolerance of one part in ten million allows for rounding in the probabilities).
p_obs <- prob[A == a]
if (a > k * m / (m + n)) {
tail <- "(upper tail)"
p1 <- upper[A == a]
} else {
tail <- "(lower tail)"
p1 <- lower[A == a]
}
p2 <- sum(prob[prob <= p_obs * (1 + 1e-7)])
cat("One sided ", tail, " ", pv(p1), " (doubled: ", pv(min(1, 2 * p1)), ")\n",
sep = "")
cat("Two sided (by summation)", pv(p2), "\n")
# Mid-P: the one sided P less half the probability of the observed table (Armitage and
# Berry 1994), and twice that, at most 1, for the two sided mid-P
mid_p <- p1 - p_obs / 2
cat("One sided mid-", pv(mid_p), "\n", sep = "")
cat("Two sided mid-", pv(min(1, 2 * mid_p)), "\n", sep = "")