The evidence contained in the P-value is context dependent
Abstract
In a recent opinion article, Muff et al. recapitulate well-known objections to the Neyman-Pearson null-hypothesis significance testing (NHST) framework and call for reforming our practices in statistical reporting. We agree with them on several important points: the significance threshold P < 0.05 is only a convention, chosen as a compromise between type I and II error rates; transforming the p-value into a dichotomous statement leads to a loss of information; and p-values should be interpreted together with other statistical indicators, in particular effect sizes and their uncertainty. In our view, a lot of progress in reporting results can already be achieved by keeping these three points in mind. We were surprised and worried, however, by Muff et al.’s suggestion to interpret the p-value as a “gradual notion of evidence”: Muff et al. recommend, for example, that a P-value greater than 0.1 should be reported as “little or no evidence” and a P-value of 0.001 as “strong evidence” in favor of the alternative hypothesis H1.
What the paper shows and why it matters (AI-generated)
Responding to a call to treat p-values as a fixed scale of evidence — p = 0.001 always meaning “strong evidence,” for instance — this short commentary argues that how much a given p-value actually tells you depends on the study design and the plausibility of the hypothesis going in, not on the number alone. Replacing one oversimplified statistical convention with a different oversimplified one doesn't fix ecology's reporting practices, it just moves the problem.