Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

Jump to content

Talk:Statistical hypothesis test

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia
Latest comment: 9 days ago by Jpeccoud in topic Error in decision criterion


Common test statistics

[edit]

I corrected the erroneous last test, ("regression t-test") to a correct F-test. Harald Lang, 2015-11-29.

Poorly organized

[edit]

I suggest defining the null hypothesis before citing its historical importance.159.83.248.41 (talk) 22:46, 29 August 2024 (UTC)Reply

Very hard to read

[edit]

I think this article needs significant reorganising & lots of junk removing.

  • "Choice of null hypothesis" does not belong in "history".
  • "History" and "modern origins" are not mutually exclusive.
  • The actual definition of a hypothesis or significance test is not clearly given.
  • Examples appear very very far down the article.
  • There are lots of paragraphs about the difference between significance tests and hypothesis tests but they are circumspect and do not clearly explain the difference.
  • In several places it is implied that hypothesis tests are "more rigorous" than significance tests, but I don't think this is accurate (I haven't been able to find sources that support it).
  • Very repetitive.

Danielittlewood (talk) 11:35, 17 May 2026 (UTC)Reply

Perhaps one way to deal with the repetition would be to try imposing some top-down structure on the article. Probably the first section should be something like "Definition and Terminology" where all the relevant terms can be defined ahead of time. Then maybe some examples, and then history? Danielittlewood (talk) 11:39, 17 May 2026 (UTC)Reply

Error in decision criterion

[edit]

I had this corrected before, but the error is back again in this article.

The significance level “alpha” is defined as the risk of rejecting a true null hypothesis (risk of type 1 error, or false positive). The p-value is defined as the probability of getting a test statistic at least as extreme as observed (not more extreme), under the null hypothesis. The page says one should reject the null hypothesis when the p-value is less than alpha. This rule contradicts the two definitions. If we reject H0 only when a sample yields a p-value that is strictly lower than alpha, the rejection rate of a true H0 might be lower than alpha, while it should equal alpha, by definition.

To illustrate: H0 is “this coin is fair” and H1 is “there is a probability >1/2 of getting a head” (one-sided test). We toss the coin 10 times. Our test statistic X is the number of heads observed in 10 trials. X follows Bi(10, 1/2) under H0. We get 5 heads. The p-value is P(X ≥ 5) = 0.6230469. You can check with R using binom.test(5, 10, 1/2, “greater”).

If we chose alpha = P(X ≥ 5) = 0.6230469, and decide to reject H0 when the p-value is strictly lower than alpha, we would reject H0 only if there are 6 heads of more, because if we get 5 heads, the p-value equals alpha. Getting 6 heads or more under H0 has a probably P(X ≥ 6) = 0.3769531. This is the rate at which we would reject the true H0. As you can see, it does not equal alpha.

The article must state that one should reject H0 when the p-value is equal or less than alpha. Jpeccoud (talk) 09:26, 24 September 2026 (UTC)Reply