Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a221326689b03191

Jump to content

Talk:Probability distribution

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia
Latest comment: 2 months ago by Johnjbarton in topic First Two Sentences

Let X e = (X1, X2, . . . , Xn) be a random sample from the normal distribution with mean µ and variance σ 2 , find the maximum likelihood estimators of µ and σ 2  Preceding unsigned comment added by 197.220.137.224 (talk) 08:06, 22 March 2024 (UTC)Reply

Lead is confusing

[edit]

"random variable may be attributed to a function defined on a state space equipped with a probability distribution that assigns a probability to every subset ... of its state space..." A function defined on a state space should assign something to points of the state space. If it assigns something to sets, it is rather a SET FUNCTION. Also, if the state space is already equipped with a probability distribution, what is the role of the random variable?

"A random variable then defines a probability measure on the sample space by assigning a subset of the sample space the probability of its inverse image in the state space." Really? The two states are swapped here. On the sample space a measure is given from the beginning; on the state space it appears due to the random variable.

"In other words the probability distribution of a random variable is the push forward measure of the probability distribution on the state space." I do not understand this phrase; once again, on which space the measure is given from the beginning? Boris Tsirelson (talk) 16:03, 12 December 2008 (UTC)Reply

Organization of the article: subsection Terminology etc

[edit]

Perhaps subsection Terminology with its current content is superfluous. Repetition of discrete vs continuous distribution seems confusing (the reader could think he overlooked sth he didn't). The notion of support should appear in the article on measures. Measure, probability and possibly a few other concepts should be referenced at the bottom of the article. More thoughts are needed about organization of the vast legion of important links to related articles.

As a continuous counterpart of formula for probability distribution should be used that with integral using density function (since resembles discrete version very well - just think of density as discrete histogram with dense set of values). After presenting discrete and absolutely continuous distributions in separate sections, there should appear the third section unifying both into the general theory with formula using Lebesgue integral w.r.t. axiomatic probability measure P (this formula which is now horribly serving as definition in the special case of absolute continuous random variable - inacceptable for general purpose encyclopedia). Then some properties, examples and graphs should follow (independently of famous and important distributions described in dedicated articles). Compare e.g. with some good stylistic and logistic ideas applied to the article on expected value. --Megaloxantha (talk) 02:37, 31 December 2008 (UTC)Reply

New "Some Properties" section

[edit]

Can we do something sensible about the stuff presently in this section, which says:

  • The probability density function of the sum of two independent random variables is the convolution of each of their density functions.
  • The probability density function of the difference of two random variables is the cross-correlation of each of their density functions.

In the first bullet point, there is a need to say this for general distributions, not just those which have densities ...unfortunately the article on convolution does not seem to give a sensible formula in terms of cumulative distribution functions. In the second bullet point, there is again the problem of dealing with general distributions but, in addition, the use of the word "cross-correlation" would need to be given the interpretation in the article pointed-to, which is very different from the common one in statistics.

Melcombe (talk) 10:38, 31 December 2008 (UTC)Reply

Proposal to archive discussion

[edit]

Given that this article has had a major revamp, much of the old discussion is irrelevant to the present content. I have reordered the threads a little to put newer stuff towards the end, but I am suggesting that all the stuff now before the section headed "Lead is confusing" be archived. Any thoughts? Melcombe (talk) 12:15, 31 December 2008 (UTC)Reply

Discussion archived as above. Melcombe (talk) 14:21, 20 May 2009 (UTC)Reply

Observation Space

[edit]

In the formal definition, we have:

"A random variable is defined as a measurable function X from a probability space to its observation space ."

Can we have someone put up a definition of the "observation space"? I'd do it myself but I'm not able enough. --WestwoodMatt (talk) 15:41, 20 March 2010 (UTC)Reply

Formally, it is just a measurable space. Informally, it is the set of all possible values (or a larger set, if more convenient). I also wonder, is "observation space" a standard word, a neologism, or what?Boris Tsirelson (talk) 16:30, 20 March 2010 (UTC)Reply
I'm assuming it's the same thing as the image of . But I'm not used to coming at this from the direction of defining the image as a measurable space - the only treatments I'm familiar with regard the image of as just being a subset of . Hence I'm slightly out of my depth, and although I could work through it myself step by step, I'd be unsure as to whether I'd done it right - I'm new to measure theory. --WestwoodMatt (talk) 23:33, 20 March 2010 (UTC)Reply
"Observation space" is used sometimes, see for example here. I did not find the definition, but probably it means the codomain, not just the image. Boris Tsirelson (talk) 07:01, 21 March 2010 (UTC)Reply
An explanation added; please look now. Boris Tsirelson (talk) 10:13, 21 March 2010 (UTC)Reply
I suspect "Observation space" is not standard in the literature, and should be omitted. At least as a PhD student in Stochastic Analysis with decent background in probability and measure theory, I've never seen used before. -- Some random passerby
Also I, an old professor-probabilist, did not see it before (that is before 21 March 2010). Probably, for now it is used mostly by non-mathematicians. But does it mean that it should be omitted? Boris Tsirelson (talk)
I think that the use of non-standard terminology is unfortunate because it makes it harder to read and understand. Especially so in definitions. I also think measure-theoretic definitions are mostly of interest to mathematicians, and thus that terminology should be taken from there. -- The same random passerby.

See also Wikipedia talk:WikiProject Mathematics#Codomain of a random variable: observation space?. Boris Tsirelson (talk) 16:50, 27 March 2010 (UTC)Reply

The term seems self-explanatory when used the way it's used in the article. Someone used the word "unfortunate" above without rigorously defining it. I don't have a problem with that. Michael Hardy (talk) 18:50, 27 March 2010 (UTC)Reply

When you're in a section called "formal definition" it is necessary to define stuff down to this level. Okay then, so although I can make a stab at defining "unfortunate" to a no-mathematically-inclined person (it describes an observation of a random variable whose outcome is contrary to the desires of one's motivational consciousness) I would not be able (in this context) to determine rigorously what an "observation space" is. When something is described as "self-explanatory" I always suspect that this is because the person so describing it lacks the ability to define it. As a logician I can not accept this as an answer. And as a statistician (i.e. I'm not one, I'm just learning this stuff as I go along) I don't understand what an "observation space" is in the terms of the mathematical objects that a probability space is defined in. --WestwoodMatt (talk) 21:35, 27 March 2010 (UTC)Reply
Formal definitions are supposed to be formal. Also, adding definitions not used in the literature is probably a breach of the "no original research" criteria for wikipedia articles.

Enough is enough. Observation space is gone.

Add "mode", "tail", "inflection", etc. to terminology?

[edit]

Shouldn't the basic terminology used to discuss/describe a distribution be included here? It would not only fill out the basic presentation of the material but also provide an anchor for references to such terms in other articles. Jojalozzo 03:02, 25 July 2011 (UTC)Reply

Strange terminology

[edit]

As far as I know, "probability distribution" is, in general, a probability measure (rather than this or that function). In some (but not all) cases it can be described by a cumulative distribution function. Sometimes also by the probability mass function; sometimes also by the probability density function. But the lead says "a probability mass, probability density, or probability distribution is a function..." Or is it meant that a measure is also a kind of function (namely, a set function)? But no, this is not written in the sections. Boris Tsirelson (talk) 05:57, 2 July 2012 (UTC)Reply

I have rewritten the lead to try to overcome this problem. But the rest of the article is extremely short of anything understandable about probability measure. Melcombe (talk) 19:20, 3 July 2012 (UTC)Reply
Nice! Much better than before. (But it would be enough, to link "probability measure" in the lead only once.) Boris Tsirelson (talk) 21:09, 3 July 2012 (UTC)Reply

Normal distribution

[edit]

Why does this article refer to the Gaussian distribution as "the most important distribution"? Isn't this kind of arbitrary? — Preceding unsigned comment added by 2607:4000:200:13:1A03:73FF:FEB3:B07C (talk) 05:46, 27 August 2012 (UTC)Reply

Good question but no it's not arbitrary: there's a theorem somewhere that says that given a large enough sample size, all distributions tend towards the Gaussian in the limit. --Matt Westwood 09:46, 27 August 2012 (UTC)Reply
A quote:
"Gaussian random variables and processes always played a central role in the probability theory and statistics. The modern theory of Gaussian measures combines methods from probability theory, analysis, geometry and topology and is closely connected with diverse applications in functional analysis, statistical physics, quantum field theory, financial mathematics and other areas."
R. Latala, "On some inequalities for Gaussian measures". Proceedings of the International Congress of Mathematicians (2002), 813-822. arXiv:math.PR/0304343.
Boris Tsirelson (talk) 12:41, 27 August 2012 (UTC)Reply
An example: The normal distribution in Rn is the unique (up to scaling) rotation-invariant probability measure with independent components. This purely mathematical result is fundamental to the Kinetic theory of gases. --Rainald62 (talk) 21:36, 29 April 2013 (UTC)Reply
Yes... You mean Maxwell–Boltzmann distribution. Boris Tsirelson (talk) 05:49, 30 April 2013 (UTC)Reply
What User:WestwoodMatt did say, is a nice rule of thumb, but not true in its full generality. There are many flavors of the central limit theorem, which pose conditions like e. g. pairwise independence and identical distribution of the random variables (or freeness and identical distribution of the non-commutative random variables) and especially something like $IE[X^2] < \infty$$. Thus, for example the Cauchy distribution does not fulfill the presuppositions of the central limit theorem. Nevertheless, most of the comon distributions tend to the Gaussian one.
Note in addition that the Gaussian distribution (of non-commutative random variables) is the uniquely determined distribution with all moments up to the first and second one being 0. Isn't it cool? --Mathelerner (talk) 10:14, 29 August 2022 (UTC)Reply

generic statistical distributions (especially sample distributions)

[edit]

Right now there seem to be somewhat strange redirects. What article should I link to for a plain old distribution of occurrences as observed in a finite sample?

Nanite (talk) 18:45, 21 January 2014 (UTC)Reply

(Non-)Zero Probabilities for Continuous random variables

[edit]

I believe "In contrast, when a random variable takes values from a continuum, probabilities can be nonzero only if they refer to intervals" is incorrect. Imagine a random variable whose image is [0,1]. Let the point 1 have a 50% chance of occurring, and the rest of the probability mass is uniformly distributed amongst [0,1). Is there anything wrong with this counter-example? Note that obviously this counter-example extends to letting any finite number of points of the image of a continuous random variable have non-zero probability. 38.88.227.194 (talk) 17:23, 23 September 2014 (UTC)Reply

Right. But that phrase occurs in Introduction, and probably is not meant to be a theorem. Maybe "In contrast, when a random variable takes values from a continuum then, typically, probabilities can be nonzero only if they refer to intervals"? Boris Tsirelson (talk) 20:23, 23 September 2014 (UTC)Reply
"For example, consider measuring the weight of a piece of ham in the supermarket, and assume the scale has many digits of precision. The probability that it weighs exactly 500 g is zero, as it will most likely have some non-zero decimal digits". If the ham is zero-probable to weigh exactly 500g, it will definitely (not 'most likely') have non-zero digits. And wouldn't the scale need to assume infinite (not just 'many') digits? Freddie Orrell (talk) 22:42, 1 June 2023 (UTC)Reply

Probability in function spaces

[edit]

I feel that this article is incomplete because it deals only with 'random variables' (scalar or not). However, the notion of a 'random element' can be applied to any set. Specifically, a probability measure on a functional space is a well defined concept under the Kolmogorov measure theoretic point of view. Nevertheless, there are several items on the article that does not apply easly to the case of random functions. For exemple, the concept of a probability density for random functions does not exist ([1]); the concept of cumulative distribution function is replaced, in most presentations, by a infinite hierarchy of cumulative distribution functions; etc. Crodrigue1 (talk) 21:36, 12 October 2016 (UTC)Reply

True. This is shortly mentioned in "Kolmogorov definition" section. This is more often called "probability measure" than "probability distribution". And, specifically for function spaces, this is usually treated in "Random processes". Boris Tsirelson (talk) 06:40, 13 October 2016 (UTC)Reply

References

  1. Novikov, 1965; Functionals and the random force method in turbulence theory; Soviet Physics JETP, 20(5), pp. 1290 - 1294

Mode: a doubt

[edit]

See Talk:Mode (statistics)#Different treatment of discrete and continuous? Boris Tsirelson (talk) 18:58, 26 June 2017 (UTC)Reply

Continuous

[edit]

There is some confusion about the term 'continuous r.v.'. A continuous r.v. is in this article defined as having a continuous cumulative distribution function, hence it doesn't need to have a density. In the article probability density function it is said a density belongs to a continuous r.v. Madyno (talk) 08:49, 23 July 2017 (UTC)Reply

My suggestion

[edit]

I would suggest to combine this page with Probability_density_function, because a probability density function is simply the mathematical representation of a given probability function. Mimigdal (talk) 17:13, 31 January 2018 (UTC)Reply

Not quite so; "Mathematicians call distributions with probability density functions absolutely continuous"; not all distributions are absolutely continuous; some are discrete, and some are singular. Boris Tsirelson (talk) 18:06, 31 January 2018 (UTC)Reply

Terminology: "Measure theoretic" or "Kolmogorov"?

[edit]

Why "Measure theoretic formulation" for "Discrete probability distribution", but "Kolmogorov definition" for "Continuous probability distribution"? Boris Tsirelson (talk) 07:42, 10 May 2019 (UTC)Reply

= Question on "General Definition" -Section

[edit]

Could someone please check this section?

Maybe I am somewhat confused, but I think what is stated as axioms of Kolmogorov here is wrong. Condition 2 has to state that the probability to be in the whole space equals $1$.

The conditions given here are, for instance, satisfied by the trivial assignment $P(X\in A)$ identical to $0$ for all considered events $A$.  Preceding unsigned comment added by 2003:DC:DF37:9300:7D0B:9E28:D2F3:91D8 (talk) 17:59, 21 March 2021 (UTC)Reply

A quibble on the examples of random phenomena in the opening paragraph?

[edit]

Giving the the fraction of male students in a school as an example of a random phenomenon only makes sense in a coeducational school; otherwise the fractions are fixed! Wprlh (talk) 05:08, 1 July 2021 (UTC)Reply

Inconsistency between die and dice?

[edit]

For example, the second paragraph of the Introduction has both “throwing a fair die” and “the dice rolls” Wprlh (talk) 05:15, 1 July 2021 (UTC)Reply

Continuous probability distribution

[edit]

The definition given of a continuous RV is actually of an absolutely continuous RV and is potentially confusing. There can exist an RV for which the CDF is continuous in the classical sense of continuity but is not absolutely continuous. Hence what is called continuous RV should be called as an absolutely continuous RV. This is explicitly done in Valentin V. Petrov's book, 'Limit theorems of probability theory: sequences of independent random variables' on Page 2:

The distribution of the random variable is said to be continuous if for any finite or countable set of points of the real line. It is said to be absolutely continuous if for all Borel sets of Lebesgue measure zero.

Theorem 31.7 in Billingsley's classic book 'Probablity and Measure' proves that the above definition is equivalent to being absolutely continuous. That book does not mention continuous RV at all, it only mentions absolutely continuous distribution function. Rosenthal's book 'A First Look at Rigorous Probability Theory' explicitly talks about absolutely continuous RV and not about continuous RV (pg 1,70). In conclusion, the cited source (ie Ross' book) is sowing some confusion when it uses the term continuous RV.

Hence for this reason, I am being bold and replacing continuous with absolutely continuous. I have added the sentence 'Some authors however use the term "continuous distribution" to denote all distributions whose cumulative distribution function is absolutely continuous, i.e. refer to absolutely continuous distributions as continuous distributions.' to make this confusion (hopefully) clear to the reader. - Abdul Muhsy talk 05:29, 18 March 2022 (UTC)Reply

General probability definition - a circular definition?

[edit]

In the passage starting

"The concept of probability function is made more rigorous by defining it as the element of a probability space ..."

"it" stands for "probability function". But if this is a definiton for the "probability function" then it contains the definiendum in the definiens by

"... and P is the probability function, ...",

i.e. it is a circular definition. Jyyb (talk) 09:18, 19 July 2022 (UTC)Reply

[edit]

In this edit MrOllie removed a link to an open source probability distribution site with an edit summary pointing to Links normally to be avoided. None of the 19 criteria on that list matched so I'd like to learn why the link was removed. (I didn't originally add it, I just verified that it seemed not to fit the list before moving it to External links). Johnjbarton (talk) 00:14, 25 February 2024 (UTC)Reply

WP:ELNO #4, since this was being systematically added across many articles by a single purpose editor. - MrOllie (talk) 03:30, 25 February 2024 (UTC)Reply
Thanks! Johnjbarton (talk) 03:51, 25 February 2024 (UTC)Reply

Figure 1

[edit]

Isn't the caption wrong? I think it should say: "The RIGHT graph shows a probability density function. The LEFT graph shows the cumulative distribution function." 99.190.32.88 (talk) 21:51, 22 January 2025 (UTC)Reply

I disagree. If we imagine chopping the graph in to 100 equal vertical slices then accumulating then left to right, the last 10 steps will tell the tale. The last ten steps of the left hand side are all small: accumulating them will have a small effect as we see at the right hand end of the right and side. On the other hand, the last ten steps of the right side are close to 1: accumulating them will have a big effect something we do not see on the left side. The caption looks correcct. Johnjbarton (talk) 02:16, 23 January 2025 (UTC)Reply

Probability measure, probability function

[edit]

The section "General probability definition" starts with an impenetrably dense and seemingly contradictory math paragraph. On the one hand it relates "probability distribution" to some kind of transformation of a probability measure. Then it says "Any probability distribution is a probability measure..." followed by saying that measure differs from the one earlier in the paragraph.

I don't think this paragraph is helpful to anyone who is not already an expert in the topic. Specifically it does not address the plain and simple question of the relationship among probability distribution, probability measure, and probability function. By simple I mean a qualitative and useful conceptual separation, rather than one that requires knowing the topic in advance. Johnjbarton (talk) 17:55, 6 May 2025 (UTC)Reply

Hi @Johnjbarton,
I agree that this paragraph could be improved, and to me the terminology used is a bit weird; but there nothing wrong / no contradiction. So before discussing how to improve the paragraph, let me clarify the terminology:
  • probability distribution and probability measure are synonyms; "probability law" is yet another one. These terms all refer to a measure with total mass 1.
  • in practice, the words distribution and law are mostly used to refer to probability measures induced by random variables. Say that variable is X: we then talk about the distribution of X (or, more rarely, the law of X).
  • The reason why I said that the terminology a bit weird is that the probability distribution of X is not something I hear or read too often: usually people simply talk about the distribution of X. Note however that it is very natural to talk about probability distributions if there is nothing else indicating what kind of "distributions" we are considering (as in "list of probability distributions").
  • probability function is not something I would ever expect to hear or read; the standard terminology is probability mass function. I also quite frequently hear PMF — even in languages other than English.
I hope that clarifies "the relationship among probability distribution, probability measure, and probability function".
Now the question is: how to improve the article (1) without getting into irrelevant details about the subtleties of the terminology (2) while remaining mathematically correct. Unfortunately, the distribution of a random variable is indeed the push-forward of the probability measure of the space on which that random variable is defined... And in order to be precise one has to introduce that probability space. I am convinced there are ways to be less technical, though; for instance, one could write something "the probability measure μ induced by X on X(Ω), through μ(A) = Pr(XA)", instead of "the pushforward of Pr by X, that is, ".
Malparti (talk) 21:21, 6 May 2025 (UTC)Reply
@Malparti Thanks! Do you have any sources that describe any two of these together?
AFAIK, "measure" in the math sense was not discussed in normal scientific math books. It's not very complex but since it's not described "These terms all refer to a measure with total mass 1." means nothing to me. And "mass" has other meanings ;-).
Similarly "induced" is a word that must come naturally to mathematicians but is unclear to me. Using this technical word in a definition fails at the goal of definition.
Your point about "probability distribution" makes sense to me. Maybe this article would be better called "Probability distributions" (explicitly plural) with a focus on common characteristics, differences, and applications. Then Probability measure would handle the generalized math and be summarized here. If, as you claim, these are synonyms we don't need both, right?
Probability function is used in
  • Dekking, F. M. (2005). A Modern Introduction to Probability and Statistics: Understanding why and how. Springer Science & Business Media.
As far as I can tell this source is describing a set function with properties similar (identical?) to probability measure, but does not use any similar terminology. I think this part of the issue: non-math people have different terminology because they are trying to use the concepts rather than prove their properties.
I would really like to have a source that explains these are similar/same. Johnjbarton (talk) 22:29, 6 May 2025 (UTC)Reply
@Johnjbarton
"Do you have any sources that describe any two of these together?"
→ I am not sure what you are asking for, but in any introductory math textbook on probability theory you'll find:
  1. A (more or less detailed) exposition of measure theory that will at the very least recall the following points of vocabulary: σ-algebra, measurable space, measurable function, measure, measure space. And most likely pushforward measure.
  2. Something along the lines of "a probability space is a measure space such that ; the measure is then called a probability measure".
  3. Something along the lines of "a random variable on a probability space is a measurable function from to some measurable space ; the pushforward measure of by is called the distribution of .
Were you looking for specific textbook recommendations? For that, I'd need to know what your goal is. If it is merely to understand the terms, honestly I think Wikipedia is your best option. :)
"Similarly "induced" is a word that must come naturally to mathematicians but is unclear to me. Using this technical word in a definition fails at the goal of definition."
induced is a non-technical word that is used whenever there is a "natural" way to define an object from another: for instance, if you have a random variable ­— which by definition comes with a measure space and a measurable space — then there is a very natural way to define a probability measure on . So, in this example, we can say that induces a measure on . However this is merely a non-rigorous description of things meant to carry the general idea, without getting into technical details. In other words: induces is a short word for defines in a canonical way; its exact meaning is going to depend on the context.
"Probability function is used in [...]"
→ Ah, but see — that's precisely why this term should not be used: it is confusing. In the text you are mentioning, the term probability function does not refer to a probability mass function; it refers to a probability measure.
"I would really like to have a source that explains these are similar/same."
→ Do you mean for yourself, or for the article? In case that is for yourself: to understand the difference between a probability mass function and a probability measure, and why we need probability measures and not just probability mass function, one has to:
  1. First, get familiar with the basic of probability in finite sample spaces: understand that we can assign a "weight" to each face of a die, and use those weights to compute expected values, etc. Those weights define a probability mass function, and that's pretty much all we need to do probability theory on finite sample spaces. Provided you know about series, all of that carries over to countable sample spaces — and, again, all you need is the probability mass function.
  2. After this, spend a bit of time thinking how to sample at random in an uncountable space. Once you start thinking about this, you'll realize that the probability mass function that was so useful in the discrete case is now completely useless; and that this is a really hard problem. After thinking a bit about it, you should realize that a good first step would be being able to assign length to subsets of , but that even that is really hard.
  3. If your name is Emile Borel (or, allegedly, Alexandre Grothendieck), you can end up solving this problem on your own. But most of us learn the answer by following a course on measure theory. ;) The bottom line is that we need to define collections of "measurable sets" known as sigma-algebras, and non-negative functions on those sets known as measures. Of course, all of that also works in countable spaces:
  • If you have a countable sample space and a probability mass function , then you get a probability measure by setting
  • If you have a probability measure a sample space , you get a probability mass function by setting . However, unless your sample space is countable, this probability mass function is going to be completely useless since you'll have for all but at most a countable number of 's; so contains no information about .
Malparti (talk) 00:15, 7 May 2025 (UTC)Reply
Thanks again!
  • I come into this topic from physics. I am not against math definitions, but "General probability definition" is opaque and unsourced. With sources I could learn more. So you could say my goal is learnging by adding sources to wikipedia ;-) Specific textbook recommendations would be great, preferably Cambridge University Press or similarly available via wikipedia library.
  • Your description of what I might read in a math textbook is great and I generally follow it. However, strictly speaking it is "off topic" for this article because no where is "probability distribution" mentioned. It seems like a discussion about probability measure.
  • Induced. Perhaps you consider it non-technical because it is not formally defined, but for a non-mathematician I'll assert it is technical: it takes you a paragraph to explain it ;-)
  • Probability function. Perhaps it should not be used but in fact it is used. And not just in the source but in this article 7 times. With a source we could say "probability function can be is a synonym for probability measure, but the terminology is easily confused with probability mass function".
  • Sources. I think we can agree I'm not Emile Borel ;-) But your description of why probability measures vs probability mass functions is exactly the kind of content I'm looking for. Understanding the purpose of a concept (probability measure) is much more valuable than a formal definition. (I add a lot of science history for this reason).
  • Here is another example of the the kind of content that I think makes a huge difference for the same reason:
    • The concept of independence was an important invention in probability theory. It shaped the theory at an early stage and is considered one of the main features specifying the place of probability theory within more general measure theory. Suhov, Y., & Kelbert, M. (2014). Probability and statistics by example: Volume 1, basic probability and statistics. Cambridge University Press. pg 18
  • Now can formulate a specific suggestion for this article. But I will start it in a new thread.
Johnjbarton (talk) 01:20, 7 May 2025 (UTC)Reply
Based on this great feedback, here is a proposal to improve the article:
  1. Rename "General probability definition" to "Definition". The current title is confusing.
  2. Delete the first, highly technical paragraph, or perhaps move it to a section "Formal definition"
  3. Move the section "Basic terms" to the start of Definition, add and subtract to include enough terms to define "probability distribution" Now that I see this section I understand that "probability function" is a synonym for "probability measure".
  4. Change "probability function" -> "probability measure" after the terminology.
However this all assumes that "Probability distribution" is really a thing different from Probability measure. If these are not different, the entire section should be reconsidered and presented as a summary of the other article instead. Johnjbarton (talk) 01:40, 7 May 2025 (UTC)Reply
Two quick remarks:
  1. The terminology probability function is not standard among probabilists and statisticians, and when it is used it usually refers to the probability mass function. Of course you'll also find people who use it to refer to a probability measure, or to the probability density -- I'm pretty sure if you search enough you'll find someone somewhere who use it to refer to random variables, or to the expected value -- see my third point.
  2. Probability distribution and probability measure are indeed synonym — in the sense that they refer to the same mathematical object. However, there tends to be a difference in usage — namely, probability measure (or simply probability) is more common when talking about a probability measure on an abstract space; whereas probability distribution is more common when referring to a probability measure on a concrete space: most people will endow a measurable space with a probability measure; but say that the normal distribution is a probability distribution on . Another way to say this is that the more likely a probability measure is to correspond to the probability distribution of a random variable of interest, the more likely it is to be called a probability distribution.
  3. The more widely used a subject is, the more variation you are going to find in the terminology. Probability theory is mostly being used and written about by people with very different background and who aren’t mathematicians. As someone interested in applied probability, I see all sorts of terminologies, all the time — that doesn't make them "standard" / worth mentioning on Wikipedia.
Malparti (talk) 02:59, 7 May 2025 (UTC)Reply
  • "Probability distribution and probability measure are indeed synonym" Just FWIW, my brain sees "distribution" as a 2D object with height and width, but for "measure" I see a function or mapping which only becomes a "distribution" after visiting its domain. We speak of the "width" of a normal distribution. This goes back to your previous comment that "probability distribution" is more of a category than a mathematical definition. (I also understand your view that "probability distribution" is a concrete instance of the abstraction "probability measure".)
  • Terminology. I think it is clear we disagree on this point. Readers like me will come to Wikipedia to learn about "probability function" because that term is clearly used and can be sourced. We need to say and source that it is often used as equivalent to "probability measure." Or "probability". Or "probability distribution". Or mistaken for "probability mass distribution". But we can't just ignore it because it is inconvenient.
I guess I will have to set this aside as I don't have access to sources now. Johnjbarton (talk) 16:41, 7 May 2025 (UTC)Reply
[edit]

In the second paragraph of the lead, I recently wikilinked "coin toss" to coin flipping and "the probability distribution of X" to Bernoulli distribution, but these links were removed with reference to MOS:OVERLINK. It seems to me that MOS:UNDERLINK is more applicable—the Bernoulli distribution is a very basic example of a probability distribution, and coin flipping is the basic example of a probabilistic experiment. Hence, "Relevant connections [...] that help readers understand the article more fully". What part of MOS:OL can make the case against these links? —St.Nerol (talk, contribs) 19:50, 16 September 2025 (UTC)Reply

The part of OVERLINK that makes the case against those links is "making it difficult to identify those likely to aid a reader's understanding" + the fact that those links do not actually help understand the article.
  • The article Coin flipping is mostly concerned about the history of coin-flipping, its use in conflict resolution, in politics, etc — none of which are relevant here. Coin flipping is, precisely, used as a self-contained example. By adding a link on it, you distract the reader because you suggest that there are some aspects of coin-flipping that should be mastered before going on with what follows. At this rate, you might as well link "0.5" to One half or "outcome" to Outcome (probability).
  • The same goes for the link the the Bernoulli distribution: at this stage it doesn't help to know what this specific distribution is called. Also, that specific link was misleading: a link on "the probability distribution of X" should point to the probability distribution of a random variables. Links should be explicit, not implicit.
Malparti (talk) 20:41, 16 September 2025 (UTC)Reply
In a paragraph about coin flipping and the Bernoulli distribution, I have a hard time seeing how linking the key terms would be distracting and could not help understanding the article. Of course, it's not the case that the reader must click and read every link before reading on; see MOS:NOFORCELINK. It's just that these are relevant terms of the sort we usually link, since a reader 'may want to read more about the main thing presently being discussed. The argument that at this stage it doesn't help to know what this specific distribution is called assumes to much—it sounds like the reader is a blank slate and we are curriculum planning. —St.Nerol (talk, contribs) 09:08, 17 September 2025 (UTC)Reply
I think the link to coin toss is ok in "the outcome of a coin toss" This term maybe cultural and the words have a meaning completely different from their literal interpretation.
I don't think "the probability distribution of X " should ever be linked in probability distribution, that's just confusing. Johnjbarton (talk) 15:23, 17 September 2025 (UTC)Reply
@Johnjbarton That "coin toss" could mean something completely different for English-speaking readers with different cultural background seems unlikely to me — but that is indeed a possibility. So this is a good argument.
I still think that linking everyday expressions that are most likely crystal-clear to 99% of readers if a bit silly; but if the argument for adding that link is "better safe than sorry — for all we know, some people might interpret 'coin toss' as asking a jinn to make their wish come true or as hurling a coin at someone to hurt them", then I'll gladly accept the consensus. Malparti (talk) 15:55, 17 September 2025 (UTC)Reply
My argument is still that it's "particularly relevant to the context" (MOS:OL), not that it's likely to be misunderstood. I guess we can have a pragmatic consensus with different individual reasons, though it's not optimal. —St.Nerol (talk, contribs) 17:44, 17 September 2025 (UTC)Reply
And I still disagree with your argument. But, as you say, here we can have a practical consensus — so there probably isn't much sense in prolonging this discussion: (1) I can't express my position much more clearly than "links are supposed to help readers understand the article and/or provide relevant background information, none of what's in the article 'Coin flipping' achieves that" and (2) you're not going to convince me that it would be a "good" thing for someone trying to understand what a probability distribution is to go read "Coin flipping".
Cheers, Malparti (talk) 18:40, 17 September 2025 (UTC)Reply

Maybe correct but not useful

[edit]

@Asdfaqwefvgrtwrewytre The opening sentence is now

This opening is bad for readers: it introduces a topic which many readers without deep mathematical training could understand using terminology way over our heads. Per It gives the basics in a nutshell, introduces the article, and cultivates interest in reading on—though not by teasing the reader or hinting at what follows. It should be written in a clear, accessible style with a neutral point of view., this sentence fails. The opening sentence is not the place for a technically correct jargonized definition. There is plenty of space for that later. Johnjbarton (talk) 00:20, 23 December 2025 (UTC)Reply

@Johnjbarton: You raise good points. Using the overly technical term "measure" in the first sentence may confuse beginners. I made an effort here to clarify my edits and to compromise.
@Malparti: reverted the opening sentence to
This revert contains no technical mistakes, but is not specific enough. It sounds like a description of a probability measure, not a probability distribution. I wanted to clarify the difference from the beginning. A probability distribution is defined in reference to a specific random variable. I tried to improve specificity without using the overly technical description of a push forward measure. Also, although a measure is a function (from a sigma algebra), I originally changed the word "function" because beginners often think of a function as being between sets of numbers rather than between more abstract objects; this may mislead beginners into thinking that a random variable's probability distribution is the same thing as its random variable's PDF, PMF, or CDF, which is one of the other main confusions I was trying to clear up in my recent series of edits.
Perhaps a good sentence which captures how a probability distribution differs from other mathematical objects (like probability measures, PDFs, PMFs, and CDFs) would look something more like
This moves the jargon "measure" to an easily-skippable parenthetical remark. It also references the relation of a probability distribution to its random variable while leaving a more specific technical discussion for later. Asdfaqwefvgrtwrewytre (talk) 20:17, 23 December 2025 (UTC)Reply
@Asdfaqwefvgrtwrewytre the distinction between "probability measure" and "probability distribution" isn't fundamental: those are the exact same object. It is true that probabilists tend to make a distinction between the two terms — in particular, given a probability space , the pushforward of by a random variable is always going to be referred to as its distribution (or as its law); and never as its "measure" (by the way, I believe this has already been discussed extensively on this talk page).
However, the term "distribution" isn't strictly reserved to this — a lot of authors, including probabilists, often use "probability distribution" to refer to a probability measure that is not associated with any specific random variable. This is especially true when the space in question is a simple space such as or .
Finally, and more importantly, probability theory doesn't belong to probabilists, and I would think that outside of probability theory the term "probability distribution" is almost always used... Most people reading this article probably aren't probabilists, and that is something to take into account. Which is why I think it is better not to confuse people with a somewhat artificial distinction that is meant to reflect subtle differences in how probabilists use terms (and not a fundamental mathematical difference).
By the way, thanks for your work on the article. :)
Cheers, Malparti (talk) 21:06, 23 December 2025 (UTC)Reply
Based solely on this discussion and the fact that "probability measure" appears multiple times in the article, maybe a sentence further down in the introduction along the lines of "Mathematically a probability distribution is related to a probability measure..." would appropriate. This way the introductory nature of the first lines won't be interrupted. Johnjbarton (talk) 23:36, 23 December 2025 (UTC)Reply
@Malparti: You're right that the distinction is mostly terminological. The main question I have is: given that separate articles exist for probability distribution and probability measure, how should the lead of this article acknowledge that separation in a way a layperson understands and an expert accepts? I don't think most readers will go to the talk page to understand the relationship between the two articles. (I only recently learned that the talk page even exists.)
I propose changing the first sentence (not by much) to something like:
Later in the lead, I can add something along the lines of:
The distinction between a probability distribution and a probability measure is one of emphasis: a probability distribution is often associated with a random variable (cite Durrett).
(Durrett has a pdf of his book up for free on his university's website, so its content is verifiable to more editors. So I'll probably cite him.) Asdfaqwefvgrtwrewytre (talk) 01:01, 24 December 2025 (UTC)Reply
@Asdfaqwefvgrtwrewytre, a few quick replies:
  • "given that separate articles exist for probability distribution and probability measure" → the way I understand it, the reason for having two separate articles is that they serve different purposes:
  1. Probability distribution is the term that is the most likely to be searched, and as such that article aims is aimed at a wide audience with no specific background. The emphasis should be on probability measures on or , on how to characterize them (cdf, pmf, generating functions, etc) and on concrete examples.
  2. Probability measure is aimed at readers who have already heard about measure theory, i.e. a much more restricted audience. This is the right place to talk about general properties of probability measures (for instance, this is where I would expect to basic results such as the fact that if two probability measures agree on a pi-system that generates the sigma-algebra, then they agree on the sigma-algebra)
In other words, what justifies having two articles is that people with very different backgrounds (and different needs) want to know about probability measures/distributions; not the fact that there is a subtle difference between how probabilists use the terms.
  • "I don't think most readers will go to the talk page to understand the relationship between the two articles." → that is not the goal of the talk page anyway: the talk pages are here for editors to discuss how to improve articles; not to help readers understand them.
  • "I propose changing the first sentence" → I agree that the link between probability distributions and random variables should be mentioned early on in the lead; in fact, that's what is done in the second paragraph, where a random variable X is introduced. The only minor comment I have is that I'm not a big fan of the phrasing "and is usually associated with a random variable": for you and I, what is meant here is clear; but for someone who has no idea what this is all about, this can be confusing: what do you mean it is usually associated with a random variable? Associated in what way? Unfortunately, I do not have a good alternative to suggest. That's because I think the current first sentences of the lead are a "not perfect, but decent" solution to the very tricky problem of how to be mathematically correct without being overly precise.
  • On a related note, concerning the paragraph you added to the lead that starts with "A random variable's probability distribution is not the same as its cumulative distribution function, its probability mass function, or its probability density function, which are distinct mathematical objects useful for describing a probability distribution.": I agree that at some point, that distinction becomes important. But I'm not sure it is what the lead should focus on (especially since many authors, probabilists included, will happily write "let be a probability distribution on ". In my opinion, what the reader should understand from the lead is "Ever heard of CDF, PMF or density? That's what we are talking about!" — not "You'd have to be dumb to think a probability distribution is the same thing as a probability density". So I might tweak your sentence a bit to put less emphasis on the fact that those things are different and more emphasis on the fact that they are related.
  • Durrett is a great reference, please go ahead and use it. (PS: I'm not Durrett and have no interest in promoting his work ;))
Malparti (talk) 10:59, 24 December 2025 (UTC)Reply
I understand.
I appreciate the clarification of the differences between the articles. Probability measure assumes a measure theory background; probability distribution is aimed at a more general audience.
I think that the lead should clarify this difference. That way, readers can quickly decide which article(s) they prefer to read. Asdfaqwefvgrtwrewytre (talk) 00:26, 25 December 2025 (UTC)Reply
  • "Probability measure assumes a measure theory background; probability distribution is aimed at a more general audience." → At least that's my understanding... But please note that many editors insist that every article should be aimed at a broad audience (I think this point of view is not widely shared among math editors, though).
  • "I think that the lead should clarify this difference." → I agree, though I don't know what's the best way to do it. Maybe a sentence like "In probability theory, probability distributions are represented by probability measures, and the term distribution is often used in reference to probability measures associated with random variables". But one should probably think a bit more about the phrasing, where to put this in the article, etc.
Cheers, Malparti (talk) 12:32, 25 December 2025 (UTC)Reply
I like your sentence. I really like the usage of the word "represented" because it suggests that probability distributions have a status as their own things which may go beyond the bounds of their formalization. This justifies the maintenance of two articles and is philosophically cool. I would just add the word "probability" before "distribution" for clarity. I suggest adding your modified sentence to the last paragraph of the lead: after the sentence on CDFs, PMFs, and PDFs and before the one on probability distributions with specific names. @Johnjbarton: What do you think? Asdfaqwefvgrtwrewytre (talk) 20:05, 25 December 2025 (UTC)Reply
I agree, with probability measures wikilinked. Johnjbarton (talk) 23:05, 25 December 2025 (UTC)Reply

First Two Sentences

[edit]

Currently, the article's first two sentences are:

"In probability theory and statistics, a probability distribution is a function that gives the probabilities of occurrence of possible events for an experiment. It is a mathematical description of a random phenomenon in terms of its sample space and the probabilities of events (subsets of the sample space)."

From the comments of previous contributors, I can see that this language is much improved over earlier renderings. However, the current language remains problematic because a probability distribution is not a mathematical function. Rather, it is a mathematical model or description of the outcome of a random process that typically is formalized by a function such as a cumulative distribution function, probability mass function, or probability density function.

As an alternative, I'd suggest replacing the first two sentences with the following:

"In probability theory and statistics, a probability distribution is a mathematical description of the random outcome of an experiment or naturally occurring phenomenon. It is defined in terms of both the sample space of all possible outcomes and the probability that any given event (i.e., subset of the sample space) is realized."

As a newcomer, I would like to hear other opinions before proceeding with a revision. Pdfs&Pmfs (talk) 19:05, 24 April 2026 (UTC)Reply

Thanks. Given that we've been around this bush a few times I think the appropriate thing is to insure that a Definition is included containing sourced content. If one definition appears across sources, we can work on a summary. If several reliable sources give conflicting definitions we'll have to work out how to summarize them in the intro. Definitions in math and physics may not agree, IDK.
The Introduction section has essential no sources. Much of the rest of the article is too formal to be useful for an introductory definition.
The current content is sourced to
  • Everitt, B. S., Skrondal, A. (2010). The Cambridge Dictionary of Statistics. United Kingdom: Cambridge University Press.
and it says
  • Probability distribution: For a discrete random variable, a mathematical formula that gives the probability of each value of the variable. See, for example, binomial distribution and Poisson distribution. For a continuous random variable, a curve described by a mathematical formula which specifies, by way of areas under the curve, the probability that the variable falls within a particular interval. Examples include the normal distribution and the exponential distribution. In both cases the term probability density may also be used.
Is there really a difference between a "formula" and a "function"?
The second source
  • Ash, R. B. (2008). Basic Probability Theory. United States: Dover Publications.
does not seem to define "probability distribution" directly. It should be removed, but this is a sign that we may have difficulty in finding an agreed definition. At the bottom of page 95 it talks about "some way of calculating...", which is consistent with "formula" above. Johnjbarton (talk) 20:43, 24 April 2026 (UTC)Reply
@Pdfs&Pmfs "However, the current language remains problematic because a probability distribution is not a mathematical function." → you are mistaken. A probability distribution is a function — namely, a function that maps certain subsets of the sample space to their probability. The other functions you describe (CDF, PMF, PDF) are merely alternative ways to describe that function (in the case of real-valued random variables for the CDF; in the case of discrete random variables for the PMF; in very specific cases for the PDF).
You are also slightly mistaken when you say "event (i.e., subset of the sample space)": although events are indeed subsets of the sample space, not all subsets of the sample space are events. To be fair, the current version is also slightly misleading because it can be interpreted as meaning that "event" is synonymous with "subset of the sample space".
Malparti (talk) 01:24, 25 April 2026 (UTC)Reply
I agree that the phrase "probability distribution" is sometimes used to refer to a specific function. However, that typically occurs in the context of discrete distributions, where "probability distribution" is used synonymously with "probability mass function" (e.g., one might say "Consider the Bernoulli probability distribution, f(x) = p^x * (1-p)^(1-x) for x in {0,1}"). On the whole, I believe it is much more common to reserve "probability distribution" for references to either: (1) specific probability models or laws, such as "the Bernoulli probability distribution", "the Normal probability distribution", etc.; or (2) general categories of probability models/laws, such as "discrete probability distributions", "continuous probability distributions", etc.
My principal concern with the current first sentence -- which links directly to Wikipedia's "Function (mathematics)" article -- is that it gives the reader the impression that [probability distribution] = [function], which leads to the natural (and currently unanswered) question: What specific mathematical function are we talking about? Pdfs&Pmfs (talk) 02:54, 25 April 2026 (UTC)Reply
@Pdfs&Pmfs: the function I'm talking about maps events to probabilities. It is neither the CMF (which maps real numbers to probabilities) nor the PMF (which maps elements of a countable sample space to probabilities) nor the PDF (which maps real numbers to non-negative real numbers).
As I said previously, these other functions are just alternative ways that can be used to encode some probability distributions: for instance, on , every probability distribution can be encoded by its CDF; on a countable space, every probability distribution can be encoded by its PMF. On a measure space, a very restricted class of probability distributions can be encoded by a PDF; on , every probability distribution can be encoded by its characteristic function; etc. But some probability distribution cannot be described by any of these functions.
As a simple example, consider a stochastic process such as Brownian motion. Formally, this is a function-valued random variable. How would you describe the distribution of this random variable?
Malparti (talk) 03:41, 25 April 2026 (UTC)Reply
@Pdfs&Pmfs to take your example of the Bernoulli(p) distribution, this function — say — is the following:
Malparti (talk) 03:54, 25 April 2026 (UTC)Reply
@Malparti I think I see your point; but feel free to correct me if I’m wrong. One can view the term “probability distribution” primarily as a synonym for “the probability measure induced by the random variable X”, which is equivalent to “the pushforward probability measure” (i.e., μX = PX-1). Clearly, this measure is a set function (mapping measurable subsets of the random variable’s state space to [0,1]), and the understanding that [probability distribution] = [pushforward probability measure] = [type of function] is indeed consistent with the article’s “General probability definition” section.
My concern is that “probability distribution” also refers to a probability model or probability law more broadly (e.g., “the Bernoulli probability distribution” or “the normal probability distribution”), without referencing any specific functional form (such as P, μX, or the CDF, PMF, PDF, etc.). This latter use forms the basis for the article’s “Introduction” section, and – to my mind – is the more common meaning for general readers familiar with probability theory at the introductory level (where set functions typically are mentioned only in passing, if at all).
To address both of these perspectives, would the language below be adequate?
"In
probability theory
and
statistics
, a
probability distribution
is a mathematical formulation of the random outcome of an experiment or naturally occurring phenomenon. It can be used to mean both: (1) a specific
set function
mapping all
measurable subsets
(i.e.,
events
) from the
sample space
of possible random outcomes to a
probability
in the interval [0,1]; and (2) a probability model or law implied by such a set function, but more commonly described by a
cumulative distribution function
or other function defined on the sample space itself."
Pdfs&Pmfs (talk) 21:09, 25 April 2026 (UTC)Reply
@Pdfs&Pmfs "My concern is that “probability distribution” also refers to a probability model or probability law more broadly" → this meaning is not broader, because every probability law is the distribution of a random variable (namely the identity). Mathematically, "probability measure", "probability distribution" and "probability law" are synonymous (in the sense that they all refer to the same object). Which term you should use depends on the context and on personal taste.
The situation is analogous to what happens with "expected value", "expectation", "first moment" and "mean": these terms all refer to the same thing, but in some contexts most people will tend to favor one of them over the others.
Malparti (talk) 22:09, 25 April 2026 (UTC)Reply
@Malparti I recognize that, in measure-theoretic probability, the term “probability distribution” can be used synonymously with both “probability measure” and “probability law”. I also understand that there is a direct and immediate connection between the “probability distribution” of any random variable X and the unique pushforward measure, μX, a set function on the state space. However, my point is that the term “probability distribution” has more than one usage.
In mathematics, it is very common for a term to denote both a precise mathematical object and, by metonymy, the broader structure, model, or theory organized around that object. For example, a “metric” is literally a distance function (d), but we also use the same term to refer to the geometry or structure induced by d. Similarly, a “differential equation” is literally a displayable equation, but in practice the term often refers to the entire set of ideas and methods associated with that equation, including its relevant domain, boundary conditions, and applicable solution concepts (consider, e.g., “the Bernoulli equation”, “the Riccati equation”, etc.).
“Probability distribution” works the same way. In measure theory, one can equate the “probability distribution” of X specifically with the pushforward measure. But in common mathematical and pedagogical usage (e.g., “the exponential distribution”, “the normal distribution”, etc.) the term refers not only to the specific measure, but also to the associated model or family, together with its various functional representations (e.g., the CDF, PMF, CF, etc.) and familiar theoretical motivations (e.g., the exponential distribution for modeling intervals between Poisson events, the normal distribution as the limiting law in the central limit theorem, etc.).
This second usage of “probability distribution” is not incompatible with the first. Moreover, it is broader – not in the sense that it denotes a more general formal object, but in that it invokes the full structure of representations, properties, and theory built around the formal probability measure, including the measure itself. Most importantly, however, the second usage is likely more familiar to readers, and therefore should be stated at the beginning of the article. Pdfs&Pmfs (talk) 15:12, 27 April 2026 (UTC)Reply
Here are comments on the version you suggested:
  • "In probability theory and statistics, a probability distribution is a mathematical formulation of the random outcome of an experiment or naturally occurring phenomenon."
→ This has two problems: (1) "a mathematical formulation of the random outcome" may be misleading because it suggest that a probability distribution models the realizations of the experiment (as a random variable does), which is not the case: it only describes their statistical properties; (2) "an experiment or naturally occurring phenomenon" is confusing, because contrasting "experiment" and "naturally occurring phenomenon" gives the impression that "experiment" refers to a lab experiment, whereas here the word has a very specific meaning.
  • "It can be used to mean both: (1) a specific set function mapping all measurable subsets (i.e., events) from the sample space of possible random outcomes to a probability in the interval [0,1]; and (2) a probability model or law implied by such a set function, but more commonly described by a cumulative distribution function or other function defined on the sample space itself."
→ The main problem to me is that one should not oppose the two as if those were different meanings, when in reality they are essentially the same thing. In my opinion it is best to say that a probability distribution is a description of certain aspects of random experiment (namely, its statistical properties, as opposed to its outcome), and that the way we formalize this is by using a function that assigns "probabilities" to various outcomes — in other words, a probability measure. In my opinion the current version of the lead is closer to this than the version you suggest.
In addition to this, three other minor problems: (1) I don't think its good to talk about measurable sets in the second sentence of the lead of "wide-audience" article like this one; (2) "a probability model or law implied by such a set function" is vague and, again, law and distribution are perfect synonyms in that context; (3) I'm nitpicking a bit here, but "or other function defined on the sample space itself" is not ideal, because one of the main tools to describe probability distributions in practice are generating / characteristic functions, which are not defined on the sample space itself.
Malparti (talk) 16:15, 27 April 2026 (UTC)Reply
The discussion above is missing the key aspect needed for articles in Wikipedia: sources to ensure verifiability. Unless you plan to get consensus without sources I encourage you to discuss them. Johnjbarton (talk) 16:21, 27 April 2026 (UTC)Reply
Statements in the lead such as basic definitions of objects typically do not need to be sourced (provided they are detailed in the body of the article and sourced there). Also note that I'm not trying to add anything that would require a source to the article: I'm trying to point out what I believe to be problems in a proposed modification.
Now, regarding sources: the article contains a lot of them. Among these is, e.g, Durrett's "Probability: Theory and Examples", which — as I already mentioned on this talk page — is as a classic recommendation for a modern introduction to probability. The definition of a probability measure is on pages 1-2, and on p.10 you'll find "If is a random variable, then induces a probability measure on called its distribution by setting for Borel sets ."
As I already mentioned to you earlier on this very page, you can open any classic probability textbook and you'll find the exact same thing. For instance,
  • on p.32 of Borovkov "Probability Theory", you'll find: "Hence one can define a probability on the measurable space which generates the probability space . [...] The probability [sic] is called the distribution of the random variable ";
  • for Kallenberg's "Fondations of Modern Probability" this is p.83: "The set function is a probability measure on the range space of , called the distribution or law of ."
In fact, right after his definition Kallenberg adds:
"We often use the term distribution as synonymous to probability measure, even when no generating random element has been introduced."
Malparti (talk) 18:23, 27 April 2026 (UTC)Reply
Thanks. The intro does not need citations, but it needs to summarize content that is sourced. Thus the logical starting point is not careful wording of the intro but sourced definitions.
Sourced definitions would allow an editor like myself to verify the content. Unfortunately your sources are vague to me and contradictory. They describe a "distribution", not a "probability distribution". Maybe you think that is trivial but I assume mathematicians typically try to be precise.
  • Durrett is unclear, but I think "its distribution" means "the random variable distribution". Durrett uses "probability distribution" in examples but does not define it as far as I can tell.
  • Borovkov has "the distribution of the random variable"; Kallenberg same.
  • If we assume that "distribution" must be "probability distribution" then the final Kallenberg quote means this article should redirect to probability measure right?
  • Ash, R. B. (2008). Basic Probability Theory. United States: Dover Publications. Again talks about the "distribution function" of a random variable (pg 52) and says that the values of that function are "probabilities".
So none of these sources say "probability distribution". Rather than coming away thinking that these sources verify definitions, I come way wondering. I'm not saying you are wrong. My point is practical: how can we verify content in the encyclopedia?
I have the same issue with the comments of @Pdfs&Pmfs. Both arguments make some sense but don't help with the practical issue. Johnjbarton (talk) 19:14, 27 April 2026 (UTC)Reply
"My point is practical: how can we verify content in the encyclopedia?" → well you could take the time to actually read the sources that I kindly linked for you. Then you'd see that Borovkov explicitly says, at the very beginning of his book (p.17 out of ~700):
"A probability on is also sometimes called a probability distribution on or just a distribution on (on )."
You'd also see that all of the books mentioned above (Borovkov, Kallenberg, Durrett and Ash) use the term "probability distribution" at least once to refer to something which they otherwise call a distribution. Malparti (talk) 23:12, 27 April 2026 (UTC)Reply
Also see https://www.randomservices.org/random/dist/index.html for a clear definition, together with a list of classic textbooks. Malparti (talk) 23:55, 27 April 2026 (UTC)Reply
I assumed you already provided the appropriate info for the topic. Johnjbarton (talk) 02:23, 28 April 2026 (UTC)Reply
To the point about different, looser definitions, I looked around for less rigorous sources which I outline below. I like the last one a lot.
This source:
  • Evans, M. J., Rosenthal, J. S. (2004). Probability and Statistics: The Science of Uncertainty. United Kingdom: W. H. Freeman.
is yet another source that does not define "probability distribution" but uses it in sentences like this one about Maxwell (pg 80)
  • The discovery that velocities and speeds of the individual particles followed a certain type of probability distribution enabled him to describe many of the basic physical properties of gases...
In the Glossary of that source it refers to normal distribution and Poisson distribution as "curves used to estimate the probability of events for a certain types of random processes".
This source has a page called "What is a Probability Distribution",
  • Guthrie, William F. (2020), NIST/SEMATECH e-Handbook of Statistical Methods (NIST Handbook 151), doi:10.18434/M32189, retrieved 2026-04-27
and yet it never defines exactly that term. It talks about "discrete probability function" and a "continuous probability function".
This glossary
has two definitions
  • The assignment of a probability to the possible outcomes (realizations) of a random variable.
  • A function that assigns a probability to each measurable subset of the possible outcomes of a random variable.
This source:
Kruschke, John K. (2015). "What is This Stuff Called Probability?". Doing Bayesian Data Analysis. Elsevier. doi:10.1016/b978-0-12-405888-0.00004-0. ISBN 978-0-12-405888-0.
at least provides a definition.
  • A probability distribution is simply a list of all possible outcomes and their corresponding probabilities
This source
  • Wackerly, Dennis; Mendenhall, William; Scheaffer, Richard L. (2008). Mathematical Statistics with Applications (7th ed.). Belmont, CA: Duxbury Press. ISBN 978-0-495-11081-1.
  • ...the probability for each value of a random variable. This collection of probabilities is called the probability distribution
This one is great:
Devore, Jay L. (2016). Probability and Statistics for Engineering and the Sciences (9th ed.). Boston, MA: Cengage Learning. ISBN 978-1-305-25180-9.
The probability distribution of X says out the total probability of 1 is distributed among (allocated to) the various possible X values. Johnjbarton (talk) 23:11, 27 April 2026 (UTC)Reply
All of these definitions except for "a function that assigns a probability to each measurable subset of the possible outcomes of a random variable" only make sense when there is a countable number of outcomes, and therefore do not apply to, e.g, the normal distribution — which is arguably both the most well-known and the most important example of probability distribution there is. Malparti (talk) 23:54, 27 April 2026 (UTC)Reply
Well that definition is useless for our purposes since "measurable subset" isn't a commonly known thing. Almost all of these sources split into discrete / continuous quickly. Maybe we simply need to have two or more defintions. Johnjbarton (talk) 02:28, 28 April 2026 (UTC)Reply
@Johnjbarton Applied probability/statistics texts provide examples of the "broader" usage. For example, the language below is copied from pp. 108-109 of
Elston, Robert C. and Johnson, William D. (2008) Basic Biostatistics for Geneticists and Epidemiologists. Chichester, UK: John Wiley & Sons. ISBN: 978-0-470-02489-8
"Here we shall be concerned with probability distributions. A probability distribution is a model for a random variable, describing the way the probability is distributed among the possible values the random variable can take on. … Mathematically, the concepts of ‘probability distribution’ and ‘random variable’ are interrelated, in that each implies the existence of the other; …. If we know the probability distribution of a random variable, we have at our disposal information that can be extremely useful in studying the patterns and tendencies of data associated with that random variable. Many mathematical models have been developed to describe the different shapes a probability distribution can have. One broad class of such models has been developed for discrete random variables; a second class is for continuous random variables." Pdfs&Pmfs (talk) 03:20, 28 April 2026 (UTC)Reply
Yes, that one is also good. It includes "distributed among" like Devore, which I think helps cement the meaning in a compact way. Johnjbarton (talk) 15:13, 28 April 2026 (UTC)Reply
@Johnjbarton the fact that it's too technical doesn't make it "useless", it just means it has to be explained. The problem of having a common definition for countable vs uncountable spaces has been solved a long time ago: describe the probability that the random variable falls various regions of its state space, rather that the probability that it takes on a specific value; not to have two definitions.
In a textbook, it makes sense to start with a countable (or even finite) state space, because (1) that makes it easier to start (2) there are already a lot of interesting things to say in that context and (3) there is a great pedagogical value in understanding why the "naive approach" does not work for uncountable state spaces, and then seeing how that problem can be "fixed". However, Wikipedia is not a textbook.
@Pdfs&Pmfs To me the specific passage you copied is not "broader" than the standard definition, it's just a loose, non-formal description of it. I do not have an issue with that text: I cannot see anything clearly wrong / confusing, and to me it is at about the right level of discourse for the beginning of a textbook called "Basic Biostatistics for Geneticists and Epidemiologists". So we can reuse some of these ideas in the lead (I never said I'm opposed to changing the lead, I'm just opposed to modifications that would suggest that there are two notions of a "probability distribution", when in reality there is just one clear concept — "the way the probability is distributed among the possible values [a] random variable can take on" to quote the book you mentioned — and different ways to formalize/encode/manipulate that information. However, also note that the passage you quoted (1) is not a definition and that (2) Wikipedia is not a textbook, and the lead has to describe the subject more precisely than what would typically be done in the first few paragraphs of a textbook.
Malparti (talk) 06:16, 28 April 2026 (UTC)Reply
Wikipedia is not a textbook but it also not a math reference work. As an encyclopedia we have made a decision that wikipedia articles should be written for the widest possible general audience. Specifically there is no need for precision in the intro, but rather a need for understandable, plain language per the page linked.
A formal, and even extensive, mathematical definition is suitable for the body of the article, where all of the technical terms can be explained. In the intro that can be boiled down and expressed along the lines of "in mathematical treatments probability distributions are equivalent to probability measures" (I'm not proposing language, just the character of the level of detail). Broader, vaguer, easier to follow definitions can also be included and described as such, eg "In applied statistics, the concept of a probability distribution is described as ...". We can include as many definitions as are broadly represented in sources, but my sense now is that these two cases would do. Johnjbarton (talk) 15:30, 28 April 2026 (UTC)Reply

This is a common problem. There is a choice between initially giving a precisely correct definition or explanation of a subject which is basically unintelligible to a person not familiar with the field, and then bring them up to speed with the details and concepts necessary for them to understand it, or, alternatively, to give an incomplete, oversimplified explanation which is intuitively accessible to a person not familiar with the field, with warnings that it is imprecise, and then bring them up to speed by repairing the incompleteness and over-simplification of the initial definition. People already familiar with the field tend to favor the first approach, while people who are not, favor the second. I favor the second approach, which is more or less a "teaching" approach and is accessible to a larger audience, yet ultimately does not do violence to the precise definition of the subject. I understand, people whose eyes glaze over as the details begin to be discussed and walk away, will only leave with a little imperfect knowledge which can be a bad thing, but those who read the first sentence and throw up their hands and walk away will have no knowledge, and that can also be a bad thing. Also, I understand that this violates the extremely strict interpretation of "Wikipedia is not a textbook" (WINAT), but one has to ask oneself, do I wish to be strictly correct while being unintelligible to everyone except a fellow expert, preaching to the choir so to speak, or do I wish to impart knowledge as per the prime directive? If it's the latter, you have to lighten up on the WINAT. PAR (talk) 09:12, 28 April 2026 (UTC)Reply

@PAR I actually also think this article should start with an intuitive and not entirely rigorous description of the topic, as long as it doesn't convey common misconceptions and it correct enough that it doesn't leave someone who knows a bit about the topic wondering "What the heck is that?". Having that in mind,
  • such descriptions exist -- the description from "Basic Biostatistics for Geneticists and Epidemiologists" quoted by @Pdfs&Pmfs is an example, although to me it's better suited for a textbook than for Wikipedia because it is a bit too verbose and misses some important information (namely, that knowng the distribution of a random variable means we know the probability that it falls in any "relevant" region of the sample space), so I'm not necessarily advocating that we go for this — but this is to give an example.
  • one common misconception is the idea that a probability distribution is simply gives you the probability of all outcomes of the random variable. That is something which is true in some basic cases, but not in cases encountered in practice, even in applied maths or in other sciences. This is also something that many 1st-to 3rd-year college students routinely get confused about.
  • in my opinion, the current version of the lead — although it could be improved a lot — is still better than anything that has been suggested here so far. I'm not advocating that we should keep this version at all cost: if you re-read my participation in this discussion from the start (which I do not encourage you to do: there are better things to do in life), you'll see that I've mostly tried to dismiss common misconceptions, from "a probability distribution is not a mathematical function" to "a probability distribution maps outcomes of a random variable to probabilities".
Malparti (talk) 13:01, 28 April 2026 (UTC)Reply
I think we actually agree about @Pdfs&Pmfs suggested "Basic Biostatistics for Geneticists and Epidemiologists" as a source for the intro definition.
Can you say more these what you mean by these "misconceptions"?
  • "probability distribution is simply gives you the probability of all outcomes of the random variable"
  • "a probability distribution is not a mathematical function".
  • "a probability distribution maps outcomes of a random variable to probabilities".
What is an "outcome of a random variable"? Is it just "possible value of a random variable"? In physics we would use "distribution" to mean the observed outcomes of a process believed to have random aspects and we would use "probability distribution" to also mean a function, formula or model that predicts those outcomes. So the words "gives you" and "maps" is puzzling here. A probability distribution is a model, a function that predicts, not a process that generates. Johnjbarton (talk) 15:48, 28 April 2026 (UTC)Reply
  • "probability distribution [is simply] the probability of all outcomes of the random variable"
→ This is a misconception because if you know the probabilities for all in the sample space, you do not necessarily know the distribution of . Thus, a probability distribution contains strictly more information than the probabilities of all outcomes.
  • "a probability distribution is not a mathematical function"
→ This is a misconception because a probability distribution is, in fact, a function. This is true in two senses: (1) there is a mathematical object called a "probability distribution", and that object is a function and (2) the information contained in that object coincides exactly with what is meant in loose speech when talking about a "probability distribution".
  • "a probability distribution maps outcomes of a random variable to probabilities"
→ This is correct; however these probabilities often carry no information. So this is not a misconception: the misconception is that a probability distribution is a map "outcome probability".
  • "What is an "outcome of a random variable"? Is it just "possible value of a random variable"?"
→ Yes, at least that what I see some people use it for. A more standard term is "realization".
Malparti (talk) 21:18, 28 April 2026 (UTC)Reply
Thank you! Johnjbarton (talk) 23:38, 28 April 2026 (UTC)Reply