Skip to content
Kudos AI
Lire en français
Probability Foundations

Probability from Zero: The Language of Uncertainty

Build probability from the ground up: possible worlds, the sample space, the two basic axioms, and the addition and multiplication rules, each derived rather than asserted, with worked numeric examples.

7 min readKudos AI
Thirty-six ordered pairs laid out as a grid, the complement rule derived in two lines, and the King of Hearts caught being counted twice.

Almost every idea on this site eventually rests on a single sentence: we do not know what will happen, but we can say how likely each outcome is. This article makes that sentence precise. We build probability from two axioms, derive the familiar rules rather than assuming them, and finish with the tools you need for every later article in this track.

Nothing here is taken on faith. Each rule below is derived from the two axioms, so by the end you will know not just what the formulas say but why they must be true.

A. Possible worlds and the sample space

Russell & Norvig frame probability in terms of possible worlds. A probabilistic assertion says how likely each world is, where a logical assertion would say only which worlds are ruled out.

The set of all possible worlds is the sample space, written Ω\Omega (uppercase omega). An individual world is written ω\omega. Two properties define it:

  • The worlds are mutually exclusive - two cannot both be the case.
  • The worlds are exhaustive - one of them must be the case.

For a single roll of an ordinary die, the sample space is

Ω={1,2,3,4,5,6}.\Omega = \{1, 2, 3, 4, 5, 6\}.

For a roll of two distinguishable dice there are 36 possible worlds: (1,1),(1,2),…,(6,6)(1,1), (1,2), \dots, (6,6). This is worth pausing on, because the count is where beginners most often go wrong: the worlds are ordered pairs, so (2,5)(2,5) and (5,2)(5,2) are two different worlds, not one.

A probability model assigns a number P(ω)P(\omega) to every possible world.

B. The two axioms

Everything follows from two requirements on that assignment:

0≤P(ω)≤1for every ω,and∑ω∈ΩP(ω)=1.0 \le P(\omega) \le 1 \quad \text{for every } \omega, \qquad\text{and}\qquad \sum_{\omega \in \Omega} P(\omega) = 1 .

In words: no probability is negative or greater than one, and the probabilities of all the possible worlds add up to exactly one - something must happen.

That is the whole foundation. These axioms trace to Kolmogorov's Foundations of the Theory of Probability (1950), and every formula in the rest of this article is a consequence of them.

If the two dice are fair and do not interfere with each other, symmetry forces each of the 36 worlds to carry the same probability, and since they must sum to 1, each is 1/361/36.

C. Events: from worlds to propositions

We rarely care about a single world. We care about a proposition such as "the roll is even". An event is the set of worlds where the proposition holds, and its probability is the sum of their probabilities:

P(a)=∑ω∈aP(ω).P(a) = \sum_{\omega \in a} P(\omega).

For one fair die and a=a = "even" ={2,4,6}= \{2, 4, 6\}:

P(a)=16+16+16=36=0.5.P(a) = \tfrac{1}{6} + \tfrac{1}{6} + \tfrac{1}{6} = \tfrac{3}{6} = 0.5 .

Why "favourable over total" is a special case, not the definition. Counting outcomes and dividing works only when every world is equally likely. A loaded die still has a perfectly good probability model; it simply assigns unequal P(ω)P(\omega). The summation above is the real definition and always applies.

D. Deriving the complement rule

Here is the machinery working. The worlds where aa holds and the worlds where ¬a\neg a ("not aa") holds are disjoint, and together they are all of Ω\Omega. So splitting the second axiom's sum:

∑ω∈aP(ω)  +  ∑ω∈¬aP(ω)=∑ω∈ΩP(ω)=1,\sum_{\omega \in a} P(\omega) \;+\; \sum_{\omega \in \neg a} P(\omega) = \sum_{\omega \in \Omega} P(\omega) = 1 ,

which is exactly

P(a)+P(¬a)=1⟹P(¬a)=1−P(a).P(a) + P(\neg a) = 1 \quad\Longrightarrow\quad P(\neg a) = 1 - P(a).

If a model gives an event probability 0.050.05, its non-occurrence has probability 0.950.95 - and now you know that as a theorem, not an intuition.

E. The addition rule, and the trap it avoids

For two events that may overlap, adding probabilities double-counts the shared worlds. Removing the overlap exactly once gives

P(a or b)=P(a)+P(b)−P(a and b).P(a \text{ or } b) = P(a) + P(b) - P(a \text{ and } b).

Worked example. Draw one card from a standard 52-card deck. Let aa = "it is a heart" and bb = "it is a King".

  • P(a)=13/52P(a) = 13/52 - thirteen hearts.
  • P(b)=4/52P(b) = 4/52 - four Kings.
  • P(a and b)=1/52P(a \text{ and } b) = 1/52 - exactly one King of Hearts.
P(a or b)=1352+452−152=1652=413≈0.3077.P(a \text{ or } b) = \frac{13}{52} + \frac{4}{52} - \frac{1}{52} = \frac{16}{52} = \frac{4}{13} \approx 0.3077 .

Naïvely adding would have given 17/52≈0.326917/52 \approx 0.3269, counting the King of Hearts twice. The overlap term is not a technicality; it is the difference between a right and a wrong answer.

Interactive: count the worlds, then count the overlap once

Every world is equally likely.

♥♦♣♠A2345678910JQKA♥2♥3♥4♥5♥6♥7♥8♥9♥10♥J♥Q♥K♥A♦2♦3♦4♦5♦6♦7♦8♦9♦10♦J♦Q♦K♦A♣2♣3♣4♣5♣6♣7♣8♣9♣10♣J♣Q♣K♣A♠2♠3♠4♠5♠6♠7♠8♠9♠10♠J♠Q♠K♠
A onlyB onlyboth: counted twice by the naive sumneither
Event A
Event B
P(A)
13/52 = 0.2500
P(B)
4/52 = 0.0769
P(A and B)
1/52 = 0.0192
P(A or B)
16/52 = 0.3077
Naive P(A) + P(B)
17/52 = 0.3269
P(A) × P(B)
0.0192

Counting the worlds directly, A or B holds in 16 of 52. The addition rule gets there from the other three counts: 13 + 4 - 1 = 16, so P(A or B) = 4/13. The naive sum counts 17, and the extra 1 is exactly the highlighted overlap, counted once for A and again for B. These two events are also independent: P(A and B) equals P(A) × P(B), so knowing one tells you nothing about the other.

F. Conditional probability

Learning something changes what we should expect. The probability of aa given that bb holds is defined as

P(a∣b)=P(a and b)P(b),P(b)>0.P(a \mid b) = \frac{P(a \text{ and } b)}{P(b)}, \qquad P(b) > 0 .

The intuition: restrict attention to the worlds where bb holds, then ask what fraction of that restricted world-set also has aa. The division by P(b)P(b) is what renormalises the restricted set back to total probability 1.

Rearranging gives the multiplication rule:

P(a and b)=P(b) P(a∣b).P(a \text{ and } b) = P(b)\, P(a \mid b).

Worked example. In a population, 10% of members belong to a high-risk group. Within that group, an event occurs with probability 20%. The probability that a randomly chosen member is both high-risk and experiences the event is

P(high risk and event)=0.10×0.20=0.02,P(\text{high risk and event}) = 0.10 \times 0.20 = 0.02 ,

that is, 2% of the whole population.

G. Independence, and why it must be earned

Two events are independent when knowing one tells you nothing about the other - formally, when P(a∣b)=P(a)P(a \mid b) = P(a). Substituting into the multiplication rule collapses it to the version most people remember:

P(a and b)=P(a) P(b)(independent events only).P(a \text{ and } b) = P(a)\,P(b) \qquad \text{(independent events only).}

For two genuinely unrelated events each of probability 0.030.03, the chance both occur is 0.03×0.03=0.00090.03 \times 0.03 = 0.0009, or 0.09%0.09\%.

The most expensive mistake in applied probability. Multiplying probabilities is valid only under independence. When a common cause drives many outcomes at once - a shared failure mode, a market-wide shock, a single upstream dependency - the events are strongly dependent, and multiplying wildly understates the chance that many happen together. Independence is an assumption to be checked against the data-generating process, never assumed for free.

H. Verifying the die-and-card arithmetic

The numbers above are small enough to check by exhaustive enumeration, which is worth doing once so you trust the rules rather than the arithmetic:

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Running it prints 16 16 4/13: the union counted directly and the union computed by the addition rule agree exactly, and the probability is the 4/134/13 derived above.

Key takeaways

  • A probability model is a sample space Ω\Omega of mutually exclusive, exhaustive possible worlds, each carrying a probability.
  • The two axioms are that probabilities lie in [0,1][0, 1] and that they sum to 1 over Ω\Omega. Everything else is derived.
  • The probability of an event is the sum over the worlds where it holds; this is the definition, and "favourable over total" is only its equally-likely special case.
  • "Or" needs the addition rule with the overlap subtracted; "and" needs the multiplication rule with a conditional probability.
  • Independence is the special case where conditioning changes nothing. It must be verified, never assumed.

What's next

Conditional probability has a consequence far more useful than it first appears: it can be reversed, letting you turn "how the evidence behaves given a cause" into "how likely the cause is given the evidence". That reversal is Bayes' theorem, and it produces results that reliably surprise people. It is the subject of Bayes' Theorem and Belief Updating.

References & further reading

  • Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

6 min readProbability Foundations

Bayes' Theorem and Belief Updating

Derive Bayes' theorem from the definition of conditional probability, then work the base-rate example that fools almost everyone, twice: once with the formula and once by pure counting.

ProbabilityMathematicsArtificial Intelligence
7 min readStatistical Learning Foundations

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

StatisticsMachine LearningMathematics
5 min readProbabilistic Reasoning

The Week That Cannot Have Happened

Take the most likely state on each day and write them down in order, and you have a report the model assigns probability exactly zero: on a four-day machine-monitoring example the day-by-day answer is healthy, healthy, failed, failed, and healthy to failed is a transition that cannot occur. What the two questions actually are, why smoothing and Viterbi answer different ones, and what the 0.411 posterior on the best path means for anyone who has to act on it.

Artificial IntelligenceProbability
← Back to all articles