Learn probability distributions with 23 worked examples and exercises, clear formulas, tables, and binomial, Poisson, and normal graphs.

Probability Distributions: Guide and Solved Exercises

Probability Distributions: Guide and Solved Exercises

PROBABILITY • STEP-BY-STEP STUDY GUIDE

Probability Distributions: Guide and Solved Exercises

Probability distributions turn uncertainty into numbers you can calculate, compare, and explain. For example, they help you estimate defective output, count incoming calls, or measure the chance that an observation falls within a range. This guide builds the subject from simple events to binomial, Poisson, and normal models. Along the way, you will work through every numbered example and exercise in the supplied pages.

First, you will learn how to describe outcomes and choose the correct probability rule. Next, you will solve the 14 introductory examples and exercises 3.1 through 3.9. Finally, you will compare the models, review common mistakes, and apply the ideas to business decisions. Each solution explains the assumptions as well as the arithmetic.

Probability foundations: outcomes, events, and uncertainty

An experiment produces an outcome. For example, rolling a die produces one of six face values. The sample space contains all possible outcomes, while an event collects the outcomes that meet a condition. Thus, “an even result” corresponds to the event {2, 4, 6}.

A probability ranges from zero to one. In a finite model, zero describes an impossible event and one describes certainty. However, continuous models need a more careful interpretation: a single exact value can have probability zero without being an impossible numerical value.

0 ≤ P(A) ≤ 1    and    P(S) = 1

Classical probability and equally likely outcomes

When a finite sample space contains equally likely outcomes, count the outcomes that satisfy the event and divide by the total. Therefore, a fair die assigns the same probability to each face. A biased die requires different probabilities, so simple counting alone no longer gives the answer.

P(A) = Number of favorable outcomesTotal number of equally likely outcomes = nAN

For example, the event “even result” contains three of the six die faces. Consequently, its probability equals 3/6, or 0.5. The word favorable means that an outcome belongs to the event; it does not mean that the outcome is desirable.

Empirical probability and subjective probability

Empirical probability uses observed results. For instance, 27 defective items among 1,000 inspected items produce an estimated defect probability of 0.027. However, that estimate describes the observed sample and may not perfectly represent future production.

Subjective probability expresses a degree of belief based on available evidence. For example, an analyst might assess a new product’s chance of meeting a sales target before comparable data exist. Nevertheless, a useful assessment should state its assumptions and change when new evidence arrives.

A compact notation guide

SymbolMeaningHow to read it
SSample spaceAll possible outcomes
AcComplement of AA does not occur
A ∪ BUnionA or B, including both
A ∩ BIntersectionBoth A and B
P(B | A)Conditional probabilityProbability of B given A
XRandom variableNumerical outcome of interest
μ and σMean and standard deviationCenter and spread
p̂Observed proportionEstimated probability from data

The complement rule

An event and its complement cover the entire sample space without overlap. Therefore, their probabilities add to one. This rule often makes a calculation shorter, especially when the question asks for “at least one” or “not.”

P(Ac) = 1 − P(A)

For example, the complement of “at least one defect” is “no defects.” Consequently, subtracting the probability of no defects from one gives the probability of at least one. Always define the event before choosing its complement.

Addition, multiplication, and conditional probability

Use addition for “or”

The general addition rule combines two events and removes their overlap. Otherwise, outcomes that belong to both events enter the calculation twice. For example, a king of spades belongs to both the spade event and the king event.

P(A ∪ B) = P(A) + P(B) − P(A ∩ B)

Mutually exclusive events cannot occur together in the same experiment. Therefore, their intersection has probability zero, and the addition rule simplifies. Rolling a two and rolling a three on one die roll provide a straightforward example.

If A ∩ B = ∅:   P(A ∪ B) = P(A) + P(B)

Use conditional probability for new information

A conditional probability restricts attention to outcomes consistent with known information. For instance, removing a king from a deck changes both the remaining king count and the deck size. As a result, the second draw requires an updated probability.

P(B | A) = P(A ∩ B)P(A)   for P(A) > 0

Use multiplication for “and”

The general multiplication rule works with dependent or independent events. First, calculate the probability of the initial event. Then multiply by the probability of the next event conditional on the first.

P(A ∩ B) = P(A) × P(B | A)

Independent events satisfy P(A ∩ B) = P(A)P(B). Thus, when P(A) is positive, learning that A occurred leaves the probability of B unchanged. Independent coin tosses illustrate this relationship, provided the experiment genuinely supports that assumption.

For independent events:   P(A ∩ B) = P(A) × P(B)

Further reading: OpenStax: Two Basic Rules of Probability.

Why mutually exclusive does not mean independent

Mutual exclusivity describes whether events can occur together. In contrast, independence describes whether information about one event changes the probability of another. These questions differ, so the terms cannot replace each other.

If two mutually exclusive events both have positive probability, they must be dependent. For example, knowing that one die roll produced a two rules out a three. However, the general statement needs a zero-probability exception: disjoint events can be independent if at least one has probability zero.

Further reading: OpenStax: Independent and Mutually Exclusive Events.

Solved Examples 1–8: basic probability rules

Example 1: one fair coin toss

Question: Find the probabilities of heads and tails.

First, list the two outcomes: H and T. Because the coin is fair, each outcome receives half the total probability.

P(H) = 1/2;   P(T) = 1/2;   P(H ∪ T) = 1

Answer: Heads: 50%; tails: 50%; heads or tails: 100%. These results assume the usual two-outcome model.

Example 2: one fair die roll

Question: Find the probability of each face and the probability of not rolling a one.

There are six equally likely outcomes. Therefore, any specified face has probability 1/6. Next, use the complement rule to count every face except one.

P(j) = 1/6 for j = 1,…,6;   P(not 1) = 1 − 1/6 = 5/6

Answer: Each face: approximately 16.6667%; not one: approximately 83.3333%.

Example 3: a jack and a non-diamond

Question: Draw one card from a standard 52-card deck without jokers. Find the chance of a jack and the chance of a card that is not a diamond.

A standard deck contains four jacks and 13 diamonds. Thus, 39 cards are not diamonds. Treat these as separate questions rather than as a joint event.

P(J) = 4/52 = 1/13;   P(Dc) = 1 − 13/52 = 3/4

Answer: Jack: approximately 7.6923%; non-diamond: 75%.

Example 4: 53 heads in 100 tosses

Question: Compare the observed frequency of heads with the theoretical probability for a fair coin.

First, divide 53 by 100 to obtain the observed proportion. Then compare that result with the fair-coin probability of 0.5. Random variation can create a difference even when the model is correct.

p̂ = 53/100 = 0.53;   p̂ − p = 0.53 − 0.50 = 0.03

Answer: The observed frequency is 53%, which is 3 percentage points above 50%. More tosses do not guarantee that every successive estimate moves closer to 50%.

Example 5: several possible die faces

Question: Find P(2 or 3) and P(2 or 3 or 4) on one fair die roll.

These face events do not overlap. Therefore, add their individual probabilities, or count the relevant faces directly.

P(2 or 3) = 2/6 = 1/3
P(2 or 3 or 4) = 3/6 = 1/2

Answer: The probabilities are approximately 33.3333% and exactly 50%, respectively.

Example 6: a spade or a king

Question: Find the probability that one card is a spade or a king.

First, count 13 spades and four kings. However, the king of spades appears in both groups. Subtract that one-card overlap before dividing by 52.

P(S ∪ K) = 13/52 + 4/52 − 1/52 = 16/52 = 4/13

Answer: Approximately 30.7692%. Here, “or” includes the king of spades.

Example 7: consecutive heads

Question: Find the probabilities of two heads in two tosses and three heads in three tosses of a fair coin.

Assume the tosses are independent. Consequently, multiply one factor of 1/2 for each required head.

P(HH) = (1/2)2 = 1/4
P(HHH) = (1/2)3 = 1/8

Answer: Two heads: 25%; three heads: 12.5%. These events specify heads on every toss, not merely a head somewhere in the sequence.

Example 8: a specified king followed by another king

Question: Draw the king of diamonds first, then another king, without replacement. What is the joint probability?

The first event has probability 1/52. Next, after that card leaves the deck, three kings remain among 51 cards. Therefore, multiply 1/52 by 3/51.

P(KD then K) = 152 × 351 = 1884

Answer: Approximately 0.00113122, or 0.113122%. The conditional second-draw probability alone is 3/51, or 5.88235%.

Notice the difference between a specified first king and any first king. If the question asked for any two kings, the first factor would be 4/52. Consequently, that different event would have probability 1/221, four times the answer above.

Understanding probability distributions

A random variable assigns a number to an outcome. For example, two coin tosses produce sequences such as HT and HH, while X can count the number of heads. A probability distribution then describes how probability spreads across the possible values of X.

Discrete probability distributions

A discrete random variable has a finite or countably infinite set of possible values. Thus, a binomial count has finitely many values, while a Poisson count can take any nonnegative integer. The word discrete does not require a finite list.

p(x) = P(X = x) ≥ 0    and    ∑x p(x) = 1

For a discrete variable, the probability mass function gives the probability at each value. Meanwhile, the cumulative distribution function adds all probabilities at or below a threshold. Therefore, F(2) means P(X ≤ 2), rather than P(X = 2).

Continuous probability distributions

A continuous model uses a probability density function. Instead of adding isolated point probabilities, calculate areas over intervals. Consequently, a curve’s height is a density, while the area beneath the curve represents probability.

P(a ≤ X ≤ b) = ∫ab f(x) dx

For an absolutely continuous distribution, P(X = a) = 0. Therefore, including or excluding a single endpoint does not change an interval probability. A density can exceed one on a narrow interval, but its total area must equal one.

Expected value and spread

The expected value is a probability-weighted average. For example, a count can have an expectation of 43.2 even though each observed count is an integer. Variance measures squared deviations from that mean, while standard deviation returns the spread to the original units.

E(X) = ∑x x p(x)
Var(X) = ∑x (x − μ)2p(x);   σ = √Var(X)

Binomial probability distributions: counting successes

Use a binomial model when you count successes in a fixed number of independent trials, each with the same success probability. Every trial must have two categories: success and failure. However, “success” simply names the category you count, so a defective item can count as a success mathematically.

X ∼ Binomial(n, p)
P(X = k) = n!k!(n − k)! pk(1 − p)n − k

The factorial term counts the arrangements that contain exactly k successes. For example, four heads in six tosses can occur in several different orders. Therefore, the formula includes both the probability of one arrangement and the number of eligible arrangements.

μ = np;   σ² = np(1 − p);   σ = √[np(1 − p)]

Further reading: NIST/SEMATECH: Binomial Distribution.

Example 9: heads in two tosses

Question: Construct the probability distribution for the number of heads in two independent fair tosses.

First, enumerate TT, TH, HT, and HH. Each sequence has probability 1/4. However, two sequences produce exactly one head, so that count has twice the probability of either endpoint.

P(X = 0) = 1/4;   P(X = 1) = 2/4;   P(X = 2) = 1/4

Answer: The probabilities are 0.25, 0.50, and 0.25. Their sum equals one.

Heads, xEligible sequencesP(X = x)F(x)
0TT0.250.25
1TH, HT0.500.75
2HH0.251.00

As an additional check, the weighted mean equals 0(0.25) + 1(0.50) + 2(0.25) = 1. Similarly, the binomial formula gives np = 2(0.5) = 1. Both methods describe the same distribution.

Example 10: four heads in six tosses

Question: Find the probability of exactly four heads in six independent fair tosses, plus the mean and standard deviation.

Set n = 6, k = 4, and p = 0.5. Next, calculate the number of arrangements: 6!/(4!2!) = 15. Each six-toss sequence has probability 1/64.

P(X = 4) = 15(1/2)4(1/2)2 = 15/64 = 0.234375
μ = 6(0.5) = 3;   σ = √1.5 ≈ 1.224745

Answer: Exactly four heads: 23.4375%; expected heads: 3; standard deviation: approximately 1.224745 heads.

For comparison, “at least four heads” includes four, five, and six. Therefore, its probability is (15 + 6 + 1)/64 = 22/64 = 0.34375. The word exactly changes the event, even though the model stays the same.

Heads0123456
Probability1/646/6415/6420/6415/646/641/64
2026-10-09T03:56:33.626452 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 0 1 2 Number of heads 0.0 0.1 0.2 0.3 0.4 0.5 Probability 2 fair coin tosses 0 1 2 3 4 5 6 Number of heads 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Probability 6 fair coin tosses
Binomial probability distributions for two and six fair tosses. Orange highlights exactly four heads in six tosses.

Poisson probability distributions: counts within an interval

A Poisson model describes a count in a specified interval. For example, the interval may represent one hour, one kilometer, or one inspected surface. The parameter λ gives the expected count for that interval, so changing the interval requires changing λ.

X ∼ Poisson(λ);   P(X = k) = e−λλkk!;   k = 0, 1, 2, …
μ = λ;   σ² = λ;   σ = √λ

For a homogeneous Poisson process, counts in disjoint intervals are independent and the rate stays constant. However, clustered arrivals or changing demand may violate those assumptions. Check the process before treating a count as Poisson merely because it measures events.

Further reading: NIST/SEMATECH: Poisson Distribution.

Example 11: two calls in an hour

Question: A department averages five calls per hour. Under a Poisson model, find the probability of exactly two calls in one hour.

The interval is one hour, so λ = 5. Next, substitute k = 2 and use 2! = 2. Keep the exponential term unrounded until the final step.

P(X = 2) = e−55²2! ≈ 0.0842243375

Answer: Approximately 8.4224%. Using the book’s rounded e−5 ≈ 0.00674 gives 0.08425 instead.

For a half-hour interval, the expected count becomes 2.5 under the same constant-rate assumption. Consequently, you would use λ = 2.5 for that different question. The mean of five calls per hour does not mean that every hour contains five calls.

2026-10-09T03:56:33.778953 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 Calls in one hour 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 Probability Poisson model: five calls per hour
Poisson probabilities with λ = 5. Orange marks two calls; the chart shows counts 0–15, with a small remaining tail beyond 15.

Normal probability distributions: areas under a curve

A normal distribution is symmetric and bell-shaped. The mean μ locates its center, while the positive standard deviation σ controls its spread. Unlike a discrete mass, a normal probability comes from an area between values.

X ∼ N(μ, σ²);   f(x) = 1σ√(2π) e−(x − μ)²/(2σ²)

Standardization expresses distance from the mean in standard deviations. Thus, z = 2 means two standard deviations above the mean, while z = −1 means one below it. For normal X, the standardized variable has a standard normal distribution.

Z = X − μσ;   Z ∼ N(0, 1)

Further reading: NIST/SEMATECH: Normal Distribution.

Read the correct type of normal table

Some tables give the cumulative area Φ(z) = P(Z ≤ z). Others report the area between zero and a positive z-score. Therefore, check the table heading before using its numbers. A value near 0.4750 at z = 1.96 refers to the mean-to-z area, whereas the cumulative area is approximately 0.9750.

P(a < X < b) = Φ((b − μ)/σ) − Φ((a − μ)/σ)

Example 12: the area from zero to 1.96

Question: Find P(0 < Z < 1.96) for a standard normal variable.

First, obtain the cumulative probability at 1.96. Then subtract the 0.5 area to the left of zero. By symmetry, the interval from −1.96 to zero has the same probability.

P(0 < Z < 1.96) = Φ(1.96) − 0.5 ≈ 0.475002

Answer: Approximately 47.5002%, or 47.50% using a four-decimal table.

Example 13: an interval from 8 to 12

Question: Let X be normal with μ = 10 and variance σ² = 4. Find P(8 < X < 12).

First, take the square root of the variance: σ = 2. Next, standardize both bounds to obtain −1 and 1. Finally, subtract the cumulative probabilities.

z₁ = (8 − 10)/2 = −1;   z₂ = (12 − 10)/2 = 1
P(8 < X < 12) = Φ(1) − Φ(−1) ≈ 0.682689

Answer: Approximately 68.2689%. The source’s rounded table calculation gives 68.26%.

Example 14: an interval from 7 to 14

Question: Using the same normal model, find P(7 < X < 14) and the probability outside that interval.

The standard deviation remains two. Therefore, the lower and upper z-scores are −1.5 and 2. Next, subtract the cumulative probabilities, then use the complement for the two outside tails.

P(7 < X < 14) = Φ(2) − Φ(−1.5) ≈ 0.910443
P(X ≤ 7 or X ≥ 14) = 1 − 0.910443 ≈ 0.089557

Answer: Inside: approximately 91.0443%; outside: approximately 8.9557%. Four-decimal table arithmetic gives 91.04% and 8.96%.

2026-10-09T03:56:33.867745 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 4 6 7 8 10 12 14 16 Value of X 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 Probability density Shaded probability 91.0443% Normal model: mean 10, standard deviation 2
The shaded area from 7 to 14 equals approximately 0.910443 under N(10, 4). Curve height measures density, not probability.

The 68–95–99.7 rule

For a normal distribution, about 68.27% of values fall within one standard deviation of the mean. Similarly, about 95.45% fall within two and 99.73% within three. These are normal-model probabilities, so do not apply them automatically to skewed or heavy-tailed data.

IntervalStandard normal eventProbability
μ ± σ−1 ≤ Z ≤ 168.2689%
μ ± 2σ−2 ≤ Z ≤ 295.4500%
μ ± 3σ−3 ≤ Z ≤ 399.7300%

Complete exercise solutions: Problems 3.1–3.9

The following solutions retain the source’s exercise numbers so you can compare each result with the photographs. However, the explanations use fresh wording and show the reasoning explicitly. Treat each subpart as its own event unless the question requests a joint probability.

Problem 3.1: three approaches to probability

Part (a): Classical probability assigns probabilities through a model of equally likely outcomes. Empirical probability estimates them from observed relative frequencies. In contrast, subjective probability expresses a degree of belief informed by evidence and judgment.

ApproachIllustrationMain limitation
ClassicalA fair die gives P(3) = 1/6.Equal likelihood requires justification.
Empirical106 threes in 600 rolls gives 106/600.Sampling noise and changing conditions affect estimates.
SubjectiveAn analyst estimates the chance of meeting a launch target.Judgment can reflect bias or incomplete information.

Part (b): The classical method becomes difficult when outcomes lack symmetry or known probabilities. Meanwhile, empirical estimates require relevant observations and can change across samples. Subjective estimates can differ across people, so transparent assumptions and calibration matter.

Part (c): We study probability to make decisions under uncertainty. For example, a production manager can estimate expected waste, while a service manager can plan staffing. Probability also supports statistical inference by linking observed data to possible underlying processes.

The classical method is not limited to games. For instance, a genuinely uniform random selection from a finite list also supports equally likely counting. Likewise, empirical probability need not converge to a classical fair-device value if the underlying device is biased.

Problem 3.2: coin, die, and complements

Part (a): A fair coin has two equally likely outcomes. Therefore, heads and tails each have probability 1/2. Their union covers every outcome in the model.

P(H) = 1/2;   P(T) = 1/2;   P(H or T) = 1

Part (b): A two occupies one of the six die faces. In contrast, “not two” includes five faces. Adding the event and its complement returns the full sample space.

P(2) = 1/6;   P(not 2) = 5/6;   P(2 or not 2) = 1

Problem 3.3: five card probabilities

Assume a well-shuffled standard deck with 52 cards and no jokers. First, identify the number of eligible cards for each event. Then divide each count by 52, using complements where convenient.

PartEventCalculationProbability
(a)A king4/52 = 1/137.6923%
(b)A spade13/52 = 1/425%
(c)The king of spades1/521.9231%
(d)Not the king of spades1 − 1/52 = 51/5298.0769%
(e)King of spades or its complement1/52 + 51/52100%

Notice that a named card represents one outcome, whereas a rank or suit represents several. Consequently, “a king” and “the king of spades” have different probabilities. The final part is certain because an event and its complement exhaust all possibilities.

Problem 3.4: colored balls and odds

An urn contains five red balls, three blue balls, and two green balls. Assume each of the ten balls has an equal chance of selection. Therefore, each color probability equals its count divided by ten.

PartRequested eventWorkingAnswer
(a)Red5/100.5 = 50%
(b)Blue3/100.3 = 30%
(c)Green2/100.2 = 20%
(d)Not blue1 − 0.30.7 = 70%
(e)Not green1 − 0.20.8 = 80%
(f)Green or not green0.2 + 0.81 = 100%
(g)Odds in favor of blue3 blue : 7 non-blue3:7
(h)Odds against blue7 non-blue : 3 blue7:3

Odds compare favorable outcomes with unfavorable outcomes. In contrast, probability compares favorable outcomes with all outcomes. Thus, odds of 3:7 correspond to probability 3/(3 + 7) = 0.3, not 3/7.

Odds in favor of A = P(A) : [1 − P(A)]

Problem 3.5: 106 threes in 600 rolls

Part (a): Divide the observed number of threes by the number of rolls. Then compare this empirical result with the fair-die model. The difference concerns observed frequency, not a change in the theoretical probability.

p̂ = 106/600 = 53/300 ≈ 0.176667
p = 1/6 ≈ 0.166667
p̂ − p = 0.01

The observed frequency equals approximately 17.6667%, which is one percentage point above 16.6667%. Moreover, a fair die produces an expected 600/6 = 100 threes in 600 rolls. The actual count of 106 is six above that expectation.

Part (b): Under independent rolls with a stable fair-die probability, the observed proportion converges toward 1/6 as the number of rolls grows. However, the approach need not be monotonic. A longer sample may temporarily move farther from 1/6 before later results bring it closer.

If the die is biased, repeated rolls instead reveal its actual probability of a three. Therefore, the law of large numbers does not make an unfair die fair. It concerns stabilization around the underlying probability.

Problem 3.6: defective production

Part (a): The observed defect rate equals 27 divided by 1,000. Therefore, the empirical probability of a defective item is 0.027, or 2.7%.

p̂ = 27/1000 = 0.027

Part (b): Apply that rate to 1,600 items, assuming the rate remains relevant. Consequently, the expected daily number of defects is 43.2. Keep the decimal when reporting the mathematical expectation.

E(D) = 1600 × 0.027 = 43.2

For a whole-item planning estimate, round to about 43 defective items per day. However, no individual day must produce exactly 43 or 44. The expectation summarizes average output across comparable days; independence is unnecessary for this expectation if every item has the same marginal defect probability.

Problem 3.7: classify event relationships

Part (a), mutually exclusive: A single die roll cannot equal both two and three. Therefore, these two events have an empty intersection. Similarly, one selected ball cannot be both red and blue under the stated color categories.

Part (b), not mutually exclusive: A card can be both an ace and a club. In particular, the ace of clubs belongs to both events. Thus, an addition calculation must account for their overlap.

Part (c), independent: In an independent coin-toss model, heads on the first toss does not change the second toss’s head probability. Consequently, the probability of two heads equals the product of the two marginal probabilities.

Part (d), dependent: Drawing cards without replacement changes the deck composition. For example, after drawing one ace, only three aces remain among 51 cards. The chance of another ace therefore changes from 4/52 to 3/51.

Problem 3.8: Venn diagrams and independence

Part (a): Draw two separate circles inside a rectangle to represent mutually exclusive events. The rectangle represents the sample space. Because the circles do not overlap, no outcome belongs to both events.

Part (b): Draw intersecting circles for events that share outcomes. Their overlap represents A ∩ B. However, the diagram alone does not establish independence, because independence depends on the probabilities.

2026-10-09T03:56:33.945063 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ A B Mutually exclusive events A B A ∩ B Overlapping events
Problem 3.8: separate circles represent disjoint events; overlapping circles represent shared outcomes. Areas are schematic, not numerical probabilities.

Part (c): If both event probabilities are positive, mutually exclusive events are dependent. To see why, compare the joint probability of zero with the positive product P(A)P(B). Since the two quantities differ, independence fails.

P(A ∩ B) = 0 ≠ P(A)P(B)   when P(A) > 0 and P(B) > 0

For completeness, a zero-probability event creates an exception. If either marginal probability equals zero, disjoint events can satisfy the product definition of independence. Therefore, the positive-probability condition matters in a precise answer.

Problem 3.9: addition with disjoint outcomes

Each part combines outcomes that cannot occur together in the stated single draw or roll. Therefore, add the favorable counts and divide by the total. Check the strict inequalities carefully: “less than three” excludes three, while “more than three” also excludes three.

PartQuestionEligible outcomesCalculationAnswer
(a)Die result less than 31 or 22/61/3 ≈ 33.3333%
(b)A heart or a club13 hearts + 13 clubs26/521/2 = 50%
(c)A red or blue ball5 red + 3 blue8/104/5 = 80%
(d)Die result greater than 34, 5, or 63/61/2 = 50%

As a quick check on part (c), the only excluded color is green. Consequently, the complement method gives 1 − 2/10 = 0.8. Agreement between direct counting and the complement provides a useful arithmetic check.

How to choose probability distributions and avoid mistakes

A practical model-selection table

Question structureUseful model or ruleCheck first
One event in a finite uniform sample spaceClassical countingAre outcomes equally likely?
A or BAddition ruleDo the events overlap?
A and BMultiplication ruleDo you need a conditional probability?
Number of successes in n trialsBinomialFixed n, independent trials, constant p
Count within an intervalPoissonAppropriate interval and count-process assumptions
Continuous measurement with a bell-shaped modelNormalPlausible shape, mean, and standard deviation
Success count in a finite sample without replacementHypergeometricPopulation size and number of successes

Start with the random variable rather than a familiar formula. For example, “number of customers who purchase among 20 contacts” suggests a different model from “number of arrivals in an hour.” Next, check the assumptions against how the observations arise.

Sampling without replacement and the binomial model

Sampling without replacement usually creates dependence in a finite population. Therefore, the exact count of selected successes follows a hypergeometric model when the population composition is fixed and the sample is uniform. A binomial approximation may work when the sample is small relative to the population, but it remains an approximation.

An unfair coin does not automatically require a hypergeometric distribution. If its tosses are independent and its head probability stays constant, a binomial model still applies with p different from 0.5. Thus, the sampling mechanism matters more than fairness alone.

Approximations require care

A Poisson distribution can approximate a binomial count when n is large and the success probability is small, with λ = np. If failure is rare instead, apply the approximation to the failure count. However, a convenient rule of thumb cannot guarantee accuracy for every tail probability.

A normal approximation to a discrete count also needs adequate spread and a continuity correction. For example, approximate P(X ≤ k) with a normal area ending at k + 0.5. Whenever exact software calculations are available, compare the approximation before relying on a borderline case.

Further reading: Penn State STAT 414: Approximations for Discrete Distributions.

Expected values support planning, not guarantees

Suppose a service team expects five calls per hour. That average helps plan capacity, but it does not describe the busiest hour. Similarly, an expected defect count of 43.2 says little about extreme daily outcomes without a model for variability.

For a business decision, pair the expected value with a relevant probability. For instance, a manager may care about the chance that demand exceeds available capacity. Consequently, the most useful calculation often concerns a tail, rather than the mean alone.

Common errors and their fixes

MistakeWhy it failsBetter approach
Adding P(A) and P(B) without checking overlapShared outcomes count twice.Subtract P(A ∩ B).
Multiplying marginal probabilities for dependent eventsThe second probability may change.Use P(B | A).
Using variance in a z-score denominatorThe denominator requires standard deviation.Take the square root first.
Treating density height as probabilityContinuous probability comes from area.Integrate or subtract cumulative areas.
Confusing odds with probabilityTheir denominators differ.Convert odds a:b into a/(a+b).
Rounding intermediate values heavilySmall errors accumulate.Round at the final reporting step.
Reading exactly as at leastThe event includes different outcomes.Write the event symbolically.

Chebyshev’s inequality: a broader spread guarantee

The normal percentages require a normal model. In contrast, Chebyshev’s inequality applies to any distribution with a finite mean and finite, positive variance. For k greater than one, it guarantees at least 1 − 1/k² of the probability within k standard deviations, using the corresponding strict-inside form.

P(|X − μ| < kσ) ≥ 1 − 1k²   for k > 1

For example, k = 2 gives a lower bound of 75%, while k = 3 gives 88.8889%. These bounds are weaker than normal-model percentages because they require much less information. Therefore, do not substitute a bound for an exact probability when a specific distribution is known.

Further reading: The Book of Statistical Proofs: Chebyshev’s inequality.

Frequently asked questions about probability distributions

What is the difference between probability and a probability distribution?

A probability describes the chance of one event. In contrast, a distribution describes the probabilities across the possible values of a random variable. You can use that distribution to calculate many different event probabilities.

Can an expected count contain a decimal?

Yes. An expected count is a weighted average, so it need not be a possible single observation. For example, 43.2 expected defects can summarize integer-valued daily counts over many comparable days.

When should I add probabilities instead of multiplying them?

Use the addition rule when the event asks for A or B. For an A-and-B event, use the multiplication rule. However, check overlap before adding and dependence before replacing conditional probabilities with marginal ones.

Does a long run of tails make heads more likely next?

No, provided the fair-coin tosses remain independent. The next head probability stays at 1/2. Therefore, the law of large numbers does not create a short-run force that compensates for earlier outcomes.

Why do my normal answers differ slightly from a printed table?

Printed tables often round areas to four decimal places. Consequently, their final percentages can differ slightly from software results that retain more precision. State your rounding method and compare the event definitions before assuming a substantive disagreement.

Can I assume that every measurement follows a normal distribution?

No. Many measurements show skewness, multiple peaks, or heavy tails. Therefore, inspect the data and the process before selecting a normal model. Standardizing a variable changes its location and scale, but it does not make an arbitrary distribution normal.

Putting probability into practice

Probability distributions become easier when you separate three tasks: define the event, choose a defensible model, and interpret the result. First, write down exactly what the question asks. Next, check equal likelihood, overlap, dependence, and the meaning of each parameter. Finally, report the probability with appropriate precision and a sentence that explains it.

The worked examples show how those habits connect elementary counting to binomial, Poisson, and normal calculations. Moreover, the exercises demonstrate why complements, conditional probabilities, and expected values matter beyond the classroom. Use the solution tables as a review aid, then redo the calculations without looking to check your understanding.

References and further study

The supplied chapter photographs provide the numerical prompts for Examples 1–14 and Problems 3.1–3.9. The sources below support further study of the probability rules and distribution formulas. All solutions, tables, and plotted calculations in this guide were independently developed for this article.

  1. OpenStax. Independent and Mutually Exclusive Events.
  2. OpenStax. Two Basic Rules of Probability.
  3. NIST/SEMATECH. Binomial Distribution.
  4. NIST/SEMATECH. Poisson Distribution.
  5. NIST/SEMATECH. Normal Distribution.
  6. Penn State. Approximations for Discrete Distributions.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *