Master probability distributions with 18 solved exercises on Bayes, binomial models, expected value, and hypergeometric sampling.

Probability Distributions: 18 Solved Exercises

Probability Distributions: 18 Solved Exercises
PROBABILITY · STEP-BY-STEP STUDY GUIDE

Probability Distributions: 18 Solved Exercises

Probability distributions help us turn uncertain outcomes into numbers we can compare, explain, and use. This complete guide connects basic probability rules with conditional probability, Bayes’ theorem, counting methods, expected value, and discrete distributions. Along the way, you will work through 18 exercises with clear calculations, complete tables, and visual explanations.

First, you will learn how to translate everyday wording into mathematical events. Next, the worked examples show why the same question can require different methods when sampling rules change. Finally, you will compare the binomial and hypergeometric models and learn how to check your answers.

The exercise numbers follow the supplied textbook pages, from 3.10 through 3.27. Each problem below uses a concise restatement and an independently explained solution. Calculations retain precision until the final step, so some answers are more precise than the rounded values on the photographed pages.

Understanding probability distributions and basic rules

A probability assigns a number between zero and one to an event. For example, a probability of 0.25 means 25%, or one chance in four under the stated model. However, a probability does not promise a fixed result in a small number of trials.

A sample space contains all possible outcomes. An event contains the outcomes that satisfy a particular condition. Therefore, when individual outcomes have equal probabilities, you can divide the number of favorable outcomes by the total number of outcomes.

P(A) = number of favorable outcomes / number of equally likely outcomes
SymbolMeaningPractical question
A ∪ BUnion: A or B, including bothDoes at least one event occur?
A ∩ BIntersection: A and BDo both events occur?
AᶜComplement: not ADoes A fail to occur?
P(A | B)Conditional probabilityHow likely is A after we know B?
E(X) or μExpected valueWhat is the long-run average?
Var(X) or σ²VarianceHow much spread does the model have?

Choose the rule before substituting numbers

P(A ∪ B) = P(A) + P(B) − P(A ∩ B)
P(A ∩ B) = P(A)P(B | A)
P(Aᶜ) = 1 − P(A)

The addition rule handles an inclusive “or.” In contrast, the multiplication rule handles an “and.” When events are independent, knowing the first outcome does not change the probability of the second, so P(B | A) = P(B).

Mutually exclusive events cannot happen together. Independent events do not change one another’s probabilities. Consequently, two events with positive probabilities cannot be both mutually exclusive and independent: their intersection would need to be both zero and positive.

Further reading: Penn State STAT 500: Probability.

Throughout the card exercises, use a standard 52-card deck with four aces and 13 cards per suit. Unless a problem specifies replacement, leave each drawn card out of the deck. For the urn exercises, the visible calculations establish 10 balls, including five red and three blue; the remaining two are neither red nor blue. The earlier problem that originally introduces the urn is outside the supplied pages.

Exercise 3.10: an ace or a club

Draw one card and find the probability that it is an ace or a club. Then explain why the addition rule subtracts an overlap term.

First, count four aces and 13 clubs. However, the ace of clubs belongs to both groups, so adding those counts includes it twice. Subtract that single overlapping card before dividing by 52.

P(ace ∪ club) = 4/52 + 13/52 − 1/52 = 16/52 = 4/13 ≈ 0.307692

Thus, the probability is approximately 30.77%. The negative term removes double counting; it does not remove the overlapping card from the event itself. Without the correction, the numerator would be 17 even though only 16 distinct cards qualify.

Exercise 3.11: overlapping economic events and card categories

(a) Inflation or recession

Suppose P(I) = 0.30, P(R) = 0.20, and P(I ∩ R) = 0.06. Because the events overlap, include the intersection correction.

P(I ∪ R) = 0.30 + 0.20 − 0.06 = 0.44 = 44%

These are hypothetical exercise inputs, not estimates of current economic conditions. As an additional check, 0.30 × 0.20 = 0.06, so this particular pair of event probabilities also satisfies the independence condition. Nevertheless, independence does not eliminate the overlap term in an “or” calculation.

(b) An ace, a club, or a diamond

There are four aces, 13 clubs, and 13 diamonds. However, the ace of clubs and ace of diamonds each appear twice in those counts. Clubs and diamonds have no common card, so there is no further intersection to subtract.

P(ace ∪ club ∪ diamond) = (4 + 13 + 13 − 1 − 1)/52
= 28/52 = 7/13 ≈ 0.538462

Alternatively, count all 26 clubs and diamonds, then add the two aces in the other suits. Both approaches give a probability of approximately 53.85%.

Exercise 3.12: multiplication for independent trials

(a) Two sixes in two rolls

A fair die has probability 1/6 of showing six on each roll. Since the rolls are independent, multiply the two probabilities.

P(two sixes) = (1/6)(1/6) = 1/36 ≈ 0.027778

(b) A six on each of two dice

Rolling two independent fair dice simultaneously creates the same probability structure. Therefore, the answer remains 1/36, or approximately 2.78%. The timing of the rolls does not change their independence.

(c) Two blue balls with replacement

The urn contains three blue balls among 10 balls. After the first draw, return the ball and mix the urn thoroughly. Consequently, the blue-ball probability stays at 3/10 for the next draw.

P(two blue) = (3/10)² = 9/100 = 0.09 = 9%

(d) Three girls in three births

Use the exercise’s simplified assumptions: independent births and a probability of 0.5 for a girl on each birth. Under that model, multiply three identical probabilities.

P(three girls) = (0.5)³ = 0.125 = 12.5%

This calculation illustrates an independent-trial model. It does not claim that the simplifying assumptions describe every real family or population exactly.

Exercise 3.13: the sample space for two dice

(a) All 36 ordered outcomes

Label the dice first and second so that their outcomes remain distinguishable. Each first-die result pairs with six second-die results. Therefore, the sample space contains 6 × 6 = 36 equally likely ordered pairs.

First die / second die123456
1(1, 1)(1, 2)(1, 3)(1, 4)(1, 5)(1, 6)
2(2, 1)(2, 2)(2, 3)(2, 4)(2, 5)(2, 6)
3(3, 1)(3, 2)(3, 3)(3, 4)(3, 5)(3, 6)
4(4, 1)(4, 2)(4, 3)(4, 4)(4, 5)(4, 6)
5(5, 1)(5, 2)(5, 3)(5, 4)(5, 5)(5, 6)
6(6, 1)(6, 2)(6, 3)(6, 4)(6, 5)(6, 6)

An ordered pair such as (1, 4) differs from (4, 1), even though both produce a total of five. As a result, the totals from two through 12 do not have equal probabilities.

(b) A total of five

The qualifying pairs are (1, 4), (2, 3), (3, 2), and (4, 1). Hence, four of the 36 outcomes meet the condition.

P(sum = 5) = 4/36 = 1/9 ≈ 0.111111

(c) A total of four or less, and more than four

A sum of two has one outcome, a sum of three has two, and a sum of four has three. Therefore, six outcomes produce a total no larger than four. Use the complement for the opposite event.

P(sum ≤ 4) = (1 + 2 + 3)/36 = 1/6
P(sum > 4) = 1 − 1/6 = 5/6
2026-10-11T12:58:09.346804 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 2 3 4 5 6 7 8 9 10 11 12 Sum of the two dice 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 Probability 0.0278 0.0556 0.0833 0.1111 0.1389 0.1667 0.1389 0.1111 0.0833 0.0556 0.0278 Two fair dice: distribution of the total
Two fair dice: distribution of the total. Bars show probability at each integer count.

The chart peaks at seven because six ordered pairs produce that sum. In contrast, only one pair produces two and only one produces 12. This counting pattern explains the triangular shape.

Exercise 3.14: conditional draws without replacement

Start with five red and five nonred balls. Here, the question gives information about earlier draws. Therefore, calculate each probability from the balls that remain, rather than multiplying by the probability of the known history.

(a) Red on the second draw after a red first draw

P(R₂ | R₁) = remaining red / remaining total = 4/9 ≈ 0.444444

The first red draw removes one red ball and one ball overall. As a result, four of the remaining nine balls are red.

(b) Red on the second draw after a nonred first draw

P(R₂ | R₁ᶜ) = 5/9 ≈ 0.555556

A nonred first draw leaves all five red balls in the urn. Consequently, the second-draw probability exceeds the original one-half probability.

(c) Red on the third draw after one red and one nonred

P(R₃ | one red and one nonred previously drawn) = 4/8 = 1/2

The first two draws remove one ball from each color group. Thus, four red and four nonred balls remain, regardless of the order of the first two colors.

Further reading: Penn State STAT 414: Conditional Probability.

Exercise 3.15: joint probabilities without replacement

Unlike Exercise 3.14, these questions ask for the probability of an entire sequence. First, calculate the initial draw probability. Then multiply by each conditional probability along the required path.

(a) Two red balls

P(R₁ ∩ R₂) = (5/10)(4/9) = 2/9 ≈ 0.222222

After one red ball leaves the urn, only four red balls remain. Consequently, using (5/10)² would incorrectly treat the draws as independent.

(b) Two aces

P(A₁ ∩ A₂) = (4/52)(3/51) = 1/221 ≈ 0.004525

The first ace reduces both the ace count and the deck size. Therefore, the probability is approximately 0.4525%, rather than 0.004525%.

(c) Ace of clubs, followed by a spade

P(A♣ first, spade second) = (1/52)(13/51) = 1/204 ≈ 0.004902

Removing the ace of clubs leaves every spade in the deck. Accordingly, the second factor has numerator 13 and denominator 51.

(d) A spade, followed by the ace of clubs

P(spade first, A♣ second) = (13/52)(1/51) = 1/204 ≈ 0.004902

The reversed sequence has the same probability here. However, the two descriptions represent different ordered outcomes; equal probabilities do not make the events identical.

(e) Three red balls without replacement

P(three red) = (5/10)(4/9)(3/8) = 1/12 ≈ 0.083333

(f) Three red balls with replacement

P(three red) = (5/10)³ = 1/8 = 0.125

Replacement raises the probability from approximately 8.33% to 12.5%. Each successful red draw otherwise makes the next red draw less likely. Thus, the sampling rule directly changes the answer.

Exercise 3.16: production shifts and total probability

A factory’s morning shift makes 1,000 items and its evening shift makes 600. Historical defect rates are 200 per 100,000 for the morning shift and 500 per 100,000 for the evening shift. Assume those rates apply to the production period under study.

ShiftOutputShare of totalDefect probabilityExpected defects
Morning10000.6250.0022
Evening6000.3750.0053
Total16001.0000.003125 overall5

First, convert each historical defect rate to a probability: 200/100,000 = 0.002 and 500/100,000 = 0.005. The photographed solution shows “20” in one morning-rate numerator, but that conflicts with both the stated 200 and the decimal 0.002. This solution consistently uses 200.

(a) Morning shift and defective

P(M ∩ D) = P(M)P(D | M) = (0.625)(0.002)
= 0.00125 = 0.125%

(b) Evening shift and defective

P(E ∩ D) = (0.375)(0.005) = 0.001875 = 0.1875%

(c) Evening shift and not defective

P(E ∩ Dᶜ) = (0.375)(1 − 0.005)
= (0.375)(0.995) = 0.373125 = 37.3125%

(d) Defective, regardless of shift

An item comes from exactly one of the two shifts. Therefore, add the two disjoint routes to a defective item.

P(D) = P(M)P(D | M) + P(E)P(D | E)
= 0.00125 + 0.001875 = 0.003125 = 0.3125%

Equivalently, the model predicts two morning defects and three evening defects on average. Dividing five expected defects by 1,600 gives the same overall probability. However, the actual batch need not contain exactly five defective items: an expectation is not an observed count.

A simple average of 0.002 and 0.005 would give 0.0035, which is incorrect here. Since the shifts produce different quantities, the overall rate must weight each defect rate by its output share.

Exercise 3.17: deriving and applying Bayes’ theorem

(a) Derive the formula

Write the same intersection in two ways. Because A ∩ B and B ∩ A describe the same event, their probabilities must agree.

P(A ∩ B) = P(A)P(B | A) = P(B)P(A | B)

Next, divide by P(B), provided P(B) > 0. This step gives Bayes’ theorem.

P(A | B) = P(A)P(B | A) / P(B)

The prior P(A) describes your probability before learning B. The likelihood P(B | A) describes how compatible the evidence is with A. Finally, the posterior P(A | B) updates the probability after you observe B.

(b) Which shift produced a defective item?

P(M | D) = (0.625 × 0.002)/0.003125 = 0.40 = 40%
P(E | D) = (0.375 × 0.005)/0.003125 = 0.60 = 60%

Although the morning shift produces 62.5% of all items, it accounts for only 40% of defective items under this model. Conversely, the evening shift produces fewer items but has a higher defect rate. That difference explains its larger share of defects.

As a check, the two posterior probabilities sum to one. Also, 40% differs sharply from the morning defect rate of 0.2%: P(M | D) and P(D | M) answer different questions.

Further reading: Penn State STAT 414: Bayes’ Theorem.

Exercise 3.18: combinations and permutations

A club has eight members and needs a three-person committee. The correct counting method depends on whether individual roles matter.

(a) A committee with no distinct offices

Choosing the same three people in another order does not create a new committee. Therefore, use a combination.

C(n, r) = n! / [r!(n − r)!]
C(8, 3) = (8 × 7 × 6)/(3 × 2 × 1) = 56

(b) A president, treasurer, and secretary

There are eight choices for president, seven remaining choices for treasurer, and six for secretary. Because the roles differ, each assignment counts separately.

P(n, r) = n! / (n − r)!
P(8, 3) = 8 × 7 × 6 = 336

Alternatively, each of the 56 committees permits 3! = 6 role assignments. Thus, 56 × 6 = 336 confirms the result. The key question is whether exchanging two selected people changes the outcome you want to count.

Exercise 3.19: random variables and probability distributions

(a) What is a random variable?

A random variable assigns a numerical value to each outcome of a random experiment. For example, let X equal the number showing on one fair die. Then X can take the values one through six.

(b) What makes a random variable discrete?

A discrete random variable has a finite or countably infinite set of possible values. Counts of heads, applicants, or defective products are examples. However, a discrete variable does not need to have only finitely many possible values; a waiting-time count can extend indefinitely.

(c) What is a discrete probability distribution?

A probability mass function, or PMF, assigns a probability to each possible discrete value. Every probability must be nonnegative, and the probabilities must sum to one. For a fair die, P(X = x) = 1/6 for x = 1, 2, 3, 4, 5, 6.

p(x) = P(X = x), with p(x) ≥ 0 and Σ p(x) = 1

(d) Probability versus observed relative frequency

A probability distribution describes a model. An empirical relative-frequency distribution summarizes observed data, using the count of each outcome divided by the total count. For example, 17 sixes in 100 rolls give an empirical relative frequency of 0.17, while the fair-die model assigns 1/6.

Under independent, identically distributed sampling, relative frequencies converge to the model probabilities as the number of observations grows. Nevertheless, they do not have to match exactly in a finite sample. Probability models can also use estimated parameters, so “theoretical” does not mean that data play no role in constructing them.

Further reading: Penn State STAT 414: Discrete Random Variables.

Exercise 3.20: expected value and variance formulas

(a) Expected value as a probability-weighted average

For observed values with frequencies fᵢ, the ordinary weighted mean is Σxᵢfᵢ/N. Rewrite that expression as Σxᵢ(fᵢ/N). Consequently, replacing observed proportions with model probabilities gives the expected value.

E(X) = μ = Σ x p(x)

The expectation summarizes the distribution’s long-run center. However, it need not equal any possible individual outcome. A fair die has expected value 3.5 even though no face shows 3.5.

(b) Variance and the computational shortcut

Variance averages squared distances from the mean. Therefore, large deviations receive more weight than small deviations.

Var(X) = E[(X − μ)²] = Σ (x − μ)²p(x)

Next, expand the square and use E(X) = μ. Since expectations preserve sums and constant multipliers, the expression simplifies.

E[(X − μ)²] = E(X²) − 2μE(X) + μ²
= E(X²) − 2μ² + μ²
Var(X) = E(X²) − [E(X)]²
SD(X) = √Var(X)

Variance has squared units, while standard deviation uses the original units. For count data, this distinction makes standard deviation easier to interpret. These formulas require the relevant moments to exist, which they do for all finite distributions in this guide.

Exercise 3.21: daily job applications

An agency records the number of applications it processes over 100 days. Treat the observed proportions as an empirical probability distribution. Then calculate the mean, variance, and standard deviation.

Applications xDays fp(x) = f/100x p(x)x²x²p(x)
7100.10.7494.9
8100.10.8646.4
10200.22.010020.0
11300.33.312136.3
12200.22.414428.8
14100.11.419619.6
Total1001.010.6—116.0

Step 1: Calculate the expected number

E(X) = 0.7 + 0.8 + 2.0 + 3.3 + 2.4 + 1.4
= 10.6 applications per day

Step 2: Calculate variance

E(X²) = 4.9 + 6.4 + 20.0 + 36.3 + 28.8 + 19.6 = 116
Var(X) = 116 − 10.6² = 3.64 applications²

Step 3: Take the square root

SD(X) = √3.64 = 1.907878 ≈ 1.91 applications

The agency’s average workload is 10.6 applications, even though every daily count is an integer. Meanwhile, the standard deviation describes the spread around that average. Forecasting future days from these frequencies requires the additional assumption that the recorded period remains representative.

This exercise asks for variance of the empirical probability distribution, so the weights sum to one. If you instead sought the unbiased sample variance for an underlying population, you would use the correction 100/99, giving approximately 3.676768. That is a different statistical target.

2026-10-11T12:58:09.406485 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 7 8 10 11 12 14 Applications processed per day 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Probability 0.1000 0.1000 0.2000 0.3000 0.2000 0.1000 Daily applications: empirical probability distribution
Daily applications: empirical probability distribution. Bars show probability at each integer count.

How binomial probability distributions work

A binomial variable counts successes in a fixed number of independent trials with the same success probability. Here, “success” simply means the outcome you choose to count. Therefore, a defective product can count as a success mathematically even though it is undesirable operationally.

X ∼ Binomial(n, p)
P(X = k) = C(n, k)pk(1 − p)n − k

The factor pk(1 − p)n − k gives the probability of one particular sequence with k successes. Meanwhile, C(n, k) counts the different positions those successes can occupy. Multiplying the two quantities includes every qualifying sequence.

Further reading: NIST/SEMATECH e-Handbook: Binomial Distribution.

Exercise 3.22: heads in five coin flips

(a) Conditions for the binomial model

  1. Fix the number of trials in advance.
  2. Classify each trial into success or failure.
  3. Require independence across trials.
  4. Keep the success probability constant.

For five fair, independent coin flips, n = 5 and p = 0.5. Accordingly, the number of heads follows a binomial distribution.

(b) Exactly three heads

P(X = 3) = C(5, 3)(0.5)³(0.5)²
= 10/32 = 0.3125 = 31.25%

(c) Fewer than three heads

“Fewer than three” includes zero, one, and two. Therefore, sum those three probabilities, without including three itself.

P(X < 3) = P(0) + P(1) + P(2)
= 1/32 + 5/32 + 10/32 = 0.5 = 50%
Heads kC(5, k)P(X = k)
010.03125
150.15625
2100.31250
3100.31250
450.15625
510.03125

As a second check, the distribution is symmetric around 2.5 heads. Consequently, zero through two heads together have the same probability as three through five heads together.

Exercise 3.23: exactly half and at least three

(a) Three blond-haired children among six

Use the stated model with n = 6, p = 1/4, and independent outcomes. Exactly half of six means k = 3. The hair-color example is a simplified probability exercise, not a general genetic model.

P(X = 3) = C(6, 3)(1/4)³(3/4)³
= 20 × 27/4096 = 135/1024
= 0.1318359375 ≈ 13.18%

The combination factor counts the 20 ways to choose which three children have the specified characteristic. Without it, you would calculate only one particular arrangement.

(b) At least three hits in four attempts

Assume independent attempts, each with hit probability 0.30. “At least three” includes exactly three and exactly four. Thus, add those two binomial probabilities.

P(X = 3) = C(4, 3)(0.3)³(0.7) = 0.0756
P(X = 4) = (0.3)⁴ = 0.0081
P(X ≥ 3) = 0.0756 + 0.0081 = 0.0837 = 8.37%

If fatigue or practice changes the hit probability between attempts, the simple binomial assumptions may fail. Nevertheless, under the exercise’s constant-probability model, 8.37% is the required result.

Exercise 3.24: quality-control probabilities

(a) No more than two defective tubes in a sample of 10

A very large shipment has a 20% defect rate. Let X count defective tubes in a sample of 10. Since the sample is tiny relative to the shipment, use the binomial model as an approximation if sampling occurs without replacement.

X ∼ Binomial(10, 0.2)
P(X ≤ 2) = Σk=02 C(10, k)(0.2)k(0.8)10 − k
k defectiveCalculationProbability
0C(10, 0) × 0.2^0 × 0.8^100.1073741824
1C(10, 1) × 0.2^1 × 0.8^90.2684354560
2C(10, 2) × 0.2^2 × 0.8^80.3019898880
P(X ≤ 2) = 0.1073741824 + 0.2684354560 + 0.3019898880
= 0.6777995264 ≈ 67.78%

“No more than two” includes two itself. In contrast, “fewer than two” would include only zero and one, giving a smaller answer.

(b) Exactly 10 acceptable items among 15

An item is acceptable with probability 0.85. Assuming independent items and a stable process, let X count acceptable items among 15. Then n = 15, k = 10, and p = 0.85.

P(X = 10) = C(15, 10)(0.85)10(0.15)⁵
= 3003(0.85)10(0.15)⁵
= 0.0448953007 ≈ 4.49%

Alternatively, count exactly five defective items with probability 0.15. That complementary count produces the same event and the same answer. However, the event is exactly 10 acceptable items, not at least 10.

Exercise 3.25: complete binomial tables and graphs

(a) Heads in four fair coin flips

Let X count heads across four independent fair flips. First, list every possible count from zero through four. Next, apply the binomial formula at each count.

Heads kC(4, k)ProbabilityPercentage
010.06256.25%
140.250025.00%
260.375037.50%
340.250025.00%
410.06256.25%
Total—1.0000100%
P(X = k) = C(4, k)/16
ΣP(X = k) = (1 + 4 + 6 + 4 + 1)/16 = 1
2026-10-11T12:58:09.452467 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 0 1 2 3 4 Number of heads 0.0 0.1 0.2 0.3 0.4 Probability 0.0625 0.2500 0.3750 0.2500 0.0625 Four fair coin flips: number of heads
Four fair coin flips: number of heads. Bars show probability at each integer count.

The distribution is symmetric, with its largest probability at two heads. Moreover, zero heads and four heads have equal probabilities, as do one head and three heads.

(b) Defective items among five, with p = 0.30

Now let X count defective items among five independent items. The defect probability is 0.30 on each trial. Consequently, use n = 5 and p = 0.30.

Defective items kProbabilityPercentage
00.1680716.807%
10.3601536.015%
20.3087030.870%
30.1323013.230%
40.028352.835%
50.002430.243%
Total1.00000100%
2026-10-11T12:58:09.498319 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 0 1 2 3 4 5 Number of defective items 0.0 0.1 0.2 0.3 0.4 Probability 0.1681 0.3601 0.3087 0.1323 0.0283 0.0024 Five items: number defective when p = 0.30
Five items: number defective when p = 0.30. Bars show probability at each integer count.

The most likely count is one defective item, with probability 0.36015. However, the mean is 1.5, so the most likely count and expected value differ. Since p is below one-half, the distribution has positive skewness and a longer right-side tail.

Exercise 3.26: binomial means, standard deviations, and skewness

For a binomial count, the mean is np and the variance is np(1 − p). Therefore, the standard deviation is the square root of np(1 − p). Apply these formulas to the four models from Exercises 3.23 and 3.24.

μ = np
σ² = np(1 − p)
σ = √[np(1 − p)]
Skewness = (1 − 2p) / √[np(1 − p)]
Part / modelnpMeanVarianceSDShape
(a) Blond-haired children60.251.51.1251.060660Right-skewed
(b) Target hits40.31.20.840.916515Right-skewed
(c) Defective tubes100.221.61.264911Right-skewed
(d) Acceptable items150.8512.751.91251.382932Left-skewed

Substitution for all four parts

(a) μ = 6(0.25) = 1.5; σ = √[6(0.25)(0.75)] ≈ 1.060660
(b) μ = 4(0.30) = 1.2; σ = √[4(0.30)(0.70)] ≈ 0.916515
(c) μ = 10(0.20) = 2; σ = √[10(0.20)(0.80)] ≈ 1.264911
(d) μ = 15(0.85) = 12.75; σ = √[15(0.85)(0.15)] ≈ 1.382932

For 0 < p < 1, the sign of 1 − 2p determines the direction of binomial skewness. Thus, p below 0.5 gives right skewness, p above 0.5 gives left skewness, and p equal to 0.5 gives symmetry. A large n can make a skewed distribution look nearly symmetric, so direction alone does not measure the strength of the asymmetry.

As a practical interpretation, 12.75 acceptable items is the average across repeated samples of 15. It does not describe a fractional product in one sample. Meanwhile, a standard deviation near 1.38 describes variation in the acceptable-item count.

Exercise 3.27: hypergeometric probability distributions

A group has 10 people, including five men. Select six people uniformly at random without replacement and find the probability of choosing exactly two men. Since the sample removes 60% of the population, dependence between draws matters greatly.

(a) The exact hypergeometric calculation

Let N denote the population size, K the number of successes in the population, and n the sample size. The random variable X counts successes in the sample. Then the hypergeometric model gives:

P(X = k) = C(K, k)C(N − K, n − k) / C(N, n)

First, choose two of the five men in C(5, 2) ways. Next, choose four of the five other people in C(5, 4) ways. Finally, divide by the number of possible six-person samples.

P(X = 2) = C(5, 2)C(5, 4)/C(10, 6)
= (10 × 5)/210 = 5/21 ≈ 0.238095 = 23.81%

(b) What would the binomial calculation give?

Pbinomial(X = 2) = C(6, 2)(0.5)²(0.5)⁴
= 15/64 = 0.234375 = 23.4375%

The binomial answer understates this particular probability by approximately 0.3720 percentage points. However, the closeness at k = 2 does not validate the model. In fact, the binomial model assigns positive probability to zero men or six men, even though both outcomes are impossible in this group.

Men in sample kExact hypergeometricBinomial comparison
00.0000000.015625
10.0238100.093750
20.2380950.234375
30.4761900.312500
40.2380950.234375
50.0238100.093750
60.0000000.015625
2026-10-11T12:58:09.551862 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ 0 1 2 3 4 5 6 Number of men in the sample 0.0 0.1 0.2 0.3 0.4 0.5 Probability Sampling six from ten: exact model versus binomial Hypergeometric: exact Binomial: comparison
Sampling six from ten: exact model versus binomial. Bars show probability at each integer count.

Why sampling without replacement changes the spread

Both models have mean three. Nevertheless, sampling without replacement reduces the variance because each draw changes the remaining population. The finite-population correction quantifies that reduction.

Var(Xhypergeometric) = np(1 − p)(N − n)/(N − 1)
= 6(0.5)(0.5)(4/9) = 2/3
Var(Xbinomial) = 6(0.5)(0.5) = 1.5

A common rule of thumb allows a binomial approximation when the sample is at most about 5% of a finite population. However, this guideline is not a universal guarantee of accuracy. When the population size and composition are known, the hypergeometric model supplies the exact count probabilities for simple random sampling without replacement.

Further reading: Penn State STAT 414: The Binomial Distribution, including sampling-model distinctions.

Quick answer key for all 18 exercises

ExerciseResults
3.10(a) 4/13; (b) subtract the overlap to avoid double counting
3.11(a) 0.44; (b) 7/13
3.12(a) 1/36; (b) 1/36; (c) 0.09; (d) 0.125
3.13(a) 36 ordered pairs; (b) 1/9; (c) 1/6 and 5/6
3.14(a) 4/9; (b) 5/9; (c) 1/2
3.15(a) 2/9; (b) 1/221; (c) 1/204; (d) 1/204; (e) 1/12; (f) 1/8
3.16(a) 0.00125; (b) 0.001875; (c) 0.373125; (d) 0.003125
3.17(a) Bayes’ formula; (b) morning 0.40, evening 0.60
3.18(a) 56 committees; (b) 336 officer assignments
3.19Random variable, discrete support, PMF, and empirical frequencies: see definitions above
3.20E(X) = Σxp(x); Var(X) = E(X²) − [E(X)]²
3.21Mean 10.6; variance 3.64; SD ≈ 1.907878
3.22(a) fixed n, binary outcomes, independence, constant p; (b) 0.3125; (c) 0.5
3.23(a) 0.1318359375; (b) 0.0837
3.24(a) 0.6777995264; (b) 0.0448953007
3.25Complete distributions and both plots appear above
3.26Means: 1.5, 1.2, 2, 12.75; SDs: 1.060660, 0.916515, 1.264911, 1.382932
3.27(a) 5/21 ≈ 0.238095; (b) 15/64 = 0.234375

How to avoid common probability mistakes

Translate the wording before calculating

WordingEvent for an integer-valued X
Exactly kX = k
Fewer than kX < k, or X ≤ k − 1
At most k / no more than kX ≤ k
At least kX ≥ k
More than kX > k, or X ≥ k + 1

First, define what X counts. Then translate the wording into an equality or inequality before entering numbers. This small step prevents most off-by-one errors in cumulative probability questions.

Keep conditional and joint probabilities separate

A conditional question treats the given information as established. In contrast, a joint question asks for the probability that the full combination occurs. Consequently, 4/9 in Exercise 3.14(a) and 2/9 in Exercise 3.15(a) are both correct because they answer different questions.

Check model assumptions and arithmetic

Before using a binomial formula, verify independence and a constant success probability. Next, check whether draws involve replacement or a negligible sampling fraction. Finally, make sure your probability lies between zero and one and that a complete distribution sums to one.

Retain several digits during intermediate calculations. Otherwise, repeated rounding can distort the final result or make table totals appear inconsistent. To convert a decimal probability into a percentage, multiply by 100 exactly once.

Frequently asked questions about probability distributions

When should I add probabilities rather than multiply them?

Add probabilities when combining alternative routes to an event, with an overlap correction where necessary. Multiply probabilities when describing successive or joint conditions. For dependent events, use conditional probabilities in the product.

Can an expected count contain a decimal?

Yes. Expected value is an average over the distribution, so it can lie between possible outcomes. For example, an expected workload of 10.6 applications is compatible with integer daily counts.

Does “without replacement” always rule out a binomial calculation?

It rules out an exact binomial model for the usual nondegenerate finite-population count because the draws are dependent. However, a binomial approximation can work well when the sample is sufficiently small relative to the population. Use the hypergeometric distribution when an exact finite-population calculation is appropriate.

Are an event and its reverse conditional equally likely?

Usually not. Bayes’ theorem relates P(A | B) to P(B | A), but the prior probabilities also matter. Therefore, always read the expression after the conditioning bar before interpreting a result.

Why does the binomial formula contain a combination?

One sequence of successes and failures represents only one ordering. The combination C(n, k) counts all orderings with exactly k successes. Thus, it converts one sequence probability into the probability of the complete count event.

Conclusion: build the model, then solve the problem

Probability distributions become easier to use when you separate the event, the assumptions, and the calculation. First, identify whether the question asks for a union, an intersection, a conditional event, or a count. Next, choose a model that matches the sampling process. Finally, verify the answer against the sample space, the units, and the wording.

The 18 solved exercises show how a small set of ideas connects card draws, dice totals, production quality, committee selection, and count data. With those connections in place, you can approach new problems systematically and explain why each formula fits.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *