PROBABILITY • STEP-BY-STEP STUDY GUIDE
Probability Distributions: Guide and Solved Exercises
Probability distributions turn uncertainty into numbers you can calculate, compare, and explain. For example, they help you estimate defective output, count incoming calls, or measure the chance that an observation falls within a range. This guide builds the subject from simple events to binomial, Poisson, and normal models. Along the way, you will work through every numbered example and exercise in the supplied pages.
First, you will learn how to describe outcomes and choose the correct probability rule. Next, you will solve the 14 introductory examples and exercises 3.1 through 3.9. Finally, you will compare the models, review common mistakes, and apply the ideas to business decisions. Each solution explains the assumptions as well as the arithmetic.
Probability foundations: outcomes, events, and uncertainty
An experiment produces an outcome. For example, rolling a die produces one of six face values. The sample space contains all possible outcomes, while an event collects the outcomes that meet a condition. Thus, “an even result” corresponds to the event {2, 4, 6}.
A probability ranges from zero to one. In a finite model, zero describes an impossible event and one describes certainty. However, continuous models need a more careful interpretation: a single exact value can have probability zero without being an impossible numerical value.
Classical probability and equally likely outcomes
When a finite sample space contains equally likely outcomes, count the outcomes that satisfy the event and divide by the total. Therefore, a fair die assigns the same probability to each face. A biased die requires different probabilities, so simple counting alone no longer gives the answer.
For example, the event “even result” contains three of the six die faces. Consequently, its probability equals 3/6, or 0.5. The word favorable means that an outcome belongs to the event; it does not mean that the outcome is desirable.
Empirical probability and subjective probability
Empirical probability uses observed results. For instance, 27 defective items among 1,000 inspected items produce an estimated defect probability of 0.027. However, that estimate describes the observed sample and may not perfectly represent future production.
Subjective probability expresses a degree of belief based on available evidence. For example, an analyst might assess a new product’s chance of meeting a sales target before comparable data exist. Nevertheless, a useful assessment should state its assumptions and change when new evidence arrives.
A compact notation guide
| Symbol | Meaning | How to read it |
|---|---|---|
| S | Sample space | All possible outcomes |
| Ac | Complement of A | A does not occur |
| A ∪ B | Union | A or B, including both |
| A ∩ B | Intersection | Both A and B |
| P(B | A) | Conditional probability | Probability of B given A |
| X | Random variable | Numerical outcome of interest |
| μ and σ | Mean and standard deviation | Center and spread |
| p̂ | Observed proportion | Estimated probability from data |
The complement rule
An event and its complement cover the entire sample space without overlap. Therefore, their probabilities add to one. This rule often makes a calculation shorter, especially when the question asks for “at least one” or “not.”
For example, the complement of “at least one defect” is “no defects.” Consequently, subtracting the probability of no defects from one gives the probability of at least one. Always define the event before choosing its complement.
Addition, multiplication, and conditional probability
Use addition for “or”
The general addition rule combines two events and removes their overlap. Otherwise, outcomes that belong to both events enter the calculation twice. For example, a king of spades belongs to both the spade event and the king event.
Mutually exclusive events cannot occur together in the same experiment. Therefore, their intersection has probability zero, and the addition rule simplifies. Rolling a two and rolling a three on one die roll provide a straightforward example.
Use conditional probability for new information
A conditional probability restricts attention to outcomes consistent with known information. For instance, removing a king from a deck changes both the remaining king count and the deck size. As a result, the second draw requires an updated probability.
Use multiplication for “and”
The general multiplication rule works with dependent or independent events. First, calculate the probability of the initial event. Then multiply by the probability of the next event conditional on the first.
Independent events satisfy P(A ∩ B) = P(A)P(B). Thus, when P(A) is positive, learning that A occurred leaves the probability of B unchanged. Independent coin tosses illustrate this relationship, provided the experiment genuinely supports that assumption.
Further reading: OpenStax: Two Basic Rules of Probability.
Why mutually exclusive does not mean independent
Mutual exclusivity describes whether events can occur together. In contrast, independence describes whether information about one event changes the probability of another. These questions differ, so the terms cannot replace each other.
If two mutually exclusive events both have positive probability, they must be dependent. For example, knowing that one die roll produced a two rules out a three. However, the general statement needs a zero-probability exception: disjoint events can be independent if at least one has probability zero.
Further reading: OpenStax: Independent and Mutually Exclusive Events.
Solved Examples 1–8: basic probability rules
Example 1: one fair coin toss
Question: Find the probabilities of heads and tails.
First, list the two outcomes: H and T. Because the coin is fair, each outcome receives half the total probability.
Answer: Heads: 50%; tails: 50%; heads or tails: 100%. These results assume the usual two-outcome model.
Example 2: one fair die roll
Question: Find the probability of each face and the probability of not rolling a one.
There are six equally likely outcomes. Therefore, any specified face has probability 1/6. Next, use the complement rule to count every face except one.
Answer: Each face: approximately 16.6667%; not one: approximately 83.3333%.
Example 3: a jack and a non-diamond
Question: Draw one card from a standard 52-card deck without jokers. Find the chance of a jack and the chance of a card that is not a diamond.
A standard deck contains four jacks and 13 diamonds. Thus, 39 cards are not diamonds. Treat these as separate questions rather than as a joint event.
Answer: Jack: approximately 7.6923%; non-diamond: 75%.
Example 4: 53 heads in 100 tosses
Question: Compare the observed frequency of heads with the theoretical probability for a fair coin.
First, divide 53 by 100 to obtain the observed proportion. Then compare that result with the fair-coin probability of 0.5. Random variation can create a difference even when the model is correct.
Answer: The observed frequency is 53%, which is 3 percentage points above 50%. More tosses do not guarantee that every successive estimate moves closer to 50%.
Example 5: several possible die faces
Question: Find P(2 or 3) and P(2 or 3 or 4) on one fair die roll.
These face events do not overlap. Therefore, add their individual probabilities, or count the relevant faces directly.
P(2 or 3 or 4) = 3/6 = 1/2
Answer: The probabilities are approximately 33.3333% and exactly 50%, respectively.
Example 6: a spade or a king
Question: Find the probability that one card is a spade or a king.
First, count 13 spades and four kings. However, the king of spades appears in both groups. Subtract that one-card overlap before dividing by 52.
Answer: Approximately 30.7692%. Here, “or” includes the king of spades.
Example 7: consecutive heads
Question: Find the probabilities of two heads in two tosses and three heads in three tosses of a fair coin.
Assume the tosses are independent. Consequently, multiply one factor of 1/2 for each required head.
P(HHH) = (1/2)3 = 1/8
Answer: Two heads: 25%; three heads: 12.5%. These events specify heads on every toss, not merely a head somewhere in the sequence.
Example 8: a specified king followed by another king
Question: Draw the king of diamonds first, then another king, without replacement. What is the joint probability?
The first event has probability 1/52. Next, after that card leaves the deck, three kings remain among 51 cards. Therefore, multiply 1/52 by 3/51.
Answer: Approximately 0.00113122, or 0.113122%. The conditional second-draw probability alone is 3/51, or 5.88235%.
Notice the difference between a specified first king and any first king. If the question asked for any two kings, the first factor would be 4/52. Consequently, that different event would have probability 1/221, four times the answer above.
Understanding probability distributions
A random variable assigns a number to an outcome. For example, two coin tosses produce sequences such as HT and HH, while X can count the number of heads. A probability distribution then describes how probability spreads across the possible values of X.
Discrete probability distributions
A discrete random variable has a finite or countably infinite set of possible values. Thus, a binomial count has finitely many values, while a Poisson count can take any nonnegative integer. The word discrete does not require a finite list.
For a discrete variable, the probability mass function gives the probability at each value. Meanwhile, the cumulative distribution function adds all probabilities at or below a threshold. Therefore, F(2) means P(X ≤ 2), rather than P(X = 2).
Continuous probability distributions
A continuous model uses a probability density function. Instead of adding isolated point probabilities, calculate areas over intervals. Consequently, a curve’s height is a density, while the area beneath the curve represents probability.
For an absolutely continuous distribution, P(X = a) = 0. Therefore, including or excluding a single endpoint does not change an interval probability. A density can exceed one on a narrow interval, but its total area must equal one.
Expected value and spread
The expected value is a probability-weighted average. For example, a count can have an expectation of 43.2 even though each observed count is an integer. Variance measures squared deviations from that mean, while standard deviation returns the spread to the original units.
Var(X) = ∑x (x − μ)2p(x); σ = √Var(X)
Binomial probability distributions: counting successes
Use a binomial model when you count successes in a fixed number of independent trials, each with the same success probability. Every trial must have two categories: success and failure. However, “success” simply names the category you count, so a defective item can count as a success mathematically.
P(X = k) = n!k!(n − k)! pk(1 − p)n − k
The factorial term counts the arrangements that contain exactly k successes. For example, four heads in six tosses can occur in several different orders. Therefore, the formula includes both the probability of one arrangement and the number of eligible arrangements.
Further reading: NIST/SEMATECH: Binomial Distribution.
Example 9: heads in two tosses
Question: Construct the probability distribution for the number of heads in two independent fair tosses.
First, enumerate TT, TH, HT, and HH. Each sequence has probability 1/4. However, two sequences produce exactly one head, so that count has twice the probability of either endpoint.
Answer: The probabilities are 0.25, 0.50, and 0.25. Their sum equals one.
| Heads, x | Eligible sequences | P(X = x) | F(x) |
|---|---|---|---|
| 0 | TT | 0.25 | 0.25 |
| 1 | TH, HT | 0.50 | 0.75 |
| 2 | HH | 0.25 | 1.00 |
As an additional check, the weighted mean equals 0(0.25) + 1(0.50) + 2(0.25) = 1. Similarly, the binomial formula gives np = 2(0.5) = 1. Both methods describe the same distribution.
Example 10: four heads in six tosses
Question: Find the probability of exactly four heads in six independent fair tosses, plus the mean and standard deviation.
Set n = 6, k = 4, and p = 0.5. Next, calculate the number of arrangements: 6!/(4!2!) = 15. Each six-toss sequence has probability 1/64.
μ = 6(0.5) = 3; σ = √1.5 ≈ 1.224745
Answer: Exactly four heads: 23.4375%; expected heads: 3; standard deviation: approximately 1.224745 heads.
For comparison, “at least four heads” includes four, five, and six. Therefore, its probability is (15 + 6 + 1)/64 = 22/64 = 0.34375. The word exactly changes the event, even though the model stays the same.
| Heads | 0 | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|---|
| Probability | 1/64 | 6/64 | 15/64 | 20/64 | 15/64 | 6/64 | 1/64 |
Poisson probability distributions: counts within an interval
A Poisson model describes a count in a specified interval. For example, the interval may represent one hour, one kilometer, or one inspected surface. The parameter λ gives the expected count for that interval, so changing the interval requires changing λ.
μ = λ; σ² = λ; σ = √λ
For a homogeneous Poisson process, counts in disjoint intervals are independent and the rate stays constant. However, clustered arrivals or changing demand may violate those assumptions. Check the process before treating a count as Poisson merely because it measures events.
Further reading: NIST/SEMATECH: Poisson Distribution.
Example 11: two calls in an hour
Question: A department averages five calls per hour. Under a Poisson model, find the probability of exactly two calls in one hour.
The interval is one hour, so λ = 5. Next, substitute k = 2 and use 2! = 2. Keep the exponential term unrounded until the final step.
Answer: Approximately 8.4224%. Using the book’s rounded e−5 ≈ 0.00674 gives 0.08425 instead.
For a half-hour interval, the expected count becomes 2.5 under the same constant-rate assumption. Consequently, you would use λ = 2.5 for that different question. The mean of five calls per hour does not mean that every hour contains five calls.
Normal probability distributions: areas under a curve
A normal distribution is symmetric and bell-shaped. The mean μ locates its center, while the positive standard deviation σ controls its spread. Unlike a discrete mass, a normal probability comes from an area between values.
Standardization expresses distance from the mean in standard deviations. Thus, z = 2 means two standard deviations above the mean, while z = −1 means one below it. For normal X, the standardized variable has a standard normal distribution.
Further reading: NIST/SEMATECH: Normal Distribution.
Read the correct type of normal table
Some tables give the cumulative area Φ(z) = P(Z ≤ z). Others report the area between zero and a positive z-score. Therefore, check the table heading before using its numbers. A value near 0.4750 at z = 1.96 refers to the mean-to-z area, whereas the cumulative area is approximately 0.9750.
Example 12: the area from zero to 1.96
Question: Find P(0 < Z < 1.96) for a standard normal variable.
First, obtain the cumulative probability at 1.96. Then subtract the 0.5 area to the left of zero. By symmetry, the interval from −1.96 to zero has the same probability.
Answer: Approximately 47.5002%, or 47.50% using a four-decimal table.
Example 13: an interval from 8 to 12
Question: Let X be normal with μ = 10 and variance σ² = 4. Find P(8 < X < 12).
First, take the square root of the variance: σ = 2. Next, standardize both bounds to obtain −1 and 1. Finally, subtract the cumulative probabilities.
P(8 < X < 12) = Φ(1) − Φ(−1) ≈ 0.682689
Answer: Approximately 68.2689%. The source’s rounded table calculation gives 68.26%.
Example 14: an interval from 7 to 14
Question: Using the same normal model, find P(7 < X < 14) and the probability outside that interval.
The standard deviation remains two. Therefore, the lower and upper z-scores are −1.5 and 2. Next, subtract the cumulative probabilities, then use the complement for the two outside tails.
P(X ≤ 7 or X ≥ 14) = 1 − 0.910443 ≈ 0.089557
Answer: Inside: approximately 91.0443%; outside: approximately 8.9557%. Four-decimal table arithmetic gives 91.04% and 8.96%.
The 68–95–99.7 rule
For a normal distribution, about 68.27% of values fall within one standard deviation of the mean. Similarly, about 95.45% fall within two and 99.73% within three. These are normal-model probabilities, so do not apply them automatically to skewed or heavy-tailed data.
| Interval | Standard normal event | Probability |
|---|---|---|
| μ ± σ | −1 ≤ Z ≤ 1 | 68.2689% |
| μ ± 2σ | −2 ≤ Z ≤ 2 | 95.4500% |
| μ ± 3σ | −3 ≤ Z ≤ 3 | 99.7300% |
Complete exercise solutions: Problems 3.1–3.9
The following solutions retain the source’s exercise numbers so you can compare each result with the photographs. However, the explanations use fresh wording and show the reasoning explicitly. Treat each subpart as its own event unless the question requests a joint probability.
Problem 3.1: three approaches to probability
Part (a): Classical probability assigns probabilities through a model of equally likely outcomes. Empirical probability estimates them from observed relative frequencies. In contrast, subjective probability expresses a degree of belief informed by evidence and judgment.
| Approach | Illustration | Main limitation |
|---|---|---|
| Classical | A fair die gives P(3) = 1/6. | Equal likelihood requires justification. |
| Empirical | 106 threes in 600 rolls gives 106/600. | Sampling noise and changing conditions affect estimates. |
| Subjective | An analyst estimates the chance of meeting a launch target. | Judgment can reflect bias or incomplete information. |
Part (b): The classical method becomes difficult when outcomes lack symmetry or known probabilities. Meanwhile, empirical estimates require relevant observations and can change across samples. Subjective estimates can differ across people, so transparent assumptions and calibration matter.
Part (c): We study probability to make decisions under uncertainty. For example, a production manager can estimate expected waste, while a service manager can plan staffing. Probability also supports statistical inference by linking observed data to possible underlying processes.
The classical method is not limited to games. For instance, a genuinely uniform random selection from a finite list also supports equally likely counting. Likewise, empirical probability need not converge to a classical fair-device value if the underlying device is biased.
Problem 3.2: coin, die, and complements
Part (a): A fair coin has two equally likely outcomes. Therefore, heads and tails each have probability 1/2. Their union covers every outcome in the model.
Part (b): A two occupies one of the six die faces. In contrast, “not two” includes five faces. Adding the event and its complement returns the full sample space.
Problem 3.3: five card probabilities
Assume a well-shuffled standard deck with 52 cards and no jokers. First, identify the number of eligible cards for each event. Then divide each count by 52, using complements where convenient.
| Part | Event | Calculation | Probability |
|---|---|---|---|
| (a) | A king | 4/52 = 1/13 | 7.6923% |
| (b) | A spade | 13/52 = 1/4 | 25% |
| (c) | The king of spades | 1/52 | 1.9231% |
| (d) | Not the king of spades | 1 − 1/52 = 51/52 | 98.0769% |
| (e) | King of spades or its complement | 1/52 + 51/52 | 100% |
Notice that a named card represents one outcome, whereas a rank or suit represents several. Consequently, “a king” and “the king of spades” have different probabilities. The final part is certain because an event and its complement exhaust all possibilities.
Problem 3.4: colored balls and odds
An urn contains five red balls, three blue balls, and two green balls. Assume each of the ten balls has an equal chance of selection. Therefore, each color probability equals its count divided by ten.
| Part | Requested event | Working | Answer |
|---|---|---|---|
| (a) | Red | 5/10 | 0.5 = 50% |
| (b) | Blue | 3/10 | 0.3 = 30% |
| (c) | Green | 2/10 | 0.2 = 20% |
| (d) | Not blue | 1 − 0.3 | 0.7 = 70% |
| (e) | Not green | 1 − 0.2 | 0.8 = 80% |
| (f) | Green or not green | 0.2 + 0.8 | 1 = 100% |
| (g) | Odds in favor of blue | 3 blue : 7 non-blue | 3:7 |
| (h) | Odds against blue | 7 non-blue : 3 blue | 7:3 |
Odds compare favorable outcomes with unfavorable outcomes. In contrast, probability compares favorable outcomes with all outcomes. Thus, odds of 3:7 correspond to probability 3/(3 + 7) = 0.3, not 3/7.
Problem 3.5: 106 threes in 600 rolls
Part (a): Divide the observed number of threes by the number of rolls. Then compare this empirical result with the fair-die model. The difference concerns observed frequency, not a change in the theoretical probability.
p = 1/6 ≈ 0.166667
p̂ − p = 0.01
The observed frequency equals approximately 17.6667%, which is one percentage point above 16.6667%. Moreover, a fair die produces an expected 600/6 = 100 threes in 600 rolls. The actual count of 106 is six above that expectation.
Part (b): Under independent rolls with a stable fair-die probability, the observed proportion converges toward 1/6 as the number of rolls grows. However, the approach need not be monotonic. A longer sample may temporarily move farther from 1/6 before later results bring it closer.
If the die is biased, repeated rolls instead reveal its actual probability of a three. Therefore, the law of large numbers does not make an unfair die fair. It concerns stabilization around the underlying probability.
Problem 3.6: defective production
Part (a): The observed defect rate equals 27 divided by 1,000. Therefore, the empirical probability of a defective item is 0.027, or 2.7%.
Part (b): Apply that rate to 1,600 items, assuming the rate remains relevant. Consequently, the expected daily number of defects is 43.2. Keep the decimal when reporting the mathematical expectation.
For a whole-item planning estimate, round to about 43 defective items per day. However, no individual day must produce exactly 43 or 44. The expectation summarizes average output across comparable days; independence is unnecessary for this expectation if every item has the same marginal defect probability.
Problem 3.7: classify event relationships
Part (a), mutually exclusive: A single die roll cannot equal both two and three. Therefore, these two events have an empty intersection. Similarly, one selected ball cannot be both red and blue under the stated color categories.
Part (b), not mutually exclusive: A card can be both an ace and a club. In particular, the ace of clubs belongs to both events. Thus, an addition calculation must account for their overlap.
Part (c), independent: In an independent coin-toss model, heads on the first toss does not change the second toss’s head probability. Consequently, the probability of two heads equals the product of the two marginal probabilities.
Part (d), dependent: Drawing cards without replacement changes the deck composition. For example, after drawing one ace, only three aces remain among 51 cards. The chance of another ace therefore changes from 4/52 to 3/51.
Problem 3.8: Venn diagrams and independence
Part (a): Draw two separate circles inside a rectangle to represent mutually exclusive events. The rectangle represents the sample space. Because the circles do not overlap, no outcome belongs to both events.
Part (b): Draw intersecting circles for events that share outcomes. Their overlap represents A ∩ B. However, the diagram alone does not establish independence, because independence depends on the probabilities.
Part (c): If both event probabilities are positive, mutually exclusive events are dependent. To see why, compare the joint probability of zero with the positive product P(A)P(B). Since the two quantities differ, independence fails.
For completeness, a zero-probability event creates an exception. If either marginal probability equals zero, disjoint events can satisfy the product definition of independence. Therefore, the positive-probability condition matters in a precise answer.
Problem 3.9: addition with disjoint outcomes
Each part combines outcomes that cannot occur together in the stated single draw or roll. Therefore, add the favorable counts and divide by the total. Check the strict inequalities carefully: “less than three” excludes three, while “more than three” also excludes three.
| Part | Question | Eligible outcomes | Calculation | Answer |
|---|---|---|---|---|
| (a) | Die result less than 3 | 1 or 2 | 2/6 | 1/3 ≈ 33.3333% |
| (b) | A heart or a club | 13 hearts + 13 clubs | 26/52 | 1/2 = 50% |
| (c) | A red or blue ball | 5 red + 3 blue | 8/10 | 4/5 = 80% |
| (d) | Die result greater than 3 | 4, 5, or 6 | 3/6 | 1/2 = 50% |
As a quick check on part (c), the only excluded color is green. Consequently, the complement method gives 1 − 2/10 = 0.8. Agreement between direct counting and the complement provides a useful arithmetic check.
How to choose probability distributions and avoid mistakes
A practical model-selection table
| Question structure | Useful model or rule | Check first |
|---|---|---|
| One event in a finite uniform sample space | Classical counting | Are outcomes equally likely? |
| A or B | Addition rule | Do the events overlap? |
| A and B | Multiplication rule | Do you need a conditional probability? |
| Number of successes in n trials | Binomial | Fixed n, independent trials, constant p |
| Count within an interval | Poisson | Appropriate interval and count-process assumptions |
| Continuous measurement with a bell-shaped model | Normal | Plausible shape, mean, and standard deviation |
| Success count in a finite sample without replacement | Hypergeometric | Population size and number of successes |
Start with the random variable rather than a familiar formula. For example, “number of customers who purchase among 20 contacts” suggests a different model from “number of arrivals in an hour.” Next, check the assumptions against how the observations arise.
Sampling without replacement and the binomial model
Sampling without replacement usually creates dependence in a finite population. Therefore, the exact count of selected successes follows a hypergeometric model when the population composition is fixed and the sample is uniform. A binomial approximation may work when the sample is small relative to the population, but it remains an approximation.
An unfair coin does not automatically require a hypergeometric distribution. If its tosses are independent and its head probability stays constant, a binomial model still applies with p different from 0.5. Thus, the sampling mechanism matters more than fairness alone.
Approximations require care
A Poisson distribution can approximate a binomial count when n is large and the success probability is small, with λ = np. If failure is rare instead, apply the approximation to the failure count. However, a convenient rule of thumb cannot guarantee accuracy for every tail probability.
A normal approximation to a discrete count also needs adequate spread and a continuity correction. For example, approximate P(X ≤ k) with a normal area ending at k + 0.5. Whenever exact software calculations are available, compare the approximation before relying on a borderline case.
Further reading: Penn State STAT 414: Approximations for Discrete Distributions.
Expected values support planning, not guarantees
Suppose a service team expects five calls per hour. That average helps plan capacity, but it does not describe the busiest hour. Similarly, an expected defect count of 43.2 says little about extreme daily outcomes without a model for variability.
For a business decision, pair the expected value with a relevant probability. For instance, a manager may care about the chance that demand exceeds available capacity. Consequently, the most useful calculation often concerns a tail, rather than the mean alone.
Common errors and their fixes
| Mistake | Why it fails | Better approach |
|---|---|---|
| Adding P(A) and P(B) without checking overlap | Shared outcomes count twice. | Subtract P(A ∩ B). |
| Multiplying marginal probabilities for dependent events | The second probability may change. | Use P(B | A). |
| Using variance in a z-score denominator | The denominator requires standard deviation. | Take the square root first. |
| Treating density height as probability | Continuous probability comes from area. | Integrate or subtract cumulative areas. |
| Confusing odds with probability | Their denominators differ. | Convert odds a:b into a/(a+b). |
| Rounding intermediate values heavily | Small errors accumulate. | Round at the final reporting step. |
| Reading exactly as at least | The event includes different outcomes. | Write the event symbolically. |
Chebyshev’s inequality: a broader spread guarantee
The normal percentages require a normal model. In contrast, Chebyshev’s inequality applies to any distribution with a finite mean and finite, positive variance. For k greater than one, it guarantees at least 1 − 1/k² of the probability within k standard deviations, using the corresponding strict-inside form.
For example, k = 2 gives a lower bound of 75%, while k = 3 gives 88.8889%. These bounds are weaker than normal-model percentages because they require much less information. Therefore, do not substitute a bound for an exact probability when a specific distribution is known.
Further reading: The Book of Statistical Proofs: Chebyshev’s inequality.
Frequently asked questions about probability distributions
What is the difference between probability and a probability distribution?
A probability describes the chance of one event. In contrast, a distribution describes the probabilities across the possible values of a random variable. You can use that distribution to calculate many different event probabilities.
Can an expected count contain a decimal?
Yes. An expected count is a weighted average, so it need not be a possible single observation. For example, 43.2 expected defects can summarize integer-valued daily counts over many comparable days.
When should I add probabilities instead of multiplying them?
Use the addition rule when the event asks for A or B. For an A-and-B event, use the multiplication rule. However, check overlap before adding and dependence before replacing conditional probabilities with marginal ones.
Does a long run of tails make heads more likely next?
No, provided the fair-coin tosses remain independent. The next head probability stays at 1/2. Therefore, the law of large numbers does not create a short-run force that compensates for earlier outcomes.
Why do my normal answers differ slightly from a printed table?
Printed tables often round areas to four decimal places. Consequently, their final percentages can differ slightly from software results that retain more precision. State your rounding method and compare the event definitions before assuming a substantive disagreement.
Can I assume that every measurement follows a normal distribution?
No. Many measurements show skewness, multiple peaks, or heavy tails. Therefore, inspect the data and the process before selecting a normal model. Standardizing a variable changes its location and scale, but it does not make an arbitrary distribution normal.
Putting probability into practice
Probability distributions become easier when you separate three tasks: define the event, choose a defensible model, and interpret the result. First, write down exactly what the question asks. Next, check equal likelihood, overlap, dependence, and the meaning of each parameter. Finally, report the probability with appropriate precision and a sentence that explains it.
The worked examples show how those habits connect elementary counting to binomial, Poisson, and normal calculations. Moreover, the exercises demonstrate why complements, conditional probabilities, and expected values matter beyond the classroom. Use the solution tables as a review aid, then redo the calculations without looking to check your understanding.
References and further study
The supplied chapter photographs provide the numerical prompts for Examples 1–14 and Problems 3.1–3.9. The sources below support further study of the probability rules and distribution formulas. All solutions, tables, and plotted calculations in this guide were independently developed for this article.
