Descriptive Statistics: 34 Solved Exercises
Descriptive statistics turns a table of numbers into a useful explanation of a dataset. However, correct answers require more than entering values into a formula. You also need to choose the right denominator, preserve decimal precision, and understand what grouped observations can reveal.
This complete guide works through 34 exercises, from 2.15 through 2.48. Along the way, you will calculate variance, standard deviation, relative variability, skewness, kurtosis, covariance, averages, and percentiles. In addition, you will construct histograms, frequency polygons, and ogives from two frequency distributions.
Each solution includes the method, numerical substitution, and interpretation. Moreover, the calculations distinguish genuine errors from differences that come from rounding or class-boundary conventions. As a result, you can use the article as both a study guide and a practical reference for checking your own work.
Descriptive statistics foundations: choose the right formula
Population and sample choices in descriptive statistics
A population includes every observation in the group you want to describe. By contrast, a sample contains only some members of a wider population. Therefore, the same numerical observations can produce different variance estimates depending on your purpose.
For a population variance, divide the squared deviations by N. By comparison, the usual unbiased sample variance estimator divides by n − 1. Under independent sampling from a common distribution with finite variance, this correction makes the expected sample variance equal to the population variance. However, taking its square root does not automatically create an unbiased estimator of the population standard deviation.
With a frequency table, multiply each contribution by its frequency f. Thus, the total observation count equals Σf, not the number of rows. When rows represent intervals, use a representative value, usually the class midpoint, and label the result as an estimate.
Reference: OpenStax: Measures of the Spread of the Data · https://openstax.org/books/introductory-statistics-2e/pages/2-7-measures-of-the-spread-of-the-data
Class boundaries in grouped descriptive statistics
The grade calculations use the supplied representative scores 2 through 10. Similarly, the wage calculations retain the supplied representatives $3.55 through $4.25. The photographed pages do not include the earlier raw-data table, so the grade calculations use its displayed frequency reconstruction and totals.
For gasoline, the class $1.00–$1.04 has midpoint $1.02. Assuming prices rounded to the nearest cent, its continuous boundaries are $0.995 and $1.045. Therefore, the next class begins at $1.045, and every class has width $0.05. The gasoline graphs and interpolated quantiles consistently use those boundaries.
For family income, this guide uses $10,000–under $12,000, then $12,000–under $14,000, and so on. Consequently, the midpoints are $11,000, $13,000, and subsequent odd thousands. Treating the printed inclusive whole-dollar endpoints literally would shift all midpoints and interpolated locations downward by $0.50. That constant shift leaves central moments and absolute deviations unchanged.
A compact descriptive statistics formula sheet
| Measure | Formula | Main purpose |
|---|---|---|
| Grouped mean | Σfx / Σf | Locate the arithmetic center |
| Population variance | Σf(x − μ)² / N | Measure squared spread |
| Sample variance | Σf(x − x̄)² / (n − 1) | Estimate population variance |
| Mean absolute deviation | Σf|x − mean| / Σf | Average absolute distance from the mean |
| Coefficient of variation | 100 × standard deviation / mean | Measure relative spread |
| Grouped quantile | L + [(pN − C) / f]h | Interpolate a percentile |
| Moment skewness | g₁ = m₃ / m₂^(3/2) | Describe asymmetry |
| Pearson kurtosis | β₂ = m₄ / m₂² | Describe standardized fourth-moment behavior |
| Population covariance | Σ(x − μₓ)(y − μᵧ) / N | Describe joint variation |
Here, L denotes the lower boundary of the target class, C the cumulative frequency before it, and h its width. Also, p represents the desired fraction, such as 0.25 for the first quartile. Finally, mᵣ = Σf(x − mean)ʳ / Σf defines a central moment with the observation count in the denominator.
Exercise 2.15: Descriptive statistics for grades: variance and spread
Use scores 2 through 10 with frequencies 3, 3, 5, 5, 6, 8, 4, 4, and 2. First, calculate N = 40 and Σfx = 240. Therefore, the mean score equals 240 / 40 = 6.
| Score x | Frequency f | fx | x − 6 | f(x − 6)² |
|---|---|---|---|---|
| 2 | 3 | 6 | -4 | 48 |
| 3 | 3 | 9 | -3 | 27 |
| 4 | 5 | 20 | -2 | 20 |
| 5 | 5 | 25 | -1 | 5 |
| 6 | 6 | 36 | 0 | 0 |
| 7 | 8 | 56 | 1 | 8 |
| 8 | 4 | 32 | 2 | 16 |
| 9 | 4 | 36 | 3 | 36 |
| 10 | 2 | 20 | 4 | 32 |
| Total | 40 | 240 | 192 |
For part (a), the reconstructed raw scores have the same sum of squared deviations, 192. In part (b), the frequency calculation counts the same values without writing each repeated score separately. Consequently, both approaches produce identical results for this reconstruction.
For part (c), standard deviation has a practical advantage: it uses the original measurement unit. Thus, a spread of 2.19 points is easier to interpret than 4.8 squared points. Nevertheless, variance remains essential for algebraic decompositions and many statistical models.
Answer: Population variance = 4.8 points²; population standard deviation ≈ 2.19089 points.
Exercise 2.16: Sample variance and standard deviation of wages
The wage table contains 25 observations. First, use its representative values to obtain Σfx = $98.75 and x̄ = $3.95. Next, subtract that mean from each representative wage and square the deviations.
| Wage representative x | f | x − 3.95 | f(x − 3.95)² |
|---|---|---|---|
| $3.55 | 1 | -0.40 | 0.16 |
| $3.65 | 2 | -0.30 | 0.18 |
| $3.75 | 2 | -0.20 | 0.08 |
| $3.85 | 4 | -0.10 | 0.04 |
| $3.95 | 5 | 0.00 | 0.00 |
| $4.05 | 6 | 0.10 | 0.06 |
| $4.15 | 3 | 0.20 | 0.12 |
| $4.25 | 2 | 0.30 | 0.18 |
| Total | 25 | 0.82 |
Because the exercise treats these observations as a sample, the denominator is 24. Moreover, the calculation estimates the variance from grouped representatives rather than original wages. Keep the extra digits for later calculations, even though the final standard deviation rounds to $0.18.
Answer: Sample variance ≈ 0.03416667 dollars²; sample standard deviation ≈ $0.184842.
Exercise 2.17: Derive the computational variance formulas
Start with the sum of squared deviations. First, expand the square; then use Σx = Nμ. This substitution removes the cross-product term and yields a shorter computational expression.
For the sample version, the numerator follows the same algebra. However, the denominator remains n − 1 rather than n.
For frequency data, expand f(x − mean)² and sum over classes. Since Σf = N and Σfx = Nμ, the frequency-weighted population identity follows immediately. Similarly, replace N and μ with n and x̄ for the sample numerator.
These identities simplify hand calculations. However, subtracting two nearly equal large numbers can lose numerical precision in software. For that reason, a direct centered calculation or a stable variance algorithm is preferable for difficult numerical datasets.
Answer: Both shortcuts follow by expanding the square and substituting the definition of the mean.
Exercise 2.18: Check descriptive statistics with a variance shortcut
The grade frequency table gives Σfx² = 1,632. Also, N = 40 and μ = 6. Substitute these totals into the identity from Exercise 2.17.
For the ungrouped reconstruction, Σx² also equals 1,632. Therefore, part (a) and part (b) agree with Exercise 2.15. This second method provides a useful arithmetic check without changing the definition of variance.
Answer: Both calculations give variance 4.8 and standard deviation approximately 2.19089.
Exercise 2.19: Check wage variance with the shortcut
For the grouped wage values, Σfx² = 390.8825. Meanwhile, n = 25 and x̄ = 3.95, so nx̄² = 390.0625. Subtracting these values reproduces the centered sum of squares.
Thus, both the direct method and the shortcut produce the same sample estimate. In practice, a disagreement usually points to an incorrect frequency, premature rounding, or a denominator mismatch.
Answer: Sample variance ≈ 0.03416667 dollars²; sample standard deviation ≈ $0.184842.
Exercise 2.20: Relative variation in descriptive statistics
The coefficient of variation in descriptive statistics expresses standard deviation as a percentage of the mean. Therefore, it allows a relative comparison when the measurement scale has a meaningful zero and the mean is positive.
Numerically, the grades have greater spread relative to their mean. However, interpreting a CV for test scores requires care: a score of zero may not represent a true absence of the underlying skill. By comparison, a positive wage has a clearer ratio-scale interpretation.
For part (c), the CV helps compare relative variability across scales or groups with different means. Nevertheless, it becomes unstable near a zero mean and is generally unsuitable for interval scales such as Celsius temperature.
Reference: NIST: Coefficient of Variation · https://www.itl.nist.gov/div898/software/dataplot/refman2/auxillar/coefvari.htm
Answer: Grades: 36.51%; wages: 4.68%. The printed grade percentage is incorrect.
Exercise 2.21: Pearson’s second coefficient of skewness
Pearson’s second coefficient compares the mean and median, scaled by standard deviation. For this exercise, retain the grouped medians supplied on the photographed page: 6.17 for grades and $3.97 for wages. These are inputs from the earlier grouped-data work, not medians recomputed from the midpoint reconstruction.
Both coefficients are negative because each supplied median exceeds its mean. Therefore, this particular measure suggests left asymmetry. However, Pearson’s median-based coefficient and third-moment skewness are different statistics, so their numerical values need not match.
Answer: Using the supplied grouped medians: grades ≈ −0.2328; wages ≈ −0.3246.
Exercise 2.22: Correct third-moment skewness for grades and wages
For moment skewness, first calculate central moments using the same denominator. Specifically, define m₂ = Σf(x − mean)² / N and m₃ = Σf(x − mean)³ / N. Then standardize the third moment by m₂ raised to the power 3/2.
For grades, the weighted cubed deviations sum to −42. Therefore, m₃ = −42 / 40 = −1.05, while m₂ = 4.8.
For wages, the weighted cubed deviations sum to −0.054. Similarly, m₃ = −0.054 / 25 = −0.00216 and m₂ = 0.82 / 25 = 0.0328.
The wage distribution has more negative moment skewness in this comparison. Crucially, using the raw sum of cubed deviations without dividing by the count inflates the result. That omission explains the large negative coefficients printed in the source exercise.
Some software reports an adjusted sample skewness instead. In that convention, G₁ = √[n(n − 1)]g₁ / (n − 2). For example, the wage sample gives G₁ ≈ −0.38725. Always identify the convention before comparing outputs.
Answer: Unadjusted moment skewness: grades ≈ −0.099845; wages ≈ −0.363616.
Exercise 2.23: Kurtosis in descriptive statistics
Kurtosis uses the fourth central moment. First, divide the fourth-power sum by the observation count. Next, divide that moment by the square of the second central moment.
Consequently, the excess kurtosis values are approximately −0.825521 and −0.508923. Under the Pearson convention, both values lie below the normal-distribution benchmark of 3. Therefore, neither midpoint distribution supports the extreme leptokurtic description printed in the exercise.
Kurtosis concerns standardized fourth-moment behavior and tail extremity; a peak alone does not determine it. Moreover, sample adjustments can change the reported statistic. The calculations here consistently use unadjusted empirical moments.
Reference: NIST: Measures of Skewness and Kurtosis · https://www.itl.nist.gov/div898/handbook/eda/section3/eda35b.htm
Answer: Pearson kurtosis: grades ≈ 2.174479; wages ≈ 2.491077. Excess values equal those results minus 3.
Exercise 2.24: Covariance between wages and schooling
Now consider ten paired observations of hourly wages X and years of schooling Y. First, calculate x̄ = 11.775 and ȳ = 13.8. Next, multiply each wage deviation by the schooling deviation for the same employee.
| Employee | Wage X | Schooling Y | X − 11.775 | Y − 13.8 | Deviation product | XY |
|---|---|---|---|---|---|---|
| 1 | 8.50 | 12 | -3.275 | -1.8 | 5.895 | 102.00 |
| 2 | 12.00 | 14 | 0.225 | 0.2 | 0.045 | 168.00 |
| 3 | 9.00 | 10 | -2.775 | -3.8 | 10.545 | 90.00 |
| 4 | 10.50 | 12 | -1.275 | -1.8 | 2.295 | 126.00 |
| 5 | 11.00 | 16 | -0.775 | 2.2 | -1.705 | 176.00 |
| 6 | 15.00 | 16 | 3.225 | 2.2 | 7.095 | 240.00 |
| 7 | 25.00 | 18 | 13.225 | 4.2 | 55.545 | 450.00 |
| 8 | 12.00 | 18 | 0.225 | 4.2 | 0.945 | 216.00 |
| 9 | 6.50 | 12 | -5.275 | -1.8 | 9.495 | 78.00 |
| 10 | 8.25 | 10 | -3.525 | -3.8 | 13.395 | 82.50 |
| Total | 117.75 | 138 | 0 | 0 | 103.550 | 1,728.50 |
This answer treats the ten pairs as the entire descriptive population, as the exercise does. Alternatively, the usual sample covariance estimate divides by nine.
The positive sign indicates that higher wages and longer schooling tend to occur together in these observations. However, this association does not establish that education caused the wage differences. The covariance unit is hourly-wage dollars multiplied by years of schooling.
Answer: Population-style descriptive covariance = 10.355; sample covariance ≈ 11.505556.
Exercise 2.25: Covariance with the alternative formula
Using the same paired data gives ΣXY = 1,728.5. Therefore, the population-style covariance equals the mean of the products minus the product of the means.
For the sample calculation, first subtract n times the product of the means from ΣXY. Then divide the remaining 103.55 by n − 1. Thus, the shortcut agrees with the deviation-product calculation in Exercise 2.24.
Answer: The alternative formula gives the same descriptive covariance, 10.355.
Exercise 2.26: Visualize descriptive statistics for gasoline prices
The gasoline distribution contains 48 stations in six equal-width price classes. First, compute each relative frequency by dividing its count by 48. Next, accumulate the frequencies from the lowest class upward.
| Price class ($) | Midpoint | f | Relative frequency | Cumulative f | Upper boundary |
|---|---|---|---|---|---|
| 1.00–1.04 | 1.02 | 4 | 8.3333% | 4 | 1.045 |
| 1.05–1.09 | 1.07 | 6 | 12.5000% | 10 | 1.095 |
| 1.10–1.14 | 1.12 | 10 | 20.8333% | 20 | 1.145 |
| 1.15–1.19 | 1.17 | 15 | 31.2500% | 35 | 1.195 |
| 1.20–1.24 | 1.22 | 8 | 16.6667% | 43 | 1.245 |
| 1.25–1.29 | 1.27 | 5 | 10.4167% | 48 | 1.295 |
For the frequency histogram, draw touching bars over the continuous class boundaries. Their heights are 4, 6, 10, 15, 8, and 5. For the relative-frequency histogram, use the same boundaries and replace the counts with percentages.
Next, construct the frequency polygon by connecting the midpoint-count pairs. In addition, place zero-frequency anchors at $0.97 and $1.32 to close the line. For the ogive, begin at ($0.995, 0) and plot each upper boundary against its cumulative count.
The fourth class has the highest count, so it forms the modal interval. Meanwhile, the ogive reaches 48 at the upper boundary $1.295. Because the classes have equal widths, count heights support direct comparisons across the histogram.
Reference: OpenStax: Histograms and Frequency Polygons · https://openstax.org/books/introductory-statistics-2e/pages/2-2-histograms-frequency-polygons-and-time-series-graphs
Answer: All four graphs appear above; relative frequencies total 100%, and cumulative frequency ends at 48.
Exercise 2.27: Four graphs for family incomes
The family-income distribution contains 100 observations. Therefore, each class count also equals its relative frequency expressed as a percentage. Use a class width of $2,000 and the rounded continuous intervals described earlier.
| Income interval ($) | Midpoint ($) | f / percent | Cumulative f |
|---|---|---|---|
| 10,000–under 12,000 | 11,000 | 12 / 12% | 12 |
| 12,000–under 14,000 | 13,000 | 14 / 14% | 26 |
| 14,000–under 16,000 | 15,000 | 24 / 24% | 50 |
| 16,000–under 18,000 | 17,000 | 15 / 15% | 65 |
| 18,000–under 20,000 | 19,000 | 13 / 13% | 78 |
| 20,000–under 22,000 | 21,000 | 7 / 7% | 85 |
| 22,000–under 24,000 | 23,000 | 6 / 6% | 91 |
| 24,000–under 26,000 | 25,000 | 4 / 4% | 95 |
| 26,000–under 28,000 | 27,000 | 3 / 3% | 98 |
| 28,000–under 30,000 | 29,000 | 2 / 2% | 100 |
For the two histograms, use class counts and percentages respectively. Next, connect midpoint frequencies for the polygon, adding zero anchors at $9,000 and $31,000. Finally, start the ogive at ($10,000, 0) and plot cumulative counts at the successive upper boundaries.
The highest bar covers $14,000–under $16,000. However, the distribution extends farther to the right of that class than to the left. This pattern anticipates the positive skewness calculated later.
Answer: The four income graphs use the same 100 families; the final ogive point is ($30,000, 100).
Exercise 2.28: Mean, median, and mode of gasoline prices
First, calculate the mean using class midpoints. The weighted sum is Σfx = 55.36, so the estimated mean price is $1.153333 per gallon.
For the median, locate observation position N/2 = 24. Because the cumulative count rises from 20 to 35 in the fourth class, that class contains the median. Its lower boundary is 1.145, its frequency is 15, and its width is 0.05.
For the interpolated mode, use the modal frequency 15 and neighboring frequencies 10 and 8. Accordingly, the two frequency differences are 5 and 7.
Rounded to cents, the three answers are $1.15, $1.16, and $1.17. However, preserve the unrounded estimates when calculating later measures. Otherwise, small rounding changes can noticeably alter a skewness coefficient.
Reference: OpenStax: Measures of the Center of the Data · https://openstax.org/books/introductory-statistics-2e/pages/2-5-measures-of-the-center-of-the-data
Answer: Mean ≈ $1.153333; median ≈ $1.158333; interpolated mode ≈ $1.165833.
Exercise 2.29: Descriptive statistics for income: mean, median, and mode
The income midpoints produce Σfx = $1,700,000. Therefore, the estimated arithmetic mean equals $17,000. For the median, the cumulative frequency reaches 50 at the end of the $14,000–under $16,000 class.
The same class contains the modal frequency of 24. Its neighboring frequencies are 14 and 15. Consequently, the interpolated mode sits slightly above the class midpoint.
Here, mean exceeds median, and median exceeds the interpolated mode. That ordering is consistent with the distribution’s rightward extension. Nevertheless, such ordering is a descriptive clue rather than a universal theorem about every possible dataset.
Answer: Mean ≈ $17,000; median ≈ $16,000; mode ≈ $15,052.63, using rounded income intervals.
Exercise 2.30: Find both means by coding
For grouped descriptive statistics, coding simplifies arithmetic by shifting and rescaling the representatives. Let u = (x − A)/h, where A is a convenient reference and h is the class width. Then the original mean equals A + h(Σfu / N).
For gasoline, choose A = 1.12 and h = 0.05. The coded values are −2, −1, 0, 1, 2, and 3. Their frequency-weighted sum equals 32.
For incomes, choose A = 17,000 and h = 2,000. The coded values run from −3 through 6, and their weighted sum equals zero. Therefore, the chosen reference value already equals the grouped mean.
Answer: Coding reproduces the gasoline mean $1.153333 and the income mean $17,000.
Exercise 2.31: Weighted average hourly wage
A firm pays 5/12 of its workers $5 per hour, 1/3 of its workers $6, and 1/4 of its workers $7. First, verify that the workforce shares sum to one. Then multiply each wage by the corresponding share.
The answer represents an employee-weighted average wage. However, if workers supply different numbers of hours, a payroll-per-hour calculation requires hour weights instead. Thus, choosing the weight is part of defining the question.
Answer: The employee-weighted average wage is approximately $5.83 per hour.
Exercise 2.32: Descriptive statistics of investment returns
The three annual rates are 1%, 4%, and 16%. First, calculate their arithmetic mean. Next, calculate the geometric mean of the three positive rate numbers, as requested by the exercise.
These calculations answer two different mathematical questions. However, the geometric mean of the rate numbers is not the compound annual investment return. For a reinvested investment, multiply growth factors instead.
The wording specifies the same amount of capital in each year. Under that equal-principal interpretation, the arithmetic mean of 7% correctly summarizes the average annual return on that fixed principal. By contrast, if the investment compounds through all three years without external cash flows, 6.808111% is the appropriate constant annual growth rate.
Therefore, the exercise’s preference for 4% does not provide a valid compound-return interpretation. For example, $100 compounded at the three stated rates grows to $121.8464, which a 4% annual rate would not reproduce.
Reference: Investor.gov: How compound growth works · https://www.investor.gov/introduction-investing/investing-basics/save-and-invest/small-savings-add-big-money
Answer: Arithmetic mean = 7%; geometric mean of the positive rates = 4%; compound annual return ≈ 6.808111%. Use the measure that matches the capital assumption.
Exercise 2.33: Average speed over unequal distances
A plane travels 200 miles at 600 miles per hour and another 100 miles at 500 miles per hour. Average speed equals total distance divided by total travel time. Therefore, calculate time for each segment before combining them.
Equivalently, this is a distance-weighted harmonic mean of the two speeds. However, the ordinary harmonic mean assumes equal distances, which this problem does not have. The simple arithmetic mean of 550 mph also answers a different question.
Answer: Average speed = 562.5 mph.
Exercise 2.34: Average gasoline price for equal spending
A driver spends $10 at $0.90 per gallon and another $10 at $1.10 per gallon. First, calculate the gallons purchased at each price. Then divide total spending by total gallons.
Because spending is equal, the driver purchases more gallons at the lower price. As a result, the effective price falls below the arithmetic midpoint of $1.00. In this equal-spending case, the harmonic mean gives the correct result.
Answer: The effective average gasoline price is exactly $0.99 per gallon.
Exercise 2.35: Gasoline quartiles, fourth decile, and 70th percentile
Use the cumulative frequencies 4, 10, 20, 35, 43, and 48. First, multiply each target proportion by 48 to locate its position. Next, interpolate within the class that contains that position.
| Measure | Position pN | Substitution | Estimate ($) |
|---|---|---|---|
| Q₁ | 12 | 1.095 + [(12 − 10)/10] × 0.05 | 1.105000 |
| Q₂ | 24 | 1.145 + [(24 − 20)/15] × 0.05 | 1.158333 |
| Q₃ | 36 | 1.195 + [(36 − 35)/8] × 0.05 | 1.201250 |
| D₄ | 19.2 | 1.095 + [(19.2 − 10)/10] × 0.05 | 1.141000 |
| P₇₀ | 33.6 | 1.145 + [(33.6 − 20)/15] × 0.05 | 1.190333 |
For example, the first quartile lies two observations into a class containing ten stations. Therefore, it sits one-fifth of the way across that class. Similarly, the 70th percentile lies 13.6 observations into the fourth class under linear interpolation.
Using printed lower class limits instead of continuous boundaries raises each estimate by $0.005. Consequently, that alternative gives D₄ = $1.146 and P₇₀ ≈ $1.195333. This difference reflects a boundary convention; it does not justify mixing conventions within one calculation.
Answer: Q₁ = $1.105; Q₂ ≈ $1.158333; Q₃ = $1.20125; D₄ = $1.141; P₇₀ ≈ $1.190333.
Exercise 2.36: Income quantiles in descriptive statistics
For 100 families, the requested percentile positions are 25, 75, 30, and 60. Next, identify the class containing each position from the cumulative table. Apply the same interpolation formula with width $2,000.
| Measure | Substitution | Estimate ($) |
|---|---|---|
| Q₁ | 12,000 + [(25 − 12)/14] × 2,000 | 13,857.14 |
| Q₃ | 18,000 + [(75 − 65)/13] × 2,000 | 19,538.46 |
| D₃ | 14,000 + [(30 − 26)/24] × 2,000 | 14,333.33 |
| P₆₀ | 16,000 + [(60 − 50)/15] × 2,000 | 17,333.33 |
For instance, the first quartile falls in the second income class. Its position is 13 observations beyond the first class’s cumulative total of 12. Therefore, it lies 13/14 of the way across that interval.
Answer: Q₁ ≈ $13,857.14; Q₃ ≈ $19,538.46; D₃ ≈ $14,333.33; P₆₀ ≈ $17,333.33.
Exercise 2.37: Range and the limits of grouped data
Range equals the largest observed value minus the smallest observed value. However, a grouped table generally hides both exact values. Therefore, distinguish the span of the displayed intervals from the actual sample range.
The outer continuous gasoline boundaries span $0.30, from $0.995 to $1.295. Similarly, the rounded income intervals span $20,000. These describe the coverage of the bins, not exact observed ranges.
Because both end classes contain observations, the gasoline data imply an observed range between $0.21 and $0.29 under the printed cent-valued limits. Likewise, the inclusive whole-dollar income classes imply a range between $16,001 and $19,999. The individual observations would determine the exact value.
Answer: Displayed class-limit spans: gasoline $0.29; income $19,999, approximately $20,000. Exact observed ranges are unknown.
Exercise 2.38: Interquartile range and quartile deviation
In descriptive statistics, the interquartile range measures the distance between the third and first quartiles. Next, divide it by two to obtain the quartile deviation, also called the semi-interquartile range.
Thus, the middle half of the income distribution spans approximately $5,681.32 under grouped interpolation. The printed income answer of $476 does not follow from the quartiles given in the exercise. Moreover, shifting all interval boundaries by the same constant does not change the interquartile range.
Answer: Gasoline: IQR = $0.09625, QD = $0.048125. Income: IQR ≈ $5,681.32, QD ≈ $2,840.66.
Exercise 2.39: Mean absolute deviation from the mean
Average deviation in this exercise means the mean absolute deviation about the arithmetic mean. First, calculate each absolute deviation. Then multiply by its frequency and divide the total by the observation count.
The gasoline calculation uses its full-precision mean, 1.153333333. However, using the rounded mean 1.15 produces $0.0575, which explains the printed answer. Thus, the discrepancy arises from early rounding rather than a different definition.
Also, mean absolute deviation differs from median absolute deviation. The latter takes a median of distances from a median and serves a different purpose. Always state the center and averaging operation when the abbreviation MAD could be ambiguous.
Answer: Mean absolute deviation: gasoline ≈ $0.05694444; income = $3,520.
Exercise 2.40: Descriptive statistics of gasoline-price dispersion
Treat the 48 stations as the descriptive population specified for this example. First, subtract the full-precision mean 1.153333333 from each midpoint. The frequency-weighted squared deviations total approximately 0.231666667.
Therefore, the midpoint-based standard deviation is approximately 6.95 cents per gallon. This number summarizes spread around the mean across the stations. However, it does not describe the uncertainty of the estimated mean, which is a separate concept.
Answer: Population variance ≈ 0.0048263889 dollars²; standard deviation ≈ $0.06947222.
Exercise 2.41: Sample variance and standard deviation of incomes
The income exercise explicitly describes a sample of 100 families. Therefore, use the sample denominator 99. The squared deviations from the grouped mean $17,000 sum to 1,976,000,000.
By contrast, dividing by 100 gives a descriptive population variance of 19,760,000 and standard deviation $4,445.222154. Those are the values printed in the exercise, despite its sample notation. Both computations have uses, but they require different labels.
Answer: Sample variance ≈ 19,959,595.96 dollars²; sample standard deviation ≈ $4,467.62.
Exercise 2.42: Gasoline variance with the computational formula
The gasoline table gives Σfx = 55.36 and Σfx² = 64.0802. First, divide the second total by 48. Then subtract the square of the unrounded mean.
As expected, the shortcut matches Exercise 2.40. In addition, this agreement checks the frequency weighting and the midpoint totals. A variance cannot become negative through correct exact arithmetic, so a negative result would signal an error or numerical precision problem.
Answer: Variance ≈ 0.0048263889 dollars²; standard deviation ≈ $0.06947222.
Exercise 2.43: Income sample variance with the computational formula
For incomes, Σfx² = 30,876,000,000 and x̄ = 17,000. Therefore, the centered sum of squares equals that total minus 100 times the squared mean. Divide by 99 for the sample estimate.
Taking the square root gives $4,467.616362. Thus, the direct calculation and the shortcut agree. The result differs from the printed sample answer because that answer uses the population denominator.
Answer: Sample variance ≈ 19,959,595.96 dollars²; sample standard deviation ≈ $4,467.62.
Exercise 2.44: Compare descriptive statistics across two distributions
For gasoline, use the population standard deviation and its grouped mean. Meanwhile, the income sample requires the sample standard deviation and its grouped mean. Then multiply both ratios by 100 to express them as percentages.
Therefore, family incomes have substantially greater relative variability. Using the descriptive population denominator for incomes would instead give 26.148366%. However, that alternative does not change which distribution has the larger relative spread.
Answer: Gasoline CV ≈ 6.02%; income sample CV ≈ 26.28%. Incomes show greater relative dispersion.
Exercise 2.45: Pearson skewness for gasoline and incomes
Use Pearson’s second coefficient with the medians calculated earlier. First, preserve the unrounded gasoline mean and median. Otherwise, their small difference can change substantially after rounding to cents.
Thus, the gasoline measure is negative and the income measure is positive. The printed gasoline answer near −0.43 follows from rounded inputs and differs from the consistent full-precision calculation. Moreover, changing the median’s boundary convention without changing the rest of the model creates another source of disagreement.
Answer: Pearson skewness: gasoline ≈ −0.215914; income sample ≈ 0.671499.
Exercise 2.46: Third-moment skewness for gasoline and incomes
Return to the empirical-moment convention from Exercise 2.22. First, compute both central moments with denominator N, even when the observations constitute a sample. That convention defines the unadjusted descriptive coefficient g₁; an adjusted estimator requires a separate correction.
| Dataset | m₂ | m₃ | g₁ = m₃ / m₂^(3/2) |
|---|---|---|---|
| Gasoline | 0.004826388889 | −0.000061342593 | −0.182948 |
| Income | 19,760,000 | 66,720,000,000 | 0.759584 |
These results support mild negative asymmetry for gasoline and positive asymmetry for income. However, their magnitudes differ from Pearson’s median-based measure because the formulas summarize different features. The printed income value 755 is not a correctly normalized moment coefficient for these data.
Answer: Moment skewness: gasoline ≈ −0.182948; income ≈ 0.759584.
Exercise 2.47: Kurtosis for gasoline and incomes
First, calculate fourth central moments using full-precision means. Next, divide by the square of each second central moment. This produces Pearson kurtosis, whose normal benchmark is 3.
Therefore, gasoline has excess kurtosis approximately −0.628694. Income has excess kurtosis approximately 0.000377, which is extremely close to zero under this midpoint approximation. However, matching the normal kurtosis benchmark does not establish a normal distribution, particularly because income still has positive skewness.
The printed coefficients 177 and 300 do not match the normalized calculations. In particular, the income result near 300 reflects a fourth-power sum without the necessary division by 100. Correct normalization brings that result close to 3.
Answer: Pearson kurtosis: gasoline ≈ 2.371306; income ≈ 3.000377. Excess kurtosis: −0.628694 and 0.000377.
Exercise 2.48: Interpret the sign and bounds of covariance
Positive covariance indicates that the centered variables tend to move together. Conversely, negative covariance indicates opposite movement around their respective means. A covariance of zero indicates no linear association as measured by this statistic.
| Relationship | Covariance sign | Interpretation |
|---|---|---|
| Positive linear association | Cov(X,Y) > 0 | Above-average values tend to occur together |
| Negative linear association | Cov(X,Y) < 0 | Above-average values tend to pair with below-average values |
| No linear association | Cov(X,Y) = 0 | Centered cross-products balance on average |
However, zero covariance does not imply independence. For example, let X take −1, 0, and 1 with equal probabilities, and set Y = X². Then Y depends completely on X, yet E(X) = 0 and E(XY) = E(X³) = 0, so their covariance equals zero.
Unlike correlation, covariance has no universal scale-free limits of −1 and 1. Nevertheless, for fixed finite standard deviations, the Cauchy–Schwarz inequality provides a bound.
Reference: Penn State STAT 414: The Correlation Coefficient · https://online.stat.psu.edu/stat414/Lesson18
Answer: Positive association: covariance > 0; negative association: covariance < 0; zero linear association: covariance = 0. Independence is a stronger condition.
Descriptive statistics answer corrections at a glance
The table below separates substantive corrections from choices about denominators, precision, and interval boundaries. Therefore, it helps explain why a sound calculation may differ from a printed answer key. Whenever you compare results, first check the definition and inputs on both sides.
| Exercise | Issue in the supplied answer | Recalculated result or clarification |
|---|---|---|
| 2.20(a) | Incorrect grade CV arithmetic | 36.51% |
| 2.22 | Unnormalized third-power sums | g₁: −0.099845 grades; −0.363616 wages |
| 2.23 | Unnormalized fourth-power sums | β₂: 2.174479 grades; 2.491077 wages |
| 2.32 | Rate geometric mean confused with growth | 4% rate GM; 6.808111% compounded growth; 7% equal-principal mean |
| 2.35 | Lower-limit versus continuous-boundary convention | Main estimates use cent-based boundaries |
| 2.37 | Class span treated as exact observed range | Raw extrema remain unknown |
| 2.38(b) | Income IQR inconsistent with quartiles | IQR ≈ $5,681.32; QD ≈ $2,840.66 |
| 2.39(a) | Mean rounded before absolute deviations | MAD ≈ $0.05694444 |
| 2.41 and 2.43 | Population denominator labeled as sample | s² ≈ 19,959,595.96; s ≈ $4,467.62 |
| 2.45(a) | Sensitivity to rounded location values | Pearson coefficient ≈ −0.215914 |
| 2.46 | Incorrect moment-skewness answers | g₁ ≈ −0.182948 gasoline; 0.759584 income |
| 2.47 | Incorrect kurtosis normalization | β₂ ≈ 2.371306 gasoline; 3.000377 income |
How to use descriptive statistics in economics and business
For additional background, review the complete guide to statistical measures and examples. Then return to these exercises to apply the concepts step by step.
The wage and income exercises illustrate why a single average rarely tells the full story. For example, two labor markets may share the same mean wage but differ substantially in wage dispersion. Consequently, a useful comparison often reports the mean, median, standard deviation, and selected percentiles together.
Similarly, the gasoline example separates the typical price from price consistency across stations. A low mean helps describe affordability, while a small standard deviation indicates that prices cluster closely. However, a historical teaching table cannot establish current market conditions or explain why stations charge different prices.
Within descriptive statistics, covariance adds information about relationships between variables. For instance, the wage-schooling data show positive joint variation, but they do not isolate education from experience, occupation, or other influences. Therefore, causal analysis requires a research design and assumptions beyond descriptive statistics.
For readers moving into econometrics, these distinctions matter. Grouped-data approximations, sample definitions, and measurement units can all influence later analysis. As a result, documenting the construction of each statistic is as important as reporting its numerical value.
Common questions about descriptive statistics
Why do sample and population standard deviations differ?
They divide the same centered sum of squares by different quantities. Specifically, population variance uses N, while the conventional sample estimate uses n − 1. Consequently, the sample standard deviation is slightly larger when both calculations use the same observations.
Can a grouped-data answer be exact?
A calculation can be exact for the chosen representatives while still approximating the original data. For example, the wage table treats every observation within a class as its representative wage. Therefore, the resulting variance need not equal the variance of the unavailable individual wages.
Why can two skewness answers both be valid?
Pearson’s second coefficient uses the mean and median, whereas moment skewness uses cubed deviations. In addition, some software adjusts moment skewness for sample size. Thus, compare coefficients only after confirming the formula, denominator, precision, and median convention.
Does kurtosis equal 3 or 0 for a normal distribution?
Pearson kurtosis equals 3 for a normal distribution. By contrast, excess kurtosis subtracts 3 and therefore equals zero. Always label the convention so readers can interpret the number correctly.
Should intermediate results use rounded values?
For reliable descriptive statistics, keep full precision during calculations and round only the final reported result. Otherwise, errors can become substantial when subtracting nearby means or raising deviations to high powers. The gasoline skewness exercises demonstrate this sensitivity particularly clearly.
Does zero covariance prove that variables are unrelated?
No. Zero covariance rules out linear co-movement as summarized by that measure, but a nonlinear relationship may remain. For example, the relationship Y = X² can have zero covariance when X has a symmetric distribution around zero.
Which average should I use for rates?
First, identify the quantity that the problem holds constant. Equal time intervals support time-weighted arithmetic speed averages, while equal travel distances lead to harmonic averages. Similarly, investment compounding requires geometric averaging of growth factors, not simply the percentage-rate numbers.
A reliable workflow for solving descriptive statistics exercises
- Define the data: identify observations, units, frequencies, and whether the task concerns a population or sample.
- Set the conventions: state class boundaries, representative values, and any interpolation assumption.
- Check totals: confirm Σf, Σfx, and cumulative counts before calculating more complicated measures.
- Preserve precision: retain full numerical values until the final rounding step.
- Verify the definition: distinguish population and sample variance, alternative skewness measures, and Pearson versus excess kurtosis.
- Interpret the answer: explain what the statistic reveals and which conclusions the data cannot support.
These 34 descriptive statistics exercises show how careful definitions turn routine arithmetic into reliable analysis. Together, the solutions cover center, spread, distribution shape, and relationships between variables. Most importantly, they demonstrate how to check an answer rather than relying on a printed result.
When you approach a new dataset, begin with its structure and measurement scale. Next, choose formulas that match the question and document any approximation. Finally, report the numerical answer alongside a clear interpretation so the reader can understand what the data actually show.
