Descriptive Statistics: 14 Solved Exercises
Descriptive statistics turns a list of numbers into information you can explain and use. For example, a frequency table reveals common quiz scores, while a median identifies the center of a wage distribution. This guide explains the main methods and solves 14 exercises step by step. Along the way, you will learn when an answer is exact, when grouping creates an estimate, and why different averages answer different questions.
The examples cover frequency distributions, histograms, mean, median, mode, weighted averages, geometric and harmonic means, quartiles, percentiles, and dispersion. Moreover, every solution includes the reasoning behind the calculation. You can therefore use the article as a study guide or as a practical reference for introductory economics and statistics.
What does descriptive statistics tell you?
A useful statistical summary answers three questions: What values occur? Where is the center? How much do the observations vary? First, tables and charts show the distribution. Next, averages describe its location. Finally, measures of dispersion describe the distances between observations or around a center.
Descriptive analysis summarizes the observations you actually have. However, a summary alone does not establish causation or prove that a sample represents a larger population. A wage average for 25 workers, for instance, describes those workers. Generalizing to an entire region requires a suitable sampling design and additional analysis.
| Question | Useful tool | Example interpretation |
|---|---|---|
| How often does a value occur? | Frequency table | Eight students scored 7. |
| What share falls in a category? | Relative frequency | 20% of students scored 7. |
| What is the center? | Mean or median | The average quiz score is 6. |
| What value occurs most often? | Mode | The most common score is 7. |
| How dispersed are the data? | Range, IQR, mean absolute deviation | The middle half spans 3.5 score points. |
| How do values accumulate? | Ogive or cumulative table | 75% of students scored 7 or less. |
Frequency tables: the foundation of descriptive statistics
Start with a clear definition of one observation. In the quiz example, one student contributes one score. In the wage example, one worker contributes one hourly wage. Consequently, the total frequency must equal the number of students or workers, not the number of distinct values.
Here, fi counts observations in one value or class, and n counts all observations. Therefore, absolute frequencies sum to n, while relative frequencies sum to 1. Cumulative frequencies add successive classes and end at n. Small rounding differences may affect displayed percentages, but they should not affect the underlying counts.
Reference: OpenStax: frequency tables and relative frequencies.
Class limits, class boundaries, and midpoints
Class limits describe the displayed endpoints of an interval. For example, a wage class of $3.50–$3.59 includes observed cent values from $3.50 through $3.59. A half-open interval such as [3.50, 3.60) gives an especially clear rule: include 3.50 and exclude 3.60. Thus, an observation cannot fall into two adjacent classes.
Alternatively, if wages represent measurements rounded to the nearest cent, real boundaries can fall halfway between cents. The same displayed class then has boundaries 3.495 and 3.595. Its midpoint equals 3.545. By comparison, the half-open interval [3.50, 3.60) has midpoint 3.55. Both conventions can support analysis, but you must state which one you use.
The supplied wage solutions use convenient midpoints of 3.55, 3.65, and so on. Accordingly, this guide uses half-open intervals with width 0.10 for its main wage calculations. A later comparison explains how cent-based real boundaries change the estimates. Keeping one convention throughout prevents hidden shifts in the results.
Exercise 2.1: organize the quiz scores
Forty students completed a quiz. First, sort the scores from smallest to largest. The following array preserves all 40 observations, including repeated values. As a result, you can check the median, quartiles, and frequency counts directly.
| Positions | Sorted scores |
|---|---|
| 1–10 | 2, 2, 2, 3, 3, 3, 4, 4, 4, 4 |
| 11–20 | 4, 5, 5, 5, 5, 5, 6, 6, 6, 6 |
| 21–30 | 6, 6, 7, 7, 7, 7, 7, 7, 7, 7 |
| 31–40 | 8, 8, 8, 8, 9, 9, 9, 9, 10, 10 |
Build the complete frequency distribution
| Score / midpoint | Class interval | Frequency | Relative frequency | Percent | Cumulative count | Cumulative percent |
|---|---|---|---|---|---|---|
| 2 | [1.5, 2.5) | 3 | 0.075 | 7.5% | 3 | 7.5% |
| 3 | [2.5, 3.5) | 3 | 0.075 | 7.5% | 6 | 15.0% |
| 4 | [3.5, 4.5) | 5 | 0.125 | 12.5% | 11 | 27.5% |
| 5 | [4.5, 5.5) | 5 | 0.125 | 12.5% | 16 | 40.0% |
| 6 | [5.5, 6.5) | 6 | 0.150 | 15.0% | 22 | 55.0% |
| 7 | [6.5, 7.5) | 8 | 0.200 | 20.0% | 30 | 75.0% |
| 8 | [7.5, 8.5) | 4 | 0.100 | 10.0% | 34 | 85.0% |
| 9 | [8.5, 9.5) | 4 | 0.100 | 10.0% | 38 | 95.0% |
| 10 | [9.5, 10.5) | 2 | 0.050 | 5.0% | 40 | 100.0% |
For example, the score 7 occurs eight times. Its relative frequency is 8 ÷ 40 = 0.20, or 20%. Furthermore, the cumulative frequency through score 7 equals 30. Therefore, 30 ÷ 40 = 75% of the class scored 7 or below.
The intervals center on integer scores, so the midpoint represents each observation exactly. Consequently, the frequency table loses no numerical information about these scores. This property explains why later frequency-weighted calculations reproduce the raw-data mean and mean absolute deviation.
Draw the four requested graphs
The frequency histogram shows counts, whereas the relative-frequency histogram shows proportions. Meanwhile, the frequency polygon joins the midpoint-frequency pairs. To close that polygon, add zero-frequency points at scores 1 and 11. The ogive starts at (1.5, 0) and plots cumulative counts at the successive upper boundaries.
Thus, the ogive points are (2.5, 3), (3.5, 6), (4.5, 11), (5.5, 16), (6.5, 22), (7.5, 30), (8.5, 34), (9.5, 38), and (10.5, 40). Each point answers a cumulative question. For instance, the point at 6.5 means that 22 students scored below 6.5, which means 6 or less for integer grades.
Reference: NIST: histograms, relative frequencies, and density normalization.
Exercise 2.2: organize and graph the hourly wages
The wage sample contains 25 workers. First, arrange the observations in ascending order. Then create eight classes of width $0.10, starting at $3.50. Because the highest observation is $4.26, the final interval must extend beyond that value.
| Positions | Sorted hourly wages ($) |
|---|---|
| 1–5 | 3.55, 3.60, 3.65, 3.75, 3.78 |
| 6–10 | 3.80, 3.85, 3.85, 3.88, 3.90 |
| 11–15 | 3.95, 3.95, 3.95, 3.96, 4.00 |
| 16–20 | 4.05, 4.05, 4.05, 4.06, 4.08 |
| 21–25 | 4.10, 4.15, 4.18, 4.25, 4.26 |
| Wage interval ($) | Midpoint ($) | Frequency | Relative frequency | Cumulative count | Cumulative percent |
|---|---|---|---|---|---|
| [3.50, 3.60) | 3.55 | 1 | 0.04 | 1 | 4% |
| [3.60, 3.70) | 3.65 | 2 | 0.08 | 3 | 12% |
| [3.70, 3.80) | 3.75 | 2 | 0.08 | 5 | 20% |
| [3.80, 3.90) | 3.85 | 4 | 0.16 | 9 | 36% |
| [3.90, 4.00) | 3.95 | 5 | 0.20 | 14 | 56% |
| [4.00, 4.10) | 4.05 | 6 | 0.24 | 20 | 80% |
| [4.10, 4.20) | 4.15 | 3 | 0.12 | 23 | 92% |
| [4.20, 4.30) | 4.25 | 2 | 0.08 | 25 | 100% |
For example, the interval [4.00, 4.10) contains six workers. Those workers earn 4.00, 4.05, 4.05, 4.05, 4.06, and 4.08. Therefore, the class accounts for 6 ÷ 25 = 24% of the sample. The final cumulative count equals 25, which confirms that the table includes every worker.
The wage histogram reaches its largest bar in [4.00, 4.10). However, a modal interval does not identify every raw-data mode. Grouping combines different wages in one class, so it hides some repeated values. This distinction matters when we compare the raw and grouped modes in Exercise 2.4.
How to read a histogram without misleading yourself
Equal-width bins allow straightforward comparisons of bar heights. However, unequal-width classes require extra care: a wider class can collect more observations simply because it covers more values. In that case, use frequency density or probability density so bar area represents the relevant count or proportion. Always inspect the horizontal scale before comparing bars.
Also, a histogram differs from a category bar chart. A histogram follows a numerical axis and adjacent intervals. A category chart compares separate labels, such as industries or payment methods. Consequently, changing histogram boundaries can change the apparent shape even when the observations remain identical.
Mean, median, and mode: three different centers
The arithmetic mean distributes the total equally across observations. The median identifies the middle of an ordered list. In contrast, the mode identifies the most frequent value. Each statistic answers a different question, so the best choice depends on the purpose of the analysis.
Use μ when the observations form the complete population you intend to describe. Use x̄ when they form a sample from a larger population. Nevertheless, both arithmetic means use the same operation: sum the observations and divide by their count. The notation identifies the role of the data rather than a different averaging rule.
An exact-value frequency table retains each distinct xi. By comparison, a grouped table substitutes a midpoint mi for every observation in an interval. Therefore, a grouped mean usually estimates the raw-data mean. It becomes exact only when the substitutions preserve the total, as they do for the integer quiz scores.
Reference: NIST: measures of location, including mean, median, and mode.
Exercise 2.3: find the center of the quiz scores
Raw-data mean, median, and mode
Because the class contains 40 students, the median uses positions 20 and 21. Both positions contain 6. Consequently, the median equals (6 + 6) ÷ 2 = 6. The score 7 appears eight times, more than any other score, so the mode equals 7.
Grouped interpolation and why it changes the answer
Here, L is the lower boundary of the median class, Fbefore is the cumulative count before it, fmedian is its frequency, and h is its width. For this exercise, those values are 5.5, 16, 6, and 1. Substitution gives the textbook-style interpolation below.
For the modal class, d1 equals its frequency minus the previous frequency. Similarly, d2 equals its frequency minus the next frequency. Here, L = 6.5, d1 = 8 − 6 = 2, and d2 = 8 − 4 = 4.
Exercise 2.4: find the center of the wage distribution
Calculate the exact statistics first
The sorted wage list has an odd number of observations. Therefore, the median occupies position (25 + 1) ÷ 2 = 13. That worker earns $3.95, so the median equals $3.95. Meanwhile, both $3.95 and $4.05 occur three times, which makes the raw distribution bimodal.
Estimate the statistics from the grouped table
| Midpoint m ($) | Frequency f | Product fm ($) |
|---|---|---|
| 3.55 | 1 | 3.55 |
| 3.65 | 2 | 7.30 |
| 3.75 | 2 | 7.50 |
| 3.85 | 4 | 15.40 |
| 3.95 | 5 | 19.75 |
| 4.05 | 6 | 24.30 |
| 4.15 | 3 | 12.45 |
| 4.25 | 2 | 8.50 |
| Total | 25 | 98.75 |
The grouped mean exceeds the exact mean by $0.004. Although both round to $3.95 at the cent level, they are not identical before rounding. Consequently, retain sufficient precision during calculations and round only when reporting the final result.
The grouped median lies in [3.90, 4.00), since its cumulative count rises from 9 to 14. The modal interval is [4.00, 4.10), with frequency 6. Thus, interpolation produces one modal estimate even though the original data have two modes. That difference illustrates the information cost of grouping.
What changes under cent-based real boundaries?
If you interpret the wages as rounded continuous measurements, the first boundaries become 3.495 and 3.595. Therefore, each midpoint and lower boundary shifts downward by $0.005. With the same frequencies and width, the grouped mean becomes $3.945, the median becomes $3.965, and the interpolated mode becomes $4.020. The raw-data mean, median, and modes do not change.
Exercise 2.5: compare mean, median, and mode
| Measure | Main strength | Main limitation | Typical use |
|---|---|---|---|
| Mean | Uses the magnitude of every observation and relates directly to totals | Extreme observations can move it substantially | Average costs, output, or total payroll per worker |
| Median | Describes the middle rank and resists large changes in a few extreme values | Does not directly recover a total from the count | Typical wages, income, and housing prices |
| Mode | Identifies the most frequent value and works with categorical data | May have ties; grouping can change the apparent peak | Most common product, response, or exact price |
For example, one exceptionally high wage can raise a mean without changing the middle worker’s wage. In that situation, report the median alongside the mean to show both perspectives. However, a median does not reveal the full distribution either. Two groups can share a median while having very different upper and lower tails.
Open-ended classes create another limitation. A class such as “$50 and above” lacks a finite upper boundary, so it has no defined midpoint for an ordinary grouped-mean estimate. Nevertheless, a grouped median may remain estimable when its class has known boundaries. The same logic applies to a modal class: identify what the table actually supports before choosing a formula.
Exercise 2.6: calculate the grouped mean by coding
Coding simplifies arithmetic when class midpoints follow a regular pattern. First, choose a convenient reference value A. Next, divide each midpoint’s difference from A by the common width h. For these wages, choose A = 3.85 and h = 0.10, which produces small integer codes.
| Midpoint ($) | Frequency | Code u | Product fu |
|---|---|---|---|
| 3.55 | 1 | -3 | -3 |
| 3.65 | 2 | -2 | -4 |
| 3.75 | 2 | -1 | -2 |
| 3.85 | 4 | 0 | 0 |
| 3.95 | 5 | 1 | 5 |
| 4.05 | 6 | 2 | 12 |
| 4.15 | 3 | 3 | 9 |
| 4.25 | 2 | 4 | 8 |
| Total | 25 | 25 |
This result exactly matches the direct midpoint calculation. However, coding does not restore information that grouping removed. It simply rewrites the same grouped arithmetic with easier numbers. Therefore, the answer remains an estimate of the raw-data mean of $3.946.
Exercise 2.7: calculate the weighted average wage
A firm pays 25 workers $4 per hour, 15 workers $6 per hour, and 10 workers $8 per hour. An ordinary average of the three wage rates would give each rate equal importance. Instead, weight each rate by the number of workers who receive it.
| Hourly wage ($) | Workers | Hourly payroll ($) |
|---|---|---|
| 4 | 25 | 100 |
| 6 | 15 | 90 |
| 8 | 10 | 80 |
| Total | 50 | 270 |
By comparison, (4 + 6 + 8) ÷ 3 = $6 treats the three categories as equally large. They are not: half the workforce earns $4. Therefore, $5.40 correctly describes average hourly pay across workers. If the workers contribute different numbers of hours and you need payroll per hour worked, use hours as the weights instead.
Exercise 2.8: geometric mean and compound inflation
The stated annual inflation rates are 2%, 5%, and 12.5%. This exercise requires a careful distinction between the geometric mean of the three positive rate numbers and the constant annual rate that reproduces their compounded price change. Those calculations answer different questions.
The geometric mean requested literally
Thus, 5% is the geometric mean of the three positive percentage magnitudes. Meanwhile, their arithmetic mean equals (2 + 5 + 12.5) ÷ 3 = 6.5%. Neither operation, applied in this way, directly calculates the equivalent compound annual inflation rate.
The correct constant annual compound rate
For compounded growth, first convert each percentage into a growth factor. The factors are 1.02, 1.05, and 1.125. Then multiply the factors, take the cube root, and subtract 1. Consequently, the economically relevant constant annual compound rate differs from 5%.
For example, a price index starting at 100 would rise to 102 after year one, 107.10 after year two, and 120.4875 after year three. Thus, cumulative inflation equals 20.4875%. A constant annual rate of approximately 6.4096% produces that same final index. By contrast, three years at 5% would produce only 115.7625.
Reference: SciPy: mathematical definition of the geometric mean.
Reference: U.S. Bureau of Labor Statistics: calculating percentage changes in price indexes.
Exercise 2.9: use the harmonic mean for average speed
A commuter travels 10 miles at 60 mph and another 10 miles at 15 mph. Because the two distances are equal, the harmonic mean gives the average speed. Nevertheless, the most direct method always starts with total distance divided by total time.
The arithmetic average, 37.5 mph, gives equal weight to the two speeds. However, the commuter spends 40 minutes at 15 mph and only 10 minutes at 60 mph. Therefore, the slower segment has much more influence on the journey’s average speed. The correct answer must reflect those unequal times.
For equal time intervals, the arithmetic mean of speeds would be appropriate. For unequal distances, use total distance divided by total time, or a distance-weighted harmonic mean. In either case, begin with the physical quantities rather than selecting an average by habit.
Reference: SciPy: mathematical definition of the harmonic mean.
Exercise 2.10: calculate quartiles, deciles, and percentiles
Quartiles divide an ordered distribution into four parts. Similarly, deciles divide it into ten parts, while percentiles divide it into 100 parts. These measures describe position rather than frequency or spread. However, different accepted algorithms can produce different values for a finite data set.
Quiz-score quantiles using the exercise convention
To match the supplied exercise, locate a quantile at np. When np is an integer, average observations at positions np and np + 1. Otherwise, use the observation at the next integer position. For Q₁, Q₂, and Q₃ here, the relevant positions are therefore 10.5, 20.5, and 30.5.
| Measure | Calculation / positions | Answer |
|---|---|---|
| Q₁ | (x₍₁₀₎ + x₍₁₁₎) / 2 = (4 + 4) / 2 | 4 |
| Q₂ | (x₍₂₀₎ + x₍₂₁₎) / 2 = (6 + 6) / 2 | 6 |
| Q₃ | (x₍₃₀₎ + x₍₃₁₎) / 2 = (7 + 8) / 2 | 7.5 |
| D₃ = P₃₀ | (x₍₁₂₎ + x₍₁₃₎) / 2 = (5 + 5) / 2 | 5 |
| P₆₀ | (x₍₂₄₎ + x₍₂₅₎) / 2 = (7 + 7) / 2 | 7 |
The photographed solution labels some positions informally, but the tied values make the final D₃ and P₆₀ answers unchanged. Nevertheless, Q₃ does depend on the algorithm. For example, a common linear interpolation rule based on 1 + (n − 1)p gives 7.25. State your convention so a software result does not appear to contradict a hand calculation.
Grouped wage quantiles
First, multiply the sample size by the desired proportion p. Next, find the first class whose cumulative frequency reaches that target. Finally, interpolate within that interval. This method assumes an even spread within the selected class, so it produces an estimate rather than an exact observed wage.
| Measure | Target np | Substitution | Estimate ($) |
|---|---|---|---|
| Q₁ | 6.25 | 3.80 + ((6.25 − 5) / 4) × 0.10 | 3.83125 |
| Q₂ | 12.5 | 3.90 + ((12.5 − 9) / 5) × 0.10 | 3.97000 |
| Q₃ | 18.75 | 4.00 + ((18.75 − 14) / 6) × 0.10 | 4.07917 |
| D₃ | 7.5 | 3.80 + ((7.5 − 5) / 4) × 0.10 | 3.86250 |
| P₆₀ | 15 | 4.00 + ((15 − 14) / 6) × 0.10 | 4.01667 |
For instance, Q₁ falls in [3.80, 3.90), because the cumulative count moves from 5 to 9 there. The target 6.25 lies 1.25 observations beyond the previous cumulative count. Therefore, the calculation travels 1.25 ÷ 4 of the way across that class. This explains both the numerator and denominator in the interpolation formula.
Exercise 2.11: calculate and interpret the range
For the quiz scores, the maximum is 10 and the minimum is 2. Therefore, R = 10 − 2 = 8 points. For the raw wages, R = 4.26 − 3.55 = $0.71. These results use actual observed extremes rather than class endpoints.
The photographed grouped wage table runs from the displayed limit $3.50 through $4.29. Subtracting those printed limits gives $0.79. However, the half-open classes used here cover [3.50, 4.30), a total span of $0.80. Neither value recovers the exact observed range of $0.71 when you only have grouped counts.
| Quantity | Calculation | Result |
|---|---|---|
| Exact quiz range | 10 − 2 | 8 points |
| Exact wage range | 4.26 − 3.55 | $0.71 |
| Span between printed wage limits | 4.29 − 3.50 | $0.79 |
| Coverage of half-open wage classes | 4.30 − 3.50 | $0.80 |
The range offers a fast summary of total spread. However, it uses only two observations and can change sharply when one extreme value changes. Consequently, pair it with a measure that describes more of the data. Open-ended classes also prevent a finite class-span calculation unless you add assumptions.
Exercise 2.12: calculate the interquartile range
For the quiz scores, use the same quartile convention as Exercise 2.10. Thus, IQR = 7.5 − 4 = 3.5 points. The quartile deviation, also called the semi-interquartile range, equals 3.5 ÷ 2 = 1.75 points. This measure describes half the width of the middle 50% interval.
If you round the quartiles first to $4.08 and $3.83, you obtain an IQR of $0.25 and QD of $0.125. The photographed solution follows that rounded route. However, the more precise estimates above avoid intermediate rounding. At ordinary reporting precision, the IQR is about $0.248 and the quartile deviation is about $0.124.
The IQR focuses on the middle half of the data, so very distant tail observations do not directly determine it. Nevertheless, it does not describe the entire distribution. For example, two wage distributions can share an IQR while having very different maximum wages. Use it with a center and a graph for a fuller summary.
Reference: NIST: measures of scale and statistical dispersion.
Exercise 2.13: calculate mean absolute deviation for scores
Mean absolute deviation measures the average distance from the mean. First, subtract the mean from every observation. Then take absolute values and average those distances. Without the absolute values, positive and negative deviations from the arithmetic mean would cancel.
| Score x | Frequency f | Distance |x − 6| | Weighted distance |
|---|---|---|---|
| 2 | 3 | 4 | 12 |
| 3 | 3 | 3 | 9 |
| 4 | 5 | 2 | 10 |
| 5 | 5 | 1 | 5 |
| 6 | 6 | 0 | 0 |
| 7 | 8 | 1 | 8 |
| 8 | 4 | 2 | 8 |
| 9 | 4 | 3 | 12 |
| 10 | 2 | 4 | 8 |
| Total | 40 | 72 |
Thus, a score lies 1.8 points from the mean on average when you measure distance with absolute values. The raw-data and frequency-table calculations agree because each score value remains exact. However, this result does not mean that every student scored between 4.2 and 7.8. It summarizes an average distance, not a guaranteed interval.
Exercise 2.14: calculate mean absolute deviation for wages
Use the grouped mean $3.95 and the same midpoints as before. Next, compute the absolute distance between each midpoint and $3.95. Multiply each distance by its class frequency, then divide the total by 25. Because midpoints replace individual wages, this calculation gives a grouped estimate.
| Midpoint ($) | Frequency | Absolute distance ($) | Weighted distance ($) |
|---|---|---|---|
| 3.55 | 1 | 0.40 | 0.40 |
| 3.65 | 2 | 0.30 | 0.60 |
| 3.75 | 2 | 0.20 | 0.40 |
| 3.85 | 4 | 0.10 | 0.40 |
| 3.95 | 5 | 0.00 | 0.00 |
| 4.05 | 6 | 0.10 | 0.60 |
| 4.15 | 3 | 0.20 | 0.60 |
| 4.25 | 2 | 0.30 | 0.60 |
| Total | 25 | 3.60 |
For comparison, the original wage observations have mean $3.946. Their absolute distances sum to $3.70, so the exact raw-data mean absolute deviation equals $3.70 ÷ 25 = $0.148. Therefore, grouping understates this measure by $0.004 in this particular example. Another data set could produce an error in the opposite direction.
Descriptive statistics answer key: all 14 exercises
| Exercise | Main result | Important distinction |
|---|---|---|
| 2.1 | Counts: 3, 3, 5, 5, 6, 8, 4, 4, 2 | Graphs must match these verified counts. |
| 2.2 | Counts: 1, 2, 2, 4, 5, 6, 3, 2 | Eight wage intervals, each $0.10 wide. |
| 2.3 | Exact mean 6; median 6; mode 7 | Interpolated median 6.1667; mode 6.8333. |
| 2.4 | Exact mean $3.946; median $3.95; modes $3.95 and $4.05 | Grouped mean $3.95; median $3.97; mode $4.025. |
| 2.5 | Choose mean, median, or mode by the question | No single center describes every feature. |
| 2.6 | Coded grouped mean $3.95 | Coding simplifies the same midpoint calculation. |
| 2.7 | Weighted mean $5.40 per hour | Use worker counts as weights. |
| 2.8 | Geometric mean of rate numbers 5% | Equivalent compound annual inflation 6.4096%. |
| 2.9 | Average speed 24 mph | Equal distances call for a harmonic mean. |
| 2.10 | Quiz: Q₁ 4, Q₂ 6, Q₃ 7.5, D₃ 5, P₆₀ 7 | Grouped wage quantiles appear in the detailed table. |
| 2.11 | Quiz range 8; raw wage range $0.71 | Grouped limits do not reveal actual extremes. |
| 2.12 | Quiz IQR 3.5; QD 1.75 | Grouped wage IQR ≈ $0.247917; QD ≈ $0.123958. |
| 2.13 | Quiz mean absolute deviation 1.8 | Exact from both raw data and score counts. |
| 2.14 | Grouped wage mean absolute deviation $0.144 | Raw-data value is $0.148. |
Common mistakes and how to avoid them
Treating every average as interchangeable
An arithmetic mean, weighted mean, geometric mean, and harmonic mean serve different purposes. Therefore, identify the quantity you need before choosing a formula. Worker counts determine an average across workers, growth factors determine compound growth, and total time determines average travel speed. A familiar formula can still answer the wrong question.
Losing track of exact values and estimates
A frequency table of exact quiz scores differs from a table of wage intervals. The first retains each distinct observed value, while the second replaces a collection of values with a class. Consequently, a midpoint estimate can look precise without being exact. Label grouped results clearly and preserve raw observations whenever possible.
Rounding too early
Keep several decimal places in means, interpolated quantiles, and intermediate products. Otherwise, later calculations can accumulate avoidable rounding errors. The wage quartile deviation illustrates this issue: half of the unrounded IQR differs slightly from half of the rounded IQR. Report the final precision that suits the question, but calculate with the underlying values.
Ignoring units and class definitions
Quiz deviations use score points, whereas wage deviations use dollars per hour. Relative frequencies have no physical unit, and percentages express proportions on a scale of 100. Moreover, class widths must use the same units as their boundaries. Clear labels make errors easier to detect before you interpret the results.
Assuming software must reproduce one quartile convention
Different programs and functions may select different finite-sample quantile definitions. Therefore, record the method as well as the answer when quartiles matter. If a result differs, compare the positions and interpolation rules before assuming a calculation is wrong. In this guide, the stated textbook convention controls the quiz answer key.
Applying descriptive statistics to economics and business
In economics, distribution matters because people and firms experience different outcomes. For example, a mean wage summarizes payroll per worker, while a median wage locates the middle worker. Meanwhile, a frequency table reveals how many workers cluster in each wage bracket. Using all three perspectives gives a more useful account than one isolated average.
In business, weighted averages help combine groups of different sizes. A company with many low-volume stores and a few large stores must decide whether it wants an average per store, per sale, or per customer. Each question implies different weights. Consequently, the denominator deserves as much attention as the numerator.
Growth analysis requires similar care. Annual changes accumulate through multiplication of growth factors, so a simple average of percentage rates does not generally reproduce the observed endpoint. Conversely, a descriptive average of annual rates can still answer a valid question if you label it correctly. State whether you want an average reported rate or an equivalent compound rate.
Finally, measures of spread help reveal how much variation an average hides. Two departments may share the same average wage but differ in wage range and IQR. As a result, a useful report combines a center, a spread measure, and a distribution chart. This combination supports interpretation without suggesting more certainty than the data provide.
Frequently asked questions about descriptive statistics
Can the mean and median be equal while the mode differs?
Yes. In the quiz data, the mean and median both equal 6, while the mode equals 7. Therefore, equality between two center measures does not require equality with the third. It also does not, by itself, prove that the entire distribution is symmetric.
Does a cumulative frequency table always end at 100%?
The cumulative count ends at the total number of observations. By contrast, cumulative relative frequency ends at 1, or 100% when you express it as a percentage. Thus, the final quiz count is 40 and the final wage count is 25, but both final cumulative percentages equal 100%.
Why does grouping change the median?
Grouping hides individual positions within a class. Therefore, an interpolation formula assumes a pattern inside the interval instead of observing the exact middle value. In the wage example, the raw median is $3.95, but the grouped estimate is $3.97. The difference reflects information loss and the interpolation assumption.
Is the geometric mean of inflation rates always useful?
A geometric mean of positive rate numbers is mathematically defined, but it does not generally express compound inflation. Instead, use the geometric mean of 1 plus each decimal rate, then subtract 1. This method also handles negative rates as long as all growth factors remain positive. A zero rate causes no difficulty for that growth-factor approach.
Is mean absolute deviation the same as standard deviation?
No. Mean absolute deviation averages absolute distances from a center. Standard deviation instead uses squared deviations and then takes a square root, with its denominator depending on the population or sample formula. Consequently, the two statistics can react differently to unusually large deviations. Do not substitute one value for the other.
What should I report when my data have two modes?
Report both modes and describe the raw distribution as bimodal. For these wages, the modes are $3.95 and $4.05. However, distinguish repeated-value modes from peaks in a grouped histogram. Changing bin boundaries may merge or separate apparent peaks without changing the underlying observations.
Conclusion: use the method that matches the question
Descriptive statistics works best when you connect each calculation to a clear question. First, organize and check the data. Next, choose a center and a spread measure that suit the distribution. Finally, explain the units, conventions, and limitations of the result.
These 14 solved exercises show why careful interpretation matters as much as arithmetic. Exact score counts preserve information, wage intervals introduce approximation, and compounded inflation requires growth factors. Therefore, a reliable solution does more than produce a number: it explains what that number measures and how the data support it.
