Learn descriptive statistics with clear formulas, worked examples, frequency tables, mean, variance, skewness, and covariance.

Descriptive Statistics: Complete Guide With Examples

Descriptive statistics help you turn raw data into useful information. They show where values cluster, how much observations vary, and what patterns deserve attention. For example, a business can summarize customer spending, while an economist can compare incomes across regions. However, a single average rarely tells the whole story.

Imagine two delivery companies that both report an average delivery time of 30 minutes. One usually delivers within a narrow window. In contrast, the other alternates between very fast deliveries and long delays. Their averages match, yet customers experience different levels of reliability.

This guide explains how to organize data, calculate common measures, and interpret the results. In addition, it covers grouped data, distribution shape, covariance, and practical mistakes. Worked examples connect the formulas to decisions in education, business, economics, and everyday life.

The examples use small, illustrative datasets so you can check every calculation. Therefore, they teach the methods without requiring specialized software or advanced mathematics.

What Are Descriptive Statistics?

Descriptive statistics summarize the observations you have collected. A table can organize their frequencies, a chart can reveal their distribution, and a numerical measure can describe their center or spread.

For a useful starting point, see the OpenStax introduction to descriptive statistics.

The Main Questions a Data Summary Should Answer

Before calculating anything, identify the question you want to answer. Otherwise, you may produce a correct number that contributes little to the analysis.

A useful descriptive report addresses these questions:

  • How many valid observations do we have?
  • Which values or categories occur most often?
  • What represents a typical observation?
  • How widely do observations vary?
  • Does the distribution have a long tail or several clusters?
  • How do two variables move together?

For example, an online store might want to understand order value. Its report could include the number of orders, median spending, mean spending, and the share of orders above a chosen threshold. Meanwhile, a histogram could show whether a few expensive purchases account for much of the revenue.

Descriptive Statistics vs. Inferential Statistics

Descriptive analysis reports what the observed data show. In contrast, inferential analysis uses a sample to draw conclusions about a wider population while accounting for uncertainty.

Suppose you survey 80 customers and find that 52 prefer online checkout. The observed proportion equals 65%. Reporting that result describes the respondents. However, claiming that 65% of all customers share the preference requires additional reasoning about sampling and uncertainty.

Therefore, separate the observed result from the broader claim. A larger dataset does not automatically eliminate selection bias, measurement errors, or missing responses.

Understand Your Data Before Choosing a Formula

The same calculation can have different meanings in different settings. Consequently, define the population, unit of observation, variable, and measurement period before summarizing the data.

Population and Sample

A population contains every unit within the scope of your question. A sample contains a subset of those units.

For example, all orders your shop received in September form a population if your question concerns that month alone. However, those orders may serve as a sample when you study the shop’s longer-term purchasing process.

This distinction matters for variance and standard deviation. Later sections show both the population formula and the usual sample estimator.

Categorical and Numerical Variables

Categorical variables identify groups, such as payment method, product category, or country. Numerical variables express quantities, such as income, weight, or delivery time.

In addition, numerical data can be discrete or continuous. A count of purchases takes whole-number values, while a duration can take fractional values.

Consider a customer satisfaction scale from 1 to 5. The numbers indicate ordered categories, but equal numerical gaps do not necessarily represent equal psychological differences. Therefore, show response frequencies alongside any average rating.

Units, Missing Values, and Duplicates

Always record the measurement unit. For instance, 20 dollars, 20 kilograms, and 20 minutes describe different quantities.

Next, distinguish a missing value from a genuine zero. A customer with no recorded income does not necessarily earn nothing. Similarly, duplicate records can inflate counts and alter averages.

For a small data audit, check these items:

  1. Confirm that each row represents the intended unit.
  2. Verify dates and measurement units.
  3. Count missing entries for each variable.
  4. Inspect unexpected zeros and extreme values.
  5. Remove duplicates only when you can identify them reliably.

As a result, your summary will describe a clearer and more defensible dataset.

Frequency Distributions: Organize the Observations

A frequency distribution counts how often values or categories occur. Instead of scanning a long list, you can inspect a compact table.

Absolute, Relative, and Cumulative Frequency

Absolute frequency gives the count in a value or class. Relative frequency divides that count by the total number of observations. Meanwhile, cumulative frequency adds the counts through an ordered value or class.

For class j, write:

Relative frequencyj=fjn

Here, fj represents the class count and n represents the total count. Multiply the result by 100 to express it as a percentage.

Consider these ten quiz scores:

6, 7, 6, 8, 5, 7, 6, 9, 10, 6

Score Absolute frequency Relative frequency Percentage Cumulative count
5 1 0.10 10% 1
6 4 0.40 40% 5
7 2 0.20 20% 7
8 1 0.10 10% 8
9 1 0.10 10% 9
10 1 0.10 10% 10
Total 10 1.00 100% 10

For example, four scores equal 6, so their relative frequency equals 4/10=0.40. In addition, seven scores fall at or below 7. Therefore, 70% of these quiz results do not exceed 7.

The cumulative column applies because scores have a meaningful order. In contrast, a cumulative total across unordered categories, such as payment methods, usually adds little value.

Grouped Frequency Distributions

When numerical data contain many distinct values, group them into intervals. However, define intervals that neither overlap nor leave unintended gaps.

For example, these intervals cover consecutive values:

  • 10 to less than 15.
  • 15 to less than 20.
  • 20 to less than 25.

A value of 15 belongs in the second interval. Consequently, each observation enters exactly one class.

You can summarize a class with its midpoint:

mj=lower endpoint+upper endpoint2

The midpoint of 10 to less than 15 equals 12.5. Nevertheless, the actual observations within that class need not equal 12.5. This difference explains why many calculations from grouped intervals produce estimates.

How Many Classes Should You Use?

Start with a reasonable number of intervals, then inspect the result. Too few classes can hide clusters, while too many can make random variation look meaningful.

For a small classroom dataset, a simple frequency table may work better than extensive grouping. Conversely, hundreds of transaction values usually benefit from intervals.

Rather than treating one class count as mandatory, compare a few choices. If your conclusion changes sharply, report that sensitivity.

Choose the Right Chart for Descriptive Statistics

Charts make some patterns easier to notice than a table does. However, match the visual to the variable and question.

Histogram

A histogram divides numerical values into intervals and displays the observations in each interval. It can reveal asymmetry, several clusters, and unusually distant values. See the NIST explanation of histograms.

For equal-width classes, bar heights can show counts or proportions. With unequal widths, use frequency density so bar areas represent counts or proportions correctly. Otherwise, a wider interval may appear more prominent simply because it covers more possible values.

For example, compare delivery times using equal five-minute intervals. Then inspect whether most deliveries cluster near the target or extend into a long right tail.

Bar Chart

A bar chart compares categories, such as payment methods or departments. The categories remain distinct, so gaps between bars usually make sense.

For example, compare the number of customers who paid with cash, a card, or a mobile wallet. Do not interpret their arrangement as a continuous numerical distribution. For more guidance, see OpenStax on bar graphs and other displays.

Frequency Polygon and Ogive

A frequency polygon connects class midpoints to their frequencies. It offers a convenient way to compare distributions when the groups use compatible intervals.

An ogive displays cumulative frequency against class boundaries. Consequently, it helps answer questions about how many observations fall below a threshold.

For example, a cumulative delivery-time display can show the share of orders that arrive within 25, 30, or 40 minutes. Meanwhile, the ordinary histogram shows how deliveries distribute within the intervals. See OpenStax on histograms and frequency polygons.

Box Plot and Time-Series Graph

A box plot summarizes the median and quartiles. Many versions also flag values beyond specified fences, although whisker conventions vary. Therefore, explain the convention when the distinction matters. See the OpenStax guide to box plots.

By comparison, a time-series graph preserves chronological order. Use it for monthly sales, weekly visits, or annual unemployment. A histogram can summarize the same measurements, but it cannot show when their changes occurred.

Measures of Central Tendency: Mean, Median, and Mode

Central tendency describes a distribution’s center. The mean, median, and mode answer different questions, so choose the measure that fits your purpose.

The NIST guide to measures of location explains these alternatives and their sensitivity to extreme observations.

Arithmetic Mean

Calculate the arithmetic mean by adding the observations and dividing by their count.

For a population:

μ=∑i=1NxiN

For a sample:

x¯=∑i=1nxin

The symbols μ and x¯ distinguish population and sample means. However, both calculations follow the same arithmetic procedure.

For the ten quiz scores, the sum equals 70. Therefore:

x¯=7010=7

The average score equals 7 points. In addition, the mean connects directly to the total: ten observations with a mean of 7 have a total of 70.

Median

To find the median, sort the observations. For an odd count, select the middle value. For an even count, average the two central values.

The sorted quiz scores are:

5, 6, 6, 6, 6, 7, 7, 8, 9, 10

The fifth value equals 6, while the sixth equals 7. Consequently:

Median=6+72=6.5

Notice that 6.5 does not occur in the original dataset. Nevertheless, it correctly identifies the midpoint between the two central observations.

Mode

The mode identifies the most frequent value. Here, 6 appears four times, so the mode equals 6.

A dataset can have several tied modes. Moreover, when every observed value appears once, analysts commonly report no useful unique mode.

The mode also applies to categories. For instance, a retailer can identify its most common shoe size or its most frequently selected shipping option.

For additional explanations, see OpenStax on measures of center.

When the Mean and Median Tell Different Stories

Consider these hypothetical annual incomes:

$30,000, $35,000, $40,000, $45,000, $250,000

Their total equals $400,000. Therefore, the mean equals $80,000, while the median equals $40,000.

The mean describes the amount per person if you divide total income equally. In contrast, the median locates the middle person. Neither number is inherently incorrect, but each serves a different purpose.

For a report about a typical household, the median may offer the clearer headline. However, an analysis of aggregate purchasing power should also consider totals and means.

Weighted Mean

A weighted mean gives observations different levels of importance:

x¯w=∑iwixi∑iwi

Suppose a course assigns 20% to homework, 30% to a midterm, and 50% to a final exam. A student scores 80, 70, and 90, respectively.

Then:

x¯w=0.20(80)+0.30(70)+0.50(90)=82

By comparison, the unweighted mean equals 80. Therefore, ignoring the assessment weights would produce the wrong course average.

The same principle helps combine subgroup means. If one store serves 20 customers and another serves 200, their average transaction values should not automatically receive equal weight.

Geometric and Harmonic Means

The geometric mean suits multiplicative changes. For positive values:

G=(∏i=1nxi)1/n

For example, an index rises 10% and then falls 10%. Its growth factors equal 1.10 and 0.90, so their product equals 0.99. Consequently, the index ends 1% below its starting level.

The equivalent average growth rate over the two periods equals:

0.99−1≈−0.5013%

The harmonic mean suits certain rate problems with equal quantities in the numerator:

H=n∑i1/xi

Suppose a driver travels 60 miles at 30 mph and another 60 miles at 60 mph. The trip takes three hours over 120 miles. Therefore, the average speed equals 40 mph, which matches the harmonic mean of the two speeds.

In contrast, equal time spent at each speed would call for an arithmetic mean. Always identify what remains equal before choosing a rate average.

Measures of Dispersion: Describe the Spread

Two datasets can share a mean while displaying different variation. Consequently, descriptive statistics should usually pair a center measure with a spread measure.

Range

The range equals the maximum minus the minimum:

R=xmax−xmin

For the quiz scores:

R=10−5=5

This result captures the distance between the extremes. However, it says little about the arrangement of the remaining scores.

Quartiles and Interquartile Range

Quartiles describe positions within sorted data. The interquartile range, or IQR, equals:

IQR=Q3−Q1

Different software packages use different quantile conventions. Therefore, specify your method when small samples produce different answers. The OpenStax explanation of percentiles and quartiles provides a helpful foundation.

For the quiz scores, use the median-of-halves method. The lower half is 5, 6, 6, 6, 6, while the upper half is 7, 7, 8, 9, 10.

Thus, Q1=6 and Q3=8. As a result:

IQR=8−6=2

Unlike the range, this measure concentrates on the central portion of the observations.

Mean Absolute Deviation

Mean absolute deviation measures the average absolute distance from the mean:

AAD=∑i|xi−x¯|n

The quiz mean equals 7. Therefore, the absolute deviations are:

1, 0, 1, 1, 2, 0, 1, 2, 3, 1

Their sum equals 12, so:

AAD=1210=1.2

On average, a quiz score lies 1.2 points from the mean in absolute distance.

Use the full measure name when possible. The abbreviation MAD can also refer to median absolute deviation, which uses a different calculation. The NIST guide to measures of scale distinguishes these measures.

Population Variance

Variance uses squared deviations from the mean. When the observations form the full population of interest:

σ2=∑i(xi−μ)2N

For the quiz scores, the squared deviations total 22. Consequently:

σ2=2210=2.2

The unit is points squared. This squared unit makes variance less intuitive for direct reporting, although it remains useful in statistical calculations.

Sample Variance and the n − 1 Denominator

The usual sample variance estimator is:

s2=∑i(xi−x¯)2n−1

Under independent, identically distributed sampling with finite variance, this denominator makes the estimator unbiased for population variance. Estimating the mean from the same observations removes one degree of freedom.

For the ten quiz scores:

s2=229≈2.4444

However, dividing by n can still describe the empirical distribution of the observed sample. The important step is to distinguish that descriptive calculation from the usual unbiased population-variance estimator.

See OpenStax on measures of spread for further discussion.

Standard Deviation

Standard deviation takes the square root of variance:

σ=σ2,s=s2

Therefore, the population standard deviation of the quiz results equals:

σ=2.2≈1.4832

The sample standard deviation equals:

s=22/9≈1.5635

Both use points rather than points squared. Consequently, they communicate spread more naturally than variance does.

Do not interpret standard deviation as the maximum possible deviation. In addition, do not assume a fixed percentage of observations falls within one standard deviation unless the distribution supports that claim.

Coefficient of Variation

For positive ratio-scale measurements with a meaningful zero, the coefficient of variation expresses spread relative to the mean:

CV=sx¯×100%

Suppose process A fills packages with mean 200 grams and standard deviation 4 grams. Its CV equals 2%. Meanwhile, process B has mean 500 grams and standard deviation 5 grams, giving a CV of 1%.

Although process B has a larger absolute standard deviation, it has less variation relative to its mean. Therefore, the CV provides a useful comparison in this setting.

Avoid it when the mean approaches zero or the scale lacks a meaningful zero. For example, Celsius temperature ratios do not support this interpretation. Negative means also complicate its meaning.

A Complete Worked Example of Descriptive Statistics

The quiz dataset connects the main calculations in one place. Treat the ten quizzes as the complete population for this example.

Score Frequency Deviation from 7 Absolute deviation Squared deviation Frequency × squared deviation
5 1 −2 2 4 4
6 4 −1 1 1 4
7 2 0 0 0 0
8 1 1 1 1 1
9 1 2 2 4 4
10 1 3 3 9 9
Total 10 22

The weighted absolute deviations total 2+4+0+1+2+3=12. Meanwhile, the weighted signed deviations total zero.

Measure Result
Number of observations 10
Mean 7
Median 6.5
Mode 6
Minimum 5
Maximum 10
Range 5
IQR, median-of-halves method 2
Mean absolute deviation 1.2
Population variance 2.2
Population standard deviation 1.4832

These results suggest a center around 6 to 7 points. In addition, the largest scores pull the mean above the median. However, ten observations provide only a small descriptive picture, so avoid drawing broad conclusions about future performance from this table alone.

Calculate Descriptive Statistics for Grouped Data

Grouped intervals preserve counts but lose the exact position of each observation. Consequently, midpoint calculations approximate the original data.

Estimate the Mean With Class Midpoints

Consider this hypothetical distribution of purchase amounts:

Purchase amount Midpoint Frequency Frequency × midpoint
$10 to less than $20 15 2 30
$20 to less than $30 25 5 125
$30 to less than $40 35 3 105
Total 10 260

The grouped mean estimate equals:

x¯grouped=∑jfjmjn=26010=26

Thus, the midpoint-based estimated mean equals $26. Nevertheless, customers within the first interval may have spent $11 or $19 rather than exactly $15.

A frequency table of exact values differs from interval grouping. If a row represents exactly $15, using its frequency produces an exact weighted calculation rather than a midpoint approximation.

Estimate the Grouped Median

A common interpolation formula is:

Mediangrouped≈L+(n/2−Fbeforefm)h

Here, L is the lower boundary of the median class, Fbefore is the cumulative count before it, fm is its frequency, and h is its width.

In the purchase example, the cumulative counts equal 2, 7, and 10. Therefore, the $20 to less than $30 interval contains the middle of the distribution.

Using L=20, Fbefore=2, fm=5, and h=10:

Mediangrouped≈20+(5−25)10=26

This interpolation assumes a uniform distribution within the median interval. However, the frequency table alone cannot verify that assumption or identify the exact sample median.

Estimate the Grouped Mode

For equal-width classes, a common modal interpolation is:

Modegrouped≈L+d1d1+d2h

The differences d1 and d2 compare the modal frequency with the preceding and following class frequencies.

Here, the modal class has frequency 5. Therefore, d1=5−2=3 and d2=5−3=2.

Consequently:

Modegrouped≈20+33+2(10)=26

The three estimates happen to match in this example. Nevertheless, grouped means, medians, and modes generally differ.

Estimate Grouped Variance

Use midpoints in place of the missing exact observations:

sgrouped2≈∑jfj(mj−x¯grouped)2n−1

For the purchase table:

∑jfj(mj−26)2=2(15−26)2+5(25−26)2+3(35−26)2=490

Thus:

sgrouped2≈4909=54.4444

The corresponding estimated sample standard deviation equals about $7.38. However, the underlying exact values could produce a different result.

Class Boundaries vs. Printed Class Limits

Suppose weights round to the nearest 0.1 ounce. A printed class of 19.2–19.4 usually corresponds to continuous boundaries of 19.15 and 19.45.

Therefore, use the actual class boundary in interpolation formulas. Do not silently combine a printed lower limit with a continuous class width. Otherwise, your grouped median or mode may shift by half the measurement increment.

Distribution Shape: Skewness and Kurtosis

Center and spread do not fully describe a distribution. Shape provides another layer of information.

Skewness: Identify the Direction of Asymmetry

Positive skewness indicates a longer right tail, while negative skewness indicates a longer left tail. For example, a few unusually expensive purchases can extend a spending distribution to the right.

However, familiar orderings such as mean greater than median greater than mode are useful patterns, not universal laws. Zero skewness also does not prove that a distribution is symmetric.

Pearson’s Second Skewness Coefficient

One simple coefficient compares the mean and median:

SkP=3(x¯−Median)s

Using the population standard deviation for the complete quiz population:

SkP=3(7−6.5)1.4832≈1.01

This coefficient is positive. Nevertheless, it differs from the moment-based measure that many software packages report.

Moment-Based Skewness

Define empirical central moments as:

mk=1n∑i(xi−x¯)k

Then an unadjusted moment coefficient is:

g1=m3m23/2

For the quiz scores, m2=2.2 and m3=2.4. Therefore, g1≈0.7355.

Some programs apply a sample-size correction. Consequently, always identify the definition before comparing reported values. See NIST on skewness and kurtosis.

Kurtosis: Look Beyond Peak Height

Kurtosis uses the standardized fourth central moment:

β2=m4m22

Ordinary kurtosis equals 3 for a normal distribution. Excess kurtosis subtracts 3, so its normal benchmark equals zero.

Although some introductions emphasize peakedness, tail behavior and extreme observations offer a more useful interpretation. Furthermore, sample corrections and reporting conventions vary across software.

For the quiz results, m4=12.2. Thus:

β2=12.22.22≈2.5207

The unadjusted excess kurtosis equals approximately −0.4793. However, this small dataset does not establish the tail behavior of a wider population.

In addition, use normalized frequencies when computing moments from a frequency table. Raw counts require division by the total count. Without that normalization, the formula changes with dataset size rather than describing the intended shape.

Covariance and Correlation: Describe Joint Movement

So far, the examples describe one variable at a time. Covariance and correlation extend the analysis to paired numerical observations.

Population and Sample Covariance

Population covariance is:

Cov(X,Y)=∑i(xi−μX)(yi−μY)N

The usual sample covariance estimator is:

sXY=∑i(xi−x¯)(yi−y¯)n−1

Positive covariance indicates that above-mean values tend to occur together. In contrast, negative covariance indicates that an above-mean value for one variable tends to accompany a below-mean value for the other.

Because covariance depends on measurement units, its magnitude lacks a universal comparison scale.

Worked Covariance Example

Consider four paired observations:

Observation X Y X − 2.5 Y − 5 Product of deviations
1 1 2 −1.5 −3 4.5
2 2 4 −0.5 −1 0.5
3 3 6 0.5 1 0.5
4 4 8 1.5 3 4.5
Total 10

The means equal 2.5 and 5. Therefore, population covariance equals 10/4=2.5, while sample covariance equals 10/3≈3.3333.

Both calculations indicate positive joint movement. Nevertheless, state which convention you use.

Pearson Correlation

Pearson correlation standardizes covariance:

r=sXYsXsY

For the paired example, Y=2X exactly. Consequently, the correlation equals 1.

A correlation near zero indicates little linear association, but a nonlinear relationship may remain. For instance, consider X=−2,−1,0,1,2 and Y=X2. Their covariance equals zero even though Y follows a precise rule based on X.

Likewise, correlation does not identify cause. A business may observe higher advertising spending during periods of strong sales, but both variables could reflect a seasonal shopping cycle.

Apply Descriptive Statistics to Economics and Business

The following hypothetical applications show how to translate calculations into useful questions. They illustrate analytical choices rather than claims about actual businesses or economies.

Income and Wage Comparisons

Suppose region A reports mean income of $70,000 and median income of $45,000. Region B reports mean income of $60,000 and median income of $50,000.

The higher mean in region A does not imply that its middle household earns more. Therefore, compare both measures and inspect the distribution.

In addition, define whether the records concern people, households, jobs, or tax returns. Those units produce different interpretations even when the currency matches.

Website Performance

Imagine that a website publishes 50 articles. Forty-nine attract 100 monthly page views each, while one attracts 10,000.

Total page views equal 14,900. Consequently, the mean equals 298 views per article, while the median equals 100.

The headline average hides the concentration in one article. Therefore, report total traffic, median article traffic, and the leading article’s share, which equals approximately 67.1%.

Next, compare topics or publication cohorts. However, give newer articles enough time before comparing their performance with older pages.

Retail and Customer Spending

Suppose five order values equal $20, $25, $30, $35, and $190. Their mean equals $60, while their median equals $30.

The mean helps connect order count to revenue. Meanwhile, the median describes the middle transaction more directly.

For inventory or promotion decisions, add a frequency table by spending band. As a result, managers can distinguish frequent small purchases from occasional large baskets.

Manufacturing and Service Reliability

Consider two processes with measurements 9, 10, 11 and 5, 10, 15. Both have mean 10.

However, their population variances equal 2/3 and 50/3, respectively. The second process varies much more widely.

Therefore, a manager should compare spread and specification limits alongside the average. A process that meets its target on average can still produce many unacceptable outcomes.

Comparing Performance Over Time

A monthly sales histogram summarizes the distribution of sales values. In contrast, a time-series graph shows the sequence of those values.

For example, the series 10, 20, 30, 40 and 40, 30, 20, 10 share the same descriptive distribution. Nevertheless, they show opposite trends.

Consequently, keep chronological order whenever timing affects the decision. Distribution summaries and time-series displays answer complementary questions.

A Practical Workflow for Descriptive Analysis

Follow a consistent process so your report moves from a clear question to an interpretable result.

Step 1: Define the Question

Write the decision or comparison in one sentence. For example: “How consistent were delivery times for completed orders last month?”

Then specify the population, period, and observation unit. This prevents an accidental shift from describing last month’s orders to predicting all future deliveries.

Step 2: Inspect Data Quality

Check missing values, duplicate records, measurement units, and impossible entries. In addition, document any exclusion rules.

For delivery data, a negative duration usually deserves investigation. However, an unusually long delivery may reflect a real service problem rather than an error.

Step 3: Show the Distribution

Create a frequency table and an appropriate chart. Next, inspect clusters, asymmetry, and extreme values.

For numerical data, try a histogram or box plot. For categories, compare counts and percentages.

Step 4: Pair Center With Spread

Report a suitable center measure together with a spread measure. For example, use median and IQR to describe a highly asymmetric duration distribution.

Meanwhile, mean and standard deviation may offer useful additional information. The choice should follow the question rather than a fixed rule that applies to every dataset.

Step 5: Compare Relevant Groups

Break the results into meaningful categories when the question requires it. However, retain group counts so readers can assess how much information supports each comparison.

Suppose one driver completed 8 orders and another completed 150. Their median times may be comparable descriptively, but they rest on very different amounts of evidence.

Step 6: Write the Interpretation

Translate each result into a sentence with units and context. Instead of writing only “SD = 6,” explain that completed delivery times had a sample standard deviation of 6 minutes.

Finally, distinguish observations from explanations. The data may show longer delays on weekends, but the summary alone cannot establish their cause.

Common Mistakes That Distort the Results

Reporting Only an Average

An average can hide a long tail or two separate clusters. Therefore, include a spread measure and inspect a chart.

For example, a workforce containing many junior employees and a few executives may have a mean salary that represents neither group well.

Averaging Percentages Without Their Denominators

Suppose campaign A records 10 purchases from 100 visits, while campaign B records 1 purchase from 10 visits. Both conversion rates equal 10%, so combining them is straightforward.

However, unequal conversion rates require careful weighting. If A records 20 purchases from 100 visits and B records 1 from 10, the overall conversion rate equals 21/110≈19.09%, rather than the unweighted average of 20% and 10%.

Confusing Standard Deviation With Standard Error

Standard deviation describes observation-level spread. By contrast, standard error describes the sampling variability of an estimator.

For independent observations under the usual assumptions, the estimated standard error of the mean equals s/n. Therefore, increasing the sample size can reduce this uncertainty without reducing the underlying spread of individual outcomes.

Removing Outliers Automatically

A distant value may represent an error or a legitimate rare event. Investigate it before deleting it. The NIST discussion of outliers emphasizes the distinction between identifying unusual observations and deciding how to handle them.

For a business example, a very large sale may matter precisely because it is unusual. Therefore, explain its influence and consider reporting results with and without it when that comparison serves the question.

Assuming Symmetry Means Normality

Consider values −10, −10, 10, 10. Their distribution is symmetric around zero, yet the observations form two separated clusters.

Consequently, symmetry alone cannot justify a normal model. Similarly, matching a mean and median does not establish the rest of a distribution’s shape.

Treating Grouped Estimates as Exact Results

Midpoint calculations replace unknown values with representative class values. Therefore, label their output as an estimate.

If you still have the raw observations, calculate the exact summaries directly. Grouping remains useful for communication, but it need not replace available information.

Rounding Too Early

Keep enough precision during intermediate calculations. Then round the final output to a level that matches the measurement quality and reporting purpose.

For example, using 1.48 instead of 1.4832 in a follow-up formula changes the result slightly. Repeated early rounding can create avoidable discrepancies.

Descriptive Statistics Formula Reference

This table collects the main formulas. It assumes equally weighted raw observations unless the row states otherwise.

Measure Formula Interpretation or qualification
Mean x¯=∑xi/n Arithmetic average
Weighted mean ∑wixi/∑wi Average with specified weights
Relative frequency fj/n Proportion in a category or class
Range xmax−xmin Distance between extremes
IQR Q3−Q1 Spread of the central portion
Mean absolute deviation ∑|xi−x¯|/n Average absolute distance from the mean
Population variance ∑(xi−μ)2/N Population spread in squared units
Sample variance ∑(xi−x¯)2/(n−1) Usual unbiased variance estimator
Standard deviation s=s2 Spread in original units
CV 100s/x¯ Relative spread for suitable positive ratio-scale data
Grouped mean estimate ∑fjmj/n Midpoint approximation for intervals
Empirical central moment mk=∑(xi−x¯)k/n Normalized moment about the mean
Moment skewness m3/m23/2 Unadjusted asymmetry coefficient
Ordinary kurtosis m4/m22 Unadjusted standardized fourth moment
Excess kurtosis m4/m22−3 Kurtosis relative to normal benchmark
Sample covariance ∑(xi−x¯)(yi−y¯)/(n−1) Joint linear movement in product units
Pearson correlation sXY/(sXsY) Standardized linear association

If all observations equal the same value, variance equals zero. Consequently, skewness, kurtosis, and correlation formulas that divide by a zero spread do not yield a defined result.

Practice Exercises With Worked Answers

Try each calculation before reading its answer. In addition, explain what the result means rather than stopping at the number.

Exercise 1: Mean, Median, and Mode

Use the dataset 2, 4, 4, 6, 9.

The sum equals 25, so the mean equals 25/5=5. Meanwhile, the middle observation equals 4 and the most frequent value also equals 4.

Therefore, mean = 5, median = 4, and mode = 4. The largest observation raises the mean above the median.

Exercise 2: Population Variance and Standard Deviation

Use the dataset 3, 5, 7, treating these values as the complete population.

The mean equals 5. Next, the squared deviations equal 4, 0, and 4, giving a sum of 8.

Thus, population variance equals 8/3≈2.6667. The population standard deviation equals 8/3≈1.6330.

If you instead use the usual sample estimator, variance equals 8/2=4. Consequently, sample standard deviation equals 2.

Exercise 3: Relative Frequency

In 40 orders, 18 customers choose standard shipping, 14 choose express shipping, and 8 choose pickup.

Divide each count by 40. Therefore, the relative frequencies equal 45%, 35%, and 20%, respectively.

Their sum equals 100%. However, a cumulative percentage would depend on an arbitrary ordering of these categories and would not offer a natural numerical threshold interpretation.

Exercise 4: Grouped Mean

Suppose five observations fall between 0 and less than 10, three fall between 10 and less than 20, and two fall between 20 and less than 30.

The midpoints equal 5, 15, and 25. Therefore:

x¯grouped≈5(5)+3(15)+2(25)10=12

The estimated mean equals 12. Nevertheless, the exact mean remains unknown without the original observations.

Exercise 5: Why an Average Can Mislead

Compare 10, 10, 10, 10 with 1, 1, 1, 37.

Both totals equal 40, so both means equal 10. However, their medians equal 10 and 1, respectively.

The first dataset has population variance zero. In contrast, the second has squared deviations totaling 81+81+81+729=972, giving population variance 972/4=243.

Therefore, equal averages can conceal very different centers and levels of variation.

Frequently Asked Questions About Descriptive Statistics

What Are the Main Types of Descriptive Statistics?

Common categories include frequencies, measures of center, measures of spread, and measures of shape. In addition, covariance and correlation describe relationships between paired variables.

Choose the categories that answer your question. A categorical survey may need frequencies and a mode, while a numerical process often needs center, spread, and a chart.

Which Is Better: Mean or Median?

Neither measure is always better. The mean connects totals to counts, while the median identifies the middle ordered observation.

For example, median order value can describe the middle purchase. Meanwhile, mean order value helps explain total revenue per order. Reporting both can clarify an asymmetric distribution.

Can Descriptive Statistics Prove a Hypothesis?

A descriptive summary alone does not establish a population-wide claim or causal relationship. However, it can reveal patterns that motivate further analysis.

For example, two groups may show different average outcomes. Before explaining that difference, examine how you selected the groups and what other factors differ between them.

What Does a Large Standard Deviation Mean?

A large standard deviation indicates substantial spread relative to the measurement scale. However, whether it counts as large depends on the context.

For instance, a standard deviation of 5 grams may be substantial for a small tablet but minor for a heavy package. Therefore, interpret the number with the mean, units, and acceptable limits.

Why Do Software Results Sometimes Differ?

Programs can use different quantile methods, population or sample denominators, and skewness or kurtosis corrections. Moreover, missing-value handling can change the observations included in each calculation.

Consequently, compare definitions before assuming a result is wrong. Record the method when reproducibility matters.

Should Every Report Include Skewness and Kurtosis?

These measures can help characterize shape, but every report does not need them. For many practical questions, a histogram, median, mean, and spread measure provide a clearer starting point.

In addition, small samples can produce unstable shape estimates. Use them as part of the overall picture rather than as automatic labels.

Conclusion: Turn Summaries Into Better Questions

Descriptive statistics provide a structured way to understand the observations in front of you. Start by defining the dataset, then organize its frequencies and inspect its distribution. Next, pair a suitable measure of center with a measure of spread.

When the question involves two variables, examine their joint movement while keeping causal claims separate. In addition, label grouped estimates and explain any software conventions that affect the results.

A useful statistical summary does more than list formulas. It shows what the measurements describe, why the chosen measures matter, and which questions remain open. Therefore, combine accurate calculations with clear interpretation whenever you turn raw data into an article, report, or business decision.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *