Standard Deviation Calculator: How It Works
Standard deviation measures spread — how far values typically sit from their mean. Two datasets can share an identical average and be completely different, and standard deviation is the number that tells them apart.
Population or sample
This is the choice that changes the answer, and the one most people get wrong.
Population: σ = √[ Σ(x − μ)² ÷ N ]
Sample: s = √[ Σ(x − x̄)² ÷ (n − 1) ]
Use the population formula only when your data is the entire group — every student in one class, every transaction last month. Use the sample formula whenever the data is a subset used to estimate something larger, which covers almost all real analysis.
Why n − 1
A sample's own mean is closer to its own values than the true population mean would be, so squared deviations from the sample mean systematically underestimate the real spread. Dividing by n − 1 instead of n corrects that bias. The correction is called Bessel's correction, and it matters most on small samples: with n = 5 it raises the result by about 12%, and by n = 100 the difference is under 0.5%.
Worked example
Data: 4, 8, 6, 5, 3 (n = 5)
- Mean = 26 ÷ 5 = 5.2
- Deviations: −1.2, 2.8, 0.8, −0.2, −2.2
- Squared: 1.44, 7.84, 0.64, 0.04, 4.84 → sum = 14.8
- Sample variance = 14.8 ÷ 4 = 3.7
- Sample SD = √3.7 = 1.92
Using the population formula instead gives 14.8 ÷ 5 = 2.96, and an SD of 1.72 — a 12% difference from the same five numbers.
Reading the result
Standard deviation is in the same units as your data, which makes it directly interpretable: an SD of 1.92 marks in a test means typical scores sit about two marks from the average. Variance, being squared, is in squared units and is rarely interpretable on its own — it exists because it behaves well mathematically.
For roughly bell-shaped data, the empirical rule applies:
| Within | Contains about |
|---|---|
| ±1 SD | 68% of values |
| ±2 SD | 95% of values |
| ±3 SD | 99.7% of values |
This only holds for approximately normal distributions. Skewed or multi-peaked data breaks it, sometimes badly.
Comparing across different scales
An SD of 5 is large for exam marks out of 10 and negligible for annual salaries. To compare spread across different units, use the coefficient of variation: SD ÷ mean, expressed as a percentage. It is unitless, which makes it the right tool for comparing the variability of, say, rainfall and revenue.
Where it misleads
- Outliers dominate it. Squaring deviations gives extreme values disproportionate weight. One data-entry error can double an SD.
- It assumes a meaningful mean. For strongly skewed data — incomes, waiting times — the interquartile range describes spread more honestly.
- It says nothing about shape. A uniform distribution and a two-peaked one can have identical mean and SD while being entirely different.
Always look at the data as well as the summary. A histogram takes seconds and reveals things no single statistic will.