Statistics Calculator: How It Works
Descriptive statistics compress a dataset into a handful of numbers describing its centre, its spread and its shape. This calculator produces the full set at once — and this page covers how to read them together, since any one in isolation can mislead.
The measures
| Statistic | Describes | Outlier-sensitive? |
|---|---|---|
| Mean | Arithmetic centre | Highly |
| Median | Middle value | No |
| Mode | Most frequent value | No |
| Range | Max − min | Extremely |
| Variance | Average squared deviation | Highly |
| Standard deviation | Typical distance from the mean | Highly |
| Quartiles & IQR | Spread of the middle 50% | No |
The right-hand column is the practical guide. When a dataset contains extreme values, the resistant measures — median, quartiles, IQR — describe it more honestly than mean, range and standard deviation.
Reading the five-number summary
Minimum, Q1, median, Q3 and maximum together describe a distribution's shape without assuming anything about it. This is what a box plot draws.
- Median near the centre of Q1 and Q3 — roughly symmetric.
- Median closer to Q1 — right-skewed, with a long upper tail. Typical of incomes and durations.
- Median closer to Q3 — left-skewed.
- A whisker far longer than the box — outliers on that side.
Identifying outliers
The standard rule flags a value as an outlier if it falls below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR. The 1.5 multiplier is a convention, not a law — with normally distributed data it flags roughly 0.7% of values.
Finding an outlier is the start of an investigation, not a licence to delete. A value can be extreme because it was mistyped, because a sensor failed, because a unit was wrong, or because something genuinely unusual happened and is the most important thing in the dataset. Remove points only with a documented reason, and report that you did.
Skewness and kurtosis
Skewness measures asymmetry: positive means a long right tail, negative a long left tail, and near zero is symmetric. Kurtosis measures tail weight — high kurtosis means extreme values are more common than a normal distribution would predict, which matters greatly in risk analysis, where underestimating tail events is expensive.
Why summaries alone are not enough
Anscombe's quartet is four datasets with nearly identical means, variances, correlations and regression lines. Plotted, they are completely different: one linear, one curved, one linear with a single outlier, one essentially a vertical line plus one point. The Datasaurus Dozen extends the demonstration to a set that includes a dinosaur.
The lesson is practical: compute the summary, then plot the data. A histogram takes seconds and catches bimodality, truncation, gaps and impossible values that no set of statistics will show you.
Before trusting the numbers
- Check the count matches what you expected — missing rows are silent.
- Check min and max are physically possible. An age of 200 or a negative price is a data problem.
- Compare mean and median. A large gap means skew.
- Plot it.