Compress PDF: How It Works
PDF file size is almost always dominated by embedded images. Compressing one means deciding how much image quality to trade for size — and knowing which of the four size drivers actually applies to your document.
What makes a PDF large
| Cause | Typical share | Fix |
|---|---|---|
| High-resolution images | Usually most of the file | Downsample to the DPI actually needed |
| Uncompressed or lossless images | Large | Re-encode as JPEG |
| Fully embedded fonts | Moderate | Subset to characters used |
| Retained edit history and metadata | Small but real | Save a clean copy rather than incremental saves |
A scanned document is essentially a stack of photographs, which is why scans dominate the large-PDF problem. A text-only PDF exported from a word processor is usually already small, and compressing it achieves very little.
Choosing a target resolution
| Purpose | DPI | Effect on a 20 MB scan |
|---|---|---|
| Screen viewing and email | 150 | ≈ 2–4 MB |
| Office printing | 200–300 | ≈ 5–8 MB |
| Professional print | 300 | ≈ 8–12 MB |
| Archival | Do not downsample | Unchanged |
Below 150 DPI, small print becomes difficult to read and OCR accuracy falls sharply. If the document may need to be searched or re-read later, 200 DPI is a safer floor than 150.
Greyscale and black-and-white
A scanned text document in full colour carries three channels of data to represent what is essentially black marks on white paper. Converting to greyscale typically halves the size with no loss of legibility, and converting to true black-and-white can reduce it by 90% — at the cost of losing any coloured stamps, signatures or highlighting, which sometimes matter legally.
Font subsetting
Embedding a complete font can add several hundred kilobytes per typeface. Subsetting embeds only the characters actually used, usually a fraction of the size, and the document still renders identically. This is standard in most export pipelines; a document that has not been subsetted is usually one that has been through several editing tools.
Compression is not reversible
Downsampling images destroys detail permanently. Compressing a file that has already been compressed compounds the loss, and repeatedly compressing a scanned document produces mushy, hard-to-read text. Keep the original and compress a copy — this is the single most useful habit with any lossy process.
When it does not help
If a PDF is large and contains almost no images, compression will barely move it. Look instead for embedded attachments, retained revision history from incremental saves, or an unsubsetted font. Saving a clean copy rather than accumulating incremental saves often reduces such a file more than any compression setting.
Compression here runs in your browser; your documents are not uploaded.