Text Tools: How It Works
This page collects the small text transformations that come up constantly in development and content work — the ones that are trivial individually but tedious to do by hand, and risky to paste into a random online tool when the text is confidential.
What is included
| Operation | Typical use |
|---|---|
| Case conversion | Matching a code convention or style guide |
| Trim and normalise whitespace | Cleaning data pasted from spreadsheets or PDFs |
| Remove line breaks | Turning wrapped text back into flowing paragraphs |
| Add or strip line numbers | Preparing text for reference or removing them after copying |
| Reverse text or lines | Debugging, puzzles, checking order |
| Find and replace | Bulk edits across a block |
| Strip HTML tags | Extracting readable text from markup |
| Count words and characters | Checking against a limit |
The whitespace problem
Text copied from PDFs, spreadsheets and websites arrives carrying invisible baggage: trailing spaces, non-breaking spaces, tabs where spaces are expected, and multiple consecutive spaces. None of it is visible, and all of it breaks comparisons, deduplication and sorting.
The non-breaking space (U+00A0) is the worst of them. It looks exactly like a space, and it is a different character — so a search for a normal space will not find it, a split on whitespace may not split on it, and two strings that look identical compare as different. Normalising whitespace before any other processing prevents a whole category of unexplained failures.
Smart quotes and dashes
Word processors silently replace straight quotes with typographic ones and hyphens with en and em dashes. This is desirable in prose and destructive in code and data:
- A curly apostrophe in a code snippet is a syntax error.
- A curly quote in a CSV breaks parsing.
- An em dash in an SMS switches the message to Unicode encoding, cutting the segment limit from 160 characters to 70.
If you have ever pasted code from a document and had it fail inexplicably, this is almost certainly why.
Stripping HTML
Removing tags to extract readable text is a common need, and worth two cautions. Block elements should become line breaks — otherwise paragraphs run together into a single wall. And HTML entities need decoding: & should become an ampersand, not remain as literal text. Stripping tags alone leaves a mess.
Order of operations
When cleaning data, sequence matters:
- Normalise line endings (Windows to Unix).
- Normalise whitespace, including non-breaking spaces.
- Trim each line.
- Remove empty lines if appropriate.
- Convert case if needed.
- Deduplicate.
- Sort.
Deduplicating before trimming leaves phantom duplicates; sorting before normalising strands whitespace-prefixed lines at the top.
Everything here runs in your browser. Nothing you paste is transmitted — which is the relevant property when the text is a customer list, an API response or an internal document.