Menu

Text Tools

Hash Generator (SHA-1 / SHA-256 / SHA-512)

Text Tools: How It Works

This page collects the small text transformations that come up constantly in development and content work — the ones that are trivial individually but tedious to do by hand, and risky to paste into a random online tool when the text is confidential.

What is included

OperationTypical use
Case conversionMatching a code convention or style guide
Trim and normalise whitespaceCleaning data pasted from spreadsheets or PDFs
Remove line breaksTurning wrapped text back into flowing paragraphs
Add or strip line numbersPreparing text for reference or removing them after copying
Reverse text or linesDebugging, puzzles, checking order
Find and replaceBulk edits across a block
Strip HTML tagsExtracting readable text from markup
Count words and charactersChecking against a limit

The whitespace problem

Text copied from PDFs, spreadsheets and websites arrives carrying invisible baggage: trailing spaces, non-breaking spaces, tabs where spaces are expected, and multiple consecutive spaces. None of it is visible, and all of it breaks comparisons, deduplication and sorting.

The non-breaking space (U+00A0) is the worst of them. It looks exactly like a space, and it is a different character — so a search for a normal space will not find it, a split on whitespace may not split on it, and two strings that look identical compare as different. Normalising whitespace before any other processing prevents a whole category of unexplained failures.

Smart quotes and dashes

Word processors silently replace straight quotes with typographic ones and hyphens with en and em dashes. This is desirable in prose and destructive in code and data:

If you have ever pasted code from a document and had it fail inexplicably, this is almost certainly why.

Stripping HTML

Removing tags to extract readable text is a common need, and worth two cautions. Block elements should become line breaks — otherwise paragraphs run together into a single wall. And HTML entities need decoding: & should become an ampersand, not remain as literal text. Stripping tags alone leaves a mess.

Order of operations

When cleaning data, sequence matters:

  1. Normalise line endings (Windows to Unix).
  2. Normalise whitespace, including non-breaking spaces.
  3. Trim each line.
  4. Remove empty lines if appropriate.
  5. Convert case if needed.
  6. Deduplicate.
  7. Sort.

Deduplicating before trimming leaves phantom duplicates; sorting before normalising strands whitespace-prefixed lines at the top.

Everything here runs in your browser. Nothing you paste is transmitted — which is the relevant property when the text is a customer list, an API response or an internal document.

Frequently Asked Questions

Why do two identical-looking lines not match?
Usually a trailing space or a non-breaking space, which looks exactly like a normal space but is a different character. Normalise whitespace before comparing, deduplicating or sorting.
Why does code copied from a document fail?
Word processors replace straight quotes with curly typographic ones and hyphens with dashes. A curly apostrophe is a syntax error in every programming language. Convert them back to straight characters before running the code.
What is a non-breaking space?
Character U+00A0, used to prevent a line break between two words. It renders identically to a normal space but is a distinct character, so searches, splits and comparisons treat it differently.
In what order should I clean text?
Normalise line endings, then whitespace, then trim, then remove blank lines, then convert case, then deduplicate, then sort. Deduplicating before trimming leaves phantom duplicates behind.
Does stripping HTML tags give clean text?
Not on its own. Block elements should be replaced with line breaks or paragraphs run together, and HTML entities need decoding so & becomes an ampersand rather than remaining as literal text.
Is my text uploaded?
No. Every operation runs in your browser and nothing is transmitted or stored, which matters when the text is customer data, an API response or an internal document.

Related Developer Tools

Browse all Developer tools →