Regex Tester: How It Works
A regular expression describes a pattern rather than a literal string. A tester lets you build one against real input and watch what it matches, which is far faster than reasoning about it and far more reliable than hoping.
The building blocks
| Pattern | Matches |
|---|---|
. | Any character except newline |
\d \w \s | Digit, word character, whitespace |
\D \W \S | The negations of each |
[abc] [^abc] | Any of / none of those characters |
* + ? | Zero or more, one or more, zero or one |
{2,5} | Between two and five repetitions |
^ $ | Start and end of string (or line, with the m flag) |
\b | Word boundary |
(...) | Capturing group |
(?:...) | Group without capturing |
a|b | Either alternative |
Greedy versus lazy
Quantifiers are greedy by default — they consume as much as possible and then give back only what they must. Against <b>bold</b> and <i>italic</i>, the pattern <.+> matches the entire string, not the first tag. Adding ? makes the quantifier lazy: <.+?> matches each tag individually.
This single behaviour accounts for a large share of regexes that 'almost work'.
Flags
| Flag | Effect |
|---|---|
g | Find all matches, not just the first |
i | Case-insensitive |
m | ^ and $ match at each line |
s | . also matches newlines |
u | Full Unicode handling |
Catastrophic backtracking
Nested quantifiers over overlapping patterns — (a+)+b is the textbook example — can take exponential time on input that nearly matches. A 30-character string can hang a process for years. If a regex runs against user-supplied input, avoid nested quantifiers, prefer possessive or atomic constructs where your engine supports them, and impose a length limit on the input before matching.
What regex is bad at
Nested structures. HTML, XML and JSON are recursive, and regular expressions describe regular languages, which by definition cannot express arbitrary nesting. A regex can extract a value from a known, stable format; it cannot reliably parse a document. Use a parser for those — the regex approach works on your test cases and fails on real input in ways that are difficult to diagnose.
Practical habits
- Name your groups —
(?<year>\d{4})— so extraction does not depend on counting parentheses. - Anchor when validating. Without
^and$, a validation pattern matches a substring, so\d{5}happily accepts 'abc12345xyz'. - Comment complex patterns using the extended flag where your language supports it. A regex you cannot read in six months is a liability.
- Do not write an email validator. The full specification is famously unmatchable in practice. Check for an
@with something either side, then send a confirmation email — which is the only real validation anyway.