Menu

Regex Tester

/ /

Regex Tester: How It Works

A regular expression describes a pattern rather than a literal string. A tester lets you build one against real input and watch what it matches, which is far faster than reasoning about it and far more reliable than hoping.

The building blocks

PatternMatches
.Any character except newline
\d \w \sDigit, word character, whitespace
\D \W \SThe negations of each
[abc] [^abc]Any of / none of those characters
* + ?Zero or more, one or more, zero or one
{2,5}Between two and five repetitions
^ $Start and end of string (or line, with the m flag)
\bWord boundary
(...)Capturing group
(?:...)Group without capturing
a|bEither alternative

Greedy versus lazy

Quantifiers are greedy by default — they consume as much as possible and then give back only what they must. Against <b>bold</b> and <i>italic</i>, the pattern <.+> matches the entire string, not the first tag. Adding ? makes the quantifier lazy: <.+?> matches each tag individually.

This single behaviour accounts for a large share of regexes that 'almost work'.

Flags

FlagEffect
gFind all matches, not just the first
iCase-insensitive
m^ and $ match at each line
s. also matches newlines
uFull Unicode handling

Catastrophic backtracking

Nested quantifiers over overlapping patterns — (a+)+b is the textbook example — can take exponential time on input that nearly matches. A 30-character string can hang a process for years. If a regex runs against user-supplied input, avoid nested quantifiers, prefer possessive or atomic constructs where your engine supports them, and impose a length limit on the input before matching.

What regex is bad at

Nested structures. HTML, XML and JSON are recursive, and regular expressions describe regular languages, which by definition cannot express arbitrary nesting. A regex can extract a value from a known, stable format; it cannot reliably parse a document. Use a parser for those — the regex approach works on your test cases and fails on real input in ways that are difficult to diagnose.

Practical habits

Frequently Asked Questions

Why does my pattern match too much?
Quantifiers are greedy by default and consume as much as they can. Add a question mark after the quantifier to make it lazy — .+? instead of .+ — so it matches the shortest possible span.
Can I parse HTML with a regex?
Not reliably. HTML is a nested structure and regular expressions cannot express arbitrary nesting. Extracting one value from a known stable format can work; parsing a document will fail on real-world input. Use a proper parser.
What is catastrophic backtracking?
Nested quantifiers over overlapping patterns can cause exponential time on input that nearly matches, hanging the process. Avoid patterns like (a+)+ on untrusted input, and limit input length before matching.
Why does my validation accept invalid input?
Most likely the pattern is not anchored. Without ^ at the start and $ at the end, the engine looks for the pattern anywhere in the string, so extra characters either side are ignored.
Are regex flavours the same across languages?
The basics are, but lookbehind, named groups, Unicode property escapes and possessive quantifiers vary. JavaScript, PCRE, Python and Go's RE2 all differ. Test in the engine you will actually deploy to.
How should I validate an email address?
Loosely. Check there is an @ with plausible content either side, then send a confirmation message. Strict pattern validation rejects valid addresses and accepts undeliverable ones — only delivery proves an address works.

Related Developer Tools

Browse all Developer tools →