Skip to content

Regex cheat sheet with real examples, and how to avoid catastrophic backtracking

By · Regex Tools · 6 min read · Updated

This is the regex reference I wish I had kept next to my editor years ago: the syntax on one screen, the patterns that come up again and again, and the one failure mode that can take down a production service. Everything uses JavaScript syntax, which matches Java, Python, Go, .NET and PCRE for everything except the rare features noted. Every example was run in Node.js; paste any of them into the Regex Tester to experiment.

Cheat sheet: syntax

SyntaxMeaningExample
.Any character except a line break (any at all with the s flag)a.c matches "abc", "a-c"
\d \w \sDigit, word character [A-Za-z0-9_], whitespace\d\d matches "42"
\D \W \SThe opposites\S+ matches a run of non-space
[abc] [a-z] [^"]One character from a set, a range, or not in a set[^,]* matches one CSV field
* + ?Zero or more, one or more, optionalcolou?r
{3} {2,5} {2,}Exactly, between, at least\d{4}
*? +? ??Lazy versions: match as little as possible<.+?>
^ $Start and end of input (of each line with m)^\d{5}$
\bWord boundary\bcat\b does not match inside "concatenate"
(…) (?:…)Capturing group, non-capturing group(?:ab)+
(?<name>…)Named group(?<year>\d{4})
a|bAlternation^(GET|POST)\s
(?=…) (?!…)Followed by, not followed by\d+(?=px)
(?<=…) (?<!…)Preceded by, not preceded by(?<=\$)\d+
\1 \k<name>Backreference to a group(\w)\1 matches "ll" in "hello"
\p{L} \p{Lu}Unicode letter, uppercase letter (needs u)\p{Lu} matches "É"

Escape these to match them literally: . ^ $ * + ? ( ) [ ] { } | \ /. The classic slip is the dot: /example.com/ happily matches "exampleXcom". Write example\.com.

Cheat sheet: flags

FlagEffect
gAll matches, not just the first. Makes the regex object stateful (see below).
iCase-insensitive
m^ and $ match at every line
s. also matches line breaks
u / vUnicode mode: code points, \p{…}, stricter syntax. v adds set operations.
ySticky: match only at lastIndex, used by tokenisers

Real examples

Each of these was tested against the input shown. They are pragmatic, not RFC-complete. That is usually what you want.

// Leading and trailing whitespace
'  hi  '.replace(/^\s+|\s+$/g, '')                        // 'hi'

// ISO date shape, with named groups, reordered
'2026-10-09'.replace(/(?<y>\d{4})-(?<m>\d{2})-(?<d>\d{2})/, '$<d>/$<m>/$<y>')
                                                          // '09/10/2026'

// key=value pairs from a log line
[...'level=info user=42 path=/api'.matchAll(/(?<key>\w+)=(?<value>\S*)/g)]
  .map((m) => m.groups)                                    // [{key:'level',value:'info'}, …]

// Semantic version with optional pre-release
/^v?(\d+)\.(\d+)\.(\d+)(?:-([\w.]+))?$/.exec('v1.10.3-rc.1')   // 1, 10, 3, 'rc.1'

// Pixel value without the unit (lookahead)
'width: 16px'.match(/\d+(?=px)/)[0]                       // '16'

// Amount after a dollar sign (lookbehind)
'total $42.50'.match(/(?<=\$)\d+(\.\d{2})?/)[0]          // '42.50'

// Log level and message, one entry per line
/^(?<level>INFO|WARN|ERROR)\s+(?<msg>.*)$/m.exec('DEBUG x\nWARN disk at 91%').groups
                                                          // { level: 'WARN', msg: 'disk at 91%' }

// UUID anywhere in text, any case
/\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b/i

// Hex colour, 3 or 6 digits
/#[0-9a-f]{3}(?:[0-9a-f]{3})?\b/i                         // '#FFaa00'

// Password policy: 12+ chars, at least one digit and one lowercase letter
/^(?=.*\d)(?=.*[a-z]).{12,}$/

Greedy, lazy and negated classes

Against <b>bold</b>, the greedy <.+> matches the whole string, because .+ runs to the end and backs off only as far as the last >. The lazy <.+?> matches <b>. The negated class <[^>]+> also matches <b>, and is usually the better choice: it says exactly what may appear, and it cannot backtrack across a >. Remember that last property, because it is the key to the next section.

Anchors make validation work

/\d{5}/.test('abc123456xyz') is true, because it finds "12345" inside. For validation you almost always want ^…$: /^\d{5}$/ rejects it. Also, shape is not validity. ^\d{4}-\d{2}-\d{2}$ accepts 2026-13-45. Check ranges in code after the regex has extracted the parts.

Catastrophic backtracking

JavaScript, Java, Python, .NET, Ruby and PCRE use backtracking engines. When part of a pattern fails, the engine goes back and tries every other way the earlier parts could have matched. Usually there are only a few. But if a pattern can match the same text in many different ways, a failing input makes the engine try all of them, and the number grows exponentially with input length.

The textbook case is a nested quantifier. Run /^(a+)+$/ against a string of as followed by !. Each a can belong to the inner or outer repetition, so n characters can be split in 2n−1 ways, and the engine has to try them all before it can report "no match". Measured in Node.js:

Input/^(a+)+$//^a+$/
24 × "a" + "!"89 ms0 ms
26 × "a" + "!"352 ms0 ms
28 × "a" + "!"1,448 ms0 ms

Every two extra characters quadruple the time. Extrapolating, 40 characters would take well over an hour. Node.js runs regexes on the main thread, so while that one request is backtracking, the process serves no other requests. That is a ReDoS (regular expression denial of service), and it has taken down large production services.

It hides in innocent-looking patterns

Nobody writes (a+)+ on purpose. They write this, to match "words separated by spaces":

/^(\w+\s?)*$/

Because the \s? is optional, \w+\s? can match "word" as one piece, as "wo" + "rd", as "w" + "ord", and so on: the same ambiguity as (a+)+. On 'word word word word word word word word word!', 45 characters ending in a character the pattern does not allow, this took 2.1 seconds. With twelve words, 60 characters, it took 204 seconds. Matching input is fast; only the failure explodes, which is why these patterns pass every test with valid data.

Another common one is /\s+$/ to find trailing whitespace. It is not exponential, but on a string with a long run of spaces in the middle it is quadratic: 20,000 spaces followed by a letter took 194 ms. That is enough to matter when an attacker can send a megabyte.

How to fix it

  1. Make each character matchable in only one way. Make the separator mandatory: /^\w+(?:\s\w+)*$/ runs in 0 ms on the same input, even with 1,000 words. For trailing whitespace, trim with String.prototype.trimEnd() instead.
  2. Never nest quantifiers over overlapping sets, as in (x+)+, (x*)*, (\w|\d)+ or (.*,)*.
  3. Prefer negated classes to .*: "[^"]*" instead of ".*?", and [^,]* for a CSV field.
  4. Limit input length before matching. A 254-character cap on an email field removes most of the risk.
  5. Use a linear-time engine for untrusted patterns or input. Go's regexp and Rust's regex crate guarantee linear time by design (no backreferences or lookarounds). In Node.js the re2 package wraps Google's RE2; .NET lets you pass a match timeout.
  6. Lint for it. ESLint plugins such as eslint-plugin-regexp flag many super-linear patterns automatically.

Gotchas worth remembering

  • Global regexes are stateful. With g or y, test() and exec() continue from lastIndex. A shared const re = /a/g returns true, then false, for re.test('a') called twice. Do not reuse global regex objects across calls; use matchAll or a fresh literal.
  • $ in replacement strings is special. $& inserts the match and $1 a group, so replacing with user-supplied text needs a function: s.replace(re, () => userText).
  • \w is ASCII-only, even with u. Use [\p{L}\p{N}_] for international names.
  • Regex is not a parser. HTML, JSON, URLs and full email grammar need real parsers, such as new URL() (or the URL Parser), JSON.parse and DOMParser. Use regex on the values they give you.

A workflow that works

Collect real samples, including lines that must not match and at least one long malformed input. Build the pattern a piece at a time in the Regex Tester. Anchor it. Then commit it with a unit test containing those samples, and a comment stating the intent in plain words. The next person to edit it, often you, will need both.