Regex Tester

Test a regex and inspect every match and capture group.

Input
TEXT input
Output
Result
Options

Without surrounding slashes.

About this tool


Enter a pattern and some text, and every match is listed with its line number, character index and the contents of each capture group, numbered and named. Seeing exactly what each group captured is usually what turns a nearly-working pattern into a correct one.

The engine is your browser's own, so the results are precisely what your JavaScript will do, not an approximation from a different regex flavour. That matters, because Python, PCRE and JavaScript differ in lookbehind support, named group syntax and Unicode handling.

How to use it

  1. Paste or upload your textDrop a file onto the input pane, use the file picker, or paste the text directly.
  2. Adjust the options if neededThe defaults suit most input; open Options to change the behaviour.
  3. TestPress Test, or use Ctrl+Enter (Cmd+Enter on macOS).
  4. Copy or downloadCopy the result, or download it as a .txt file.

Worked examples


Each example below is executed against this tool by the test suite, so what you see is what the tool actually produces.

Numbered capture groups

Input

order a1 and order b2

Output

2 matches for /([a-z])(\d)/g

Match 1, line 1, index 6: "a1"
  group 1: "a"
  group 2: "1"
Match 2, line 1, index 19: "b2"
  group 1: "b"
  group 2: "2"

Pattern: ([a-z])(\d) with the global flag. Each match reports its index and both groups.

What to watch for


The details that decide whether a conversion is correct, and where information can be lost without any error being raised.

Writing the pattern
Enter the pattern without surrounding slashes, \d{3}-\d{4} rather than /\d{3}-\d{4}/. Flags are set with the checkboxes rather than written after a closing slash.
What each flag does
Global (g) finds every match instead of stopping at the first. Ignore case (i) is self-explanatory. Multiline (m) makes ^ and $ match at line boundaries rather than only at the start and end of the whole string. Dot-all (s) lets . match a newline, which it otherwise never does. Unicode (u) enables \p{...} property escapes and makes the pattern operate on code points, so an emoji counts as one character rather than two.
Capture groups, numbered and named
Each parenthesised group is reported by number in order of its opening bracket, and named groups written (?<name>...) are reported by name as well. A group that took part in no match shows explicitly as "(no match)", which is a common cause of an unexpected undefined in code, distinguishing "matched an empty string" from "did not participate" is often the bug.
Patterns that could hang are rejected
A pattern like (a+)+ against text that nearly matches can take exponential time, and JavaScript regex execution cannot be interrupted once started. It would freeze the page outright. Patterns with a quantifier applied to an already-quantified group, or a repeated alternation with overlapping branches, are therefore refused with an explanation. Nested quantifiers are almost always unintentional: (a+)+ matches exactly what a+ matches.
Common pitfalls worth knowing
Inside a character class most metacharacters lose their meaning, so [.] matches a literal dot. The dot never matches a newline unless the s flag is set. Quantifiers are greedy by default, so .* takes as much as possible, add ? to make it lazy. And \b depends on the definition of a word character, which is ASCII-only without the u flag.

Limitations


  • Uses the JavaScript engine, so results may differ from other regex flavours.
  • Patterns with nested quantifiers are rejected, because they cannot be executed safely.
  • Reporting stops at 1,000 matches.
  • Processing happens in your browser, so very large inputs are bounded by available memory. Files above roughly 10 MB are handled but will feel slower, and multi-hundred-megabyte files are better suited to a command-line tool.

Questions


Should I include the slashes around my pattern?
No. Enter just the pattern and use the checkboxes for flags. Including slashes would make them literal characters to match.
Why was my pattern rejected as unsafe?
It contains a nested quantifier such as (a+)+, which can take exponential time and would freeze the page, since JavaScript cannot interrupt a running regex. Simplify it, (a+)+ is equivalent to a+.
Will these results match my Python or PHP regex?
Not necessarily. This uses the JavaScript engine, and flavours differ on lookbehind, named group syntax and Unicode handling. The results are exactly right for JavaScript.
Why does a capture group show "(no match)"?
The group is in a branch that did not participate, typically inside an alternation or an optional section. In code that group is undefined rather than an empty string, which is a frequent source of bugs.