Why regex looks scarier than it is
A block of regex like ^[\w.+-]+@[\w-]+\.[a-zA-Z]{2,}$ looks like line noise the first time you see it. But almost all everyday regex use is just a handful of building blocks, repeated and combined: match a digit, match a letter, match this many times, stop here. You don't need to memorize the full specification to validate an email format or pull every phone number out of a block of text.
Part of the intimidation is purely visual. Regex has no spaces, no keywords, no obvious structure to a first-time reader, just a dense string of symbols that all mean something different depending on where they sit. That's a real learning curve, but it's a shallow one. Once you can recognize a dozen or so symbols on sight, most patterns stop looking like noise and start reading almost like a sentence: "start of string, one or more digits, an at-sign, one or more word characters."
Most developers use regex the same way they use fifteen keyboard shortcuts out of the hundred available. You learn the few patterns that solve 90% of the problems you actually run into, and you look up or test the rest when you need them. Nobody has the entire regex specification memorized, including people who use it daily. What they have is a small working vocabulary and a fast way to check anything outside it.
There's also a common misconception worth clearing up early: regex isn't really about validating that data is correct, it's about checking that data has the right shape. A pattern that confirms something looks like an email address doesn't know or care whether that address actually receives mail. That distinction matters because it explains why the "simple" patterns further down this page are intentionally not exhaustive. They check shape, not truth, and that's usually exactly what you need.
That's really the whole approach this page takes: a short reference table, three patterns worth memorizing, one gotcha that trips up almost everyone at least once, and a way to check your work without guessing.
The building blocks, as a quick reference table
These cover the large majority of everyday regex. Once these are familiar, most patterns you'll encounter are just combinations of them, stacked in different orders to describe a different shape of text.
A few things are worth noticing as you read the table. First, the lowercase-vs-symbol pairs are opposites: where \d, \w, and \s match something, their uppercase versions, \D, \W, and \S, match the opposite (anything that is not a digit, word character, or whitespace). Second, quantifiers like +, *, and ? never stand alone. They always apply to whatever came immediately before them, whether that's a single character, a character class, or a group in parentheses.
| Pattern | Matches | Example |
|---|---|---|
\d | Any digit (0-9) | \d\d\d matches "123" |
\w | Any word character (letters, digits, underscore) | \w+ matches "hello_2" |
\s | Any whitespace (space, tab, newline) | a\sb matches "a b" |
. | Any single character except a newline | c.t matches "cat" or "cut" |
+ | One or more of the previous character or group | \d+ matches "7" or "70042" |
* | Zero or more of the previous character or group | ab* matches "a" or "abbb" |
? | Zero or one of the previous character or group (optional) | colou?r matches "color" or "colour" |
^ | Start of the string (or line, in multiline mode) | ^Hi matches "Hi there" but not "Say Hi" |
$ | End of the string (or line, in multiline mode) | bye$ matches "goodbye" but not "bye now" |
[] | Any one character from a set | [aeiou] matches any single vowel |
| | Either the pattern on the left or the right (or) | cat|dog matches "cat" or "dog" |
Three patterns you'll actually reuse
These aren't meant to be bulletproof against every edge case in the spec they're named after. They're the practical versions people actually paste into their code, because they cover the real-world shape of the data without turning into an unreadable wall of escape characters.
- A simple email-format check. Not a fully RFC-compliant pattern, just something that confirms the input looks like an email: letters or numbers (plus a few common symbols like dots, plus signs, and hyphens), an @ symbol, and a domain with at least one dot and two letters after it.
^[\w.+-]+@[\w-]+\.[a-zA-Z]{2,}$This will happily reject "not-an-email" and accept "name@example.com". It will also accept some strings that aren't real, deliverable addresses, because confirming an address is real requires actually sending mail to it, not a regex.
- A basic phone number matcher. Digits with optional separators like spaces, dashes, or parentheses, loose enough to cover most of the ways people actually type phone numbers.
^[\d\s()+-]{7,15}$The
{7,15}part sets a minimum and maximum length rather than an exact one, since phone number length varies by country and by whether someone includes a country code. - Stripping extra whitespace. Collapse repeated spaces, tabs, or line breaks down to a single space, useful for cleaning up pasted or user-submitted text before you store or display it.
text.replace(/\s+/g, ' ')The
gflag at the end matters here. Without it, only the first run of whitespace gets replaced; with it, every run in the string gets collapsed.
Notice all three lean on the same building blocks from the table above, just arranged differently. That's the pattern behind the patterns: once you're comfortable with ten or so symbols, writing a new one is mostly about deciding which symbols describe the shape of the text you already have in front of you.
Greedy vs lazy matching, briefly
This is the single most common regex surprise, and it catches beginners and experienced developers alike. Quantifiers like * and + are greedy by default, meaning they try to match as much text as possible before backing off. Against a string with more than one instance of what you're looking for, that can go badly wrong in a way that's hard to spot just by reading the pattern.
Take a pattern like ".*" run against the text "first" and "second". You might expect it to match just "first". Instead, because .* is greedy, it matches all the way from the first quote to the very last one: "first" and "second", swallowing everything in between, including the word "and" and the second pair of quotes.
This isn't a bug, it's the defined behavior of the quantifier. Greedy matching starts by trying to consume the entire rest of the string, then backs off one character at a time until the rest of the pattern (in this case, the closing quote) can still match. Since there's a closing quote way at the end of the string, that's where it stops backing off.
The fix is adding a ? right after the quantifier, turning it lazy (sometimes called non-greedy). .*? matches as little as possible instead of as much as possible, expanding only when the rest of the pattern can't match otherwise. Run against the same text, ".*?" stops at the very first closing quote it finds, correctly matching just "first".
The same ? trick works after other quantifiers too: +? is a lazy version of +, and {2,5}? is a lazy version of {2,5}. You won't need lazy matching for every pattern, but it's worth trying the moment a greedy match grabs more text than you expected.
Greedy grabs everything it can and gives back only what it must. Lazy grabs the minimum and only takes more when forced to. When a match looks suspiciously long, greedy quantifiers are almost always the reason.
Testing a pattern without guessing
Reading a regex pattern and knowing exactly what it will and won't match is a skill that takes a while to build, and even experienced developers get it wrong on anything longer than a few characters. Rather than guessing, it's faster to just run it against real examples and look at what actually matched.
Our own free Regex Tester highlights every match live as you type your pattern and your test string, so you see immediately whether your regex is too greedy, too strict, or exactly right. There's nothing to install and no sign-up, just a pattern field, a test string, and instant visual feedback on what matched, updated on every keystroke.
That immediate feedback loop is the real shortcut here. Instead of memorizing edge cases, you can just try a pattern against real examples of the text you're trying to match, watch which parts light up as a match, and adjust the pattern until the highlighting covers exactly what you meant it to. If a match runs too long, that's your cue to check for a greedy quantifier that needs a lazy version instead. If nothing highlights at all, that usually means a character class or an escaped symbol isn't quite what you think it is.
Testing this way also makes it much easier to build confidence in a pattern before you drop it into production code, where a regex that's subtly wrong can silently reject valid input or, worse, accept input it should have caught. A minute spent testing against a handful of realistic examples, including the awkward edge cases you'd rather not think about, is usually cheaper than debugging a support ticket later.