How to Use AI to Write a Regex You Can Trust
AI writes a working regex fast, but hides which edge cases it silently skipped. Verify any pattern with this three-step process before it ships.
How to Use AI to Write a Regex You Can Trust
AI writes a working regex on the first try more often than most people write one on the fifth. The problem is not accuracy, it is confidence: a model hands you a pattern with no hint of which edge cases it silently ignored. Treat the output as a draft that needs three specific checks, not a finished answer, and the process below takes about the same time as writing a mediocre regex by hand.
Why the failure mode is different from normal AI mistakes
A wrong SQL query usually errors loudly. A wrong regex usually does not. It matches almost everything you wanted, misses a handful of inputs nobody tested, and ships. Three patterns show up constantly in AI-generated regex:
Greedy where it should be lazy. `<.*>` will span across multiple tags in a string with more than one, because `.*` grabs as much as it can. The fix is `<.*?>`, and a model asked for "match an HTML tag" will pick either one depending on phrasing.
Missing anchors. A pattern meant to validate a whole string (a username, a postcode) but written without `^` and `$` will happily match a valid substring inside garbage input.
Catastrophic backtracking. Nested quantifiers like `(a+)+b` can take exponential time on the wrong input. This is the one that turns into an incident, not just a bug, because it looks fine in testing and then hangs in production on a longer string.
None of these show up by reading the regex. They show up by running it against inputs designed to break it.
The three-step process
1. Ask for the pattern and the reasoning, not just the pattern
A prompt that asks for output only gets output. Ask the model to state its assumptions:
Write a regex that matches [what you need].
Then list, as bullet points:
- What counts as a valid match, with one example
- What it will incorrectly match (false positives)
- What it will incorrectly miss (false negatives)
- Whether it needs case-insensitive or multiline flagsReading the false-positive and false-negative list is usually where you catch the real gap, because it forces the model to articulate a boundary it would otherwise leave implicit. If the list is empty or vague, that itself is a signal to push back and ask again.
2. Build your own adversarial test cases
Do not trust the model's own test cases, since a model that got the pattern wrong will often get its self-check wrong the same way. Write five to eight inputs yourself, aimed at the exact failure modes above:
Test input type | What it catches |
|---|---|
Empty string | Missing anchors or unhandled edge case |
Minimal valid match | Confirms the happy path actually works |
Valid match with unusual but legal characters | Overly narrow character classes |
Almost-valid string that should fail | Missing anchors, wrong boundaries |
Very long repeated pattern (50+ chars) | Catastrophic backtracking |
Multiple matches in one string | Greedy vs lazy quantifier bugs |
Run these against the pattern in your actual language's regex engine, not a general-purpose regex playground, because flavors differ: JavaScript, PCRE, and Python's `re` module disagree on lookbehind support, named groups, and a few escape sequences.
3. Time it on a long adversarial string
For anything that runs on user input rather than a fixed internal string, test the pattern against a long, repetitive, almost-matching string and check it does not take more than a few milliseconds:
Input: "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa!"If a pattern with nested quantifiers takes noticeably longer on 60 characters than on 30, it is exponential, and doubling the input length again will make that obvious fast. Ask the model directly whether a given pattern is vulnerable to catastrophic backtracking. It is often right when asked, even when it was wrong by default.
A worked example
Asked for "a regex for a valid product SKU, letters and numbers, 6 to 10 characters," a model commonly returns:
[A-Za-z0-9]{6,10}This looks right and passes an obvious test. It also matches a 6-character substring inside a 40-character garbage string, because there is no `^` and `$`. The corrected version:
^[A-Za-z0-9]{6,10}$is a one-character difference that the three-step process catches on the first "almost-valid string" test, and would not have been caught by eyeballing the pattern.
Validation, extraction, and replacement need different scrutiny
Not every regex carries the same risk, and it helps to know which category you are in before deciding how hard to test.
Validation (does this string match a shape, like a SKU or a postcode) is the highest-stakes category, because a false negative rejects a real user and a false positive lets bad data through. This is where missing anchors bite hardest, and where the test table above matters most.
Extraction (pull the phone number out of this block of text) tolerates missing a few matches better than it tolerates grabbing the wrong span. Ask the model for a pattern with named capture groups so the extracted piece is unambiguous:
(?<area_code>\d{3})-(?<number>\d{3}-\d{4})Test extraction patterns against text with more than one plausible match nearby, since that is exactly where greedy quantifiers grab too much or a capture group binds to the wrong span.
Replacement (find this pattern and swap it for something else) is the category where a scope-too-wide pattern does the most damage, because it silently rewrites text that only coincidentally matched. Before running a replacement regex anywhere near production data, run it in dry-run or preview mode first and read every match it found, not just a sample.
When to skip the regex entirely
Some things that look like regex problems are better solved by a parser or a library, particularly emails, URLs, and phone numbers, where the real specification is far messier than intuition suggests. If you are prompting AI for an email regex "good enough for a signup form," say that explicitly so the model does not hand you an attempt at full RFC 5322 compliance, which is genuinely enormous and not what most forms need.
This same discipline, generating a plausible first draft and then verifying it against cases you design rather than cases the model designs, applies directly to the rest of an AI-written codebase. We cover the general version of this in how to debug AI-generated code, and the test-first variant in how to get an AI coding agent to write tests first. If a pattern fails after it was working, how to prompt AI to fix a flaky test covers the closely related case of intermittent failures. For the broader question of why confident-sounding output is not the same as correct output, see what is an AI hallucination.
FAQ
Can AI write a regex that handles every edge case automatically?
No. It writes a pattern that handles the cases implied by your prompt. Vague prompts produce patterns with vague boundaries. Specific prompts, including the characters and lengths you actually expect, produce tighter patterns.
Is it safe to use an AI-generated regex on user-facing forms without testing?
Not for anything beyond a quick internal script. Untested regex on user input is how both false rejections (valid input blocked) and catastrophic backtracking (a denial-of-service on your own form) reach production.
Which AI models are best at writing regex?
Any current frontier model handles common patterns well. The gap between models shows up on obscure syntax, like lookbehind assertions in engines that support them inconsistently, and on catching their own backtracking risk when asked directly.
How do I test a regex for catastrophic backtracking?
Feed it a long string, 40 to 80 characters, built from the same repeated character the pattern's quantifiers operate on, and time the match. A pattern with nested quantifiers like `(a+)+` will show a sharp jump in execution time as the string grows, well before it becomes unusable.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


