Unit Testing a Regex Pattern

Pin a pattern down with a small regression suite of should-match and should-not-match cases, then share the whole suite as a single link a teammate can open and verify.

Unit Testing a Regex Pattern

A pattern that passes on the one string you tried it against is not tested; it is only unrefuted. The gap between those two states closes the moment you write down a handful of cases the pattern must handle correctly, positive and negative alike, and check them as a group instead of trusting a single example to represent every input your code will ever see.

Run this check yourself in the Regex Tester & Match Debugger.

Open in the tool →

Why one passing example is not the same as a tested pattern

Regular expressions fail in a specific, predictable way: they usually match more than the author intended, not less. A pattern like <[A-Za-z0-9]+> written to match HTML tags also matches something like <1>, which is not a valid tag at all, because a plus quantifier happily repeats past the point the author actually meant to stop.1 A pattern written and tested against exactly one happy-path string almost never gets caught matching too broadly, because the one string you tried never exercised the ambiguous edge the pattern actually leaves open.

Unit Tests mode exists to close that gap by treating a pattern the way you would treat a function: define a small set of inputs, decide what the correct output should be for each one, and run all of them together every time the implementation changes.2 A product SKU pattern like ^[A-Z]{2}\d{4}$ looks obviously correct at a glance, yet a test batch immediately reveals whether it correctly rejects a lowercase variant or an extra trailing digit, the exact kinds of malformed input real user entry actually produces. Anchoring is what makes that verdict trustworthy: ^ matches the position before the first character and $ matches right after the last, so ^ and $ together force the entire string to satisfy the pattern rather than just some substring of it3.

Writing negative cases that earn their place in the suite

A negative test case, one explicitly marked should-not-match, is doing real work only when it targets something the pattern plausibly could get wrong, not an input so obviously invalid it was never in question. For the SKU pattern, ab1234 tests the case-sensitivity assumption directly, and AB12345 tests whether the digit count is actually being enforced rather than merely suggested by the pattern's shape.

Reading the pass count as you refine the pattern

Every edit to the pattern field re-runs the entire test batch in one pass, and the summary line above the rows reports a simple fraction: how many of your defined cases currently pass. That single number turns "I think I fixed it" into a checkable claim the moment you finish typing, rather than something you have to verify by re-reading each row individually. Each row's verdict is the boolean returned by a match test, true or false, and that boolean is the same signal a real application gets when it asks a pattern whether the input is acceptable4.

Iterating without losing track of which case is failing

Each row keeps its own PASS or FAIL badge visible at all times, so tightening the pattern to fix one failing case while accidentally breaking a previously-passing one shows up immediately as a badge flipping color, not as a silent regression you discover later. That immediate feedback loop is what makes iterating on a tricky pattern here faster than iterating inside application code, where you would otherwise need to re-run a whole test file just to see the effect of one small change.

Once every row shows PASS, you have something stronger than a pattern that "seems to work": you have a pattern with a specific, written-down definition of correct behavior that it currently satisfies, case by case. That written definition is what makes the pattern maintainable months later, when whoever edits it next has a concrete list of behavior to preserve instead of a single vague memory of what it was originally supposed to do.

Sharing the whole suite instead of just the pattern

A pattern shared on its own loses the context that makes it trustworthy: what it was actually tested against, and what a correct result was supposed to look like for each case. Clicking Share encodes the pattern, its flags, the test text, and every row in your Unit Tests suite, expectation included, into a single link and copies that link to your clipboard. Writing the encoded state into the address bar this way updates only what the browser displays; it never triggers a page load or a check that anything at that address actually exists.5

What a teammate sees the moment they open the link

Sending that link to a teammate hands them not just your pattern but the evidence it works, already loaded and ready to inspect the moment they open it. They see the same pass count you saw, can add a case you missed, and can confirm your fix actually addressed the specific input that was failing before, all without re-typing a single test string from a description in a chat message.

Because the entire suite lives inside the link itself rather than on a server, opening it later still reproduces the exact same state, pattern, flags, test text, and every row's expectation, even if the pattern in your actual codebase has since moved on. That makes a shared link a durable snapshot of "here is what I verified," worth attaching directly to a pull request or a code review comment instead of describing the test cases in prose.

When to use this

Use this guide whenever you want to pin a pattern down with a small regression suite of should-match and should-not-match cases before relying on it in code.

Examples

Building a regression suite for a product SKU pattern

Before
Pattern: ^[A-Z]{2}\d{4}$
Test strings: AB1234, ab1234, AB12345
After
AB1234 is marked should match and passes. ab1234 and AB12345 are marked should not match and both correctly fail, so the summary line reads 3 of 3 passing.
Sources
  1. 1.

    Jan Goyvaerts, "Regex Tutorial: Repetition with Star and Plus," regular-expressions.info, accessed August 2026. https://www.regular-expressions.info/repeat.html

  2. 2.

    Martin Fowler, "Unit Test," martinfowler.com, May 2014. https://martinfowler.com/bliki/UnitTest.html

  3. 3.

    Jan Goyvaerts, "Regex Tutorial: Start and End of String or Line Anchors," regular-expressions.info, accessed August 2026. https://www.regular-expressions.info/anchors.html

  4. 4.

    MDN Web Docs, "RegExp.prototype.test()," developer.mozilla.org, accessed August 2026. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/RegExp/test

  5. 5.

    MDN Web Docs, "History: replaceState() method," developer.mozilla.org, accessed August 2026. https://developer.mozilla.org/en-US/docs/Web/API/History/replaceState

Testing a URL-Slug Pattern

A slug generator that mostly works is worse than one that obviously does not, because the failures only show up once a stray capital letter or double hyphen ships into a live URL. Pinning down exactly what counts as a valid slug, lowercase letters and digits, single hyphens as separators, nothing else, before that pattern reaches a routing layer or a database constraint, catches the failure while it still costs nothing to fix.

Run this check yourself in the Regex Tester & Match Debugger.

Open in the tool →

What a strict slug pattern actually needs to enforce

A URL slug has a narrower job than most strings a validator checks: lowercase letters, digits, and single hyphens used only as separators between words, with no leading, trailing, or doubled hyphen anywhere in the string. Letters, digits, hyphens, and a small handful of other characters are the "unreserved" set the URI specification exempts from percent-encoding, which is exactly why a slug built only from that set survives every URL context untouched.1 That narrowness is exactly what makes it worth testing carefully, since a pattern that is slightly too permissive can let a genuinely broken slug slip straight through into a URL.

The tester's built-in URL slug snippet, available from the Snippets menu, encodes that rule as ^[a-z0-9]+(?:-[a-z0-9]+)*$. Anchored at both ends so a partial match at the start or end of a string cannot pass, the pattern reads as one or more lowercase alphanumeric characters, optionally followed by any number of hyphen-plus-alphanumeric groups. A capital letter, a leading hyphen, or two hyphens in a row all break that structure and correctly fail the match. The same lowercase-letters-digits-hyphens shape shows up well beyond URL slugs; Google Cloud's own resource-naming rules for Compute Engine require names to match a near-identical pattern, since infrastructure identifiers and SEO-friendly slugs both converge on the same practical, encoding-safe character set.2

Checking one candidate at a time in Match mode

Before committing to a batch of test cases, pasting a single candidate slug into the test text field and watching whether it highlights confirms the pattern behaves the way you expect on an obvious case. my-product-name should highlight as a full match; My-Product-Name should not, since the pattern has no case-insensitive flag active by default, and the engine only folds letter case when that flag is explicitly turned on.3 That quick check takes seconds and catches an obviously wrong pattern before you invest time building out a larger batch of edge cases around it.

Moving from one check to a whole batch with Unit Tests mode

Checking slugs one at a time in Match mode works for a quick sanity check, but a slug generator needs to handle dozens of edge cases correctly, not just the one you happened to think of first. Unit Tests mode is built exactly for that shift: each row holds one candidate string and an explicit should-match or should-not-match expectation, and every edit to the pattern re-runs the entire batch at once.

Building a batch that actually exercises the edge cases

A useful test batch for a slug pattern includes the obvious positive case, my-product-name, alongside the specific failures that matter most: a version with a capital letter, a version with a leading hyphen, a version with a trailing hyphen, and a version with a doubled hyphen from a bad string-replace somewhere upstream. Marking each row's expectation correctly and watching the pass count update live turns "the pattern looks right" into a verified claim backed by five or six concrete assertions.

Every PASS or FAIL badge sits directly next to its row, so a single failing case among a dozen passing ones is immediately visible rather than buried in a wall of text output. That visibility is what makes Unit Tests mode faster than re-reading the pattern by eye once your test batch grows past two or three candidates.

Deciding what happens to a slug that fails validation

A validation pattern only tells you a candidate string is invalid; it does not fix it. Most slug pipelines pair this kind of check with a normalization step upstream, lowercasing the input, replacing spaces and punctuation with single hyphens rather than underscores, since search engine crawling guidance specifically recommends hyphens as word separators over underscores, and trimming any leading or trailing hyphen the replacement step introduced, so that a failure here is rare in production rather than routine.4

Treating a validator failure as a signal about the pipeline

If your validator is failing regularly on real user-submitted titles, that is usually a sign the normalization step upstream needs work, not that the validation pattern itself is too strict. Re-checking a single candidate by hand means anchoring the pattern again: ^ asserts the beginning of input and $ asserts the end of input, so the same string returns a different verdict the moment either anchor is dropped5. Test the normalized output against this pattern, not the raw title text a user actually typed, since the two are supposed to represent different stages of the same pipeline and testing the wrong one gives you a misleading picture of where the real problem sits.

When to use this

Use this guide when you need to validate a lowercase, hyphen-separated slug pattern against a batch of candidate strings before wiring it into a routing layer or database constraint.

Examples

Testing a slug pattern against several candidate strings at once

Before
Pattern: ^[a-z0-9]+(?:-[a-z0-9]+)*$
Test strings: my-product-name, My-Product-Name, -leading-hyphen, double--hyphen
After
my-product-name is marked should match and passes. The other three are marked should not match and all correctly fail, for a pass count of 4 of 4.

This is the same "URL slug" pattern available from the Snippets menu inside the tool.

Sources
  1. 1.

    T. Berners-Lee, R. Fielding, and L. Masinter, "Uniform Resource Identifier (URI): Generic Syntax," RFC 3986, IETF, January 2005. https://www.rfc-editor.org/rfc/rfc3986

  2. 2.

    Google Cloud, "Naming resources," docs.cloud.google.com, accessed August 2026. https://docs.cloud.google.com/compute/docs/naming-resources

  3. 3.

    MDN Web Docs, "RegExp.prototype.ignoreCase," developer.mozilla.org, accessed August 2026. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/RegExp/ignoreCase

  4. 4.

    Google Search Central, "URL Structure Best Practices for Google Search," developers.google.com, accessed August 2026. https://developers.google.com/search/docs/crawling-indexing/url-structure

  5. 5.

    MDN Web Docs, "Assertions," developer.mozilla.org, accessed August 2026. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Guide/Regular_expressions/Assertions

FAQ

The most common cause is a doubled hyphen or a leading/trailing hyphen left over from a normalization step upstream, both of which are invisible at a glance but fail a strict slug pattern immediately. Paste the exact string your pipeline produced, not a retyped version, since retyping tends to accidentally fix the very issue you are trying to catch.

Search engine crawling guidance recommends hyphens over underscores as word separators, but the final call is still yours to enforce consistently. If your routes only ever use hyphens as word separators, keeping underscores out of the pattern catches an inconsistent slug before it reaches your database; if your project already mixes both, adjust the character class to match your actual convention.

Yes. Unit Tests mode is built for exactly that: add a row per candidate, mark each one's expectation, and CapyToolkit re-runs every row against the current pattern the moment you edit it, so a batch of twenty candidates checks just as fast as one.

It handles the common convention, lowercase alphanumeric characters separated by single hyphens, which covers most blog, product, and article URLs. If your project allows numbers-only slugs, multi-byte characters, or a different separator, adjust the snippet rather than assuming it fits every project unmodified.

No. The tester compiles your pattern with JavaScript's native RegExp engine, the same one your Node.js or browser code already runs, so a batch of candidates that passes here behaves identically once the pattern reaches your validation function.

FAQ

There is no fixed number, but a useful minimum is one obvious positive case plus one negative case for each assumption the pattern makes, such as case sensitivity, length, or an allowed character set. Three to six focused cases usually catch more real bugs than a dozen redundant variations of the same string.

Yes. Every keystroke in the pattern field re-runs the full batch of test rows and refreshes both the individual PASS/FAIL badges and the summary count, so you always see the current state of every case without clicking a separate 'run tests' button.

The pattern, its active flags, the test text, and every row in your Unit Tests suite including each row's should-match or should-not-match expectation. CapyToolkit doesn't store any of that server-side; the entire state lives inside the URL itself, so it reaches only the people you send the link to.

Yes. Opening a shared link loads the suite exactly as it was shared, and you can add, edit, or remove rows freely from there. Clicking Share again afterward generates a new link reflecting your changes; the original link a teammate already has stays exactly as you sent it.

It depends on the pattern, but negative cases usually deserve more attention, since a regex is far more likely to accidentally match too much than too little. A single positive case plus several targeted negative cases, each testing a different way the pattern could be too permissive, tends to catch more real bugs than the reverse ratio.