What is URL percent-encoding
Every URL is constrained to a limited set of printable ASCII characters. Characters outside that set,
including spaces, non-ASCII letters, and many punctuation marks, cannot appear directly in a URL without
causing ambiguity or transmission errors. Percent-encoding solves this by replacing each unsafe byte
with a percent sign followed by two uppercase hexadecimal digits. A space becomes %20, a
hash becomes %23, an at-sign becomes %40, and a non-ASCII character like
the accented letter é becomes the multi-byte sequence %C3%A9.1
The specification is defined in RFC 3986. It separates characters into three categories. Unreserved characters (letters, digits, hyphens, dots, underscores, and tildes) are safe to use as-is anywhere in a URL. Reserved characters (such as /, ?, #, &, and =) have structural meaning and must be percent-encoded when they appear as data values rather than structural delimiters. All other characters must always be percent-encoded.2
The é example from the opening paragraph generalizes. A URL is a byte stream, and percent-encoding escapes bytes, so a character whose UTF-8 encoding spans several bytes becomes several escapes: é takes two bytes and therefore two escapes, and an emoji takes four and therefore four.1 One decode restores the character, because the escapes return the original bytes and those bytes are then read back together as UTF-8 text. Counting escapes is therefore a rough way to count bytes: four escapes in a row usually means one four-byte character, not four separate ones.
Why percent-encoded strings show up everywhere
Browser DevTools, server logs, and API responses frequently display percent-encoded strings. Knowing how
to read and reverse them is a practical skill for debugging HTTP requests, constructing URLs
programmatically, and understanding what a link is actually requesting before you follow it.
Spotting the difference between a literal ampersand and an encoded %26, for example, can explain why a
URL behaves differently than it looks in a log line.
Double encoding and the %25 pattern
Double encoding is the most common percent-encoding bug. A literal percent sign must itself be escaped
as %25, so running an encoder over a string that was already encoded turns %20 into
%2520 and %3D into %253D. The doubled artifact appears whenever two layers of a
pipeline each encode, when a library encodes a value the caller already encoded, or when an encoded
string is pasted into a field that encodes on save. The result is a URL that decodes to another encoded
string instead of the value you meant to send, so the server receives a different value than the one
you intended.
Spotting it here takes one decode pass. Decode a suspect string once: if the output still contains percent escapes, that is the tell, because the first pass only peeled the outer shell. Copy that output, paste it into the input, and decode again to recover the original. The code-side fix is to encode exactly once, at the boundary that builds the URL, and never to re-encode data that arrived already encoded. This tool performs one decode per interaction and never re-decodes automatically, which is what makes the two-step check reliable and repeatable.
Percent-Encoding URLs: When and How
Knowing how percent-encoding works is half the job; the other half is knowing which part of the URL you are encoding, because the path, the query string and the fragment each allow different characters. The guide below shows which characters to encode in each component, the encoding functions for each language and context, the mistakes that lead to double encoding and how fragments and mailto links differ from ordinary URLs.
Percent-encoding converts characters that are not allowed in a URL component into a %XX hexadecimal representation. Every byte of the UTF-8 representation of the character is encoded separately: the space character becomes %20, and the euro sign (€, U+20AC) becomes %E2%82%AC1. Encoding rules differ by component because characters that are safe in a path segment may be delimiters in a query string, requiring encoding only in that context.
Most URL-related bugs involving encoding fall into two categories: double-encoding (encoding an already-encoded string, turning %20 into %25202) and context mismatch (encoding a full URL string instead of individual values, which turns / into %2F and breaks routing). Using context-aware encoding functions prevents both.
Which characters to encode and where
Unreserved characters (A–Z, a–z, 0–9, hyphen, period, underscore, tilde) never require encoding in any URL component. Reserved characters have special meaning in URLs (:/?#[]@!$&'()*+,;=) and must be encoded when used as literal data rather than as delimiters1. Consequently, path segments may contain unreserved characters, :, @, !, $, &, ', (, ), *, +, ,, ;, and = without encoding, but spaces and most other special characters must be encoded as %20. Query parameter values must encode all reserved characters except ~ and the safe set. Building on this, fragments follow the same rules as query parameters, so the same encoding function can be applied to both components without modification.
Encoding functions by language and context
JavaScript provides two encoding functions for URL components: encodeURIComponent() encodes everything except unreserved characters2 and is correct for path segments and query parameter values. encodeURI() is less strict because it preserves :/?#@!$&'()*+,;=, so it is suitable for encoding a complete URL that is already partially formed but should not be re-encoded. For query parameter values in Python, urllib.parse.quote(value, safe='') encodes per RFC 3986 (%20 for spaces)3. Building on this, quote_plus(value) encodes for form data format (+ for spaces), which is the correct choice for HTML form submission query strings. Never use quote() on a complete URL string.
Common mistakes and double-encoding
Double-encoding occurs when an already-encoded value is encoded again: a%20space encoded again produces a%2520space. Always decode before re-encoding: URLDecoder.decode(value) in Java, unquote(value) in Python, decodeURIComponent(value) in JavaScript. Consequently, when a server receives a query parameter and passes it to another URL, decode the received value and re-encode it for the new context. Building on this, + in a query string means a space only in application/x-www-form-urlencoded format, while in an RFC 3986 URL + is a literal plus sign4. Using the wrong decoder produces subtle data corruption when values contain plus signs, so always verify which format the source URL uses before choosing a decoder.
Percent-encoding in the fragment component
The fragment component follows the same percent-encoding rules as the query string, but with one key difference: the application/x-www-form-urlencoded convention (+ for spaces) does not apply. A fragment of #q=hello+world means the literal string "hello+world", not "hello world". This distinction matters for single-page applications that store search state in the fragment: if your JavaScript reads window.location.hash and splits on + expecting spaces, you will misinterpret the user's input. Always use %20 for spaces in fragment values, and decode with decodeURIComponent() rather than a form-specific decoder.
Fragment encoding in OAuth implicit grants
The OAuth 2.0 implicit grant delivers the access token in the fragment: https://app.example.com/callback#access_token=abc123&token_type=Bearer5. The fragment is never sent to the server, which protects the token from appearing in server logs. However, the fragment is visible to any JavaScript on the page, including third-party analytics scripts. This is why the implicit grant is deprecated in favor of the authorization code flow with PKCE: the code flow delivers tokens via a back-channel POST that is invisible to page-level JavaScript. If you must use the implicit grant, ensure no third-party scripts have access to the callback page.
Because the token lives in the fragment only for the lifetime of the page, extract it immediately on load and clear the fragment with history.replaceState before any analytics or routing code runs. Treat the token as already exposed the moment it appears, since the fragment persists in browser history and can be read by any script that later loads on that origin. Removing it from the URL as early as possible limits the window in which a compromised library could capture it.
Encoding differences between URL components in practice
Each URL component has its own encoding context, and a character that is safe in one component may need encoding in another. The @ character is valid in the userinfo portion of the authority (user@host) but must be encoded as %40 in a path segment. The / character is a valid path delimiter but must be encoded as %2F inside a single path segment. The ? character starts the query component but is valid (and must be encoded as %3F) inside a path segment. Context-aware encoding functions handle these rules: encodeURIComponent() in JavaScript encodes everything except unreserved characters, making it safe for any component.
Encoding slashes in REST API paths
A common REST API design question is how to handle resource identifiers that contain slashes. A file path like documents/2026/report.pdf cannot appear directly in a URL path segment without being interpreted as three separate segments. The standard solution is to encode the slash as %2F: /files/documents%2F2026%2Freport.pdf. However, some web servers (Apache with AllowEncodedSlashes, nginx with merge_slashes) decode %2F before routing, which means the server sees the decoded path. Test your server's behavior with encoded slashes before relying on this pattern, or use a query parameter instead: /files?path=documents/2026/report.pdf.
URL encoding in email and mailto links
Mailto links require percent-encoding for special characters in the subject and body parameters. A mailto: link with a subject containing an ampersand or question mark breaks the URL structure unless those characters are encoded: mailto:[email protected]?subject=Hello%20%26%20Goodbye&body=Line%201%0ALine%2026. The %0A encodes a newline in the body. Not all mail clients handle complex mailto links correctly; some truncate at the first &, others ignore the body parameter entirely. For reliable email composition, keep mailto links simple (address only) and use a contact form for messages that require formatting.
Encoding URLs in HTML attributes
When embedding URLs in HTML attributes (href, src, action), the URL must be both percent-encoded (for URL syntax) and HTML-escaped (for HTML syntax). An ampersand in a query parameter must be percent-encoded as %26 for the URL layer, and the entire attribute value must escape & as & for the HTML layer, since a raw ampersand breaking inside an href attribute is what happens when only one of those two layers gets applied. The correct form is: <a href="https://example.com/search?q=hello%20world&page=2">. The & in the HTML source becomes & when the browser parses the HTML, and the resulting URL contains the literal & that separates query parameters. Forgetting the HTML-level escaping is one of the most common bugs in hand-written HTML with query parameters7.
Apply percent-encoding whenever you interpolate user-supplied values, file names, or non-ASCII text into a URL component such as a path segment, query parameter value, or header value that contains a URL.
Encode path segment and query value correctly in JavaScript
// Wrong: encodeURI applied to a component value leaves & and = unencoded
const category = "food & drink";
const url = `/items?category=${encodeURI(category)}`;
// → /items?category=food%20&%20drink (& breaks the query!) // Correct: encodeURIComponent encodes & and = in values
const url2 = `/items?category=${encodeURIComponent(category)}`;
// → /items?category=food%20%26%20drink Avoid double-encoding a URL parameter in Python
# Wrong: value already encoded encoding again produces %25 from urllib.parse import quote received = "hello%20world" # came from a URL query param encoded = quote(received, safe="") # → "hello%2520world" (double-encoded!)
# Correct: decode first then re-encode for the new URL from urllib.parse import unquote, quote decoded = unquote(received) # → "hello world" re_encoded = quote(decoded, safe="") # → "hello%20world"
- 1.
T. Berners-Lee, R. Fielding, and L. Masinter, "Uniform Resource Identifier (URI): Generic Syntax," RFC 3986, IETF, January 2005. https://www.rfc-editor.org/rfc/rfc3986.html
- 2.
Mozilla Developer Network, "encodeURIComponent() — JavaScript," developer.mozilla.org, accessed June 2026. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/encodeURIComponent
- 3.
Python Software Foundation, "urllib.parse — Parse URLs into components," docs.python.org, accessed June 2026. https://docs.python.org/3/library/urllib.parse.html
- 4.
WHATWG, "URL Standard — application/x-www-form-urlencoded," url.spec.whatwg.org, accessed June 2026. https://url.spec.whatwg.org/#application-x-www-form-urlencoded
- 5.
D. Hardt, "The OAuth 2.0 Authorization Framework," RFC 6749, IETF, October 2012. https://datatracker.ietf.org/doc/html/rfc6749
- 6.
M. Duerst, L. Masinter, and J. Zawinski, "The 'mailto' URI Scheme," RFC 6068, IETF, October 2010. https://www.rfc-editor.org/rfc/rfc6068.txt
- 7.
Stack Overflow, "Do I encode ampersands in <a href...>?," stackoverflow.com, accessed June 2026. https://stackoverflow.com/questions/3705591/do-i-encode-ampersands-in-a-href
In RFC 3986 URLs, %20 is the percent-encoded form of a space character and + is a literal plus sign. In the application/x-www-form-urlencoded format (HTML form submissions), + also represents a space. When in doubt, use %20 because it is unambiguous in all contexts. Use + only when explicitly building form-encoded query strings, and document which convention your API expects so that clients do not have to guess.
encodeURI() is designed for encoding a complete URL because it preserves all characters that have meaning in URL structure (:/?#@). Use it only when the input is a full, already-formed URL that you want to embed in another context. encodeURIComponent() encodes everything except unreserved characters, so use it for encoding individual path segments and query parameter values. CapyToolkit offers a URL Parser tool that runs both functions on any input so you can compare their output side by side.
Split the path into segments, encode each segment with encodeURIComponent() (JavaScript) or urllib.parse.quote(segment, safe='') (Python), and join them with /. Never encode the entire path string at once because that would encode the / separators, breaking the path structure.
RFC 3986 Section 6.2.2.1 specifies that uppercase is the canonical form. %2F and %2f refer to the same character, but %2F is normalized. Treat them as equivalent when decoding; emit %2F (uppercase) when encoding for consistency and compatibility with signature algorithms that compare the raw encoded strings.
Convert the character to its UTF-8 byte sequence, then percent-encode each byte. The letter é (U+00E9) encodes as the two UTF-8 bytes 0xC3 0xA9, producing %C3%A9. The euro sign € (U+20AC) encodes as 0xE2 0x82 0xAC, producing %E2%82%AC. All modern URL functions handle this automatically: you pass a Unicode string and the function handles UTF-8 encoding and percent-encoding internally.
encodeURIComponent vs encodeURI
JavaScript exposes two built-in functions for URL encoding, and choosing the wrong one is a common source of bugs because each function leaves a different set of characters untouched. Picking the wrong function can corrupt a query string or leave unsafe characters in a URL, which is why understanding the difference between them matters before you encode anything.
When to use encodeURI
encodeURI is designed for complete URLs. It preserves all characters that have structural
meaning in a URL: the protocol colon and slashes, question marks, hash signs, ampersands, equals signs,
and domain punctuation. Passing a full URL like
https://example.com/search?q=hello world through encodeURI safely encodes the
space as %20 while leaving the protocol, domain, path separators, and query string
structure intact.3 Because encodeURI leaves the equal sign and
ampersand untouched, it is the wrong choice for encoding a single value that will later be inserted into
a query string.
encodeURIComponent is designed for individual URL components such as a query parameter
value, a path segment, or a hash fragment. It encodes everything except letters, digits, and the four
characters -_.~. This makes it safe to embed as a value inside a query string, since it
will encode the & and = characters that would otherwise break the
query string structure. Passing hello world through encodeURIComponent gives
hello%20world, which is safe to append as ?q=hello%20world.4
The practical rule: use encodeURIComponent when building URLs from parts by encoding each
value. Use encodeURI only when you already have a complete URL and need to make it safe
for an HTML attribute like href. Never use encodeURI on user-supplied query
parameter values or you will produce a URL that cannot be reliably decoded.
Percent-encoding beyond JavaScript
The standard is language-neutral. RFC 3986 defines which bytes need escaping and the
%XX form every escape takes, so every mainstream language and
HTTP client ships a percent-encoder and only the function names differ. When any of them escapes a
character, the byte sequence it emits matches what the rows in this tool's encode panel produce for
the same input. Differences between languages show up in which characters stay literal, not in what an
escaped byte becomes. You can paste the same value here and read the three rows as the reference
output.
The boundary is worth stating plainly. This tool speaks the two JavaScript functions shown above plus
the form-encoding row. Other languages' full-string encoders may keep a slightly different set of
characters literal, the way encodeURI keeps & and =, so when you
port an encoded value between languages, compare outputs here rather than assuming the functions match
one-to-one. A query string that survived one stack can arrive with broken separators in another for
exactly this reason. The bytes never lie; the literal sets do. Encoding the value here after the
transfer shows whether the structure survived intact.
Form encoding and the + character
The application/x-www-form-urlencoded format is the default encoding for HTML form
submissions. It is nearly identical to standard percent-encoding but uses a plus sign (+)
rather than %20 to represent spaces. This convention traces back to early web forms and
remains common today in GET request query strings generated by HTML forms.5
When a browser submits a form, the request body or appended query string uses this format. Many
server-side frameworks and web APIs decode + back to a space automatically when parsing
form data. However, calling decodeURIComponent directly on a form-encoded string will not
convert plus signs to spaces. If you see unexpected + characters after decoding, the
input was form-encoded. Replace each + with %20 before passing it to
decodeURIComponent, or use a dedicated form parser.
TIP The encode panel on this builder shows the form-encoded variant (+ for spaces) alongside the two standard JavaScript encoding functions so you can compare all three outputs at once and see exactly how a single input string differs depending on which encoding rules are applied. Comparing them side by side is the fastest way to spot whether a plus sign in a decoded value is a real space or a leftover form-encoding artifact.
Choosing the right decoder for the format you have
Before decoding, identify which format the source string uses. A query string pulled from an HTML form
submission likely uses the plus-for-space convention, so it needs each + changed to
%20 before decodeURIComponent returns the intended spaces.5 A path segment or a
value encoded by JavaScript code is more likely to be standard percent-encoding. Decoding here
follows decodeURIComponent exactly, so a plus sign survives as a plus sign no matter which
convention produced the string. When you cannot tell which convention a string uses, the ENCODE
panel's form row is the comparison point: re-encoding the value shows exactly which characters each
convention would produce.
Parsing and editing query strings
Query strings follow a consistent structure: key-value pairs separated by &, with
keys and values joined by =. The query string begins after the ? in a URL.
A URL like https://example.com/search?q=cats&page=2&sort=date contains three
parameters: q, page, and sort.2 The
same structure appears in OAuth redirects, analytics tracking links, and API endpoints, so being able to
read and modify query strings by hand is a useful debugging skill.
Editing values without breaking the structure
When you paste a URL with a query string in decode mode, this tool parses the parameters and shows them in an editable table. Each row displays the decoded key and an editable field for the decoded value. Changing a value immediately rebuilds the full encoded query string shown below the table. This makes it straightforward to debug API calls, construct test requests, or understand what a URL is requesting without manually working through percent-encoded characters.
The rebuilt query string uses encodeURIComponent on every key and value, so the output
is safe to append directly to any URL. The original input field is not overwritten when you edit the
table, which prevents re-parse loops and lets you compare the original and modified strings side by
side.
URL encoding is not HTML encoding
The two systems solve different parsers' problems. HTML escaping turns & into
& and < into < so a value can sit inside markup without
the markup parser reading it as tags; percent-encoding turns & into %26 so a value
can sit inside a URL without the URL parser reading it as a separator. They protect different layers,
and a value that travels through both layers needs each escape applied at its own layer. Neither
escape substitutes for the other. Confusing the two produces a recognizable symptom, and the next
paragraph names it.
The classic symptom is an & reaching a server inside a URL. That means HTML escaping
leaked into a URL context, usually a template that escaped a value before building the query string.
The correct order runs the other way: percent-encode the value when the URL is assembled, then
HTML-escape the finished attribute when it lands in markup.6
Decoding a suspect string here helps you confirm which layer went wrong, because percent-decoding an
already-HTML-escaped value leaves the entity intact and visible. The entity text surviving a
percent-decode is the fingerprint of the markup layer.
- 1.
WHATWG, "URL Standard," url.spec.whatwg.org, June 2026. https://url.spec.whatwg.org/
- 2.
T. Berners-Lee, R. Fielding, and L. Masinter, "Uniform Resource Identifier (URI): Generic Syntax," RFC 3986, IETF, January 2005. https://www.rfc-editor.org/rfc/rfc3986
- 3.
Mozilla Developer Network, "encodeURI()," developer.mozilla.org, July 2025. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/encodeURI
- 4.
Mozilla Developer Network, "encodeURIComponent()," developer.mozilla.org, October 2025. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/encodeURIComponent
- 5.
WHATWG, "HTML Standard," html.spec.whatwg.org, August 2026. https://html.spec.whatwg.org/multipage/form-control-infrastructure.html#application/x-www-form-urlencoded-encoding-algorithm
- 6.
OWASP Foundation, "Cross Site Scripting Prevention Cheat Sheet," cheatsheetseries.owasp.org, accessed September 2026. https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html