Guide

URL encoding explained

A URL may contain only a restricted set of ASCII characters, and some of those characters have structural meaning, `?` starts a query, `&` separates parameters, `#` starts a fragment. Percent-encoding is how a value that contains one of them travels inside a URL without being mistaken for structure.

The mechanism is simple: a byte becomes a `%` followed by its two hexadecimal digits, so a space becomes `%20`. The difficulty is knowing which characters need it, because the answer depends on where in the URL the value sits.

Open the URL Encoder

Use this when


  • Building a query string from values that may contain spaces or symbols
  • Putting a URL inside another URL, such as a redirect parameter
  • Debugging a link that works for some inputs and breaks for others
  • Handling non-ASCII text (names, search terms) in a URL
  • Reading a log line or an API request that is full of percent sequences

Reserved, unreserved, and why position matters


The unreserved characters (letters, digits, hyphen, period, underscore and tilde) never need encoding anywhere. The reserved characters do have structural meaning, and whether they need encoding depends on whether you mean that meaning.

A slash in a path is a path separator; a slash inside a query parameter value is just a character and is usually left alone. An ampersand between parameters is a separator; an ampersand inside a value must be `%26` or everything after it becomes a separate parameter. The question is never "does this character need encoding" but "is this character data or structure here".

This is why encoding a whole URL at once is almost always wrong. It escapes the slashes and the question mark that make the URL a URL, producing a string that is no longer a link. Encode the values you are inserting, not the URL you are inserting them into.

The space problem


A space can be `%20` or `+`, and which one is correct depends on where it appears. In a path, only `%20` is valid. In a query string, both are accepted by convention, because HTML form submissions historically encoded spaces as `+` under the `application/x-www-form-urlencoded` rules.

`%20` is the safer choice everywhere, and the one to use if you are generating URLs. It is unambiguous in every position, whereas `+` in a path is a literal plus sign and will not be read as a space.

When decoding, the direction of the ambiguity reverses and it becomes a real trap: a literal plus sign in a query value must be sent as `%2B`, because a bare `+` will be decoded as a space. This is exactly how Base64 strings and phone numbers get corrupted in query parameters, the `+` arrives, is read as a space, and the value is wrong with no error anywhere.

Non-ASCII characters


Percent-encoding escapes bytes, not characters, so non-ASCII text has to be encoded to bytes first. UTF-8 is the modern standard and what every current browser and server expects.

One character can therefore become several percent sequences: `é` is two bytes in UTF-8 and encodes as `%C3%A9`, while an emoji is four bytes and becomes eight hex digits. Seeing a long run of percent sequences for a short piece of text is normal, not a symptom.

Domain names work differently. Internationalised domains use Punycode rather than percent-encoding, so a Cyrillic or Chinese hostname becomes an ASCII string beginning `xn--`. Percent-encoding a hostname does not work, which surprises people who try it.

Double encoding


Encoding an already-encoded string escapes its percent signs: `%20` becomes `%2520`, because `%` itself encodes to `%25`. Decoding once then yields `%20` rather than a space, so the value is wrong in a way that looks almost right.

This is the most common URL bug in real systems, and it comes from layers each doing their job without knowing the others did too, a framework helpfully encoding a parameter that was already encoded, or a redirect URL encoded once when constructed and again when appended.

The diagnostic is easy once you know it: `%25` in a URL where you did not intend a literal percent sign means double encoding. The fix is to find which layer is encoding redundantly, not to decode twice, decoding twice happens to produce the right answer here and will produce a wrong one, or a security hole, on input that contains a genuine `%25`.

Choosing the right function


JavaScript offers two encoders and they are not interchangeable. `encodeURIComponent` escapes everything reserved, which is what you want for a single value being placed into a query parameter or path segment. `encodeURI` deliberately leaves the structural characters alone, which is only appropriate for an entire URL that is already correctly assembled.

Using `encodeURI` on a value is the usual mistake: it leaves `&`, `?`, `/` and `=` unescaped, so a value containing one of them silently corrupts the URL structure around it. A search term containing an ampersand becomes two query parameters.

Neither function escapes `!`, `'`, `(`, `)` or `*`, which are legal in a URL but occasionally awkward in other contexts such as OAuth signatures. When building URLs from scratch in JavaScript, `URLSearchParams` is usually the better instrument than either function, because it handles the joining and escaping together and cannot be applied to the wrong part of the URL.

Common problems


Each of these is something that actually happens, with the cause rather than a generic suggestion to check your input.

A plus sign in a value arrives as a space

Cause
In a query string a bare `+` decodes to a space, by form-encoding convention.
Fix
Encode a literal plus as `%2B`. This is what corrupts phone numbers and Base64 in query parameters.

The URL contains %2520 and the value is wrong

Cause
Double encoding: an already-encoded string was encoded again, so `%` became `%25`.
Fix
Find the layer encoding redundantly and remove it. Do not decode twice as a workaround. That breaks on values containing a real `%25`.

Everything after part of a value became separate parameters

Cause
The value contained `&` or `=` unescaped, probably from `encodeURI` rather than `encodeURIComponent`.
Fix
Encode each value with `encodeURIComponent`, or build the query with `URLSearchParams`.

The whole URL was encoded and is no longer a link

Cause
The full URL went through a component encoder, escaping the `://` and the slashes that give it structure.
Fix
Encode only the values being inserted. A URL is assembled from encoded parts; it is not encoded as a whole.

A path with a space does not resolve

Cause
The space was encoded as `+`, which is only a space by convention in a query string.
Fix
Use `%20`. It is correct in every position, which is why it is the better default everywhere.

Questions


Should a space be %20 or +?
`%20`. It is valid everywhere. A `+` means a space only in a query string, by form-encoding convention, and is a literal plus sign in a path.
What is the difference between encodeURI and encodeURIComponent?
`encodeURIComponent` escapes the reserved characters and is what you want for a single value. `encodeURI` leaves `&`, `?`, `/` and `=` alone because it is meant for a complete URL. Using `encodeURI` on a value is how a URL silently gains extra parameters.
Why is one character several percent sequences?
Because percent-encoding escapes bytes, and non-ASCII characters are several bytes in UTF-8. `é` is two bytes and becomes `%C3%A9`; an emoji is four and becomes eight hex digits.
How do I tell if a value is double-encoded?
Look for `%25` where no literal percent sign was intended, or `%2520` where you expected a space. Both mean an encoded string was encoded a second time.