URL Encoder: How It Works
URL encoding replaces characters that have structural meaning in a web address with a percent sign and their hexadecimal byte value. Knowing which of the two encoding functions to use, and where, prevents a whole class of bugs that only appear with unusual input.
Reserved characters
These have a job in a URL, so a literal one must be escaped when it appears in data.
| Char | Encoded | Its structural role |
|---|---|---|
| space | %20 (or + in query strings) | Terminates a URL in many contexts |
? | %3F | Starts the query string |
& | %26 | Separates parameters |
= | %3D | Separates key from value |
# | %23 | Starts the fragment |
/ | %2F | Path separator |
+ | %2B | Means space in form encoding |
The two functions
JavaScript offers encodeURI and encodeURIComponent, and choosing wrongly is the standard bug.
encodeURIencodes a whole URL, leaving structural characters such as/ ? & = #intact so the address still works.encodeURIComponentencodes a single value, escaping those characters too because inside a parameter they are data rather than structure.
Encoding a parameter with encodeURI leaves any & in the value unescaped, which silently splits it into two parameters. A search for 'fish & chips' becomes a search for 'fish ' plus an unexpected parameter named 'chips'. Use encodeURIComponent for values, always.
The plus-sign ambiguity
Two conventions coexist. In application/x-www-form-urlencoded data — HTML form submissions — a space is encoded as +. In URL paths and in the modern percent-encoding standard, a space is %20 and + is a literal plus.
The practical consequence: an email address containing a plus, such as user+tag@example.com, arrives as user tag@example.com if the receiving code treats the query string as form-encoded. Always encode a literal plus as %2B.
Double encoding
Encoding an already-encoded string escapes the percent signs themselves: %20 becomes %2520. The result decodes once to %20 and only twice to a space. This appears whenever a URL passes through two layers that each helpfully encode it — a redirect service, a tracking wrapper, a framework's router. If you see %25 sequences in production logs, something is encoding twice.
Unicode
Percent-encoding operates on bytes, so text must first be encoded to bytes — in practice UTF-8. A single non-ASCII character often becomes several percent sequences: é is %C3%A9, and an emoji can be four. Truncating an encoded URL by character count can therefore split a multi-byte sequence and produce an undecodable string.
Where to encode
- Query parameter values, always with the component function.
- Path segments containing user data — note that
/must be escaped there. - Fragment identifiers.
- Not the whole URL after assembling it, which double-encodes what you already escaped.
Encoding is not sanitisation. It makes a value safe to carry in a URL; it does not make it safe to insert into HTML, SQL or a shell command, each of which needs its own escaping at its own boundary.