URL anatomy and percent-encoding done right
URLs are so familiar that developers rarely think about their structure until something breaks: an OAuth redirect that does not match, a search query that loses everything after an ampersand, a file name with a space that returns 404, or a value that arrives as %2520. This guide explains every part of a URL, the encoding rules for each, and the bugs that follow from getting them wrong.
The parts of a URL
https://user:[email protected]:8443/v1/users/42?fields=name,email&page=2#profile
\___/ \_______/ \_____________/ \__/\__________/ \_____________________/ \_____/
scheme userinfo host port path query fragment
- Scheme: the protocol, such as
https,http,ftpormailto. It is case-insensitive and normally written in lowercase. - Userinfo: an optional username and password. Browsers increasingly ignore or block it because it is used in phishing links like
https://[email protected]/, where the real host isevil.example. - Host: a domain name or IP address. Domain names are case-insensitive. Internationalised names such as
bücher.exampleare converted to an ASCII form called Punycode,xn--bcher-kva.example. - Port: optional. When omitted, the default for the scheme is used: 443 for HTTPS, 80 for HTTP.
- Path: segments separated by
/. Unlike the host, the path is case-sensitive on most servers. - Query: everything after
?, conventionallykey=valuepairs joined by&. - Fragment: everything after
#. Browsers never send it to the server.
The scheme, host and port together form the origin, which browsers use for security boundaries such as CORS, cookies and storage. https://example.com and https://api.example.com are different origins.
Paste any URL into the URL Parser to see it split into these components, with every query parameter decoded.
Percent-encoding
URLs may only contain a limited set of ASCII characters. Anything else, and any character that would otherwise have a special meaning, is percent-encoded: each byte of its UTF-8 representation is written as % followed by two hexadecimal digits.
space → %20
é → %C3%A9 (two UTF-8 bytes)
& → %26
/ → %2F
😀 → %F0%9F%98%80
Characters that never need encoding are called unreserved: letters, digits, -, ., _ and ~. Reserved characters such as : / ? # [ ] @ ! $ & ' ( ) * + , ; = delimit parts of the URL, so they must be encoded when they appear inside a value rather than as a delimiter.
Different parts, different rules
Which characters must be encoded depends on where they appear:
- In a path segment,
/must be encoded as%2Fif it is part of a name, otherwise it splits the segment.?and#must be encoded because they would end the path. - In a query value,
&,=,#and+must be encoded. An unencoded&starts a new parameter, soq=salt&peppersendsq=saltand a second, empty parameter calledpepper. - In a fragment, most characters are allowed, but encoding is still safest for anything unusual.
The plus sign problem
HTML forms submitted with application/x-www-form-urlencoded encode spaces as + rather than %20. Many server frameworks therefore decode + in query strings as a space. That leads to two classic bugs:
- A literal plus sign that was not encoded, like a phone number
+4930123or an email address[email protected], arrives as a space. - Base64 values contain
+, so they are corrupted in query strings unless encoded. That is one reason Base64URL exists.
The fix is to always encode values with a proper function: a literal plus becomes %2B.
Use the right encoding function
In JavaScript, there are several functions, and the difference matters:
encodeURIComponent()encodes everything except unreserved characters. Use it for individual path segments and query values.encodeURI()leaves reserved characters alone, because it assumes you pass a complete URL. Using it on a single value leaves&,=and?unencoded.URLSearchParamsbuilds query strings correctly for you, and uses+for spaces, which servers decode correctly.new URL()parses and normalises URLs using the WHATWG URL standard, the same rules browsers apply.
const url = new URL('https://api.example.com/search');
url.searchParams.set('q', 'salt & pepper');
url.searchParams.set('tag', 'c++');
url.toString();
// https://api.example.com/search?q=salt+%26+pepper&tag=c%2B%2B
Other languages have equivalents, such as Python's urllib.parse.urlencode and Java's URLEncoder combined with URI. The rule is the same everywhere: never build URLs by concatenating unencoded strings.
Double encoding and double decoding
If a value is encoded twice, the % itself gets encoded as %25, so a space becomes %2520. This typically happens when one layer encodes a value and a second layer, such as an HTTP client or a redirect, encodes the whole URL again. The opposite also happens: decoding twice turns a legitimately encoded %2520 into a space. Encode exactly once, when building the URL, and decode exactly once, when reading a parameter. Double decoding has also been the root of security bugs, where a check runs on the once-decoded value and the application uses the twice-decoded one.
The fragment is client-only
Because browsers do not send the fragment, the server cannot see it, log it or redirect based on it. Single-page applications use fragments for client-side routes, and some OAuth flows return tokens in the fragment precisely so they do not appear in server logs. If a server-side redirect needs to preserve state, put it in the query string instead.
Relative URLs
A relative reference is resolved against a base URL. With a base of https://example.com/docs/guide/:
introresolves tohttps://example.com/docs/guide/intro.../apiresolves tohttps://example.com/docs/api./pricingresolves tohttps://example.com/pricing.//cdn.example.com/x.jskeeps the scheme but changes the host.
Note how the trailing slash on the base matters: without it, intro would replace guide rather than go inside it.
Checklist for URLs in code
- Build URLs with a URL API, not string concatenation.
- Encode each value once, with the function meant for that part of the URL.
- Keep tokens, passwords and personal data out of query strings; they end up in logs, browser history and Referer headers.
- Compare OAuth redirect URIs exactly, including scheme, port, path and trailing slash.
- Choose one trailing-slash convention for your site and redirect the other form to it.