Skip to content

URL anatomy and percent-encoding done right

By · URL Tools · 5 min read · Published

URLs are so familiar that developers rarely think about their structure until something breaks: an OAuth redirect that does not match, a search query that loses everything after an ampersand, a file name with a space that returns 404, or a value that arrives as %2520. This guide explains every part of a URL, the encoding rules for each, and the bugs that follow from getting them wrong.

The parts of a URL

https://user:[email protected]:8443/v1/users/42?fields=name,email&page=2#profile
\___/   \_______/ \_____________/ \__/\__________/ \_____________________/ \_____/
scheme  userinfo       host       port    path              query          fragment
  • Scheme: the protocol, such as https, http, ftp or mailto. It is case-insensitive and normally written in lowercase.
  • Userinfo: an optional username and password. Browsers increasingly ignore or block it because it is used in phishing links like https://[email protected]/, where the real host is evil.example.
  • Host: a domain name or IP address. Domain names are case-insensitive. Internationalised names such as bücher.example are converted to an ASCII form called Punycode, xn--bcher-kva.example.
  • Port: optional. When omitted, the default for the scheme is used: 443 for HTTPS, 80 for HTTP.
  • Path: segments separated by /. Unlike the host, the path is case-sensitive on most servers.
  • Query: everything after ?, conventionally key=value pairs joined by &.
  • Fragment: everything after #. Browsers never send it to the server.

The scheme, host and port together form the origin, which browsers use for security boundaries such as CORS, cookies and storage. https://example.com and https://api.example.com are different origins.

Paste any URL into the URL Parser to see it split into these components, with every query parameter decoded.

Percent-encoding

URLs may only contain a limited set of ASCII characters. Anything else, and any character that would otherwise have a special meaning, is percent-encoded: each byte of its UTF-8 representation is written as % followed by two hexadecimal digits.

space   →  %20
é       →  %C3%A9   (two UTF-8 bytes)
&       →  %26
/       →  %2F
😀      →  %F0%9F%98%80

Characters that never need encoding are called unreserved: letters, digits, -, ., _ and ~. Reserved characters such as : / ? # [ ] @ ! $ & ' ( ) * + , ; = delimit parts of the URL, so they must be encoded when they appear inside a value rather than as a delimiter.

Different parts, different rules

Which characters must be encoded depends on where they appear:

  • In a path segment, / must be encoded as %2F if it is part of a name, otherwise it splits the segment. ? and # must be encoded because they would end the path.
  • In a query value, &, =, # and + must be encoded. An unencoded & starts a new parameter, so q=salt&pepper sends q=salt and a second, empty parameter called pepper.
  • In a fragment, most characters are allowed, but encoding is still safest for anything unusual.

The plus sign problem

HTML forms submitted with application/x-www-form-urlencoded encode spaces as + rather than %20. Many server frameworks therefore decode + in query strings as a space. That leads to two classic bugs:

  1. A literal plus sign that was not encoded, like a phone number +4930123 or an email address [email protected], arrives as a space.
  2. Base64 values contain +, so they are corrupted in query strings unless encoded. That is one reason Base64URL exists.

The fix is to always encode values with a proper function: a literal plus becomes %2B.

Use the right encoding function

In JavaScript, there are several functions, and the difference matters:

  • encodeURIComponent() encodes everything except unreserved characters. Use it for individual path segments and query values.
  • encodeURI() leaves reserved characters alone, because it assumes you pass a complete URL. Using it on a single value leaves &, = and ? unencoded.
  • URLSearchParams builds query strings correctly for you, and uses + for spaces, which servers decode correctly.
  • new URL() parses and normalises URLs using the WHATWG URL standard, the same rules browsers apply.
const url = new URL('https://api.example.com/search');
url.searchParams.set('q', 'salt & pepper');
url.searchParams.set('tag', 'c++');
url.toString();
// https://api.example.com/search?q=salt+%26+pepper&tag=c%2B%2B

Other languages have equivalents, such as Python's urllib.parse.urlencode and Java's URLEncoder combined with URI. The rule is the same everywhere: never build URLs by concatenating unencoded strings.

Double encoding and double decoding

If a value is encoded twice, the % itself gets encoded as %25, so a space becomes %2520. This typically happens when one layer encodes a value and a second layer, such as an HTTP client or a redirect, encodes the whole URL again. The opposite also happens: decoding twice turns a legitimately encoded %2520 into a space. Encode exactly once, when building the URL, and decode exactly once, when reading a parameter. Double decoding has also been the root of security bugs, where a check runs on the once-decoded value and the application uses the twice-decoded one.

The fragment is client-only

Because browsers do not send the fragment, the server cannot see it, log it or redirect based on it. Single-page applications use fragments for client-side routes, and some OAuth flows return tokens in the fragment precisely so they do not appear in server logs. If a server-side redirect needs to preserve state, put it in the query string instead.

Relative URLs

A relative reference is resolved against a base URL. With a base of https://example.com/docs/guide/:

  • intro resolves to https://example.com/docs/guide/intro.
  • ../api resolves to https://example.com/docs/api.
  • /pricing resolves to https://example.com/pricing.
  • //cdn.example.com/x.js keeps the scheme but changes the host.

Note how the trailing slash on the base matters: without it, intro would replace guide rather than go inside it.

Checklist for URLs in code

  1. Build URLs with a URL API, not string concatenation.
  2. Encode each value once, with the function meant for that part of the URL.
  3. Keep tokens, passwords and personal data out of query strings; they end up in logs, browser history and Referer headers.
  4. Compare OAuth redirect URIs exactly, including scheme, port, path and trailing slash.
  5. Choose one trailing-slash convention for your site and redirect the other form to it.