Base64 explained: what it is, what it is not, and the URL-safe variant
Base64 shows up everywhere: email attachments, data URIs, JWTs, Kubernetes Secrets, Basic authentication headers and binary fields in JSON. It is also one of the most misunderstood encodings. People mistake it for encryption, use it where it only adds overhead, and lose hours to the difference between two nearly identical alphabets. This guide explains how Base64 works, what it is not, and how to handle it correctly in code.
The problem it solves
Many channels were built to carry text, not arbitrary bytes. Email historically assumed 7-bit ASCII. JSON strings must be valid Unicode. HTTP headers allow a limited character set, and environment variables cannot contain null bytes. To push an image, a key or a compressed blob through one of these, you need to represent any byte sequence using only characters that survive the trip.
Base64 uses 64 such characters, A-Z, a-z, 0-9, + and /, plus = for padding. That is the whole idea: a reversible mapping from bytes to boring text.
How the encoding works
Base64 takes three bytes at a time. Three bytes are 24 bits, which split exactly into four 6-bit groups. Each group is a number from 0 to 63, which picks one character from the alphabet. The text Hi! is the bytes 72, 105 and 33:
01001000 01101001 00100001 the 3 bytes
010010 000110 100100 100001 regrouped into 6-bit values
18 6 36 33 as numbers
S G k h looked up in the alphabet
So Hi! encodes to SGkh. When the input is not a multiple of three bytes, the last group is padded: Hi becomes SGk= and H becomes SA==. The padding tells the decoder how many real bytes the final group holds.
Why the output is a third larger
Every 3 bytes become 4 characters, so output is about 33% larger than input. A 3 MB file becomes a 4 MB string, and a 50 MB upload sent as Base64 inside JSON becomes roughly 67 MB that both sides must hold in memory and push through a JSON parser. MIME email also inserts a line break every 76 characters, adding a little more.
What Base64 is not
It is not encryption
There is no key. Anyone can decode Base64 instantly, and plenty of people can do it by eye for short strings. The header Authorization: Basic dXNlcjpwYXNz is simply user:pass, which is why Basic auth is only acceptable over HTTPS. Kubernetes Secrets are Base64-encoded so they can hold binary data, not to protect it. Anyone who can read the manifest, or run kubectl get secret -o yaml, can read the secret. If something has to stay confidential, encrypt it, and treat the Base64 form as exactly as sensitive as the original.
It is not compression
It always makes data bigger. If size matters, compress first (gzip, Brotli), then encode the result if you need text.
It is not a hash or a checksum
It does not detect tampering, and two encodings of the same bytes can differ in whitespace or padding. If you need integrity, sign or hash the bytes; see the hashing guide.
Text must become bytes first
Base64 encodes bytes, not characters, so text has to be converted with a character encoding first, and the choice changes the result:
Buffer.from('café', 'utf8').toString('base64') // 'Y2Fmw6k='
Buffer.from('café', 'latin1').toString('base64') // 'Y2Fm6Q=='
Use UTF-8 unless a protocol says otherwise. This is where the browser's btoa() trips people up. It accepts only characters in the Latin-1 range, so btoa('café ☕') throws InvalidCharacterError. Worse, btoa('café') succeeds and returns the Latin-1 encoding Y2Fm6Q==, which a UTF-8 decoder on the server turns into garbage. The correct browser code goes through TextEncoder:
function toBase64(text) {
const bytes = new TextEncoder().encode(text); // UTF-8
let bin = '';
for (const b of bytes) bin += String.fromCharCode(b);
return btoa(bin);
}
function fromBase64(b64) {
const bin = atob(b64);
return new TextDecoder().decode(Uint8Array.from(bin, (c) => c.charCodeAt(0)));
}
toBase64('café ☕') // 'Y2Fmw6kg4piV'
Newer runtimes, including browsers released since late 2025 and current Node.js, have this built in: new TextEncoder().encode(text).toBase64() and Uint8Array.fromBase64(str). In Node.js, Buffer.from(text).toString('base64') has always defaulted to UTF-8. The Base64 Encoder uses UTF-8 too, so its output matches what a server expects.
Base64URL, the URL-safe variant
The standard alphabet's + and / mean something in URLs, and = separates query parameters. Base64URL (RFC 4648, section 5) swaps + for - and / for _, and usually drops the padding. JWTs, OAuth PKCE challenges, WebAuthn and most "random token in a URL" schemes use it.
The same three bytes in each variant:
Buffer.from([251, 255, 191]).toString('base64') // '+/+/'
Buffer.from([251, 255, 191]).toString('base64url') // '-_-_'
Converting by hand when your platform only offers one variant:
const toBase64Url = (b64) => b64.replace(/\+/g, '-').replace(/\//g, '_').replace(/=+$/, '');
const fromBase64Url = (s) =>
s.replace(/-/g, '+').replace(/_/g, '/') + '='.repeat((4 - (s.length % 4)) % 4);
Mixing the variants causes one of the more annoying bugs to track down, because it only fails on some inputs: those whose encoding happens to contain +, /, - or _. Each output character has only a 1 in 32 chance of being one of the two that differ, and plain ASCII text rarely produces them, so a test with a few short strings passes while production fails now and then. Agree on the variant explicitly on both sides of an integration, and test with random binary input, not just text.
Decoders disagree about bad input
Not all decoders are equally strict. Node's Buffer.from(s, 'base64') silently skips characters outside the alphabet, so 'SGkh!!' decodes to Hi! without complaint. The browser's atob throws on the same input. Java's Base64.getDecoder() rejects line breaks, while getMimeDecoder() accepts them. Lenient decoding hides corruption. Strict decoding rejects data that another system considers valid. When a value round-trips in one language and fails in another, this is usually why.
Good uses and bad uses
Base64 is the right tool for small binary values inside text formats: a public key or signature in JSON, a certificate in an environment variable, an email attachment, or a random token encoded with Base64URL. It also works for tiny data URIs, such as an icon under a few kilobytes inlined into CSS to save a request.
It is the wrong tool for file uploads (use multipart or a pre-signed object storage URL), large inline images (they bloat HTML and CSS, block rendering and cannot be cached separately), and anything you are trying to hide.
Debugging a value that will not decode
- Variant. Does it contain
-or_? It is Base64URL. - Padding. Is the length a multiple of four? If not, padding was stripped; add
=until it is. - Whitespace. Line breaks from MIME wrapping or a terminal copy break strict decoders.
- Character set. Accents decode as
é? The text was encoded with one charset and decoded with another. - Double encoding. If the decoded result looks like Base64 again, something encoded it twice. JWT payloads inside query strings are a favourite place for this.
Paste the value into the Base64 Encoder to test each of these quickly, or into the JWT Decoder if it has two dots in it.