Guide

Base64 encoding explained

Base64 turns arbitrary bytes into text using 64 characters that survive almost any transport: A–Z, a–z, 0–9, plus and slash. It exists because a great deal of infrastructure was built for text and mangles raw bytes, email headers, JSON strings, URLs, HTTP headers, XML documents.

It is worth being precise about what it is for, because Base64 is one of the most consistently misunderstood tools in this list. It provides no security whatsoever. Anything Base64-encoded can be read back by anyone, instantly, with no key and no effort.

Open the Base64 Encoder

Use this when


  • Embedding a small image directly in CSS or HTML as a data URI
  • Putting binary data inside a JSON field, which can only hold text
  • Sending credentials in an HTTP Basic authentication header
  • Reading the contents of a JWT, whose segments are Base64url
  • Attaching a file to an email, where MIME requires it
  • Passing a binary value through a system that only accepts ASCII

How it works


Base64 processes three bytes at a time. Three bytes are 24 bits, which divide evenly into four groups of six, and each six-bit group indexes into the 64-character alphabet. Three input bytes therefore always become four output characters, which is why encoded data is about a third larger than what went in, 33% overhead, exactly.

When the input length is not a multiple of three, the final group is short. Base64 pads it with zero bits to complete the last character and then appends `=` signs to record how much padding was added: one `=` means the input had one leftover byte pair, two means one leftover byte. That is the entire meaning of the trailing equals signs.

Because the mapping is fixed and public, decoding needs nothing but the alphabet. There is no key, no secret and no variation between implementations of the standard alphabet, which is what makes it reliable, and also what makes it useless for concealment.

It is not encryption, and this matters


Base64 is a representation change, like writing a number in hexadecimal. Encoding a password does not protect it in any sense: the original is recoverable by anyone who sees the string, and every browser has a one-line function that does it.

This has real consequences. HTTP Basic authentication sends `username:password` as Base64, which is why Basic auth over plain HTTP transmits credentials in effectively plain text, the Base64 wrapper stops nothing. The security in that scheme comes entirely from TLS.

The same applies to the signed compact JWS form commonly used for JWTs. Its header and payload are Base64url-encoded, not encrypted, so anyone holding it can read the claims. A signature provides integrity only after verification; it does not hide the contents. Encrypted JWE tokens are different: their payload is not readable just by reversing Base64url, and this site’s JWT decoder does not decrypt them.

Standard versus URL-safe


The standard alphabet uses `+` and `/` for its final two characters, and both are a problem in a URL. A plus sign in a query string means a space, and a slash is a path separator, so a standard Base64 string dropped into a URL either breaks or silently changes meaning.

The URL-safe variant, defined in the same RFC, substitutes `-` for `+` and `_` for `/`. Nothing else changes, and the two are trivially interconvertible. URL-safe Base64 also usually omits the `=` padding, since the length can be inferred and `=` needs escaping in a URL as well.

JWTs use Base64url throughout. If you are decoding a token segment by hand with a standard Base64 decoder, converting the two characters back and re-adding padding to a multiple of four is exactly what you need to do, and it is why a JWT segment often fails to decode in a generic tool.

Data URIs, and their cost


A data URI embeds a file directly in a document: `data:image/png;base64,` followed by the encoded bytes. The benefit is one fewer HTTP request, which for a tiny icon in a stylesheet can genuinely be worth having.

The costs are easy to underestimate. The data is 33% larger than the file, it cannot be cached separately from the document that contains it, and it is re-downloaded on every change to that document. A large image inlined into CSS makes the stylesheet render-blocking and delays the entire page.

The rough working rule is that inlining pays below a few kilobytes and stops paying quickly above it. An SVG icon, yes. A photograph, no, and an SVG can usually be inlined as markup rather than Base64, which avoids the 33% penalty entirely.

Unicode, and the classic JavaScript failure


Base64 encodes bytes, not characters. Text has to be turned into bytes first, and that requires choosing an encoding, UTF-8 in practice. Encoding the same string as UTF-16 produces entirely different Base64, so both ends must agree.

In the browser, this is the reason `btoa()` throws on non-Latin-1 text. `btoa` expects each character to be a single byte, so anything above U+00FF (an emoji, Cyrillic, Chinese, or a curly quote) raises an InvalidCharacterError. The fix is to encode to UTF-8 bytes first with `TextEncoder`, then Base64 those bytes.

The subtler version of the bug is a string that encodes without error and decodes to mojibake, which happens when one side used UTF-8 and the other Latin-1. If decoded text looks like the right characters wearing the wrong ones, an encoding mismatch rather than a Base64 problem is the cause.

Common problems


Each of these is something that actually happens, with the cause rather than a generic suggestion to check your input.

Decoding fails with an invalid-character or length error

Cause
Usually URL-safe input in a standard decoder: the `-` and `_` characters are not in the standard alphabet, and the padding has often been stripped.
Fix
Convert `-` to `+` and `_` to `/`, then pad with `=` until the length is a multiple of four.

A Base64 string breaks when put in a URL

Cause
Standard Base64 contains `+` and `/`. In a query string `+` means a space.
Fix
Use the URL-safe variant, or percent-encode the whole value. Do not do both, double-encoding is its own bug.

btoa() throws on text with emoji or non-Latin characters

Cause
`btoa` handles only characters up to U+00FF, one byte each.
Fix
Encode to UTF-8 bytes with `TextEncoder` first, then Base64 the byte array.

Decoded text is garbled but roughly recognisable

Cause
An encoding mismatch, one side used UTF-8, the other Latin-1.
Fix
Standardise on UTF-8 at both ends. The Base64 itself is fine; the bytes are being interpreted with the wrong encoding.

A Base64 string from an email or config file will not decode

Cause
Line breaks. MIME wraps Base64 at 76 characters, and those newlines are not part of the data.
Fix
Strip all whitespace before decoding. A tolerant decoder does this automatically.

Questions


Is Base64 encryption?
No. It is a reversible representation with no key, and anyone can decode it instantly. It protects nothing, never use it to conceal a password, a token or any other secret.
Why does Base64 end in one or two equals signs?
Padding. Base64 works on three-byte groups, and when the input length is not a multiple of three the final group is incomplete. The equals signs record how many bytes were missing, so a decoder knows how many to discard.
What is the difference between Base64 and Base64url?
Two characters. Base64url uses `-` and `_` where standard Base64 uses `+` and `/`, because those two are unsafe in URLs, and it usually drops the `=` padding. JWTs use Base64url.
How much larger does Base64 make my data?
Exactly one third, plus up to two padding characters: every three bytes become four characters. A 3 KB image becomes 4 KB of text. This is worth remembering before inlining anything large as a data URI.
Is my input sent to a server?
No. Encoding and decoding both run in your browser. That matters more here than for most tools, since people do paste tokens and credentials into Base64 decoders. The tool does not transmit your input or output; encoding and decoding do not use an upload endpoint.