MD5 vs SHA-256 vs bcrypt: which hash to use for what
"Which hash should I use?" has three different answers, because people use hash functions for three different jobs: detecting changed data, proving a message came from someone holding a key, and storing passwords. MD5, SHA-256 and bcrypt each belong to one of these jobs, and most hashing mistakes in code review come from using one where another belongs. This guide explains what each guarantees, and gives working Node.js code for each job. Every example and timing was measured on an ordinary laptop.
The short answer
| Job | Use | Never use |
|---|---|---|
| Checksum a file or detect corruption | SHA-256 (MD5 only to match an existing value) | - |
| Content addressing, cache keys, ETags, dedup | SHA-256 | MD5 or SHA-1 where an attacker controls the input |
| Sign a webhook or API request | HMAC-SHA256 | A plain hash of secret + message |
| Store passwords | Argon2id, bcrypt or scrypt | MD5, SHA-1, SHA-256, even salted |
| Hash table or sharding key | xxHash, MurmurHash or your language's built-in | - |
What a cryptographic hash guarantees
A hash function maps input of any size to a fixed-size digest: 128 bits for MD5 (32 hex characters), 256 bits for SHA-256 (64 hex characters). A cryptographic hash should be:
- Deterministic: the same bytes give the same digest everywhere.
- Avalanching:
helloandHellohave completely unrelated SHA-256 digests (2cf24dba…and185f8db3…). - Preimage-resistant: given a digest, you cannot find an input for it except by guessing.
- Collision-resistant: you cannot find two inputs with the same digest.
Hashes are one-way. There is no "decrypt". Sites that "reverse" MD5 look the digest up in a table of hashes of common strings, which is exactly why fast hashes are useless for passwords.
MD5 and SHA-1: broken, but for what?
MD5 collisions have been practical since 2004, and today a pair of different files with the same MD5 can be produced in seconds. Attackers used MD5 collisions to forge a trusted CA certificate in 2008, and the Flame malware used one to fake a Microsoft code-signing certificate. SHA-1 fell in 2017, when Google published two different PDFs with the same SHA-1.
Be precise about what "broken" means. These are collision attacks: someone who controls the input can craft two inputs that match. They do not let anyone reverse a digest, and they do not make random corruption go unnoticed. So:
- Still fine: detecting a corrupted download or disk copy, deduplicating your own trusted data, and matching systems that require MD5, such as S3 ETags for single-part uploads, Gravatar URLs and many legacy APIs. The MD5 Generator is for exactly these cases.
- Not fine: signatures, certificates, anything where an attacker supplies or influences the input, and passwords.
For anything new, use SHA-256. It is about as fast as MD5 in practice. In Node.js on short inputs, both were limited by call overhead, at about 600,000 per second each. It also removes the need to argue about the threat model.
Job 1: checksums
import crypto from 'node:crypto';
import fs from 'node:fs';
const sha256 = (data) => crypto.createHash('sha256').update(data).digest('hex');
sha256('hello'); // '2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824'
// Large files: stream, so memory stays flat
const fileSha256 = (path) => new Promise((resolve, reject) => {
const h = crypto.createHash('sha256');
fs.createReadStream(path).on('data', (c) => h.update(c))
.on('end', () => resolve(h.digest('hex'))).on('error', reject);
});
From a shell: sha256sum file on Linux, shasum -a 256 file on macOS, Get-FileHash file in PowerShell. Know what a checksum proves: your file matches the published digest. If an attacker can replace both the file and the checksum on the same server, it proves nothing. That is why projects sign their checksum files (GPG, Sigstore) or publish them separately.
When two digests "should" match but do not, the input differs, not the algorithm. echo hello | md5sum hashes hello\n and gives b1946ac9…, not the 5d41402a… of hello. Use printf or echo -n. Also check CRLF versus LF line endings, UTF-8 versus UTF-16 (PowerShell's default for some commands), hex versus Base64 output, and JSON with different key order or whitespace.
Job 2: proving who sent a message (HMAC)
A plain hash proves integrity only if the digest itself arrives safely. An attacker who can change a webhook body can recompute its SHA-256 too. HMAC mixes a secret key into the hash, so only key holders can produce a valid tag. GitHub, Stripe, Slack and Shopify all sign webhooks with HMAC-SHA256, and JWTs signed with HS256 are HMAC too; the JWT Decoder can check one against a secret.
function verifySignature(rawBody, signatureHex, secret) {
const expected = crypto.createHmac('sha256', secret).update(rawBody).digest();
const received = Buffer.from(signatureHex, 'hex');
return received.length === expected.length && crypto.timingSafeEqual(received, expected);
}
Three details matter. Compute the HMAC over the raw request bytes, before any JSON parsing and re-serialising, which changes whitespace and breaks the match. Compare with timingSafeEqual, not ===, so response timing does not leak how many leading bytes were right; it throws on unequal lengths, hence the length check. And never invent sha256(secret + message). With SHA-256, MD5 and SHA-1, that construction is open to length-extension attacks, where an attacker appends data and computes a valid tag without the secret. HMAC exists precisely to prevent that. Most providers also include a timestamp in the signed content. Reject old ones to stop replays.
Job 3: storing passwords
General-purpose hashes are designed to be fast, and that is exactly wrong for passwords. If your user table leaks, the attacker hashes guesses and compares. A single modern GPU computes billions of SHA-256 or MD5 hashes per second, so every common password and most eight-character ones fall within hours. A per-user salt stops precomputed tables, but it does not slow down guessing.
Password hashing functions are deliberately slow, and Argon2id and scrypt are also memory-hard, which cripples GPUs. Each includes a random salt and its parameters in the output string, so you store one value:
import bcrypt from 'bcrypt';
const stored = await bcrypt.hash(password, 12); // ~200 ms here; '$2b$12$…'
const ok = await bcrypt.compare(attempt, stored); // true or false
import * as argon2 from '@node-rs/argon2';
const stored = await argon2.hash(password, { memoryCost: 19456, timeCost: 2, parallelism: 1 });
// '$argon2id$v=19$m=19456,t=2,p=1$…'
const ok = await argon2.verify(stored, attempt);
Measured here: bcrypt at cost 10 took 52 ms and cost 12 took 200 ms. Every +1 doubles the time. Argon2id at the OWASP baseline settings above (19 MiB of memory, 2 iterations) took 14 ms, and the memory requirement is what makes it expensive on GPUs. Node's built-in crypto.scrypt needs no dependency; with N = 2^17 it took 328 ms. Hashing the same password twice gives different strings, because each has its own salt.
Which one?
- Argon2id is the current first choice for new systems (OWASP, RFC 9106).
- bcrypt is fine and everywhere, with one trap: it uses only the first 72 bytes of input. Hashing
"a" × 72 + "X"and comparing against"a" × 72 + "Y"returnedtrue. That matters for long passphrases, and especially for "pepper" or pre-hash schemes that concatenate strings. Some libraries now reject inputs over 72 bytes instead. - scrypt is memory-hard and built into Node.js and Python's
hashlib. - PBKDF2 with a high iteration count (OWASP suggests 600,000 for HMAC-SHA256) when FIPS compliance requires it.
Tune the cost so a hash takes roughly 100-500 ms on your servers. Raise it as hardware gets faster: on login, if the stored hash uses old parameters, re-hash the password you just verified and save the new string. Limit login attempts too, because a slow hash protects a leaked database, not a login form that allows unlimited guesses.
Migrating away from MD5 or SHA-1 passwords
If you inherit a table of unsalted MD5 password hashes, you do not need everyone's password to fix it. Wrap the old hashes: store bcrypt(md5_hash) for every row immediately, and mark them as wrapped. At login, compute bcrypt.compare(md5(attempt), stored) for wrapped rows, and on success replace the row with a plain bcrypt(attempt). The weak hashes disappear from the database on day one, not when the last user logs in.
Summary
MD5 for compatibility checksums, SHA-256 for integrity and identifiers, HMAC-SHA256 when a secret must vouch for a message, and a slow, salted password hash for passwords. When you see a fast hash near the word "password" in a code review, stop and ask which of the three jobs it is really doing. Related reading: how JWT signatures work and why Base64 is not encryption.