Why a Password Hash Takes 50ms Locally and 400ms in Production
You tuned Argon2id on your laptop until a login cost about 50ms. You deployed it, and the login endpoint now takes 400ms, with the same code and the same parameters. Nothing is broken. Argon2 is designed to eat memory and CPU, so it is unusually sensitive to every difference between your laptop and a container on a shared host.
There are a handful of usual suspects. Most can be measured in a couple of minutes, so let's measure rather than guess.
Start with a benchmark you can run in both places
Run the same tiny program on the laptop and inside the production container (same image, same limits). It prints the CPU count Go sees, then times the hash a few times.
package main
import (
"crypto/rand"
"fmt"
"runtime"
"time"
"golang.org/x/crypto/argon2"
)
func main() {
salt := make([]byte, 16)
if _, err := rand.Read(salt); err != nil {
panic(err)
}
fmt.Println("NumCPU:", runtime.NumCPU(), "GOMAXPROCS:", runtime.GOMAXPROCS(0))
for i := 0; i < 5; i++ {
start := time.Now()
argon2.IDKey([]byte("correct horse"), salt, 2, 64*1024, 4, 32)
fmt.Println(time.Since(start))
}
}
The arguments after the salt are time cost (2 passes), memory in KiB (64 MiB), threads (4 lanes) and key length. Keep them identical in both places. Now read the output against the causes below.
Suspect 1: your threads parameter is not free parallelism
The "threads" value in Argon2 is really the number of lanes. Each lane can be computed independently, so on a laptop with 8 cores, four lanes run side by side. On a container with one vCPU, the same four lanes run one after another.
The arithmetic is blunt: four lanes on one core take roughly four times the wall-clock time of four lanes on four idle cores. Total work is unchanged; only the elapsed time moves. If your laptop shows 50ms and production has one usable core, 200ms is the expected result before anything else goes wrong.
Check the first line of the benchmark output. If production reports NumCPU: 1 or a small number, that alone may account for much of the gap.
Suspect 2: the CPU quota throttles you mid-hash
This one is sneakier. A container limit such as "500m" CPU is usually enforced as a quota per scheduling period (commonly 100ms). Half a CPU means 50ms of CPU time per 100ms window. Your hash burns through CPU time from several threads at once, hits the quota, and every thread is frozen until the next period.
So a hash needing 60ms of total CPU across four threads can stall for most of a period even though the container is nowhere near "busy" on average. Averaged CPU graphs hide this completely.
Quick detour, because it is worth knowing: the runtime and the quota are separate mechanisms. Go decides how many threads to run from the CPU count (and, in recent Go versions, may take cgroup limits into account; check your version's release notes rather than assuming). The kernel enforces the quota regardless. You can have a process that happily runs eight threads under a limit of one CPU.
To see whether it is happening, read the cgroup counters from inside the container (cgroup v2):
cat /sys/fs/cgroup/cpu.max
cat /sys/fs/cgroup/cpu.stat
cpu.max shows the quota and period. In cpu.stat, look at nr_throttled and throttled_usec. Run the benchmark, read them again, and if they jumped, you have found your 400ms. The kernel documentation for cgroup v2 describes these fields.
Suspect 3: 64 MiB of memory has to come from somewhere
Argon2 allocates its whole memory matrix up front and then walks it in a data-dependent pattern. On a fresh process, those pages are untouched, so the first touch of each 4 KiB page is a page fault handled by the kernel. That is thousands of faults for 64 MiB.
On a long-lived server the Go runtime may recycle memory, so later hashes can look different from the first. The benchmark shows this: if run one is slower than runs two to five, you are paying allocation and fault costs. If every run is uniformly slow, look elsewhere.
A memory limit that is too tight makes it worse. Ten concurrent logins at 64 MiB each is 640 MiB of live matrix, before your application's own heap. Near the limit you get garbage collection pressure, reclaim work, or an OOM kill. Latency climbs long before anything crashes.
Suspect 4: the hardware is genuinely slower per core
A modern laptop can boost a single core to a high clock speed for short bursts and has fast memory sitting close by. A cloud vCPU is often a hyperthread on a server part with a lower sustained clock, sharing caches and memory bandwidth with other tenants.
Argon2 is memory-hard on purpose: it hammers memory bandwidth and cache misses, which is exactly where shared servers hurt most. A noisy neighbour does not need to steal your CPU; it only needs to compete for memory bandwidth.
You cannot fix this from code. You can only measure it, which is why the benchmark should run on the production instance type, at a quiet time and at a busy one.
Suspect 5: concurrency multiplies everything above
A single-request benchmark is the best case. In production, logins arrive together. Each one wants its lanes, its 64 MiB and its CPU time, and they all queue on the same quota.
If you have one effective core and ten simultaneous logins, the tenth waits behind nine others. Per-request timing looks fine in a test and terrible under load. Latency here is a queueing problem as much as a hashing problem, so measure with concurrent callers too.
Bounding concurrency also protects you from a cheap denial of service, since anyone can hit the login endpoint. A small semaphore around the hash call, sized to the cores you actually have, keeps memory use predictable and makes the queue explicit:
var hashSlots = make(chan struct{}, 2)
func hash(pw, salt []byte) []byte {
hashSlots <- struct{}{}
defer func() { <-hashSlots }()
return argon2.IDKey(pw, salt, 2, 64*1024, 4, 32)
}
Choose the slot count from your real CPU allowance, not from a guess. In a real handler you would also want to give up when the request context is cancelled, rather than blocking forever on the channel.
What to change once you know which one it is
- Lane count too high for the cores: set threads to the CPU you reliably get in production, not what your laptop has.
- Throttling: raise the CPU limit (or remove the limit and keep a request), so bursts are not frozen mid-hash.
- Memory pressure: raise the memory limit above concurrency times matrix size plus your heap, or lower concurrency.
- Slow hardware: retune the parameters on that hardware, not on your laptop.
Retuning has a wrinkle. Lowering the cost weakens the hash, so decide on a defensible floor first, and change the lane count before you cut memory, because lanes affect timing far more than they affect the attacker's cost model. Old hashes keep verifying if the parameters are stored in the encoded string, and you can upgrade them at the next login.
Where the 50ms number really comes from
The lesson is that "50ms" was never a property of the algorithm. It was a property of your laptop's cores, clock, memory and idle load at that moment. Tune on the target environment, with the target limits, under concurrent load, and record the environment next to the number. Then the next person to see 400ms will know which knob to check first.