Blog / Security

  • totp
  • mfa
  • ntp
  • time-synchronisation
  • go
  • debugging

Why Your TOTP Codes Stop Working When the Server Clock Drifts

The classic symptom: a user types the six digits from their authenticator app, the digits are correct, and the server says no. Nothing is wrong with the secret or the app. The server just thinks it is a different time from everyone else.

A TOTP code is not a message from the server or a stored value. It is a hash of a shared secret and a counter, and the counter is the clock. If two clocks disagree, the two sides compute different codes, and both are right by their own reckoning.

The counter is just the time divided by 30

In RFC 6238, the counter is floor(unix_time / 30) with the default 30 second step. That counter goes into HOTP (an HMAC, then dynamic truncation to six digits). There is no other input.

So a server that is 40 seconds slow is computing codes for a counter one step behind the phone. Every code the phone shows is for a step the server considers to be in the future. The check fails purely because of the clock, and it fails every time, not intermittently.

Here is the whole thing in Go, using only the standard library:

package totp

import (
	"crypto/hmac"
	"crypto/sha1"
	"crypto/subtle"
	"encoding/binary"
	"fmt"
	"time"
)

func code(secret []byte, step int64) string {
	var msg [8]byte
	binary.BigEndian.PutUint64(msg[:], uint64(step))
	mac := hmac.New(sha1.New, secret)
	mac.Write(msg[:])
	sum := mac.Sum(nil)
	off := sum[len(sum)-1] & 0x0f
	n := binary.BigEndian.Uint32(sum[off:off+4]) & 0x7fffffff
	return fmt.Sprintf("%06d", n%1000000)
}

// Verify checks input against steps now-window..now+window and
// reports which offset matched, so the caller can log drift.
func Verify(secret []byte, input string, now time.Time, window int64) (offset int64, ok bool) {
	cur := now.Unix() / 30
	for d := -window; d <= window; d++ {
		want := code(secret, cur+d)
		if subtle.ConstantTimeCompare([]byte(want), []byte(input)) == 1 {
			offset, ok = d, true
		}
	}
	return offset, ok
}

Note the loop does not return early. Every candidate is always checked, so timing does not reveal which step matched.

What a window actually buys you

RFC 6238 itself suggests accepting a step either side of the current one, to cover network delay and a user typing slowly at the end of a step. With a window of 1 you accept three steps, so a code stays valid for up to 90 seconds.

That also gives you a little clock tolerance. A skew of 30 seconds or less always works with a window of 1, and anything up to 60 seconds works depending on where in the step you land. Past that, you fail every time.

Quick detour, because it explains why the failures look odd: a skew of, say, 45 seconds is flaky rather than broken. Users succeed when the phone's step happens to sit within one of the server's, and fail otherwise. "It works if I wait a few seconds and try again" is the signature of skew sitting right at the edge of your window.

Do not just widen the window

Widening to 10 steps makes the symptom go away and makes the security worse. Each extra step is another valid code at any moment, so an attacker guessing six digits gets more chances, and a phished code stays usable for longer.

A better order of operations:

  1. Fix the server clock.
  2. Keep the window at 1 (or 0 if your users are happy with it).
  3. Rate-limit attempts per account, whatever the window.
  4. Reject a code whose step is not newer than the last one accepted.

That last point matters: RFC 6238 says a verifier must not accept the second attempt of an OTP once it has been validated. Store the last accepted step per user and require the new one to be greater.

Check the server clock first

Servers drift more often than phones, because phones sync with the network constantly. Common causes:

  • No time daemon running, or it is installed but cannot reach any upstream (firewalled UDP 123).
  • A VM or container host that was paused, snapshotted or suspended; the guest clock jumps or lags until corrected.
  • A laptop or WSL environment resuming from sleep.
  • A container reading a host clock that is itself wrong (containers share the host kernel clock, so fix the host).

Compare against a trusted source rather than trusting your eyes:

date -u +%s
timedatectl status
chronyc tracking

Look for "System clock synchronized: yes" in timedatectl. With chrony, "System time" in chronyc tracking shows the current offset from NTP time. Anything near a second is already worth a look, and anything near 30 seconds is your bug.

Log the offset and you can see it coming

This is why Verify above returns the matched offset. If successful logins mostly show offset 0 and then you start seeing -1, -1, -1, your server clock is creeping behind, or a particular user's device is. You find out from a log line instead of a support ticket.

RFC 6238 goes further and describes resynchronisation: when a code validates at an offset, the verifier can remember that drift for that token and centre the window on it next time. That suits hardware tokens whose clocks cannot be set. For phone apps it mostly hides a problem that the phone's own network time would have fixed, so I would log it and alert on it rather than silently compensate.

A note on leap seconds

Unix time ignores leap seconds, and TOTP uses Unix time, so a leap second does not shift your counter. If a smearing time source is in play, the server and phone may differ by a fraction of a second for a while, which is far below one step and does not matter here.

One more thing worth testing: make your verifier take now as a parameter, as above, rather than calling time.Now() inside. Then you can write a test that feeds a code generated for step N-2 and checks it is rejected with window 1 but accepted with window 2. It takes a few lines and it proves the boundary does what you think.