Threat Modelling a Webhook Receiver Before You Write the Handler
A webhook receiver is a public URL that changes your system's state when someone sends it JSON. That is the whole threat model in one sentence. Anyone on the internet can POST to it, and the only thing separating "payment succeeded" from "a stranger says payment succeeded" is whatever you decide to check.
Most people write the handler first, parse the JSON, do the thing, and bolt security on later. This post does it the other way round: list what can go wrong, decide what the handler must do about each, and only then write code. The code turns out to be short.
Start with what the endpoint is allowed to cause
Before any attacker enters the picture, write down what a valid event does. Marks an order paid? Triggers a deploy? Creates a user? The worse the side effect, the more the rest of this list matters.
Then list who is talking to you. There is the provider (honest, but retries aggressively), the attacker (who knows your URL, because URLs leak into logs, repos and browser histories), and your own future self (who will rotate a secret at 2am and forget one consumer). Three actors is enough for a short pass.
The six things that go wrong
- Spoofing: someone forges an event.
- Tampering: someone alters a genuine event in transit or after capture.
- Replay: someone resends a genuine, correctly signed event.
- Duplicates: the honest provider delivers the same event twice.
- Exhaustion: huge bodies, floods, or slow handlers that tie up workers.
- Leakage: payloads or secrets end up in logs and error pages.
Notice that replay and duplicates look identical from the handler's side. They have different causes but the same fix, which is handy.
Spoofing and tampering: sign the raw bytes
The standard answer is an HMAC over the request body using a secret shared with the provider. You recompute it and compare. Two details matter more than the algorithm.
- Verify the raw bytes you received, not a re-serialised version. Parse then re-encode and key order or whitespace changes, so a valid signature fails (or worse, you "fix" it by loosening the check).
- Compare with
hmac.Equal, never==. It is constant time, and it is one function call.
Quick detour, because it is a tempting shortcut: IP allowlisting. Providers publish source ranges, and filtering on them feels solid. It is a useful extra layer, but not a replacement. Ranges change, proxies muddy what "source IP" even means, and an allowlist says nothing about whether the body was altered. Treat it as noise reduction, not authentication.
Replay: a valid signature is not a fresh one
An HMAC over the body alone proves the provider sent that body at some point. It does not prove they sent it now. Capture one signed "refund issued" event and you can post it forever.
Many providers fix this by signing a timestamp together with the body. Stripe, for example, signs the timestamp, a full stop, then the payload. You then reject anything outside a tolerance window. Five minutes is a common choice, but it is a trade-off against clock skew, so pick it deliberately.
The window only shrinks the replay opportunity; it does not close it. Closing it needs the next section.
Duplicates: make the handler idempotent
Providers retry when you time out or return a 5xx, so duplicates are normal behaviour, not an attack. Most deliver a unique event ID. Record it, and if you have seen it, return success without doing the work again.
Hang on, why return success rather than an error for a duplicate? Because an error invites yet another retry. From the provider's side, "already handled" and "handled now" are the same good news.
The ID check has to be atomic with the side effect, or two concurrent deliveries both pass it. A unique constraint on the event ID in your database does this for free: insert first, and treat a conflict as "seen".
Exhaustion: cap, acknowledge, defer
Three cheap rules cover most of it:
- Cap the body size before reading it, so a 2 GB POST cannot fill memory.
- Verify the signature before doing anything expensive such as database lookups or JSON parsing.
- Return quickly and process later. If the work takes thirty seconds, the provider will time out and retry, and now you have a duplicate problem and a load problem.
Note the ordering: the cheapest check goes first, and the most trusting step (acting on the content) goes last.
Leakage: the boring one that bites
Webhook bodies often contain email addresses, names and amounts. Logging the full payload on failure is tempting for debugging, and it quietly builds a store of personal data with no retention plan. Log the event ID, type and the reason for rejection instead. Never log the signature header together with the body, since together they are exactly what a replay needs.
Also keep the signing secret out of the repo, and support two secrets at once during rotation so you can change one without dropping events.
The handler that falls out
With the list done, the code is mostly a checklist. This is a simplified sketch: the Store and queue are stand-ins for a real database table with a unique constraint and a real job queue, and the header names follow a Stripe-like scheme, so adapt them to your provider.
package webhook
import (
"crypto/hmac"
"crypto/sha256"
"encoding/hex"
"errors"
"io"
"log/slog"
"net/http"
"strconv"
"time"
)
const (
maxBody = 1 << 20
tolerance = 5 * time.Minute
)
type Receiver struct {
Secrets [][]byte // current first; old one kept during rotation
// Claim returns false if the event ID was already recorded.
Claim func(id string) (bool, error)
Enqueue func(body []byte) error
}
func verify(secret []byte, ts string, body []byte, sigHex string, now time.Time) error {
sent, err := strconv.ParseInt(ts, 10, 64)
if err != nil {
return errors.New("bad timestamp")
}
age := now.Sub(time.Unix(sent, 0))
if age > tolerance || age < -tolerance {
return errors.New("timestamp outside window")
}
got, err := hex.DecodeString(sigHex)
if err != nil {
return errors.New("bad signature encoding")
}
mac := hmac.New(sha256.New, secret)
mac.Write([]byte(ts))
mac.Write([]byte("."))
mac.Write(body)
if !hmac.Equal(got, mac.Sum(nil)) {
return errors.New("signature mismatch")
}
return nil
}
func (rc *Receiver) ServeHTTP(w http.ResponseWriter, r *http.Request) {
r.Body = http.MaxBytesReader(w, r.Body, maxBody)
body, err := io.ReadAll(r.Body)
if err != nil {
http.Error(w, "body too large", http.StatusRequestEntityTooLarge)
return
}
ts, sig := r.Header.Get("X-Timestamp"), r.Header.Get("X-Signature")
var verr error
for _, s := range rc.Secrets {
if verr = verify(s, ts, body, sig, time.Now()); verr == nil {
break
}
}
if verr != nil {
slog.Warn("webhook rejected", "reason", verr)
http.Error(w, "unauthorised", http.StatusUnauthorized)
return
}
fresh, err := rc.Claim(r.Header.Get("X-Event-ID"))
if err != nil {
http.Error(w, "try again", http.StatusServiceUnavailable)
return
}
if fresh {
if err := rc.Enqueue(body); err != nil {
http.Error(w, "try again", http.StatusServiceUnavailable)
return
}
}
w.WriteHeader(http.StatusAccepted)
}
One wrinkle: if Enqueue fails after Claim succeeds, the retry will look like a duplicate and be dropped. In a real system, claim and enqueue should be one transaction, or the claim should be released on failure. That is exactly the sort of gap the threat list is meant to expose before it reaches production.
Check that it actually rejects things
A signature check that has never been seen to fail is a guess. Write tests for the unhappy paths: a flipped body byte, a timestamp ten minutes old, a malformed hex signature, an oversized body, and the same event ID sent twice. Five small cases, and each one maps to a line in the list at the top.