A webhook is an HTTP callback. The producer (Stripe, GitHub, your application) sends an HTTP POST to a consumer-provided URL when something happens. The consumer responds with HTTP 2xx to acknowledge.
Webhooks are the standard pattern for server-to-server event notifications. The mechanics are simple; the operational details — delivery guarantees, retries, signatures — are where the complexity lives.
This page is about the patterns from both sides: designing webhooks that work, and consuming webhooks reliably.
Webhooks are at-least-once. The producer sends; if delivery fails (timeout, 5xx), the producer retries. Eventually most events are delivered; sometimes events are delivered multiple times.
Consumers must handle:
Senders should not assume:
Standard retry approach for senders:
Backoff schedule example: 1m, 5m, 15m, 1h, 6h, 24h, 48h, 72h. After ~3 days of retries, mark as failed.
Document the retry policy. Consumers need to know how long to expect retries to continue.
Webhooks must be signed; consumers must verify. Without signatures, anyone can POST to the webhook URL and trigger consumer logic.
Standard pattern (Stripe, GitHub, Slack):
1. Producer creates HMAC-SHA256 of payload using shared secret
2. Producer includes signature in HTTP header
3. Consumer recomputes HMAC and verifies equality
POST /webhooks/orders
X-Signature: sha256=<hex-encoded-hmac>
Content-Type: application/json
{ "event": "order.shipped", ... }
Consumer code:
expected = hmac.new(secret, payload, hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, received):
return 401
compare_digest is constant-time; == is timing-attack-vulnerable.
Even with signatures, an attacker who captures a valid webhook can replay it. Defenses:
Stripe's pattern: signature includes timestamp; consumer rejects if |now - timestamp| > 300s.
Reliable webhook consumers need:
Respond with 2xx as soon as the webhook is received and stored — typically under 5 seconds. Producers often timeout faster than that.
def handle_webhook():
verify_signature()
enqueue_for_processing() # async work
return 200 # ack now
Don't do the actual work synchronously; enqueue and process asynchronously. Otherwise slow processing causes producer retries (and duplicate work).
Same event delivered twice should not double-process. Use the event ID:
if not seen_event_ids.exists(event_id):
seen_event_ids.add(event_id, ttl=7days)
process(event)
See IdempotencyPatterns.
Don't process events directly from the HTTP handler. Persist to a queue (database, Kafka, SQS) and process from there. If processing fails, the event is preserved for retry.
If your webhook endpoint is down, events are eventually lost (after producer retry exhaustion). Mitigations:
Even with at-least-once delivery, include an event ID:
{
"event_id": "evt_8d4f...",
"event_type": "order.shipped",
"data": {...}
}
Consumers use the event ID for deduplication.
Within a single resource (one order's events), events should be sequential. Events to different resources can be parallel.
If strict ordering matters, include a sequence number or rely on retry-with-backoff to maintain order.
Provide a UI or API for:
The Stripe/GitHub-style management UI sets the bar.
Events may evolve. Version them:
version field in payloadOr: namespace event types per major version.