How to Build a Reliable Webhook Handler with Retries and Idempotency
- Author
- Vishal Maurya
- Published on
- Reading time
- 4 min read
Overview
Webhooks let one system notify another when an event occurs. Payment providers, CRMs, email platforms, and SaaS integrations use them to deliver updates without requiring constant polling.
Webhook delivery is rarely a guarantee that each event arrives once, in order, and without delay. A sender may retry after a timeout, deliver events out of order, or send an event your application cannot process yet. A robust handler plans for those conditions.
1. Verify the Sender Before Trusting the Payload
Use the provider's documented signature-verification process. Many providers sign the raw request body with a secret, and changing the body before verification can invalidate the signature.
Do not trust a user ID, payment status, or event type merely because it appears in JSON. A public endpoint that accepts arbitrary webhook payloads can become a path to forged state changes.
Use constant-time comparison when implementing signature verification yourself, and prefer the provider's maintained SDK when available. Keep signing secrets server-side and support rotation according to the provider's guidance.
2. Acknowledge Quickly and Move Slow Work to a Queue
A webhook sender may stop waiting after a short deadline and retry delivery. Avoid running lengthy document processing, sending multiple emails, or making several slow API calls before returning the acknowledgement.
A common pattern is:
- Read and verify the request.
- Validate the event envelope.
- Persist the event or enqueue it durably.
- Return the provider's expected success response.
- Process the business operation asynchronously.
The persistence step matters. Returning success before the event has been safely recorded can lose work if the application crashes immediately afterward. Choose a durable queue or database-backed inbox appropriate to the system's reliability requirements.
3. Make Duplicate Delivery Harmless
Many providers retry delivery when they do not receive a successful response. Your handler should therefore be idempotent.
Store the provider's event ID under a uniqueness constraint. When the same event arrives again, detect the duplicate and avoid repeating the business action. A database transaction can help ensure that event registration and the corresponding state change remain consistent.
A simple check-then-insert without a database constraint can race when two duplicate requests arrive at the same time. Use a unique key and handle the conflict explicitly.
4. Do Not Assume Events Arrive in Order
An update event may arrive after a later event, or an old event may be retried. Where possible, fetch the authoritative current object from the provider or compare event versions and timestamps using the provider's documented semantics.
Do not blindly overwrite newer application state with whichever payload arrived last. Event ordering rules vary by provider; use its documentation rather than assuming timestamps are globally reliable.
5. Separate Delivery Status from Business Status
A webhook can be successfully received but fail during business processing. Track those states separately.
Useful fields include provider event ID, event type, received time, processing status, attempt count, last error category, and completion time. Avoid storing unnecessary sensitive payload data indefinitely; define retention and redaction policies.
Provide a controlled way to inspect and replay failed events. A replay should pass through the same idempotency checks as the original delivery.
6. Retry with Limits and Backoff
Retry temporary network errors and recoverable dependency failures with bounded exponential backoff. Do not retry invalid signatures or malformed payloads as if they were transient. After the retry limit, route the event to an operational review queue or dead-letter workflow.
Also consider rate limits on the provider side. A retry storm can worsen an outage, so add jitter and avoid multiple layers of retries multiplying each other.
7. Test the Failure Cases
Include tests for invalid signatures, duplicate event IDs, simultaneous duplicate requests, worker crashes, dependency timeouts, malformed payloads, and events that arrive out of order. Test the case where the database commits but the HTTP response is lost; that is a common reason a provider sends the same event again.
Conclusion
A webhook integration is more than a POST endpoint. It needs signature verification, durable receipt, idempotent processing, recovery controls, and clear operational status.
If your product integrates with payment, CRM, messaging, or third-party SaaS webhooks, I can help build the receiving endpoint and the processing workflow around it. Contact me with the provider and event flow you need to support.
Additional Resources
- Stripe webhook documentation (useful as an example of provider-specific delivery semantics)
- GitHub webhook best practices