What Is a Webhook?
Scrapeless Scraping API can send a task result to a configured webhook endpoint after an asynchronous job completes.
TL;DR
- A webhook is an event notification sent as an HTTP request. A producer calls a receiver URL when a subscribed event occurs.
- A receiver needs a delivery contract. The request method, payload, success response, and failure behavior must come from the provider documentation.
- Receiving a request is not the same as trusting it. Authenticate the source using the mechanism the provider actually supports and validate the event schema.
- Duplicate and out-of-order events are normal design cases. Use a stable event or task identifier and separate receipt from business processing.
A webhook lets a service tell your application that something has happened without waiting for your application to ask again. A task service may report completion, a payment service may report a settled charge, or a deployment platform may report a finished build. The producer initiates an HTTP request to a URL controlled by the consumer. Your application receives the message, decides whether to accept it, and then updates its own state.
The short definition hides several engineering choices. An event can arrive after the original user action has ended, and a network acknowledgement can be lost even when processing succeeded. A useful webhook design therefore covers authentication, storage, deduplication, ordering, and recovery. This guide follows one notification from registration through receipt and explains where a receiver should trust evidence rather than assumptions.
The Producer, Event, and Receiver Contract
Three roles make a webhook exchange understandable. The producer owns the event, the event record describes what occurred, and the receiver exposes a reachable endpoint. The producer needs an endpoint URL and often a subscription that specifies which events to send. The receiver needs to know the expected request method, headers, payload shape, and acknowledgement response. The HTTP semantics specification defines the request and response framework; the provider contract supplies the event meaning.
Imagine a scraping job identified by a task ID. Submission may return before collection finishes. Later, a completion message can carry that ID and a status. The receiver uses the ID to join the event with the stored job record. It should not assume that every payload embeds the full final dataset; some services send a pointer, while others include the result. Persist the original task ID at submission so this later association does not depend on a guessed URL or display label.
The current AI Scraper task lifecycle documents an optional webhook URL and the pushed task ID, status, input, and optional task result. That is a concrete contract, not a universal webhook schema. A receiver built for another product must check that product’s documentation. Even within one platform, two event families can use different field names and completion states.
Webhook Delivery Versus Polling
Polling asks a service for state on a schedule. A webhook changes the direction of the first move: the source contacts your endpoint when it has a relevant event. This can reduce empty checks and shorten the delay before your workflow notices a change. It also adds a public receiver, network delivery uncertainty, and new security work. Neither pattern makes the underlying business event transactional across two independently operated systems.
A practical architecture often combines a webhook for quick reaction with a periodic status check for reconciliation. The scheduled check compares local jobs with the provider’s authoritative task state and finds missing notifications or local processing mistakes. Keep that check narrow: query only jobs whose local state has not reached a trusted terminal state. The receiver and the reconciler should both call the same idempotent state transition function rather than create two conflicting paths.
The choice depends on the time sensitivity and the provider’s features. If a result must be processed shortly after completion and the provider documents callbacks, a webhook is useful. If the consumer has no reachable HTTPS endpoint or the provider has no delivery mechanism, polling may be simpler. For interactive back-and-forth messages, a persistent connection can fit better. The name webhook does not promise any particular latency or durability.
How to Accept a Notification Safely
Treat an incoming HTTP request as untrusted input. First limit its size and method, require the intended route, and parse only the documented media type. Then apply the provider’s authentication method. If the provider signs payloads, verification must use the exact raw bytes and the documented signature headers before a JSON parser changes representation. The Standard Webhooks specification describes a common signed-event model, but a producer must actually implement it before the receiver can rely on that model.
A signature check answers whether a holder of the signing secret produced those bytes. It does not prove that the event is fresh, that the payload matches your subscription, or that the task belongs to your account. Check timestamp or replay metadata when the provider offers them; map the task ID to a local job; validate required fields and allowed state transitions. Keep secrets in a managed secret store and use constant-time comparison when your chosen verification scheme calls for comparing message authentication values.
Be precise about Scrapeless behavior. The documented AI Scraper lifecycle establishes the callback payload and HTTPS URL requirement, but it does not by itself establish that this callback carries a particular signature header. Build the receiver from the selected product’s current delivery contract. If authentication details are not documented for that event family, ask support or use an application-controlled secret in the receiver URL only if the provider explicitly supports that design and your security review accepts its exposure risks.
Receipt, Queueing, and Idempotent Processing
The HTTP handler should do enough work to make receipt durable and then acknowledge using the provider’s documented success response. A common sequence is validate request, record an event ID or task ID with the raw payload and receipt time, enqueue a processing job, then respond. Long data transformations inside the request path raise the chance that the producer sees a timeout even if your database update eventually succeeded. That ambiguity is why business processing needs a duplicate guard.
Choose a deduplication key from the provider contract. A stable event ID is ideal when it exists. If the producer only exposes a task ID plus terminal status, construct an application key from those documented fields and the subscription or account scope. Add a database uniqueness constraint around that key. The consumer can then accept the same notification more than once while applying the business effect only once. Do not rely on an in-memory set that disappears when the process restarts.
A task can also change state before messages arrive. For example, a completion notice might be processed after a separate status check has already marked the job complete. The transition function should compare current and incoming states and decline any change that would move a completed job backward. Store the original payload and decision so an operator can explain why a message was accepted, ignored as duplicate, or rejected as invalid.
Registration, Failures, and Observability
Register only a receiver URL you control. Keep the path purpose-specific and avoid exposing internal addresses as callback destinations. An application that lets arbitrary users register webhook endpoints can become a server-side request forgery channel if it fetches private network targets. Validate scheme and destination policy at registration, and recheck resolved addresses when delivery occurs if your system is itself the producer. The OWASP SSRF guidance explains the network boundary behind this risk.
Record delivery evidence at the consumer: receive time, provider task or event identifier, schema version when supplied, authentication decision, acknowledgement status, queue ID, and final processing state. Exclude secrets and personal data that are not necessary for operations. A dashboard that says only “webhook succeeded” cannot distinguish HTTP delivery from successful downstream storage. Separate those measurements and alert on stalled accepted events.
Test the workflow with a harmless event before depending on it. Confirm that the receiver is reachable through HTTPS, that a valid payload takes the expected state transition, that malformed payloads are rejected, and that repeated delivery changes no additional business state. Simulate an out-of-order message if the provider can send several events for one object. These checks exercise the receiver contract without assuming the producer guarantees order or exactly-once delivery.
When a Webhook Is the Wrong Tool
A webhook is a good fit for discrete changes that matter to another system. It is a poor fit for reading a current resource on demand. If a user opens a dashboard and asks for the latest account balance, an ordinary API query is clearer than waiting for a past notification. A webhook may signal that the balance changed, while the API remains the place to reconcile the authoritative value.
High-volume event streams can also need broker features that an isolated HTTP receiver does not provide, such as partitioned ordering and replay by offset. A webhook producer might offer its own delivery logs, but that is a provider-specific feature, not part of the basic concept. Use the smallest mechanism that satisfies the event volume, delay tolerance, and recovery requirement.
For Scrapeless jobs, keep the callback tied to a documented asynchronous task. Scraping API is a product surface for structured web data, while the related actor workflow guide explains why a task identifier and result handling matter. Define exactly what local action a completed task triggers, and retain a status lookup path for auditing that action later.
Conclusion
A webhook is an event-triggered HTTP call from a producer to a receiver. The reliable unit is the complete contract: registration, authenticated receipt, durable acknowledgement, duplicate control, state transition, and reconciliation. Start with one documented event and a measurable local outcome before extending the receiver to additional event families.
Build an Event-Driven Collection Flow
Create a supported task, retain its identifier, and connect its completion to a receiver you can validate and observe.
Sign up today and get $5 in free credit — no credit card required.
Claim Your $5 Credit →FAQ
Is a webhook the same as an API?
A webhook is one way to use an HTTP API boundary for event notifications. The producer initiates the call after an event, while an ordinary API client usually initiates a query or command when it needs a result. A system can use both patterns for the same task: the callback signals completion and a status endpoint confirms authoritative state.
Can a webhook arrive twice?
Yes. A delivery acknowledgement can be lost, or a producer can send more than one notification for the same object. The receiver should record a stable identifier and make the business transition idempotent. A repeated message may still be acknowledged after the consumer recognizes it as a duplicate.
Should the receiver verify a signature?
Verify a signature when the provider documents one for that event family. Use the raw request body and the exact verification procedure the provider specifies. Do not invent a header name or assume every webhook has a signature. Also validate freshness, task ownership, and payload shape where the contract supports those checks.
What response should a receiver send?
Send the success or failure status defined by the producer’s webhook contract after durable acceptance or rejection. Keep expensive downstream work outside the request path when the contract permits. The acknowledgement only reports receipt; your own job record should separately show whether business processing completed.
Can a webhook replace polling entirely?
A webhook can eliminate frequent empty checks, but a periodic reconciliation query remains useful for important state. It can detect lost notifications, application outages, or local processing errors. The correct interval depends on the task’s tolerance for stale state and the provider’s documented status interface.