The Meeting Record

Webhook and Event-Driven Delivery for Meeting Transcription Results

Webhooks beat polling for transcription pipelines by eliminating wasted requests and latency.

Editor at Large · · 13 min read
Cover illustration for “Webhook and Event-Driven Delivery for Meeting Transcription Results”
Meeting Data Pipelines · September 29, 2026 · 13 min read · 2,826 words

Webhook and Event-Driven Delivery for Meeting Transcription Results.

Why polling is the wrong default for transcription pipelines

Webhook delivery is the production-grade pattern for meeting transcription pipelines, and the reason comes down to a simple design tradeoff that polling never resolves. Polling means the client asks "is it done yet?" on some fixed interval, burning a request each time regardless of whether anything changed, and paying for that rhythm in both wasted calls and added latency equal to whatever interval was chosen. For an API that answers in milliseconds, this barely matters. Transcription doesn't work that way.

Transcription jobs run for minutes, not milliseconds, and that duration breaks the math behind polling entirely. A short interval causes a service to burn through requests checking on a job that was never going to finish that fast. A long interval leaves results sitting ready on the server while the client remains ignorant, adding delay that has nothing to do with actual processing time. Any interval choice sacrifices one problem to favor the other.

Prototypes and CLI scripts can get away with polling, but backend services processing audio at any real volume cannot. A script running on a developer's machine, checked by a human watching the terminal, tolerates the inefficiency because the cost is measured in a few wasted calls. A backend service handling concurrent jobs from real users multiplies that inefficiency by every job in flight, and the wasted requests turn into real infrastructure cost.

Register a URL, the platform POSTs the moment a job completes, and the client does zero waiting. That's the whole idea. Nothing about it is exotic, but the discipline it requires (configuring endpoints correctly, reading the payload right, verifying who's actually sending the request, and handling the retries that inevitably occur) is what separates an integration that happens to work in a demo from one that holds up in production. The rest of this piece walks through each of those pieces in turn.

The webhook delivery model in a transcription context

The mechanism is straightforward once it's named clearly. An event occurs, a meeting ends or a job finishes, and the platform POSTs a structured payload to whatever HTTPS URL the developer registered ahead of time; the receiving service then acts on that payload. Every implementation, regardless of vendor, is built from the same handful of components repeated in different arrangements.

There's the event source, which is the transcription platform itself detecting that a job has finished. There's the event, the specific trigger being reported, such as a completed transcript, a failed job, or a meeting that has ended. There's the endpoint, the publicly reachable HTTPS URL the developer registers with the platform in advance. There's the payload, the JSON body carrying the event's details, and what that body actually contains varies quite a bit from one platform to the next. Layered on top of all of that is a security concern, since something needs to confirm the request is genuinely coming from the platform and not from an imposter who found the URL.

A useful way to explain this model to someone unfamiliar with it: instead of the client repeatedly pulling for updates, the server pushes the update out on its own schedule. Some engineers call it a reverse API, and the label holds up. The client that once had to ask now simply waits to be told.

Within meeting transcription specifically, this pattern splits into two distinct temporal modes that solve genuinely different problems. One mode delivers incremental segments while the meeting is still happening. The other delivers a single, complete payload once the whole job, start to finish, has wrapped up. Conflating these two is where a lot of integration work goes sideways, so they deserve to be treated separately.

Real-time versus post-meeting delivery: two different problems

Real-time delivery means the platform is emitting events while the meeting is still live, streaming finalized speech segments as they're produced rather than waiting for the whole session to end. Recall.ai's approach illustrates this well: it emits transcript.data webhook events carrying finalized segments, and developers who need lower latency can also configure transcript.partial_data for results that arrive before a segment is fully finalized. Delivery through this path runs at sub-second latency, which is the entire point of choosing it over waiting for a post-meeting job.

Zoom's RTMS model takes a different architectural approach to the same real-time goal. A webhook signals that a streaming session is starting, the developer's server then opens a WebSocket connection directly to Zoom's media servers, and from that point on, live audio, video, transcript data, and participant events stream straight from Zoom's infrastructure. Notably, there's no bot sitting in the meeting room capturing anything, since the data comes from Zoom's own servers rather than from a recording participant. That distinction matters operationally: fewer moving parts inside the meeting itself, and no bot avatar for participants to notice or object to.

Real-time delivery earns its architectural complexity when the product genuinely needs to react while the meeting is still running. Live captioning is the obvious case, but the same logic applies to in-meeting AI assistants, compliance monitoring that needs to flag something before a conversation ends, and analytics dashboards updating in real time. AssemblyAI's streaming architecture offers a hybrid: the session delivers responses over a WebSocket while it's live, and once the streaming session ends, a webhook fires to deliver the complete transcript through a standard HTTP callback. That's a sensible middle path for products that want live feedback during the session but still want a clean, complete artifact afterward.

Post-meeting delivery is simpler because it asks less of the architecture. A single webhook fires once the transcript job is fully done: one event, one delivery, no need to reconcile a stream of partial updates. AssemblyAI's asynchronous model sets a webhook_url at the moment the job is submitted (a decision made per transcript, not through some global registration), and the payload that eventually arrives contains nothing but a transcript_id and a status; the receiver fetches the actual content afterward with a separate GET request. Zoom's Cloud Recording pipeline behaves differently again Zylos AI Research.

The decision rule that falls out of all this is not complicated. If the product needs to act during the meeting, real-time delivery is a requirement, and the added architectural weight (WebSocket management, partial-segment handling, session lifecycle tracking) comes with the territory. If the product only cares about the finished artifact, post-meeting asynchronous delivery is simpler to build, simpler to reason about, and sufficient on its own. For Nylas Notetaker, processing after a meeting ends typically takes 2–5 minutes depending on recording duration, and the webhook fires when processing is complete Nylas Notetaker API.

Configuring endpoints across the major transcription platforms

In the Python SDK, the pattern is to build a TranscriptionConfig, call set_webhook(url) on it, and submit the job with submit() rather than transcribe(), since submit() returns immediately instead of blocking until the job finishes. Because the webhook URL is attached per transcript rather than registered globally, there's no central webhook registry to maintain; each job simply carries its own callback destination.

Recall.ai takes a bot-based approach across meeting platforms. A webhook_url gets passed at the moment a bot is created, or configured through the dashboard instead, and bot status change events are delivered through Svix's infrastructure. Real-time transcription isn't automatic; it has to be explicitly enabled in the bot's recording_config, specifying which events, such as transcript.data and transcript.partial_data, should be in the realtime_endpoints array. The appeal of this approach is breadth: a single API abstracts Google Meet, Zoom, Microsoft Teams, Slack, and others, so a developer isn't rewriting integration code for each platform separately.

Out of five available trigger events, two key ones developers commonly subscribe to are notetaker.media and notetaker.meeting_state. When notetaker.media fires with state marked as available, the payload carries URLs for the recording, the transcript, a summary, and any action items pulled from the conversation.

MeetingBaaS configures webhooks with a webhook_url, an array of subscribed events (meeting.started, meeting.completed, meeting.failed, transcription.available), a shared secret, and an enabled flag. It integrates with Gladia's Whisper-Zero ASR, tuned for the kind of messy audio real enterprise meetings actually produce.

Zoom offers two genuinely different products under one brand name, and conflating them is a common mistake. RTMS is the other product entirely: the flow starts with an RTMS event arriving by webhook, which prompts the developer's server to open a WebSocket to Zoom's media servers, from which live per-participant audio, video, transcripts, screen share, and participant events all stream directly. It covers Meetings, Webinars, Video SDK sessions, and Zoom Contact Center, ships with a Node.js SDK supporting webhook-based, client-based, and singleton connection approaches, and, again, requires no bot sitting inside the meeting.

Scriberr, an open-source reference point in this space, added an authenticated webhook configuration UI along with a CRUD API in a September 2026 pull request, giving it versioned lifecycle events like recording.uploaded, transcription.completed, and summary.failed, each carrying a schema_version field for safe evolution over time. Developers can set webhook_url in the body of POST /v2/transcript at the base URL https://api.assemblyai.com, with authentication via a plain authorization header and no Bearer prefix. It supports over 30 languages via Recall Transcription. It also functions as an ingestion-simplification layer on top of Zoom RTMS for teams that want bot-free infrastructure access without building raw RTMS integration. Nylas Notetaker API. The bot joins at meeting start, captures video and audio, and produces diarized transcripts that are speaker-labeled, powered by AssemblyAI integration. A critical operational detail is that media URLs in notetaker.media webhooks expire after 60 minutes, so files must be downloaded immediately on receipt, and if later access is needed, the Download Notetaker Media endpoint should be used for fresh URLs Nylas Notetaker API. For endpoint verification, Nylas sends a GET request with a challenge query parameter, and the server must return the exact challenge value in a 200 OK within 10 seconds Nylas Notetaker API AssemblyAI. Zoom Cloud Recording (native). For recording.completed, the file_type attribute values include TRANSCRIPT, CC, and TIMELINE. The recording.transcript_completed event requires a paid Zoom plan, specifically Pro, Business, Education, or Enterprise. Processing delay can reach up to 24 hours under load, making it unsuitable when near-real-time delivery matters Zylos AI Research. Zoom RTMS is real-time and infrastructure-level.

Payload contents and handling steps

The most important distinction to internalize about webhook payloads in this space is that the webhook tells the receiver when something happened; it doesn't necessarily tell it what happened. AssemblyAI's asynchronous payload is the sparest example of this pattern. It contains exactly two fields: transcript_id and status, either completed or error. There's no transcript text in that payload, and no error detail either; both require a follow-up GET request to /v2/transcript/{transcript_id}. The webhook is a notification, not a delivery mechanism, and building a receiver that expects otherwise is a fast way to end up confused about why the transcript field is missing.

Nylas takes a richer approach but pairs it with a catch: the receiver has to actually retrieve the linked resources before they expire. The notetaker.media payload includes URLs for the recording, the transcript, a summary, and action items directly, sparing the receiver a follow-up call in most cases. MeetingBaaS and Scriberr both include an event type field and a schema version in their payloads, letting a receiver branch cleanly on event type and validate against a known schema before doing anything else with the data.

First, verify that the request is actually authentic, a step detailed in the next section; second, acknowledge receipt with a 2xx response immediately, without doing any real processing inline; third, drop the transcript_id or event data onto a queue for asynchronous handling; and fourth, let a separate worker process fetch the full data, run whatever processing the application needs, and update state from there.

That second step carries more weight than it might first appear to. Processing a transcript inline, parsing it, running it through a summarizer, writing it to a database, all before responding, risks blowing past that window, and a timeout reads to the sending platform as a failed delivery. That failure triggers a retry, and the retry arrives while the first request might still be finishing its work, a duplicate-processing bug that's painful to track down after the fact. URLs expire after 60 minutes, meaning immediate download is not optional but required Nylas Notetaker API. Step 2 matters so much because the endpoint must return a 2xx within 10 seconds under AssemblyAI's limit, Nylas verification also requires a response within 10 seconds, and processing inline risks timing out and triggering retries Nylas Notetaker API.

Verifying that a webhook request is authentic

Skipping verification isn't a shortcut, it's a vulnerability. A spoofed "transcription completed" event could push an application to fetch data that was never actually processed, or corrupt state, or fire off a downstream action that should never have run. None of this is exotic; it's the ordinary risk that comes with running an endpoint that accepts arbitrary internet traffic.

AssemblyAI's pre-recorded transcription webhooks don't use an HMAC signing scheme, relying instead on three complementary controls. A custom auth header can be set at submission time through webhook_auth_header_name and webhook_auth_header_value, checked on every incoming request. Requests can also be restricted at the network layer by allowlisting AssemblyAI's known source IPs, 44.238.19.20 for US traffic and 54.220.25.36 for EU traffic. And the payload's webhook_status_code field offers a further signal to inspect. None of these three is sufficient on its own, and using more than one control together is recommended.

AssemblyAI's Voice Agents and streaming sessions use a proper HMAC scheme instead. Every delivery is signed using a secret the developer chooses at setup, one the API never returns again once it's been created. The scheme, labeled v1, computes a hex-encoded HMAC-SHA256 over the string formed by the timestamp followed by a period and then the raw bytes of the request body, keyed with the subscription secret. A receiver has to recompute that same signature and compare it against what arrived, using a constant-time comparison rather than ordinary string equality, since naive comparison leaks timing information an attacker can exploit to guess the correct signature byte by byte. Any delivery whose timestamp sits more than 300 seconds from the receiver's own clock should be rejected outright, since that window is the entire defense against an attacker replaying a captured, previously valid request.

Nylas documents its own signature verification separately, and its endpoint registration flow adds a challenge-response step: returning the exact challenge value in a 200 OK closes off the possibility of a spoofed endpoint being registered in the first place. MeetingBaaS accepts a secret field as part of its webhook configuration for the same purpose.

Recall.ai sidesteps a chunk of this work for bot status webhooks by routing delivery through Svix, which builds enterprise-grade signature verification into the delivery infrastructure itself, so a developer consuming those events is working with an already-verified stream rather than hand-rolling the cryptography. Verification is non-negotiable because an unverified endpoint can be triggered by anyone who discovers the URL, and spoofed events can corrupt application state or trigger unintended actions. General best-practice layering follows the May 2026 guidance in the research. HMAC signature verification involves validating the SHA-256 hash in the request header using constant-time comparison. IP whitelisting restricts inbound traffic to the provider's published IP ranges. Schema validation enforces a JSON Schema on every payload before processing.

Retry behavior across providers and the duplicate-delivery problem it creates

Delivery failure is the norm to plan around, not the exception. Roughly 15% of initial webhook delivery attempts fail on the first try. Networks drop packets, receiving servers restart mid-request, load balancers time out connections that were about to succeed. None of that is a defect in the platform sending the webhook; it's just what happens at scale across the open internet, and every provider covered above builds retry logic around that fact rather than pretending it away.

The problem retries solve is also the problem they create. A provider that resends a webhook after a timeout or a non-2xx response has no reliable way of knowing whether the receiver's earlier attempt actually processed the request before failing to respond in time. The receiver might have written the transcript to a database, fired off a notification email, and only then hit whatever caused the timeout, so the retry arrives to redo all of it. The fix has to live on the receiving side, since the provider's retry behavior isn't something a developer can turn off simply by asking nicely. That means treating every incoming event as a key to check against before acting, so a second delivery of the same event gets acknowledged with a 2xx rather than triggering retries. It's a small piece of bookkeeping, but it separates an integration that merely works most of the time from one that holds up under the ordinary, unavoidable unreliability of pushing data across the internet.

Sources

  1. Transcription Webhooks & Callbacks: The Complete Guide
  2. feat: add webhooks by jordanmarchetto · Pull Request #487 · rishikanthc/Scriberr
  3. assemblyai.com
  4. meetingbaas.com
  5. recall.ai
  6. activepieces.com
  7. developers.zoom.us

More in Meeting Data Pipelines