Deepgram vs AssemblyAI for Meeting Transcription Accuracy
Speaker attribution matters more than word accuracy when choosing a meeting transcription system.

Raw word error rate tells you almost nothing useful about whether a meeting transcript will hold up once someone acts on it. The real test of a meeting transcription system runs through three dimensions WER was never built to measure: who said what, how the system holds up once the room gets loud or the language switches mid-sentence, and whether it can be told in advance who's in the meeting and what they're likely to say.
The wrong scorecard for meeting transcription
Picture a sales call transcript with a low word error rate, a number that would make most engineering teams comfortable shipping. Now look closer: the system has attributed the client's pricing objection to the account manager. The summary that gets generated from that transcript, and the follow-up email that gets drafted from that summary, both carry the error forward. The WER score never caught it, because WER treats every token the same way. "The" and a participant's name count identically in the formula, even though losing one costs nothing and losing the other costs a deal.
This is the structural blindness built into word error rate. Catching that requires a speaker-aware scoring method called cpWER. The next two sections unpack it in full.
A model can post a strong aggregate WER while its mistakes cluster on exactly the words a meeting depends on, the names, the dollar figures, the dates, the commitments that get copied into a CRM or a follow-up email. A transcript can get the overwhelming majority of its words exactly right and still misrepresent the small share that mattered. That asymmetry is why some researchers have started building semantic WER, a newer evaluation approach that uses a large language model as a judge of whether meaning survived. The gap between those two outcomes is the gap between a metric built for academic benchmarking and one built to tell you whether the transcript actually works.
The metrics that predict whether a meeting transcript is usable are attributed word error rate (cpWER), entity accuracy on names and numbers, and how the system holds up under real acoustic noise. All three get flattened out when a vendor reports a single aggregate WER number. That number is the wrong place to start an evaluation.
How benchmark audio differs from real meeting audio
Published accuracy numbers come from somewhere, and that somewhere is rarely a conference room. Benchmark datasets are built overwhelmingly from clean studio recordings or phone-quality audio, read in controlled conditions by speakers who know they're being recorded for a test. A real meeting looks nothing like that, and the conditions it introduces compound in ways no single number can summarize.
Real meetings stack several problems on top of each other at once. The audio itself arrives already degraded: Zoom, Teams, and Meet recordings pass through lossy compression before any speech-to-text model ever touches the file, so the model works from a lower-fidelity signal than the recordings used to build most published benchmarks. Accented speech, speakers switching languages mid-sentence, and the short, fragmentary, filler-laden style of actual conversation, so different from the complete sentences used in read-speech benchmarks, widen the gap between a leaderboard score and production performance fast.
AssemblyAI has a submission on the Open ASR Leaderboard; Deepgram does not appear on it. Where the two vendors' WER figures do show up on other independent benchmarks, they fall in a broadly similar range. None of those benchmark sets are built from meeting audio, so a leaderboard functions as a way to narrow the field of contenders worth testing further, not as a basis for a final decision.
What follows from this is simple to state and easy to skip in practice: evaluation has to run on the reader's own meeting recordings, not on either vendor's published headline figure. Accuracy and speed are properties of a complete workload: audio quality, number of speakers, accents in the room, the vocabulary being used, the background noise level, how the audio was encoded, and whether it's live or pre-recorded all shift the result. Two teams using the same vendor can get different outcomes depending on how their meetings actually sound.
Speaker diarization is where meeting transcription breaks or holds together
For a transcription system handling a meeting, "was this word transcribed correctly?" is the easier question. The harder one, and the one that determines whether the output is usable, is whether the word got assigned to the right person. Those two questions are measured by different metrics, and they diverge sharply once audio stops being clean.
Diarization Error Rate, or DER, is the metric most commonly quoted in academic comparisons, but it measures speaker segmentation in isolation from the transcript itself. A system can post a low DER by merging short speaker turns into whoever is dominating the conversation, since DER doesn't penalize that behavior the way a transcript-aware metric does. Concatenated minimum-permutation word error rate, cpWER, closes that gap. It scores transcription accuracy and speaker attribution together, penalizing both failures at once, which makes it a tougher standard to meet and a far more useful signal for anyone deciding whether a transcript can be trusted. The failure modes cpWER catches are the ones that actually cause damage: a participant's reported concern attributed to the facilitator, a client's objection credited to the account manager, an action item assigned to the wrong name on the call.
Diarization degrades in specific, predictable ways once meeting conditions get involved. A third problem appears specifically in compressed conferencing audio, where two participants of similar vocal range get mixed together because the codec has already flattened the spectral detail a speaker model relies on to tell voices apart.
Phantom Farm treats diarization as the central engineering problem in meeting transcription rather than a feature bolted onto a transcription engine after the fact. Its approach to speaker attribution is built around sustaining identity across an entire meeting rather than re-identifying speakers turn by turn, which is the design choice that matters most for the short-turn and overlapping-speech failures described above: a system that has to re-derive who's speaking from a two-second clip of "exactly" will fail far more often than one that carries forward a running model of each participant's voice built from the full length of the call. Any reader evaluating Phantom Farm should test a recording with several speakers, at least one rapid back-and-forth exchange, and at least one moment of true overlap, scored on cpWER.
AssemblyAI's diarization keeps a rolling memory of the conversation by default, so the model retains context from earlier in a call, which keeps speaker identity consistent across a long meeting, and the same category of error occurs in business meetings whenever a client's words get attributed to the account team.
None of this settles which vendor wins. It settles what the test has to measure. Run both APIs against a recording with at least four speakers, at least one exchange built from short turns, and at least one overlapping segment, and score the result on cpWER. A vendor that performs well on a clean two-person demo call tells you almost nothing about how it will hold up on the kind of audio real meetings actually produce.
How noisy audio separates benchmark performance from production performance
Noise is closer to the default operating condition for meeting transcription, and it's the dimension on which two models that look nearly identical on clean audio tend to pull apart most sharply.
Meeting noise takes several forms, each with its own effect on a transcription model. Room echo and reverb occur in in-person recordings captured with a single microphone picking up a whole conference room. Codec artifacts come from the conferencing platform itself: Teams compresses audio at 16 kHz, Zoom at 32 kHz, Meet at 48 kHz, and each of those compression rates strips information out of the signal before any speech model gets a chance to process it. Hybrid meetings compound all of it at once, since a single recording might concatenate a crisp in-room feed with a remote participant's compressed, bandwidth-limited connection, producing wildly different channel quality within the same file.
Accuracy shifts dramatically the moment accented speech, background noise, code-switching, and the general messiness of real conversation enter the picture, and that shift applies as much to a meeting recording as it does to any other production audio a speech model is asked to handle.
Testing for this honestly means building a representative sample. Use recordings from three distinct channel types: an in-person room microphone, a laptop microphone on a video call, and a phone dial-in. Score entity capture, meaning whether named participants, action items, and numbers come through correctly, separately from overall word error rate, because that separation is what reveals where each model's mistakes actually concentrate. A model with an unremarkable overall WER might still catch every name and number correctly, while a model with a better aggregate score might be losing exactly the entities that matter. Running the same audio set, the same normalization rules, and the same error analysis across whichever models a team is actually considering is the only way to get a result that means anything; a published test design from a vendor or a blog can help structure that process, but it can't substitute for a controlled test on a team's own recordings.
Code-switching and multilingual meetings as a growing accuracy differentiator
Code-switching, meaning speakers alternating between languages within a single sentence or across a meeting, is the condition that breaks models trained and tested on monolingual benchmarks most completely, and it's an increasingly ordinary feature of international business meetings.
The difficulty here goes beyond simply handling two languages. A model configured for one language will try to force the phonemes of the other language into its existing vocabulary, which tends to produce garbled text, invented words, or silent deletions rather than any honest signal that the model is uncertain. Standard WER benchmarks make this almost impossible to spot in advance, because they're built to be monolingual by design, so a model's code-switching performance stays largely invisible in its leaderboard score on mainstream monolingual tests.
Phantom Farm's handling of multilingual meetings is built around the same principle that shapes its diarization work: sustaining context across the full meeting rather than re-evaluating language and speaker identity sentence by sentence. For meetings where participants switch between languages mid-sentence, that continuity is what keeps a language switch from also triggering a false speaker change, since the system isn't re-deriving who's talking and what language they're using from an isolated clip. Teams running multilingual meetings should test this directly against the language pairs they actually use in practice, rather than relying on documentation describing supported languages in the abstract.
Evaluating either AssemblyAI or Deepgram on this dimension means checking the specific model, language, locale, and feature combination in question, because multilingual capability on both platforms is model- and plan-specific. Accuracy on code-switching and accented speech swings widely enough that AssemblyAI's own guidance points toward verifying support against the specific model tier, Universal-3.5 Pro versus Universal-3.6 Pro Realtime, rather than assuming identical performance across both. The only reliable way to know how a given system handles a specific language pair is to test it with real recordings drawn from that pair, not the vendor's published language list.
Context-injection and domain vocabulary handling for meeting-specific accuracy
A generic transcription model and a meeting-ready one can post similar raw WER scores and still behave completely differently once a meeting introduces its own vocabulary. The distinction that actually separates them is whether the model can be told, in advance, who's in the room and what they're likely to say.
Product names, internal project codenames, client organization names, and technical jargon are all phonetically rare from a model's point of view. Without some prior signal telling it to expect those words, a model defaults to the most statistically common word that sounds similar, which is how a client's product name turns into a generic phrase that means nothing to anyone reading the transcript later. Participant names carry the same risk in a more personal form: "Sean" against "Shawn," "Reza" against "Resa," and any name drawn from outside the model's dominant training data are all vulnerable to being guessed wrong, and a wrong name at the top of a transcript can cascade into every downstream attribution built on top of it.
This is the architectural layer that separates a transcription engine built to transcribe audio in general from one built to handle a specific meeting. Deepgram's Keyterm Prompting, available on its Nova-3 and Flux configurations, lets a team supply the names and terms a given call is likely to include before the audio is processed, directly addressing the phonetic-rarity problem for named participants and project terms. The value of this kind of context injection is the gap between a transcript that spells a client's name correctly once and carries that correct spelling through every reference afterward, and one that renders it a different way each time it comes up. Any team evaluating a vendor for recurring meetings with the same participants and the same recurring vocabulary should test this capability specifically, feeding in the actual names and terms that will appear, rather than judging accuracy on vocabulary the model would never encounter in production.
Sources
- Word Error Rates (WER) for AI Transcription: What Do They Tell Us? - University Transcription Services
Provided background on how word error rate is calculated and its structural limitations for evaluating practical transcription quality.
- Evaluation of Automatic Speech Recognition Using Generative Large Language Models
Supported the discussion of using large language models as semantic judges of transcription accuracy, informing the article's mention of semantic WER.
- MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
Informed the treatment of multiparty conversation evaluation, including short-turn and overlapping-speech failure modes in meeting transcription.
- Towards Quantifying Benchmark Optimization in ASR Models
Supported the article's discussion of how published benchmark scores can diverge from real-world production performance for ASR systems.
- Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
Provided evidence on how code-switching in multilingual speech exposes gaps in models trained and evaluated on monolingual benchmarks.
- A Review of Speaker Diarization: Recent Advances with Deep Learning
Provided technical grounding for the discussion of speaker diarization metrics, including DER and cpWER, and their respective failure modes.
- Best Practices for building Meeting Notetakers - AssemblyAI
Informed the section on AssemblyAI's rolling conversational memory for maintaining speaker identity across long meetings.
- Changelog
Supplied details on AssemblyAI's specific model tiers referenced in the article, including Universal-3.5 Pro and Universal-3.6 Pro Realtime.


