
Melia hits a 206x speed factor on Artificial Analysis's independent leaderboard — ahead of AssemblyAI, ElevenLabs, Soniox, Gladia, Amazon and Rev AI.
A full hour of audio comes back in under 20 seconds (30 minutes in under 10, 10 minutes in under 5) — in one pass, not chunked.
That turnaround unlocks near-real-time workflows for contact centers, compliance, healthcare, education, media and finance, without trading away accuracy or re-architecting for streaming.
Switching is a one-line config change: set model to melia-1 and language to multi.
Some audio processing can't wait. The moment the recording is finished, someone is waiting on that transcript. A clinician confirming a session summary before the patient has left the room. A video editor cropping footage the second it's imported. A compliance reviewer flagging a call before the next one starts. A CCaaS platform pushing post-call analytics back before the agent takes the next call.
And the wait is getting more expensive. Workflows keep getting more complex with summarization, sentiment, and other AI tooling stacked on top of the transcription, each step taking valuable time. You can't keep absorbing that and still give your end customer the experience they expect. Melia takes the processing time you've historically had and shrinks it, at the same quality, so you can add functionality without the workflow or the customer paying for it.
Once a recording is finished, getting the transcript back fast has historically meant picking between options that don't quite fit:
Drop to a lower-quality model for the sake of speed. You get the turnaround, and you pay for it in the accuracy everything downstream depends on.
Micro-batch it yourself. Split the recording and stitch the snippets back together, and you lose valuable information like speaker attribution — plus the workflow gets more complex.
Re-architect around streaming. Standing up a persistent real-time connection, just to get the final segment back within a second post recording, is a big architectural shift.
None of these are really feasible since each one trades accuracy, engineering time, or your architecture for speed.
Melia solves this. Melia outpaces AssemblyAI, ElevenLabs and Soniox and others on Artificial Analysis's own leaderboard while covering all 55+ languages in one model, so a multilingual transcription doesn't need a language selected up front.

Artificial Analysis independently and continuously benchmarks speed factor (audio seconds transcribed per second of processing) across dozens of speech-to-text models.
Melia is at a speed factor of 206x, ahead of AssemblyAI (up to 105x), Soniox (up to 37x), ElevenLabs (up to 55x), Gladia (up to 81x), Amazon (up to 18x), Rev AI (up to 13x) and others.
Melia processes a recording in one pass rather than chunking it, so the turnaround holds up on a single long file, not just on short clips.
The difference shows up the moment you stop needing to plan around a wait.
Where it matters | What's now possible |
|---|---|
Contact center analytics | Score a call before the agent logs the next one |
Compliance monitoring | Flag a non-compliant call before the next one starts |
Podcast and audio production | Hit a publish deadline without a processing buffer |
Video and content editing | Start editing on import, not after a wait |
Healthcare consultations | Confirm the summary before the patient leaves the room |
University lectures | Turn a lecture into notes and captions before students leave the hall |
Earnings calls | Quote the call in a note the same hour it ends |
Contact center analytics. Scoring has always been retrospective, a sample of calls reviewed hours or days later. Melia returns the transcript in seconds, so downstream processing like sentiment analysis, turn detection and topic extraction can run on every call, while the agent still remembers it.
Compliance monitoring. A bank transcribing every call to catch the non-compliant ones doesn't get to skip the calls that switch from English into Spanish or Arabic mid-conversation. Melia returns each as one labelled transcript, fast enough for a reviewer to flag a policy violation before the next call starts.
Podcast and audio production. A podcast that has to go out on schedule doesn't have minutes to spare on transcription after the episode wraps. Melia returns a full episode's transcript, captions and show notes fast enough to hit that publish time without building in a processing buffer.
Video and content editing. A creator importing raw footage doesn't want to wait three or four minutes before they can start editing through the transcription output. With Melia, that wait disappears. The transcript is back in seconds, so editing starts on import.
Healthcare consultations. There's no gap between appointments to sit and wait on a transcript. Melia transcribes the full consultation in one pass, so an hour-long session comes back in under 20 seconds, fast enough to check the summary before the patient leaves.
University lectures. Lecture capture usually means the recording lands first and the transcript follows, overnight if automated, days if outsourced. Melia returns an hour-long lecture in under 20 seconds, which can be used to generate captions, summaries, chapters and interactive transcripts for the student, in whatever language it was taught in.
Earnings calls. The hour after a call is when it matters, and that's exactly when analysts are working off their own notes. Melia returns the full call in under 20 seconds: searchable, quotable, and easy to compare against last quarter.
Switching over is a config change, not a new integration. In your transcription config, set model to melia-1 and language to multi. Melia covers all 55+ supported languages, so there's no language pack to pick, but if you know which ones to expect, language hints nudge recognition toward them.
Jobs are processed asynchronously by default: submit, then check status or wait for a webhook notification. If you'd rather block for the result in a single request, see synchronous transcription in the docs. Runnable examples are in the GitHub Academy: Melia Multilingual for the model itself, and Call Analytics for a full contact-center build.

Speechmatics is now live on Zapier. Connect industry-leading speech-to-text to 8,000+ apps with no code.
With 93% accuracy, our new model is twice as good as the nearest competitor.
![[alt: Speaker lock blog image]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2Fue7vVoLyWYL8hohG7JNz6%2F0df9783d84d0e93b5b04f8ecbe33d8f0%2FSpeaker_lock-blog-v1_-_Header_16-9.webp&w=3840&q=75)
Because in the real world, conversations are messy, and Voice AI needs to keep up.

The move brings speech recognition directly into CATalyst VP eliminating the challenges of running multiple applications, making it the first seamless solution for the voice reporting industry.
A founder's account of taking a hackathon side-project live with 30 sales teams, and the four real-time integration bugs that stood between a demo and a product people could trust.
![[alt: Pattern of blue coins with "Cr", circuit symbols, and sparkles on a dark teal background, creating a crypto-themed design.]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2F5wtHENHGzT1GK6tG4M6USf%2Fa380442e95ea4f30ad0f5e6c42ea7ee8%2Fpricing-change-wide-carousel_2x_1.webp&w=3840&q=75)
From 1 August 2026, Speechmatics moves to credit-based billing: one credit balance across every product, same prices, no action needed.