
Speechmatics is now live on Zapier, connecting speech-to-text to 8,000+ apps with no code.
Two actions: Get Transcript from Audio (synchronous) and Submit Audio (Webhook) for longer files (asynchronous).
Configure speaker diarization and output format (text, JSON, or SRT) directly inside the Zap step.
Runs on the Batch API, so Melia's multilingual code-switching transcription is available out of the box.
Currently in beta — a native "transcript ready" trigger is coming.
Say it. Zap it. Ship it.
Plenty of things can stall a voice-powered idea. One of the most avoidable is glue code: the webhook you have to write to move a transcript into Slack, the polling script that checks if a job's finished, the custom endpoint that pushes call notes into your CRM. None of that is the interesting part of the build, but it's often where the time goes.
We're closing that gap. Speechmatics is now live on Zapier, connecting our speech-to-text engine directly to the 8,000+ apps builders already use, no API client, no server, no code required.
Zapier exists to remove the plumbing between tools. Bringing Speechmatics into that ecosystem means the same industry-leading accuracy that powers production voice AI systems is now available to anyone who can drag a trigger into an action, no engineering sprint required.
Concretely, that means:
Ship in minutes, not sprints. Wire up a transcription workflow the same afternoon you think of it, without provisioning infrastructure or writing integration code.
Prototype before you commit. Validate a voice-powered workflow with Zapier first, then graduate to the API directly once it's proven out, with no throwaway work in between.
Put transcription wherever your team already works. Speechmatics output can land in a spreadsheet, a CRM, a chat channel, or a ticketing system without a single custom integration.
Free up engineering time for the hard problems. Let the people closest to a workflow (ops, support, product) build and iterate on it themselves, instead of queuing a request for engineering.
Speechmatics plugs into Zapier as two actions, the steps that run after something else kicks off your Zap.
The two actions:
Get Transcript from Audio: synchronous. Pass an audio URL, get the transcript back immediately, inside the same Zap. Built for shorter clips where you want the text right away, ready to drop straight into the next step.
Submit Audio (Webhook): asynchronous. Submit an audio URL and a webhook URL, and Speechmatics posts the finished transcript to that webhook once it's done. Built for longer recordings, where you don't want a Zap hanging around waiting for a job to finish.
Because there's no native trigger for "transcript ready," the webhook version needs a second, linked Zap that picks up the POST once Speechmatics delivers it, then carries on with whatever happens next.
Inside either action, you can configure:
Speaker diarization (on/off): separates the transcript by speaker, so downstream steps know who said what, not just what was said.
Output format: plain text, JSON (with word-level timing and confidence), or SRT subtitles, so the transcript lands in your next app already shaped the way you need it.
Connect a trigger's file URL to the Audio URL field, configure diarization and output format, test the step, and publish. That's the whole integration.
The integration is currently in beta, so expect the details, including a native trigger, to keep evolving as we roll it out further.
Because almost anything can be the trigger, the workflows aren't limited to one team or use case:
Meeting and call notes: a new Zoom, Google Meet, or Teams recording lands, gets transcribed, and a summary is auto-posted to Slack or dropped into a Google Doc or Notion page. A sales call recording gets transcribed and appended as a note on the matching HubSpot or Salesforce record.
Content repurposing: a new podcast episode or webinar recording is transcribed and turned into a draft blog post in WordPress (or another CMS), or into social snippets. A video upload comes back as SRT output and the captions get attached to the video asset automatically.
Customer support and compliance: a support call is transcribed with diarization intact and logged to Zendesk, Intercom, or another ticketing system with speaker labels preserved. Regulated teams in finance, healthcare, or legal archive searchable transcripts in Drive or SharePoint to build an audit trail.
Research and qualitative workflows: user interview recordings come back as diarized transcripts and get pushed into a research repo (Notion, Airtable, or a Dovetail-style database), tagged by speaker.
Accessibility: any new audio or video content generates SRT subtitles automatically as part of the publishing pipeline, rather than as a manual step someone has to remember.
Voicemail and IVR: voicemail audio becomes a plain-text transcript and gets emailed or Slacked to the right person, instead of them having to dial in and listen.
Long-form and batch processing: long recordings, earnings calls, all-day workshops, get submitted via the webhook action, with a separate Zap catching the result and triggering summarization, distribution, or storage once processing completes.
That's the point of putting speech-to-text on Zapier: the integration doesn't assume a use case. It just removes the cost of trying one.
Wiring up a Zap is the easy part — picking the engine underneath it is the decision that matters. Here's how the leading speech-to-text APIs compare on accuracy, languages, and deployment.
Over half the world speaks more than one language, and real conversations move between them mid-sentence. That's exactly what our Melia model is built for: a single model that handles code-switching across all 56+ languages we support, so a recording that drifts from English to Spanish to Mandarin and back comes home as one continuous transcript, no language to pick in advance, no separate model per market.
Because the Zapier integration runs on our Batch API, it's a natural home for Melia: drop in a multilingual recording from any trigger app, and Speechmatics transcribes it in one pass rather than routing it through a stack of single-language models.
That matters most for teams who weren't well served by voice AI before: global contact centres, multilingual broadcast captioning, and compliance teams reviewing calls across markets, where a code-switched conversation used to mean a gap in the transcript, not just an accent to handle.
Melia Real-Time is dropping soon, so keep an eye on our feed!
A no-code front door doesn't mean a lighter-weight engine. Every Zap runs on the same speech-to-text behind Speechmatics' API and SDKs, with diarization, flexible output formats, and multilingual transcription available as first-class capabilities, not afterthoughts.
Read the full setup guide: Zapier integration docs
With 93% accuracy, our new model is twice as good as the nearest competitor.
![[alt: Speaker lock blog image]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2Fue7vVoLyWYL8hohG7JNz6%2F0df9783d84d0e93b5b04f8ecbe33d8f0%2FSpeaker_lock-blog-v1_-_Header_16-9.webp&w=3840&q=75)
Because in the real world, conversations are messy, and Voice AI needs to keep up.

The move brings speech recognition directly into CATalyst VP eliminating the challenges of running multiple applications, making it the first seamless solution for the voice reporting industry.
A founder's account of taking a hackathon side-project live with 30 sales teams, and the four real-time integration bugs that stood between a demo and a product people could trust.
![[alt: Pattern of blue coins with "Cr", circuit symbols, and sparkles on a dark teal background, creating a crypto-themed design.]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2F5wtHENHGzT1GK6tG4M6USf%2Fa380442e95ea4f30ad0f5e6c42ea7ee8%2Fpricing-change-wide-carousel_2x_1.webp&w=3840&q=75)
From 1 August 2026, Speechmatics moves to credit-based billing: one credit balance across every product, same prices, no action needed.