Processing modes
Processing modes
Get the very best performance whether you're transcribing pre-recorded files or live audio.
Deployment options
Run Speechmatics wherever your infrastructure and compliance needs requireDeployment options
Choose how Speechmatics runs — in the cloud, on your own infrastructure, or directly on-device — deployed however suits you.
On-Prem
Meet architecture, security and compliance needs by hosting our API in your own environment. Deploy with Docker containers, Kubernetes, or a preconfigured virtual appliance.
On-Device
Run Speechmatics directly on your devices for ultra-low latency and maximum data privacy — nothing sent over a network. Get cloud-parity accuracy (within 5%), full speaker diarization, and sub-second latency, on Mac or Windows. Ideal for use cases where connectivity is limited and data must stay local.
Models
Models
Every model runs on the same API. Pick the one that matches your priorities.
Melia 1: automatic language detection, labelling, and seamless code-switching in one API, across all 56+ supported languages. Available now for batch transcription, with real-time support on the way.
Model overview
Model | Best for | Real-time | Batch | Multilingual |
|---|---|---|---|---|
Enhanced | Highest accuracy, one language per file or stream | ✓ Yes | ✓ Yes | ✗ No |
Standard | Strong accuracy with faster turnaround and lower cost | ✓ Yes | ✓ Yes | ✗ No |
Melia 1 | Audio that switches between languages mid-conversation, across all 56+ supported languages | Coming soon | ✓ Yes | ✓ Yes |
Feature support by model
Feature | Enhanced | Standard | Melia 1 |
|---|---|---|---|
Speaker & channel diarization | ✓ Yes | ✓ Yes | ✓ Yes |
Word-level timestamps & punctuation | ✓ Yes | ✓ Yes | ✓ Yes |
Custom dictionary | ✓ Yes | ✓ Yes | Not yet |
Confidence scores | ✓ Yes | ✓ Yes | Not yet |
Translation, summaries, chapters | ✓ Yes | ✓ Yes | Not yet |
Code-switching across 56+ languages | ✗ No | ✗ No | ✓ Yes |
Diarization ships on every model, including Melia 1. It's one of the areas Speechmatics is strongest, and it's not something every multilingual model gets right.
See pricing for model costs, or the developer docs for configuration syntax.
Transcription Features
Everything you need to hit the highest accuracy possibleTranscription Features
Our customization options allow you to finely tune your set up to achieve high accuracy with even the most unique words and phrases.
Custom Dictionary
Boost accuracy for proper nouns, acronyms or industry-specific terms by providing a list of custom words.
Speaker & Channel Diarization
Track who said what and when with speaker labelling for each word, available for both batch and real-time transcription.
Numeral Formatting
Identify and correctly format numbers, dates and currencies automatically to improve transcript readability and enable effective post-processing.
Profanity & Disfluency Detection
Aid comprehensibility and compliance by detecting and optionally removing words that are considered profanities or hesitations.
File Formats
Minimize the resource needed to prepare audio or video files with support for all major audio and video formats along with automatic sample rate detection.
Advanced Features
Easily push a variety of media formats to the APIAdvanced Features
Easily push a variety of media formats to the API and get a rich set of metadata to support your post processing needs.
Confidence Scores
Collect confidence scores for every word in the transcript to enable efficient human review and editing.
Word Timings
Get accurate timestamps for every word in the transcript to allow for post-processing and improved end user experience.
Advanced Punctuation & Casing
Improve readability with language-specific capitalization and punctuation including commas, question marks and exclamation marks.
Languages
Partner with Speechmatics to maximize your total addressable marketLanguages
We deliver for multilingual, multicultural and multinational businesses, with coverage of 56+ languages across a range of dialects and accents.
Accuracy tuning extends to specific industries too, including finance terminology and a dedicated medical model reaching 93% accuracy with 50% fewer keyword errors than the next-best system.
Language Coverage
56+ languages, with code-switching at the word level.
Accents and dialects
Whether you need Brazilian Portuguese or Canadian French, we have you covered with a single language model that supports all associated accents and dialects.
Translation
Transcribe and translate audio to and from English for over 30 languages using a single API call.
Language Identification
Simplify integration and ensure accurate transcription with automatic detection of the language spoken.
Speech understanding
Speech understanding
Transcription pairs with summaries, sentiment, topics, and chapters in the same API call.
Translation
With automatic translation with a single API call, you can translate media and provide captions for over half the world’s population.
Certifications & compliance
Built to meet enterprise security and compliance requirementsCertifications & compliance
Speechmatics meets the security and compliance bar enterprises need, wherever you deploy. See speechmatics.com/security for full detail.
ISO/IEC 27001:2022
Accredited to the ISO/IEC 27001:2022 information security standard.
SOC 2 Type II
SOC 2 Type II-certified, independently audited on security controls over time, not just a point in time.
HIPAA
Fully compliant with the Health Insurance Portability and Accountability Act, for healthcare and clinical use cases.
GDPR
Compliant with GDPR and other regional privacy directives.
Encryption
Data encrypted at rest (AES-256) and in transit (TLS 1.2+).
Data residency
Processing environments in the US, EU, or Australia, to meet data sovereignty requirements.
Resources for features and deployments

Speechmatics versus Whisper: how Adobe Premiere's on-device speech engine got rebuilt
Quantization was the key to fitting a cloud-grade model on a laptop. Getting the full optimization chain to cooperate around it was the hard part.
![[alt: Orange gradient background with "Melia" centrally placed, highlighting multilingual support with code-like symbols scattered.]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2F1xWHUZrv79DXnCdWmyusmJ%2F20c92142a1007065d31f1d1b7656bc95%2Fmelia-WideCarousel-1200x480.webp&w=3840&q=75)
Introducing Melia, our new multilingual speech-to-text model
A multilingual speech-to-text model from Speechmatics, with code-switching across 56+ languages. Available today in production preview, starting with batch transcription.

The Adobe story: How we made cloud-grade AI work on your laptop
Behind the build: what it takes to make cloud-grade speech recognition work inside Adobe Premiere, and why Whisper raised the stakes.

How to build a microbatching workflow with the Speechmatics API
Build a cleaner path between batch and real time. Learn when micro-batching makes sense, how to chunk audio, submit jobs, stitch JSON, and scale safely with the Speechmatics API.

Adobe and Speechmatics deliver cloud-grade speech recognition on-device for Premiere
Adobe Premiere users can run the most accurate on-device transcription locally; efficient enough for a laptop, powerful enough for professional work.

Best speech-to-text AI guide: APIs, platforms and services compared
Speech-to-text has moved from novelty to enterprise infrastructure. Here's how the leading platforms stack up in 2026 — and how to pick the right one.
![[alt: Two healthcare professionals, wearing blue scrubs, engage in conversation in a hospital in Sweden]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2F6rSJH52PtxeumhJ0QFTQ6S%2Fd5e564d487d2972f63aa28d098e19bd3%2FImage_fx__4_-wide-carousel.webp&w=3840&q=75)
Speechmatics launches new Swedish medical model, cutting transcription errors by 40%
Expanding a Nordic medical lineup with 3.91% KWER model that delivers sub-second latency across Swedish, Finnish, Danish, and Norwegian clinical workflows.

What is Speaker Diarization and why does it matter in voice AI?
The breakthrough technology helping AI understand conversations like humans do.