Today marks an exciting new chapter for Speechmatics.
We’re bringing the expertise that has allowed us to lead the speech recognition market for more than a decade into the world of Voice AI and voice agents.
Voice agents enable companies to build responsive, natural speech interactions into their products. They combine accurate, low-latency speech recognition with speaker identification, helping technology understand interruptions, multiple speakers and conversations in noisy, real-world environments.
The range of potential applications is vast. Voice AI can power assistants, customer service agents, healthcare tools, educational products and any other experience that becomes easier when people can simply speak.
Why is Speechmatics moving into this space? Does it change our mission to understand every voice?
For us, the answer is simple. Our mission remains the same, but the possibilities for our technology are growing.

Why do we need speech interactions?
There is a simple and compelling reason to add speech to the ways we use technology: it is deeply intuitive.
We have been using our voices for 100,000 years, far longer than any tool, let alone a keyboard and mouse.
It is the default way that we, as humans, communicate with each other.
As technology becomes increasingly powerful, products that meet us where we are will ultimately be far more valuable than those that require specialist knowledge to use.
Technology we can work alongside to perform tasks and solve problems, without lifting a finger, could transform the way we work and play.
For many, our current methods of using technology seem straightforward. A touchscreen displaying information and images, or a mouse and keyboard used to complete tasks.
Simple enough.
Except, for millions, it is not.
Worldwide, there are millions of people for whom these options present significant barriers:
A World Economic Forum study in 2017 revealed that almost a quarter of respondents did not know how to operate a computer.
Three in ten Americans struggle to use the internet.
Ten percent of the world’s population has dyslexia. That is around 780 million people.
Forty million people are blind, with a further 250 million experiencing visual impairment.
Seven percent of working adults have dexterity issues and may struggle to use a keyboard and mouse.
Alongside the opportunity to give people a more natural way to interact with technology, Voice AI has the potential to create more inclusive products that reduce traditional barriers to adoption.
Many of the leading AI companies are beginning to implement speech as an interface to their intelligence.
People may already use Siri or Google Assistant, or have experimented with speaking to ChatGPT.
Some may believe that speech has already been solved.
They would be wrong.
An easy way to demonstrate this is to imagine the Turing test.
First proposed in 1950, the Turing test evaluates a machine’s ability to exhibit intelligent behaviour equivalent to, or indistinguishable from, that of a human.
Another way of thinking about it is that if a machine can convince someone that they are interacting with another person, the test has been passed.
It could be argued that we have reached this point with some text-based interactions. When chatting with an LLM, it can be easy to forget that there is no person on the other side.
But imagine trying to pass this test using speech.
Could someone conduct a conversation with an AI over the phone and come away believing they had spoken to another person?
We believe the answer is still no.
Live human conversation is messy.
Picture yourself sitting in a bustling coffee shop, messaging a group of friends.
The grammar of the interaction is well understood. Someone writes a sentence or two and sends it. You read their message, decide how to respond, write a few words and wait.
Each message usually represents a complete thought. The interaction moves slowly and steadily.
Now imagine sitting in that same coffee shop with the same friends, having an animated conversation.
It is faster, more dynamic and filled with subtle cues.
Face-to-face conversations contain emphasis, intonation, accents, slang, dialects, interruptions, crosstalk, pauses, stammers, laughter, background noise, sarcasm and more.
Everyone at the table can also recognise each other’s voices. In theory, you could all close your eyes and continue the conversation.
You can distinguish the people taking part from the background chatter elsewhere in the coffee shop.
These characteristics are deeply intuitive to humans, but incredibly difficult for technology to understand.
Imagine making a voice agent part of that conversation.
Would it blend naturally into the group?
Or would its basic functionality break down under the number of variables being thrown its way?
We still have a long way to go.
The speech-input stage of any Voice AI interaction is also one of the most important links in the chain.
Even the most advanced i
ntelligence and realistic synthetic voice cannot deliver a natural interaction if the system cannot understand what is being said.
Solving this challenge is fundamental to the widespread adoption of Voice AI.
Only when technology can pass the Turing test using speech will voice become a truly ubiquitous way to interact with products and services.
Watch Speechmatics and Pipecat demonstrate a responsive voice agent that can handle interruptions, identify speakers and keep pace with real human speech.
So, how should we build the next generation of Voice AI?
Our approach builds on the strengths that make Speechmatics a leading choice for speech-to-text: high accuracy, powerful real-time performance and flexible deployment options.
We believe three principles should guide its development.
Voice AI should respond in the way we expect another person to respond.
That means actively listening, understanding and replying without long or unnatural pauses.
Low latency is vital, but responsiveness also means recognising interruptions, understanding when someone has finished speaking and responding appropriately to how something has been said.
For Voice AI to move beyond novelty, it must be usable by everyone.
Unlike previous generations of interfaces, this is not a question of training people to use the technology. The responsibility sits with the technology to understand the person speaking.
Given the rich variety of human voices, that means understanding different languages, accents, dialects, speech patterns, intonations and ways of expressing meaning.
Voice AI cannot be built around a single supposedly standard voice.
For Voice AI to become widely used, organisations need to be able to integrate it confidently into their products.
Businesses have different privacy, security, infrastructure and integration requirements. There cannot be a one-size-fits-all approach to the speech technology that underpins these experiences.
An educational assistant, for example, may need to access specific course materials and student information without exposing that information more broadly.
Providing flexible, enterprise-ready deployment options is therefore essential. Without them, the number of companies able to build speech-powered products will remain limited.
Voice AI can perform impressively in a controlled demonstration. Real-world speech presents a much harder test.
People interrupt each other. They speak over background noise, use unfamiliar terms and switch between speakers. They pause, hesitate, laugh, change direction and leave sentences unfinished.
For a voice agent to respond intelligently, it must first understand what was actually said.
Speechmatics’ real-time speech recognition is designed to understand speech as it happens. This helps voice agents respond quickly, handle interruptions and keep pace with natural speech.
It can also provide powerful speaker identification and diarization, helping systems understand who is speaking and when.
Depending on the use case, a voice agent may need to address multiple people, follow a particular speaker or distinguish the main interaction from surrounding voices.
Outside a controlled environment, these capabilities are essential. Voice experiences should not break down because someone interrupts, another person speaks nearby or background noise enters the audio.
Accuracy is equally fundamental. Voice AI must understand a broad vocabulary spoken across different accents and dialects because there is no single average voice.
Together, accuracy, speed and speaker awareness create the listening foundation that natural Voice AI depends on.
See how real-time speaker diarization helps Voice AI distinguish between different people and follow multi-speaker interactions.
CTA: Watch the speaker diarization demo
Our ambition is clear: to help create Voice AI that can pass the Turing test with speech.
Reaching that point will require significant breakthroughs across the voice technology stack. But accurate listening remains the foundation.
Even the world’s greatest intelligence and most realistic voice synthesis cannot deliver a convincing experience if the system mishears a word, misses an interruption or responds to the wrong speaker.
Every improvement we make to real-time speech recognition brings Voice AI closer to understanding people as intuitively as people understand one another.
This ambition is also consistent with our mission to understand every voice. Advances made for Voice AI strengthen our speech recognition, while advances in speech recognition create better voice experiences.
The journey has already begun.
We cannot wait to see what developers build next.
Build responsive, seamless and inclusive voice experiences with technology designed to understand every voice.
CTA button: Build Voice Agents
With 98% accuracy, our new model is twice as good as the nearest competitor.
![[alt: Speaker lock blog image]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2Fue7vVoLyWYL8hohG7JNz6%2F0df9783d84d0e93b5b04f8ecbe33d8f0%2FSpeaker_lock-blog-v1_-_Header_16-9.webp&w=3840&q=75)
Because in the real world, conversations are messy, and Voice AI needs to keep up.

The move brings speech recognition directly into CATalyst VP eliminating the challenges of running multiple applications, making it the first seamless solution for the voice reporting industry.
![[alt: Pattern of blue coins with "Cr", circuit symbols, and sparkles on a dark teal background, creating a crypto-themed design.]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2F5wtHENHGzT1GK6tG4M6USf%2Fa380442e95ea4f30ad0f5e6c42ea7ee8%2Fpricing-change-wide-carousel_2x_1.webp&w=3840&q=75)
From 1 August 2026, Speechmatics moves to credit-based billing: one credit balance across every product, same prices, no action needed.
A founder's account of taking a hackathon side-project live with 30 sales teams, and the four real-time integration bugs that stood between a demo and a product people could trust.

How Wellcom Health uses real-time transcription, medical-grade accuracy and diarization to turn messy Dutch consultations into validated clinical reports.
