- Use Cases
- Legal Transcription
Legal speech‑to‑text built for the courtroom
Speech recognition software built for court reporters, legal professionals, and law firms who need unmatched accuracy across every accent, dialect, and speaker — in real time.

Trusted by Legal Technology Leaders
Industry-first on-device speech recognition for CATalyst VP
How Stenograph brought industry-first, on-device speech recognition into CATalyst VP for real-time legal voice reporting.Industry-first on-device speech recognition for CATalyst VP
How Stenograph brought industry-first, on-device speech recognition into CATalyst VP for real-time legal voice reporting.Transcription trade-off that puts cases at risk
Overlapping speakers, accents, and legal terminology break accuracy when it matters most.
Stenographers are declining. Demand is rising. Final transcripts are expected faster with fewer resources.
When the record is wrong, the costs are measured in lost cases, not just dollars.
Expensive human court reporters with a long turnaround, or generic speech recognition that miss critical terminology.
Overlapping speakers, accents, and legal terminology break accuracy when it matters most.
Stenographers are declining. Demand is rising. Final transcripts are expected faster with fewer resources.
When the record is wrong, the costs are measured in lost cases, not just dollars.
Expensive human court reporters with a long turnaround, or generic speech recognition that miss critical terminology.
When every word is evidence, accuracy isn't optional.
When every word is evidence, accuracy isn't optional.
Real-time and file-based AI transcription for the highest-stakes conversations — from courtrooms to depositions to criminal evidence.
Unbeatable legal transcription
Sub-second latency - transcripts appear as words are spoken
Accent-agnostic - consistent accuracy across dialects and non-native speakers
Noise-resilient - reliable in challenging audio from busy courtrooms to body cam recordings
55+ languages - multilingual transcription with dialect comprehension across every language
The speech engine powering legal transcription
The speech engine powering legal transcription
Deployment flexibility means legal technology partners meet any client data security requirement without changing providers

AI-Assisted Court Reporting
Generate real-time draft transcripts during live court proceedings. Court reporters get a reliable starting point to review, edit, and deliver final transcripts in hours instead of days.

Deposition & Discovery Transcription
Batch transcription of deposition libraries with speaker diarization and timestamps. Custom vocabularies capture party names and case-specific terminology, reducing costs and turnaround.

Audio Evidence Transcription
Process body cam footage, paramedic calls, jail calls, interrogation recordings, and surveillance audio at scale. Accurate, reliable transcripts that hold up to scrutiny in trial proceedings.
Your data stays where your compliance demands
Your data stays where your compliance demands
Speechmatics offers a secure platform with full deployment flexibility - cloud, on-premises, or on-device - so law firms and legal technology providers meet any data sovereignty requirement.

Privacy by Design
No data logging by default. Sensitive client data - testimony, depositions, evidence recordings - stays protected. You control your legal documentation.

SOC2 | ISO | HIPAA | GDPR
Fully compliant AI that meets the high levels of security legal organizations demand. Deployment practices align with the strictest industry regulations.
Resources for healthcare & medical

What Word Error Rate Is Acceptable for Legal Transcription?
Word error rate for legal transcription has no single acceptable threshold. But knowing how accuracy, audio quality, and review obligations connect to real legal risk is what separates a reliable transcript from a costly one.

The court reporter shortage crisis: data, causes, and what legal teams are doing about it
The court reporter shortage is reshaping litigation. Explore data, causes, and how legal teams are using digital reporting and AI transcription to adapt.
Legal Speech-to-Text FAQs
AI-powered legal transcription uses automatic speech recognition (ASR) models trained on legal vocabulary to convert spoken audio into structured text in real time or from recorded files. The system processes audio through acoustic and language models, applies custom legal dictionaries, and outputs formatted transcripts with speaker labels and timestamps — ready for attorney review.
Speech-to-text can generate real-time draft transcripts during live proceedings, giving court reporters an accurate starting point to review and certify rather than transcribing from scratch. It also handles batch transcription of depositions, hearings, and evidence recordings — dramatically reducing turnaround time and backlog.
Traditional stenography relies on a trained human using a shorthand machine to capture speech at speed, then translate and format the transcript manually. AI transcription captures audio directly and produces a structured draft in seconds. The key difference is speed and scalability — AI handles volume that would require multiple stenographers, though a certified reporter still reviews and certifies the final output.
Yes, when combined with human review. Leading ASR systems achieve word error rates well below 5% on legal audio with clear conditions and custom vocabularies. The workflow pairs AI-generated drafts with certified reporter review, meeting or exceeding accuracy standards while significantly reducing the time required.
For a draft transcript used as a starting point, a WER of under 3–5% is considered acceptable in most legal contexts. For the final certified transcript, the standard is effectively zero — every word must be accurate. The AI draft gets the reporter to near-perfect quickly; human review closes the remaining gap.
Audio is captured via microphone or a recording interface and streamed to the ASR engine, which returns transcript text with a latency typically under one second. The live transcript appears on screen for the reporter and, optionally, for counsel and participants. Corrections can be made in real time, and the session is simultaneously saved as a recording for post-proceeding verification.
Real-time transcription processes audio as it is spoken, delivering a live text feed during the proceeding. Batch transcription processes completed recordings after the fact — useful for depositions, archived evidence, or high-volume discovery work. Both use the same underlying models; the choice depends on whether a live transcript is operationally required.
Speaker diarization segments the audio stream into turns and assigns each segment to a distinct speaker identity. In legal settings, speakers can be pre-enrolled by voice profile so the system labels turns as "Judge," "Plaintiff Counsel," "Witness," etc., rather than generic Speaker A/B labels. This ensures every statement is correctly attributed before the reporter reviews the draft.
Yes, with voice enrollment. Participants' voice profiles are registered at the start of a case or matter, and the system maps incoming speech to those profiles throughout the proceeding. For recurring participants — judges in a particular court, for example — profiles can be stored and reused, improving attribution accuracy over time.