- Speech To Text
- English Mandarin Malay Tamil
English-mandarin-malay-tamil speech to text transcription API
Convert English-mandarin-malay-tamil voice into accurate text in seconds. Whether you need English-mandarin-malay-tamil speech to text for real-time applications, voice recordings, or multilingual content, our transcription API delivers fast, secure, and accurate results. Trusted for English-mandarin-malay-tamil voice to text and transcription use cases, integrate high-quality English-mandarin-malay-tamil ASR into your product.
- •High-accuracy transcription of standard English-mandarin-malay-tamil and dialects
- •Supports real-time and batch processing
- •Easy to integrate with our developer-friendly API
- •Built for global enterprise scale, with secure and private processing.
- High-accuracy transcription of standard English-mandarin-malay-tamil and dialects
- Supports real-time and batch processing
- Easy to integrate with our developer-friendly API
- Built for global enterprise scale, with secure and private processing.
English-mandarin-malay-tamil transcription accuracy
Understands every accent Trained for variations of dialects and accents. Get accurate transcriptions, no matter the region. Ready for real-time scale Our API handles live recorded and live English-mandarin-malay-tamil audio at scale – with secure cloud, on-prem or on-device deployment. Built for the real world Noisy calls, fast speakers, crosstalk – our tech thrives in messy audio. Experience English-mandarin-malay-tamil transcription that works
Try our live English-mandarin-malay-tamil transcription for yourself
Speak into your mic and watch real-time English-mandarin-malay-tamil transcription in action. Fast, accurate, and built for natural conversations.
Four languages supported by one pack
Combined, the four official languages of Singapore cover the everyday speech of over 50 million people in Singapore, Malaysia, Brunei, and the diaspora across Southeast Asia and beyond.
English (en). Includes Singaporean English and Singlish.
Mandarin (cmn). Accents from China, Taiwan, Singapore, and Malaysia.
Malay (ms). Includes Bahasa Melayu spoken in Singapore and Malaysia.
Tamil (ta). Includes Singapore Tamil and Tamil across the diaspora.

Everything you need for accurate, scalable English-mandarin-malay-tamil speech to text.
Built for real-world use cases and global applications.Everything you need for accurate, scalable English-mandarin-malay-tamil speech to text.
AI speech to text transcription in 55+ languages
Frequently Asked Questions — English, Mandarin, Malay & Tamil
The English, Mandarin, Malay and Tamil multilingual pack, language code cmn_en_ms_ta, is a single speech recognition model that transcribes audio containing any mix of the four languages. It handles code-switching mid-sentence as part of natural speech, unlike general-purpose multilingual ASR, which tends to treat each language as separate.
The pack was built for Southeast Asia, where the four languages are the official languages of Singapore and are spoken across Malaysia, Brunei, and the diaspora. It powers contact center transcription, voice agents, broadcast captioning, and government workflows across the region.
Read more about how the models were built in The real language of business.
Yes. Singlish and code-switching are the default case, not an exception.
Singlish, the English-based creole spoken across Singapore, blends English grammar with Malay, Hokkien, Cantonese, and Tamil vocabulary. Code-switching, the practice of moving between languages within one sentence, is how millions of people in Singapore and Malaysia actually speak. The multilingual pack was trained on conversational audio that reflects this. It transcribes what was said, without breaking at the switch.
General-purpose multilingual ASR supports a long list of languages individually. Each one is handled by a separate model. When a speaker switches languages mid-sentence, most of these systems drop words, produce broken output, or force the developer to make multiple API calls per audio file.
The Speechmatics multilingual pack uses one model and one API call for all four languages, including the switch points. Fewer errors on Singaporean English. Fewer errors on code-switched audio. One transcript, in order, per request.
The English, Mandarin, Malay and Tamil pack runs on two operating points. Enhanced returns the lowest word error rate. It's the default choice for compliance, quality monitoring, and other accuracy-critical work. Standard trades a small amount of accuracy for faster throughput, useful for high-volume batch jobs or cost-sensitive workloads.
Both operating points handle the same four-language code-switching. Set operating_point to enhanced or standard in your transcription config.
Real-time transcription streams audio to the API over a WebSocket connection and returns transcript results as the audio arrives. Partial transcripts land in under a second. Final transcripts follow within two seconds.
Speechmatics supports real-time transcription for the full 55+ language range, including the English, Mandarin, Malay and Tamil pack. The system is built for spontaneous speech, interruptions, background noise, and language switching. For pre-recorded audio and video, batch transcription runs the same model at higher throughput.
Read the real-time overview or the real-time quickstart in docs.
The API integrates multilingual transcription into applications, platforms, and internal systems.
You can:
Transcribe audio and video files programmatically across all four languages.
Stream live audio for real-time transcription with code-switching support.
Return structured transcripts with timestamps and speaker identification.
Prepare text for analytics, subtitling, and downstream workflows.
Add custom dictionary entries for local terms, place names, and technical vocabulary.
The API is production-grade and supports cloud, hybrid, on-premise, and on-device deployment.
The English, Mandarin, Malay and Tamil pack is used across:
Customer interaction analysis and quality monitoring in contact center solutions.
Voice automation in AI voice agents.
Subtitle creation and accessibility in media distribution and captioning.
Collaboration and discussion capture in meeting platforms.
Brand and news monitoring in media monitoring.
Organizations with data residency or scale requirements deploy through enterprise speech recognition.
Upload the file to the Speechmatics portal or send it through the Batch API. Set the language to cmn_en_ms_ta. The system returns a transcript with timestamps and speaker labels. Export as text, JSON, or SRT.
Yes. Create an account and you get eight hours of free transcription every month, across every supported language, including the English, Mandarin, Malay and Tamil pack. See pricing for volume rates.
Yes. The pack runs on CPU and GPU containers, Kubernetes, and air-gapped environments. Same model, same accuracy, your infrastructure. Common for organizations with strict data residency requirements across Singapore, Malaysia, and the wider region.
Yes. The model is trained on real-world audio, including phone-quality audio at 8 kHz, background noise, and cross-talk typical of contact center recordings.
WAV, MP3, AAC, OGG, MPEG, AMR, M4A, MP4, and FLAC.
Speechmatics supports multilingual audio in two ways.
Specialist bilingual and multilingual language packs — These cover a fixed set of languages that you select in advance. In addition to the four-language cmn_en_ms_ta pack, the pack range includes:
Mandarin and English (cmn_en)
Malay and English (en_ms)
Tamil and English (en_ta)
Arabic and English (ar_en)
Spanish and English
Melia — Melia is Speechmatics' broader multilingual model. It handles code-switching across all 55+ supported languages in a single model, without requiring you to select or manage individual packs. Melia is in production preview for Batch, with real-time support on the roadmap.
See the full language coverage in docs or the Melia model docs.
Contact centers, voice agent platforms, government and public-sector workflows, media and broadcasting, meeting and collaboration tools, and accessibility workflows.
![[alt: Industry-leading transcription accuracy in 55+ languages]](/_next/image?url=https%3A%2F%2Fimages.ctfassets.net%2Fyze1aysi0225%2F1dGuTnCrsPeC1XuiYZHdJx%2F854dfedb68eee0749d5b5f2521030fd6%2F9e3ae9aeb3cd6c9da26f9068fe1a29ce1098b1f9.png&w=3840&q=75)