Skip to content
AgentRadar
FunASR

FunASR

Verified

Industrial-grade open-source speech recognition: 170x realtime, 50+ languages, diarization & emotion

Voice Agents Open Source

Released

2023

Country

China

API

Available

Self-Host

Yes

GitHub Stars

18,043

Last Updated

2026-06

About FunASR

FunASR is an industrial-grade, open-source speech recognition toolkit developed by the ModelScope team at Alibaba DAMO Academy. It is built for production voice workloads where accuracy, speed, and language coverage matter: the toolkit delivers up to 170x realtime transcription speed, supports more than 50 languages with particularly strong Chinese performance, and goes well beyond plain text with speaker diarization (telling who spoke when), emotion recognition, punctuation restoration, and streaming capabilities for live audio. What sets FunASR apart from typical ASR demos is its end-to-end production orientation. It ships pretrained models that are ready to deploy, offers an OpenAI-compatible API so it can drop into existing stacks, and provides both local and cloud deployment paths. This makes it suitable for real applications — call-center transcription, meeting captioning, voice assistants, media subtitling, and compliance monitoring — rather than just experimentation. The toolkit is part of the broader ModelScope model ecosystem, which gives it access to a wide library of continuously improved speech models. FunASR has attracted a large community of developers (tens of thousands of GitHub stars), especially among teams building Chinese-language voice products who need an open, self-hostable alternative to proprietary speech APIs. Because the models and runtime are open, teams can run everything on their own infrastructure for data sovereignty and predictable cost. It is aimed at developers and engineering teams who need accurate, fast, multilingual speech recognition they can self-host and integrate into products — particularly those with heavy Chinese-language requirements.

Verdict

The open-source speech recognition engine for production, especially in Chinese. FunASR's speed, language coverage, and beyond-text features (diarization, emotion) make it a compelling self-hostable alternative to commercial ASR — strongest where Chinese-language accuracy is critical.

Features

Up to 170x realtime transcription
50+ languages (strong Chinese)
Speaker diarization
Emotion recognition
Streaming / live audio
OpenAI-compatible API
Punctuation restoration

Detailed Ratings

Ease of Use
7.4
Value for Money
8.4
Features
8.2
Support
7.4
Performance
8.2
Overall Rating
8.0 /10

Pros & Cons

Pros

  • Industrial-grade accuracy and speed (up to 170x realtime)
  • Excellent Chinese-language support plus 50+ languages
  • Goes beyond text — diarization, emotion, punctuation
  • Open-source and fully self-hostable for data sovereignty
  • OpenAI-compatible API drops into existing stacks

Cons

  • Requires technical setup and adequate hardware for self-hosting
  • Less polished documentation than commercial speech APIs
  • Strongest advantages are in Chinese-language scenarios

Use Cases

Call-center & meeting transcriptionLive captioning & streaming audioVoice assistant input pipelinesMedia subtitlingCompliance & monitoring transcription

Who Is It For?

Developers and engineering teams who need accurate, fast, multilingual speech recognition they can self-host, especially for Chinese-language workloads

#speech-recognition#asr#open-source#chinese#self-hosted#streaming#modelscope

Frequently Asked Questions

What is FunASR?

FunASR is an open-source, industrial-grade speech recognition toolkit from Alibaba DAMO Academy's ModelScope team. It transcribes audio at up to 170x realtime across 50+ languages, with speaker diarization, emotion recognition, and streaming support.

Is FunASR free?

Yes. FunASR is free and open-source, and you can self-host it on your own infrastructure. You only pay for the compute hardware or cloud GPUs you choose to run it on.

How does FunASR compare to Whisper or commercial speech APIs?

FunASR is production-oriented with strong Chinese performance, diarization, emotion detection, and an OpenAI-compatible API. Whisper is a strong general model but lighter on production features. Commercial APIs are more polished but closed and usage-billed.

Can FunASR run on my own servers?

Yes. FunASR is fully self-hostable. The pretrained models and runtime are open, so you can run the entire speech pipeline on your own infrastructure for data sovereignty and predictable cost.

Related Agents

Top Alternatives

Compare with these similar tools

Links & Resources

AR Reviewed by AgentRadar · Reviewed on · How we rate

Profiles are AI-assisted from public information — not independently hands-on tested.