FunASR
VerifiedIndustrial-grade open-source speech recognition: 170x realtime, 50+ languages, diarization & emotion
Released
2023
Country
China
API
Available
Self-Host
Yes
GitHub Stars
18,043
Last Updated
2026-06
About FunASR
Verdict
The open-source speech recognition engine for production, especially in Chinese. FunASR's speed, language coverage, and beyond-text features (diarization, emotion) make it a compelling self-hostable alternative to commercial ASR — strongest where Chinese-language accuracy is critical.
Features
Detailed Ratings
Pros & Cons
Pros
- Industrial-grade accuracy and speed (up to 170x realtime)
- Excellent Chinese-language support plus 50+ languages
- Goes beyond text — diarization, emotion, punctuation
- Open-source and fully self-hostable for data sovereignty
- OpenAI-compatible API drops into existing stacks
Cons
- Requires technical setup and adequate hardware for self-hosting
- Less polished documentation than commercial speech APIs
- Strongest advantages are in Chinese-language scenarios
Use Cases
Who Is It For?
Developers and engineering teams who need accurate, fast, multilingual speech recognition they can self-host, especially for Chinese-language workloads
Frequently Asked Questions
What is FunASR?
FunASR is an open-source, industrial-grade speech recognition toolkit from Alibaba DAMO Academy's ModelScope team. It transcribes audio at up to 170x realtime across 50+ languages, with speaker diarization, emotion recognition, and streaming support.
Is FunASR free?
Yes. FunASR is free and open-source, and you can self-host it on your own infrastructure. You only pay for the compute hardware or cloud GPUs you choose to run it on.
How does FunASR compare to Whisper or commercial speech APIs?
FunASR is production-oriented with strong Chinese performance, diarization, emotion detection, and an OpenAI-compatible API. Whisper is a strong general model but lighter on production features. Commercial APIs are more polished but closed and usage-billed.
Can FunASR run on my own servers?
Yes. FunASR is fully self-hostable. The pretrained models and runtime are open, so you can run the entire speech pipeline on your own infrastructure for data sovereignty and predictable cost.
Related Agents
Leon
Voice Agents
Open-source personal assistant you can self-host, with voice and a growing agentic skill system
AgentGPT
AI Assistants
Assemble, configure, and deploy autonomous AI agents right in your browser
AnythingLLM
Productivity Agents
The full-stack AI application for chatting with your documents and building private LLM workspaces
Deepgram
Voice Agents
AI speech recognition and understanding platform
Top Alternatives
Compare with these similar tools
Links & Resources
Profiles are AI-assisted from public information — not independently hands-on tested.