funasr
#39 of 54 by Downloads / month among PyPI packages
OpenAI-compatible speech recognition toolkit with WebSocket streaming, vLLM acceleration, and llama.cpp/GGUF edge runtime.
About
Industrial speech recognition. 170x faster than Whisper. 50+ languages.
No local setup? Open the Colab quickstart to transcribe a public sample or upload your own audio in a browser.
Flagship model — Fun-ASR-Nano (LLM-ASR, 31 languages; the default recommendation, needs a GPU):
On CPU (or for multilingual + emotion in one pass), use SenseVoice — which also returns speaker diarization and timestamps:
Output — structured text with speaker labels, timestamps, and punctuation:
That's it. One model, one call — VAD segmentation, speech recognition, punctuation, speaker diarization all happen automatically.
At scale, accelerate Fun-ASR-Nano with vLLM (batch processing):
Deploy as API…
Excerpted from pypi.org/project/funasr/
Latest metrics
| Downloads / month | 469.9k | 2026-08-02 |
|---|---|---|
| Downloads / week | 116.8k | 2026-08-02 |
| Downloads / day | 14.2k | 2026-08-02 |