fish-speech
View on GitHubSOTA Open Source TTS
Fish Speech (Fish Audio S2 Pro) is an open-source 4B multilingual TTS system using a Dual-AR transformer plus RVQ codec, GRPO alignment and inline [tag] prosody control. It supports voice cloning, streaming inference via SGLang, and ships a Gradio WebUI and API server.
Use Cases
Multilingual text-to-speech generation across 80+ languagesRapid voice cloning from 10-30s reference audioEmotion and prosody control with inline [tag] syntaxMulti-speaker and multi-turn dialogue audio generationLow-latency streaming TTS serving via SGLang/vLLMNarrating audiobooks, dubbing and podcast contentBuilding voice assistant or chatbot speech backendsBatch speech synthesis through the HTTP API server and WebUI
Built With
- Language
- Python
- Frameworks
- PyTorch · Transformers · PyTorch Lightning · Hydra · Gradio · SGLang · vLLM · Docker · uv · TensorBoard · Weights & Biases · React · Vite · FastAPI · uvicorn
Tags
tts · text-to-speech · voice-cloning · multilingual · dual-ar · transformer · audio-codec · rvq · streaming · grpo · speech-synthesis · generative-audio · webui · inference-server · emotion-control · multi-speaker