Vibe Coding Discover

AI Tools

fish-speech

View on GitHub

SOTA Open Source TTS

★ 33K2,835 forksPythonCustomfishaudio

Fish Speech (Fish Audio S2 Pro) is an open-source 4B multilingual TTS system using a Dual-AR transformer plus RVQ codec, GRPO alignment and inline [tag] prosody control. It supports voice cloning, streaming inference via SGLang, and ships a Gradio WebUI and API server.

Use Cases

Multilingual text-to-speech generation across 80+ languagesRapid voice cloning from 10-30s reference audioEmotion and prosody control with inline [tag] syntaxMulti-speaker and multi-turn dialogue audio generationLow-latency streaming TTS serving via SGLang/vLLMNarrating audiobooks, dubbing and podcast contentBuilding voice assistant or chatbot speech backendsBatch speech synthesis through the HTTP API server and WebUI

Built With

Language
Python
Frameworks
PyTorch · Transformers · PyTorch Lightning · Hydra · Gradio · SGLang · vLLM · Docker · uv · TensorBoard · Weights & Biases · React · Vite · FastAPI · uvicorn

Tags

tts · text-to-speech · voice-cloning · multilingual · dual-ar · transformer · audio-codec · rvq · streaming · grpo · speech-synthesis · generative-audio · webui · inference-server · emotion-control · multi-speaker