CosyVoice
★ 24KCosyVoice is an LLM-based text-to-speech system (Fun-CosyVoice3 / CosyVoice2 / CosyVoice1) supporting multilingual zero-shot voice cloning, streaming output down to ~150ms latency, instruct control and voice conversion. Ships inference, fine-tuning and deployment paths (vLLM, TensorRT-LLM, gRPC/FastAPI, Gradio demo).
AI Tools | Python · text-to-speech · tts
View Project →