Vibe Coding Discover

AI Tools

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.

★ 10K853 forksPythonMITopen-mmlab

OpenMMLab toolkit for audio, music and speech generation: TTS, singing voice synthesis/conversion, voice conversion and text-to-audio, with neural codecs, vocoders and evaluation metrics. Ships reproducible PyTorch recipes, pretrained checkpoints and the Emilia dataset.

Use Cases

text-to-speech synthesiszero-shot voice cloningvoice conversionsinging voice synthesissinging voice conversiontext-to-audio generationmusic generationneural vocoder inferencespeech codec token extractionaccent conversionspeech dataset preprocessing for trainingaudio generation evaluation metricsreproducing speech generation research paperstraining custom TTS/VC models

Built With

Language
Python
Frameworks
PyTorch · Hugging Face Transformers · Hugging Face Hub · Hugging Face Datasets · PyTorch Lightning · ModelScope · ONNX

Tags

text-to-speech · speech-synthesis · voice-conversion · singing-voice-conversion · singing-voice-synthesis · text-to-audio · music-generation · audio-generation · neural-audio-codec · vocoder · zero-shot-tts · multilingual-speech · reproducible-research · speech-dataset · evaluation-metrics · pytorch