Back to Catalog

OmniVoice

Model

OmniVoice is an open multilingual text-to-speech model developed by the k2-fsa organization. The model supports speech synthesis in 646 languages, including Slovak, and is based on a diffusion language model architecture. It enables speech generation from text input as well as zero-shot voice cloning from a reference recording without the need for additional training. In addition to voice cloning, it supports synthesis style control, not only based on a reference recording but also through text instructions describing the desired delivery style (e.g., emotion, pace, or intonation).

Contributors
Xiaomi Corp.
License
Apache License 2.0
Language
sk, eng, other
Modality
Speech/audio
Task
Text-to-speech