Duration-controllable, long-form zero-shot TTS (with Meta co-authors): progress-monitoring rotary embeddings let the model hit a target duration and extrapolate beyond training lengths. In-window successor to David Harwath and Puyuan Peng's VoiceCraft (arXiv:2403.16973, ACL 2024, ~8K★), the SOTA zero-shot speech-editing/TTS model that predates the window by four months.

Paper

audiogenerationopen-weight