Speech Technology
1.72K subscribers
140 photos
4 videos
1 file
2.24K links
Download Telegram
Quite an obvious but so ignored by industry before. There is certainly no need to clone from 3 seconds

You can now create an instant voice clone from up to 2 minutes of reference audio, compared with 20 seconds before.

The longer sample gives the model much more information about the voice, including phonemes, intonation, pauses, rhythm, pacing, emphasis, emotion, and speaking style.

The result:
• Higher voice similarity
• Greater consistency, especially across longer sentences
• Better pronunciation coverage
• More natural prosody
• Better preservation of speaking style

https://www.linkedin.com/posts/soniox_soniox-texttospeech-voiceai-activity-7501594413660401664-p9Cq
https://arxiv.org/abs/2609.01246v1

Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation

Thibaut Thonet, Jos Rozen, Laurent Besacier

Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet poorly suited for spoken delivery via Text-to-Speech (TTS). In this work, we study how to make LLMs natively generate TTS-friendly text, which we frame as a preference alignment problem: instead of relying on downstream rewriting modules, we directly align LLMs to generate text optimized for spoken delivery. We introduce two preference datasets spanning different target domains, CORA and Recipe, which contain paired TTS-friendly and TTS-unfriendly responses. We further propose an evaluation suite combining a pattern-based heuristic metric, a TTS→ASR evaluation pipeline, and a MUSHRA listening study with human judges. Our experiments compare the recently proposed Feature-aware Sampling and Tuning (FaST) framework -- leveraging interpretable features instead of a black-box reward model -- against an array of alignment baselines on the TTS-friendly generation task. Notably, we found that FaST achieves the best overall tradeoff between TTS-friendliness and helpfulness across various settings. We also identified a strong correlation between our different metrics, highlighting the ability to reliably assess TTS-friendliness via an efficient heuristic.
Bodhan AI together with AI4Bharat recently released a great update on Indic ASR

https://bodhan.ai/research/blogs/indic-transcribe
https://huggingface.co/tencent/AuK

AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface.

https://arxiv.org/abs/2609.08936
Suplime results on Russian telephony data, good results, first place, large model is better than Diarizen

A bit slow though
My friend Miro recommended me Orukeet model

https://github.com/Oruk-AI/orukeet

https://arxiv.org/abs/2609.10054

it is indeed a good parakeet improvement, about 10% better. Interesting that people left scaling and return back to in-depth architecture analysis.

Oruk.AI does some other nice things, for example a visualization of emotion representation in different layers of speech models

https://x.com/OrukLabs/status/2073457781018087473

https://oruk.ai/research/how-models-represent-speech
https://x.com/unilightwf/status/2098261200480174123

https://arxiv.org/abs/2603.14328

Cross-lingual cloning TTS is still a big problem, some accent metrics demo interesting results

Also worth checking

https://iwslt.org/2026/voice-cloning

with some useful data

https://huggingface.co/datasets/ymoslem/acl-6060
We see a huge decline of the interest in speech technology in China but a rise in Europe and India. Recently I discussed it with one of my Chinese friends - he confirms that nobody is interested anymore in plain speech. It has to be multimodal - video, music generation, etc.

"百闻不如一见" — hearing something a hundred times is worse than seeing it once

China is ahead of time here.
Serious issues with Qwen3-TTS stability, long texts and audio prompts break things

https://arxiv.org/abs/2609.16989

https://x.com/RmdW_W/status/2102032894600401231

Taming Long-form Text-to-Speech
Rongxiang Wang, Berkin Durmus, Aysegul Orhon, Eduardo Pacheco, Atila Orhon
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and speaker similarity (SIM) on short-form prompts but significantly deteriorate when used with long-form prompts. We propose Localized Attention-Constrained Inference (LACI), an inference-only method to detect TTS errors in near real-time, roll back to the error onset and regenerate with temporary guardrails, adding negligible computational overhead. Using LACI, we improve worst-of-N WER across 10 RNG seeds for Qwen3-TTS-0.6B from 35.2% to 3.4% on prompts longer than 1500 words, even surpassing its short-form reliability of 5.4\% on prompts with fewer than 500 words. To demonstrate the efficacy of LACI on voice cloning reliability, we propose a sliding-window version of the SIM metric that we call wSIM. wSIM exposes several novel failure patterns that are not captured by SIM. LACI improves worst-of-N wSIM from 0.01 to 0.47 on 120 seconds of reference audio while reducing the rate of catastrophic generations with WER above 30% from 26% to below 1%