Quite an obvious but so ignored by industry before. There is certainly no need to clone from 3 seconds
You can now create an instant voice clone from up to 2 minutes of reference audio, compared with 20 seconds before.
The longer sample gives the model much more information about the voice, including phonemes, intonation, pauses, rhythm, pacing, emphasis, emotion, and speaking style.
The result:
• Higher voice similarity
• Greater consistency, especially across longer sentences
• Better pronunciation coverage
• More natural prosody
• Better preservation of speaking style
https://www.linkedin.com/posts/soniox_soniox-texttospeech-voiceai-activity-7501594413660401664-p9Cq
You can now create an instant voice clone from up to 2 minutes of reference audio, compared with 20 seconds before.
The longer sample gives the model much more information about the voice, including phonemes, intonation, pauses, rhythm, pacing, emphasis, emotion, and speaking style.
The result:
• Higher voice similarity
• Greater consistency, especially across longer sentences
• Better pronunciation coverage
• More natural prosody
• Better preservation of speaking style
https://www.linkedin.com/posts/soniox_soniox-texttospeech-voiceai-activity-7501594413660401664-p9Cq
LinkedIn
Soniox TTS v2: Improved Voice Cloning with Longer Audio Samples | Soniox posted on the topic | LinkedIn
We’ve shipped a major improvement to voice cloning in Soniox TTS v2.
You can now create an instant voice clone from up to 2 minutes of reference audio, compared with 20 seconds before.
The longer sample gives the model much more information about the voice…
You can now create an instant voice clone from up to 2 minutes of reference audio, compared with 20 seconds before.
The longer sample gives the model much more information about the voice…
https://arxiv.org/abs/2609.01246v1
Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation
Thibaut Thonet, Jos Rozen, Laurent Besacier
Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet poorly suited for spoken delivery via Text-to-Speech (TTS). In this work, we study how to make LLMs natively generate TTS-friendly text, which we frame as a preference alignment problem: instead of relying on downstream rewriting modules, we directly align LLMs to generate text optimized for spoken delivery. We introduce two preference datasets spanning different target domains, CORA and Recipe, which contain paired TTS-friendly and TTS-unfriendly responses. We further propose an evaluation suite combining a pattern-based heuristic metric, a TTS→ASR evaluation pipeline, and a MUSHRA listening study with human judges. Our experiments compare the recently proposed Feature-aware Sampling and Tuning (FaST) framework -- leveraging interpretable features instead of a black-box reward model -- against an array of alignment baselines on the TTS-friendly generation task. Notably, we found that FaST achieves the best overall tradeoff between TTS-friendliness and helpfulness across various settings. We also identified a strong correlation between our different metrics, highlighting the ability to reliably assess TTS-friendliness via an efficient heuristic.
Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation
Thibaut Thonet, Jos Rozen, Laurent Besacier
Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet poorly suited for spoken delivery via Text-to-Speech (TTS). In this work, we study how to make LLMs natively generate TTS-friendly text, which we frame as a preference alignment problem: instead of relying on downstream rewriting modules, we directly align LLMs to generate text optimized for spoken delivery. We introduce two preference datasets spanning different target domains, CORA and Recipe, which contain paired TTS-friendly and TTS-unfriendly responses. We further propose an evaluation suite combining a pattern-based heuristic metric, a TTS→ASR evaluation pipeline, and a MUSHRA listening study with human judges. Our experiments compare the recently proposed Feature-aware Sampling and Tuning (FaST) framework -- leveraging interpretable features instead of a black-box reward model -- against an array of alignment baselines on the TTS-friendly generation task. Notably, we found that FaST achieves the best overall tradeoff between TTS-friendliness and helpfulness across various settings. We also identified a strong correlation between our different metrics, highlighting the ability to reliably assess TTS-friendliness via an efficient heuristic.
arXiv.org
Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation
Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet poorly suited for spoken delivery via...
Scicom from Malaysia tries Ascend 910B3
https://github.com/Scicom-AI-Enterprise-Organization/TTS-API-Neucodec/blob/main/ASCEND_910B3_PRECISION_REPORT.md
https://github.com/Scicom-AI-Enterprise-Organization/TTS-API-Neucodec/blob/main/ASCEND_910B3_PRECISION_REPORT.md
GitHub
TTS-API-Neucodec/ASCEND_910B3_PRECISION_REPORT.md at main · Scicom-AI-Enterprise-Organization/TTS-API-Neucodec
TTS API OpenAI compatible on top of Neucodec Speech Tokenizer LLM - Scicom-AI-Enterprise-Organization/TTS-API-Neucodec
Bodhan AI together with AI4Bharat recently released a great update on Indic ASR
https://bodhan.ai/research/blogs/indic-transcribe
https://bodhan.ai/research/blogs/indic-transcribe
ParsVoice, the largest open-source Persian speech dataset, along with a TTS model and an open-source processing pipeline is released.
The paper has also been accepted as a main conference paper at EMNLP 2026.
Paper: https://arxiv.org/abs/2510.10774
Dataset: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice
TTS model: https://huggingface.co/MohammadJRanjbar/ParsVoice-XTTS
Code & pipeline: https://github.com/MohammadJRanjbar/ParsVoice
The paper has also been accepted as a main conference paper at EMNLP 2026.
Paper: https://arxiv.org/abs/2510.10774
Dataset: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice
TTS model: https://huggingface.co/MohammadJRanjbar/ParsVoice-XTTS
Code & pipeline: https://github.com/MohammadJRanjbar/ParsVoice
arXiv.org
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for...
Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech...
https://huggingface.co/tencent/AuK
AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface.
https://arxiv.org/abs/2609.08936
AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface.
https://arxiv.org/abs/2609.08936
huggingface.co
tencent/AuK · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
My friend Miro recommended me Orukeet model
https://github.com/Oruk-AI/orukeet
https://arxiv.org/abs/2609.10054
it is indeed a good parakeet improvement, about 10% better. Interesting that people left scaling and return back to in-depth architecture analysis.
Oruk.AI does some other nice things, for example a visualization of emotion representation in different layers of speech models
https://x.com/OrukLabs/status/2073457781018087473
https://oruk.ai/research/how-models-represent-speech
https://github.com/Oruk-AI/orukeet
https://arxiv.org/abs/2609.10054
it is indeed a good parakeet improvement, about 10% better. Interesting that people left scaling and return back to in-depth architecture analysis.
Oruk.AI does some other nice things, for example a visualization of emotion representation in different layers of speech models
https://x.com/OrukLabs/status/2073457781018087473
https://oruk.ai/research/how-models-represent-speech
GitHub
GitHub - Oruk-AI/orukeet: Orukeet: multilingual ASR with fitted, frozen Gabor kernels and native inference
Orukeet: multilingual ASR with fitted, frozen Gabor kernels and native inference - Oruk-AI/orukeet
https://x.com/unilightwf/status/2098261200480174123
https://arxiv.org/abs/2603.14328
Cross-lingual cloning TTS is still a big problem, some accent metrics demo interesting results
Also worth checking
https://iwslt.org/2026/voice-cloning
with some useful data
https://huggingface.co/datasets/ymoslem/acl-6060
https://arxiv.org/abs/2603.14328
Cross-lingual cloning TTS is still a big problem, some accent metrics demo interesting results
Also worth checking
https://iwslt.org/2026/voice-cloning
with some useful data
https://huggingface.co/datasets/ymoslem/acl-6060
10000xRT for Zipformer on A100, 18000xRT on H200 with specialized CUDA tricks
https://github.com/SoundsGoodAI/fast-gpu-asr
https://github.com/SoundsGoodAI/fast-gpu-asr
GitHub
GitHub - SoundsGoodAI/fast-gpu-asr: High-throughput batched ASR for NVIDIA GPUs. Run Zipformer and NVIDIA NeMo Parakeet models…
High-throughput batched ASR for NVIDIA GPUs. Run Zipformer and NVIDIA NeMo Parakeet models through one Python API with TensorRT, custom CUDA plugins, GPU beam search, and word timestamps - up to 25...
2.1k parameters VAD
https://github.com/AydinAdnan/PulseVAD
reimplementation of KiloVAD
https://arxiv.org/abs/2607.25870
https://github.com/AydinAdnan/PulseVAD
reimplementation of KiloVAD
https://arxiv.org/abs/2607.25870
GitHub
GitHub - AydinAdnan/PulseVAD: Ultra-tiny 2.1k streaming voice activity detector (2.1 KB INT8). 200 ms window for microcontrollers…
Ultra-tiny 2.1k streaming voice activity detector (2.1 KB INT8). 200 ms window for microcontrollers, edge devices, and DSPs. ships with ONNX, TorchScript, and standalone C header. - AydinAdnan/Puls...
We see a huge decline of the interest in speech technology in China but a rise in Europe and India. Recently I discussed it with one of my Chinese friends - he confirms that nobody is interested anymore in plain speech. It has to be multimodal - video, music generation, etc.
"百闻不如一见" — hearing something a hundred times is worse than seeing it once
China is ahead of time here.
"百闻不如一见" — hearing something a hundred times is worse than seeing it once
China is ahead of time here.
Another small 260k params TTS, hard to believe into UTMOS 4.0 based on examples but ok
https://github.com/lab-emi/GrainSpeech
https://lab-emi.github.io/GrainSpeech/
https://arxiv.org/abs/2609.18856
https://github.com/lab-emi/GrainSpeech
https://lab-emi.github.io/GrainSpeech/
https://arxiv.org/abs/2609.18856
GitHub
GitHub - lab-emi/GrainSpeech: Official repository for “GrainSpeech: Less Context, More Detail for Compact Speech Synthesis.”
Official repository for “GrainSpeech: Less Context, More Detail for Compact Speech Synthesis.” - lab-emi/GrainSpeech
An advanced dictation model from Whispr, an interesting part is GPRO over user corrections
https://wisprflow.ai/canto
https://wisprflow.ai/canto
wisprflow.ai
Canto: a speech model built for the real world | Wispr Flow
Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation. On an evaluation of real-world dictations, Canto achieved the lowest word error rate among all the models we tested.
Voice agents related benchmarks somehow passed our attention, here are few recent ones:
https://github.com/sierra-research/tau2-bench
https://github.com/ServiceNow/eva
tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
https://arxiv.org/abs/2603.13686
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
https://arxiv.org/abs/2605.13841
https://github.com/sierra-research/tau2-bench
https://github.com/ServiceNow/eva
tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
https://arxiv.org/abs/2603.13686
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
https://arxiv.org/abs/2605.13841
GitHub
GitHub - sierra-research/tau2-bench: τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains - sierra-research/tau2-bench