Bodhan AI together with AI4Bharat recently released a great update on Indic ASR
https://bodhan.ai/research/blogs/indic-transcribe
https://bodhan.ai/research/blogs/indic-transcribe
ParsVoice, the largest open-source Persian speech dataset, along with a TTS model and an open-source processing pipeline is released.
The paper has also been accepted as a main conference paper at EMNLP 2026.
Paper: https://arxiv.org/abs/2510.10774
Dataset: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice
TTS model: https://huggingface.co/MohammadJRanjbar/ParsVoice-XTTS
Code & pipeline: https://github.com/MohammadJRanjbar/ParsVoice
The paper has also been accepted as a main conference paper at EMNLP 2026.
Paper: https://arxiv.org/abs/2510.10774
Dataset: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice
TTS model: https://huggingface.co/MohammadJRanjbar/ParsVoice-XTTS
Code & pipeline: https://github.com/MohammadJRanjbar/ParsVoice
arXiv.org
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for...
Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech...
https://huggingface.co/tencent/AuK
AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface.
https://arxiv.org/abs/2609.08936
AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface.
https://arxiv.org/abs/2609.08936
huggingface.co
tencent/AuK · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
My friend Miro recommended me Orukeet model
https://github.com/Oruk-AI/orukeet
https://arxiv.org/abs/2609.10054
it is indeed a good parakeet improvement, about 10% better. Interesting that people left scaling and return back to in-depth architecture analysis.
Oruk.AI does some other nice things, for example a visualization of emotion representation in different layers of speech models
https://x.com/OrukLabs/status/2073457781018087473
https://oruk.ai/research/how-models-represent-speech
https://github.com/Oruk-AI/orukeet
https://arxiv.org/abs/2609.10054
it is indeed a good parakeet improvement, about 10% better. Interesting that people left scaling and return back to in-depth architecture analysis.
Oruk.AI does some other nice things, for example a visualization of emotion representation in different layers of speech models
https://x.com/OrukLabs/status/2073457781018087473
https://oruk.ai/research/how-models-represent-speech
GitHub
GitHub - Oruk-AI/orukeet: Orukeet: multilingual ASR with fitted, frozen Gabor kernels and native inference
Orukeet: multilingual ASR with fitted, frozen Gabor kernels and native inference - Oruk-AI/orukeet
https://x.com/unilightwf/status/2098261200480174123
https://arxiv.org/abs/2603.14328
Cross-lingual cloning TTS is still a big problem, some accent metrics demo interesting results
Also worth checking
https://iwslt.org/2026/voice-cloning
with some useful data
https://huggingface.co/datasets/ymoslem/acl-6060
https://arxiv.org/abs/2603.14328
Cross-lingual cloning TTS is still a big problem, some accent metrics demo interesting results
Also worth checking
https://iwslt.org/2026/voice-cloning
with some useful data
https://huggingface.co/datasets/ymoslem/acl-6060
10000xRT for Zipformer on A100, 18000xRT on H200 with specialized CUDA tricks
https://github.com/SoundsGoodAI/fast-gpu-asr
https://github.com/SoundsGoodAI/fast-gpu-asr
GitHub
GitHub - SoundsGoodAI/fast-gpu-asr: High-throughput batched ASR for NVIDIA GPUs. Run Zipformer and NVIDIA NeMo Parakeet models…
High-throughput batched ASR for NVIDIA GPUs. Run Zipformer and NVIDIA NeMo Parakeet models through one Python API with TensorRT, custom CUDA plugins, GPU beam search, and word timestamps - up to 25...
2.1k parameters VAD
https://github.com/AydinAdnan/PulseVAD
reimplementation of KiloVAD
https://arxiv.org/abs/2607.25870
https://github.com/AydinAdnan/PulseVAD
reimplementation of KiloVAD
https://arxiv.org/abs/2607.25870
GitHub
GitHub - AydinAdnan/PulseVAD: Ultra-tiny 2.1k streaming voice activity detector (2.1 KB INT8). 200 ms window for microcontrollers…
Ultra-tiny 2.1k streaming voice activity detector (2.1 KB INT8). 200 ms window for microcontrollers, edge devices, and DSPs. ships with ONNX, TorchScript, and standalone C header. - AydinAdnan/Puls...
We see a huge decline of the interest in speech technology in China but a rise in Europe and India. Recently I discussed it with one of my Chinese friends - he confirms that nobody is interested anymore in plain speech. It has to be multimodal - video, music generation, etc.
"百闻不如一见" — hearing something a hundred times is worse than seeing it once
China is ahead of time here.
"百闻不如一见" — hearing something a hundred times is worse than seeing it once
China is ahead of time here.
Another small 260k params TTS, hard to believe into UTMOS 4.0 based on examples but ok
https://github.com/lab-emi/GrainSpeech
https://lab-emi.github.io/GrainSpeech/
https://arxiv.org/abs/2609.18856
https://github.com/lab-emi/GrainSpeech
https://lab-emi.github.io/GrainSpeech/
https://arxiv.org/abs/2609.18856
GitHub
GitHub - lab-emi/GrainSpeech: Official repository for “GrainSpeech: Less Context, More Detail for Compact Speech Synthesis.”
Official repository for “GrainSpeech: Less Context, More Detail for Compact Speech Synthesis.” - lab-emi/GrainSpeech
An advanced dictation model from Whispr, an interesting part is GPRO over user corrections
https://wisprflow.ai/canto
https://wisprflow.ai/canto
wisprflow.ai
Canto: a speech model built for the real world | Wispr Flow
Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation. On an evaluation of real-world dictations, Canto achieved the lowest word error rate among all the models we tested.
Voice agents related benchmarks somehow passed our attention, here are few recent ones:
https://github.com/sierra-research/tau2-bench
https://github.com/ServiceNow/eva
tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
https://arxiv.org/abs/2603.13686
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
https://arxiv.org/abs/2605.13841
https://github.com/sierra-research/tau2-bench
https://github.com/ServiceNow/eva
tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
https://arxiv.org/abs/2603.13686
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
https://arxiv.org/abs/2605.13841
GitHub
GitHub - sierra-research/tau2-bench: τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains - sierra-research/tau2-bench
Serious issues with Qwen3-TTS stability, long texts and audio prompts break things
https://arxiv.org/abs/2609.16989
https://x.com/RmdW_W/status/2102032894600401231
Taming Long-form Text-to-Speech
Rongxiang Wang, Berkin Durmus, Aysegul Orhon, Eduardo Pacheco, Atila Orhon
https://arxiv.org/abs/2609.16989
https://x.com/RmdW_W/status/2102032894600401231
Taming Long-form Text-to-Speech
Rongxiang Wang, Berkin Durmus, Aysegul Orhon, Eduardo Pacheco, Atila Orhon
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and speaker similarity (SIM) on short-form prompts but significantly deteriorate when used with long-form prompts. We propose Localized Attention-Constrained Inference (LACI), an inference-only method to detect TTS errors in near real-time, roll back to the error onset and regenerate with temporary guardrails, adding negligible computational overhead. Using LACI, we improve worst-of-N WER across 10 RNG seeds for Qwen3-TTS-0.6B from 35.2% to 3.4% on prompts longer than 1500 words, even surpassing its short-form reliability of 5.4\% on prompts with fewer than 500 words. To demonstrate the efficacy of LACI on voice cloning reliability, we propose a sliding-window version of the SIM metric that we call wSIM. wSIM exposes several novel failure patterns that are not captured by SIM. LACI improves worst-of-N wSIM from 0.01 to 0.47 on 120 seconds of reference audio while reducing the rate of catastrophic generations with WER above 30% from 26% to below 1%
arXiv.org
Taming Long-form Text-to-Speech
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models...
https://github.com/SamsungLabs/samsone
Samsone is a family of open Small Audio Language Models (SALMs) for efficient audio understanding. The release includes Samsone-99M, Samsone-134M, and Samsone-356M checkpoints for both on-device and server-side inference.
https://arxiv.org/abs/2609.21666
Samsone: A Family of Open Small Audio Language Models for On-Device Inference
Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski
Samsone is a family of open Small Audio Language Models (SALMs) for efficient audio understanding. The release includes Samsone-99M, Samsone-134M, and Samsone-356M checkpoints for both on-device and server-side inference.
https://arxiv.org/abs/2609.21666
Samsone: A Family of Open Small Audio Language Models for On-Device Inference
Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone.
GitHub
GitHub - SamsungLabs/samsone: Samsone: A Family of Open Small Audio Language Models for On-Device Inference
Samsone: A Family of Open Small Audio Language Models for On-Device Inference - SamsungLabs/samsone
Parakeet v3 ternary quantization with good accuracy (178mb) 113xRT on CPU
https://huggingface.co/moondream/parakeet-redux
https://huggingface.co/moondream/parakeet-redux
huggingface.co
moondream/parakeet-redux · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark
Benchmark for the speech recognition with background speech
Data. 265 recordings in real offices, four call-center floors and cars. Each work and call-center recording follows a fixed structure: primary alone, both speakers, secondary alone, primary alone. Speakers read scripts to give exact ground truth; environments, devices and background speech are real, with no synthetic mixing. Segments are hand-labeled primary / mix / secondary / noise.
Evaluation. 11 STT configurations across 9 engines, streaming and batch, run on raw audio and after voice isolation. WER is computed with jiwer 4.0.0 at corpus level (errors pooled across files, not averaged). Reference and hypothesis go through the same normalization: regex splitting of alphanumeric tokens, then NVIDIA NeMo WFST text normalization, then lowercasing, contraction expansion, punctuation and filler removal. Perceptual quality is scored with DNSMOS-C on primary-speaker segments only.
Results.
Corpus WER: 23.29% → 6.26%
All 11 configurations improve
Engine spread narrows from 17–37% to 4–8%
Clean phone audio regresses slightly: 3.48% → 3.91%
Limitation. On clean narrowband audio there's little to remove, and isolation removes some of the primary signal. We report it rather than filtering it out. Segment-level labels also let you check for deletions specifically: a model that correctly outputs silence and one that drops primary-speaker words can post similar corpus WER.
Dataset: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset
Model outputs: https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark
Benchmark for the speech recognition with background speech
Data. 265 recordings in real offices, four call-center floors and cars. Each work and call-center recording follows a fixed structure: primary alone, both speakers, secondary alone, primary alone. Speakers read scripts to give exact ground truth; environments, devices and background speech are real, with no synthetic mixing. Segments are hand-labeled primary / mix / secondary / noise.
Evaluation. 11 STT configurations across 9 engines, streaming and batch, run on raw audio and after voice isolation. WER is computed with jiwer 4.0.0 at corpus level (errors pooled across files, not averaged). Reference and hypothesis go through the same normalization: regex splitting of alphanumeric tokens, then NVIDIA NeMo WFST text normalization, then lowercasing, contraction expansion, punctuation and filler removal. Perceptual quality is scored with DNSMOS-C on primary-speaker segments only.
Results.
Corpus WER: 23.29% → 6.26%
All 11 configurations improve
Engine spread narrows from 17–37% to 4–8%
Clean phone audio regresses slightly: 3.48% → 3.91%
Limitation. On clean narrowband audio there's little to remove, and isolation removes some of the primary signal. We report it rather than filtering it out. Segment-level labels also let you check for deletions specifically: a model that correctly outputs silence and one that drops primary-speaker words can post similar corpus WER.
Dataset: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset
Model outputs: https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark
huggingface.co
VoiceIsolation Benchmark - a Hugging Face Space by Krisp-AI
This web page lets you see benchmark results for Krisp’s Voice Isolation on speech‑to‑text systems. You can view tables, charts, and listen to audio clips before and after the isolation is applied,...