Parakeet v3 ternary quantization with good accuracy (178mb) 113xRT on CPU
https://huggingface.co/moondream/parakeet-redux
https://huggingface.co/moondream/parakeet-redux
huggingface.co
moondream/parakeet-redux · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark
Benchmark for the speech recognition with background speech
Data. 265 recordings in real offices, four call-center floors and cars. Each work and call-center recording follows a fixed structure: primary alone, both speakers, secondary alone, primary alone. Speakers read scripts to give exact ground truth; environments, devices and background speech are real, with no synthetic mixing. Segments are hand-labeled primary / mix / secondary / noise.
Evaluation. 11 STT configurations across 9 engines, streaming and batch, run on raw audio and after voice isolation. WER is computed with jiwer 4.0.0 at corpus level (errors pooled across files, not averaged). Reference and hypothesis go through the same normalization: regex splitting of alphanumeric tokens, then NVIDIA NeMo WFST text normalization, then lowercasing, contraction expansion, punctuation and filler removal. Perceptual quality is scored with DNSMOS-C on primary-speaker segments only.
Results.
Corpus WER: 23.29% → 6.26%
All 11 configurations improve
Engine spread narrows from 17–37% to 4–8%
Clean phone audio regresses slightly: 3.48% → 3.91%
Limitation. On clean narrowband audio there's little to remove, and isolation removes some of the primary signal. We report it rather than filtering it out. Segment-level labels also let you check for deletions specifically: a model that correctly outputs silence and one that drops primary-speaker words can post similar corpus WER.
Dataset: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset
Model outputs: https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark
Benchmark for the speech recognition with background speech
Data. 265 recordings in real offices, four call-center floors and cars. Each work and call-center recording follows a fixed structure: primary alone, both speakers, secondary alone, primary alone. Speakers read scripts to give exact ground truth; environments, devices and background speech are real, with no synthetic mixing. Segments are hand-labeled primary / mix / secondary / noise.
Evaluation. 11 STT configurations across 9 engines, streaming and batch, run on raw audio and after voice isolation. WER is computed with jiwer 4.0.0 at corpus level (errors pooled across files, not averaged). Reference and hypothesis go through the same normalization: regex splitting of alphanumeric tokens, then NVIDIA NeMo WFST text normalization, then lowercasing, contraction expansion, punctuation and filler removal. Perceptual quality is scored with DNSMOS-C on primary-speaker segments only.
Results.
Corpus WER: 23.29% → 6.26%
All 11 configurations improve
Engine spread narrows from 17–37% to 4–8%
Clean phone audio regresses slightly: 3.48% → 3.91%
Limitation. On clean narrowband audio there's little to remove, and isolation removes some of the primary signal. We report it rather than filtering it out. Segment-level labels also let you check for deletions specifically: a model that correctly outputs silence and one that drops primary-speaker words can post similar corpus WER.
Dataset: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset
Model outputs: https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark
huggingface.co
VoiceIsolation Benchmark - a Hugging Face Space by Krisp-AI
This web page lets you see benchmark results for Krisp’s Voice Isolation on speech‑to‑text systems. You can view tables, charts, and listen to audio clips before and after the isolation is applied,...
Facebook's Muse Transcribe is 1st on AA leaderboard and 25th on Huggingface leaderboard
https://x.com/EricBezzam/status/2102368894929514711
https://x.com/EricBezzam/status/2102368894929514711
X (formerly Twitter)
Eric Bezzam (@EricBezzam) on X
25th overall, and 27th on our private sets
Some recent advanced TTS evaluation
Blog by @altsoph from Inworld
https://altsoph.substack.com/p/wtf-is-voice-steering
J-HARD-TTS-Eval, a benchmark designed to evaluate the robustness of autoregressive Japanese Text-To-Speech (TTS) models (V2 is coming)
https://github.com/Parakeet-Inc/J-HARD-TTS-Eval
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
https://arxiv.org/abs/2505.23009
Blog by @altsoph from Inworld
https://altsoph.substack.com/p/wtf-is-voice-steering
J-HARD-TTS-Eval, a benchmark designed to evaluate the robustness of autoregressive Japanese Text-To-Speech (TTS) models (V2 is coming)
https://github.com/Parakeet-Inc/J-HARD-TTS-Eval
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
https://arxiv.org/abs/2505.23009
Substack
WTF is Voice Steering?
TLDR: We trained our new TTS model to follow free-form voice instructions, then needed a way to measure how well it followed them.
Yodas3 was shared yesterday. 55TB, 1.1M hours of speech data
https://huggingface.co/datasets/espnet/yodas3
from previous experience not very useful to be honest
https://huggingface.co/datasets/espnet/yodas3
from previous experience not very useful to be honest
huggingface.co
espnet/yodas3 · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Highly efficient audio inference engine
https://github.com/pegainfer-project/pega-omni
custom kernel for Mimi codec, etc, 128 Personaplex streams in realtime on GB300, should be like 30 streams on 4090
https://github.com/pegainfer-project/pega-omni
custom kernel for Mimi codec, etc, 128 Personaplex streams in realtime on GB300, should be like 30 streams on 4090
GitHub
GitHub - pegainfer-project/pega-omni: OpenAI-compatible speech serving in Rust: streaming TTS and full-duplex voice over GPT-Live.…
OpenAI-compatible speech serving in Rust: streaming TTS and full-duplex voice over GPT-Live. 128 concurrent PersonaPlex sessions on one GPU. - pegainfer-project/pega-omni
Interspeech 2026 starts today
https://www.isca-archive.org/interspeech_2026/
Let us be brave to read all this. Surprisingly, very few agentic papers.
https://www.isca-archive.org/interspeech_2026/
Let us be brave to read all this. Surprisingly, very few agentic papers.
Good Paper from Interspeech and dataset for game developers
https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset
https://ncai-official.github.io/speech/publications/designed-vocalizations-dataset/index.html
Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations such as monster growls and robotic voices underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, created by applying professional vocal effects processing to diverse vocal sources to produce paired original and effect-modified audio. We further provide a standardized test set with explicit seen/unseen splits over source types and preset styles to assess generalization under controlled conditions, together with baseline benchmark results for reproducible evaluation.
https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset
https://ncai-official.github.io/speech/publications/designed-vocalizations-dataset/index.html
Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations such as monster growls and robotic voices underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, created by applying professional vocal effects processing to diverse vocal sources to produce paired original and effect-modified audio. We further provide a standardized test set with explicit seen/unseen splits over source types and preset styles to assess generalization under controlled conditions, together with baseline benchmark results for reproducible evaluation.
huggingface.co
NCSOFT/Designed-Vocalizations-Dataset · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
New update from PocketTTS, different objective improves things
https://kyutai.org/blog/2026-09-28-pocket-tts-drifting/
https://kyutai.org/blog/2026-09-28-pocket-tts-drifting/
kyutai.org
Building a Pocket TTS with a drifting objective
Our mission is to build and democratize artificial general intelligence through open science.
You can inject directly into KV cache ;)
Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
https://arxiv.org/abs/2609.30784
Popular thing in modern LLMs, Mostik is doing similar research, also
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
https://arxiv.org/abs/2510.03215
Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
https://arxiv.org/abs/2609.30784
Popular thing in modern LLMs, Mostik is doing similar research, also
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
https://arxiv.org/abs/2510.03215
arXiv.org
Symbiotic Architecture for Post-Hoc Audio Extension of Frozen...
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs...
The Conversational AI Reading Group is back for Fall 2026!
Over the past two years, we've hosted more than 50 talks from researchers in academia and industry working on speech, audio, and conversational AI, and we're excited to continue this season.
We're kicking off this week with Sreyan Ghosh who will talk about building open audio-language models:
Towards Fully Open General Audio Intelligence
Thursday, October 8 | 11:00 AM – 12:00 PM ET
Speaker: Sreyan Ghosh - Google DeepMind
Details, Zoom link, and the full schedule of upcoming talks: https://poonehmousavi.github.io/rg.html
Youtube: https://www.youtube.com/@CONVAI_RG
Over the past two years, we've hosted more than 50 talks from researchers in academia and industry working on speech, audio, and conversational AI, and we're excited to continue this season.
We're kicking off this week with Sreyan Ghosh who will talk about building open audio-language models:
Towards Fully Open General Audio Intelligence
Thursday, October 8 | 11:00 AM – 12:00 PM ET
Speaker: Sreyan Ghosh - Google DeepMind
Details, Zoom link, and the full schedule of upcoming talks: https://poonehmousavi.github.io/rg.html
Youtube: https://www.youtube.com/@CONVAI_RG
Linkedin
Sreyan Ghosh - Google DeepMind | LinkedIn
Website: https://sreyan88.github.io/
I am Sreyan Ghosh, a Research Scientist at… · Experience: Google DeepMind · Education: University of Maryland · Location: United States · 500+ connections on LinkedIn. View Sreyan Ghosh’s profile on LinkedIn, a professional…
I am Sreyan Ghosh, a Research Scientist at… · Experience: Google DeepMind · Education: University of Maryland · Location: United States · 500+ connections on LinkedIn. View Sreyan Ghosh’s profile on LinkedIn, a professional…
Our friend @RND_RandoM recently released a cool full-duplex voice agent setup which that rivals GPT-Live
https://github.com/speakrail/speakrail
The pipeline consists of:
STT: Voxtral Realtime with an attached turn head (HF), running on our audio.cpp fork.
LLM: Gemma 4 12B QAT with microturn finetuning (HF). It is chosen because it fits the GPU quite well, has vision support (I want to test it soon), and in general, the Gemma models perform well in real-life tasks, general chatting, etc.
TTS: Breeze TTS 2, patched to run at int8 (GitHub fork), although it can be replaced by any streaming TTS.
The harness itself: it is the glue between all the components, and has many latency-saving measures, like speculative LLM+TTS firing (inspired by HF speech-to-speech).
The cool thing is that Danil took inspiration from several "think while talking" papers (e.g. SHANKS): while you are talking, a base Gemma 4 12B int4 writes thinking notes, which are then passed to the talker. It helps with harder tasks that require more reasoning.
Please try it out
https://github.com/speakrail/speakrail
The pipeline consists of:
STT: Voxtral Realtime with an attached turn head (HF), running on our audio.cpp fork.
LLM: Gemma 4 12B QAT with microturn finetuning (HF). It is chosen because it fits the GPU quite well, has vision support (I want to test it soon), and in general, the Gemma models perform well in real-life tasks, general chatting, etc.
TTS: Breeze TTS 2, patched to run at int8 (GitHub fork), although it can be replaced by any streaming TTS.
The harness itself: it is the glue between all the components, and has many latency-saving measures, like speculative LLM+TTS firing (inspired by HF speech-to-speech).
The cool thing is that Danil took inspiration from several "think while talking" papers (e.g. SHANKS): while you are talking, a base Gemma 4 12B int4 writes thinking notes, which are then passed to the talker. It helps with harder tasks that require more reasoning.
Please try it out
GitHub
GitHub - speakrail/speakrail: Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090
Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090 - speakrail/speakrail