Speech Technology
1.72K subscribers
140 photos
4 videos
1 file
2.24K links
Download Telegram
Some recent advanced TTS evaluation

Blog by @altsoph from Inworld
https://altsoph.substack.com/p/wtf-is-voice-steering

J-HARD-TTS-Eval, a benchmark designed to evaluate the robustness of autoregressive Japanese Text-To-Speech (TTS) models (V2 is coming)
https://github.com/Parakeet-Inc/J-HARD-TTS-Eval

EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
https://arxiv.org/abs/2505.23009
Interspeech 2026 starts today

https://www.isca-archive.org/interspeech_2026/

Let us be brave to read all this. Surprisingly, very few agentic papers.
Good Paper from Interspeech and dataset for game developers

https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset

https://ncai-official.github.io/speech/publications/designed-vocalizations-dataset/index.html

Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations such as monster growls and robotic voices underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, created by applying professional vocal effects processing to diverse vocal sources to produce paired original and effect-modified audio. We further provide a standardized test set with explicit seen/unseen splits over source types and preset styles to assess generalization under controlled conditions, together with baseline benchmark results for reproducible evaluation.
You can inject directly into KV cache ;)

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
https://arxiv.org/abs/2609.30784

Popular thing in modern LLMs, Mostik is doing similar research, also

Cache-to-Cache: Direct Semantic Communication Between Large Language Models
https://arxiv.org/abs/2510.03215
The Conversational AI Reading Group is back for Fall 2026!

Over the past two years, we've hosted more than 50 talks from researchers in academia and industry working on speech, audio, and conversational AI, and we're excited to continue this season.

We're kicking off this week with Sreyan Ghosh who will talk about building open audio-language models:

Towards Fully Open General Audio Intelligence
Thursday, October 8 | 11:00 AM – 12:00 PM ET
Speaker: Sreyan Ghosh - Google DeepMind
Details, Zoom link, and the full schedule of upcoming talks: https://poonehmousavi.github.io/rg.html
Youtube: https://www.youtube.com/@CONVAI_RG
Our friend @RND_RandoM recently released a cool full-duplex voice agent setup which that rivals GPT-Live

https://github.com/speakrail/speakrail

The pipeline consists of:

STT: Voxtral Realtime with an attached turn head (HF), running on our audio.cpp fork.
LLM: Gemma 4 12B QAT with microturn finetuning (HF). It is chosen because it fits the GPU quite well, has vision support (I want to test it soon), and in general, the Gemma models perform well in real-life tasks, general chatting, etc.
TTS: Breeze TTS 2, patched to run at int8 (GitHub fork), although it can be replaced by any streaming TTS.
The harness itself: it is the glue between all the components, and has many latency-saving measures, like speculative LLM+TTS firing (inspired by HF speech-to-speech).

The cool thing is that Danil took inspiration from several "think while talking" papers (e.g. SHANKS): while you are talking, a base Gemma 4 12B int4 writes thinking notes, which are then passed to the talker. It helps with harder tasks that require more reasoning.

Please try it out