Speech: TTS and ASR
Text-to-speech, speech recognition and voice models you can run yourself.
73 links, newest first.
- Speech: TTS and ASRPost on X
Packed FlashAttention Speeds Up a Transformer Vocoder
A post reports that waveform reconstruction in a transformer vocoder, not autoregressive decoding, was the TTS latency bottleneck. A packed FlashAttention path with variable-length packing, local-causal attention, and RoPE reportedly improved QPS by 48.64% and reduced latency by 32.71%.
It highlights why profiling the full TTS pipeline can reveal optimization targets that differ from LLM decoding bottlenecks.
Open-source real-time voice demo
Hugging Face and Cerebras built a real-time voice demo with open-source models and code. The demo supports WebSocket or WebRTC.
Engineers can test the demo and inspect or modify its code and models.
- Speech: TTS and ASRPost on X
Artificial Analysis introduces a Speech-to-Speech model index
The index combines Big Bench Audio, a Full Duplex Bench subset, and τ-Voice with equal weighting to assess speech reasoning, conversational dynamics, and agentic performance. Models need valid results on all three datasets to be included.
It gives engineers a combined benchmark for comparing native Speech-to-Speech models across three capabilities.
- Speech: TTS and ASRPost on X
LocalVQE Pi V1 targets audio cleanup on Raspberry Pi 5
The post announces a 49k-parameter model for acoustic echo cancellation, noise suppression, and dereverberation, claiming 21× real-time performance on one Raspberry Pi 5 core.
It may interest engineers building local speech-audio processing on Raspberry Pi hardware.
- Speech: TTS and ASRRepository
VoiceGate adds cross-lingual dubbing workflows to ComfyUI
VoiceGate is a video dubbing project built with VoxCPM2 and ComfyUI. It combines ASR, LLM translation, multilingual TTS, and SRT timestamp-based audio alignment.
Engineers can inspect a workflow for coordinating transcription, translation, speech synthesis, and audio-video synchronization.
- Speech: TTS and ASRArticle
Ubuntu introduces Project Myna for desktop dictation
Project Myna aims to bring speech-to-text dictation to Ubuntu Desktop, with a planned Ubuntu 26.10 workflow that inserts spoken text into the active application. The project description says processing will run on local hardware.
Engineers can follow development of an upcoming, locally run dictation feature integrated with the Ubuntu desktop.
- Speech: TTS and ASRPost on X
Cohere announces open-source speech recognition model
Cohere says its open-source speech recognition model, Cohere Transcribe, ranks first on Hugging Face’s new Far-Field ASR benchmark.
The benchmark result may help engineers evaluating speech recognition models.
- Speech: TTS and ASRArticle
Gradium updates TTS pronunciation for structured text
Gradium says its upgraded TTS model improves pronunciation of spelling, acronyms, emails, phone numbers, and codes. The company reports benchmark results against its previous model and several real-time competitors.
Accurate readback of structured text is important for voice agents handling contact details and verification codes.
- Speech: TTS and ASRArticle
Live translation with the Gemini Live API
Google’s Gemini API documentation covers live translation with the Gemini Live API.
Useful to engineers exploring API-based live translation for multilingual experiences.
- Speech: TTS and ASRRepository
UNITTS unifies inference and benchmarking for open-source TTS
UNITTS is a toolkit for running inference and benchmarking across open-source TTS models. The post says it supports Chatterbox and Fish Audio s2-pro.
A unified interface can make it easier to compare and switch between open-source TTS engines.
- Speech: TTS and ASRPost on X
Open ASR Leaderboard Adds Private Evaluation Data
The Open ASR Leaderboard now includes private evaluation data from Appen and DataoceanAI, aiming to reduce test-set contamination and overfitting in speech recognition benchmarks.
Private evaluation data can make ASR benchmark results harder to optimize through test-set leakage.
- Speech: TTS and ASRPost on X
Qwen3-ASR-Enhanced v0.1 Beta Checkpoint Announced
The author describes Qwen3-ASR-Enhanced v0.1 as an early beta checkpoint. Support for nonverbal tags is not yet stable, and a revised version is planned.
Engineers evaluating ASR models can track an early checkpoint and its known limitation with nonverbal tags.
- Speech: TTS and ASRRepository
FireRedVAD supports voice activity detection in 100+ languages
FireRedVAD is a GitHub project for voice activity detection and audio event detection. Its description says it supports more than 100 languages.
Engineers building speech pipelines can evaluate it for voice activity and audio event detection.
- Speech: TTS and ASRPost on X
Microsoft VibeVoice-ASR supports long-form transcription
Microsoft says VibeVoice-ASR can transcribe 60-minute audio in one pass, with speaker diarization, timestamps, and hotwords. It supports more than 50 languages without requiring a language setting.
The stated capabilities may make it useful for engineers building multilingual speech-recognition pipelines.
- Speech: TTS and ASRRepository
C Inference for Qwen3-ASR Transcription Models
The antirez/qwen-asr repository provides C inference for the Qwen3-ASR 0.6B and 1.7B transcription models.
It offers a C-based option for running speech recognition models in applications such as transcription bots and voice-driven interfaces.
- Speech: TTS and ASRArticle
vLLM Adds Streaming Inputs and a Realtime WebSocket API
The vLLM blog describes support for streamable inputs and a Realtime WebSocket API for audio, video, robotics, and low-latency applications that process inputs incrementally.
Engineers building low-latency audio systems can review vLLM’s streaming interface and examples.
- Speech: TTS and ASRRepository
Local Studio PR adds ROCm and speech support
A pull request for Local Studio describes ROCm support and TTS/STT call mode using ACE-Step and Whisper v3 Large. It also mentions image and video support.
Engineers can review the proposed speech and AMD GPU support in the pull request.
- Speech: TTS and ASRRepository
Hugging Face Open-Source Speech-to-Speech Repository
Hugging Face’s speech-to-speech repository is for building voice agents with open-source models.
It gives engineers a starting point for building voice agents with open-source models.
- Speech: TTS and ASRPost on X
VoxCPM: tokenizer-free text-to-speech system
VoxCPM is a tokenizer-free TTS system built on the MiniCPM-4 backbone, using an end-to-end diffusion-autoregressive architecture. The post describes zero-shot voice cloning and expressive speech generation.
Engineers exploring TTS can assess an approach that generates continuous speech representations without discrete tokens.
vLLM adds support for Qwen3-ASR
vLLM has day-one support for Qwen3-ASR. The linked usage guide covers running the model with vLLM.
The guide gives engineers a starting point for serving Qwen3-ASR with vLLM.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor










