Skip to content
EN

Speech: TTS and ASR

Text-to-speech, speech recognition and voice models you can run yourself.

73 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Speech: TTS and ASR

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Packed FlashAttention Speeds Up a Transformer Vocoder

    A post reports that waveform reconstruction in a transformer vocoder, not autoregressive decoding, was the TTS latency bottleneck. A packed FlashAttention path with variable-length packing, local-causal attention, and RoPE reportedly improved QPS by 48.64% and reduced latency by 32.71%.

    It highlights why profiling the full TTS pipeline can reveal optimization targets that differ from LLM decoding bottlenecks.

  2. Open-source real-time voice demo

    Hugging Face and Cerebras built a real-time voice demo with open-source models and code. The demo supports WebSocket or WebRTC.

    Engineers can test the demo and inspect or modify its code and models.

  3. Artificial Analysis introduces a Speech-to-Speech model index

    The index combines Big Bench Audio, a Full Duplex Bench subset, and τ-Voice with equal weighting to assess speech reasoning, conversational dynamics, and agentic performance. Models need valid results on all three datasets to be included.

    It gives engineers a combined benchmark for comparing native Speech-to-Speech models across three capabilities.

  4. LocalVQE Pi V1 targets audio cleanup on Raspberry Pi 5

    The post announces a 49k-parameter model for acoustic echo cancellation, noise suppression, and dereverberation, claiming 21× real-time performance on one Raspberry Pi 5 core.

    It may interest engineers building local speech-audio processing on Raspberry Pi hardware.

  5. VoiceGate adds cross-lingual dubbing workflows to ComfyUI

    VoiceGate is a video dubbing project built with VoxCPM2 and ComfyUI. It combines ASR, LLM translation, multilingual TTS, and SRT timestamp-based audio alignment.

    Engineers can inspect a workflow for coordinating transcription, translation, speech synthesis, and audio-video synchronization.

  6. Ubuntu introduces Project Myna for desktop dictation

    Project Myna aims to bring speech-to-text dictation to Ubuntu Desktop, with a planned Ubuntu 26.10 workflow that inserts spoken text into the active application. The project description says processing will run on local hardware.

    Engineers can follow development of an upcoming, locally run dictation feature integrated with the Ubuntu desktop.

  7. Cohere announces open-source speech recognition model

    Cohere says its open-source speech recognition model, Cohere Transcribe, ranks first on Hugging Face’s new Far-Field ASR benchmark.

    The benchmark result may help engineers evaluating speech recognition models.

  8. Gradium updates TTS pronunciation for structured text

    Gradium says its upgraded TTS model improves pronunciation of spelling, acronyms, emails, phone numbers, and codes. The company reports benchmark results against its previous model and several real-time competitors.

    Accurate readback of structured text is important for voice agents handling contact details and verification codes.

  9. Live translation with the Gemini Live API

    Google’s Gemini API documentation covers live translation with the Gemini Live API.

    Useful to engineers exploring API-based live translation for multilingual experiences.

  10. UNITTS unifies inference and benchmarking for open-source TTS

    UNITTS is a toolkit for running inference and benchmarking across open-source TTS models. The post says it supports Chatterbox and Fish Audio s2-pro.

    A unified interface can make it easier to compare and switch between open-source TTS engines.

  11. Open ASR Leaderboard Adds Private Evaluation Data

    The Open ASR Leaderboard now includes private evaluation data from Appen and DataoceanAI, aiming to reduce test-set contamination and overfitting in speech recognition benchmarks.

    Private evaluation data can make ASR benchmark results harder to optimize through test-set leakage.

  12. Qwen3-ASR-Enhanced v0.1 Beta Checkpoint Announced

    The author describes Qwen3-ASR-Enhanced v0.1 as an early beta checkpoint. Support for nonverbal tags is not yet stable, and a revised version is planned.

    Engineers evaluating ASR models can track an early checkpoint and its known limitation with nonverbal tags.

  13. FireRedVAD supports voice activity detection in 100+ languages

    FireRedVAD is a GitHub project for voice activity detection and audio event detection. Its description says it supports more than 100 languages.

    Engineers building speech pipelines can evaluate it for voice activity and audio event detection.

  14. Microsoft VibeVoice-ASR supports long-form transcription

    Microsoft says VibeVoice-ASR can transcribe 60-minute audio in one pass, with speaker diarization, timestamps, and hotwords. It supports more than 50 languages without requiring a language setting.

    The stated capabilities may make it useful for engineers building multilingual speech-recognition pipelines.

  15. C Inference for Qwen3-ASR Transcription Models

    The antirez/qwen-asr repository provides C inference for the Qwen3-ASR 0.6B and 1.7B transcription models.

    It offers a C-based option for running speech recognition models in applications such as transcription bots and voice-driven interfaces.

  16. vLLM Adds Streaming Inputs and a Realtime WebSocket API

    The vLLM blog describes support for streamable inputs and a Realtime WebSocket API for audio, video, robotics, and low-latency applications that process inputs incrementally.

    Engineers building low-latency audio systems can review vLLM’s streaming interface and examples.

  17. Local Studio PR adds ROCm and speech support

    A pull request for Local Studio describes ROCm support and TTS/STT call mode using ACE-Step and Whisper v3 Large. It also mentions image and video support.

    Engineers can review the proposed speech and AMD GPU support in the pull request.

  18. Hugging Face Open-Source Speech-to-Speech Repository

    Hugging Face’s speech-to-speech repository is for building voice agents with open-source models.

    It gives engineers a starting point for building voice agents with open-source models.

  19. VoxCPM: tokenizer-free text-to-speech system

    VoxCPM is a tokenizer-free TTS system built on the MiniCPM-4 backbone, using an end-to-end diffusion-autoregressive architecture. The post describes zero-shot voice cloning and expressive speech generation.

    Engineers exploring TTS can assess an approach that generates continuous speech representations without discrete tokens.

  20. vLLM adds support for Qwen3-ASR

    vLLM has day-one support for Qwen3-ASR. The linked usage guide covers running the model with vLLM.

    The guide gives engineers a starting point for serving Qwen3-ASR with vLLM.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor