Speech: TTS and ASR
Text-to-speech, speech recognition and voice models you can run yourself.
73 links, newest first.
- Speech: TTS and ASRPost on X
Qwen3-TTS announced as open source
The post announces that Qwen3-TTS is being open-sourced and mentions voice cloning and voice design. It does not include a repository or other link.
Engineers evaluating self-hosted TTS may want to investigate the model and its voice features.
- Speech: TTS and ASRModel
Step Audio R1.1 for Real-Time Speaking Dialogue
The post describes Step Audio R1.1 as a real-time, interactive speaking dialogue model under Apache 2.0. It claims the model can think while speaking and uses a dual-brain architecture for low latency.
Engineers can inspect the model page and evaluate its license and architecture claims for speech applications.
- Speech: TTS and ASRPost on X
Pocket TTS: 100M-Parameter TTS for Laptops
Kyutai Labs introduces Pocket TTS, a 100M-parameter text-to-speech model that the post says supports voice cloning and runs on a laptop without a GPU. The post describes it as open-source.
Engineers can evaluate a self-hostable TTS model designed to run without a GPU.
- Speech: TTS and ASRArticle
LEMAS releases a 150K-hour multilingual speech dataset and models
LEMAS describes a 150K-hour multilingual audio suite and generative speech models. The post says the dataset covers 10 languages with word-level timestamps and names LEMAS-TTS and LEMAS-Edit.
Engineers can explore speech data and models for multilingual TTS and speech editing.
- Speech: TTS and ASRRepository
Meta releases Omnilingual ASR for 1,600+ languages
Meta’s Omnilingual ASR is an open-source speech recognition model suite and dataset for more than 1,600 languages, including 500 described as previously unsupported.
Engineers can explore an open-source ASR system aimed at broad multilingual coverage.
- Speech: TTS and ASRRepository
Step-Audio-EditX for Prompt-Based Audio Editing
StepFun says Step-Audio-EditX is a 3B-parameter model for audio editing, zero-shot multilingual TTS, and prompt-based control of emotion, speaking style, and vocal elements such as breaths and laughs. The post says it supports single-GPU deployment and uses Apache 2.0.
Engineers can evaluate a single-model approach to speech generation and iterative audio editing.
- Speech: TTS and ASRArticle
Inworld TTS 1 Max leads the Artificial Analysis Speech Arena
Artificial Analysis reports that Inworld TTS 1 Max leads its Speech Arena leaderboard, where users compare generated speech and choose their preferred output without seeing the model names. Inworld says TTS Max costs $10 per million characters.
The leaderboard describes a human-preference evaluation across four real-world speech categories.
- Speech: TTS and ASRPost on X
Inworld TTS 1 Max leads the Artificial Analysis Speech Arena
The post reports that Inworld TTS 1 Max leads the Speech Arena leaderboard, which ranks TTS models through blind human comparisons. It describes both Inworld models’ 12-language support, short-audio voice cloning, and voice tags.
The comparison and feature details may help engineers assess TTS options.
- Speech: TTS and ASRRepository
LongCat-Audio-Codec for Speech LLMs
Meituan open-sourced an audio tokenizer and detokenizer optimized for Speech LLMs. The post describes parallel semantic and acoustic tokens at 16.7 Hz, low-bitrate encoding, and a low-latency streaming decoder.
Engineers can evaluate its audio tokenization and streaming decoder for Speech LLM applications.
- Speech: TTS and ASRPost on X
DiaMoE-TTS uses IPA and MoE for dialect speech synthesis
DiaMoE-TTS is an IPA-based dialect TTS framework with a dialect-aware Mixture-of-Experts and LoRA and Conditioning Adapters for adaptation. The post reports zero-shot synthesis on unseen dialects, including Peking Opera, with a few hours of data.
Its phonetic representation and adaptation approach may help engineers build TTS systems for dialects with limited data.
- Speech: TTS and ASRPost on X
Gemini 2.5 Native Audio Thinking on Big Bench Audio
Artificial Analysis reports that Gemini 2.5 Native Audio Thinking scored 92% on its Big Bench Audio benchmark, which uses 1,000 audio questions adapted from Big Bench Hard. It reports 3.87 seconds average time to first token for the thinking model and 0.63 seconds for its non-thinking equivalent.
The results offer a reasoning benchmark and latency comparison for speech-to-speech models.
- Speech: TTS and ASRPost on X
FireRedChat: Open-Source Full-Duplex Voice Interaction
FireRedChat is a pluggable voice interaction system with cascaded and semi-cascaded implementations. The post says it supports interruption handling, endpoint detection, real-time responses, emotion perception, and expressive speech synthesis.
Engineers can evaluate its architecture and voice interaction features for conversational assistants.
- Speech: TTS and ASRRepository
Qwen3-Omni open-sources multimodal models with speech I/O
Qwen says it has open-sourced three Qwen3-Omni models for instruction following, reasoning, and captioning. The model combines text, image, audio, and video, with speech input and output.
Engineers can explore an open-source model that handles speech alongside other modalities.
- Speech: TTS and ASRPost on X
Mini-Omni-Reasoner interleaves reasoning and speech
Mini-Omni-Reasoner uses a hierarchical Thinker–Talker architecture to interleave silent reasoning with spoken tokens. The post reports a new Spoken-Math-Problems-3M dataset and gains in arithmetic reasoning and contextual understanding.
The approach and reported results may inform designs for speech models that reason while speaking.
- Speech: TTS and ASRModel
FireRedTTS 2 open-source text-to-speech model
FireRedTTS 2 is an open TTS model on Hugging Face. The post describes long-form conversational speech, multilingual and cross-lingual voice cloning, low-latency output, and random timbre generation.
Engineers can evaluate a self-hostable TTS model for multilingual speech generation and voice cloning.
- Speech: TTS and ASRRepository
Qwen3-ASR-Toolkit handles long-audio transcription
Qwen3-ASR-Toolkit is a Python CLI for the Qwen3-ASR API. It supports long audio and video transcription using VAD splitting, parallel API calls, multiple media formats, and automatic resampling.
It provides a way to process long recordings through the Qwen3-ASR API with parallel calls and media-format support.
- Speech: TTS and ASRPost on X
FireRedTTS-2 Uses Streaming Tokens for Multi-Speaker Dialogue
FireRedTTS-2 is described as a long-form streaming TTS system using a 12.5 Hz speech tokenizer, a dual-transformer architecture, and interleaved text–speech input for context-aware multi-speaker dialogue. The post claims it outperforms several systems on intelligibility, speaker turns, and…
Its streaming design and multi-speaker format may be relevant to engineers building conversational speech generation.
- Speech: TTS and ASRArticle
Deploy streaming Kyutai STT on Modal
Modal provides an example for deploying a streaming audio transcription service with Kyutai STT.
It offers engineers a concrete reference for running streaming speech recognition.
- Speech: TTS and ASRPost on X
Zonos TTS Beta and Cross-Platform Local Launcher
Zonos announced a beta expressive TTS model with voice cloning, releasing transformer and SSM-hybrid models under Apache 2.0. The post also presents a one-click launcher for Mac, Windows, and Linux.
Engineers can explore a locally runnable voice-cloning TTS model and its available model architectures.
- Speech: TTS and ASRRepository
Hibiki: a model for streaming speech translation
Hibiki is a decoder-only model for simultaneous French–English speech translation. Its synthetic training corpus uses translations from MADLAD, the authors’ TTS, and a simple lag rule.
The repository offers engineers a model for streaming speech translation to evaluate and run.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor








