Training and fine-tuning
Pre-training, post-training, LoRA and the recipes behind better models.
108 links, newest first.
- Training and fine-tuningRepository
Pocket TTS Releases Its Training Stack
Kyutai open-sourced Pocket TTS’s training stack, including its data pipeline, training recipes, and evaluations. The project is designed to train a text-to-speech model and run it on CPU.
Engineers can inspect and adapt an end-to-end TTS training workflow, from data preparation through evaluation.
- Training and fine-tuningPost on X
Marin exposes details of its model training
The post says Marin provides views of its pre-training mixture by domain, sampled documents, live training loss, configs, and scaling laws.
These materials can help engineers inspect a model training run and its underlying choices.
Slides on Post-Training Methods and Real-World Adaptation
Slides from a talk introducing economists to post-training outline SFT, offline and online methods, RL environments, distillation, and world adaptation. They raise questions about anticipating real-world adaptation and scaling human oversight.
Useful overview of post-training methods and open research questions for deploying agents in changing environments.
- Training and fine-tuningArticle
A Training Log for Pretraining a Mini Kimi K3
A free book documents an attempt to pretrain a 1.02B-parameter Mini Kimi K3 on one H200. It covers the model design, data pipeline, training bugs, and attempts to improve throughput.
Its detailed failure logs and training decisions can help engineers understand the practical challenges of pretraining MoE models.
- Training and fine-tuningRepository
datatrove 0.10.0 adds Hugging Face Jobs pipeline execution
datatrove 0.10.0 adds a JobsPipelineExecutor for running pipelines on Hugging Face Jobs, support for HF storage buckets as a DataFolder, and preservation of reasoning outputs in inference results.
Engineers building data pipelines can run multi-stage jobs with retries and resume without a Slurm cluster.
ResidencyRL trains clinical agents in simulated patient encounters
The paper describes reinforcement learning in simulated clinical environments, modeling residency through patient encounters and feedback. The preview notes that clinical reasoning involves gathering history, refining diagnostic hypotheses, and making decisions under uncertainty.
It explores a training setup for developing clinical reasoning beyond static medical benchmarks.
EnvACE trains agents through internal world rehearsal
EnvACE is an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. Its policy alternates between generating tool calls and rehearsal.
It presents an approach to reduce reliance on costly environments or difficult-to-ground external simulators.
- Training and fine-tuningPost on X
Ostris AI Toolkit Adds LoRA Training for MiniMax H3
Ostris AI Toolkit supports training LoRAs for MiniMax H3, currently limited to text-to-video and image-to-video. The post says it uses Comfy quantized weights.
Engineers working on video-model fine-tuning can check the toolkit’s current support and limitations.
- Training and fine-tuningPost on X
A First-Principles Guide to JEPA’s SIGReg
The post describes a guide to JEPA’s SIGReg that covers complex numbers, Euler’s formula, the Cramér–Wold theorem, and a working training loop.
It offers a step-by-step explanation of concepts and a training loop relevant to model training.
- Training and fine-tuningPost on X
The Stack v3 releases 5T deduplicated code tokens
The Stack v3 is an open code dataset with about 5T tokens of deduplicated and filtered source code across 713 languages. The post describes it as excluding restrictively licensed code.
Engineers training code models can evaluate a larger dataset with broad language coverage.
- Training and fine-tuningDataset
The Stack v3 releases a 5T-token code dataset
The Stack v3 is a code dataset with about 5 trillion tokens across 713 filtered languages and 224 million repositories. It embeds source code directly and is based on a GitHub re-crawl completed by August 2025.
Engineers can use the dataset as a source of code for training and fine-tuning models.
A Study of Layer-Wise Reinforcement Learning for LLMs
The preprint examines how reinforcement-learning adaptation is distributed across transformer layers during LLM post-training. It challenges the assumption that all layers contribute similarly to RL gains.
It may inform decisions about which model parameters to update during RL post-training.
RLSD uses self-distillation to scale token-level RLVR updates
The RLSD paper combines environment rewards with a self-distillation signal to scale token updates within a trajectory. On Qwen3-VL-8B, it reports higher mean accuracy than the base model and GRPO across five multimodal reasoning benchmarks.
The approach offers an alternative to GRPO’s uniform sequence-level advantage for training reasoning models.
- Training and fine-tuningPost on X
Dual On-Policy Distillation Routes Supervision Per Token
The post describes a paper on Dual On-Policy Distillation, which dynamically routes each token between teacher-based and self-based supervision.
It may interest engineers exploring alternatives to standard teacher distillation and self-distillation.
Lightweight Fine-Tuning for ColBERT Vector Compression
The paper presents pooling-aware fine-tuning for ColBERT models to reduce the number of token vectors per document. Its preview describes lightweight k-means pooling-aware fine-tuning as enabling compression with no accuracy loss.
It may help engineers reduce vector storage and memory costs in late-interaction retrieval systems.
Tapered Language Models Allocate Capacity Across Depth
The linked paper examines uniform parameter allocation across model layers. The post reports that a cosine taper, with more capacity in early layers, improved perplexity across several architectures and model scales without extra parameters or FLOPs.
It explores whether layer-wise capacity allocation can improve model quality without increasing parameter or compute budgets.
Pruning vs. scratch training for small LLMs
The paper compares six methods for pruning Llama-3.1-8B with scratch training under two token-budget settings. Pruned models lead when retraining tokens are matched; with the full pipeline budget, results depend on pruning granularity.
The controlled comparisons help engineers assess when pruning is a useful alternative to training a small model from scratch.
- Training and fine-tuningPost on X
Local Branch Routing for Test-Time Language Model Scaling
The paper proposes previewing and branching over candidate next tokens, then routing to one branch before committing. The post reports improved math reasoning over CoT, standard RLVR, and soft-token branching.
The approach offers an alternative way to add test-time branching while keeping the model trainable.
- Training and fine-tuningPost on X
Single-Layer Training for Transformer RL Post-Training
The post describes a paper that trains or boosts high-contribution middle layers during RL post-training. It reports that this can match or outperform full-parameter RL with fewer changes.
It suggests a way to reduce the number of model parameters changed during RL post-training.
- Training and fine-tuningRepository
Online DSpark training example in Speculators
The Speculators repository includes a shell script for online DSpark training with Qwen3 0.6B and ShareGPT. The project is a library for building, evaluating, and storing speculative decoding algorithms for vLLM.
Engineers can inspect a concrete training example in a library for speculative decoding in vLLM.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor


