Skip to content
EN

Training and fine-tuning

Pre-training, post-training, LoRA and the recipes behind better models.

108 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Training and fine-tuning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Vocabulary Transfer for Sparse Retrieval with Advanced Encoders

    The paper attributes advanced encoders’ lag behind BERT-base in learned sparse retrieval to a vocabulary gap: modern tokenizers map semantic units to redundant surface forms. It proposes Vocabulary Transfer to address the mismatch.

    Engineers working on learned sparse retrieval can examine a proposed approach to tokenizer vocabulary mismatch.

  2. Effective Model Pruning Uses Effective Sample Size

    Researchers from the University of Florida and Ohio University introduce Effective Model Pruning (EMP), which uses effective sample size to calculate how many model components to keep. The post says it provides a guarantee that performance loss stays bounded.

    EMP offers an alternative to manually choosing pruning budgets for neural networks and language models.

  3. IW-OPD Reweights Tokens in On-Policy Distillation

    Researchers introduce IW-OPD, which assigns more weight to earlier tokens and less to later ones in on-policy distillation. The post says it converges faster and scores 6.9 points above standard OPD on AIME-2025.

    The method targets performance degradation from later tokens during on-policy distillation.

  4. Addressing the Training–Inference Gap in LLM Reinforcement Learning

    A paper examines training-inference mismatch in LLM post-training, where separate engines can assign inconsistent probabilities to the same trajectories. The post describes TIS and rejecting checkpoint updates when training and inference rewards diverge significantly.

    The approach highlights a way to address training-inference mismatch, while noting that rejection requires additional rollouts.

  5. Learning-rate scaling for LLM training may be nonlinear

    The paper evaluates learning-rate scaling using GPT-2-style models from 22M to 707M parameters trained on 5B to 100B tokens. It reports that the optimal learning rate develops upward ...

    Engineers extrapolating learning rates from smaller runs may need to account for nonlinearity and effective learning rate.

  6. TAPA introduces token-aware phase attention for positional encoding

    The paper presents Token-Aware Phase Attention (TAPA), a positional encoding method with a learnable phase function. It analyzes RoPE’s distance-dependent bias and compares TAPA’s long-context performance with RoPE.

    Engineers working on long-context models can evaluate an alternative positional encoding and its theoretical analysis.

  7. PostTrainBench ranks GLM 5.2 first on its leaderboard

    The post says GLM 5.2 (Max reasoning) scored 34.29% on PostTrainBench, narrowly ahead of Opus 4.8 Max at 34.08%. It reports no failed runs across 84 runs for GLM 5.2.

    The leaderboard offers a comparison of model scores and run reliability.

  8. Video Explains LoRA and Related Fine-Tuning Methods

    A video introduces LoRA and related methods: LoRA+, QLoRA, VeRA, and DoRA.

    Useful as an introduction to several parameter-efficient fine-tuning methods.

  9. Model-free training of a metasurface neural network

    A Nature Communications paper presents a physical neural network built from four programmable metasurfaces. It learns from electromagnetic interactions for wave focusing, object recognition, and localization without digital models.

    It describes a training approach for physical neural networks that avoids relying on digital models.

  10. Why long-horizon RL may favor PPO over GRPO

    A Zhihu post argues that GRPO’s sampled baseline suits short LLM reinforcement-learning tasks, while longer, noisier agentic rollouts make credit assignment harder and can make PPO’s learned value model useful.

    It outlines a tradeoff between avoiding a critic with GRPO and using value modeling for longer-horizon tasks.

  11. Turn-PPO for Multi-Turn Agentic LLM Training

    The paper introduces Turn-PPO, a turn-level advantage estimation method using PPO for multi-turn reinforcement learning in agentic LLMs. It examines limitations of GRPO on long-horizon tasks.

    It describes an alternative to GRPO for training agents on multi-turn tasks.

  12. Why PPO May Suit Long-Horizon Tasks Better Than GRPO

    The post argues that GRPO group synchronization is difficult for training infrastructure and that group-based variance reduction weakens as sequences get longer. It links an arXiv paper and refers to its orange curves.

    The trade-offs described may help engineers choose and scale reinforcement-learning training methods.

  13. ARGUS diagnoses fail-slow issues in large-scale training

    An arXiv paper describes ARGUS, a system that combines CPU stacks, framework semantics, and GPU kernel data to diagnose training slowdowns. It reports under 2% always-on overhead and kernel-event compression of about 3,700×.

    Engineers running distributed training can use its approach to identify stragglers without the overhead of always-on fine-grained profiling.

  14. slime: an open-source framework for LLM post-training

    The THUDM slime repository describes a framework for reinforcement-learning post-training and scaling of large language models.

    Engineers can inspect the code and use it as a starting point for RL post-training workflows.

  15. RL training for broadly beneficial model behavior

    OpenAI reports that reinforcement learning targeting beneficial behavior in realistic scenarios improved alignment across domains and remained effective under adversarial pressure. The post says training used a small share of behavior-focused data alongside ordinary training data.

    The results suggest behavior-focused RL may generalize beyond the domains represented in its training data.

  16. An interaction-based view of supervised fine-tuning in LLMs

    The paper studies why supervised fine-tuning can help smaller networks but have inconsistent or harmful effects in large language models. It uses changes in token interactions to explain SFT’s effects.

    The analysis may help engineers reason about SFT behavior and when to stop training.

  17. TC-JEPA Uses Captions to Guide Visual Representation Learning

    The post describes TC-JEPA, a self-supervised method that uses image captions to guide masked patch prediction. It claims the method improves training stability and visual reasoning compared with contrastive approaches.

    It highlights a way to use text conditioning in self-supervised visual model training.

  18. LeWorldModel trains a pixel-based JEPA on one GPU

    The post describes LeWorldModel, a 15M-parameter JEPA trained end-to-end from raw pixels. It says the method uses two loss terms and plans up to 48 times faster than foundation-model-based world models on control tasks.

    The reported setup and training approach may interest engineers building compact world models.

  19. SAM in Mid-Training to Reduce Fine-Tuning Forgetting

    The post claims that controlling sharpness with SAM in the final ~10% of mid-training can reduce forgetting after fine-tuning or quantization by over 35%, even if base-model quality worsens. It also suggests trying learning rates up to ~10× higher.

    The proposed mid-training recipe may help engineers trade base-model quality for better retention after fine-tuning or quantization.

  20. Fine-tuning LFM2.5-1.2B-Instruct with GRPO

    A tutorial explains GRPO and fine-tunes LFM2.5-1.2B-Instruct with Unsloth for OCR receipt extraction into JSON. It also links to a Kaggle notebook.

    Engineers can use the walkthrough to explore a GRPO fine-tuning workflow for structured extraction.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor