Reinforcement learning
Rewards, policies, environments and results, for language models and beyond.
79 links, newest first.
LED uses intermediate-layer diversity to restore exploration
The paper reports that RL post-training reduces entropy in the final-layer distribution while intermediate layers retain higher entropy. Its training-free Latent Exploration Decoding method uses intermediate-layer distributions to improve sampling in reasoning models.
Engineers can evaluate a decoding approach aimed at improving pass@n without further training.
Single-Rollout Asynchronous Optimization for Agentic RL
The linked paper examines asynchronous RL for post-training language models, where models are updated as rollouts arrive. The post says GLM 5.2 PPO uses async RL improvements including clipping adjustments, and also links to VAPO.
Engineers working on agentic RL can review approaches to asynchronous training and stability.
Privileged Self-Distillation Can Degrade Long Reasoning Traces
The paper studies self-distillation with privileged information, such as a math solution, in thinking models. It reports degradation on long reasoning traces across five Qwen3 and OLMo models evaluated on AIME.
It highlights a potential trade-off when using privileged information to improve reasoning models.
- Reinforcement learningPost on X
RubricEM decomposes research-agent policies with rubrics
RubricEM trains research agents to create rubrics before working, then use them to plan, search, review, and answer. A judge scores trajectory stages, and judged attempts become lessons for future tasks.
Stagewise rubric rewards and reusable lessons offer an approach to training agents beyond single final-score feedback.
EfficientRollout speeds up RL rollout generation for LLMs
The paper presents system-aware self-speculative decoding for RL rollouts, using a quantized copy of the model, a roofline-based activation rule, and adaptive draft lengths. It reports up to 19.6% faster rollout generation and 12.7% faster training steps.
It describes ways to reduce rollout latency without changing the model being trained.
Two-Phase Distillation for Multi-Task Agentic LLMs
The paper studies consolidating separately trained, task-specific RL experts into a multi-task model through distillation, instead of training one model on mixed tasks. The post describes using off-policy distillation for initialization and on-policy distillation for refinement.
The approach offers an alternative to mixed-task training and examines a limitation of off-policy distillation in multi-task settings.
DreamSmooth smooths sparse rewards in model-based RL
The ICLR paper proposes temporally smoothing rewards across neighboring steps, using kernels such as Gaussian or EMA, to make reward modeling easier. The post says it evaluates the method on RoboDesk and Shadow Hand.
Reward smoothing may help engineers train model-based RL agents when rewards are sparse in time.
- Reinforcement learningPost on X
AutoDecompiler Uses RL for Feedback-Driven Decompilation
The post describes AutoDecompiler as an RL-optimized model for multi-turn decompilation that uses feedback directly, unlike prior one-shot approaches.
It may interest engineers exploring reinforcement learning for iterative code generation and decompilation.
- Reinforcement learningPost on X
TMAX trains terminal agents with diverse Dockerized RL environments
The TMAX post describes 14.6k Dockerized RL environments and an outcome-only DPPO recipe for training terminal agents. It reports a 9B model reaching 27% on Terminal-Bench 2.0.
The recipe and environment design may help engineers train terminal agents for multi-turn shell tasks.
A Proposal for Training-Free Reasoning with Verified Examples
The post proposes keeping an LLM frozen and retrieving verified reasoning from an external pool instead of updating weights. It uses SymbCoT-style symbolization for checking logical structure; whether this can approximate RL gains remains open.
It outlines a possible alternative to weight updates, while highlighting verification and deployment as unresolved challenges.
- Reinforcement learningArticle
Infrastructure optimizations for RL at trillion-parameter scale
A deep dive into prime-rl 0.6.0 training trillion-parameter MoE models, covering FP8, wide expert parallelism, P/D disaggregation, router replay, and 3-D parallelism.
Engineers can review training and inference techniques for scaling RL workloads on large MoE models.
- Reinforcement learningArticle
Post claims OPD is more efficient than RL
The author claims OPD uses compute and samples more efficiently than RL, which they say finds high-reward reasoning traces. The linked PDF is titled “llm_distillation.pdf.”
The claim concerns the compute and sample efficiency of approaches to training language models.
- Reinforcement learningPost on X
A proposed test for reward-hacking mitigations
The post contrasts blocking suspicious tool calls and returning dummy information with penalizing a CoT monitor, which it says can lead to obfuscation. It asks whether a head-to-head test in the same environment has measured this difference.
The comparison could help engineers evaluate whether interventions stop reward hacking or merely make it harder to detect.
- Reinforcement learningPost on X
Why RL Scaling Laws Differ from Pretraining
The post contrasts pretraining and RL scaling laws, noting that RL compute includes sampling and policy updates and may be measured in FLOPs or GPU hours. It also distinguishes within-run and across-run extrapolation.
It highlights compute and evaluation choices that complicate comparisons and extrapolation in RL experiments.
- Reinforcement learningArticle
Overview of RL methods for reasoning LLMs
A blog post surveys RL methods including REINFORCE, PPO, RLHF, GRPO, RLOO, Dr. GRPO, DAPO, CISPO, MaxRL, DPPO, and ScaleRL.
Useful as a single starting point for engineers comparing RL methods used with language models.
- Reinforcement learningArticle
Running verl RLHF Training on AMD GPUs with ROCm 7.0
An AMD ROCm blog describes deploying verl for RLHF training on AMD GPUs, with ROCm optimization and Docker scripts. It reports throughput and convergence results.
Useful for engineers evaluating AMD GPU infrastructure and deployment options for verl-based RLHF.
CogRouter adapts reasoning depth at each agent step
CogRouter uses four hierarchical cognitive levels to adjust reasoning depth for LLM agents. Its training combines supervised fine-tuning and policy optimization; the post reports an 82.3% benchmark success rate for a 7B model.
Step-level reasoning control may help engineers trade off agent performance and token use.
Self-distillation for continual learning without reward functions
The post describes a method that uses a model conditioned on a demonstration as a teacher for the same model generating text without that demonstration. The student is trained to match the teacher’s token distributions on its generated text.
This approach may help engineers train models on new tasks while reducing catastrophic forgetting, without defining a reward function.
- Reinforcement learningArticle
Online RL for HPC Code Generation with Machine Benchmarks
The article describes using online reinforcement learning with real-machine benchmark rewards to improve LLMs’ HPC code generation. It notes that generated code’s runtime performance is not guaranteed.
Engineers can see an approach to training code-generation models using measured runtime performance.
- Reinforcement learningPost on X
MaxRL uses likelihood-based training for binary-reward RL
The post describes MaxRL, which uses additional rollouts to approximate maximum-likelihood training with non-differentiable sampling. It says the method addresses underweighting of hard prompts and reports up to 20× test-time efficiency versus GRPO.
The approach may matter to engineers evaluating reward optimization and compute scaling for RL systems.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor
