AI agents
Tools, memory, evaluation and orchestration for agents that do real work.
104 links, newest first.
- AI agentsPost on X
Microsoft releases 14B FrogMini coding agent
The post says Microsoft released FrogMini on Hugging Face, a 14B coding agent scoring 45.3% on SWE-Bench Verified. It describes a training approach in which agents unintentionally break tests while adding features.
The reported benchmark result and training approach may be relevant when evaluating coding agents.
- AI agentsPaper
MemRL learns agent memory utility through runtime reinforcement learning
MemRL is a non-parametric approach that applies reinforcement learning to episodic memory, aiming to help agents learn from experience without changing LLM weights.
It explores ranking relevant memories by learned utility rather than semantic similarity alone.
- AI agentsRepository
Use a stop hook to keep Claude running
The post points to Anthropic’s Ralph Wiggum plugin and describes using a stop hook to prompt Claude to continue when it stops.
Stop hooks can help engineers orchestrate longer-running agent work.
- AI agentsPaper
Agent-R1: Modular RL for Multi-Turn LLM Agents
Agent-R1 is a framework for training LLM agents with end-to-end reinforcement learning across multi-turn interactions. The paper focuses on agents that use tools and interact with environments over multiple rounds.
It addresses how to train agents for sequential tool use and interaction with environments.
- AI agentsPost on X
Self-play RL for software bug injection and repair
The post introduces Self-play SWE-RL (SSR), a method for training one LLM agent to alternate between injecting and repairing bugs in real-world repositories, without human-labeled issues or tests.
The approach may interest engineers exploring ways to train coding agents without labeled issue and test data.
- AI agentsPaper
PasoDoble uses dual-play to train LLM reasoning
The paper presents dual-play, an adversarial learning framework that assigns two models specialized roles to reduce reliance on external supervision. One model creates questions and the other solves them.
Engineers can explore an approach to training LLM reasoning through competition between specialized models.
- AI agentsPaper
ReCode unifies planning and action for LLM agents
ReCode is a paradigm that uses recursive code generation to unify planning and action, letting agents adjust decision granularity. Its paper argues that rigidly separating high-level planning from low-level action limits LLM agents.
The approach offers a way to build agents that shift between strategic planning and fine-grained actions.
- AI agentsPaper
A Predictive Model for Scaling Agent Systems
The paper introduces quantitative scaling principles that model how agent-system performance varies with coordination, model capability, and measurable system and task factors. It evaluates 260 configurations across six agentic benchmarks and five architectures.
Engineers can use the work to reason about how coordination and model capability affect agent-system performance.
- AI agentsPost on X
Engineer critiques an AGI paper’s definition and benchmarks
The post discusses a paper that defines AGI as replicating a model of human cognition and tracks metrics over time. The author praises its discussion of capability variation and criticizes its benchmarks and narrow view of AI.
It highlights design choices in AGI definitions and evaluations that engineers may want to scrutinize.
- AI agentsPaper
A study of data, algorithms, and reasoning modes for agentic RL
The paper investigates agentic reinforcement learning across data, algorithm design, and reasoning modes. It reports findings on trajectory data and tool use, and links to an open-source repository.
The study may help engineers make design choices when training LLM agents for reasoning and tool use.
- AI agentsPost on X
TUMIX uses diverse agents to refine answers at test time
The post describes TUMIX, a system where agents using different tools solve problems independently, share answers, and refine them over several rounds. It claims diverse agents outperform copies of one model and improve Gemini-2.5 results.
The approach highlights agent diversity and coordination as ways to improve reasoning without retraining the base model.
- AI agentsPost on X
Early experience for agent learning
The post describes early experience: using future states from an agent’s own interactions as supervision without reward signals. It examines implicit world modeling and self-reflection as two strategies.
The approach targets learning challenges in environments with unverifiable rewards or costly long-horizon rollouts.
- AI agentsPost on X
ToolUniverse connects LLMs to research tools
ToolUniverse proposes a shared interface for discovering and calling research tools, with support for local tools and remote tools using MCP. The post describes a library of 600+ tools and a drug-discovery demonstration.
A shared discovery and execution interface could reduce the need for one-off integrations in research workflows.
- AI agentsArticle
Stanford course covers self-improving AI agents
Stanford’s CS329A course, “Self-Improving AI Agents,” includes AB-MCTS, The AI Scientist, and the Darwin Gödel Machine, according to the post.
The course brings together work on self-improving agents that engineers can explore.
- AI agentsArticle
Meta releases 32B open-weights Code World Model
Meta describes Code World Model (CWM), a 32-billion-parameter open-weights LLM released for research on code generation with world models.
Engineers can examine an open-weights model aimed at code-generation research with world models.
- AI agentsPost on X
LIMI studies data curation for agent autonomy
The post describes LIMI, which uses 78 curated demos and reports 73.5% on AgencyBench, outperforming models trained on 10,000 samples. It argues that strategic curation, rather than data scale, drives autonomy.
The reported results raise a testable question about how training-data curation affects agent benchmark performance.
- AI agentsPost on X
Theory-Grounded Prompts for Generalizing Social Agents
The post describes an MIT study that uses behavioral theory to write agent prompts, tunes them on small human datasets, and validates them on related but distinct settings. It reports lower prediction error on new game variants than language-model and equilibrium baselines.
Cross-setting validation offers a way to test whether social agents generalize beyond their tuning data.
- AI agentsPaper
SFR-DeepResearch trains single agents for research tasks
The paper presents RL-trained autonomous agents for deep research that use web search, browsing, and Python. The agents can summarize prior results when context is limited.
It describes an approach to tool use, long-horizon memory, and end-to-end RL for research agents.
- AI agentsPost on X
NVIDIA’s Universal Deep Research runs research strategies as code
The system compiles natural-language research strategies into a generator function that calls search and an interchangeable LLM. It stores intermediate facts in named variables and runs strategies in a sandbox.
Engineers can inspect and customize research control flow, model use, and source rules rather than relying on a fixed workflow.
- AI agentsPost on X
RLVMR Adds Process-Level Rewards to Agent Training
Tencent researchers propose RLVMR, which combines outcome rewards with supervision for steps such as planning, exploration, and reflection. A 7B agent reportedly scored 83.6% on the toughest unseen ALFWorld and ScienceWorld tasks.
Process-level rewards may help engineers train agents to reduce redundant actions and recover from errors.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor