Skip to content
EN

AI agents

Tools, memory, evaluation and orchestration for agents that do real work.

104 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: AI agents

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. AI agentsPost on X

    Microsoft releases 14B FrogMini coding agent

    The post says Microsoft released FrogMini on Hugging Face, a 14B coding agent scoring 45.3% on SWE-Bench Verified. It describes a training approach in which agents unintentionally break tests while adding features.

    The reported benchmark result and training approach may be relevant when evaluating coding agents.

  2. MemRL learns agent memory utility through runtime reinforcement learning

    MemRL is a non-parametric approach that applies reinforcement learning to episodic memory, aiming to help agents learn from experience without changing LLM weights.

    It explores ranking relevant memories by learned utility rather than semantic similarity alone.

  3. AI agentsRepository

    Use a stop hook to keep Claude running

    The post points to Anthropic’s Ralph Wiggum plugin and describes using a stop hook to prompt Claude to continue when it stops.

    Stop hooks can help engineers orchestrate longer-running agent work.

  4. Agent-R1: Modular RL for Multi-Turn LLM Agents

    Agent-R1 is a framework for training LLM agents with end-to-end reinforcement learning across multi-turn interactions. The paper focuses on agents that use tools and interact with environments over multiple rounds.

    It addresses how to train agents for sequential tool use and interaction with environments.

  5. AI agentsPost on X

    Self-play RL for software bug injection and repair

    The post introduces Self-play SWE-RL (SSR), a method for training one LLM agent to alternate between injecting and repairing bugs in real-world repositories, without human-labeled issues or tests.

    The approach may interest engineers exploring ways to train coding agents without labeled issue and test data.

  6. PasoDoble uses dual-play to train LLM reasoning

    The paper presents dual-play, an adversarial learning framework that assigns two models specialized roles to reduce reliance on external supervision. One model creates questions and the other solves them.

    Engineers can explore an approach to training LLM reasoning through competition between specialized models.

  7. ReCode unifies planning and action for LLM agents

    ReCode is a paradigm that uses recursive code generation to unify planning and action, letting agents adjust decision granularity. Its paper argues that rigidly separating high-level planning from low-level action limits LLM agents.

    The approach offers a way to build agents that shift between strategic planning and fine-grained actions.

  8. A Predictive Model for Scaling Agent Systems

    The paper introduces quantitative scaling principles that model how agent-system performance varies with coordination, model capability, and measurable system and task factors. It evaluates 260 configurations across six agentic benchmarks and five architectures.

    Engineers can use the work to reason about how coordination and model capability affect agent-system performance.

  9. AI agentsPost on X

    Engineer critiques an AGI paper’s definition and benchmarks

    The post discusses a paper that defines AGI as replicating a model of human cognition and tracks metrics over time. The author praises its discussion of capability variation and criticizes its benchmarks and narrow view of AI.

    It highlights design choices in AGI definitions and evaluations that engineers may want to scrutinize.

  10. A study of data, algorithms, and reasoning modes for agentic RL

    The paper investigates agentic reinforcement learning across data, algorithm design, and reasoning modes. It reports findings on trajectory data and tool use, and links to an open-source repository.

    The study may help engineers make design choices when training LLM agents for reasoning and tool use.

  11. AI agentsPost on X

    TUMIX uses diverse agents to refine answers at test time

    The post describes TUMIX, a system where agents using different tools solve problems independently, share answers, and refine them over several rounds. It claims diverse agents outperform copies of one model and improve Gemini-2.5 results.

    The approach highlights agent diversity and coordination as ways to improve reasoning without retraining the base model.

  12. AI agentsPost on X

    Early experience for agent learning

    The post describes early experience: using future states from an agent’s own interactions as supervision without reward signals. It examines implicit world modeling and self-reflection as two strategies.

    The approach targets learning challenges in environments with unverifiable rewards or costly long-horizon rollouts.

  13. AI agentsPost on X

    ToolUniverse connects LLMs to research tools

    ToolUniverse proposes a shared interface for discovering and calling research tools, with support for local tools and remote tools using MCP. The post describes a library of 600+ tools and a drug-discovery demonstration.

    A shared discovery and execution interface could reduce the need for one-off integrations in research workflows.

  14. AI agentsArticle

    Stanford course covers self-improving AI agents

    Stanford’s CS329A course, “Self-Improving AI Agents,” includes AB-MCTS, The AI Scientist, and the Darwin Gödel Machine, according to the post.

    The course brings together work on self-improving agents that engineers can explore.

  15. AI agentsArticle

    Meta releases 32B open-weights Code World Model

    Meta describes Code World Model (CWM), a 32-billion-parameter open-weights LLM released for research on code generation with world models.

    Engineers can examine an open-weights model aimed at code-generation research with world models.

  16. AI agentsPost on X

    LIMI studies data curation for agent autonomy

    The post describes LIMI, which uses 78 curated demos and reports 73.5% on AgencyBench, outperforming models trained on 10,000 samples. It argues that strategic curation, rather than data scale, drives autonomy.

    The reported results raise a testable question about how training-data curation affects agent benchmark performance.

  17. AI agentsPost on X

    Theory-Grounded Prompts for Generalizing Social Agents

    The post describes an MIT study that uses behavioral theory to write agent prompts, tunes them on small human datasets, and validates them on related but distinct settings. It reports lower prediction error on new game variants than language-model and equilibrium baselines.

    Cross-setting validation offers a way to test whether social agents generalize beyond their tuning data.

  18. SFR-DeepResearch trains single agents for research tasks

    The paper presents RL-trained autonomous agents for deep research that use web search, browsing, and Python. The agents can summarize prior results when context is limited.

    It describes an approach to tool use, long-horizon memory, and end-to-end RL for research agents.

  19. AI agentsPost on X

    NVIDIA’s Universal Deep Research runs research strategies as code

    The system compiles natural-language research strategies into a generator function that calls search and an interchangeable LLM. It stores intermediate facts in named variables and runs strategies in a sandbox.

    Engineers can inspect and customize research control flow, model use, and source rules rather than relying on a fixed workflow.

  20. AI agentsPost on X

    RLVMR Adds Process-Level Rewards to Agent Training

    Tencent researchers propose RLVMR, which combines outcome rewards with supervision for steps such as planning, exploration, and reflection. A 7B agent reportedly scored 83.6% on the toughest unseen ALFWorld and ScienceWorld tasks.

    Process-level rewards may help engineers train agents to reduce redundant actions and recover from errors.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor