Skip to content
EN

AI agents

Tools, memory, evaluation and orchestration for agents that do real work.

104 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: AI agents

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. A Rejection-Sampling Approach to On-Policy LLM Alignment

    The paper proposes transferring off-policy tokens into on-policy tokens for LLM alignment, using rejection sampling so accepted tokens can be treated as on-policy. It targets variance from compounded token-level importance-sampling ratios.

    It may help engineers reason about variance and data-policy mismatch in RL post-training.

  2. A Taxonomy of Agentic Recommender Systems

    This survey reviews LLM-based agents in recommender systems and proposes a taxonomy based on autonomy. It describes three paradigms: agent-assisted recommendation, agent-as-recommender, and agent-as-user-simulator.

    The taxonomy helps engineers distinguish design patterns for adding agents to recommender systems.

  3. ReContext replays relevant evidence for long-context reasoning

    ReContext is a training-free inference harness that uses model-internal relevance signals to create a query-conditioned evidence pool and replay it before generation. The paper reports improved evidence utilization across eight 128K datasets on three model backbones.

    It offers an approach to improve evidence use in long-context tasks without training or external memory.

  4. AI agentsPost on X

    AutoMem Trains Agents to Manage Their Own Memory

    The post describes AutoMem, a paper about agents writing, searching, and cleaning up notes and learning when memory is useful. It claims this improved a 32B open model’s performance on long games by 2x to 4x.

    Agent memory management can improve long-horizon task performance without increasing the context window.

  5. Autodata uses agent loops to generate training data

    Autodata is a method for agents to create training and evaluation data. Its Agentic Self-Instruct implementation iteratively tests candidate questions with weak and strong solvers, keeping questions that separate their performance.

    The approach offers a way to generate examples targeted at a model’s current capabilities.

  6. PACE estimates agent performance from a small task subset

    The post describes PACE, which uses a regression framework over a small set of non-agentic tasks to predict full agentic performance. It reports PACE-BENCH has under 4% mean absolute error and about 85% ranking accuracy at roughly 100× lower cost.

    A cheaper proxy for agent benchmarks could help teams evaluate model changes with less time and compute.

  7. HASTE organizes reusable skills for ML engineering agents

    HASTE is a hierarchical multi-agent system that organizes skills into global, domain, and competition-specific tiers. In an ablation with 159 skills across eight competitions, tiered loading achieved a 100% medal rate, versus 62.5% with flat loading.

    The results suggest that how agents scope and load accumulated skills can affect competition performance and token use.

  8. AI agentsPost on X

    Reward signals for coding agents can drift from human intent

    The post describes four reward signals for coding tasks—test suites, scored web-page checklists, real engineer-assistant conversations, and project-level agent grading—and discusses their failure modes and repairs.

    Engineers evaluating coding agents can use these examples to spot reward hacking and design stronger checks.

  9. RoPoLL Uses Geometric Median to Aggregate LLM Judge Panels

    The paper formalizes panels of LLM evaluators under a contamination model and proposes RoPoLL, which aggregates scores with the geometric median. It reports experiments across 13 judges and corruption rates up to 50%.

    Engineers evaluating LLMs can use this work to compare robust panel aggregation with averaging.

  10. AI agentsPost on X

    Continual Harness for ARC-AGI-3 Agents

    A post about “Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3” says the benchmark’s test-time learning demands push agents to build a world model of rules and mechanics that updates with new evidence.

    It highlights the role of updating an agent’s world model during test-time learning.

  11. Dockerless verifies coding-agent patches without execution

    Dockerless is an environment-free verifier that evaluates generated code patches without executing them. The paper describes program verifiers as tools for selecting SFT trajectories and providing RL rewards.

    It explores an alternative to setting up per-repository environments for execution-based verification.

  12. EvoRec: A Multi-Agent Framework for Evolving Recommender Systems

    EvoRec is a proposed multi-agent framework for automating recommender-system optimization. It targets agents that retain methodology across iterations and explore ideas beyond a predefined optimization space.

    It describes an approach to reducing manual iteration in recommender-system optimization.

  13. The Verification Horizon for Coding-Agent Rewards

    The paper studies coding-agent reward signals, including test pass rates, LLM judges, and execution traces. It reports that each eventually stops tracking correctness and becomes vulnerable to hacking as task horizons grow.

    Engineers designing long-horizon coding agents can use the findings to assess how long a reward signal remains a reliable proxy for correctness.

  14. Towards Automating Scientific Review with Paper Assistant

    The paper frames the scaling of scientific review as a systems challenge and proposes four levels of AI-human collaboration. It also discusses Google Paper Assistant as an early tool for scaling parts of the review process.

    Engineers building research agents can learn how automated verification and human oversight are framed for scientific review.

  15. Google’s Paper Assistant for Automating Scientific Review

    The paper proposes a taxonomy for AI-assisted scientific review, motivated by the difficulty of scaling traditional peer review to match the influx of AI-assisted science. The post describes the system as focused on detecting errors and reviewing manuscripts.

    Engineers building AI agents can examine how the system frames automated verification and review.

  16. Red Queen Gödel Machine Co-Evolves Agents and Evaluators

    The paper describes a self-improvement method that evolves an agent alongside its evaluator. It uses fixed evaluators within epochs and replaces them only when they score better on held-out ground truth.

    Co-evolving evaluators may reduce benchmark gaming while preserving evaluation guarantees within each epoch.

  17. DiscoBench Evaluates Clarification-Aware Deep Search Agents

    DiscoBench is a benchmark for search agents powered by LLMs. It tests whether agents can handle vague, underspecified, or factually incorrect requests during multi-step retrieval and reasoning.

    It gives engineers a way to evaluate how agents handle ambiguity in real-world search tasks.

  18. Red Queen Gödel Machine Co-Evolves Agents and Evaluators

    The paper proposes making evaluation part of recursive self-improvement by co-evolving agents and their evaluators. It addresses methods that assume a fixed verifier, benchmark, or labeled dataset.

    Engineers building self-improving agents can consider how fixed evaluation criteria may become inadequate as agents improve.

  19. BINEVAL uses binary questions for interpretable LLM evaluation

    BINEVAL decomposes evaluation criteria into atomic yes-or-no questions and aggregates their verdicts into interpretable, multidimensional scores. The post reports training-free results matching or beating UniEval and G-Eval on three datasets.

    Inspecting individual verdicts can help engineers diagnose evaluation scores and use them to guide prompt improvements.

  20. Survey maps self-improving agents and experience-driven evolution

    A survey and accompanying list cover how deployed agents turn interaction traces into durable capabilities, from self-evolution to meta-evolution.

    Useful for engineers evaluating the infrastructure and approaches behind agents that improve from experience.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor