Skip to content
EN

AI agents

Tools, memory, evaluation and orchestration for agents that do real work.

104 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: AI agents

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. NOVA Uses Verification-Aware Agents to Evolve Recommender Architectures

    NOVA is an agent harness for architecture evolution in industrial recommender systems. Its preview describes coordinated changes to model topology, feature configuration, and interaction modules under interface, resource, and serving constraints.

    It addresses how agents can validate domain-specific constraints beyond whether generated code runs.

  2. Survey examines evaluation of agent memory systems

    The paper surveys the evolution of LLM agent memory into systems for persistent storage, retrieval, updates, consolidation, and lifecycle governance. It notes that existing evaluations often rely on end-to-end task metrics and treat memory as a black box.

    Engineers can use its framing to assess what current agent memory benchmarks do—and do not—measure.

  3. AI agentsArticle

    Ornith-1.0 releases open models for agentic coding

    Ornith-1.0 is a family of open-source models ranging from 9B to 397B parameters. Its training uses reinforcement learning to optimize both coding solutions and the task-specific scaffolds that guide them.

    Engineers can evaluate the models and their self-scaffolding approach for agentic coding workflows.

  4. AI agentsPost on X

    Autodata uses agents to build training and evaluation data

    Autodata is a method for AI agents to act as data scientists and create training and evaluation data. It includes data creation, data analysis, and meta-optimization stages.

    The staged approach may help engineers structure agent workflows for generating and improving datasets.

  5. NatureBench tests coding agents on 90 research tasks

    NatureBench contains 90 tasks based on Nature-family papers and evaluates coding agents against published results without web search or access to the original method. The post reports that the best agent surpassed the published SOTA on 17.8% of tasks.

    The benchmark offers a way to evaluate agents’ ability to select methods and reproduce research results.

  6. AI agentsPost on X

    Automated code iteration improves scientific prediction methods

    A post describes a Nature study in which researchers repeatedly generated and machine-scored code for scientific prediction tasks across six fields. Of 55 combinations of existing methods, 24 outperformed both source methods.

    The work highlights how automated coding and evaluation loops could shorten experimentation on prediction software.

  7. AI agentsArticle

    AutoResearch framework automates an RL experiment pipeline

    The AutoResearch project describes an agent that planned GPU experiments and ran RL experiments on the DeepSeek 285B model. The post says it automated experiment design, coding, execution, debugging, and conclusion summaries.

    Engineers can examine how the framework structures autonomous experimentation across an end-to-end RL workflow.

  8. AI agentsArticle

    Deli AutoResearch automates RL experiment workflows

    Deli AutoResearch is an open-source agent framework. The author says it autonomously planned GPU experiments and automated experiment design, code writing, execution, debugging, and conclusion summaries for RL runs on DeepSeek 285B.

    The project offers an example of an agent handling multiple stages of an RL research workflow.

  9. AI agentsPost on X

    Self-Harness lets agents refine their own control harnesses

    The paper explores agents mining their failures and proposing small harness edits, retaining changes that pass regression tests. It reports improved held-out pass rates on Terminal-Bench-2.0 across MiniMax, Qwen, and GLM.

    Engineers can assess an approach to iteratively improving agent prompts, tools, retries, and verification without fine-tuning.

  10. AI agentsPost on X

    Recursive Multi-Agent Systems Refine Shared Latent States

    The post describes a paper in which agents recursively refine latent thoughts and exchange hidden states, decoding text only at the end. It says experiments found more accurate collaboration, faster execution, and lower token use.

    The approach explores whether latent-state communication can improve multi-agent collaboration while reducing token use.

  11. Agentic Harness Engineering for Coding Agents

    The paper introduces a framework for automatically evolving coding-agent harnesses using observable components, condensed trajectory evidence, and decisions tested against task outcomes. It reports Terminal-Bench 2 pass@1 rising from 69.7% to 77.0% over ten iterations.

    The approach makes harness changes easier to evaluate, attribute, and revert.

  12. PARE-Bench evaluates proactive agents in stateful app simulations

    PARE models applications as finite state machines to simulate active users. PARE-Bench includes 143 tasks across communication, productivity, scheduling, and lifestyle apps, testing goal inference, intervention timing, and multi-app orchestration.

    It offers engineers a benchmark for evaluating agent behavior across stateful, sequential user interactions.

  13. StructMem adds temporal structure to LLM agent memory

    The StructMem paper proposes a hierarchical memory system for long-term conversational agents. It aims to capture relationships between events for temporal reasoning and multi-hop question answering, balancing flat memory’s efficiency against the construction costs of graph memory.

    Engineers building long-horizon agents can assess an approach to preserving event relationships without relying on costly graph construction.

  14. TACO learns context compression rules for terminal agents

    TACO is a framework that discovers and refines context compression rules from terminal-agent interaction trajectories. The paper reports tests on TerminalBench, SWE-Bench Lite, and CompileBench.

    Engineers building long-horizon terminal agents can explore an approach to reducing noisy observations in context.

  15. AI agentsArticle

    Weekly AI research roundup includes several agent papers

    The roundup lists research on Claude Code’s agent-system design, long-horizon engineering and task scaling, and memory transfer in coding agents, alongside work on distillation and other topics.

    It offers leads on agent design, memory, and coordination for engineers tracking current research.

  16. AI agentsPost on X

    Protocol for auditable self-improving agents

    The post describes a paper proposing a protocol for agents to propose, assess, and commit improvements, with auditable lineage and rollback.

    The framework’s audit and rollback mechanisms are relevant to engineers designing self-improving agent systems.

  17. AutoSOTA automates iterative AI model improvement

    AutoSOTA is an end-to-end system that uses eight specialized agents to reproduce and improve models from AI papers. The post reports 105 new state-of-the-art models across LLMs, computer vision, and time series.

    It offers an example of multi-agent orchestration for research workflows involving code, experiments, and model refinement.

  18. ML-Master 2.0 uses layered memory for long-horizon research

    The post describes ML-Master 2.0’s Hierarchical Cognitive Caching: short-, medium-, and long-term memory for research across experiments and sessions. The team reports a 56.44% medal rate on MLE-Bench after a 24-hour run.

    The layered memory design offers a concrete approach to managing agent state across extended research tasks.

  19. AI agentsRepository

    Curated papers and resources on agentic reasoning

    A GitHub list of papers and resources based on the survey “Agentic Reasoning for Large Language Models.”

    Useful for engineers looking for research and resources on reasoning in AI agents.

  20. Controlled Self-Evolution for Algorithmic Code Optimization

    The paper proposes Controlled Self-Evolution, which uses diverse initial strategies, feedback-guided mutations, and cross-task experience to improve algorithmic code through iterative evolution. It evaluates the approach on EffiBench-X.

    Its approach to exploration, feedback, and cross-task memory offers concrete design ideas for coding agents.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor