Skip to content
EN

AI agents

Tools, memory, evaluation and orchestration for agents that do real work.

104 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: AI agents

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. AI agentsPost on X

    DSH post reports a shell prompt mismatch causing command delays

    The author says DSH’s terminal-bash and tool-bash-persistent use different prompt strings, triggering a 3.5-second timeout. They report changing the prompt constant reduced a test command’s latency from 3,600 ms to 158 ms.

    The reported prompt mismatch may help engineers diagnose slow shell commands in DSH.

  2. AI agentsArticle

    Survey of Self-Evolving Coding Agents

    The survey reviews systems that learn from software-development interactions and use those lessons to modify their memories, workflows, tools, or model parameters. It covers 108 systems.

    It maps ways coding agents can adapt using lessons from prior development work.

  3. AI agentsArticle

    Benchmark compares seven self-hosted memory providers for Hermes Agent

    A community benchmark compares seven self-hosted memory providers for Hermes Agent. Each was tested with the same 71,060 conversation turns and 3,750 questions about facts that change over time.

    The comparison offers evidence for engineers choosing a memory provider for Hermes Agent.

  4. AI agentsArticle

    Argus uses a verification-gated runtime for long-horizon agents

    Argus is a general-purpose agentic reasoning runtime for long-horizon tasks with underdefined objectives and sparse feedback. The post says it makes pivots verification-gated and reports 78% on SWE-Bench Pro versus 59% for Direct Copilot.

    Engineers building long-running agents can examine how verification gates runtime state changes.

  5. AI agentsRepository

    Semantica builds graph-native infrastructure for AI agents

    Semantica is an open-source project for turning data into knowledge graphs and recording agent decisions for later tracing. The post describes it as self-hosted and MIT-licensed.

    Decision traces and provenance can help engineers inspect how an agent reached an outcome.

  6. AI agentsRepository

    Buzz publishes a memory specification

    Buzz’s NIP-AE document specifies its memory feature. The project says formal writeups are intended to make features implementable and interoperable.

    A formal memory spec can help engineers understand and interoperate with Buzz’s agent memory design.

  7. AI agentsPost on X

    Survey of Self-Improvement in Agentic Systems

    A survey of 239 papers on how AI agents self-improve by updating the model or the scaffold, including prompts, memory, and tools.

    It maps approaches to improving agents across both model updates and the systems built around them.

  8. HarnessBank for self-evolving agent harnesses

    HarnessBank is a method for evolving agent harnesses, including prompts, tools, and control loops. It uses semantic gene-bank search with gated verification to address greedy candidate selection and noisy self-generated feedback.

    It describes an approach to improving agent performance by changing the harness around the model.

  9. MemoHarness adapts agent harnesses using execution history

    MemoHarness edits six harness surfaces and retrieves lessons from similar past cases to adapt to new ones. On a shell-agent benchmark, it scores 0.806 versus 0.722 for the strongest fixed-harness baseline.

    It explores a way to improve agent control layers from their own executions without test-time labels or extra search.

  10. A Survey of Self-Improvement in Modern Agentic Systems

    The survey frames modern agents as foundation models coupled with operational scaffolds such as prompts, memory, and tools. It presents self-improvement as adaptation that turns experience into accumulated capability gains.

    The system-level framing helps engineers reason about where agent updates can occur and how to make adaptation controllable.

  11. AI agentsPost on X

    Post Reports an Eight-Day Autoresearch Experiment

    The post claims that an autoresearch agent improved itself over eight days and beat a hand-tuned harness on held-out benchmarks.

    The claim may interest engineers evaluating recursive self-improvement, though the post provides no experimental details here.

  12. The Harness Effect on Enterprise Agent Costs

    The paper argues that agent orchestration—the harness—shapes token use and enterprise agent costs. The post reports lower cost and latency with quality at parity across 22 tasks and six models when orchestration changed.

    Engineers can assess orchestration choices as a lever for agent cost and latency, not just model selection.

  13. AI agentsPost on X

    ThetaEvolve framework for test-time learning on open problems

    ThetaEvolve is an open-source framework that extends AlphaEvolve with in-context learning and reinforcement learning at test time. It uses a single LLM, a program database, batch sampling, lazy penalties, and optional reward shaping.

    Its design offers engineers concrete techniques for scaling agent exploration and learning on open optimization problems.

  14. AI agentsArticle

    OpenResearch uses agents to reproduce paper results

    OpenResearch is described as using agents to reproduce research-paper results and provision GPUs in the cloud. The author says it reproduced their paper end to end.

    Engineers can assess an agent workflow that combines research reproduction with cloud GPU provisioning.

  15. The Verification Challenge for Coding Agent Rewards

    The paper examines verification for coding agents, arguing that reliably checking generated solutions has become harder than generating them. It analyzes four reward-verification designs and finds no single solution.

    It highlights how reward signals can diverge from human intent and why agent evaluations need robust verification.

  16. AREAL2.0 proposes a system for self-evolving agents

    The paper presents AREAL2.0, an architecture for continuous agent learning from live workloads. It includes a protocol for step-by-step learning signals, a proxy for turning tasks into training data, and an automatic update trigger.

    Engineers building agent systems can examine how the architecture connects deployment workloads to policy updates.

  17. Proactive memory for long-horizon agents

    The paper introduces a separate memory agent that maintains structured memories and decides when to surface reminders to an action agent. It evaluates the approach on Terminal-Bench 2.0 and tau-squared-Bench.

    Engineers can explore an approach to keeping task state and prior decisions influential as agent trajectories grow.

  18. Study measures how orchestration affects agent costs and quality

    The study evaluates 22 tasks across six foundation models while changing only the orchestration layer. It reports lower cost, token use, and runtime, with completion quality at parity.

    It suggests orchestration choices can materially affect agent efficiency across different models.

  19. AI agentsArticle

    BrowserBC distills browsing traces into reusable agent skills

    BrowserBC turns human browsing traces into natural-language skill cards that agents can retrieve and compose to navigate unfamiliar sites. The post reports gains over baselines in task success, recovery, and cross-site generalization.

    It offers an approach for reusing human browsing behavior to help agents handle unfamiliar websites.

  20. AI agentsArticle

    Harness engineering as a path to AI self-improvement

    Lilian Weng’s article surveys recursive self-improvement and argues that designing and optimizing the model’s harness is a practical starting point. It also discusses challenges including evaluation, diversity collapse, and reward hacking.

    It outlines how harness design could support self-improvement loops and highlights their evaluation and safety challenges.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor