Skip to content
EN

AI agents

Tools, memory, evaluation and orchestration for agents that do real work.

104 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: AI agents

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. AI agentsArticle

    JAZ explores a minimalist recursive agent framework

    The post describes JAZ as using a single recursive code primitive for agent loops. It claims JAZ can outperform Letta and ACE at lower cost; the linked preview describes agents using tools and alternating through a workflow.

    The comparison may help engineers assess whether simpler agent loops can replace specialized memory and self-improvement frameworks.

  2. JAZ uses a minimal, code-driven agent harness

    JAZ is an agent framework with a single `invoke` primitive; the LLM writes code that can call it recursively. The post reports results on StuLife and AppWorld, including lower costs than the compared approaches.

    The framework explores implementing memory and self-improvement within the agent loop instead of as separate subsystems.

  3. Ablations of Coding Agent Harnesses

    The paper reports 176 ablation setups across SWE-Bench and Terminal-Bench. The post says rule-based context trimming with a short summary outperformed elaborate recovery, while explicit planning helped weaker models but not frontier models’ success rates.

    Its ablations can help engineers assess which harness components improve agent performance or reduce costs.

  4. DeepSeek describes DSec sandbox infrastructure for agent training

    The DSec paper describes elastic, stateful execution environments for large-scale agent training and evaluation, including varied isolation needs and bursty sandbox creation. The post also reports defenses against agents exploiting their training environments.

    Sandbox design affects the scale, efficiency, and security of agent training and evaluation.

  5. Agora uses Git as shared memory for research agents

    Agora stores agent contributions in an append-only Git DAG, with links to prior work and searchable views of results and verification status. In a nearly 12-day run, 13 LLM workers made 1,703 contributions to initialize a hybrid model without training data or gradient updates.

    The design offers a concrete way for parallel agents to share verified results and avoid repeating experiments.

  6. AI agentsRepository

    DeepSeek Harness adds browser and computer-use options

    DeepSeek Harness v0.1.6-alpha.1 adds a Web sidebar with persistent multi-tab terminals and MCP resource support. Experimental browser use supports Playwright MCP, Chrome DevTools MCP, and Stagehand; computer use supports Cua Driver MCP or a native driver.

    The release expands available interfaces for agents that need to use browsers, terminals, and local computers.

  7. AI agentsRepository

    Tencent Cloud’s Octop is a self-hosted, multi-user agent platform

    Octop is an open-source, self-hosted AI assistant platform described as supporting multiple users and agents. The post says it stores data locally in SQLite and integrates with chat platforms, ACP, browser automation, and terminal assistance.

    Engineers can evaluate how Octop combines local storage, team access, messaging integrations, and agent tools.

  8. Byte Models’ Scaling Trends Compared with Token Models

    The paper studies distilled byte and token models at scales up to 1 trillion training bytes. It introduces approximate and exact methods for converting token logits into byte logits and reports that byte models reach a higher ceiling as compute grows.

    The results may help engineers weigh tokenization choices and distillation methods when scaling language models.

  9. Scientific Agent Skills: 163 Research Procedures Across 16 Fields

    The paper presents an open library of 163 procedures for research agents across 16 areas, including genomics, cheminformatics, medical imaging, and study design. It focuses on procedural choices needed for defensible analyses.

    Engineers building research agents can use the library as a reference for domain-specific procedures and caveats.

  10. AI agentsPost on X

    Alibaba’s Zvec Team Releases Local Search Tool zg

    Alibaba’s Zvec team says it open-sourced zg, a local search tool for developers and AI agents. It combines semantic, BM25, hybrid, and rg search and works with popular agents.

    Engineers building agents can consider a local tool that brings several search modes together.

  11. AI agentsPost on X

    Reliability checks reduce errors in AI-generated research papers

    The post describes an arXiv paper reporting severe result hallucinations in papers produced by Agent Laboratory and Co-Scientist when reliability modules were removed. Checking manuscript claims against execution logs reduced the reported rate to 4%.

    It highlights the value of validating agent-generated research against execution evidence.

  12. WikiSkill Uses Persistent Knowledge to Evolve Agent Skills

    WikiSkill separates execution traces, a persistent knowledge wiki, and executable skills. It consolidates experience in the wiki, which informs later skill updates.

    The approach offers a way to reuse agent experience across skill iterations rather than relying on scattered optimization histories.

  13. JIT-Agent Generates Task-Adaptive Agent Harnesses

    JIT-Agent synthesizes agent harnesses on the fly for off-the-shelf agentic LLMs. Its four-module protocol covers memory, planning, action protocols, and tool orchestration; it can also repair and evolve harnesses.

    The paper explores automating harness design and adaptation, including components engineers often build and maintain manually.

  14. Recuris evolves agent memory around a frozen LLM

    Recuris describes two loops for long-horizon agents: working and experiential memory guide task execution, while a fixed meta-agent analyzes execution traces and proposes targeted memory repairs.

    It explores improving agent behavior through memory and execution feedback without changing model weights.

  15. AI agentsRepository

    Code Notebooks for Agentic Design Patterns

    Antonio Gulli's guide covers prompt chaining, routing, reflection, tool use, planning, and multi-agent systems, with code notebooks for hands-on exploration.

    Engineers can use the notebooks to explore common agent design patterns in code.

  16. Scroll manages long-horizon agent context as an executable environment

    Scroll pairs an append-only event log with a sandboxed, persistent Python kernel. Agents can search and transform session state, while an eviction index links compact landmarks to recoverable log entries.

    The design offers an alternative to committing long-running agent history to a fixed compressed memory representation.

  17. Active Inference for Agent Context Acquisition

    The paper models how interactive agents choose between acting on assumptions and spending resources on clarifying questions, retrieval, or other context-gathering actions. It frames these choices as active inference over a latent task state, with costs included in the objective.

    The framework gives engineers a way to reason about when an agent should gather more context instead of acting on uncertain assumptions.

  18. AI agentsPost on X

    Pandora’s Router Uses Cost-Aware Model Selection

    The paper frames model routing as a Pandora’s Box problem: it uses cheap, noisy scores, then pays for stronger estimates when their expected value exceeds their cost. On MATH, RAG, and EmbedLLM, it had the lowest or tied-lowest average combined routing regret and inspection cost across tested…

    Engineers building model routers can compare the approach’s tradeoff between estimation cost and routing regret.

  19. AI agentsPost on X

    AI R&D agents rarely revisit their training strategy

    A study of 1,338 post-training trajectories found agents rarely changed high-level training strategy: 74 of 3,557 adjacent experiments did. Journals, skill libraries, and evaluators improved benchmark scores but did not trigger strategy changes.

    Agent workflows may need explicit triggers to reconsider the overall strategy, not just refine its implementation.

  20. AI agentsPost on X

    AgentSysBench Studies Serving Bottlenecks in Agent Workloads

    The paper introduces AgentSysBench, covering 10 agentic applications, and reports that model inference is often not the main serving bottleneck. It evaluates task-aware serving, communication-aware placement, state offloading, and caching.

    The results suggest agent serving systems may need to optimize models, tools, memory, and communication together.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor