RAG and retrieval
Chunking, embeddings, rerankers, evaluation and what holds up in production.
66 links, newest first.
- RAG and retrievalPaper
GraphRAG for query-focused summarization
The paper presents GraphRAG, an approach for answering questions over document collections. It addresses global questions about a corpus, which standard RAG retrieval does not handle well.
It describes an approach for corpus-level questions that may not be answered by retrieving individual passages.
- RAG and retrievalRepository
Tencent Open-Sources WeKnora, an LLM Knowledge Platform
WeKnora is an open-source framework for document understanding, semantic retrieval and autonomous reasoning. The post describes source-linked answers, file-generating skills, opt-in long-term memory and self-hosting.
Engineers can inspect a RAG framework that combines document retrieval with agent skills and memory.
- RAG and retrievalRepository
ripwire builds structural repository context for coding agents
ripwire is a zero-dependency C++23 CLI and MCP server that parses 21 languages with Tree-sitter and returns task-relevant symbols, relationships, and tests in a token-budgeted response without embeddings or a vector database.
It offers an alternative to repeated repository searches for agents that need relevant code, potential change impact, and tests.
- RAG and retrievalPost on X
Weixin Open-Sources WeMM-Embedding Multimodal Model
Weixin says its Vision team developed WeMM-Embedding to support search and recommendations across text, images, and video. The 9B model is deployed across several Weixin services and ranked first on MMEB-v2 and MMEB-v3, according to the announcement.
Engineers evaluating multimodal retrieval can inspect an embedding model used in production search and recommendations.
- RAG and retrievalRepository
Tencent Releases WeMM-Embedding Multimodal Models
Tencent’s WeChat Vision team released WeMM-Embedding, a family of multimodal embedding models for understanding and retrieval. The post says the models support text, image, and video search and recommendations.
Engineers can evaluate the open-source models for multimodal retrieval workloads.
- RAG and retrievalModel
Tencent’s WeMM-Embedding-9B maps multiple modalities to a shared space
Tencent’s WeMM-Embedding-9B maps text, images, videos, and visual documents into a shared space. The post says it achieves SOTA on MMEB-v2 and MMEB-v3.
A shared embedding space may be relevant when building retrieval systems across text and visual content.
- RAG and retrievalPaper
Capacity Allocation in Hierarchical Search Agents
The paper factorizes hierarchical search agents into delegation and execution roles and examines how model capacity should be distributed. The post says decomposition capacity is the key performance bottleneck.
Engineers building search agents can consider whether delegation and execution need models of different capacities.
- RAG and retrievalPaper
NapMem treats long-term memory as an agent action space
NapMem organizes user history into a linked pyramid of raw conversations, typed records, topic tracks, and user profiles. The agent learns to choose which memory granularity to inspect using memory-tool reinforcement learning.
The approach offers an alternative to supplying agents with only evidence preselected by a retriever.
- RAG and retrievalPaper
SearchEyes trains multimodal search agents in simulated worlds
SearchEyes uses a typed knowledge graph as the backbone of a simulated search world for training multimodal agents. The paper addresses disconnected training data, search environments, and reward signals in multi-hop reasoning.
The approach connects search-world structure and training signals for multimodal search agents.
- RAG and retrievalPaper
NapMem treats long-term user memory as an action space
The paper introduces NapMem, a framework that organizes user history into a linked, multi-granularity memory pyramid. It frames memory use as structured navigation rather than passive retrieval.
Engineers evaluating conversational memory systems can compare active navigation with pre-selected retrieval.
- RAG and retrievalPaper
Theoretical Capacity of MaxSim Retrieval Models
The paper studies the representation power of MaxSim and shows it can exactly replicate inner products between non-negative k-sparse vectors. The preview notes strong empirical performance for late-interaction models but limited prior theoretical understanding.
It gives engineers a theoretical result for assessing MaxSim in late-interaction retrieval.
- RAG and retrievalPaper
CMDR benchmarks cross-page multimodal document retrieval
The paper introduces CMDR and CMDR-Bench for retrieving relevant pages from multimodal documents, addressing queries that require context across multiple pages. The post also describes an embedding model for this task.
It highlights how page-level retrieval can miss context needed to answer queries spanning a document.
- RAG and retrievalPaper
Relevance-Based Embeddings for Candidate Retrieval
The paper describes representing queries and items with embeddings to retrieve candidates efficiently when the relevance function is expensive. The post says Yandex bases these representations on relevance to selected support items or queries.
It may be useful for engineers exploring efficient candidate retrieval with expensive relevance models.
- RAG and retrievalPaper
A Study of In-Context Retrieval at Million-Token Scale
The paper studies language models as in-context retrievers on million-token corpora and examines length generalization. It introduces BlockSearch, a 0.6B-parameter retriever that the post says generalizes up to 10 times beyond its training length.
Engineers can compare in-context retrieval with vector-based retrieval at corpus scales practical systems face.
- RAG and retrievalPaper
Studying In-Context Retrieval at Million-Token Scale
The paper studies whether language models can retrieve answers directly from in-context corpora, examining million-token corpora and length generalization. The post says it finds attention dilution and proposes length-aware fixes.
It examines the limits of using long in-context corpora as an alternative to vector-based retrieval.
- RAG and retrievalPaper
LLM-Based Hard Negative Sampling for Two-Tower Retrieval
The paper proposes a self-supervised hard negative sampling technique for two-tower recommendation models. It uses LLM-derived item clusters to generate harder negatives in real time.
Harder training negatives may help engineers address a limitation of standard negative sampling in large-scale retrieval.
- RAG and retrievalPaper
Trie-based execution plans for IR pipeline experiments
The paper presents a radix-trie execution plan for PyTerrier that reuses shared pipeline prefixes. The post reports experiment-time reductions of up to 26%.
Reusing shared pipeline work may make retrieval experiments faster to run.
- RAG and retrievalPaper
Diffusion-GR2 speeds up generative reasoning reranking
Diffusion-GR2 converts an autoregressive reasoning reranker into a block-diffusion model. The post reports near-AR ranking accuracy and 2.4–3.5× faster decoding.
It explores parallel decoding as a way to reduce inference cost in reasoning-based rerankers.
- RAG and retrievalPaper
STEB benchmarks style text embeddings across 96 datasets
STEB is an open-source benchmark for evaluating style embeddings across 96 datasets and 7 languages. The post says semantic embeddings underperform on stylistic tasks and no single model dominates.
It offers a standardized way to evaluate embeddings for style-focused retrieval and related tasks.
- RAG and retrievalPaper
KbSD uses self-distillation to calibrate agentic search
KbSD proposes a self-distillation framework for agentic search, where a hint-augmented teacher provides dense, token-level supervision. It targets decisions about using model memory, retrieved evidence, or abstaining.
Engineers building retrieval agents can examine an approach to calibrating when a model relies on search or its own knowledge.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor


