Skip to content
EN

Attention and model architecture

Transformers, attention variants, state-space models and the ideas behind new architectures.

156 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Attention and model architecture

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. HySparse2 uses two-level KV sharing for sparse attention

    HySparse2 is a hybrid sparse attention method with two-level KV sharing. It targets efficient prefill and compact KV-cache storage for long-horizon, multi-turn agents.

    Engineers working on long-context models can assess an attention design aimed at reducing prefill and KV-cache demands.

  2. Memory Attention Reuses Keys in the Value Expression

    A post describes preliminary ablation results suggesting that reusing keys in V=K+M is an effective architectural choice. The author says it still needs validation at larger scales.

    The result may inform attention architecture experiments, but the post notes that larger-scale validation is still needed.

  3. Efficient Autoregressive Inference for Transformer Probabilistic Models

    An ICLR 2026 paper on efficient autoregressive inference for transformer probabilistic models, which the post says addresses a decoding bottleneck.

    Relevant to engineers working on faster autoregressive decoding for probabilistic transformer models.

  4. PolaFormer++ Adds Polarity-Aware Features to Linear Attention

    PolaFormer++ introduces a theoretical criterion for feature-map spikiness and proposes a polarity-aware, channel-wise feature map. The authors evaluate linear attention across six visual task families.

    It offers a feature-map design and theoretical framework for engineers exploring linear attention in vision models.

  5. TWT Compresses Similar Vision Transformer Layers

    Transformer-Within-Transformer replaces contiguous groups of similar Vision Transformer layers with single attention-based surrogate layers. On DINOv2, the post reports roughly half the compute with a tiny accuracy drop.

    It describes a way to reduce Vision Transformer depth and compute while preserving most of the reported accuracy.

  6. Reducing Redundancy in Looped Transformers

    The paper identifies three types of computational redundancy across loops in looped Transformers. It reports 1.65× lower latency and 6× less memory by exploiting them.

    Engineers evaluating looped Transformers can assess techniques aimed at reducing their compute and memory costs.

  7. Native Sparse Attention for Efficient Long-Context Modeling

    The paper presents NSA, a natively trainable sparse attention mechanism that combines algorithmic innovations with hardware-aligned optimizations for efficient long-context modeling.

    It describes an approach to sparse attention designed to improve efficiency while maintaining model capabilities.

  8. Sparse Layers Are Critical to Scaling Looped Language Models

    The paper compares standard and Mixture-of-Experts transformers, with and without looping. It reports that Looped-MoE models scale better than the standard baseline, while dense looped models do not.

    The results help engineers assess sparse layers and looping when designing models with adaptive depth.

  9. Memory Attention Adds Token-Indexed Memory to Attention

    Memory Attention replaces the Transformer’s learned value projection with a sum of contextual keys and layer-specific, token-indexed memory vectors. The post says the memory can reside on CPU.

    Engineers exploring attention architectures can assess an approach that adds model capacity through token-indexed memory.

  10. Memory Attention replaces the value projection with token memory

    Memory Attention combines layer-specific, token-indexed memory with contextual keys to form attention values. The post says this makes value computation largely lookup-based, with potential CPU offloading and KV-cache reduction.

    Engineers evaluating attention alternatives can assess its memory, compute, and KV-cache trade-offs.

  11. RetNet combines parallel training with recurrent inference

    The Retentive Network paper proposes an architecture for language models with a retention mechanism and parallel, recurrent, and chunkwise recurrent computation paradigms. It derives a connection between recurrence and attention.

    Its alternative sequence-modeling mechanism and inference paradigms are relevant to engineers evaluating Transformer architectures.

  12. FlashAttention-3 targets faster attention on Hopper GPUs

    The paper presents techniques to speed up attention on Hopper GPUs, including exploiting asynchrony and low-precision computation. It addresses GPU memory traffic and hardware utilization.

    Engineers working on Transformer performance can assess attention optimizations designed for Hopper GPUs.

  13. How Mamba and Transformers Connect Through State-Space Duality

    The paper develops theoretical connections between state-space models such as Mamba and variants of attention, using decompositions of structured semiseparable matrices.

    The framework helps engineers understand the relationship between SSMs and attention architectures.

  14. Megalodon: Efficient Sequence Modeling with Unlimited Context

    The paper introduces Megalodon, a sequence-modeling architecture designed for efficient pretraining and inference with unlimited context length. It builds on Mega and targets the long-sequence limitations of Transformers.

    It offers engineers an architecture to evaluate for long-context workloads and efficient sequence modeling.

  15. Differential Transformer subtracts two attention maps

    The Differential Transformer subtracts two attention maps to reduce attention noise and focus on signal, according to the post.

    The attention variant may be relevant to engineers exploring alternatives to standard Transformer attention.

  16. Memory Attention Replaces Learned Value Projections

    Memory Attention (MA) replaces the Transformer’s learned value projection with a sum of contextual keys and layer-specific, token-indexed memory vectors. The paper examines MA across several attention variants.

    The design explores an alternative way to represent and retrieve information in attention layers.

  17. HySparse2 uses KV sharing for long-context inference

    HySparse2 introduces two levels of KV sharing: cross-decoder KV bridging and reuse between sparse and full-attention layers. The post reports lower prefill FLOPs, a smaller KV cache, and better retrieval scores than MiMo-V2.6's Hybrid SWA architecture at 1M tokens.

    Its KV-sharing design targets prefill cost, cache size, and retrieval accuracy in growing-context agentic workloads.

  18. iSDFT: Information-Proximal Self-Distillation for Continual Learning

    iSDFT is a method for continual learning in LLMs, described as information-proximal on-policy self-distillation. Its GitHub repository provides training code, checkpoints, and evaluation results.

    The implementation and evaluation results offer engineers resources for exploring continual learning in LLMs.

  19. Mixture-of-Depths Dynamically Allocates Transformer Compute

    The paper presents a method for transformer language models to allocate compute to specific sequence positions across layers. It caps how many tokens participate in self-attention and MLP computations at each layer.

    Engineers can explore an approach to varying per-token compute under a layer-level budget.

  20. Quiet-STaR trains language models to think before speaking

    Quiet-STaR explores teaching language models to generate internal reasoning before producing text, building on the idea that reasoning is implicit in much written language.

    It offers engineers an approach to training models to reason before generating their visible output.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor