LLM inference
Serving, quantization, batching, speculative decoding and the cost of every token.
85 links, newest first.
- LLM inferenceRepository
ntransformer Runs Llama 70B on an RTX 3090
ntransformer is a C++/CUDA LLM inference engine whose GitHub page says it can run Llama 70B on an RTX 3090. The post describes using NVMe-to-GPU transfer while bypassing the CPU.
Engineers can examine an inference engine targeting large-model execution on a single consumer GPU.
- LLM inferenceRepository
llmfit identifies models that fit local hardware
llmfit is a GitHub tool for finding models and providers that run on your hardware. The post says it probes hardware, selects quantization for available RAM, handles MoE expert offloading, and estimates tokens per second.
It can help engineers assess local model options and expected inference speed before downloading weights.
- LLM inferencePost on X
hf-mem estimates model memory requirements
A post describes hf-mem as a tool for estimating memory requirements for new models and gives an example command using Qwen3.5-397B-A17B with experimental mode and an FP8 KV cache.
It offers a command-line approach to estimating memory needs as model architectures change.
- LLM inferenceArticle
Latency Optimization for Qwen3 on AMD MI300X
An LMSYS blog post about latency optimization for Qwen3 and Qwen3-VL on AMD MI300X GPUs.
It may offer relevant inference optimization techniques for engineers serving Qwen models on AMD hardware.
- LLM inferenceArticle
Unsloth Guide to Tool Calling with Local LLMs
Unsloth’s guide covers function calling with open models, with examples for story writing, Python execution, terminal calls and math.
Useful for engineers integrating tool calling into applications that run local open models.
- LLM inferencePost on X
LMCache reuses KV states across storage tiers
LMCache is an open-source extension for LLM serving that manages KV cache across GPUs, CPUs, and local disks. It can reuse repeated text fragments, not only prefixes.
Reusing KV states can reduce prefill work and GPU memory use in LLM serving.
- LLM inferenceRepository
Tencent Releases HPC-Ops LLM Inference Operator Library
HPC-Ops is an open-source LLM inference operator library with FusedMoE, GroupGEMM, attention, and multi-node communication support. Tencent reports production throughput gains and kernel speedups against named alternatives.
Engineers can inspect its operators and benchmarks when optimizing inference on GPUs.
- LLM inferencePost on X
CLI estimates model-loading VRAM requirements
The `do-i-have-the-vram` tool estimates how much VRAM is needed to load a model without loading it. It can be installed with `pip install do-i-have-the-vram`.
It can help engineers check whether a model fits on available GPU memory before attempting to load it.
- LLM inferencePaper
LoPA uses lookahead token ordering to speed up dLLM inference
LoPA is a training-free, plug-and-play decoding algorithm for diffusion LLMs. It uses lookahead to identify token filling orders, addressing the limited parallelism of confidence-driven decoding.
Token filling order can affect parallelism, making LoPA relevant to engineers optimizing diffusion LLM inference.
- LLM inferenceModel
GLM-4.7 185B W4A16 weights on Hugging Face
The linked Hugging Face repository is for GLM-4.7 185B W4A16 weights. The post says the weights are 92 GB and that the author ran 100 million tokens locally.
The quantized weights may be relevant to engineers evaluating local inference for a large model.
- LLM inferenceArticle
Transformers optimizations used by OpenAI gpt-oss
A Hugging Face blog post lists techniques from OpenAI’s gpt-oss model for use with Transformers, including MXFP4 quantization, tensor and expert parallelism, and dynamic sliding windows.
These techniques are relevant to engineers optimizing LLM inference with Transformers.
- LLM inferencePost on X
DEER Uses Diffusion Drafts for LLM Inference
DEER drafts tokens with diffusion models and verifies them with autoregressive models. The post claims up to 5.54× faster inference with lossless acceleration.
The approach may be relevant to engineers evaluating speculative decoding alternatives for inference acceleration.
- LLM inferencePost on X
SWAA for Efficient Long-Context LLM Inference
The post describes Sliding Window Attention Adaptation (SWAA), a set of training-free recipes for adapting full-attention LLMs to linear scaling while recovering performance.
Engineers evaluating long-context inference can consider a training-free approach to reducing attention costs.
- LLM inferencePost on X
Perplexity Builds MoE Inference Kernels for AWS EFA
Perplexity describes expert-parallel kernels for serving large MoE models across AWS GPUs using EFA. GPU dispatch and combine pack tokens into RDMA writes, while a host proxy thread coordinates transfers alongside grouped GEMM compute.
The design offers an approach to multi-node MoE serving when EFA lacks GPUDirect Async.
- LLM inferencePaper
Router-R1 Uses Reinforcement Learning to Coordinate Multiple LLMs
Router-R1 frames multi-LLM routing as a sequential decision process, alternating between reasoning and invoking models. Its reward combines format, outcome, and cost terms; the post reports results across seven QA benchmarks.
Engineers can examine an RL-based approach to balancing multi-model performance and cost.
- LLM inferencePaper
InfLLM-V2 Switches Between Dense and Sparse Attention
InfLLM-V2 is a dense-sparse switchable attention system for adapting from short to long sequences. Its paper discusses long-sequence processing bottlenecks and limitations of existing trainable sparse attention methods.
Engineers can assess an attention approach designed to support long contexts while addressing standard Transformer bottlenecks.
- LLM inferencePost on X
PHLoRA extracts LoRA adapters from fine-tuned checkpoints
PHLoRA uses the base and fine-tuned checkpoints to extract a low-rank adapter via singular value decomposition, without training data or gradients. The post reports accuracy within about 1% of full-rank models at ranks 32 or 64.
Engineers can reduce adapter loading and serving costs by converting existing full-rank fine-tunes.
- LLM inferenceArticle
Sipeed Maix4-HAT brings an AX650N NPU to Raspberry Pi
The Raspberry Pi 5 AI HAT uses an AXera AX650N NPU, rated for up to 72 TOPS at INT4 or 18 TOPS at INT8. It includes 8 GB LPDDR4x RAM and accelerates Transformer-based models for edge applications.
It offers engineers a compact platform to evaluate quantized Transformer inference at the edge.
- LLM inferenceRepository
Mirage compiles LLMs into persistent megakernels
Mirage Persistent Kernel is a compiler that transforms LLMs into optimized megakernels. The post claims this reduces latency by 1.2–6.7×.
Engineers working on LLM inference can evaluate a compiler approach to reducing latency through megakernels.
- LLM inferencePost on X
AirLLM runs large models with layer-wise inference
AirLLM uses layer-wise inference, loading one Transformer layer at a time from disk so the full model need not stay in GPU memory. The post also describes block-wise quantization and compression support.
Engineers evaluating inference on memory-constrained GPUs may find its layer-loading approach relevant.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor






