LLM inference
Serving, quantization, batching, speculative decoding and the cost of every token.
85 links, newest first.
- LLM inferenceRepository
Community LLM serving recipes for RTX 3090, 4090, and 5090
The GitHub repository provides model-agnostic community recipes for serving LLMs on RTX 3090, 4090, and 5090 GPUs, with support for vLLM, llama.cpp, and ik_llama.
Engineers can compare serving setups across GPU generations and inference engines.
- LLM inferenceArticle
A guide to 13 open-source foundation model deployment tools
Turing Post’s guide covers 13 open-source tools for deploying, serving, and running foundation models, from local LLMs to high-throughput production inference.
It helps engineers compare deployment tools across local and production inference use cases.
- LLM inferenceRepository
llama.cpp Adds Multi-Token Prediction Support
A llama.cpp pull request adds support for MTP heads, which the post says can predict multiple tokens per pass. The author reports Qwen3.6-27B reached 65 tok/s, up from 38 tok/s, on an RTX 3090 with MTP enabled.
Engineers serving supported models can evaluate MTP as a way to increase inference throughput.
- LLM inferencePost on X
MTP inference benchmarks on three GTX 1080 Ti GPUs
The post reports llama.cpp MTP throughput benchmarks for Qwen 3.6 models on three GTX 1080 Ti GPUs, with Q4_0 K/V cache and draft-MTP flags. It lists context sizes and token rates.
The reported setup and flags may help engineers compare inference performance on older GPUs.
- LLM inferencePaper
Fluxion: Hybrid Sparse Attention for Long-Context Inference
The paper presents Fluxion, a hybrid sparse-attention system for long-context inference that uses CPU-GPU parallelism.
Engineers working on long-context inference may find its CPU-GPU parallelism approach relevant.
- LLM inferencePost on X
RTX 3090 LLM efficiency tests at different power limits
The author reports tests of Qwen3.6 27B on one RTX 3090 at concurrencies of 1, 4, 8, and 16. They say 225W had the best efficiency, while 250W provided higher throughput.
Useful measurements for engineers tuning power limits and concurrency on local LLM inference.
- LLM inferencePost on X
Luce PFlash Reports 2.89× TTFT Speedup Over Ollama at 64K
The post reports a benchmark in which Luce PFlash achieved 2.89× faster time to first token than Ollama at a 64K context.
The result may be relevant when comparing inference serving performance at long context lengths.
- LLM inferenceArticle
How KV cache avoids recomputing processed tokens
The article explains that an LLM's KV cache stores the Key and Value of tokens it has already processed, so they do not need to be computed again for each new token.
Understanding KV cache helps engineers reason about memory use and repeated computation during LLM inference.
- LLM inferenceRepository
Lucebox: speculative inference server for consumer GPUs
The post claims Qwen3.6-27B reaches 120–200 tokens/s on one RTX 3090 and links to Lucebox, a speculative inference server for heterogeneous hardware and consumer GPUs.
The repository may help engineers explore speculative inference on consumer GPU hardware.
- LLM inferencePost on X
RTX 3090 LLM tests compare throughput and power efficiency
An engineer tested 8 local LLMs on a single RTX 3090 at power limits from 100W to 450W. Average throughput was 90.4 tok/s at 225W and 107.1 tok/s at 450W, while efficiency fell from 0.4167 to 0.2731 tok/s/W.
The results show the throughput and power-efficiency tradeoff when choosing a GPU power limit for local inference.
- LLM inferenceRepository
dflash-mlx v0.1.5 adds runtime and serving features
The release adds a unified CLI for serving, generation, benchmarking, diagnostics, profiling, and model listing. It also introduces RAM and SSD snapshot caching, prefix matching, and experimental long-context support.
Engineers serving models on Apple Silicon can evaluate its runtime, caching, and diagnostics features.
- LLM inferenceRepository
Luce PFlash adds prefix caching and cold-start tuning
Luce PFlash adds prefix caching, cold-start tuning, and a CUDA VMM fix. The author reports about 10× faster warm performance and 2.5× faster cold performance, with block sparse attention autotune, for Qwen3.6 27B.
The linked repository is an LLM speculative inference server for heterogeneous hardware and consumer GPUs.
- LLM inferenceRepository
Lucebox explores speculative prefill for faster LLM inference
Lucebox is an LLM speculative inference server for heterogeneous hardware and consumer GPUs. The post claims speculative prefill speeds up Qwen3.6 27B time to first token by up to 10×.
Engineers can inspect a server implementation focused on speculative inference across different hardware.
- LLM inferencePost on X
Post claims DeepSeek-V4-Flash could fit on two RTX Pro 6000s
The author says DeepSeek-V4-Flash’s experts are already 4-bit and claims the model would fit on two RTX Pro 6000s. They say vLLM or SGLang still needs SM120 support for it.
It highlights how quantization and serving-stack support may affect the hardware needed for inference.
- LLM inferencePost on X
DFlash and DDTree run Qwen3.6-27B on an RTX 3090
The post reports 73 tok/s for Qwen3.6-27B on one RTX 3090 using a DFlash and DDTree speculative decoding stack. It says the stack loads the model because its architecture string and layer and head dimensions match Qwen3.5, but notes lower throughput than on Qwen3.5.
It highlights how architecture compatibility can enable speculative decoding on consumer GPUs before dedicated upstream support arrives.
- LLM inferenceArticle
Inference Engineering covers production AI inference systems
Philip Kiely's book covers the hardware, software, techniques, and infrastructure required to run AI models in production. The post says it focuses mostly on LLMs, with some coverage of diffusion.
A broad overview can help engineers build foundational knowledge of production inference systems.
- LLM inferenceRepository
Training Neural Networks on Apple’s Neural Engine
The ANE repository describes training neural networks using reverse-engineered private APIs for Apple’s Neural Engine.
Engineers exploring neural-network training and inference on Apple hardware can examine the implementation.
- LLM inferenceArticle
Fast LLM Inference From Scratch
Andrew Chan’s article is titled “Fast LLM Inference From Scratch.”
It may offer engineers a practical entry point to implementing LLM inference.
- LLM inferenceArticle
OpenAnonymity Proposes Unlinkable Inference for AI
The post describes a layer intended to prevent AI providers from linking inference calls to individual users. It says the approach is built into open-source infrastructure and a chat app.
Engineers evaluating hosted LLMs can assess an approach to separating user identity from inference requests.
- LLM inferencePost on X
Technical tutorial on dLLM inference and training
A roughly 22-minute tutorial on dLLMs covers self-distillation to reduce diffusion steps, curriculum learning, an LLM verifier, and KV cache.
It outlines training and inference techniques relevant to engineers exploring dLLMs.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor






