LLM inference
Serving, quantization, batching, speculative decoding and the cost of every token.
85 links, newest first.
- LLM inferencePost on X
AI21 releases Jamba 1.5 Mini and Large
AI21 Labs announced Jamba 1.5 Mini and Large, MoE models combining Mamba and Joint Attention with a 256K context window. The post mentions ExpertInt8 quantization, JSON mode, tool use, and Transformers integration.
The models’ architecture and quantization are relevant to engineers evaluating LLM inference options.
- LLM inferencePost on X
Llama 3.1 8B Inference on One RTX 3090
The post says a Backprop blog benchmark of Llama 3.1 8B in fp16 on one RTX 3090 found reasonable tokens per second at 100+ concurrent requests.
It offers a reference point for evaluating single-GPU inference capacity under concurrent load.
- LLM inferenceRepository
LLM Compressor brings model compression tools to vLLM
Neural Magic released LLM Compressor, a library for applying compression algorithms including GPTQ, SmoothQuant, and SparseGPT. The post says compressed models aim to reduce inference latency while maintaining accuracy.
Engineers using vLLM can explore one library for several model-compression methods.
- LLM inferencePost on X
SGLang and vLLM throughput comparison on A6000 GPUs
An engineer reports replacing vLLM with SGLang in a local data-generation pipeline. On four A6000 GPUs running 4-bit Mistral Large 2, they report about 1,100 tokens/s with SGLang versus under 750 with vLLM; performance is similar at smaller batch sizes.
The comparison highlights how inference throughput can vary by serving engine and batch size on the same hardware and model.
- LLM inferenceArticle
PyTorch introduces torchchat for local LLM inference
Torchchat is a PyTorch library for running Llama 3, 3.1, and other large language models on laptops, desktops, and mobile devices. The post describes it as designed for seamless, performant local inference.
Engineers exploring local LLM deployment across device types can evaluate the library.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor