Skip to content
EN

LLM inference

Serving, quantization, batching, speculative decoding and the cost of every token.

85 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: LLM inference

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. LLM inferencePost on X

    AI21 releases Jamba 1.5 Mini and Large

    AI21 Labs announced Jamba 1.5 Mini and Large, MoE models combining Mamba and Joint Attention with a 256K context window. The post mentions ExpertInt8 quantization, JSON mode, tool use, and Transformers integration.

    The models’ architecture and quantization are relevant to engineers evaluating LLM inference options.

  2. LLM inferencePost on X

    Llama 3.1 8B Inference on One RTX 3090

    The post says a Backprop blog benchmark of Llama 3.1 8B in fp16 on one RTX 3090 found reasonable tokens per second at 100+ concurrent requests.

    It offers a reference point for evaluating single-GPU inference capacity under concurrent load.

  3. LLM inferenceRepository

    LLM Compressor brings model compression tools to vLLM

    Neural Magic released LLM Compressor, a library for applying compression algorithms including GPTQ, SmoothQuant, and SparseGPT. The post says compressed models aim to reduce inference latency while maintaining accuracy.

    Engineers using vLLM can explore one library for several model-compression methods.

  4. LLM inferencePost on X

    SGLang and vLLM throughput comparison on A6000 GPUs

    An engineer reports replacing vLLM with SGLang in a local data-generation pipeline. On four A6000 GPUs running 4-bit Mistral Large 2, they report about 1,100 tokens/s with SGLang versus under 750 with vLLM; performance is similar at smaller batch sizes.

    The comparison highlights how inference throughput can vary by serving engine and batch size on the same hardware and model.

  5. PyTorch introduces torchchat for local LLM inference

    Torchchat is a PyTorch library for running Llama 3, 3.1, and other large language models on laptops, desktops, and mobile devices. The post describes it as designed for seamless, performant local inference.

    Engineers exploring local LLM deployment across device types can evaluate the library.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor