Please confirm you are human

This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.

A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.

Hold with a pointer, or hold Space or Enter.

News

vLLM Recipes
recipes.vllm.ai > browse

Browse all recipes

1+ week, 2+ day ago   (14+ words) vLLM Filter 161 recipes by task, architecture, size, precision, and hardware....

vLLM
recipes.vllm.ai > Lightricks > LTX-2.5-Diffusers

Lightricks/LTX-2.5-Diffusers — 19B · DENSE

1+ week, 2+ day ago   (183+ words) 19B diffusion transformer for joint video and synchronized audio generation, served via vLLM-Omni LTX-2.5-Diffusers is a dense 19B diffusion transformer for video with synchronized 48 kHz stereo audio. vLLM-Omni supports T2V and first-frame I2V through four pipeline classes: Select the class with --model-class-name; no…...

vLLM
recipes.vllm.ai > Qwen > Qwen3.8-27B

Qwen/Qwen3.8-27B — 27B · DENSE · 256K ctx

1+ week, 4+ day ago   (222+ words) 27B-parameter dense hybrid-attention model with linear attention on 48 of 64 layers, a vision tower, a built-in MTP draft head, 262K native context window and extensible to 1M context Fits one Blackwell GPU in every precision: NVFP4 in 24.6 GiB, 6.6M KV tokens at 1M context Qwen3.8-27B is the…...

vLLM
recipes.vllm.ai > MiniMaxAI > MiniMax-H3

MiniMaxAI/MiniMax-H3 — 64B · DENSE

3+ week, 1+ day ago   (844+ words) Open-weight general-purpose multimodal generation model — jointly generates 24 FPS video with native stereo audio from text, image, video, and audio references, served via vLLM-Omni 8.7 s of 1248×768 video with synchronized stereo audio in ~87 s on 4×B300 MiniMax H3 is an open-weight, general-purpose multimodal generation…...

vLLM
recipes.vllm.ai > thinkingmachines > Inkling-Small

thinkingmachines/Inkling-Small — 276B / 12B active · MOE · 1024K ctx

3+ week, 4+ day ago   (173+ words) Natively multimodal 276B-parameter MoE from Thinking Machines Lab — 12B active parameters, text/image/audio in, text out, and up to 1M context. 276B total / 12B active Inkling architecture for lower-cost, lower-latency deployment TML Inkling-Small is a natively multimodal 276B-parameter Mixture-of-Experts model from Thinking Machines…...

vLLM
recipes.vllm.ai > moonshotai > Kimi-K3

moonshotai/Kimi-K3 — 2.8T / 16 experts/token + shared (of 896 routed) active · MOE · 1024K ctx

3+ week, 5+ day ago   (198+ words) vLLM Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model (16 of 896 experts active per token) built on the Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with a 1M-token context window and native vision. - Hardware: At least 8x GB300. Multi-node for real production traffic. - ROCm:…...

vLLM
recipes.vllm.ai > poolside > Laguna-S-2.1

poolside/Laguna-S-2.1 — 118B / 8B active · MOE · 256K ctx

1+ mon, 3+ day ago   (386+ words) Poolside's 118B total / 8B activated MoE coding model with mixed sliding-window + global attention, native interleaved reasoning, and 256K context — the larger sibling of Laguna XS-2.1, tuned for agentic coding. 118B/8B-A MoE for agentic coding with interleaved thinking and tool use Laguna S-2.1 is Poolside's…...

vLLM
recipes.vllm.ai > tencent > Hy3

tencent/Hy3 — 295B / 21B active · MOE · 256K ctx

1+ mon, 2+ week ago   (374+ words) Tencent Hy3 — scaled-up MoE language model (295B total / 21B active) with a 3.8B MTP layer for speculative decoding, 256K context, and hy_v3 tool/reasoning parsers Hy3 MoE — 295B/21B on 8×H200, 8×H20-3e(141GB), or 8×AMD MI300X/MI355X with MTP Hy3 is Tencent's latest open-source Mixture-of-Experts language model: 295B total parameters with 21B activated per token, plus a 3.8B MTP…...

vLLM
recipes.vllm.ai > poolside > Laguna-XS-2.1

poolside/Laguna-XS-2.1 — 33B / 3B active · MOE · 256K ctx

1+ mon, 3+ week ago   (115+ words) Poolside's 33B total / 3B activated MoE coding model with mixed sliding-window + global attention, native interleaved reasoning, and 256K context — designed for agentic coding. 33B/3B-A MoE for agentic coding with interleaved thinking and tool use Laguna XS-2.1 support is available in vLLM 0.21.0 and later (via…...

recipes.vllm.ai
recipes.vllm.ai > moonshotai > Kimi-K2.7-Code

moonshotai/Kimi-K2.7-Code — 1T / 32B active · MOE · 256K ctx

2+ mon, 1+ week ago   (243+ words) recipes.vllm.ai Kimi-K2.7-Code is a coding-focused agentic model built on top of Kimi-K2.6. It substantially improves real-world long-horizon coding performance — strengthening end-to-end task completion across complex software-engineering workflows while improving token efficiency, reducing thinking-token usage by ~30% versus Kimi-K2.6. It…...