a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)
-
Updated
Sep 6, 2026 - C++
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)
Benchmarking seven LLM optimization techniques across memory, latency, throughput, batching, and distributed-inference trade-offs.
A minimalist, educational deep-dive into the vLLM architecture.
To associate your repository with the continous-batching topic, visit your repo's landing page and select "manage topics."