Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.
-
Updated
Sep 29, 2026 - Markdown
Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.
Rust + cuTile research prototype for paged latent-cache LLM decode attention, validated on an RTX 4060.
A practical handbook for software engineers to learn AI, Large Language Models (LLMs), and Inference Engineering—from fundamentals to production systems.
CPU-only conversational voice assistant with VAD-based speech segmentation, streaming LLM response generation overlapped with concurrent TTS synthesis via producer-consumer queues. Batch STT (Whisper). Fully instrumented with per-stage latency benchmarking.
it's me as a repository
LLM inference benchmarking dashboard: Python FastAPI backend with async orchestration, WebSocket live TTFT/TBT/throughput comparison across configs (512/128 to 4096/1024 tokens), Grafana + Docker Compose stack, GitHub Actions CI; 21/21 pytest passing.
An interactive playground for learning inference engineering—explore LLM serving concepts, tune the stack, and graduate to production incidents.
Evidence-gated benchmark, comparison and promotion control plane for LLM inference serving changes.
Deterministic simulator for KV-cache admission, placement, movement, eviction, and recomputation policies
Offline, evidence-first bottleneck hypotheses for LLM inference traces
Trace-Aware Serving Controller: eval-gated inference policy optimization
Deep Agents and SvelteKit harness for authoring verifier-gated Bonsai workflow packs.
TypeScript inference SDK for self-hosted LLM, ASR, TTS and embeddings. GPU lifecycle coordination, model handoffs and clients for vLLM, Whisper, Chatterbox, Qwen3 and Kokoro.
Executable field guide and deterministic labs for LLM inference engineering across kernels, scheduling, KV cache, placement, and control.
Hands-on LLM inference labs on free T4 GPUs with vLLM: prefix caching, AWQ quantization, throughput knee, chunked prefill, KV-cache pressure. Every number reproducible.
Interactive inference-engineering lessons and a bilingual AI mentor, built with Next.js and the Vercel AI SDK.
A hands-on lab for how AI models get served: GPU memory, batching, the scheduler, parallelism and fleet economics. Drag sliders instead of reading the book.
To associate your repository with the inference-engineering topic, visit your repo's landing page and select "manage topics."