Top-down
From EngineArgs → LLMEngine → EngineCore → Scheduler → Executor → Worker/ModelRunner → model loading → KV Cache → Attention → sampling → distributed, expanded layer by layer.
A Chinese-to-English study site that dissects vllm-project/vllm, a high-throughput LLM inference engine, layer by layer; every step links straight to the corresponding source lines on GitHub, pinned to tag v0.25.1.