LLM & Prompting
Aleks Gordic presents a detailed breakdown of vLLM's architecture for high-throughput LLM inference, highlighting core components such as paging attention, continuous batching, prefix caching, and KV-cache management. The post traces the path from a single-GPU offline prototype to multi-GPU, multi-node online serving, and covers advanced features like chunked prefill, prefix caching, guided decoding, speculative decoding, and disaggregated prefill/decode, supported by diagrams and examples.
General
The article documents an effort to back up the firmware of an older Lego NXT brick and explores multiple approaches to extract or execute code on the device. It covers hardware-level interfaces (JTAG, bootloaders), VM IO-Maps, and a memory-exploitation path that targets a function pointer inside the firmware to achieve native code execution, culminating in a method to dump the firmware. The piece emphasizes the security implications of embedded devices, discusses research ethics, and provides a detailed narrative of exploration and learning rather than a production-ready exploit guide.