ESP32s3-LLM-Cluster
Summary
The ESP32s3-LLM-Cluster project demonstrates a distributed inference setup using seven ESP32-S3 boards to run a 0.4B LLM with 1.58-bit BitNet quantization over SPI. The master node handles tokenization and embedding while compute nodes execute attention and MLP layers, showcasing edge AI deployment across microcontrollers with on-device processing. The repo provides architecture details, workflow guidance, and a complete hardware/software stack under an MIT license.