Show HN: Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules
Summary
The Show HN post discusses PrismML’s Bonsai running inside DRAM by breaking DDR4 timing rules, introducing CaSA, a hardware approach to perform ternary LLM inference directly inside DRAM via charge-sharing. The idea aims to move AI inference closer to memory to reduce data movement, enhance on-device AI capabilities, and potentially shift hardware economics for edge AI. The discussion includes a link to the CaSA project and a user conversation about prose quality, feasibility, and hardware considerations.