llama.cpp
Summary
llama.cpp is an open-source project that enables running large language models locally on consumer hardware with no API keys or telemetry, keeping data on-device. It supports pairing with a local coding agent like pi-llama, with commands to serve a model and run locally, and emphasizes hardware-agnostic performance across GPUs and CPUs.