Petals — Run large language models at home, BitTorrent-style
Summary
Petals enables distributed LLM inference by loading parts of models on consumer GPUs and serving others over a BitTorrent-like network. It supports Llama 3.1, Mixtral, Falcon, BLOOM, with fine-tuning, PyTorch and Transformers compatibility, and offers Colab/demo access and GitHub docs.