Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
Summary
The article reveals a massive scan of Hugging Face public datasets, uncovering 7.6 petabytes across 187 million files and 221,303 live credentials in 6,003 datasets. It highlights real-world risk, including cloud keys, database access, and tokens that could enable code changes or infrastructure access, impacting both AI providers and users. The piece underscores the need for pre-publish scanning, credential rotation, and coordinated disclosure with vendors to mitigate supply-chain risks in AI training data.