Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong
Summary
This Show HN discusses Cactus Hybrid and Gemma 4 E2B Hybrid, an on-device model that ships probes to score each answer with a confidence and routes low-confidence prompts to a bigger model. The post covers benchmarking against Gemini 3.1 Flash-Lite across multiple tasks, quantization details (4-bit, 3-bit), and code examples for using Cactus, MLX, and llama.cpp backends to run the hybrid pipeline, highlighting privacy-preserving, low-latency AI at the edge.