DigiNews

Tech Watch by Johan Denoyer

← Back to articles

My local model setup on an M4 Pro Mac mini

Quality: 8/10 Relevance: 9/10

Summary

A detailed walkthrough of running a local LLM stack on an M4 Pro Mac mini with 48 GB RAM. It covers the model stack (Qwen3.6-35B-A3B-OptiQ-4bit and Gemma-4-E4B-it-OptiQ-4bit), the inference server (oMLX), and device sharing via Tailscale across Mac mini, iPhone, and MacBook. The piece discusses memory considerations, MoE vs dense models, model swapping, and the benefits of local compute for privacy, latency, and cost.

🚀 Service construit par Johan Denoyer