DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Auto-research with codex: How I achieved a 232x Faster Kernel over baseline with Codex in GPU Mode's qr_v2 problem

Quality: 8/10 Relevance: 9/10

Summary

This article is a detailed case study of auto-research using AI agents (Codex, Claude, Modals) to optimize a GPU QR decomposition kernel. It documents the workflow, prompts, and logging that enabled many iterations, culminating in a 232x speedup over the baseline and a long discussion of algorithms like blocked Householder and WY updates. It also provides insights into prompt engineering, agent steering, and idea diversification for auto-research workflows.

🚀 Service construit par Johan Denoyer