I tested 10 model/harness combinations on the same Three.js task
Summary
This article documents a benchmarking exercise comparing 10 model/harness combinations on a single Three.js task, detailing prompt setup, token usage, durations, and observed outputs. It serves as a practical data source for evaluating LLM prompting and model selection in code-generation tasks.