Can a MUD evaluate LLMs? A $99 proof of concept
Summary
The article describes cruciblebench, a research-stage benchmark placing language models in a persistent MUD to evaluate AI-agent behavior with hidden social objectives over 50 turns. It highlights phase-1 results, ongoing phase-2 work, and a design philosophy that uses constrained, persistent environments to surface measurable failure modes in AI agents. It also discusses sponsorship and collaboration aspects of the project.