Text-to-meowdio models
Summary
The article analyzes text-to-audio models and their expressive range, comparing multiple open models and prompting strategies. It introduces a methodology for exploratory data analysis, including interactive visualizations and a large-scale clip dataset, and emphasizes listening as a practical measure of output quality. It also highlights how prompts influence model behavior across different architectures and presents an interactive explorer for exploring the results.