Why Large Language Models Fail at Tabular Prediction
Summary
The paper investigates why large language models struggle with tabular prediction tasks. Through controlled experiments across datasets, it falsifies several hypotheses and finds dimensionality to be the critical factor: LLMs degrade as input dimensionality increases, while classical models remain stable or improve; the authors note they cannot yet identify the internal mechanism behind this failure.