keywords:
embodied cognition
problem solving
causal reasoning
artificial intelligence
reasoning
Large Language Models (LLMs) have shown surprising capabilities in reasoning tasks despite lacking direct physical experience with the world. We examine LLMs' ability to reason about object affordances through a tool innovation task where one must select unconventional objects to replace typical tools. In a study comparing GPT-3.5-turbo and GPT-4o with human participants (N=100), we found that while GPT-3.5 performed significantly worse than humans (38.7% vs. 85.8%), GPT-4o with chain-of-thought prompting achieved human-level performance (85.0%). Qualitative analysis revealed that both models could identify causally relevant object properties, but GPT-4o was superior in flexibly applying these properties in novel contexts. We argue that this success relies on compositional reasoning—the ability to decompose objects into abstract properties and recombine them for novel uses. Our findings suggest that LLMs’ ability to reason about object affordances has progressed substantially, highlighting the need for further mechanistic research to characterise LLMs’ underlying abilities.