keywords:
cognitive development
cognitive architectures
machine learning
neural networks
The Transformer architecture used in LLMs has garnered widespread attention due to these model’s human-like conceptual knowledge and language understanding, yet understanding how these models’ capabilities result from experience-guided learning, and connecting this learning process with the structure in their training data, can seem intractable. Here we present preliminary steps to characterizing the developmental trajectory of a minimal Transformer trained on a next-token prediction task, using a simple dataset with quantifiable uncertainty and a simple, intuitively characterizable structure that captures some aspects of natural semantic structure learned by LLMs from large datasets. We show how the dynamic learning process of this model is a predictable consequence of the structure of the training data, exhibiting attested features of human semantic development, as captured in a theory of neural network learning dynamics (Saxe et. al. 2019) previously used to capture such dynamics in a network originally introduced by Rumelhart & Todd (1993).
