This work addresses the notion of prediction within the context of contemporary cognitive science. Predictive coding has emerged as a central framework for understanding perception, cognition, and action. In parallel, prediction also plays a key operational role in transformer-based Large Language Models, which have achieved remarkable success across a wide range of linguistic and cognitive tasks. Despite the shared emphasis on prediction, the term is used in importantly different ways across these domains, raising questions about their comparability. In predictive coding, prediction is understood as a hierarchical, generative process driven by error minimization. In contrast, Transformers are typically described as performing nexttoken prediction via conditional probability distributions. However, this standard account overlooks the distinction between training and inference and underestimates the role of context-sensitive computations. We explore whether Transformer-based models can be meaningfully compared to predictive coding beyond superficial analogies. To this end, we provide a more precise articulation of both notions of prediction and introduce criteria for their comparison, focusing on dynamical elements and context sensitivity. We argue that Transformer computation involves not only final next-token prediction but also intermediate, context-dependent predictions. This allows for a more nuanced comparison with predictive coding and supports a reassessment of how predictive mechanisms are functionally implemented across these frameworks.
Prediction in Predictive Coding and Transformers. A Conceptual Comparison
Garavaglia Fabrizia Giulia
;Pinna Simone;Giunti Marco;Giuseppe Sergioli
2026-01-01
Abstract
This work addresses the notion of prediction within the context of contemporary cognitive science. Predictive coding has emerged as a central framework for understanding perception, cognition, and action. In parallel, prediction also plays a key operational role in transformer-based Large Language Models, which have achieved remarkable success across a wide range of linguistic and cognitive tasks. Despite the shared emphasis on prediction, the term is used in importantly different ways across these domains, raising questions about their comparability. In predictive coding, prediction is understood as a hierarchical, generative process driven by error minimization. In contrast, Transformers are typically described as performing nexttoken prediction via conditional probability distributions. However, this standard account overlooks the distinction between training and inference and underestimates the role of context-sensitive computations. We explore whether Transformer-based models can be meaningfully compared to predictive coding beyond superficial analogies. To this end, we provide a more precise articulation of both notions of prediction and introduce criteria for their comparison, focusing on dynamical elements and context sensitivity. We argue that Transformer computation involves not only final next-token prediction but also intermediate, context-dependent predictions. This allows for a more nuanced comparison with predictive coding and supports a reassessment of how predictive mechanisms are functionally implemented across these frameworks.I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.



