A new computational model of cognition, called IM-LEPP, integrates vision and language in a hierarchical, energy-based framework, offering a unified explanation for attention, language comprehension, and memory.
The Research
Subir Varma, an independent researcher, proposed IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing) in a paper posted on arXiv on August 8, 2026. The model extends a previous single-modality model (LEPP) to handle both visual and linguistic input. Drawing on the theory that generative neural networks serve as effective theories of cognitive dynamics, IM-LEPP represents cognition as latent states flowing through learned energy landscapes, rather than simulating neural circuitry.
The architecture follows a hub-and-spoke hierarchy, grounded in the controlled semantic cognition framework by Lambon Ralph et al. In this design, separate predictive-coding pipelines for visual objects, scenes, and language converge on a shared amodal hub modeled after the anterior temporal lobe. Each pipeline retains its own identity while being influenced by the current state of the hub, allowing every prediction to reflect the full multimodal context.
The paper demonstrates that IM-LEPP can explain attentional phenomena like inattentional blindness and Necker-cube bistability. It also accounts for established psycholinguistic findings, including surprisal theory, the N400/P600 ERP components, and garden-path sentence reanalysis. The model offers a falsifiable contrast with transformer-based language models, highlighting differences in trajectory-sensitivity for next-word prediction. Additionally, the paper discusses data-efficient language acquisition relative to LLMs, outlines a semantic/episodic memory subsystem, and situates the model against predictive coding, the free-energy principle, JEPA, and Hierarchical Temporal Memory.
Why It Matters
For the average person, this research suggests that our brain integrates information from different senses in a predictive, energy-efficient manner. Understanding this can help explain why we sometimes miss obvious things (inattentional blindness) or perceive ambiguous images differently (bistability). It also sheds light on how we process language in real time, including the brain's response to unexpected words or grammatical errors. This knowledge can inform practical strategies for improving attention, memory, and language learning.
What You Can Do
To leverage these insights, practice active observation and mindfulness to enhance your attentional focus. When learning new information, try to connect it with multiple senses (e.g., visual and verbal) to deepen encoding. Challenge your language skills by reading complex or ambiguous texts, which exercises your brain's predictive processing.
Source: arXiv q-bio.NC
Curious about your own brain? Take our free adaptive IQ test or try 306 brain training levels.