Home · Blog · Research

AI Explanations Reveal How Language Models Mirror Brain Activity

AI Explanations Reveal How Language Models Mirror Brain Activity

How can we know what a large language model (LLM) is really doing when it predicts the next word in a sentence? Researchers have found that by using explainable AI (XAI) to highlight which words matter most for predictions, they can also map those explanations onto brain activity—suggesting a deeper connection between how machines and brains process language.

The research

Maryam Rahimi, Mohammad Reza Daliri, and Yadollah Yaghoobzadeh from the Iran University of Science and Technology and the University of Tehran conducted this study. They used several LLMs, including a version of GPT-2, and applied attribution methods—which measure the contribution of each word to the model's prediction—to generate 'explanations' for each input word. They then compared these explanations with fMRI data from participants who listened to narrative stories, drawing on publicly available datasets such as those from the University of California, Berkeley and Princeton University.

The team found that gradient-based attribution methods, which compute how sensitive the prediction is to each word, robustly aligned with brain activity across a broad network of language-related regions. These explanations contributed unique variance beyond simple acoustic and word-rate features, meaning they captured something extra that basic properties of speech couldn't. Notably, in early auditory cortex—the initial stage of hearing—these attributions outperformed the model's internal representations, which are typically used to measure brain alignment.

Using a technique called conductance, the researchers extended the analysis to individual layers of the model. They discovered that early layers exhibited greater sensitivity to word type (e.g., nouns versus verbs) and aligned preferentially with auditory regions, while the final layer's attributions were dominated by positional information and showed broad cortical alignment. This suggests that different layers of the model are using different types of information, and that the brain may integrate these in a way that matches the model's overall computation.

Why it matters

This research offers a new way to interpret what LLMs are doing internally, which can help improve transparency and trust in AI. It also provides a bridge between artificial and human language processing: by seeing which explanations align best with brain activity, we gain insight into the computations that underlie human comprehension. For anyone curious about their own cognition, it highlights that language understanding is not a single process but a layered one, combining sound, meaning, and context.

What you can do

You can strengthen your own language processing by engaging in active listening and reading. Try to predict what comes next in a conversation or a story, and notice which words carry the most weight. Puzzles like crosswords or word games can also sharpen your brain's ability to weigh different types of information quickly.

Source: arXiv q-bio.NC

Curious about your own brain? Take our free adaptive IQ test or try 306 brain training levels.

Curious about your own IQ?

Take our free, scientifically designed adaptive test across 7 cognitive domains. No signup required.

Take the free test