A new study finds that a single general intelligence factor explains far less of language model performance than commonly assumed, suggesting machine intelligence is only partially interpretable.
The Research
Researchers Faiz Ghifari Haznitrama, Afrizal Hasbi Azizy, and Faeyza Rishad Ardi analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Using factor analysis—a dimension-reduction technique borrowed from psychometrics—they tested whether AI performance is organized around a general, domain-free intelligence factor, similar to fluid intelligence in humans.
The dataset was super-sparse, meaning most models were not tested on most benchmarks. To address this, the team triangulated their analysis across different data densifiers and imputation methods. A robust pattern emerged: a general intelligence factor accounted for 70.8% of variance in model performance at the most generous estimate, and far less in most solutions. Content-similar benchmarks did not necessarily cluster together, and the g factor was not dominated by any common theme. Critically, there was a lack of evidence that standard "intelligence" benchmarks well-proxy this factor.
Why It Matters
These findings challenge current efforts to define, identify, and target general intelligence as a tangible construct in language model development. If first-order abilities are partially idiosyncratic and not identifiable in practice, then targeting a single conceptual ability is unsupported. For anyone interested in cognition, this mirrors debates in human psychometrics: general intelligence is real but not the whole story. Just as human cognitive abilities are a mix of broad and specific factors, machine performance may depend on many narrow skills that do not reduce to one number. This suggests that both human and machine intelligence are better understood as profiles of strengths rather than a single rank.
What You Can Do
Instead of fixating on a single IQ score, explore your cognitive profile. Take tests that measure different domains—memory, reasoning, verbal skills, spatial ability—and notice where you excel. Then, train those specific areas with targeted exercises. This evidence-based approach respects the complexity of real intelligence, whether in brains or machines.
Source: arXiv q-bio.NC
Curious about your own brain? Take our free adaptive IQ test or try 306 brain training levels.