What’s happening with the ARC-AGI prize? Well, the 2026 contest launched in March, and applicants have until November 8 to turn in designs.
I wrote about the ARC AGI challenge a couple of years ago, and the work of Francois Chollet and others in exploring how to measure AGI. That work continues, as the singularity nears, according to many of those close to the AI race.
Today I wanted to write a little about the methodology of the people at ARC-AGI, and the underlying philosophy that drives it, because I think that it is really fascinating from a neuroscience perspective.
Explaining ARC-AGI Research
On the ARC-AGI web site, on a page detailing the approach of the team, you can find the following Francois Chollet quote, from his book On the Measure of Intelligence:
“The intelligence of a system is a measure of its skill-acquisition efficiency over a scope of tasks, with respect to priors, experience, and generalization difficulty.”
Now importantly, the page claims a distinction between two very different types of intelligence. Here’s the relevant paragraph:
“A principle underlying ARC-AGI’s design is the need to create a fair and meaningful comparison between artificial intelligence and human intelligence. To achieve this, ARC-AGI focuses on fluid intelligence (the ability to reason, solve novel problems, and adapt to new situations) rather than crystallized intelligence, which relies on accumulated knowledge and skills. This distinction is critical because crystallized intelligence, by definition, includes cultural knowledge and learned information, which would provide an unfair advantage.”
Indeed. The crystallized intelligence is just the acquisition of facts. This is why IQ tests don’t ask test-takers when the American Revolutionary War happened, or what ketchup is made out of. The ARC-AGI, similarly, looks at skills and adaptation, not crystallized knowledge.
The page continues to explain these “priors” with a reference to the work of Elizabeth Spelke, et. al., in something called “core knowledge theory,” referencing “those cognitive building blocks that are either present at birth or acquired very early in human development with minimal explicit instruction.”
I took a foray into core knowledge theory, and it got interesting.
The Gyrus and the Sulcus
First of all, here’s how something posted in Cognitive Psychology describes the discipline itself:
“Core knowledge theory, developed by Elizabeth Spelke and colleagues, proposes that human cognition is built on a foundation of innate knowledge systems that are present from early infancy. These core systems provide domain-specific knowledge about objects (physical reasoning), number (numerical reasoning), agents (goal-directed behavior), geometry (spatial navigation), and social partners (in-group preferences). Each system has a distinct evolutionary history and operates throughout life as the foundation for more complex cognitive achievements.”
But the same page then directly goes on to list a number of “key structures, that, as it turns out, represent specific brain anatomy: There’s the intraparietal sulcus, described as “a parietal sulcus containing regions critical for spatial attention, eye movements, reaching, and numerical magnitude processing, particularly in relation to number,” and the fusiform gyrus, “a cortical region on the ventral temporal surface involved in high-level visual processing of faces, words, and objects, particularly in relation to faces.”
If that all sounds like Greek to you, don’t be alarmed. Unless you’ve studied neuroscience in particular, it’s unlikely that you’ll have used the word “sulcus” before.
I looked it up: Wikipedia defines the sulcus as “a shallow depression or groove in the cerebral cortex.”
What’s my point? My point is that the theory is using actual brain anatomy and contours, as well as neural firing patterns, to construct the building blocks that it uses to diagnose cognition (my term).
What Do You Speak?
You can look up the rest of this interesting brain anatomy to understand how scientists are looking at whirls and gyrations across our gray matter to try to isolate how we think and what AI can capably imitate.
Anyway, the same ARC-AGI page goes further and makes a good point that is so often overlooked, not just in benchmarking, but in practical development of LLMs. It involves language.
“If a benchmark included, for instance, tasks involving written English,” the writers explain, “it would immediately disadvantage any AI that hadn’t been extensively pre-trained on vast text corpora. It would also disadvantage humans that did not know English. English, or any language, is a cultural artifact, not a measure of inherent cognitive ability. The same is true for any knowledge related to specific human cultures or practices.”
This is why Kiefer ended up making a sovereign AI platform for the nation of Greece. In Greek. Nations don’t want to outsource these powerful tools to another population speaking another language. We might think that English is “the language of the world,” but that’s somewhat narcissistic.
Learning Like a Young Person
So instead of that crystalized knowledge, which you tend to get as you age, the ARC-AGI prize wants to measure the kinds of thinking that take place “tabula rasa” if you will, like those of a very young person, and by that I mean, say 2-3 months. Writers for the ARC-AGI explain:
“If an AI system has access to extensive, task-specific prior knowledge that is not available to a human, its performance on that task becomes a measure of the developer’s cleverness in encoding that knowledge, not the AI’s inherent intelligence. By focusing solely on a small set of universally shared core knowledge priors, ARC-AGI ensures that success truly reflects the system’s ability to learn and generalize, placing the AI and human on a comparable footing.”
Keep it Simple?
Here’s one more way that the folks at ARC-AGI articulate the goal:
“Many AI benchmarks measure performance on tasks that require extensive training or specialized knowledge (Ph.D.-level problems). ARC Prize focuses instead on tasks that humans solve effortlessly, yet AI finds challenging, which highlight fundamental gaps in AI’s reasoning and adaptability. This approach reveals the essential qualities of intelligence – such as the ability to generalize from limited examples, synthesize symbolic rules, and flexibly apply known concepts in novel contexts – that current AI systems struggle to replicate. By emphasizing these human-intuitive tasks, we not only measure progress more clearly but also inspire researchers to pursue genuinely novel ideas, moving beyond incremental improvements toward meaningful breakthroughs.”
I’d say that’s a pretty good way to tackle benchmarking, and it makes a number of philosophical points about our relation to AI systems. And the brain anatomy thing is really next-level. So take a look, and maybe you’ll have an inclination to follow along as the 2026 contest comes to a close.

