Elena' s AI Blog

ARC-AGI benchmark and a hefty prize

15 Jun 2024 (updated: 05 Oct 2026) / 10 minutes to read

Elena Daehnhardt

AI robots play a grand puzzle, Midjourney 6.0 art, February 2024


TL;DR:
  • Information about the ARC-AGI benchmark competition on Kaggle: advancing general intelligence through puzzle-solving challenges with substantial prizes.

Previous: Part 5 — Robots and True Love

Next: Part 7 — Narrow AI, General AI, Superintelligence, and The Real Intelligence

What Is the ARC-AGI Benchmark and ARC Prize 2024?

The ARC-AGI benchmark (Abstraction and Reasoning Corpus for Artificial General Intelligence) is a program-synthesis test that measures an AI system’s ability to generalise to novel reasoning tasks it was never trained on. Recently, I received an email informing me about the Kaggle competition launching ARC Prize 2024, which is built on this benchmark. Below is some general information about the ARC-AGI puzzle benchmark: what it measures, how its tasks work, and what the prize asks of participants. What is so special about this competition?

Why the ARC-AGI Benchmark Matters

The ARC-AGI benchmark stands out for several reasons:

  1. Focus on Generalisation: Unlike many AI benchmarks that test performance on specific tasks, ARC-AGI emphasises the ability to generalise to novel problems. It assesses an AI system’s capacity to learn new skills and solve tasks it hasn’t been explicitly trained on.

  2. Measures Fluid Intelligence: ARC-AGI aims to measure general fluid intelligence similar to what humans possess. This involves abstract reasoning, pattern recognition, and problem-solving abilities applied to unfamiliar situations.

  3. Minimal Prior Knowledge: The tasks in ARC-AGI require minimal prior knowledge. They focus on core reasoning skills rather than relying on extensive domain-specific information.

  4. Human-Level Performance: Humans generally score high on ARC-AGI tasks (around 85%), while current AI systems lag significantly behind. This indicates that ARC-AGI presents a challenging frontier for AI development.

    Update: that gap closed faster than most expected. In December 2024, OpenAI’s o3-preview model scored 87% on this same benchmark at high compute, edging past the human baseline, though its huge compute cost meant it couldn’t compete for the Kaggle prize itself (ARC Prize’s analysis of o3). That breakthrough is a big part of why the ARC Prize team introduced a harder ARC-AGI-2 benchmark in 2025.

  5. Prize Competition: The ARC Prize, a $1,000,000+ competition, was launched to encourage researchers to develop AI systems that can beat the benchmark and potentially contribute to progress towards Artificial General Intelligence (AGI).

Is the ARC-AGI Benchmark a Puzzle Game?

You can try to test your human or bot intelligence with the ARC Prize website, which is very easy for humans but difficult for AI:

The ARC task is solved

The ARC task is solved

ARC-AGI itself is not a puzzle game in the traditional sense of entertainment. However, the tasks within the benchmark often resemble puzzles. They consist of input and output grids with visual patterns, and the goal is to figure out the rule or transformation that generates the output from the input.

The tasks require:

  • Pattern recognition: Identifying the underlying relationships and rules within the visual patterns.
  • Logical reasoning: Applying the identified rules to generate the correct output for new input patterns.
  • Abstraction: Understanding the core concept or principle behind the pattern transformation rather than memorising specific examples.

These cognitive skills are similar to those used in solving puzzles, making the tasks feel like puzzles. However, the purpose of ARC-AGI is not entertainment but rather to assess and advance AI capabilities in abstract reasoning and generalisation.

So, while ARC-AGI is not a puzzle game per se, the nature of its tasks often evokes a similar problem-solving experience.

Why the ARC-AGI Benchmark Is Unique

The ARC benchmark is defined in this GitHub repository (along with the dataset and testing interface):

ARC can be seen as a general artificial intelligence benchmark, as a program synthesis benchmark, or as a psychometric intelligence test. It is targeted at both humans and artificially intelligent systems that aim at emulating a human-like form of general fluid intelligence.

Overall, ARC-AGI is a unique and essential benchmark as it:

  • Challenges Current AI Limitations: Highlights the gap between current AI capabilities and human-like general intelligence.
  • Promotes Research on Generalisation: Encourages the development of AI systems that can learn and adapt to new tasks, a crucial step towards AGI.
  • Offers a Standardised Measure: This measure provides a standardised way to assess progress in developing AI with general problem-solving abilities.

ARC Prize 2024: Prize Money and Rules

ARC’s grand prize is $600,000 for the first team achieving 85% accuracy on the private evaluation set, on top of $50,000 in progress prizes and $75,000 in paper awards. The total ARC Prize 2024 pool exceeds $1,000,000. Anyone can join the competition. At the time of writing, MindsAI is leading the public leaderboard.

Update: ARC Prize 2024 closed in November 2024, and the grand prize went unclaimed — no team reached 85% on the private evaluation set. MindsAI did post the highest score overall (55.5%), but the rules required an open-source submission to collect prize money, so that entry was ineligible. The open-source winner was the ARChitects (Daniel Franzen and Jan Disselhoff), who topped the leaderboard at 53.5% (ARC Prize 2024 Winners announcement). ARC Prize has continued every year since: 2025 introduced a tougher ARC-AGI-2 benchmark with a $700,000 grand prize that also went unclaimed (ARC Prize 2025 Results and Analysis), and a 2026 edition is now running.

For further information, you can explore the following resources:

Discussion: Is AGI Possible in the Near Future?

I could not resist sharing my opinion here :)

General intelligence is about a machine’s ability to learn and adapt. The competition is about AI that does not memorise but solves open-ended problems like humans. It is a very important step in achieving progress in AGI and improving AI’s ability to acquire skills and become more inventive as it evolves.

Would AGI be possible in the near future?

In my opinion, we could imitate human reasoning and teach machines to acquire new skills and become self-learners to a certain extent in the next five to ten years. Please subscribe to get updated on my future AI predictions on this blog :)

What do we still need for AGI to become a reality sooner? Some argue that AI’s power is correlated with the number of parameters and computational resources it uses. I agree that more parameters can allow for smarter AI networks. However, more parameters are not necessarily better. Consider the possibility of AI systems memorising the data by heart and “overfitting.” (read about machine-learning overfitting in my post Bias-Variance Challenge)

We really want AI systems to become “inventive” and intelligent. If we want human-like AGI, we need more than merely ten or a hundred times more parameters — we would need to approach the scale of neurons in the human brain.

To achieve real AI intelligence, we have to move outside the scope of digital representation. Why? Because we humans do not think discretely. Our neurons work through chemical reactions, giving us far more computational and non-stochastic power than any AI, even the smartest one.

To create AGI beyond automation and content generation, such as in LLMs, we would need non-discrete systems — call it “reinventing the wheel” if you like. Should we instead focus on understanding natural intelligence first? There might be more potential in that path.

It would also help to learn more about the human brain, to understand how intelligence and creativity develop. Only once we truly understand what makes intelligence work in nature can we start modelling it, and possibly improving on it in AI.

Conclusion: The ARC-AGI Benchmark as a Measure of Generalisation

The ARC-AGI benchmark is a generalisation test that distinguishes AI systems which reason on unseen tasks from those which merely memorise training data. I hope this explanation sheds light on its significance! Will you join the ARC competition?

Try the following fantastic AI-powered applications.

I am affiliated with some of them (to support my blogging at no cost to you). I have also tried these apps myself, and I liked them.

Chatbase provides AI chatbots integration into websites.

CustomGPT.AI is a very accurate Retrieval-Augmented Generation tool that provides accurate answers using the latest ChatGPT to tackle the AI hallucination problem.

Flot.AI assists in writing, improving, paraphrasing, summarizing, explaining, and translating your text.

MindStudio.AI builds custom AI applications and automations without coding. Use the latest models from OpenAI, Anthropic, Google, Mistral, Meta, and more.

Originality.AI is very effecient plagiarism and AI content detection tool.

Did you like this post? Please let me know if you have any comments or suggestions.

Posts about AI that might be interesting for you






References

1. ARC Prize 2024

2. Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI)

3. ARC Prize website

4. Bias-Variance Challenge

5. ARC Prize 2024 competition page

6. ARC Prize 2024 Winners & Technical Report

7. ARC Prize 2025 Results and Analysis

8. ARC Prize’s analysis of OpenAI’s o3 on ARC-AGI-1

desktop bg dark

About Elena

Elena, a PhD in Computer Science, simplifies AI concepts and helps you use machine learning.

Citation
Elena Daehnhardt. (2024) 'ARC-AGI benchmark and a hefty prize', daehnhardt.com, 15 June 2024. Available at: https://daehnhardt.com/blog/2024/06/15/arc_agi_benchmark_prize/
All Posts