RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The field guide · 120 retrospective records ↗
Turntaking Review

The field guide / Evaluation & evidence

Evaluation & evidence / From the field guide · 21 November 2023 event · prepared 16 September 2026

GAIA's 2023 questions were easy for people, hard for GPT-4

GAIA's paper reports a 92% human score against 15% for GPT-4 with hand-picked plugins.

Visual for this record: GAIA's 2023 questions were easy for people, hard for GPT-4
Visual published by cdn-thumbnails.huggingface.co, shown for identification of the record. Credit: cdn-thumbnails.huggingface.co · source page ↗ Rights: owner-review-pending.

The conversation

In November 2023, researchers from Meta, Hugging Face and academic partners published GAIA, submitted 21 November 2023, a benchmark of 466 questions designed, in the authors' words, to be 'conceptually simple for humans yet challenging for most advanced AIs'. The project's Hugging Face organisation hosts the released question set and leaderboard, with answers withheld for 300 of the 466 questions specifically to support that ongoing leaderboard.

What the documents show

The paper organises its questions into three difficulty levels defined by how much tool use and reasoning they require: Level 1 needs no tool, or at most one, and no more than five steps; Level 2 typically needs five to ten steps combining different tools; and Level 3 is built to test what the authors call 'a near perfect general assistant', requiring arbitrarily long sequences of actions. Answering a GAIA question can require web browsing, reasoning across steps, and handling more than one type of media. The paper reports human respondents scored 92% on average, while GPT-4 with plugins scored 15%, a gap the authors call notable given GPT-4's stronger performance on many existing professional benchmarks. The authors flag that the GPT-4-with-plugins figure is their own 'oracle' estimate, since plugins were chosen by hand for each question rather than selected automatically, calling it not an easily reproducible result on its own.

The system boundary

GAIA questions are graded against a single verifiable final answer, not the reasoning path used to reach it, which the paper argues keeps the benchmark simple to score even though the underlying task can require many steps. The paper treats human performance, not a theoretical ceiling, as the practical target: general assistant capability should be judged by whether a system approaches the reliability an average person already shows on this kind of question, rather than by whether it wins on narrower, specialised tasks.

Where it fails

Because the reported GPT-4-with-plugins score depended on a human manually selecting which plugin fit each question, the paper's own 15% figure describes a best-case, hand-tuned configuration rather than an automated system a builder could reproduce directly. The paper's own difficulty levels also show that most of the gap concentrates in the higher levels requiring longer tool-use sequences, meaning an assistant might look competent on Level 1 questions while still failing the benchmark's harder tiers. A builder should treat a later system's GAIA leaderboard score as a separate, later submission, not a component of this original paper's own reported numbers.

  • Was a cited GAIA score produced with automatically selected tools, or with a person choosing tools for the system as GPT-4-with-plugins did here?
  • Which GAIA difficulty level does the operator's own use case most resemble, and does the assistant's score hold up at that level specifically?
  • Does the assistant fail by picking a wrong tool, or by reasoning incorrectly once results are returned, and does that distinction change what needs fixing?

GAIA's own framing keeps the target modest and concrete: matching what an ordinary competent person can already do on real, multi-step questions, not a claim about general intelligence.

Sources & reading trail

GAIA: a benchmark for General AI Assistants ↗

The paper's own question design, its three difficulty levels, and its reported 92% human versus 15% GPT-4-with-plugins scores.

Source published: 21 November 2023 · Retrieved: 16 September 2026

gaia-benchmark organisation ↗

The project's own hosting of the released question set and leaderboard, with a withheld-answer subset for scoring.

Source published: Not established · Retrieved: 16 September 2026

Documentation, rulings and incident records establish the entry; the boundary reading is Chatbot Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.