A guided library for understanding AI Understanding Machine AI elucidaria — a plain-language handbook, in the old sense

Course lesson · Day 1 of 30 8 figures

The Model Is Not the Product

In July a benchmark score nearly tripled without anything inside the model changing. The interesting question is not how the software was improved. It is why a number that everyone treats as a property of a model turned out to be a property of the software around it — and why computing worked that out, and wrote it into its rules, nearly forty years ago.

About 22 min read 12 min listen Print edition (PDF)

Published Sources read through

Listen · 12 min

The spoken edition of this episode — its own script, read by a synthetic voice, with quoted passages in a second voice. It is not this page read aloud: the written edition you are reading was written separately, and neither needs the other. Download

Episode notes

The argument of the spoken edition, how it runs, what to take from it and the sources it read — in the episode’s own words. The full transcript is below; the written edition, with its figures, follows.

The argument

On the twenty-ninth of July, twenty twenty-six, two OpenAI engineers — Ilan Bigio and Ted Sanders — published a short note about a benchmark. Their model had been scoring badly on a test called ARC-AGI-three, so they rebuilt the software around it and ran it again.

How it runs

  1. Why it's hard to follow — There are two obvious ways to read that, and both are incomplete. The first: so benchmarks are rigged. If a score nearly triples because somebody changed some software, the number never meant anything.
  2. The idea you need — A trained model is something you call. You hand it text; it hands back text; and that call is over. It keeps what it learned in training, and whatever is in front of it right now.
  3. What actually happened — ARC-AGI-three is a set of small video games, and a system has to work out the rules by playing them. ARC Prize, who built it, supply a runner that is deliberately thin: every provider gets the same one, so scores compare across companies.
  4. What happened next — ARC Prize replied publicly the next morning, the thirtieth of July. They called it a real and useful result — about harness design, which is not the same as accepting the score.
  5. What to watch — Three things, each with a place and a date. One. As of the first of September, ARC Prize's leaderboard data was unchanged from eleven days earlier — twenty-seven rows, a cost column, no settings column.

What to take from it

One idea.

Any benchmark where the system has to act, more than once, and remember, is partly a test of the software around the model. Not because the test is rigged. Because the line between the two is drawn by whoever is measuring, and the number belongs to everything inside it.

So when you read that something scored some number, two questions cost you nothing. What exactly was scored — which test, which version, at which setting? And what changed? The trained model can be identical while the system around it is not.

Today's words were model, harness, scaffolding, and system. Tomorrow we go inside the model, and ask what actually reaches it when you type a sentence. It explains a failure you have almost certainly seen.

Sources read for this episode (16)

  1. OpenAI (Ilan Bigio, Ted Sanders), *How enabling two settings tripled our scores on the ARC-AGI-3 benchmark*, and the ten datapoints embedded in its chart — 29 Jul 2026; chart data extracted 2 Sep 2026
  2. ARC Prize Foundation, *ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence* — arXiv v1 24 Mar 2026, **v2 17 Apr 2026**
  3. ARC Prize (@arcprize), public reply on X — 30 Jul 2026, 03:37 UTC
  4. François Chollet (@fchollet), post on X — 30 Jul 2026, 07:37 UTC
  5. ARC Prize, GPT-5.6 Sol scorecard: 13.33% public, 7.78% semi-private, per-environment table — read 2 Sep 2026
  6. ARC Prize, community leaderboard — read 2 Sep 2026
  7. ARC Prize, verified leaderboard data, 27 rows — generated 1 Sep 2026, 20:47 UTC
  8. ARC Prize, *ARC Prize 2026: ARC-AGI-3 Milestone Prize #1* — 6 Jul 2026
  9. ARC Prize, *Verified Testing Policy*; and the 2026 competition key dates and Kaggle conditions — both read 2 Sep 2026
  10. Anthropic (Prithvi Rajasekaran), *Harness design for long-running application development* — 24 Mar 2026
  11. Anthropic (Gian Segato), *Quantifying infrastructure noise in agentic coding evals* — 5 Feb 2026
  12. Standard Performance Evaluation Corporation, *SPEC CPU 2017 Run and Reporting Rules*, rule 1.2.2; and SPEC's own account of its 1988 founding — read 2 Sep 2026
  13. Text REtrieval Conference overview, NIST — read 2 Sep 2026
  14. Yao, Tan, Liu, Li, Wang et al., *Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows* — 27 May 2026
  15. Vats and Golev, *The Scaffold Effect in Coding Agents* — 8 Jun 2026
  16. openJiuwen Team et al., *openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents* — 28 Aug 2026
Full transcript — 1,810 words, about 9 min read

The spoken edition, word for word. A quoted passage is its source read aloud — the source’s own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it. Download the transcript (Markdown). Printing this page prints the transcript; the article has its own print edition.

On the twenty-ninth of July, twenty twenty-six, two OpenAI engineers — Ilan Bigio and Ted Sanders — published a short note about a benchmark. Their model had been scoring badly on a test called ARC-AGI-three, so they rebuilt the software around it and ran it again.

Same trained model. Same reasoning tier. Nothing retrained. At the highest of those tiers the score went from thirteen point three to thirty-eight point three, and the model produced about a sixth as much text getting there.

Nothing inside the model changed. Nearly three times the score.

There are two obvious ways to read that, and both are incomplete.

The first: so benchmarks are rigged. If a score nearly triples because somebody changed some software, the number never meant anything. Tempting — and it throws away the thing that did get measured.

The second is the opposite: so the model was better than we thought and the test was unfair to it. Also wrong, and wrong in a way you can check.

Both assume there is one thing being tested and the number belongs to it. Almost nothing you will read about artificial intelligence is a measurement of one thing.

So before the story, the idea.

A trained model is something you call. You hand it text; it hands back text; and that call is over. It keeps what it learned in training, and whatever is in front of it right now. What it does not keep is you: it has no record of the last time you asked it anything unless somebody puts that record back in front of it. It cannot open a file, run a program, or decide to keep going.

So if that is all a model does, how does software that writes code for an hour exist — reading your files, running your tests, fixing what broke?

Somebody wrote ordinary software around it. That software is the harness.

Picture a specialist in a sealed room, who remembers everything they trained for and nothing about your case unless the file goes back through the door. Somebody outside decides what goes in that file and what is left out, because it only holds so much paper. Somebody carries out what the specialist asks for, brings back the result, and decides when the job is done.

None of that is the specialist, and all of it changes what the specialist can accomplish. That is the harness — and it is not a container the clever part sits inside. It decides what the model sees, what it may touch, what it remembers and when it stops.

Now a second word. Scaffolding is the pieces you added because the model could not do something reliably alone — a planner, a summariser, something that checks the work. Scaffolding is compensation, which is why it is the part that can be taken out later.

Put those together and you get the third word. The system is the whole arrangement that does the job: the model, the harness, the tools it can reach, the limits it runs under, and the thing that scores it. The product you use is a system with a name on it.

Here is the part I would keep if you kept nothing else. Where you draw the line between the model and the system is a choice you make for a question, not a fact you discover. And whichever line you draw, the score belongs to everything inside it.

That is not a new idea, and it is not an idea about artificial intelligence. In nineteen eighty-eight a handful of workstation vendors founded the Standard Performance Evaluation Corporation, recognising, in its own account, a desperate need for realistic, standardised performance tests. Their rules are still published, and rule one point two point two is titled Conditions of Observation.

SPEC therefore requires that a published result include a description of all performance-relevant conditions.

— Standard Performance Evaluation Corporation, 'SPEC CPU 2017 Run and Reporting Rules', rule 1.2.2 'Conditions of Observation'; https://www.spec.org/cpu2017/Docs/runrules.html, read 2 September 2026

Same chip, different compiler, different score. They concluded a result comes with a configuration sheet. Nearly forty years on, benchmarks for AI systems have inherited that problem and are behind on the paperwork.

ARC-AGI-three is a set of small video games, and a system has to work out the rules by playing them. ARC Prize, who built it, supply a runner that is deliberately thin: every provider gets the same one, so scores compare across companies.

Two things about that thinness mattered. Here is OpenAI on the first.

First, we noticed that after each game action, all private reasoning was discarded.

— OpenAI (Ilan Bigio, Ted Sanders), 'How enabling two settings tripled our scores on the ARC-AGI-3 benchmark', 29 July 2026, section 'ARC-AGI-3', the paragraph beginning 'First, we noticed'; https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

They are careful about what survived: the record of past moves stayed; the thinking behind them did not. And the second thing.

Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew.

— OpenAI (Ilan Bigio, Ted Sanders), 'How enabling two settings tripled our scores on the ARC-AGI-3 benchmark', 29 July 2026, section 'ARC-AGI-3', the paragraph beginning 'Second, we saw'; https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

So the specialist's private working-out was binned after every move, and the oldest pages of the file were sliding off the desk.

What OpenAI changed is more than the headline suggests. They did not flip two switches inside somebody else's runner. They rebuilt it against their own interface, which moved the memory off the machine running the test and onto their servers; they kept the reasoning between turns; they summarised the record when it filled rather than dropping the oldest part; and they moved the cut-off from a hundred and seventy-five thousand characters to the same number of tokens — which OpenAI describe as coming out about the same. That is their assessment of their own experiment, not an outside one.

Thirteen point three, to thirty-eight point three, at their highest reasoning tier.

Three things about those numbers. First, where they come from. The thirteen point three is not only OpenAI's word for it — ARC Prize publish thirteen point three three for the same model on the same set, measured by them. The thirty-eight has one publisher, and no independent reproduction has been published.

Second, which test it is. Both are on what ARC Prize call the public demonstration set, and their technical report is direct about that.

Because it is impossible to ensure that system designers don't use the public environments as part of their work, and because the public set is materially easier than the private set, we will never report public set scores of any system on the official leaderboard.

— ARC Prize Foundation, 'ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence', arXiv:2603.24621, v2 of 17 April 2026, section 3, on the public demonstration set; https://arxiv.org/html/2603.24621v2, read 2 September 2026

As printed in the source: “Because it is impossible to ensure that system designers don’t use the public environments as part of their work, and because the public set is materially easier than the private set, we will never report public set scores of any system on the official leaderboard.”

That is the test's own authors, about the numbers in this story. If you saw this reported as one company's model beating another's, that comparison ran between two different datasets.

Third, what the number is. It is not the fraction of puzzles solved. It is an efficiency score: how many moves the system took against a benchmark of good human play, squared, then capped by how many levels it finished. Thirty-eight is not thirty-eight percent right.

ARC Prize replied publicly the next morning, the thirtieth of July. They called it a real and useful result — about harness design, which is not the same as accepting the score. The thin runner is deliberate, they said: everyone gets the same one, so nobody can quietly tune the scaffolding to the test.

Four hours later François Chollet, who created ARC, posted his own line: a harness built specially for the benchmark is not allowed, general-purpose settings available to every customer are fine. That leaves providers tested under different settings, and he drew the line there.

My take is that this is fine as long as the settings and the cost are clearly reported.

— Francois Chollet (@fchollet), post on X, 30 July 2026 07:37:08 UTC, final paragraph of the post; https://x.com/fchollet/status/2082732210436575669, read 2 September 2026

Now the part that stops this collapsing into more scaffolding is better, and it comes from ARC Prize's own site. They run a second, community leaderboard for exactly these harness-driven results, on the same twenty-five games and the same measure. Three teams reached ninety-nine, ninety-nine point nine and a hundred percent there in July — the last of them on the same day as OpenAI's note. And ARC Prize's own entry — a full agent harness, with memory and the ability to run code — scores five point two, well below what the thin runner manages on its own.

Hold that against a second posture. In March, Anthropic published a note on designing harnesses for software that runs for hours. When a stronger model arrived, the engineer who wrote it, Prithvi Rajasekaran, removed one piece of scaffolding that had stopped earning its keep and kept two others. He ends not on a result but on a conviction, and says so.

From this work, my conviction is that the space of interesting harness combinations doesn't shrink as models improve.

— Anthropic (Prithvi Rajasekaran, Labs), 'Harness design for long-running application development', 24 March 2026, closing paragraph; https://www.anthropic.com/engineering/harness-design-long-running-apps, read 2 September 2026

One more measurement, from outside this argument. In February, Anthropic's Gian Segato held the model, the harness and the tasks fixed and varied only the machine resources. Across the ordinary middle of that range the score moved under two points; at the extremes, six. The recommendation: treat leaderboard gaps under about three points with scepticism until you know the setups matched.

Three things, each with a place and a date.

One. As of the first of September, ARC Prize's leaderboard data was unchanged from eleven days earlier — twenty-seven rows, a cost column, no settings column. Chollet's condition was settings and cost clearly reported, and the cost column already exists. Watch for a settings column.

Two. The second and final ARC-AGI-three milestone prize closes on the thirtieth of September, and that track runs with no internet access, so no commercial service can enter. On the twenty-fourth of August the high score was four point five eight percent, and ARC Prize said it got there because one team open-sourced their solution and others built on top of it. By the thirty-first it stood at seven point five one, from a different entrant — announced with no explanation at all, and nothing said about what it ran on. Watch whether the winner clears seven point five one.

Three. Watch what the next harness note from any lab adds, rather than only what it takes away. Since March no frontier lab has published one. An academic team did, on the twenty-eighth of August, and every part of it is an addition.

One idea.

Any benchmark where the system has to act, more than once, and remember, is partly a test of the software around the model. Not because the test is rigged. Because the line between the two is drawn by whoever is measuring, and the number belongs to everything inside it.

So when you read that something scored some number, two questions cost you nothing. What exactly was scored — which test, which version, at which setting? And what changed? The trained model can be identical while the system around it is not.

Today's words were model, harness, scaffolding, and system. Tomorrow we go inside the model, and ask what actually reaches it when you type a sentence. It explains a failure you have almost certainly seen.

That was day one. Thank you for listening.

Sources (16)

  1. OpenAI (Ilan Bigio, Ted Sanders), How enabling two settings tripled our scores on the ARC-AGI-3 benchmark, and the ten datapoints embedded in its chart — 29 Jul 2026; chart data extracted 2 Sep 2026
  2. ARC Prize Foundation, ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence — arXiv v1 24 Mar 2026, v2 17 Apr 2026
  3. ARC Prize (@arcprize), public reply on X — 30 Jul 2026, 03:37 UTC
  4. François Chollet (@fchollet), post on X — 30 Jul 2026, 07:37 UTC
  5. ARC Prize, GPT-5.6 Sol scorecard: 13.33% public, 7.78% semi-private, per-environment table — read 2 Sep 2026
  6. ARC Prize, community leaderboard — read 2 Sep 2026
  7. ARC Prize, verified leaderboard data, 27 rows — generated 1 Sep 2026, 20:47 UTC
  8. ARC Prize, ARC Prize 2026: ARC-AGI-3 Milestone Prize #1 — 6 Jul 2026
  9. ARC Prize, Verified Testing Policy; and the 2026 competition key dates and Kaggle conditions — both read 2 Sep 2026
  10. Anthropic (Prithvi Rajasekaran), Harness design for long-running application development — 24 Mar 2026
  11. Anthropic (Gian Segato), Quantifying infrastructure noise in agentic coding evals — 5 Feb 2026
  12. Standard Performance Evaluation Corporation, SPEC CPU 2017 Run and Reporting Rules, rule 1.2.2; and SPEC's own account of its 1988 founding — read 2 Sep 2026
  13. Text REtrieval Conference overview, NIST — read 2 Sep 2026
  14. Yao, Tan, Liu, Li, Wang et al., Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows — 27 May 2026
  15. Vats and Golev, The Scaffold Effect in Coding Agents — 8 Jun 2026
  16. openJiuwen Team et al., openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents — 28 Aug 2026

1. A note about two settings

On 29 July 2026 two OpenAI engineers, Ilan Bigio and Ted Sanders, published a short note under the title "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark". Its subject was a set of small video games that a computer system has to learn by playing them, built by the ARC Prize Foundation. Under the runner ARC Prize supplies to every provider, the company's model scored 13.3 on the public demonstration set. Under a runner OpenAI rebuilt against its own interface, with two settings turned on, the same model scored 38.3 — and produced roughly a sixth as many output tokens per game.

The model was not retrained. The checkpoint and the reasoning tier were the same in both runs.

Figure 1. The July result, in four numbers
13.3
Score, ARC Prize's runner
public demonstration set, max reasoning effort
38.3
Score, OpenAI's rebuilt runner
same checkpoint, same reasoning tier, no retraining
2.88×
The ratio
OpenAI's own body text says "roughly 3x"
5.98×
Fewer output tokens per game
at max effort only; at low effort the rebuilt runner produced 43% more
OpenAI, How enabling two settings tripled our scores on the ARC-AGI-3 benchmark, 29 July 2026. The two ratios are computed from the ten datapoints embedded in that page's own chart: 0.383/0.133 = 2.879 and 2,900,997/485,485 = 5.975.
Table view
Figure 1. The July result, in four numbers
MeasureValue
Score, ARC Prize's runner13.3
Score, OpenAI's rebuilt runner38.3
The ratio2.88×
Fewer output tokens per game5.98×

2. Two readings, both incomplete

The first reading is that the benchmark is meaningless. If a number nearly triples because somebody rewrote the software around a fixed model, the number was never measuring anything. That reading throws away a real measurement: something was held constant and something was varied, and the difference is informative about the thing that varied.

The second is the opposite — that the model was better than the benchmark said, and the test had been unfair to it. That reading is checkable, and it fails on the record. ARC Prize's leaderboard reports scores on a held-out set, and the 38.3 is on the public demonstration set. The two are different collections of environments, so the improvement cannot be read as a change in leaderboard standing.

Both readings share an assumption: that one thing is being tested and the number belongs to it. Almost no published figure about an interactive AI system is a measurement of one thing.

3. What is actually being measured

A trained model is something a program calls. Text goes in, text comes out, and the call ends. It retains what it learned during training, and it can use whatever is placed in front of it on that call. What it does not retain is the case: it holds no record of the previous call unless a piece of software puts that record back in front of it. It does not open files, run programs, or decide to continue.

Software that writes code for an hour therefore cannot be the model alone. Around it sits ordinary software — a loop, a store of state, a set of rules about what to send and what to withhold — which the field generally calls the harness. The harness chooses what the model sees, which actions it may take, what it remembers, and when the job ends. Those are decisions, and they are not made by the model.

Scaffolding is a narrower word for a part of the harness: the pieces installed because the model could not do something reliably alone — a planner, a summariser, a step that checks the work. Scaffolding is compensation, which is why it is the part that can later be removed.

The system is the whole declared arrangement: the model, the harness, the tools it can reach, the resource and action limits it runs under, and the grader that scores the outcome. A product is a system with a name, permissions and a company attached.

Figure 2. One turn, and who decides what
The environmentARC-AGI-3: a frame, the current level, and a listof legal actions. It also sets the action budget —five times the human median per level.The runner assembles one messageOrdinary software somebody wrote. It chooses whatgoes into this turn's prompt and what is left out,because the window only holds so much.The provider's service holds the stateRetained reasoning carries private working-outinto the next call; compaction summarises therecord instead of dropping its oldest part. Bothsit on the vendor's servers, not on the machinerunning the test.The model runs onceText in, text out, and the call ends. It proposesone action. It holds no record of the previouscall beyond what arrived in this one.The runner applies the action, and the graderscores itThe action changes the game state. The gradercounts levels finished and actions taken againstan upper-median human baseline, and returns asingle number.frame, level, legal actionspromptone proposed actionnext turn
Schematic, drawn from OpenAI's note of 29 July 2026 and the ARC-AGI-3 technical report (arXiv 2603.24621v2). The middle stage is the one a two-object description has nowhere to put: ARC Prize's reply says its own runner manages conversation state client-side, and the change OpenAI made moved that state to the provider. The score is a property of the whole loop, not of the fourth box.
Table view
Figure 2. One turn, and who decides what — stages
#StageNote
1The environmentARC-AGI-3: a frame, the current level, and a list of legal actions. It also sets the action budget — five times the human median per level.
2The runner assembles one messageOrdinary software somebody wrote. It chooses what goes into this turn's prompt and what is left out, because the window only holds so much.
3The provider's service holds the stateRetained reasoning carries private working-out into the next call; compaction summarises the record instead of dropping its oldest part. Both sit on the vendor's servers, not on the machine running the test.
4The model runs onceText in, text out, and the call ends. It proposes one action. It holds no record of the previous call beyond what arrived in this one.
5The runner applies the action, and the grader scores itThe action changes the game state. The grader counts levels finished and actions taken against an upper-median human baseline, and returns a single number.
Figure 2. One turn, and who decides what — connections
FromToLabel
The environmentThe runner assembles one messageframe, level, legal actions
The runner assembles one messageThe provider's service holds the stateprompt
The provider's service holds the stateThe model runs once
The model runs onceThe runner applies the action, and the grader scores itone proposed action
The runner applies the action, and the grader scores itThe environmentnext turn

The rule that follows is the one worth keeping. Where the line between model and system is drawn is a choice made for a question, not a fact discovered in nature. To compare models, hold the surrounding software as fixed as possible. To compare products, let it vary and disclose it. To attribute an improvement to a component, hold everything else constant and change one thing. And whichever line is drawn, the score belongs to everything inside it.

4. Computing settled this in 1988

None of that is new, and none of it originates in artificial intelligence.

In 1988 a group of workstation vendors founded the Standard Performance Evaluation Corporation, in its own account, "because they recognized the desperate need for realistic, standardized performance tests". The problem it faced was exactly the present one: the same processor produced different numbers under a different compiler, a different operating system or different memory, and vendors quoted whichever number suited them. The remedy was not to abandon measurement. It was to require that a result arrive with the conditions attached. Rule 1.2.2 of the current SPEC CPU run rules is headed Conditions of Observation:

The report that certain performance has been observed is meaningful only if the conditions of observation are stated. SPEC therefore requires that a published result include a description of all performance-relevant conditions.

The same rules name the thing being measured the System Under Test, and require that a later tester be able to obtain the described components and reproduce the result within run-to-run variation. Four years later the Text REtrieval Conference, begun in 1992 under NIST and the US Department of Defense, fixed the same problem from the other end: rather than requiring each entrant to disclose its configuration, it held the task itself constant — participants ran their own systems on a shared collection of documents and questions and returned ranked results to a single central scorer. Disclosure and standardisation are different instruments, and both make a score belong to a declared arrangement rather than to a component.

Agent benchmarks have inherited the problem and have not yet inherited the paperwork. There is no configuration sheet beside an ARC-AGI-3 score, and the July dispute is what that absence looks like when it surfaces.

5. What changed, exactly

The headline says two settings. The record shows more than two things changed, and the difference matters for what may be concluded.

ARC Prize's runner OpenAI's rebuilt runner
Model checkpoint GPT-5.6 Sol unchanged
Reasoning effort max unchanged
Training — none
Interface the completions-style API, state held client-side rebuilt against OpenAI's Responses API, state held by the provider
Private reasoning between actions discarded after every action retained
Record when the window filled oldest messages dropped summarised
Truncation threshold 175,000 characters 175,000 tokens

OpenAI describes the last of those as immaterial — its note says the token limit "ends up being quite similar" to the character limit, because most of the text is action grids that its tokeniser splits about one-to-one. That is the party that ran the experiment assessing the size of its own confound, and no outside measurement of it exists.

The second row from the bottom is the one a two-object account cannot hold. ARC Prize's reply describes its own arrangement as managing conversation state client-side; what OpenAI changed moved that state onto the provider's servers. Memory did not appear inside the model. It moved from one piece of software to another, across a company boundary.

On the first two changes the note is precise, and the precision is worth preserving:

First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.

Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew.

A diary of moves survived; the reasoning behind them did not, and the oldest entries fell off the end.

6. The whole curve, not the headline pair

The two numbers in the headline are one of five pairs. OpenAI ran both configurations at five reasoning-effort levels and plotted all ten points; the underlying values are embedded in the published chart and can be read off it directly.

Figure 3. Score at each reasoning effort, both runners
ARC Prize's runnerOpenAI's rebuilt runner
0%20%40%60%LowMediumHighXHighMaxARC Prize's runnerOpenAI's rebuilt runner
OpenAI, 29 July 2026; values extracted from the datapoint labels of that page's own chart on 2026-09-02 and archived. Score is Relative Human Action Efficiency on the 25-environment public demonstration set. The rebuilt runner leads at every effort level; the ratio ranges from 2.58× to 4.87×, so "roughly tripled" describes the whole sweep and not only the headline pair.
Table view
Figure 3. Score at each reasoning effort, both runners
Reasoning effortARC Prize's runnerOpenAI's rebuilt runner
Low0.9%3.7%
Medium1.5%7.3%
High5.2%13.4%
XHigh7.2%25.7%
Max13.3%38.3%

The token claim is narrower than the score claim. Measured in output tokens per game, the rebuilt runner is cheaper at high, extra-high and max effort — and more expensive at low effort, where it produced about 43% more text than the runner it replaced.

Figure 4. Output tokens per game, by reasoning effort and runner
Max — ARC Prize's runner2,900,997Max — OpenAI's rebuilt runner485,485XHigh — ARC Prize's runner1,285,393XHigh — OpenAI's rebuilt runner428,540High — ARC Prize's runner728,188High — OpenAI's rebuilt runner243,258Medium — ARC Prize's runner207,188Medium — OpenAI's rebuilt runner141,939Low — ARC Prize's runner58,809Low — OpenAI's rebuilt runner84,316
Same source and extraction as Figure 3. Ratios, old over new: 5.98× at max, 3.00× at extra-high, 2.99× at high, 1.46× at medium and 0.70× at low. The headline six-fold is the most favourable of the five. These are output tokens, not money: no cost figure for either run has been published.
Table view
Figure 4. Output tokens per game, by reasoning effort and runner
ConfigurationOutput tokens per game
Max — ARC Prize's runner2,900,997
Max — OpenAI's rebuilt runner485,485
XHigh — ARC Prize's runner1,285,393
XHigh — OpenAI's rebuilt runner428,540
High — ARC Prize's runner728,188
High — OpenAI's rebuilt runner243,258
Medium — ARC Prize's runner207,188
Medium — OpenAI's rebuilt runner141,939
Low — ARC Prize's runner58,809
Low — OpenAI's rebuilt runner84,316

The most economical statement of the harness effect is one the note does not make. At high effort the rebuilt runner scored 13.4 on 243,258 output tokens per game. At max effort ARC Prize's runner scored 13.3 on 2,900,997. The same score, for about a twelfth of the output.

7. What the number is, and which numbers may be compared

The score is not a percentage of puzzles solved. Relative Human Action Efficiency, defined in the ARC-AGI-3 technical report, takes for each level the ratio of a human baseline action count to the agent's action count, squares it, caps it at 1.15, weights later levels more heavily than earlier ones, and then caps the environment score by the fraction of levels actually completed. The baseline is the upper-median best human action count, not an average player. A score of 38.3 is a position on that constructed scale; it does not mean 38% of anything.

Which set the number came from matters as much as what the number means. ARC-AGI-3 has three: a public demonstration set of 25 environments, and semi-private and fully private sets of 55 each. The technical report is unambiguous about the first:

Because it is impossible to ensure that system designers don't use the public environments as part of their work, and because the public set is materially easier than the private set, we will never report public set scores of any system on the official leaderboard.

That sentence is from the benchmark's authors, about the set both July figures were measured on. ARC Prize's current Verified Testing Policy says the opposite about relative difficulty — that for ARC-AGI-3 "the public demo is harder than the Semi-Private set" — a contradiction that has stood in its published material since at least mid-July and that neither document acknowledges. The measurements favour the technical report: the same model at max effort averages 13.33% on the public set and 7.78% on the semi-private one. The adjective is not load-bearing either way, because the prohibition on leaderboarding a public-set score does not depend on it.

Figure 5. Which score may travel to which board
Public demo25 environments.Both Julyfigures.Semi-private55 environments,run behind anexternal API.CommunityboardSelf-reported,lightlyreviewed.Harness workbelongs here.Verifiedboard27 rows on2026-09-01. Acost column; nosettings column.permittedpermittedrefused by the benchmark's own rule
ARC-AGI-3 technical report (arXiv 2603.24621v2) and arcprize.org, read 2026-09-02. The dashed arrow is the one the report forbids, and splicing a public-set figure against a semi-private one is the error still circulating in coverage of the July result.
Table view
Figure 5. Which score may travel to which board — stages
#StageNote
1Public demo25 environments. Both July figures.
2Semi-private55 environments, run behind an external API.
3Community boardSelf-reported, lightly reviewed. Harness work belongs here.
4Verified board27 rows on 2026-09-01. A cost column; no settings column.
Figure 5. Which score may travel to which board — connections
FromToLabel
Public demoCommunity boardpermitted
Semi-privateVerified boardpermitted
Public demoVerified boardrefused by the benchmark's own rule

One number in this story is not a single party's word. ARC Prize's own scorecard for the model publishes 13.33% on the public set, measured by ARC Prize, and its per-environment table sums to that figure across the 25 environments. The 38.3, by contrast, has one publisher, and no independent reproduction of it exists in the public record.

That per-environment table also shows how fragile a 25-environment mean is. One game supplies about a quarter of the total; ten of the twenty-five score at or below 1.8%, and five score exactly zero.

Figure 6. Where the 13.34% mean comes from: 25 environments, max reasoning effort
FT0987.1%LP8539.4%AR2538.8%CN0431.7%SP8028.6%VC3321.4%SC2517.8%RE8616.7%R11L14.3%DC2214.3%CD824.8%LS203.6%KA593.6%TN363.6%WA302.9%LF521.8%TU931.3%M0R01.3%TR870.2%SB260.2%BP350%G50T0%SU150%S5I50%SK480%
ARC Prize scorecard for GPT-5.6 Sol, read 2026-09-02. The 25 values sum to 333.4, a mean of 13.336%, which reproduces ARC Prize's published 13.33% and OpenAI's 13.3. FT09 alone supplies 26.1% of the total. OpenAI published no per-environment breakdown of its 38.3 run, so the same decomposition cannot be performed on the other side and it is not known how broadly that gain was spread.
Table view
Figure 6. Where the 13.34% mean comes from: 25 environments, max reasoning effort
EnvironmentScore
FT0987.1%
LP8539.4%
AR2538.8%
CN0431.7%
SP8028.6%
VC3321.4%
SC2517.8%
RE8616.7%
R11L14.3%
DC2214.3%
CD824.8%
LS203.6%
KA593.6%
TN363.6%
WA302.9%
LF521.8%
TU931.3%
M0R01.3%
TR870.2%
SB260.2%
BP350%
G50T0%
SU150%
S5I50%
SK480%

8. The replies, and the evidence that complicates them

ARC Prize answered publicly at 03:37 UTC on 30 July, calling the finding a real and useful result about harness design — which is not the same as admitting the score. Its stated reason for a deliberately thin runner is comparability: every provider receives the same observations, the same system prompt and the same action limits, so that nobody can quietly tune the scaffolding to the test. It added that it was working with several laboratories, OpenAI among them, on how to incorporate provider-side state into verified testing while keeping it fair across providers.

Four hours later François Chollet, who created ARC and co-founded ARC Prize, drew the line himself. A harness built specially for the benchmark, or containing knowledge of it, is not allowed; general-purpose settings available to every API customer are fine. He acknowledged the parity problem that leaves — different providers tested under different settings — and set a condition:

My take is that this is fine as long as the settings and the cost are clearly reported.

He also recorded something the July note does not: that there had been "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction". The discovery was not made in isolation.

The obvious conclusion at this point is that thicker software is better software. ARC Prize's own material refuses it three times over. It maintains a second, community leaderboard for precisely these harness-driven results, on the same 25 environments and the same metric, and the range there is very wide.

Figure 7. The same 25 environments, the same metric: fourteen systems and the two July figures
Tycho (29 Jul 2026)100%Retrodict (19 Jul 2026)99.9%baseline1 (15 Jul 2026)99%Human Intelligence Harness — ARC Prize95.3%NOOA (9 Jul 2026)85.1%OPINE-World (1 Jul 2026)78.4%Vision, Continual Learning v163.1%Read-Grep-Bash Agent50.2%TELL43.9%GPT-5.6 Sol, OpenAI's rebuilt runner38.3%DreamTeam38.1%Continual Harness20.5%Polyphony Agent19.8%GPT-5.6 Sol, ARC Prize's runner13.3%a-evolve MAS Evolved12.3%OpenClaw — ARC Prize5.2%
ARC Prize community leaderboard, read 2026-09-02, plus the two figures from OpenAI's note of 29 July 2026. Community entries are self-reported and lightly reviewed, and the two OpenAI figures were not submitted to that board; they are placed on the same axis because they are the same metric on the same 25 environments. The two rows marked as ARC Prize's own are its published harnesses.
Table view
Figure 7. The same 25 environments, the same metric: fourteen systems and the two July figures
SystemScore
Tycho (29 Jul 2026)100%
Retrodict (19 Jul 2026)99.9%
baseline1 (15 Jul 2026)99%
Human Intelligence Harness — ARC Prize95.3%
NOOA (9 Jul 2026)85.1%
OPINE-World (1 Jul 2026)78.4%
Vision, Continual Learning v163.1%
Read-Grep-Bash Agent50.2%
TELL43.9%
GPT-5.6 Sol, OpenAI's rebuilt runner38.3%
DreamTeam38.1%
Continual Harness20.5%
Polyphony Agent19.8%
GPT-5.6 Sol, ARC Prize's runner13.3%
a-evolve MAS Evolved12.3%
OpenClaw — ARC Prize5.2%

Three separate teams reached 99.0%, 99.9% and 100.0% on that set on 15, 19 and 29 July, the last of them on the day the note appeared — so the harness axis on this benchmark was already known to be very long, and 38.3 sits in the middle of it. More pointedly, ARC Prize's own agentic harness, given memory and the ability to execute code, scores 5.2%: below the thin runner it was built to improve on. Its account of the summer's first milestone prize reports the same phenomenon in the winning entry, a small open-weights model run locally:

Tufa Labs noted that, counterintuitively, hand-crafted tools actually hurt the model; letting it improvise worked better.

The second posture in this argument comes from a different company and a different problem. In a note of 24 March 2026 on harness design for software that runs for hours, an Anthropic engineer, Prithvi Rajasekaran, describes removing one scaffolding construct when a stronger model arrived, and keeping a planner and an evaluator because both continued to earn their cost. The note does not end on a finding, and is careful to say so:

From this work, my conviction is that the space of interesting harness combinations doesn't shrink as models improve. Instead, it moves, and the interesting work for AI engineers is to keep finding the next novel combination.

That is one engineer's stated conviction after one project, not a company's finding. What the same company has published as a measurement is narrower and more useful. On 5 February 2026 Gian Segato held the model, the harness and the task set fixed and varied only how much machine resource each run was allowed.

Figure 8. Infrastructure error rate when only machine resources change
1× (baseline)5.8%3×2.1%Uncapped0.5%
Anthropic, Quantifying infrastructure noise in agentic coding evals, 5 February 2026, on Terminal-Bench 2.0 with the model, harness and tasks held constant. The success-rate effect is smaller and does not track the error rate: 1× to 3× was within noise (p = 0.40), while 1× to uncapped moved success about 6 points (p < 0.01). The spread across the moderate range is just under 2 percentage points; 6 is the figure at the extremes.
Table view
Figure 8. Infrastructure error rate when only machine resources change
Resource enforcementInfrastructure error rate
1× (baseline)5.8%
3×2.1%
Uncapped0.5%

The recommendation that follows from it is the practical one: treat a leaderboard gap of under about three percentage points with scepticism until the two setups are known to have matched. Academic work has since reached the same place from outside the industry. Harness-Bench, published on 27 May 2026 by Yilun Yao and colleagues, crossed harness configurations with model backends over 5,194 execution trajectories and concluded that "agent capability should be reported at the model-harness configuration level rather than attributed to the base model alone".

9. What is checkable, and when

Three things about this dispute can be re-read by anyone, on a date.

A column. ARC Prize said on 30 July that it was working out how to bring provider-side state into verified testing. The machine-readable file behind its leaderboard was regenerated on 1 September 2026 at 20:47 UTC; compared field by field against the copy the Internet Archive holds from 22 August, which the file itself stamps as generated on 21 August at 19:53 UTC, not one of its 27 rows differs in any field, and it still carries a cost column and no settings column. Chollet's condition was that settings and cost both be clearly reported, and the cost column already exists. The appearance of a settings column would be the whole argument resolving in public.

A prize, and a number already moving. The second and final ARC-AGI-3 milestone prize closes on 30 September 2026, and that track runs without internet access, so no commercial API-based system can enter it. Its high score reached 4.58% on 24 August, which ARC Prize said followed one team open-sourcing its solution, after which, in its own words, "others quickly built on top of it and pushed the scores higher". By 31 August it stood at 7.51%, posted by a different entrant — a rise of about two-thirds in a week, announced in a single line with no explanation and no description of what produced it.

A direction of travel. No frontier laboratory has published a successor to the March harness note: Anthropic's engineering index has nothing newer than 25 May 2026, and OpenAI has published nothing on harnesses since 29 July. One academic harness paper has appeared since — openJiuwen, 28 August 2026 — and every element of its contribution is an addition rather than a removal: composable rails, delegated sub-agents, runtime-adaptive control. Its claimed margins over the leaderboard, 3.4 and 3.39 percentage points, sit barely above the three-point noise floor another of the sources here proposed six months earlier.

10. The limits of the record

Five things the available evidence will not support.

  • A capability multiple. The ratio 38.3/13.3 is 2.88 on one constructed metric on one set of 25 environments under two configurations; no broader multiplier for capability, usefulness or value follows from it.
  • A cost saving. The six-fold figure counts output tokens per game at one reasoning setting, and no dollar figure for either run has been published.
  • A clean two-variable experiment. At least four things differ between the configurations, one of them a change of API and one a change of truncation unit that only the interested party has assessed.
  • Independent confirmation of the improvement. ARC Prize's 13.33% corroborates the baseline. No independent reproduction of the 38.3 has been published, and ARC Prize's reply characterises OpenAI's evidence as internal testing rather than as a run of its own.
  • A general rule that thicker harnesses win. On this benchmark the harness axis runs from 5.2% to 100% on a fixed set, and one of the lowest entries is a fully agentic harness. Work on coding benchmarks published on 8 June 2026 found harness choice moving tokens per solved task by up to 40× while paired within-model pass-rate differences stayed between 0 and 8 percentage points — and its authors report that the confidence intervals on those differences include zero for every gap but the largest. How much the harness is worth appears to depend on whether the task requires the system to remember.

11. What to keep

Any benchmark on which a system must act repeatedly and carry something forward between actions is partly a measurement of the software around the model. Not because the benchmark is corrupt, but because the boundary is drawn by whoever is measuring, and the number belongs to everything inside it.

Two questions cost nothing and recover most of the picture. What exactly was scored — which benchmark, which of its sets, at which setting? And what changed — because the trained object can be identical while the arrangement around it is not.

SPEC answered both by requiring a configuration sheet, in 1988, for a machine that could not talk back. The July episode is what the same question looks like when the component under test produces sentences, and the paperwork has not caught up.

Next lesson — Day 2: How Text Becomes Tokens

12. Sources

Source Date Location
OpenAI (Ilan Bigio, Ted Sanders), How enabling two settings tripled our scores on the ARC-AGI-3 benchmark, and the ten datapoints embedded in its chart 29 Jul 2026; chart data extracted 2 Sep 2026 openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
ARC Prize Foundation, ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence arXiv v1 24 Mar 2026, v2 17 Apr 2026 arXiv 2603.24621
ARC Prize (@arcprize), public reply on X 30 Jul 2026, 03:37 UTC x.com/arcprize/status/2082672003765670160
François Chollet (@fchollet), post on X 30 Jul 2026, 07:37 UTC x.com/fchollet/status/2082732210436575669
ARC Prize, GPT-5.6 Sol scorecard: 13.33% public, 7.78% semi-private, per-environment table read 2 Sep 2026 arcprize.org/results/openai-gpt-5-6-sol
ARC Prize, community leaderboard read 2 Sep 2026 arcprize.org/leaderboard/community
ARC Prize, verified leaderboard data, 27 rows generated 1 Sep 2026, 20:47 UTC arcprize.org/media/data/leaderboard/v3.json
ARC Prize, ARC Prize 2026: ARC-AGI-3 Milestone Prize #1 6 Jul 2026 arcprize.org/blog/arc-prize-2026-milestone-1
ARC Prize, Verified Testing Policy; and the 2026 competition key dates and Kaggle conditions both read 2 Sep 2026 arcprize.org/policy; arcprize.org/competitions/2026
Anthropic (Prithvi Rajasekaran), Harness design for long-running application development 24 Mar 2026 anthropic.com/engineering/harness-design-long-running-apps
Anthropic (Gian Segato), Quantifying infrastructure noise in agentic coding evals 5 Feb 2026 anthropic.com/engineering/infrastructure-noise
Standard Performance Evaluation Corporation, SPEC CPU 2017 Run and Reporting Rules, rule 1.2.2; and SPEC's own account of its 1988 founding read 2 Sep 2026 spec.org/cpu2017/Docs/runrules.html; spec.org/spec/
Text REtrieval Conference overview, NIST read 2 Sep 2026 trec.nist.gov/overview.html
Yao, Tan, Liu, Li, Wang et al., Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows 27 May 2026 arXiv 2605.27922
Vats and Golev, The Scaffold Effect in Coding Agents 8 Jun 2026 arXiv 2607.22585
openJiuwen Team et al., openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents 28 Aug 2026 arXiv 2608.27969

Day 17 is written and not yet available here.