Nvidia just showed that the harness, not the AI model, is now the real hero
Nvidia published some interesting new research on Friday suggesting it’s the harness, more than the underlying model, that is far more important when asking an AI to do long-horizon tasks. A harness is the software wrapper around an AI model — the tools, memory management, and rules that turn a raw model into something that can act on its own.

Why It Matters
This story touches on nvidia, not, model — topics readers are actively tracking. Review and add editorial context before publishing.
Key Facts
- Fact 1: The TL;DR: Simply by using a custom harness tweaked to handle memory well and including a “supervisor” boss-like component, researchers got Claude Opus 5 to achieve a 100% score on the interactive reasoning benchmark ARC-AGI-3 — a set of 2D games with no instructions, where the model has to figure out how to play and win, similar to how a human would.
- Fact 2: (That’s a benchmark that has particularly irked rival frontier lab OpenAI.) Without the harness, Opus 5 scored 30%, which was the top result among all the models tested.
- Fact 3: For example: Microsoft published research in April that tested 19 LLMs on long-horizon tasks involving document editing and discovered that all the models, including frontier ones, filled the documents with errors.
- Fact 4: A 100% score means that the model can beat the games as well as humans.
Nvidia published some interesting new research on Friday suggesting it’s the harness, more than the underlying model, that is far more important when asking an AI to do long-horizon tasks. A harness is the software wrapper around an AI model — the tools, memory management, and rules that turn a raw model into something that can act on its own.
The TL;DR: Simply by using a custom harness tweaked to handle memory well and including a “supervisor” boss-like component, researchers got Claude Opus 5 to achieve a 100% score on the interactive reasoning benchmark ARC-AGI-3 — a set of 2D games with no instructions, where the model has to figure out how to play and win, similar to how a human would. (That’s a benchmark that has particularly irked rival frontier lab OpenAI.) Without the harness, Opus 5 scored 30%, which was the top result among all the models tested. Nvidia’s research is another indicator that, while model choice does matter, the model itself — the part that acts as the agent’s “brain” — is a smaller part of an agentic system than many AI users realize, especially for long-horizon tasks.
(Original synthesis pending human/AI review — generated by the stub provider by selecting real sentences from the source material, not by writing new analysis or commentary.)
Original source: TechCrunch
Related Stories
Instagram’s ‘First Draft’ feature aims to make editing Reels less tedious

Take a look at Microsoft’s new 25th anniversary Halo accessories

Tonight marks your last chance to save up to $300 on a TechCrunch Disrupt 2026 pass
