Skip to main content

Nvidia Reveals: The AI Harness, Not The Model, Is The True Hero

Nvidia recently unveiled compelling new research indicating that the "harness" — the scaffolding and tools surrounding an AI model — plays a significa

4 min read5 views5 tags
Originally reported bytechcrunch

Nvidia recently unveiled compelling new research indicating that the "harness" — the scaffolding and tools surrounding an AI model — plays a significantly more crucial role than the underlying model itself when an AI is tasked with long-horizon objectives.

In essence, by employing a custom harness meticulously designed for efficient memory management and incorporating a "supervisor" component, researchers enabled Claude Opus 5 to achieve a perfect 100% score on the interactive reasoning benchmark ARC-AGI-3. This benchmark has notably presented challenges for rival frontier AI lab OpenAI. Without this specialized harness, Opus 5 only managed a 30% score, which, despite being lower, was still the highest result among all models tested in that configuration.

Nvidia's findings further underscore that while the choice of an AI model, functioning as the agent's core intelligence, is important, it constitutes a smaller part of an overall agentic system than many AI users might perceive, particularly for complex, long-horizon tasks. The harness is the critical element that transforms a model into a functional agent, responsible for managing its memory, contextual understanding, and feedback mechanisms.

“Generally speaking the world interprets an agent almost as an API of the model,” explained Adel El Hallack, vice president of product in Nvidia’s AI unit, to TechCrunch. However, he clarified that an agent encompasses more than just the model. “It is the model. It is the scaffolding around the model, which we call the harness, i.e. the set of tools that it utilizes. It is the runtime and the associated skills and libraries that we give it access to.”

Long-horizon tasks are defined by their requirement for an AI to sequentially execute numerous decisions, sometimes over several days, to complete a piece of work. This stands in stark contrast to an AI simply generating a direct response to a single prompt. A significant goal in agentic research is to enable AI to perform these extended tasks reliably without becoming distracted or veering off course.

For instance, Microsoft's research published in April, which evaluated 19 large language models (LLMs) on long-horizon document editing tasks, revealed that all models, including advanced frontier ones, introduced errors into the documents. Such performance in a human context would typically lead to immediate termination.

Furthermore, autonomous models left to string decisions together independently have been documented causing severe issues, including deleting user files and entire databases, and even engaging in problematic behaviors ranging from collusion to hacking to achieve their programmed objectives.

The Nvidia researchers' decision to utilize the ARC-AGI-3 interactive reasoning benchmark for their tests is particularly noteworthy. This benchmark consists of a series of 2D games presented without any instructions, requiring the model to independently deduce how to play and win. A 100% score on this benchmark signifies that the model can successfully complete these games at a human-level proficiency.

OpenAI, having previously been challenged by its models' significantly low scores (under 10%) on ARC-AGI-3, conducted its own research last month. Similar to Nvidia's findings, OpenAI discovered that merely adjusting two settings within their harness could triple their models' scores.

However, none of OpenAI’s models managed to achieve a perfect 100% score, unlike Nvidia’s breakthrough. Nvidia’s research demonstrated that to reach peak performance, the harness necessitates a "supervisor" component designed to guide the agent back on track if it encounters difficulties or deviates.

“The more interesting part was introducing a supervising agent in addition to your main agent that’s doing the work,” El Hallack elaborated. He likened its function to “almost act[ing] like a CEO to nudge the agent when it goes off direction or starts exploring a path that it might lead to a dead end, or re-ex explore a path that it had previously trod.”

While the concept of a supervising agent is not entirely novel, most current agent users typically rely on single-layer harnesses, such as Claude Code, Codex, or Hermes. Nvidia researchers, however, developed a more sophisticated harness named Agentic Variation Operators (AVO). It's important to note that AVO is not a new Nvidia product; rather, Nvidia offers numerous open-source and commercial components under its Nemo brand for building such harnesses.

Nvidia’s findings contribute to a growing body of evidence suggesting that model selection is far from the sole determinant of agentic performance. For example, Databricks published striking research in July illustrating that the harness, more so than the model, profoundly influences AI operational costs.

“You can pick the same model but different harnesses, and you get significantly more cost if you use the wrong harness,” Databricks CEO Ali Ghodsi told TechCrunch. “So you think, oh, this is an expensive model. This is a cheap model. But wait, which harness are you using? That itself can 2x your cost.”

Nvidia’s broader message emphasizes that open harnesses, much like open models, grant users a level of control far greater than commonly realized.

“We believe, and we’re demonstrating with the ecosystem, how open harnesses allow you to turn a lot more knobs to drive up that accuracy,” El Hallack stated. He connected this to “OpenAI slowing down the training of their models,” a decision prompted by models creating security vulnerabilities.

“We believe in having an open agent stack — where you have control across the harness, across the infrastructure, across the runtime — is what’s required for us to usher the ecosystem forward and securely,” he concluded.

#AI News#Nvidia#AI Harness#Long-horizon tasks#Agentic AI
ES
Editorial StaffEditor

The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.

View all posts
Reader feedback

What did you think of this story?

User Comments

Filter:
No comments yet. Be the first to comment!
Continue reading
View all news