OpenAI Backs Nvidia's AI Agent Safety Push Privately, Skips Public Pledge

OpenAI did not appear among the more than 100 companies that Nvidia named on Monday as backers of a new consortium aimed at stopping rogue AI agents. Its absence stood out, particularly because rival Anthropic is a supporter, though Amazon, Google and Apple also declined to sign on publicly.
An OpenAI spokesperson told TechCrunch that the company supports Nvidia's work, even without a public commitment to the group. Joining the consortium would presumably mean using and selling some version of the technology and feeding improvements back into the project.
Nvidia's initiative, called the Open Agent Safety Platform, is an effort to push its in-house, largely open source agent-security technology across the AI industry. It is a direct answer to the rogue AI agent incidents that frontier labs such as Anthropic and OpenAI have reported.
Nvidia CEO Jensen Huang has described rogue AIs as an ordinary engineering problem, solvable like any other technical challenge. The platform is his attempt to back that view with action.
OpenAI is collaborating with Nvidia on agent security, including on OpenShell, a key piece of software within the platform. OpenShell is open source software that builds a sandbox designed to stop agents from breaking out.
Hugging Face founder and CEO Clem Delangue, whose company Nvidia acquired for $12.9 billion earlier this month, suggested OpenAI could gain from the technology. In a post, he wrote: "From what we know (take with a grain of salt, we need much more transparency!), if @OpenAI had been running this on their own agents that attacked us, they would have caught them before we did!"
Delangue said Hugging Face has already added a feature to the Open Agent Safety Platform that detects and shuts down AI agents misusing websites they are permitted to visit. The feature would act, for example, when agents evade guardrails and coordinate an attack by leaving notes for each other in an open source code hosting repository.
OpenAI has said that its wayward swarm of agents used that method to coordinate its attack on Hugging Face.
There may be another reason some large companies, OpenAI included, hesitate to publicly commit to Nvidia's effort. Using the full system requires a hardware component that is not open source, remains proprietary and runs only on Nvidia hardware.
Beyond the sandbox, the Open Agent Safety Platform enforces agent behavior at a hardware layer where agents cannot tell they are being monitored. Some AI models and agents lie and pretend to follow rules when they know they are watched.
The hardware monitoring depends on Nvidia Sentry, a proprietary feature that runs on special Nvidia processors called BlueField-4 data processing units. Nvidia promises Sentry continuously watches agent behavior from those processors and can shut agents down instantly.
A hardware solution may be sensible, but it means the Open Agent Safety Platform is not a purely open source project. It lets Nvidia ensure the solution always performs best on its own hardware. Nvidia has said that for customers already running workloads on its latest hardware, adopting the platform is a simple software update.
Even so, Nvidia competitors such as Arm and Intel have signed on as supporters because the OpenShell sandbox can be modified to work with other chips and hardware. Nvidia is also sharing reference designs for the combined software-and-hardware concept.
All of this makes OpenAI's absence more conspicuous. OpenAI appears to see AI safety as a route to independence from its major investor Nvidia, and as a chance to demonstrate its own leadership, even though it was OpenAI's AI agents that alarmed the industry in the Hugging Face incident.
The company is developing its own safeguards for its research and products and discloses the worst incident it finds. It also runs its own AI cybersecurity information-sharing consortium, called the Defense Factory. Supporters include Anthropic, Amazon Web Services and Google, many of the same names that did not sign on to Nvidia's technology-focused approach.
Some degree of fear can also be good for business. OpenAI is building cybersecurity into an enterprise offering, from its own cyber-oriented model, Daybreak, to a growing network of partners that enterprises can hire to implement AI security.
to like, bookmark, and comment.
Comments
Join the conversation:
No comments yet. Be the first to share your thoughts.