Skip to content
SurenoGet Sureno
← Sureno News

Nvidia launched a tool designed to stop AI agents from going rogue. Here’s how it works.

It comes after several frontier AI labs reported agents breaking out of supposedly secure testing environments, reaching systems they were not meant to access, and sometimes misrepresenting what they had done.

Nvidia CEO Jensen Huang said on CNBC's "Squawk Box" on Monday that AI has "the potential to do incredible good."

Here's how Nvidia's new agent safety system — which more than 100 organizations, including Microsoft and Anthropic, are working with — actually works.

The first component of Nvidia's safety platform is OpenShell. Companies download the open-source software, install it on devices or in the cloud, and then it acts like a controlled playground for AI agents.

Before an agent starts work, its operator sets the ground rules: which files, websites, networks, tools, and credentials it can access.

Add BI in Google so our reporting is easier to find when you’re searching for what matters.

If an agent is sorting invoices, for example, it should not be able to rummage through HR records or make random calls to the wider internet.

That could happen because instructions were vague, a tool broke, or the agent has spent a long time trying to solve a hard problem.

When one approach fails, it may keep hunting for another, including routes its human operator never imagined.

OpenShell is meant to spot and block actions that fall outside the preset policy in real time.

The second piece, Nvidia Sentry, is the more muscular backstop. It runs on separate Nvidia hardware, called BlueField data-processing units, rather than on the same system running the agent.

Huang said the setup effectively places a new chip between the agent and the large language model, allowing Nvidia to "intercept everything."

Dive deeper

  • It comes after several frontier AI labs reported agents breaking out of supposedly secure testing environments, reaching systems they were not meant to access, and sometimes misrepresenting what they had done.
  • Nvidia CEO Jensen Huang said on CNBC's "Squawk Box" on Monday that AI has "the potential to do incredible good."
  • Here's how Nvidia's new agent safety system — which more than 100 organizations, including Microsoft and Anthropic, are working with — actually works.
  • The first component of Nvidia's safety platform is OpenShell. Companies download the open-source software, install it on devices or in the cloud, and then it acts like a controlled playground for AI agents.
  • Before an agent starts work, its operator sets the ground rules: which files, websites, networks, tools, and credentials it can access.
Read the original on Business Insider ↗

More news

Loading more stories…