NVIDIA: AI Security Is an Engineering Problem, Solved With Verifiable Controls at Every Layer of the Agent Stack

Published · AI Daily — AI-assisted deep research, methodology & disclosure

In a post on the NVIDIA Blog, Saša Zdjelar argues that AI security is an engineering problem. It needs defined security requirements, enforceable controls, named owners and evidence that protections work. The post splits an AI agent into three layers: the model, the harness that organizes context, tools and workflows, and the runtime where actions execute. Each layer needs its own controls. A security boundary must hold even when the agent makes the wrong decision. The post presents NVIDIA OpenShell, an open source secure runtime, and tools from partners such as Cisco, JFrog, CrowdStrike and Palo Alto Networks.

On September 21, 2026, NVIDIA published a post on its blog, written by Saša Zdjelar, with a blunt title: AI security is an engineering problem. The central claim is that security cannot stay at the level of slogans and prompts. It has to become defined security requirements, enforceable controls, named owners and evidence that protections work. As AI grows more capable, the post says, the industry must speed up security engineering, widen access to defensive tools and share what works faster. The post starts from a plain fact: technology changes, but security fundamentals endure. The internet and cloud computing changed how software operates. The core duties stayed the same: establish identity, control access, limit exposure and verify that protections work. AI agents add new capabilities. They reason, they use tools and they adapt their actions based on the data they meet. Those capabilities do not cancel the old principles. They force teams to apply them under new operating conditions. The post is also candid about the pressure. Organizations want the productivity of AI while the practices to govern and secure these systems are still developing.

The key to the whole piece is the full-stack view. Applications depend on code, data, identities, services and infrastructure, and security depends on how those parts work together. AI agents extend that system. The post divides an agent into three layers. Models provide capabilities. Harnesses organize context, tools and workflows. Runtime environments provide the infrastructure in which actions execute. Each layer carries security duties, and controls must exist at every layer as data, instructions and actions move through the system. No single layer can be trusted to catch everything. The post makes this concrete with one scenario. An agent is updating a customer record. It finds malicious instructions in an attached document and tries to export customer data to an unauthorized destination. A network policy should block the transfer. Protected logs should capture the attempted tool call, the authorization decision and the outcome, so the security team can identify the tool used and the destination the agent tried to reach. Then comes the point about permission granularity: permission to update a customer record should not automatically extend to exporting that data. An agent can request additional access, but it cannot authorize that access itself.

That leads to the strongest sentence in the post: a security boundary has to hold even when an agent makes the wrong decision. It follows that limits cannot depend on the agent's own judgment. The environment where the agent runs must install limits on files, network destinations and processes independently of the agent's reasoning. Instructions and safeguards can guide behavior, but security also requires enforceable boundaries. The post lists several engineering requirements. Each agent needs a traceable identity and credentials limited to its assigned task. Organizations need clear policies on what information agents can reach, which systems they can change and which actions require approval. Consequential actions and permission changes still need human approval. Teams must also verify the source and integrity of the tools, skills and dependencies agents use. The post also covers the day after an incident. Protected records of tool calls, authorization decisions and outcomes help investigators reconstruct what happened. Clear procedures for revoking access and containing incidents make that evidence actionable. In short, logs exist so that a response can rest on facts. On the product side, NVIDIA points to NVIDIA OpenShell, an open source, secure runtime. It enforces policies outside the agent's reach, provides sandboxed execution and governs how agents access data, network and system resources. The post says partners in the Open Secure AI Alliance are building on OpenShell. Cisco's DefenseClaw adds a governance layer. JFrog integrates with OpenShell to scan and verify agent skills and to enforce policies on which skills agents can access. The second major theme is evidence. Before deployment, teams need proof that controls block attempts to obtain credentials beyond an agent's scope or to send sensitive data to an unauthorized destination. Testing should also cover attempts to change permissions or to interfere with monitoring, and it should be repeated after material changes to models, tools or workflows. A named owner uses those results to decide whether the system is ready and makes sure failed tests lead to corrective action. Failures found in testing or operation should be reproduced, investigated and addressed. Each finding then becomes a repeatable test, so teams can check that the fix still holds in future releases. The post cites CrowdStrike's SafeMind, which tests and strengthens defenses through repeated attack simulations, and Palo Alto Networks Prisma AIRS, which provides continuous red teaming as models and applications change.

The third theme is the tools defenders hold. Investigating failures takes capable tools suited to the task, data and environment. The post says open and closed models serve complementary needs. Closed models offer managed capabilities and services. Open models let defenders inspect relevant components, adapt strategies and work on infrastructure they control. During an incident, that control can help a team reproduce a failure, test a fix against its own systems and keep sensitive evidence inside its environment. Capable AI can help find vulnerabilities, validate fixes and investigate attacks, and its value should be judged by reproducible findings, verifiable fixes and faster response. Examples include Capital One's VulnHunter for AI-powered code security and ReversingLabs' Spectra Assure for AI-powered analysis of software packages to detect malware and tampering. The closing section, on shifting the advantage toward defenders through open work, argues for sharing evidence of what failed and which controls worked. The source excerpt we received ends mid-sentence there, so we do not describe that section further. One caveat matters. This is a framework and position post. It gives no performance benchmarks, cost figures or measured reduction in risk, so readers cannot judge from the text how much these measures help. What follows is our analysis, not a claim from the source. For developers, the most direct lesson is to put the permission model and isolation outside the agent: give each agent its own identity and least-privilege credentials, and let the runtime, not the system prompt, enforce network and file limits. For enterprises, it means naming an owner for every agent deployment, making passed red-team tests a release gate and turning every incident into a regression test. For the ecosystem, the OpenShell-plus-partners picture shows a division of labor: the runtime enforces, while governance layers, supply-chain scanning and continuous red teaming fill the other roles. The challenges are plain. First, policies must be fine-grained enough to allow 'update but not export', and fine-grained policy costs effort to maintain. Second, supply-chain risk stays: verifying the source of skills and dependencies needs shared standards across the ecosystem. Third, continuous testing means continuous cost, and the post does not say who pays. Fourth, the post comes from a vendor that sells related platforms, so readers should weigh the framework against their own environment. Still, framing security as verifiable, accountable engineering work, not a one-time promise, is a direction the whole industry can use.

Sources