The Direct Answer

Three agent security architectures are commonly presented as strong answers for AI agents that browse websites, execute code, call tools, or operate inside enterprise systems: kernel-level sentinels, policy enforcement gateways based on Open Policy Agent, and identity-centered zero-trust controls. Each addresses a real weakness in conventional application security, but none proves that an agent’s goals, tool calls, generated code, or delegated actions are safe merely because the execution environment passed a technical control. The unresolved problem is intent-aware authorization: determining whether a particular action remains appropriate after accounting for the user’s request, the agent’s plan, accumulated context, data sensitivity, and the possible consequences of failure.

Also worth reading: How Do Deterministic AI Gatekeeper Architectures Prevent LLM Hallucinations and Security Breaches in Production Systems? · How Do Enterprise Multi-Agent Governance Frameworks Actually Function in Modern AI Architectures? · How should engineering teams design secure autonomous agent architectures in production environments?

As of 28 September 2026, these architectures should therefore be treated as enforcement layers rather than complete security models. Kernel-level monitoring can stop suspicious low-level behavior, but it cannot reliably distinguish a malicious plan from a benign plan expressed in legitimate operating-system calls. Open Policy Agent, or OPA, can make authorization decisions against explicit policy, but policy authors must anticipate dangerous combinations of otherwise permitted actions. Identity and zero-trust systems can issue short-lived credentials and restrict access, but a correctly authenticated agent can still make an unsafe decision. The most defensible answer is to combine all three with scoped tools, human approval gates, runtime evidence, session limits, and continuous evaluation.

Architecture One: Kernel-Level Sentinels

A kernel-level sentinel places monitoring or enforcement close to the operating-system boundary, where an agent’s process, files, processes, sockets, and system calls become visible. Meta’s reported Muse agent architecture is associated with this approach and illustrates a broader direction toward containing agent activity beneath the application layer. The attraction is coverage: a sentinel can observe behavior that would be invisible to a prompt filter or a tool-specific wrapper. It can potentially terminate a process that attempts credential theft, unexpected privilege escalation, unauthorized file access, or communication with a prohibited destination. This makes the pattern relevant to coding agents capable of installing packages and modifying repositories.

The gap begins when the question changes from “Did this system call violate a technical rule?” to “Should this agent be taking this action at all?” A browser agent may use a shell, file system, or network socket for entirely legitimate work, while a malicious agent may stay within normal operating-system permissions. Kernel controls also create operational risks because an overly broad denial rule can break builds, interrupt package installation, or prevent expected developer workflows. Conversely, a narrow policy may miss exfiltration through an allowed service. Sentinel telemetry is valuable for evidence and rapid containment, but it does not replace application-aware checks such as transaction size, destination reputation, repository scope, or the relationship between a tool call and the user’s original objective.

Security concernKernel sentinel strengthRemaining weakness
Process behaviorBroad runtime visibility and rapid terminationCannot infer whether an allowed action serves the user’s intent
File and secret accessCan block paths and sensitive memory regionsLegitimate reads may look identical to reconnaissance
Network activityCan restrict destinations, protocols, and portsApproved endpoints can still receive harmful or private data
Privilege escalationStrong enforcement near the OS boundaryCorrectly authorized code can still cause damage
InvestigationRich host-level evidenceDoes not explain the agent’s goal or decision path
Operational controlCentral containment pointPoorly tuned rules can block normal work
A practical implementation should use the sentinel as the final containment layer, not the first semantic check. Before execution, tools should be reduced to the smallest useful set, and high-impact operations should be checked against task, identity, resource, and environment attributes. During execution, the sentinel should emit signed events for new privilege use, secret access, persistence attempts, unusual child processes, and bulk outbound transfers. A reasonable pilot threshold is to deny any process that reads common credential paths, accesses another user’s files, or creates persistence outside an ephemeral workspace unless an explicit policy exception applies. The architecture works best for high-risk coding and infrastructure agents, where operating-system behavior is a meaningful part of the threat model.

Architecture Two: Policy Enforcement Gateways

A policy gateway evaluates agent actions before tools execute, using attributes such as user identity, agent identity, task classification, tool name, requested scope, environment, and risk score. OPA is a concrete policy technology used to make these decisions outside application code, and the Cupcake project referenced in the research context applies OPA-style controls to coding-agent execution. This architecture is attractive because policies can be versioned, tested, audited, and updated without modifying every agent. It also supports deny-by-default behavior: an action proceeds only when a policy returns an explicit allow, rather than whenever no rule happens to match a prohibition.

The unresolved issue is that policy is only as expressive and correct as its assumptions. A gateway can enforce “this agent may update repository X in branch Y,” yet it may not recognize that the proposed patch silently weakens authentication, removes a test, changes a deployment setting, or introduces a dependency from an untrusted source. Combinatorial actions are another problem: 10 individually harmless searches can become reconnaissance, 20 permitted file reads can become bulk data collection, and repeated approved emails can become spam. Security teams should therefore evaluate sequences of calls, cumulative data volume, rate limits, state changes, and reversibility. The gateway must distinguish read-only planning from execution and attach a decision identifier to every approved action so later evidence can be correlated with the policy that allowed it.

A mature policy design separates at least four decisions: whether the principal may invoke a tool, whether the tool may receive the requested data, whether the action is permitted in this environment, and whether the action requires human confirmation. Risk can be calculated from signals such as production access, personally identifiable information, financial value, destructive scope, privilege level, and deviation from a known task plan. Exact numerical thresholds should be calibrated through testing rather than adopted as universal constants, but useful starting points are a 10-minute expiry for ordinary tool grants, a 1-hour expiry for repository write access, and immediate revocation after a policy, credential, or environment change. OPA-style gateways are best suited to organizations that need centralized governance across several agents and tools, but they should not be sold as proof that the underlying model reasoning is correct.

Architecture Three: Identity-Centered Zero Trust

Identity-centered zero trust gives every human, service, and agent a distinct identity and continuously verifies access instead of trusting a network location. For agent systems, this means short-lived credentials, narrowly scoped tokens, workload identity, device or workload attestation, session-level policy, and auditable delegation. The research context points to Okta’s agent-identity work, the Blueprint Alliance involving Okta, AWS, Google Cloud, and other organizations, and Outerlimit’s extension of zero-trust concepts to AI-agent activity. The commercial interest is understandable because enterprises need a way to determine which agent is acting, on whose authority, with which permissions, and under what conditions those permissions should end.

Identity solves the “who is this?” question more effectively than the “should this exact plan be safe?” question. A delegated token can correctly represent a user while still allowing an agent to send the wrong file, invoke the wrong workflow, or pursue a harmful interpretation of the request. Traditional IAM can also become confused when multiple agents call the same service under a shared service account. The design should preserve the chain of delegation from the human request through the orchestrator, planner, tool broker, and destination system. Every step should use a separate identity where practical, and credentials should never be copied into prompts, logs, generated source files, or browser content. Access should be constrained by resource, operation, data class, and time rather than only by a broad role.

The zero-trust approach is strongest when combined with purpose-bound authorization. For example, access to a customer database might be valid for answering a support question but not for training a model, exporting records, or changing account status. Token claims should therefore include audience, task identifier, permitted operations, approved data classes, and an expiration no longer than the delegated task requires. A practical review threshold is to require a new authorization decision whenever the active tool, target environment, data category, or delegated objective changes. For high-impact tools—such as issuing refunds, changing production infrastructure, sending external communications, or modifying access policy—human confirmation should remain the default even when the caller is perfectly authenticated. Identity is thus the control plane for agents, but it still requires a policy and action-safety layer.

Why the Three Architectures Remain Incomplete

The unresolved security requirement is end-to-end intent and consequence analysis across the agent’s full action chain. A prompt may be benign while retrieved content contains hostile instructions; a model may generate correct code while a dependency is malicious; a tool may behave correctly while receiving excessive data; or a user may approve one narrow action that becomes harmful when repeated. Conventional security controls generally operate on known principals, explicit APIs, and predefined objects. Agents create dynamic plans whose permissible next step depends on context assembled from prompts, memory, tool results, files, and prior actions.

There is also an economic problem. Strong approval gates reduce autonomy, while weak gates increase exposure, and no single threshold fits every task. A read-only brainstorm involving public documents may tolerate nearly continuous operation. A coding agent with repository write access, shell access, cloud credentials, and network access should operate under tighter time, scope, and approval controls. A useful classification is based on four dimensions: data sensitivity, action reversibility, privilege breadth, and external reach. Agents with low sensitivity, reversible actions, limited privilege, and no external reach can usually operate semi-autonomously. Agents that combine sensitive data with production writes, money movement, or broad distribution should use a separate, tightly governed execution profile.

The three models are complementary, but complementarity does not remove their shared failure mode. A kernel sentinel sees calls after software asks the operating system to make them. A policy gateway evaluates metadata and requested actions before or during tool invocation. An identity system verifies the principal and authorization context. None alone observes whether the complete outcome remains coherent with the user’s actual goal. Secure agent platforms should add provenance, retrieval filtering, memory boundaries, output validation, task-level evaluation, and incident reconstruction. These controls should cover both intentional attacks and accidental harm, because incorrect tool selection, stale memory, malformed data, and faulty planning can create damage without any adversary being present.

A Practical Control Pattern for Agent Platforms

A defensible deployment begins with a threat model that separates direct prompt injection, indirect injection through retrieved content, tool poisoning, credential theft, data exfiltration, unauthorized action, excessive agency, and model or orchestration failure. Each agent should then receive a declared capability profile containing allowed tools, destinations, data classes, spending or transaction limits, time limits, and prohibited actions. The profile should be enforced in a dedicated tool broker rather than described only in a system prompt. The broker should use deny-by-default access, parameterized tools, schema validation, output isolation, and separate channels for trusted instructions and untrusted retrieved material.

Execution should proceed through explicit stages: authenticate the request, classify its risk, retrieve only necessary context, generate a proposed action, evaluate that action, obtain approval when required, execute under a temporary grant, inspect the result, and commit state only after validation. Approval interfaces should show the exact target, operation, expected data, scope, and consequence rather than presenting a vague “Allow agent access?” dialog. A user should also be able to deny one step, terminate the session, inspect tool history, and revoke delegated credentials. The platform should log prompts and tool results according to data-retention policy, but sensitive values should be redacted or tokenized before storage.

For an innovation lab, the same pattern can protect concept generation without blocking exploration. Public research and ideation can use a low-risk profile with no secrets, read-only web access, sandboxed code execution, and synthetic data. Prototype validation can add write access to isolated repositories and disposable test environments. Production-like experiments should not automatically receive the permissions of the user’s everyday systems; they should use staged environments and explicit promotion criteria. Useful launch gates include 100% of privileged tools having server-side authorization, 100% of secrets being issued through a broker rather than exposed to prompts, critical actions having rollback or approval controls, and every production-capable agent completing adversarial tests before receiving persistent credentials. These are engineering targets, not evidence that a system is secure.

Comparison and Alternatives

Selecting an architecture depends on where the agent can cause harm and what evidence an organization can collect. Kernel sentinels are strongest near code execution and host activity, policy gateways are strongest for centralized action governance, and identity-centered zero trust is strongest for delegation and access lifecycle management. None should be evaluated by feature count. The relevant questions are whether controls run independently of the model, whether failures are contained, whether decisions are explainable, and whether operators can revoke access quickly.

Decision factorKernel sentinelPolicy gatewayIdentity-centered zero trust
Primary control pointOperating-system boundaryTool and action requestHuman, agent, service, and workload access
Best protectionMalicious processes and host activityScope, sequence, and action policyCredential misuse and excessive delegation
Main evidenceSystem calls, process trees, file and socket activityPolicy inputs, decisions, and tool tracesToken use, attestation, and access history
Fast deploymentModerateHigh for centralized toolsHigh in existing IAM environments
Typical weaknessSemantic intent remains unknownPolicy gaps and stale contextAuthenticated harmful action
Best initial useSandboxed coding agentsMulti-agent tool platformsEnterprise and cross-system agents
Human approval needHighest for destructive host actionsHigh for high-risk or state-changing actionsRequired for sensitive delegated privileges
Alternatives include sandbox-only execution, model-level filtering, retrieval isolation, application firewalls, secrets management, and fully manual approval. Sandboxing is necessary but can be escaped or misused if network and credential controls are weak. Model filtering is probabilistic and should not authorize destructive actions. Retrieval isolation can reduce prompt-injection exposure, but it does not protect a tool after malicious information has already entered the workspace. Application firewalls can observe known protocols but miss novel workflows. A human-in-the-loop process can provide judgment, yet reviewers frequently approve routine prompts without reading them, so approval must be specific, short-lived, and backed by enforceable limits.

Common Mistakes and Cost Trade-Offs

A common mistake is treating security as a single product category rather than a set of independent decisions. Another is giving every agent one general-purpose cloud identity because permanent credentials make implementation easy. Shared accounts also destroy attribution, encourage excessive privileges, and make revocation unreliable. Teams may similarly confuse authorization with intent, or assume that a model’s refusal rate is equivalent to application safety. Refusals should be tested against realistic tasks, adversarial variants, tool failures, and repeated sequences; a system that refuses harmful prompts but mishandles a legitimate parameter can still create operational risk.

Cost should be evaluated across engineering time, runtime infrastructure, policy maintenance, evaluation, security operations, and incident response. Prices for cloud IAM, OPA, sandboxing, logging, and model APIs vary by provider, region, volume, and contract, so universal vendor price claims would be misleading. A small pilot may use managed identity, container isolation, an existing policy service, and a few controlled model endpoints, but enterprise deployment will add dedicated gateways, telemetry, data controls, red-team testing, and 24/7 response coverage. A practical budget rule is to reserve at least 20% of the first security release for testing and remediation rather than treating controls as fixed configuration. Budgets with less than 5% allocated to continuous evaluation may not support reliable agent operations, although the appropriate percentage depends on risk and scale.

The most expensive mistake is usually overloading human reviewers. If every action creates a prompt, users either ignore approvals or abandon the product, which turns a security control into an availability failure. Better systems ask for approval only at defined boundaries, such as an external email, production deployment, payment, deletion, or privilege change. Low-risk reversible actions can proceed automatically when their scope is narrow. The architecture should be action-specific enough that a reviewer can understand the consequence in under a minute, and the platform should show what changed since the previous approval. Security spending should therefore target decision quality, containment, and observability rather than a large number of shallow alerts.

When to Act and What to Measure

Organizations should act before an agent receives write access, production credentials, personal data, or authority to communicate externally. The minimum trigger is not a particular model release or agent count; it is the first time an agent can change a system that someone else depends on. A useful pilot should include at least 20 representative tasks, 5 adversarial tool sequences, and 3 failure modes such as timeout, stale memory, and partial completion. Track unauthorized tool attempts, policy denials, approval latency, rollback success, secret exposure, data egress, and false-positive rate. If the system blocks more than 10% of routine low-risk actions without a clear security reason, review the policy before adding more autonomy.

A staged rollout can use four gates. At gate one, the agent reads public or synthetic material and runs code only in a disposable sandbox. At gate two, it can access a dedicated repository with branch restrictions and short-lived credentials. At gate three, it can modify staging infrastructure after policy checks and reversible deployment controls. At gate four, it receives narrowly approved production actions, with human confirmation and automatic expiration. Promotion should require passing security tests, documenting the tool inventory, verifying incident contacts, and testing revocation. Any change to the model, system prompt, tool schema, data source, or policy engine should trigger regression tests because a security architecture is only current for the exact configuration that was tested.

The definitive conclusion is that three major agent security architectures still leave an unresolved gap: they control execution, authorization, and identity without fully determining whether an agent’s evolving plan is appropriate in context. Kernel sentinels, OPA-based gateways, and identity-centered zero trust are useful foundations, but they need task-scoped capabilities, provenance-aware data handling, action sequence checks, human approval at consequence boundaries, and post-execution validation. The right standard is not “the agent cannot technically escape,” but “the agent can do only what was durably authorized, can be stopped quickly, and produces evidence that permits a reviewer to reconstruct what happened.” That standard is demanding, yet it is more realistic than treating any one architecture as a complete answer.