The Unified AI Risk Assessment Framework · v2.0
Agent Security Methodology.
Traditional cybersecurity and application-level LLM guardrails fail to protect modern enterprises from the unique threat vectors of autonomous AI agents. The Whitefin Methodology bridges this gap by unifying the structural topology of the enterprise AI system — the 7-Layer AI Security Stack — into a single, rigorous, quantifiable risk-scoring framework.
Five active layers, differentiated weights, twenty-four capability dimensions — producing a deterministic Total Defense Score that exposes the structural gaps current vendors leave open.
01 — The Stack
The 7-Layer AI Security Stack.
Every enterprise AI environment maps to seven distinct architectural layers — each governing a different boundary. The layers span two fundamentally different domains: the Semantic Control Plane (language and logic, where agents reason) and the Execution Infrastructure (OS and hardware, where actions actually run). The gap between these two domains is where most attacks succeed. The Bridge — L4 — is the only deterministic, real-time layer that closes it.
02 — Filtering the Stack
Two layers are prerequisites. Five form the evaluation engine.
To establish an accurate AI Agent Security score, two foundational infrastructure layers are excluded from the computation. They are strict prerequisites for enterprise operations — but they are not AI-agent-aware governance controls.
Physical / Infrastructure
Hardware, CPU/GPU, and cloud compute baselines. Standard operational prerequisites — cannot detect malicious prompt escalation or tool abuse.
Virtualization / Container
Standard container hardening and CNAPP tools. Cannot detect malicious prompt escalation or autonomous tool abuse.
03 — Evaluation Layers
Five layers. Twenty-four dimensions. Weighted by defensive leverage.
The Total Defense Score is a weighted normalization of scores (1.0 – 5.0) across 5 distinct operational layers. The 24 Capability Dimensions are distributed based on where enforcement actually occurs.
Application & Prompt Control
Evaluation of static input sanitization, structural prompt filtering, and post-generation text analysis. This layer operates at the semantic boundary where natural language enters and exits the model.
Reasoning & Intent Analysis
Evaluation of the agent's internal planning cycle and its interaction with untrusted external data. This layer addresses the Chain-of-Thought attack surface where indirect injection most frequently succeeds.
Orchestration & Tool Governance
Inventory management, schema monitoring, and configuration scanning of the agent integration plane and MCP ecosystems. This layer governs the tool surface the agent can reach — before execution.
The Process Boundary — The Bridge
The critical cryptographic translation layer. Evaluates the system's ability to bind an agent's semantic identity to a trackable, immutable OS process token. This bridge enables Causal Provenance — linking every kernel syscall back to the agent reasoning step that caused it.
Deterministic Execution Governance
The core enforcement layer. Synchronous, binary, low-latency controls at the kernel boundary via eBPF. Operates on an Assume Breach posture — acts regardless of the LLM's cognitive state. Every action inspected before execution.
04 — Total Defense Score
One formula. Deterministic. Same math for everyone.
Scores are normalized 1.0 – 5.0 per layer, where 1.0 = no capability and 5.0 = fully implemented and adversarially verified.
The Bridge — The Primary Differentiator
The Bridge score (20% weight) can only be achieved by vendors with simultaneous presence in both user-space SDK and kernel space, linked by cryptographic causal provenance. Vendors that operate only in one plane — model-side prompt filters, kernel-side syscall enforcement, identity-only governance — score 1.0 on this dimension regardless of how well they execute their primary domain.
Kernel-only vendors achieve strong L4 scores but cannot link a syscall back to the agent reasoning that caused it. Model-only vendors see the reasoning but never the action. The Bridge requires both — at the same time, in the same enforcement path. That is the architecture that defines the category.
05 — Key Architectural Insights
Four observations the framework forces you to see.
The Failure Mode Symmetry Trap
If your agent defense layer shares a failure mode with the agent itself — natural language processing — it can be manipulated using the same exploit vector. An LLM-based security monitor is vulnerable to the same prompt injection that compromises the agent it guards. Security must be enforced in a non-linguistic medium: kernel space and syscalls.
The Identity Delegation Trap
Verifying who an agent is at the orchestration layer (L5) is insufficient if that identity is lost when the tool spawns a local child process. A valid identity token can authorize a process that then executes entirely outside the governance boundary. Cryptographic identity delegation to the OS kernel — via Agent Passport — is mandatory to close this gap.
Note 8 Compliance — Deterministic-First Evaluation
Runtime evaluation must begin with the most efficient deterministic method — binary allowlists evaluated at the kernel boundary — before deferring to resource-heavy, probabilistic LLM-based reasoning. WhiteFin's Data Plane executes this deterministic step at p99 <10ms. Most competitors begin where determinism ends.
The Independence Prerequisite
Cloud providers (AWS, Azure, GCP) are the infrastructure agents act upon — they cannot simultaneously be the arbiter of what agents do to that infrastructure. LLM providers govern what the model says, not what the process executes. Existing security vendors depend on these platforms and cannot act as neutral enforcers. As Gartner states: the Guardian must be independent. Independence is not a feature — it is a prerequisite.
06 — Run It Yourself
Score any vendor with your own LLM.
We packaged this entire framework as a ready-to-paste evaluator prompt. Load it into any model and it turns into an architecture-based vendor analyst — one that scores on what a product's architecture actually enforces, not what its marketing claims. Nothing to install, no account, no data sent to us.
How to use this prompt
Copy everything into your LLM of choice — a Claude Project, ChatGPT Custom Instructions, a Gemini Gem, or any system-prompt slot.
Then ask: "Evaluate [Vendor Name] using this framework."
The model scores the vendor across 24 dimensions and 5 layers, produces a Total Defense Score, and returns a full gap analysis — using the same methodology published on this page.
Built on the Whitefin Unified AI Risk Assessment Framework. The prompt instructs the model to score on architecture, not marketing — and to score lower when a claim's enforcement mechanism is undocumented. Run it on us too.
The framework, painted
Four rows. Six columns.
Each cell is one of the 24 capability dimensions. Hover the grid; the dimensions become the painting that named our design system.
Blue Facade, 1914 — Piet Mondrian
click to restore
Use the framework against us.
We'd rather you score us than take our word for it. Bring the methodology to a call — we'll answer every dimension on the record.