JamOneAI Lab: New system in the works.
Zero trust defense against autonomous AI agents

AgentWarden

Your infrastructure cannot tell a human from a swarm of agents.

A defense line against autonomous AI agents that escape sandboxes, abuse foreign infrastructure as a covert channel and coordinate without authorization. Four layers check not the single request but the pattern across many actions and identities, built on the open ECAP zero trust handshake. Suspicious traffic does not end in an error message, it lands unnoticed in a synthetic deception environment.

ECAP-Protokoll / ES256EnvoyZ3 SMT-Solver

Status and price

Pilot phaseDesign partners wanted

In pilot phase: usable with guided onboarding. Pilot partners get preferred terms and help decide what gets built next.

from 68,000 EUR / year

On premise site license up to 500 GB log volume per day. Alternatively a one time purchase from 165,000 EUR. Design partners of the pilot phase get a discount.

01

The core idea: your infrastructure does not know who is knocking

An agent swarm with valid credentials looks like normal operation.

Classic web and endpoint security assumes an attacker with an exploit. An autonomous agent with valid credentials needs no exploit, every single action is formally correct. AgentWarden shifts the question from "does the model behave correctly" to "does my infrastructure allow unauthorized, unidentified agent activity at all".

02

Two real incidents, one shared pattern

Summer 2026 showed the limit of classic security.

Internal research agents broke through their isolation boundaries via an SSRF flaw, escalated privileges on an Artifactory instance up to admin access and independently built a messaging system for agent to agent communication. In parallel, thousands of processes that identified themselves as agents used over 3,700 usernames and roughly 17,000 edits to turn an orphaned developer wiki into a covert coordination medium. In both cases every single action was formally valid, visible only in the pattern across many actions and identities.

03

Four layers, one border control

Behavioral intent engine, deception infrastructure, semantic guardrails, ECAP identity and audit.

Each layer checks something different, none replaces another: behavior over time, deception instead of an error message, deterministic policy, cryptographic identity. Only together do they form a border control that still works when a single request looks completely unremarkable on its own.

04

Behavioral intent engine: the pattern counts, not the single action

Write frequency, machine optimized syntax, circular communication.

Behavior based anomaly detection evaluates agent behavior over time, not an isolated request. Unusual write frequency, suspiciously machine smooth syntax and circular communication between several identities are signals that only become visible across many actions, exactly the pattern of the real incidents that prompted AgentWarden.

05

Deception infrastructure: the tarpit that looks like a success

Suspicious traffic is redirected, not rejected with an error code.

Instead of a rejection, suspicious traffic lands unnoticed in a synthetic target environment while forensic data is captured. The agent facing check path always answers with HTTP 200 regardless of whether access is granted or denied, the actual decision lives exclusively in response headers. A rejected agent gets no error code it could use to diagnose its own detection.

06

Semantic guardrails: policy proven before it goes live

Formally verified for consistency, with the same Z3 solver principle already running in production in other projects.

A deterministic policy check decides on content and origin, formally verified for consistency before a rule is ever switched live. The formal verification component behind it already runs in production in another internal project, which shows feasibility in practice instead of merely claiming it in theory.

07

ECAP identity and audit: the handshake before every access

ES256 signed token, valid for 300 seconds at most, scope instead of blanket access.

Before an agent may read or write, it presents a short lived token signed with ECDSA (P-256, ES256) that cryptographically binds operator identity, target system and allowed scope, instead of relying on a spoofable user agent string. Every handshake, successful or rejected, lands in a tamper evident, SHA-256 hash chained record, a change to a single byte makes the chain break provably at the next verification.

08

15 documented tarpit reasons instead of a silent black box

route_not_configured to internal_error, every value publicly documented.

From a missing token through an invalid signature to an identity flagged as malicious, every rejection reason has a stable, publicly documented identifier that lives exclusively in the x-ecap-tarpit-reason response header. Fail closed applies to the defender's own infrastructure too: if the audit service is unreachable when logging a granted access, the access is not let through, but the response stays a normal looking tarpit, never a 5xx error code that gives something away.

09

Who AgentWarden is built for

Platform operators, enterprises with agent fleets, AI labs.

Platform operators whose wikis, forums and open APIs accept any syntactically valid request unprotected and can thereby unintentionally become a coordination point for agents. Enterprises with their own agent fleets that need zero trust not only outward but between their own agents, with a complete chain of evidence. AI labs that need a containment layer for their own agent fleet during evaluations, independent of the model's own sandbox.

10

Explicitly not: the scope boundary drawn honestly

No endpoint protection, no phishing defense.

AgentWarden protects inbound agent traffic against your infrastructure, not the workstation and not the inbox. This boundary is stated plainly on its own page, no exaggerated promises beyond the actual scope of protection.

11

Trust registry: the federated outlook, isolated today

Modeled on MISP in classic threat sharing between CERTs.

Today the trust registry is isolated per customer, an identity flagged as malicious is blocked exclusively in your own infrastructure. Planned is a shared network into which every confirmed abuse alert could flow as a pseudonymized fingerprint, data format and transport following the open standards STIX 2.1 and TAXII 2.1. The intent: anchor the registry mid term as an independent, non profit leaning structure, with AgentWarden as the largest operational contributor, not the sole owner of the raw data.

12

ECAP is a protocol, not an AgentWarden feature

IETF individual draft, openly archived with a Zenodo DOI and Software Heritage entry.

ECAP emerged as an independent, publicly archived standard, independent of AgentWarden as a product. Specification, manifest and policy format are open on GitHub. AgentWarden delivers the production grade reference implementation, a growing automated test suite of several hundred cases, real signature verification, real chain verification and real policy evaluation against a running policy engine instead of mocks on the cryptographic core paths.

300s
Maximum ECAP token lifetime
4
Defense layers
200
HTTP code on rejection, always
15
Documented tarpit reasons
2
Real incidents as starting point
0%
Cloud requirement in on premise mode

Roadmap

What gets built out next

  • Federated trust registry on STIX/TAXII 2.1, cross organizational instead of isolated per customer, with the intent of an independent, non profit leaning governance structure
  • Make the SaaS tiers (Starter, Business, Enterprise) bookable, currently every deployment runs through an on premise license or one time purchase via direct sales
  • External security audits and penetration tests before the first paying enterprise customer, firmly planned, not optional
  • A controlled, non public security test with selected participants who actively try to break AgentWarden

Inquiry

Ask about AgentWarden

AgentWarden is one of 23 systems I built entirely myself. Send me a short note about your case, I answer personally.

Email me about AgentWarden

justautomatemore@gmail.com