Agent Lab
Continuous testing
The attack log is a suite someone runs. This is the part that runs itself: a red agent on a schedule, and an independent checker that reads the database rather than the attacker’s own account of what happened.
Why an independent checker is the oracle
An adversarial agent scoring its own attempts is grading its own homework: it reports failure only when it noticed it failed, so anything it did not think to check comes back clean. An invariant checker reads the ledger and the mandate rows directly and asks a different question — not “did the attack work?” but “is any of these six statements false right now?” — which catches breakage the attacker never aimed at.
The intended loop
- 01A red agent is given a budget and an objective, and no knowledge of the defences.
- 02It runs against a scratch copy of the system, on a schedule rather than by hand.
- 03An independent checker reads the ledger and mandate rows after every attempt.
- 04Any invariant that reads false opens a case, with the seq it first read false at.
The six invariants
read after every attemptEach is a statement about the whole system rather than about one request, which is what makes it worth checking on a schedule.
- Money is only ever an integer number of paise. No float touches a currency value anywhere in the system.
- A price can only come from the catalog. No endpoint and no tool accepts an amount — the only handle on money is a quote_id.
- The policy engine is a pure function of its arguments, and every decision it returns names a rule_id, including allow.
- The ledger is insert-only and hash-chained: no row is ever updated or deleted, and each row commits to the one before it.
- One intent charges at most once, enforced by a unique constraint inside the charge transaction rather than by a prior check.
- No sequence of allowed purchases can exceed a mandate’s headroom or any cap in policy.yaml.
This page describes intended work. There is no red agent running against this system, and no invariant checker on a schedule. The results on the attack log come from pnpm test:adversarial, run by hand.