Chaperoneflight recorder for AI agents on AWSSessions
All sessions

For judges: how Chaperone works, and how to check it

CloudTrail records an AI agent's AWS calls as the developer's own. Chaperone pulls them apart: the agent session, each MCP tool call, and every AWS call that tool made.

Architecture: CloudTrail and EventBridge in every region feed an ingest Lambda and a poller; events are stored in DynamoDB; an API Lambda answers the MCP server (signed, unmasked) and the website through CloudFront (masked, cached), with Bedrock, Cloud Control and IAM Access Analyzer behind it.
Click the diagram for full size.

The diagram, in words

  1. 1

    Capture

    CloudTrail records every AWS call in all regions. EventBridge forwards each one to us-east-1 within seconds; a poller Lambda fetches MCP tool calls and sign-ins every minute.

  2. 2

    Process

    An ingest Lambda files each call under who made it (agent or person), notes the channel (MCP, Terraform, CLI, SDK, console) and classes its risk with fixed rules.

  3. 3

    Store

    One DynamoDB table, events kept per actor for 90 days. Sessions are built when read, and each MCP tool call is joined to its AWS calls by request ID.

  4. 4

    Answer

    One API Lambda answers every question, with Cloud Control (what was left running), IAM Access Analyzer (least privilege) and Bedrock (a summary written once per session).

  5. 5

    Serve

    The MCP server calls it signed and unmasked. This site calls it through CloudFront, masked and read-only. Scripts and CI can call it signed too.


One API, three ways in

MCP is one door. The same answers reach people and pipelines too.

One API, three ways in: the MCP server (for the agent, SigV4, unmasked), the web console (for people, through CloudFront, masked and read-only) and scripts or CI (SigV4, same JSON) all call one Chaperone API Lambda. It reads sessions from DynamoDB and calls Cloud Control (what was left running), IAM Access Analyzer (least privilege) and Bedrock (a summary once per session). DynamoDB is filled from CloudTrail in all regions, through EventBridge within seconds and a poller every minute, into an ingest Lambda that records who acted, the channel and the risk.
Click the diagram for full size.
MCP serverfor the agent

Claude Code or any MCP-capable agent asks about its own session ("me") before it reports back.

Web consolefor people

This site: the flight path, every change riskiest first, and a plain-English summary per session.

HTTP APIfor scripts and CI

The same JSON over a function URL with IAM auth: anything that can sign an AWS request can ask.

Chaperone API

sessions · what happened · risky calls · review · least privilege · explanations

Every AWS call in the account

CloudTrail in every region, attributed to agent or person, classed by fixed rules

  • Same answer everywhere

    The MCP tool, this site and the API read one query layer. Only the masking differs.

  • Not tied to one agent or channel

    It records MCP, Terraform, the CLI, SDKs and the console, and flags an agent working on a person's identity.

  • Safe by construction

    Signed callers get the full record. This public site gets a masked, read-only view that can't start work.


The agent checks its own work

Chaperone is also an MCP server in the agent. Here Claude Code asks it "was anything I did risky?"

A real run, sped up, identifiers masked. The same run on this site: the session with the demo queue.

Claude Code created and deleted a queue through the AWS MCP Server, then asked Chaperone "was anything I did risky?" The answer comes from CloudTrail, which the agent didn't write and can't edit: one destructive call, on a queue the same session created.

risky_calls("me") →
{
  "risky": {
    "destructive": {
      "calls": [{
        "time": "2026-09-27T22:30:13Z",
        "action": "sqs:DeleteQueue",
        "target": "https://sqs.us-east-1.amazonaws.com/111122223333/chaperone-demo-queue",
        "via": "mcp",
        "tool": "aws___run_script",
        "reasons": ["sqs:DeleteQueue removes or stops a resource"],
        "on_resource_created_this_session": true
      }]
    }
  },
  "summary": { "destructive": 1 }
}

Verify it on a session

Look at three things in this session:

  1. 1

    The summary

    The top panel says in plain English what the agent did.

  2. 2

    The flight path

    One lane per channel. Click a key icon to see why an IAM change was flagged.

  3. 3

    The access it needed

    It was allowed "Action": "*". See what it actually used.

Open the start-here session 14 agent sessions · 6,999 AWS calls · 73 MCP tool calls

Other details

A real agent with its own identity

Claude Code runs on a Linux VPS and reaches AWS through the AWS MCP Server, signed in as its own IAM Identity Center user, chaperone-agent. CloudTrail records each tool call as an event from aws-mcp.amazonaws.com whose user agent names Claude Code. Those events are the top lane of every agent session here; click one to see the AWS calls under it.

Five MCP tools, for any MCP-capable agent

  • list_sessions — recent sessions in the account, newest first, each with its riskiest call.
  • what_did_the_agent_do — one session: each MCP tool call with the AWS calls it made, and what it created and deleted.
  • risky_calls — the calls above plain writes, grouped by risk class, with the rule that flagged each one.
  • review_session — what the session left running and its monthly cost, plus IAM deny guardrails checked by Access Analyzer.
  • least_privilege — the permissions the session actually used, as a policy, compared with IAM Access Analyzer's.

Asked for session "me", the agent reviews its own latest session before it reports back. The server runs with the developer's own AWS credentials, so it can't be tried from this page; the repository has the setup steps.

Rules decide risk, not a model

  • Serverless: CloudTrail, EventBridge, Lambda, DynamoDB, CloudFront, S3, Bedrock, IAM Access Analyzer, Cloud Control.
  • All infrastructure in Terraform. Risk classes come from fixed rules, tested on real recorded events.
  • The plain-English explanations are written once per session by a Bedrock model and stored.

ARCHITECTURE.md has the decisions behind it.

This site is masked and read-only

  • CloudFront marks every request from this site as public. In that view the API masks account IDs, role suffixes, identity IDs and IP addresses, and refuses anything that starts work. Try /api/replay?id=me: it answers 403. Nothing a visitor does calls a model.
  • Chaperone records and explains; it never blocks, deletes or reverts. A recorder that acts would be one more agent to watch.
  • Events appear a few minutes after the call, as fast as CloudTrail delivers them. AWS doesn't record the scripts an agent sends through MCP, so Chaperone shows what was called, not the code.

$0.07 for the whole account in September

By Cost Explorer, up to the 27th. Explanations cost a few cents per session, once.


Source code (PolyForm Noncommercial)