Synthetic data · fictional bank
A working demo

Supervisory letter review, built from a skill and a tool

The expert's lens lives in a skill. Every number comes from a deterministic tool. And when the agent finds a gap, it improves by proposing an edit to the skill, which a person approves. The model is never retrained.

Tool deterministic code, same answer every time Model reads, classifies, writes Human owns the skill and signs off
The three parts

Skill, tools, harness

Skill

How the expert reviews a letter

A plain SKILL.md file owned by the Head of Regulatory Affairs: classify each finding, quote the obligation word for word, map it to the issue register, name one owner, never compute a date by hand.

Tools (MCP)

The bank's existing reports

    All date math lives in tested code. A wrong number is a software bug, fixed once.

    Harness

    Any agent that speaks MCP

    The agent chooses which tool to call and writes the memo. This recorded run used Claude in Claude Code. The same skill folder and MCP server can be loaded by any harness that supports skills and MCP.

    Input

    The letter

    A synthetic supervisory letter to a fictional bank: one Matter Requiring Immediate Attention on agent entitlements in payments, two Matters Requiring Attention, and an observation.

    Read the letter
    The issue register the tools read
    Run 1 · skill v1

    The agent follows the skill, and finds a gap

    Run 1 memo (full)
    Self-correction

    The skill changes, not the model

    The agent proposed one amendment: a rule for findings that repeat an issue closed after a prior letter. Here is exactly what changed in the skill file.

    Awaiting skill owner approval The business owner of the skill reviews every amendment before it is used.

    Skill v1 → v2, line by line
    Run 2 · skill v2

    Same tools, same numbers, better memo

    Run 1 summary

    Run 2 summary

    Run 2 memo (full)
    Why it's built this way

    Controls first, model second

    1. No arithmetic in the model. Days to a deadline, overdue counts and days since closure all come from code with unit tests.
    2. Institutional knowledge is a file the business owns. The review method is readable, versioned and approved by a person, like a procedure.
    3. Learning is an edit you can audit. The improvement in run 2 is a visible diff with an owner, not a change hidden inside model weights.
    4. The agent drafts, a person decides. Every memo ends "Prepared by an agent for human review. Not approved."

    Related: Jevlis ka applies the same idea to a model's answers: treat every model output like an order and check it before it acts.