Supervisory letter review, built from a skill and a tool
The expert's lens lives in a skill. Every number comes from a deterministic tool. And when the agent finds a gap, it improves by proposing an edit to the skill, which a person approves. The model is never retrained.
Tool deterministic code, same answer every timeModel reads, classifies, writesHuman owns the skill and signs off
The three parts
Skill, tools, harness
Skill
How the expert reviews a letter
A plain SKILL.md file owned by the Head of Regulatory Affairs: classify each finding, quote the obligation word for word, map it to the issue register, name one owner, never compute a date by hand.
Tools (MCP)
The bank's existing reports
All date math lives in tested code. A wrong number is a software bug, fixed once.
Harness
Any agent that speaks MCP
The agent chooses which tool to call and writes the memo. This recorded run used Claude in Claude Code. The same skill folder and MCP server can be loaded by any harness that supports skills and MCP.
Input
The letter
A synthetic supervisory letter to a fictional bank: one Matter Requiring Immediate Attention on agent entitlements in payments, two Matters Requiring Attention, and an observation.
Read the letterThe issue register the tools read
Run 1 · skill v1
The agent follows the skill, and finds a gap
Run 1 memo (full)Self-correction
The skill changes, not the model
The agent proposed one amendment: a rule for findings that repeat an issue closed after a prior letter. Here is exactly what changed in the skill file.
Awaiting skill owner approvalThe business owner of the skill reviews every amendment before it is used.
Skill v1 → v2, line by lineRun 2 · skill v2
Same tools, same numbers, better memo
Run 1 summary
Run 2 summary
Run 2 memo (full)Why it's built this way
Controls first, model second
No arithmetic in the model. Days to a deadline, overdue counts and days since closure all come from code with unit tests.
Institutional knowledge is a file the business owns. The review method is readable, versioned and approved by a person, like a procedure.
Learning is an edit you can audit. The improvement in run 2 is a visible diff with an owner, not a change hidden inside model weights.
The agent drafts, a person decides. Every memo ends "Prepared by an agent for human review. Not approved."
Related: Jevlis ka applies the same idea to a model's answers: treat every model output like an order and check it before it acts.