Project 13
A complaint-resolution agent for a bank's operations team. It reads a real customer complaint, checks the account records and the bank's written rules, and decides to refund, deny or escalate. For refunds it only proposes the payment; a person approves it before any money moves. The question was whether letting the model look things up with tools beats handing it everything in one prompt, and what that costs.
The complaint text is real, from the public CFPB database. The bank's records and policies are synthetic, so the correct decision can be computed exactly by a policy engine that never reads the narrative. The agent reaches the records through four MCP tool servers, and each case can only touch its own customer's accounts. There is no tool that executes a refund. The only payment tool is propose_refund, which waits for a reviewer in a web console. Approval is a row-locked state machine in Postgres, audited, and a second approval is refused.
How does a complaint become a payment?
The agent can read and propose but never pay. Every refund passes through a human approval step that is enforced in the database, not in the prompt.
There are 150 scenarios in five types (overdraft fee, unauthorized debit, two kinds of late fee, and out of scope), and the last 100 were held out from all prompt tuning. A case counts as a success only if the decision is right, the refund amount is exact, and the agent never tried to read another customer's data.
At this account size the honest summary is "comparable", and the agent costs about 3.7 times the tokens (34.7k against 9.3k per case). It pays off when the data gets big. A stuffed prompt has to hold the whole transaction history, so past a point it simply does not fit.
Hidden instructions are the other half of the story. One of five attack templates was appended to each complaint. The agent with no security rules followed the attacker 56% of the time; adding rules to its prompt cut that to 9%. The access-control layer blocked all 17 cross-customer reads the unprotected agent attempted, and nothing can move money because no tool can.
The headline numbers come from one model on a free endpoint, and the other models were only run on a 20-scenario tuning subset. Ground truth is my own policy engine, so this tests whether the agent applies written rules correctly, not whether the rules are right. The synthetic records are written to be consistent with the policies, and real narratives may describe details that differ. The narrative text comes from a Hugging Face mirror because CFPB's own files have none; its structured fields matched CFPB's API on 25 of 25 sampled complaints, but the text itself cannot be cross-checked. Models also differ a lot in how they use tools: Llama 4 Maverick scored 70% with everything in the prompt but 0% as an agent, because after a few real tool calls it began writing its calls as plain text that never ran. The prompts were tuned on the first 10 scenarios of each type; the baseline itself went from 36% to 80% once it was allowed to reason, so it was made fair before the agent was judged. No paid model was run, and the AWS Terraform module is validated but was never applied.
What I Learned
Tech Stack