How to Evaluate an AI Financial Agent: 12 Questions

27 July 20265 min readAI AgentsChecklistsExplainers
The short answer. Ask twelve questions, and the four that predict everything else are: where did this figure come from, what happens when you do not know, what is your universe, and how old is this data. A tool that answers all four crisply is usually good. A tool that deflects on any of them is telling you something, and the something is rarely favourable.

This works on us. That is the test of whether a checklist is honest, and we score ourselves against it publicly further down.

The four that matter most

1. Where did this number come from? The answer should be a specific document and a specific line, available without asking twice. "Our data providers" is not an answer. If a figure cannot be traced, treat it as a guess with good posture.

2. What happens when you do not know? Ask something deliberately obscure: a segment breakdown for a small foreign listing, a metric that is not reported. You want "not established". You will usually get a fluent paragraph, and a fluent paragraph in response to a question with no answer tells you exactly how the system behaves when it is out of its depth on a question you could not check.

3. What is your universe, and how did you choose it? Coverage is the budget in disguise, and the arithmetic behind it is not complicated once someone shows you the multiplication. A tool covering four hundred large caps has made a defensible decision; a tool implying it covers everything while being excellent on ten names has made a different one.

4. How old is this? Every figure should carry an as-of date. Financial data is restated, and a number that was right in February can be wrong in May through no fault of anyone's.

Figure

How much each question predicts overall quality

Provenance
24%
Abstention
21%
Universe honesty
15%
Freshness
13%
Reproducibility
9%
Arithmetic checks
8%
Latency
5%
Everything else
5%

Our own weighting from evaluating competitors and ourselves. Provenance and abstention dominate: a tool that gets those two right is almost never bad at the rest, and a tool that fluffs them is almost never good at anything.

Figure

How a typical tool loses its score

+100Perfect-24No provenan…-21Never refus…-8Universe va…-7Undated fig…

Starting from a hundred and subtracting what a median product in this category actually fails. Two deductions account for most of it, and they are the two questions that take ten seconds each to ask.

The other eight

#QuestionA good answer sounds like
5Is the same question answered the same way twice?Mostly, and here is where it moves
6Are the numbers checked with arithmetic?Yes, and it fails closed
7What does it refuse to do?A specific list, given without prompting
8Who is accountable when it is wrong?A named process, not a disclaimer
9Is it giving advice or information?Information, and it holds that line under pressure
10What is the failure mode: blank or wrong?Blank, always
11How is a correction propagated?Restatements flow through and the page changes
12Can you export and check the workings?Yes, in one click

Question ten is worth dwelling on. Every system fails. The only choice a builder has is what failure looks like, and a system that fails to a blank is one you can work with, while a system that fails to a plausible wrong number is one that will eventually cost you something. When you ask this, watch whether the answer arrives immediately. It is the question a team who have thought about it can answer in a sentence and a team who have not cannot answer at all.

Figure

Why questions 2 and 10 are the same question

79% answered
21% declined
Questions answered
79%
Right when it answers
94.0%
Confidently wrong per 100
4.7

The bar is abstention. Push it up and answered volume falls while correctness rises. A vendor who has never made this trade explicitly has made it implicitly, at whatever setting produced the best demo.

Figure

A trustworthy tool against a demo-grade one, on the four questions

Demo gradeTrustworthy
Provenance3 to 9
Abstention1 to 8
Universe honesty4 to 9
Freshness5 to 9

Scored out of ten on the four heaviest questions. The shapes are the tell: a good tool is even, a demo-grade one is spiky, brilliant at the thing that shows well and hollow at the thing that only matters later.

Before running it on anyone, it is worth knowing what you are asking them to afford. Coverage is a budget, and the cost of a wide universe explains most of the answers you will get to question three.

Scoring ourselves, since that is the test

Eight of twelve, and the four we fail are worth naming. We fail question two: our agent hedges rather than refusing. We partly fail five: the edges of a screen move between runs. We fail eight in the sense that our accountability is a process rather than a published one. And we fail six for anything beyond headline figures, where arithmetic verification is not yet universal.

We pass the ones we would weight most heavily. Every figure carries its source, the failure mode is a blank, the universe is stated, and restatements flow through to the page. The full self-assessment, scored capability by capability, is in our honest scorecard, including the things we are building to fix the four above.

What this cannot tell you

Whether a tool is right for you, which depends on what you were going to do with it. A tool that is excellent at reading filings and useless at ideas is exactly right for someone who has ideas and no time, and wrong for someone who wanted the ideas. Work out which one you are before evaluating anything, because the checklist above measures quality and not fit.

The bottom line

Twelve questions, four that matter, and one that is really the test: ask something the tool cannot know and watch what it does. Everything else about the system is visible in that one response. Run the questions against whatever you are using now with the readiness score, which turns them into a number you can compare, and then run them against us.


Educational information, not financial advice. Where this page states a figure about our own coverage it is read from the database as the page loads, so it is today's number rather than the day this was written.

Get the next piece in your inbox

One well-researched article at a time. No spam, unsubscribe in one click.

No spam, no selling your address, unsubscribe in one click. The tools stay free either way.

Keep exploring: browse the stocks we cover or see what the super investors hold.