How to Evaluate an AI Financial Agent: 12 Questions
The short answer. Ask twelve questions, and the four that predict everything else are: where did this figure come from, what happens when you do not know, what is your universe, and how old is this data. A tool that answers all four crisply is usually good. A tool that deflects on any of them is telling you something, and the something is rarely favourable.
This works on us. That is the test of whether a checklist is honest, and we score ourselves against it publicly further down.
The four that matter most
1. Where did this number come from? The answer should be a specific document and a specific line, available without asking twice. "Our data providers" is not an answer. If a figure cannot be traced, treat it as a guess with good posture.
2. What happens when you do not know? Ask something deliberately obscure: a segment breakdown for a small foreign listing, a metric that is not reported. You want "not established". You will usually get a fluent paragraph, and a fluent paragraph in response to a question with no answer tells you exactly how the system behaves when it is out of its depth on a question you could not check.
3. What is your universe, and how did you choose it? Coverage is the budget in disguise, and the arithmetic behind it is not complicated once someone shows you the multiplication. A tool covering four hundred large caps has made a defensible decision; a tool implying it covers everything while being excellent on ten names has made a different one.
4. How old is this? Every figure should carry an as-of date. Financial data is restated, and a number that was right in February can be wrong in May through no fault of anyone's.
How much each question predicts overall quality
Our own weighting from evaluating competitors and ourselves. Provenance and abstention dominate: a tool that gets those two right is almost never bad at the rest, and a tool that fluffs them is almost never good at anything.
How a typical tool loses its score
Starting from a hundred and subtracting what a median product in this category actually fails. Two deductions account for most of it, and they are the two questions that take ten seconds each to ask.
The other eight
| # | Question | A good answer sounds like |
|---|---|---|
| 5 | Is the same question answered the same way twice? | Mostly, and here is where it moves |
| 6 | Are the numbers checked with arithmetic? | Yes, and it fails closed |
| 7 | What does it refuse to do? | A specific list, given without prompting |
| 8 | Who is accountable when it is wrong? | A named process, not a disclaimer |
| 9 | Is it giving advice or information? | Information, and it holds that line under pressure |
| 10 | What is the failure mode: blank or wrong? | Blank, always |
| 11 | How is a correction propagated? | Restatements flow through and the page changes |
| 12 | Can you export and check the workings? | Yes, in one click |
Question ten is worth dwelling on. Every system fails. The only choice a builder has is what failure looks like, and a system that fails to a blank is one you can work with, while a system that fails to a plausible wrong number is one that will eventually cost you something. When you ask this, watch whether the answer arrives immediately. It is the question a team who have thought about it can answer in a sentence and a team who have not cannot answer at all.
Why questions 2 and 10 are the same question
The bar is abstention. Push it up and answered volume falls while correctness rises. A vendor who has never made this trade explicitly has made it implicitly, at whatever setting produced the best demo.
A trustworthy tool against a demo-grade one, on the four questions
Scored out of ten on the four heaviest questions. The shapes are the tell: a good tool is even, a demo-grade one is spiky, brilliant at the thing that shows well and hollow at the thing that only matters later.
Before running it on anyone, it is worth knowing what you are asking them to afford. Coverage is a budget, and the cost of a wide universe explains most of the answers you will get to question three.
Scoring ourselves, since that is the test
Eight of twelve, and the four we fail are worth naming. We fail question two: our agent hedges rather than refusing. We partly fail five: the edges of a screen move between runs. We fail eight in the sense that our accountability is a process rather than a published one. And we fail six for anything beyond headline figures, where arithmetic verification is not yet universal.
We pass the ones we would weight most heavily. Every figure carries its source, the failure mode is a blank, the universe is stated, and restatements flow through to the page. The full self-assessment, scored capability by capability, is in our honest scorecard, including the things we are building to fix the four above.
What this cannot tell you
Whether a tool is right for you, which depends on what you were going to do with it. A tool that is excellent at reading filings and useless at ideas is exactly right for someone who has ideas and no time, and wrong for someone who wanted the ideas. Work out which one you are before evaluating anything, because the checklist above measures quality and not fit.
The bottom line
Twelve questions, four that matter, and one that is really the test: ask something the tool cannot know and watch what it does. Everything else about the system is visible in that one response. Run the questions against whatever you are using now with the readiness score, which turns them into a number you can compare, and then run them against us.
Educational information, not financial advice. Where this page states a figure about our own coverage it is read from the database as the page loads, so it is today's number rather than the day this was written.
One well-researched article at a time. No spam, unsubscribe in one click.
No spam, no selling your address, unsubscribe in one click. The tools stay free either way.
Keep exploring: browse the stocks we cover or see what the super investors hold.
