Free scorecard
AI agent readiness score
Two questions predict almost everything else: where did this figure come from, and what happens when it does not know. Answer eleven, get a number you can put next to another number, and score us with it too.
Impressive on the names you first tried. Test it on something obscure before you rely on the breadth it implies.
Biggest single gap: Can it show you where a specific figure came from? Worth 12 points on its own.
Can it show you where a specific figure came from?
Good: Names a document and a line, in one or two clicks, without being asked twice.
Bad: “Our data providers”, a hover tooltip with a vendor name, or nothing at all.
What does it do when it does not know?
Good: Says the thing is not established, plainly, and stops.
Bad: A fluent hedged paragraph. This reads as an answer and is the more expensive failure.
Is the universe stated, and is it honest?
Good: A number, a scope, and an explanation of what was left out and why.
Bad: Implies everything, is excellent on ten names and vague on the eleventh.
Does every figure carry an as-of date?
Good: Dates on the figure itself, and restatements flow through to the page.
Bad: One “updated daily” line in a footer covering data of several different ages.
Does the same question give the same answer twice?
Good: Mostly, and they will tell you where it moves and why.
Bad: Never asked themselves, or claims perfect determinism from a system that reads.
Are computed numbers checked with arithmetic?
Good: Derived figures are recomputed from the retrieved inputs and fail closed.
Bad: A second model is asked whether the first looks right. That is correlation, not verification.
Is there a stated list of what it refuses to do?
Good: Offered without prompting, and it includes giving advice.
Bad: A disclaimer in the terms of service and nothing in the product.
Does it fail to a blank or to a wrong number?
Good: Blank. Answered in one sentence, immediately.
Bad: They have not thought about it, which the pause tells you.
Is there a named process when it is wrong?
Good: A correction path, a changelog, a human who owns it.
Bad: The disclaimer is the process.
Does it hold the line between information and advice?
Good: Holds it even when pushed hard in the chat.
Bad: Produces a recommendation if you phrase the question as a hypothetical.
Can you export the workings?
Good: One click, and the export contains the sources.
Bad: Screenshot the dashboard.
Score us with it. We publish our own answers, including the four we fail, in our honest scorecard. A checklist its own author will not run on themselves is a marketing asset, not a checklist.
Your score, the gaps, and the one fix worth the most points.
No spam, no selling your address, unsubscribe in one click. The tools stay free either way.
Straight answers
What is the single fastest test of an AI finance tool?
Ask it something it cannot know, such as a segment breakdown a company does not report. You want an explicit statement that this is not established. A fluent hedged paragraph in answer to an unanswerable question shows exactly how the tool behaves when out of its depth on a question you could not have checked.
Why are provenance and abstention weighted so heavily?
Because in practice they predict the rest. A tool that shows where every figure came from and will say it does not know is almost never bad at freshness, arithmetic or universe honesty, and a tool that fluffs those two is almost never good at anything else either.
Is a low score always disqualifying?
No. A tool that scores in the middle is often fine for narrowing a field, as long as you verify anything that goes into a decision. The scores that should stop you are the ones below about forty, where the tool will produce precise unsourced numbers and you will have no way to tell which are wrong.
How does SteadyShares score on its own checklist?
We pass the two heaviest, provenance and failure mode, and we fail abstention: our agent currently hedges rather than saying it does not know. The full self-assessment, scored capability by capability with the things we are building to fix it, is published in our honest scorecard.
Should an AI agent ever give investment advice?
No, and a good one holds that line even when the question is rephrased as a hypothetical. Information about a company is checkable. A recommendation from a system with no published track record is an opinion with no evidence behind it, and in most jurisdictions a regulated activity as well.
Educational information, not financial advice. Figures current as of July 2026 where dated; allowances and rates change, so check the source before acting.
