Our AI Agents, Scored Honestly: What Works and What Does Not

27 July 20265 min readAI AgentsProductTransparency
The short answer. Our agents are good at retrieval and grounding, adequate at multi-step screening, and weak at abstention and at replanning. Specifically: every figure they state is fetched and carries its source, which is the thing we would defend; they will not currently tell you they are unsure, they will hedge instead, which is worse; and coverage quality falls off outside large US listings. The next generation, in build now, targets the abstention and the replanning first.

This is the page a competitor would write about us. We would rather write it ourselves, accurately, than have the version with the numbers guessed.

A caveat before the scores, because it cuts against us: these are our own assessments against our own criteria. We do not yet publish an evaluation set you could run, which means everything below is a good faith self-report rather than a measurement. Building that set is on the roadmap and it is the single thing that would make this page trustworthy rather than merely candid.

The scorecard

CapabilityScoreWhat that means in practice
Grounding: every figure is fetched, not recalled9 / 10The failure mode is a blank, not a wrong number
Provenance: the source is shown by default9 / 10You can check any figure in two clicks
Retrieval accuracy on large US listings8 / 10Genuinely reliable
Retrieval accuracy elsewhere5 / 10Thinner filings coverage, and the agent inherits it
Multi-step screening6 / 10Runs a mostly fixed sequence, adapts inside it
Catching accounting distortions6 / 10Good at one-off inflows, poor at rescues
Abstention: saying it does not know3 / 10Hedges instead. Our worst score and our biggest risk
Replanning: choosing the next source3 / 10Largely fixed. In build
Reproducibility across runs5 / 10Edges of a list move between runs
Latency on a cold company6 / 10Seconds, not instant, because it is reading
Figure

Our agents, scored against the ten things that matter

Grounding
9
Provenance
9
Retrieval, US
8
Screening
6
Distortions
6
Retrieval, non-US
5
Reproducibility
5
Abstention
3
Replanning
3

Self-assessed, which is a real weakness of this chart and is stated above. The shape is the point: strong where the architecture does the work, weak where judgement does.

Figure

Where we are against where the next generation has to be

TodayTarget
Grounding9 to 10
Retrieval, non-US5 to 8
Screening6 to 8
Abstention3 to 9
Replanning3 to 8

Out of ten. The two bars that matter are the bottom two, and the gap in them is the entire reason the roadmap below is ordered the way it is rather than by what would demo best.

The three failures worth naming

It hedges instead of refusing. Ask about something we do not hold and you are more likely to get a carefully qualified paragraph than "we do not have this". The paragraph is worse, because a reader in a hurry reads a qualified answer as an answer. This is the failure we would most like you to catch us on, and it is first in the queue.

Non-US coverage is a different product. A FTSE mid-cap or a European industrial gets a noticeably thinner read than a US large cap, because the filings pipeline behind it is thinner. The agent does not currently say so. It should, and making it say so is a smaller job than fixing the underlying coverage, which is why it comes first.

Two runs, two lists. Ask for the same screen twice and the middle of the list is stable while the edges move. We can make it deterministic by pinning it, at the cost of the reading step that made the agent worth using. We have not solved this and neither has anyone else; we are stating it because a screen you cannot reproduce is a screen you should not build a position from without looking.

Figure

Why we cannot simply put a human on everything

Errors caught
13
Errors that ship
67
Wrong per 1,000 shipped
67
Human hours per 1,000
6.7h
Hours per error caught
0.52h

Set the error rate to the six per cent we estimate for our screening step, then drag the review coverage. The hours column is the constraint. This is the calculation behind our decision to invest in abstention rather than in review, and you can disagree with it here.

Figure

The setting we are currently wrong about

79% answered
21% declined
Questions answered
79%
Right when it answers
94.0%
Confidently wrong per 100
4.7

Our agent sits too far to the left: it answers almost everything, which keeps the answered column high and quietly keeps the third column high too. The build moves the bar right and accepts the drop in coverage, because a confident wrong answer costs more than ten missing ones.

What the next generation actually changes

Not "better". Four specific things, in this order.

Figure

The build, in order, with what each one is for

  1. 1. Calibrated refusal+1

    A real not-established path, and a visible confidence on every claim, so a hedge stops being the cheapest output.

  2. 2. Coverage honesty+2

    The agent states the depth of the data behind an answer before the answer, so a thin read announces itself.

  3. 3. Replanning+3

    The next source is chosen from what the last one said. This is what makes the rescue case work.

  4. 4. A published eval set+4

    So the scores at the top of this page stop being self-reported. The one that makes the rest checkable.

Ordered by how much wrongness each removes per unit of work, not by how good each sounds. Abstention is first because a confident wrong answer costs more than ten missing ones.

Right now we currently hold 7,992 live companies, of which 3,065 carry a rating and 2,871 carry a discounted cash flow estimate, and the deterministic layer behind all of it runs daily whatever the agents are doing. That separation is deliberate: if the agent work fell over tomorrow, the data, the fair values and the filings would all still be right.

What this cannot tell you

Whether we will ship these, and when. We are not putting dates on this page, because a date we miss is worse than no date and this is a small team building in the open. What we will do is keep this page current: a score that moves gets moved here, and a thing that ships stops saying in build.

The bottom line

Every company in this category is currently claiming the same six capabilities, and roughly two of them are real at any given vendor. Ours are grounding and provenance; our abstention is bad and we are fixing it before we build anything more impressive. The reasoning behind that ordering is set out in what an AI financial agent actually is. If you want to hold us to this, the twelve question checklist is what we score ourselves against, and it works just as well pointed at anyone else.


Educational information, not financial advice. Where this page states a figure about our own coverage it is read from the database as the page loads, so it is today's number rather than the day this was written.

Get the next piece in your inbox

One well-researched article at a time. No spam, unsubscribe in one click.

No spam, no selling your address, unsubscribe in one click. The tools stay free either way.

Keep exploring: browse the stocks we cover or see what the super investors hold.