Our AI Agents, Scored Honestly: What Works and What Does Not
The short answer. Our agents are good at retrieval and grounding, adequate at multi-step screening, and weak at abstention and at replanning. Specifically: every figure they state is fetched and carries its source, which is the thing we would defend; they will not currently tell you they are unsure, they will hedge instead, which is worse; and coverage quality falls off outside large US listings. The next generation, in build now, targets the abstention and the replanning first.
This is the page a competitor would write about us. We would rather write it ourselves, accurately, than have the version with the numbers guessed.
A caveat before the scores, because it cuts against us: these are our own assessments against our own criteria. We do not yet publish an evaluation set you could run, which means everything below is a good faith self-report rather than a measurement. Building that set is on the roadmap and it is the single thing that would make this page trustworthy rather than merely candid.
The scorecard
| Capability | Score | What that means in practice |
|---|---|---|
| Grounding: every figure is fetched, not recalled | 9 / 10 | The failure mode is a blank, not a wrong number |
| Provenance: the source is shown by default | 9 / 10 | You can check any figure in two clicks |
| Retrieval accuracy on large US listings | 8 / 10 | Genuinely reliable |
| Retrieval accuracy elsewhere | 5 / 10 | Thinner filings coverage, and the agent inherits it |
| Multi-step screening | 6 / 10 | Runs a mostly fixed sequence, adapts inside it |
| Catching accounting distortions | 6 / 10 | Good at one-off inflows, poor at rescues |
| Abstention: saying it does not know | 3 / 10 | Hedges instead. Our worst score and our biggest risk |
| Replanning: choosing the next source | 3 / 10 | Largely fixed. In build |
| Reproducibility across runs | 5 / 10 | Edges of a list move between runs |
| Latency on a cold company | 6 / 10 | Seconds, not instant, because it is reading |
Our agents, scored against the ten things that matter
Self-assessed, which is a real weakness of this chart and is stated above. The shape is the point: strong where the architecture does the work, weak where judgement does.
Where we are against where the next generation has to be
Out of ten. The two bars that matter are the bottom two, and the gap in them is the entire reason the roadmap below is ordered the way it is rather than by what would demo best.
The three failures worth naming
It hedges instead of refusing. Ask about something we do not hold and you are more likely to get a carefully qualified paragraph than "we do not have this". The paragraph is worse, because a reader in a hurry reads a qualified answer as an answer. This is the failure we would most like you to catch us on, and it is first in the queue.
Non-US coverage is a different product. A FTSE mid-cap or a European industrial gets a noticeably thinner read than a US large cap, because the filings pipeline behind it is thinner. The agent does not currently say so. It should, and making it say so is a smaller job than fixing the underlying coverage, which is why it comes first.
Two runs, two lists. Ask for the same screen twice and the middle of the list is stable while the edges move. We can make it deterministic by pinning it, at the cost of the reading step that made the agent worth using. We have not solved this and neither has anyone else; we are stating it because a screen you cannot reproduce is a screen you should not build a position from without looking.
Why we cannot simply put a human on everything
Set the error rate to the six per cent we estimate for our screening step, then drag the review coverage. The hours column is the constraint. This is the calculation behind our decision to invest in abstention rather than in review, and you can disagree with it here.
The setting we are currently wrong about
Our agent sits too far to the left: it answers almost everything, which keeps the answered column high and quietly keeps the third column high too. The build moves the bar right and accepts the drop in coverage, because a confident wrong answer costs more than ten missing ones.
What the next generation actually changes
Not "better". Four specific things, in this order.
The build, in order, with what each one is for
- 1. Calibrated refusal+1
A real not-established path, and a visible confidence on every claim, so a hedge stops being the cheapest output.
- 2. Coverage honesty+2
The agent states the depth of the data behind an answer before the answer, so a thin read announces itself.
- 3. Replanning+3
The next source is chosen from what the last one said. This is what makes the rescue case work.
- 4. A published eval set+4
So the scores at the top of this page stop being self-reported. The one that makes the rest checkable.
Ordered by how much wrongness each removes per unit of work, not by how good each sounds. Abstention is first because a confident wrong answer costs more than ten missing ones.
Right now we currently hold 7,992 live companies, of which 3,065 carry a rating and 2,871 carry a discounted cash flow estimate, and the deterministic layer behind all of it runs daily whatever the agents are doing. That separation is deliberate: if the agent work fell over tomorrow, the data, the fair values and the filings would all still be right.
What this cannot tell you
Whether we will ship these, and when. We are not putting dates on this page, because a date we miss is worse than no date and this is a small team building in the open. What we will do is keep this page current: a score that moves gets moved here, and a thing that ships stops saying in build.
The bottom line
Every company in this category is currently claiming the same six capabilities, and roughly two of them are real at any given vendor. Ours are grounding and provenance; our abstention is bad and we are fixing it before we build anything more impressive. The reasoning behind that ordering is set out in what an AI financial agent actually is. If you want to hold us to this, the twelve question checklist is what we score ourselves against, and it works just as well pointed at anyone else.
Educational information, not financial advice. Where this page states a figure about our own coverage it is read from the database as the page loads, so it is today's number rather than the day this was written.
One well-researched article at a time. No spam, unsubscribe in one click.
No spam, no selling your address, unsubscribe in one click. The tools stay free either way.
Keep exploring: browse the stocks we cover or see what the super investors hold.
