Do AI Investing Agents Actually Work? The Evidence, Honestly

27 July 20265 min readAI AgentsEvidenceExplainers
The short answer. On narrow, checkable tasks, yes: extracting figures from filings, summarising ownership changes, and flagging accounting distortions are all things current systems do well enough to be useful. On the task people actually want, picking investments that beat a market, there is no credible public evidence that any AI agent does this reliably, and the studies claiming it almost all contain a look-ahead or survivorship flaw. As of 27 July 2026, treat the first category as real and the second as unproven.

This is the question the category deserves and mostly avoids. It has two answers because it is two questions.

The tasks where the evidence is decent

TaskHow well it worksWhy it works
Pulling a stated figure out of a filingWellThe answer is in the document; the job is locating it
Summarising what changed between two filingsWellA comparison over structured text
Flagging a one-off distorting a ratioReasonablyThe pattern is visible in the cash flow statement
Explaining a term or a mechanismWellIt is a well documented thing, not a prediction
Ranking companies by a stated criterionReasonablyAs long as the criterion is actually stated
Judging whether a business is durablePoorlyRequires a view of the future
Predicting returnsNo credible evidenceRequires a view of the future, and everyone else has one too

Notice where the line falls. It is not a line between easy and hard. It is a line between questions whose answers already exist somewhere and questions whose answers do not exist yet. Everything on the good side of it is retrieval, comparison and pattern matching over documents. Everything on the bad side requires knowing something the market does not, which is a far higher bar than it sounds and is the same bar an economic moat has to clear.

Figure

Reliability by task type, roughly

Locate a stated fact
93%
Compare two filings
86%
Spot a one-off
71%
Judge durability
41%
Predict returns
0%

Our own read of where current systems sit, ours included. The scale is deliberately coarse because a precise number here would be a false one: nobody in this industry, including us, publishes an evaluation set you could check.

Figure

Benchmark scores against anything you could trade

Two years agoNow
Filing comprehension54 to 88
Figure extraction61 to 93
Multi-step reasoning29 to 67
Public forward records0 to 0

The category's benchmark numbers have improved fast and genuinely. What they measure is comprehension of documents, which is not the same thing as an edge, and the gap between the two lines has not closed at all.

Why the impressive studies are usually broken

Three failure modes account for almost all of them, and once you know the names you will see them everywhere.

Look-ahead. The model was trained on text that includes the outcome. Ask it in 2026 to pick stocks "as of 2019" and it knows what happened in 2020. This is not subtle and it invalidates most backtests of a language model on historical data. The only clean test is forward, in public, from a date stamped before the model saw anything.

Survivorship. The universe is drawn from companies that still exist. Every strategy looks good on a list of things that did not go bankrupt.

Multiple comparisons. Run forty variants, publish the one that worked. This is the oldest error in quantitative finance and adding a language model to it does not fix it.

Figure

What is left of a typical headline result

+14%Claimed edge-7%Look-ahead-3%Survivorship-3%Variant pic…-1.5%Costs

Illustrative, and unkind on purpose. The point is not the exact numbers, it is that the three standard corrections are each large, and almost no published result applies all three.

Figure

How to read any result in this category, in four steps

  1. 1. Was it forward?+1

    Dated before the model saw anything, or it proves nothing about prediction.

  2. 2. Was the universe fixed first?+2

    Otherwise you are reading a list of companies that survived.

  3. 3. How many variants were run?+3

    One published result from forty attempts is a lottery ticket held up as a method.

  4. 4. Net of costs?+4

    Spreads and fees eat the size of edge most of these papers claim to have found.

Apply these in order and most published results do not survive step two. That is not a reason to dismiss the field; it is a reason to weight forward evidence enormously more than backward evidence.

The thing that would change our mind

A forward, public, dated, unedited track record over at least a couple of years, on a universe fixed in advance, net of costs, with every position published on the day it was taken. Nobody in this category has one yet. When somebody does, it will be worth far more than any benchmark score, and you should ask us for ours too.

Figure

Why a real edge can still ruin you

100% survive
Chance of ruin
0%
Average ending bank
£NaN

Set the win rate above fifty and the edge is genuine. Now raise how much is risked on each trade. A positive expectation and a large bet size still produce ruin surprisingly often, which is why edge alone was never the interesting question.

What this cannot tell you

It cannot tell you what the systems will do in two years. The retrieval side of this has improved quickly and there is no obvious reason it stops. The prediction side runs into a wall that is not about model quality: markets price in what is knowable, and a widely available tool that finds an edge removes it. That is not pessimism about the technology, it is arithmetic about competition.

The bottom line

Buy the reading, not the picking. An agent that reads three hundred filings this week and tells you which four contain something odd is doing work you could not do and cannot easily check elsewhere, and it is honest about what it is. An agent that hands you a portfolio is making a claim nobody in this industry has yet earned, and you can check the same instinct against what the professionals actually filed last quarter rather than what anyone says about them. If you want to see what the reading looks like, our company pages show every figure with the filing row it came from, which is the checkable version.


Educational information, not financial advice. Where this page states a figure about our own coverage it is read from the database as the page loads, so it is today's number rather than the day this was written.

Get the next piece in your inbox

One well-researched article at a time. No spam, unsubscribe in one click.

No spam, no selling your address, unsubscribe in one click. The tools stay free either way.

Keep exploring: browse the stocks we cover or see what the super investors hold.