Documented test

We ran the same 10-K through four AI tools. Here is what they missed.

A repeatable test on a real annual report: identical questions, saved outputs, and every answer checked against the filing itself. The tools were fast and mostly fluent. Fluent is not the same as right.

A candlestick price chart with moving-average lines on a dark monitor
The test subject in this article is text, not price: a 200-page annual report and the tools asked to summarize it. Photo: Unsplash.
Scope note

This article is educational material about research method. It is not personalized investment advice, and nothing in it is a view on whether any company's stock is worth buying or selling.

Why a 10-K is the right stress test

If an AI tool is going to be useful in stock research, the annual report is where it has to prove itself. A 10-K is long, repetitive, legally precise, and full of sentences where one qualifier changes the meaning. It is also the document investors most often ask AI to compress, because reading it cover to cover takes hours.

That combination makes it a near-perfect test instrument. The correct answers are all in one public document. Errors have consequences that are easy to explain. And the failure modes that matter in practice (a figure attached to the wrong year, a softened risk, a confident invention) are exactly the ones a casual reader is least likely to catch, because the output reads so smoothly.

We chose the Apple Inc. annual report for fiscal year 2024, filed with the SEC on November 1, 2024. It is a well-known filing, freely available on SEC EDGAR, which means you can repeat everything below with the same source document and judge our claims yourself.

The setup, fixed in advance

Four tools were tested, anonymized here as A through D. Three are general-purpose chatbots from major AI labs; the fourth is a search-oriented assistant that emphasizes citations. They are anonymized for a practical reason: these products change weekly, and a named ranking would be stale within a month while the error patterns remain stable. What matters is the pattern, not the scoreboard.

Before opening any tool, we wrote down four questions and their correct answers, taken directly from the filing. Fixing answers in advance is the whole point of the method: it removes the temptation to grade fluent output generously.

The four test questions, with answers fixed from the filing before testing
QuestionCorrect answer (from the 10-K)What it tests
What were total net sales for fiscal 2024?$391.035 billion, fiscal year ended September 28, 2024Basic figure retrieval
How did iPhone and Services net sales compare?iPhone $201.183B; Services $96.169BSegment detail, two figures at once
Which fiscal period do these figures belong to?FY2024, with FY2023 and FY2022 shown for comparisonPeriod attribution, where errors hide
Summarize the most important risk factors.Item 1A, with the filing's own framing and qualifiersCompression of qualified language

Each tool received the same prompts in the same order, with the filing provided the same way. Tool name, access date, and full raw outputs were saved to our test library before any analysis began.

What the tools did well

Credit where due. Every tool located the correct section for every question, quickly. The headline figure, $391.035 billion in net sales, came back right from all four, usually with the surrounding table correctly understood. One tool reformatted the segment data into a cleaner comparison than the filing's own layout, which is a legitimate time save.

If your use of AI in research stops at "find where the filing talks about X" and "reformat this table," the tools are already dependable. The trouble starts one step further in.

Where they failed

Wrong fiscal period

Three of the four tools, at some point in the session, attached a figure to the wrong year. The typical shape of the error: asked to compare iPhone and Services, the tool answers with the correct FY2024 numbers and then, in a follow-up about trends, quietly mixes in an FY2023 figure while labeling it FY2024. The filing shows three fiscal years side by side; the models do not always respect which column they are reading from.

This is the most dangerous error type for an investor, because the figure itself is real. It appears in the filing. It is just not from the year the summary claims. Nothing about the sentence looks wrong.

The dropped qualifier

Asked for the most important risk factors, all four tools produced tidy lists. And all four, to different degrees, flattened the filing's careful language into something stronger. A risk the 10-K frames with conditions and context came back as a blunt statement of fact. The lawyers wrote those qualifiers because the qualifiers are where the meaning lives. Compression removed them.

The pattern, paraphrased: the filing says a risk applies "if" certain conditions occur and describes mitigating factors; the summary says the risk "is" a problem, full stop.

Observed across all four tools on the Item 1A summary task, test session July 2026.

Read only the summary and you come away with a risk picture that is simultaneously scarier and less precise than the real one.

Invented statements

Two tools, when pushed with follow-up questions the filing does not answer, produced confident statements that appear nowhere in the document. Not misreadings: inventions, phrased in the same flat declarative tone as the correct answers. This behavior, usually called hallucination, is well documented in academic work on language models; a broad survey of the phenomenon is maintained on arXiv. What our test adds is a concrete reminder that it shows up in ordinary, careful use, not only in adversarial prompts.

The defense is boring and effective: any claim that would influence a decision gets checked against the filing before it is allowed into your notes. Which leads to the practical part.

A five-minute verification pass

Here is the checkpoint routine we now run on any AI-generated filing summary before trusting it. It takes about five minutes and catches every error type described above.

  1. Pick the three numbers that matter most in the summary and find each one in the filing itself. Check the column header and the fiscal period label, not just the value.
  2. Read one risk factor in full, the original rather than the summary, and compare its qualifiers to what the AI wrote. If the summary is blunter, discount the whole risk section.
  3. Ask the tool one question the filing cannot answer (for example, something about events after the reporting date). A tool that answers it confidently has told you something important about how much to trust everything else it said.
  4. Note what the summary omitted. Ask "what is not in this summary that I would expect from a 10-K?" Omissions do not announce themselves.

Limits of this test

One filing, four questions, four tools, one point in time. This is a documented observation, not a benchmark study. Different documents, different prompts, or next month's model versions could produce different results, which is precisely why the five-minute pass matters more than any ranking. Tool access for this test used paid tiers we subscribe to ourselves; no vendor was involved or informed. Our working method describes how test files are kept, and readers can request the prompts used.

References

  1. Apple Inc., Form 10-K for fiscal year 2024, filed November 1, 2024. SEC EDGAR filing index. Accessed July 2026.
  2. Huang, L. et al., "A Survey on Hallucination in Large Language Models," arXiv:2311.05232. Accessed July 2026.
  3. U.S. Securities and Exchange Commission, investor education resources at Investor.gov. Accessed July 2026.

Further reading