A team of former Ernst & Young employees just released FAB (Finance Agents Benchmark), a public benchmark testing how well AI agents handle real financial due diligence work—using actual company data rooms and tasks drawn from professional practice.

The early results are telling: agents can retrieve relevant facts from complex document sets, but they consistently struggle to carry analysis through to complete, reliable conclusions. The gap between retrieval and reasoning is real, especially in high-stakes financial contexts where incomplete or uncertain answers create material risk.

For founders building AI tools for finance, legal, compliance, or any professional service, this benchmark is a reality check—and a roadmap.

What FAB Measures

FAB tests AI agents on tasks typical of financial due diligence: identifying revenue recognition issues, flagging related-party transactions, assessing liquidity, evaluating contractual obligations, and other work that combines document review, domain knowledge, and judgment.

The benchmark uses real company data rooms—collections of financial statements, contracts, board minutes, and other materials that mirror what analysts encounter in actual M&A or investment diligence. Tasks require agents to surface information, synthesize it across documents, and draw defensible conclusions.

The setup is deliberately challenging. Financial due diligence isn't a single-hop Q&A task; it's iterative, contextual, and unforgiving of errors. A missed contingent liability or misread covenant can sink a deal or trigger litigation.

Where Agents Succeed and Fail

Early results show agents perform well at information retrieval. Given a specific question—"What was the company's total revenue in 2023?"—agents can locate the relevant figure across multiple documents with reasonable accuracy.

But performance degrades sharply when tasks require synthesis or judgment. Agents struggle to connect facts across documents, interpret ambiguous language in contracts, or flag inconsistencies that suggest deeper issues. They surface pieces but don't reliably assemble them into coherent, complete analysis.

The failure mode isn't hallucination in the classic sense—it's incompleteness. Agents return partial answers, miss edge cases, or fail to flag uncertainty when data is ambiguous or missing. In professional services, that's as dangerous as being wrong.

What This Means for Founders Building AI Products

If you're building AI tools for finance, legal, or any domain where errors have consequences, FAB's findings should shape your product strategy and positioning:

Build for augmentation, not replacement

Products that treat AI as a junior analyst—surfacing facts, drafting summaries, flagging potential issues for review—will win adoption faster than those claiming to automate entire workflows. Experts don't want black-box answers; they want leverage on the tedious parts and confidence that nothing was missed.

Surface uncertainty explicitly

When your agent can't complete an analysis or encounters ambiguous data, flag it. Build workflows that escalate incomplete work to human review rather than returning a best-guess answer. Transparency about limitations builds trust; overconfidence destroys it.

Make outputs verifiable and traceable

Every claim your agent makes should link back to source documents with enough context for a human to verify it in seconds. Citation isn't a nice-to-have—it's the difference between a tool professionals will use and one they'll avoid.

Be honest with buyers and investors

If you're pitching enterprise customers or raising capital, don't oversell agent reliability in high-stakes contexts. Buyers in finance and legal have seen enough AI demos to be skeptical. A product that acknowledges its boundaries and demonstrates how it handles edge cases will differentiate faster than one making ambitious automation claims it can't yet deliver.

Key Takeaways

The opportunity for AI in professional services is real, but it's not automation—it's leverage. Products that help experts work faster and more thoroughly, while keeping them in control of judgment calls, will earn adoption and expand from there.

If you're building a product where reliability and speed both matter, you need an MVP that lets you test real workflows with real users before you scale. Get your MVP built in 3 days and start learning what works—without spending months on infrastructure that might not fit the problem.

Sources: https://github.com/SecondState-ai/finance-agents-benchmark