FinanceReasoning
Multi-step quantitative questions with financial context and numerical answers.
AIMultiple / September 2026 / published research
Evaluation setup
AIMultiple evaluated the 238-question hard subset of FinanceReasoning using API calls, a common prompt, model-based answer extraction, and a 0.2% numerical tolerance. These are publisher-reported results, not Finsider runs. Reasoning settings and provider snapshots should be checked in the source before reproducing the evaluation.
Reading this comparison
Ranks compare displayed scores only within this named evaluation and selected model roster. Ties share a rank. Different benchmarks are not combined into an overall ranking. A model omitted for lack of a published result is not assigned a zero. A newer release is not automatically better on every task.
Boundaries
A benchmark answer score does not establish a model's readiness to deliver a financial diligence engagement. It does not measure CPA sign-off, evidence retention, access controls, or performance on your transaction documents.
Provenance
Source snapshot compiled September 14, 2026. Evaluations were performed by the named publishers, not Finsider. CSV downloads record model versions, configurations, sources, and any arithmetic derivation. No comparable result for Finsider's deterministic math engine is available in these sources, so the engine is not ranked.
Performance Comparison
Accuracy (%)
GPT-5.6 Sol Pro
90.76
Claude Fable 5.1
90.34
GPT-5.6 Sol
90.34
Claude Opus 5
89.92
Kimi K3
88.24
Grok 4.5
88.24
GPT-6 Astra
87.39
Claude Sonnet 5
86.97
Gemini 3.5 Flash
86.97
GLM-5.2
86.13
Selected frontier models from the same 238-question evaluation. Selection includes the requested Astra, Fable 5.1, and Opus 5 releases; it is not the full publisher roster. Astra's 87.39 is from the article's financial-reasoning chart data, not its unrelated general-model index.
Rank uses displayed scores; ties share rank. Bars use a 0-100 scale. No confidence intervals are available in the cited results; small differences should not be interpreted as statistically established superiority.
Source: AIMultiple / September 2026