Finsider Labs
All leaderboards

FinanceReasoning

Multi-step quantitative questions with financial context and numerical answers.

AIMultiple / September 2026 / published research

Evaluation setup

AIMultiple evaluated the 238-question hard subset of FinanceReasoning using API calls, a common prompt, model-based answer extraction, and a 0.2% numerical tolerance. These are publisher-reported results, not Finsider runs. Reasoning settings and provider snapshots should be checked in the source before reproducing the evaluation.

Reading this comparison

Ranks compare displayed scores only within this named evaluation and selected model roster. Ties share a rank. Different benchmarks are not combined into an overall ranking. A model omitted for lack of a published result is not assigned a zero. A newer release is not automatically better on every task.

Boundaries

A benchmark answer score does not establish a model's readiness to deliver a financial diligence engagement. It does not measure CPA sign-off, evidence retention, access controls, or performance on your transaction documents.

Provenance

Source snapshot compiled September 14, 2026. Evaluations were performed by the named publishers, not Finsider. CSV downloads record model versions, configurations, sources, and any arithmetic derivation. No comparable result for Finsider's deterministic math engine is available in these sources, so the engine is not ranked.

Performance Comparison

Accuracy (%)

1

GPT-5.6 Sol Pro

90.76

2

Claude Fable 5.1

90.34

2

GPT-5.6 Sol

90.34

4

Claude Opus 5

89.92

5

Kimi K3

88.24

5

Grok 4.5

88.24

7

GPT-6 Astra

87.39

8

Claude Sonnet 5

86.97

8

Gemini 3.5 Flash

86.97

10

GLM-5.2

86.13

050100

Selected frontier models from the same 238-question evaluation. Selection includes the requested Astra, Fable 5.1, and Opus 5 releases; it is not the full publisher roster. Astra's 87.39 is from the article's financial-reasoning chart data, not its unrelated general-model index.

Rank uses displayed scores; ties share rank. Bars use a 0-100 scale. No confidence intervals are available in the cited results; small differences should not be interpreted as statistically established superiority.

Source: AIMultiple / September 2026