AIMultiple / September 2026
FinanceReasoning
Financial numerical reasoning
GPT-5.6 Sol Pro
90.76
Claude Fable 5.1
90.34
GPT-5.6 Sol
90.34
Claude Opus 5
89.92
Kimi K3
88.24
Accuracy (%) / top five in this selection
Full comparison[FRONTIER LEADERBOARDS]
Published financial AI comparisons, presented in the original ranked-chart format. Each view retains its own study, metric, model versions, and evaluation conditions.
Published model evaluations, reviewed September 14, 2026. Protocols and model versions remain separate; no Finsider product-performance claim or cross-study aggregate score.
AIMultiple / September 2026
Financial numerical reasoning
GPT-5.6 Sol Pro
90.76
Claude Fable 5.1
90.34
GPT-5.6 Sol
90.34
Claude Opus 5
89.92
Kimi K3
88.24
Accuracy (%) / top five in this selection
Full comparisonModel ML / September 2026
Financial calculations
GPT-6 Astra
93.21
Claude Fable 5.1
92.54
Gemini 3.8 Flash
92.39
Reported score (0-100) / published comparison
Full comparisonModel ML / September 2026
Multi-step financial work
Gemini 3.8 Flash
76.33
GPT-6 Astra
72.36
Claude Fable 5.1
60.00
Reported score (0-100) / published comparison
Full comparisonModel ML / September 2026
Financial document retrieval
GPT-6 Astra
86.30
GPT-5.6 Sol
82.96
Gemini 3.8 Flash
81.48
Reported score (0-100) / published comparison
Full comparisonModel ML / September 2026
Cross-document financial evidence
GPT-6 Astra
78.71
GPT-5.6 Sol
76.77
Gemini 3.8 Flash
63.23
Reported score (0-100) / published comparison
Full comparisonAIMultiple / September 2026
FinanceReasoning / open-weight selection
Kimi K3
88.24
GLM-5.2
86.13
gpt-oss-120b
81.09
Qwen3-235B-A22B-Thinking-2507
75.63
Llama 4 Maverick
75.21
Accuracy (%) / top five in this selection
Full comparisonFinsider deterministic math engine: comparative evaluation pending. These diligence-specific protocols are research plans, not measured scores, and remain separate from the published comparisons above.
Earnings quality evaluation protocolNot yet evaluated Cash proof evaluation protocolNot yet evaluated Working capital evaluation protocolNot yet evaluated Source lineage evaluation protocolNot yet evaluated Document extraction evaluation protocolNot yet evaluated Review readiness evaluation protocolNot yet evaluated