AIMultiple / September 2026
FinanceReasoning
Financial numerical reasoning
GPT-5.6 Sol Pro
90.76
Claude Fable 5.1
90.34
GPT-5.6 Sol
90.34
Claude Opus 5
89.92
Kimi K3
88.24
Accuracy (%) / top five in this selection
Full comparisonFinsider Labs is our applied AI research lab for financial intelligence. We explore source lineage, reproducible analysis, financial screening, and the role of professional judgment in AI-powered diligence.
Published model comparisons for financial reasoning and document understanding. Historical source results, not Finsider test scores.
AIMultiple / September 2026
Financial numerical reasoning
GPT-5.6 Sol Pro
90.76
Claude Fable 5.1
90.34
GPT-5.6 Sol
90.34
Claude Opus 5
89.92
Kimi K3
88.24
Accuracy (%) / top five in this selection
Full comparisonModel ML / September 2026
Financial calculations
GPT-6 Astra
93.21
Claude Fable 5.1
92.54
Gemini 3.8 Flash
92.39
Reported score (0-100) / published comparison
Full comparisonModel ML / September 2026
Multi-step financial work
Gemini 3.8 Flash
76.33
GPT-6 Astra
72.36
Claude Fable 5.1
60.00
Reported score (0-100) / published comparison
Full comparisonModel ML / September 2026
Financial document retrieval
GPT-6 Astra
86.30
GPT-5.6 Sol
82.96
Gemini 3.8 Flash
81.48
Reported score (0-100) / published comparison
Full comparisonModel ML / September 2026
Cross-document financial evidence
GPT-6 Astra
78.71
GPT-5.6 Sol
76.77
Gemini 3.8 Flash
63.23
Reported score (0-100) / published comparison
Full comparisonAIMultiple / September 2026
FinanceReasoning / open-weight selection
Kimi K3
88.24
GLM-5.2
86.13
gpt-oss-120b
81.09
Qwen3-235B-A22B-Thinking-2507
75.63
Llama 4 Maverick
75.21
Accuracy (%) / top five in this selection
Full comparisonOriginal Finsider methodology drafts on financial screening, traceable evidence, and reviewer-led M&A diligence. No benchmark results or peer review are implied.
An evidence-first framework for pre-QoE screening
Practical perspectives on financial diligence from Finsider Labs. Local editorial drafts for review.
A first-pass financial screen should identify the next diligence questions, not turn incomplete records into definitive conclusions.
A financial finding becomes useful when another reviewer can reconstruct its source, calculation, assumptions, and decision.
The transition from a screen to QoE begins with scope, evidence requirements, and professional judgment, not a change in the report title.
Cash and earnings tell related but different stories. A useful cash proof makes timing, coverage, and reconciliation differences explicit.
View allAll posts