Finsider Labs

[FRONTIER LEADERBOARDS]

Financial intelligence benchmarks

Published financial AI comparisons, presented in the original ranked-chart format. Each view retains its own study, metric, model versions, and evaluation conditions.

Published model evaluations, reviewed September 14, 2026. Protocols and model versions remain separate; no Finsider product-performance claim or cross-study aggregate score.

AIMultiple / September 2026

FinanceReasoning

Financial numerical reasoning

1

GPT-5.6 Sol Pro

90.76

2

Claude Fable 5.1

90.34

2

GPT-5.6 Sol

90.34

4

Claude Opus 5

89.92

5

Kimi K3

88.24

050100

Accuracy (%) / top five in this selection

Full comparison

Model ML / September 2026

Analytical finance

Financial calculations

1

GPT-6 Astra

93.21

2

Claude Fable 5.1

92.54

3

Gemini 3.8 Flash

92.39

050100

Reported score (0-100) / published comparison

Full comparison

Model ML / September 2026

Financial workflows

Multi-step financial work

1

Gemini 3.8 Flash

76.33

2

GPT-6 Astra

72.36

3

Claude Fable 5.1

60.00

050100

Reported score (0-100) / published comparison

Full comparison

Model ML / September 2026

Single-document analysis

Financial document retrieval

1

GPT-6 Astra

86.30

2

GPT-5.6 Sol

82.96

3

Gemini 3.8 Flash

81.48

050100

Reported score (0-100) / published comparison

Full comparison

Model ML / September 2026

Multi-document analysis

Cross-document financial evidence

1

GPT-6 Astra

78.71

2

GPT-5.6 Sol

76.77

3

Gemini 3.8 Flash

63.23

050100

Reported score (0-100) / published comparison

Full comparison

AIMultiple / September 2026

Open-weight financial reasoning

FinanceReasoning / open-weight selection

1

Kimi K3

88.24

2

GLM-5.2

86.13

3

gpt-oss-120b

81.09

4

Qwen3-235B-A22B-Thinking-2507

75.63

5

Llama 4 Maverick

75.21

050100

Accuracy (%) / top five in this selection

Full comparison

Finsider evaluation protocols

Finsider deterministic math engine: comparative evaluation pending. These diligence-specific protocols are research plans, not measured scores, and remain separate from the published comparisons above.

Earnings quality evaluation protocolNot yet evaluated Cash proof evaluation protocolNot yet evaluated Working capital evaluation protocolNot yet evaluated Source lineage evaluation protocolNot yet evaluated Document extraction evaluation protocolNot yet evaluated Review readiness evaluation protocolNot yet evaluated