Finsider Labs
All leaderboards

Analytical finance

A financial-services evaluator's reported comparison of analytical tasks.

Model ML / September 2026 / published research

Evaluation setup

Model ML reports results from its proprietary financial-services evaluation. The article does not publish the complete task set, sample sizes, prompts, or model reasoning settings. Values below are the figures available in the report, not an independently reproduced benchmark.

Reading this comparison

Ranks compare displayed scores only within this named evaluation and selected model roster. Ties share a rank. Different benchmarks are not combined into an overall ranking. A model omitted for lack of a published result is not assigned a zero. A newer release is not automatically better on every task.

Boundaries

A benchmark answer score does not establish a model's readiness to deliver a financial diligence engagement. It does not measure CPA sign-off, evidence retention, access controls, or performance on your transaction documents.

Provenance

Source snapshot compiled September 14, 2026. Evaluations were performed by the named publishers, not Finsider. CSV downloads record model versions, configurations, sources, and any arithmetic derivation. No comparable result for Finsider's deterministic math engine is available in these sources, so the engine is not ranked.

Performance Comparison

Reported score (0-100)

1

GPT-6 Astra

93.21

2

Claude Fable 5.1

92.54

3

Gemini 3.8 Flash

92.39

050100

The three model scores explicitly disclosed for this category. Other model scores are not supplied here; an omitted model is not a zero or a failed run.

Rank uses displayed scores; ties share rank. Bars use a 0-100 scale. No confidence intervals are available in the cited results; small differences should not be interpreted as statistically established superiority.

Source: Model ML / September 2026