Financial workflows
Reported outcomes for financial workflows in the Model ML Composite.
Model ML / September 2026 / published research
Evaluation setup
Model ML reports results from its proprietary financial-services evaluation. The article does not publish the complete task set, sample sizes, prompts, or model reasoning settings. Values below are the figures available in the report, not an independently reproduced benchmark.
Reading this comparison
Ranks compare displayed scores only within this named evaluation and selected model roster. Ties share a rank. Different benchmarks are not combined into an overall ranking. A model omitted for lack of a published result is not assigned a zero. A newer release is not automatically better on every task.
Boundaries
A benchmark answer score does not establish a model's readiness to deliver a financial diligence engagement. It does not measure CPA sign-off, evidence retention, access controls, or performance on your transaction documents.
Provenance
Source snapshot compiled September 14, 2026. Evaluations were performed by the named publishers, not Finsider. CSV downloads record model versions, configurations, sources, and any arithmetic derivation. No comparable result for Finsider's deterministic math engine is available in these sources, so the engine is not ranked.
Performance Comparison
Reported score (0-100)
Gemini 3.8 Flash
76.33
GPT-6 Astra
72.36
Claude Fable 5.1
60.00
Three models with values disclosed directly or recoverable from the report's stated differences. Astra: 76.33 - 3.97 = 72.36. Fable 5.1: 72.36 - 12.36 = 60.00. These two derived values retain the precision of the published differences.
Rank uses displayed scores; ties share rank. Bars use a 0-100 scale. No confidence intervals are available in the cited results; small differences should not be interpreted as statistically established superiority.
Source: Model ML / September 2026