Open-weight financial reasoning
A selected open-weight view of the same FinanceReasoning evaluation, not a separate experiment.
AIMultiple / September 2026 / published research
Evaluation setup
AIMultiple evaluated the 238-question hard subset of FinanceReasoning using API calls, a common prompt, model-based answer extraction, and a 0.2% numerical tolerance. These are publisher-reported results, not Finsider runs. Reasoning settings and provider snapshots should be checked in the source before reproducing the evaluation.
Reading this comparison
Ranks compare displayed scores only within this named evaluation and selected model roster. Ties share a rank. Different benchmarks are not combined into an overall ranking. A model omitted for lack of a published result is not assigned a zero. A newer release is not automatically better on every task.
Boundaries
A benchmark answer score does not establish a model's readiness to deliver a financial diligence engagement. It does not measure CPA sign-off, evidence retention, access controls, or performance on your transaction documents.
Provenance
Source snapshot compiled September 14, 2026. Evaluations were performed by the named publishers, not Finsider. CSV downloads record model versions, configurations, sources, and any arithmetic derivation. No comparable result for Finsider's deterministic math engine is available in these sources, so the engine is not ranked.
Performance Comparison
Accuracy (%)
Kimi K3
88.24
GLM-5.2
86.13
gpt-oss-120b
81.09
Qwen3-235B-A22B-Thinking-2507
75.63
Llama 4 Maverick
75.21
DeepSeek-R1-0528
62.18
Selected models with downloadable weights and a reported score in this evaluation. Model-specific licenses apply. Results belong to these exact versions, not newer successors; no score is invented for GLM-5.3 or DeepSeek-V4. The publisher evaluated hosted endpoints, not Finsider self-hosted deployments.
Rank uses displayed scores; ties share rank. Bars use a 0-100 scale. No confidence intervals are available in the cited results; small differences should not be interpreted as statistically established superiority.
Source: AIMultiple / September 2026