On September 16, Ant Group released a finance-focused open-weights model. Ling-3.0-flash-Fin. Ant says it built the model with financial institutions to support financial research, source checking, valuation spreadsheets, and report writing.
But the scores look puzzling at first. Intelligence Index 23. The top model scores 58, more than twice as high. Yet Ant says this smaller model can be better for finance work. Why? Because finance measures different things than general benchmarks.
Three Ways Finance Differs from General AI
1. Numbers are the answer
General writing: "roughly right is fine" → Finance: one decimal is money
2. Sources are part of the answer
General writing: plausible passes → Finance: no source means invalid
3. Format is the work
General writing: free format is OK → Finance: spreadsheets and report templates are the deliverable
Ant's phrasing is exact: "checking sources, building valuation spreadsheets and writing reports." Source checks, numeric deliverables, reports. That is all of finance RAG.
How to Read the Ling-3.0-flash-Fin Scores
| Item | Number | Reading |
|---|---|---|
| Intelligence Index | 23 | Small-tier overall |
| Finance & Accounting Index | 24 | Visible on the finance axis |
| Work knowledge accuracy | 17% | Higher than sibling VL at 11% |
| Work knowledge hallucination rate | 33% | Worse than VL at 19% (accuracy and hallucination rose together) |
| GDPval practical Elo | 1171 | Practical score on docs and spreadsheet tasks |
| Size | 124B total, 5.1B active | MoE, above the Pareto line for active size vs intelligence |
| License | MIT | Commercial use allowed |
Here is the key paradox. Finance tuning raised accuracy, but hallucination rose too. Domain training increased both knowledge and bluffing. So finance AI evaluation must look at accuracy and non-hallucination together. The principle from the hallucination vs abstention article applies most sharply in finance.
The AA Capability Indices v1.1 also show how finance evaluation is built. Work knowledge 30%, practical knowledge tasks 30%, reasoning 20%, tool use 10%, long docs 5%, non-hallucination 5%. Note that tool use (the AutomationBench finance branch) is newly added, while customer-service simulation is removed. Finance AI evaluation is shifting from "talking well" to "finishing work."
4 Places Where Finance RAG Breaks
Apply the 5-stage RAG failure guide to finance and you get this.
1. Table parsing: financial statements and spreadsheets arrive broken
→ Check the source first with the [Mistral OCR method](/en/blog/mistral-ocr-4-1-document-rag/)
2. Period and version mixing: 2024 rules and 2026 rules appear in one answer
→ Force date and version metadata into chunking
3. Missing numeric citations: the answer exists but the source cell or page does not
→ Force a source number per sentence (Hebbia-style citation recall)
4. Failed refusal: inventing numbers with no evidence
→ Allow "Not confirmed in the documents" as a valid answer
Item 4 is the lifeline in finance. Saying "I don't know" beats writing a number you do not have. Argon recording the lowest hallucination rate at 15% points the same way.
Demand Has Already Arrived
Three signals show the shift.
- ChatGPT personal finance: bank and brokerage links, 200M people asking finance questions monthly
- FactSet and S&P price moves reported after Anthropic finance agents
- AA finance index adds a tool-use axis (service simulation removed)
In short, the market is moving from "a model that knows finance" to "an agent that finishes finance work." What matters is not the model score but whether spreadsheets, reports, and source checks get completed.
CodeBridge Mini Lab: Build an Eval Set from 5 Finance Docs
1. Prepare 5 doc types (annual report, terms, notice, spreadsheet, meeting notes)
2. Write 10 questions:
- 4 numeric citations (exact figures + source pages required)
- 3 period comparisons (include year and version distinctions)
- 2 summaries (require a one-table summary)
- 1 trap (content missing from docs → refusal is correct)
3. Judge:
[ ] Numeric accuracy (down to decimals)
[ ] Source labels (page, cell, clause)
[ ] Version separation (old vs new confusion)
[ ] Trap refusal (outputs "not confirmed" or not)
Add this eval set to the structure from the internal docs chatbot article and you have minimum validation for a finance chatbot.
Conclusion: Finance AI Has Its Own Scorecard
To sum up:
It is not that 58 overall is enviable. The exam itself is different.
Finance is graded on numeric accuracy, source citation, version separation, and trap refusal. That is also what the 23 of Ling-3.0-flash-Fin suggests. Even a small model becomes useful when the test matches the domain. Build a 10-question eval set from your own team docs. That is the first step into finance AI.
Further reading
- When RAG Fails, Do Not Blame Embeddings First
- RAG Through Mistral OCR 4.1
- How Does an Internal Docs AI Chatbot Work?
References
- Artificial Analysis: Ant Group Ling-3.0-flash-Fin
- Artificial Analysis: Capability Indices v1.1
- Artificial Analysis: Finance & Accounting Index
Go deeper with a course
If you want to design production RAG from table parsing to enforced citations, from classic RAG to GraphRAG and agentic RAG, guided lessons connect directly to this finance eval set.