What does AI actually understand about money?
Goaly Research studies the financial capability of frontier AI systems: what they get right, where they quietly fail, and which businesses become possible once the gap closes. GoalyBench is the instrument we built to answer that with numbers instead of impressions.
- Graded tasks
- 12,480
- Capability domains
- 6
- Top composite score
- 93%
- Evaluation rounds
- 6
Financial competence is not a language problem.
A model can describe a balance sheet fluently and still misread one. Fluency is cheap; arithmetic under ambiguity is not. Our work starts from the assumption that the only honest way to talk about AI and money is to grade it on the work itself: real ledgers, real receipts, real edge cases, scored the way an auditor would score them.
Capability mapping
We chart where a model is genuinely reliable with money and where it is confidently wrong, one narrow task at a time.
Opportunity modelling
Every capability that clears a reliability bar unlocks a product. We size those openings before we build toward them.
Failure economics
A misfiled receipt is an annoyance. A misread tax rule is a penalty. We weight errors by what they actually cost a person.
Deployment guardrails
We write the conditions a model must meet before it is allowed anywhere near a real ledger inside Goaly.
GoalyBench
GoalyBench measures financial literacy the way it shows up in an independent earner's life. Every task is drawn from anonymised, permissioned activity and rewritten by hand, then graded blind by two reviewers with a third resolving disagreements. Models see a fresh held out set each round, so nothing can be memorised between releases.
- Same time limit for models and the human control group.
- Adversarial ledgers with duplicated, missing, and reversed entries.
- A refusal on an unanswerable task scores higher than a confident guess.
The business that sits on the other side
Each domain we grade maps to something a person currently pays for, waits for, or simply goes without. These are the openings we track.
Bookkeeping that finishes itself
Reconciliation is the largest recurring time cost for independent earners. It is also the domain where models already clear 95%.
Advice priced for people without an accountant
Roughly nine in ten independent earners never pay for financial advice. Competent models change what that service can cost.
Underwriting for thin file income
Lenders reject variable income because they cannot read it. A model that forecasts it well turns a rejection into a risk price.
Tax intelligence in the moment
Deductions are lost at the point of purchase, not at filing. This is the domain with the widest spread between models.
Make financial competence measurable, then make it useful.
Benchmarks are only worth running if something changes because of them. Ours exists to set the bar that decides which capabilities reach the people using Goaly, and in which order. Nothing ships into the product on vibes.
Give an independent earner the quality of financial thinking that used to require a firm on retainer, and be able to prove the thinking is sound.
Build the measurement
GoalyBench v1, six domains, blind grading, and a public methodology note.
Grade the frontier continuously
Every major release is scored on a held out set within two weeks of launch.
Publish the reliability bars
Per domain thresholds that state plainly what a model must reach before it touches a ledger.
Open the evaluation set
Release the adversarial ledger corpus so other teams can reproduce and challenge the results.
Round six, scored and published.
Two frontier systems were evaluated at full reasoning budget against the same held out set, alongside a human control group and a median of eleven open weight models.
Claude Opus 5
93%Strongest across every domain, and the only system to clear 90% on advisory judgment, where the correct answer is often a refusal.
- Best domain
- Forecasting, 96%
- Weakest domain
- Advisory, 90%
GPT 5.6 Sol Ultra
87%Close on mechanical work and reconciliation, with the gap opening on tax treatment and on judgment calls that reward saying nothing.
- Best domain
- Forecasting, 91%
- Weakest domain
- Tax, 84%
GoalyBench v1.2 composite score
Percentage of 12,480 graded financial tasks answered correctly, blind scored.
Score by capability domain
The composite splits into six domains. Every score sits on the same 0 to 100 scale.
The widest gap sits in tax treatment and advisory judgment, the two domains where a wrong answer costs the most money.
Composite score across evaluation rounds
Six rounds of GoalyBench, each run on a freshly held out task set.
Both systems improved steadily, and the six point gap at the top of the board has held across the last three rounds. Values are also listed in the domain table above.
Scores are the mean of three independent runs per model. Grading is blind, the task set is held out and rotated every round, and the full methodology note is available on request from the research team.
Talk to the research team.
We publish what we find, and we would rather be corrected than quoted. If you work on financial reasoning, evaluation, or the products that depend on both, get in touch.
Methodology questions, replication requests, and dataset access.
Model providers and institutions who want a system evaluated.
Interviews, figures, and citation of GoalyBench results.