Goaly Research

What does AI actually understand about money?

Goaly Research studies the financial capability of frontier AI systems: what they get right, where they quietly fail, and which businesses become possible once the gap closes. GoalyBench is the instrument we built to answer that with numbers instead of impressions.

Graded tasks
12,480
Capability domains
6
Top composite score
93%
Evaluation rounds
6
The research

Financial competence is not a language problem.

A model can describe a balance sheet fluently and still misread one. Fluency is cheap; arithmetic under ambiguity is not. Our work starts from the assumption that the only honest way to talk about AI and money is to grade it on the work itself: real ledgers, real receipts, real edge cases, scored the way an auditor would score them.

Capability mapping

We chart where a model is genuinely reliable with money and where it is confidently wrong, one narrow task at a time.

Opportunity modelling

Every capability that clears a reliability bar unlocks a product. We size those openings before we build toward them.

Failure economics

A misfiled receipt is an annoyance. A misread tax rule is a penalty. We weight errors by what they actually cost a person.

Deployment guardrails

We write the conditions a model must meet before it is allowed anywhere near a real ledger inside Goaly.

The instrument

GoalyBench

GoalyBench measures financial literacy the way it shows up in an independent earner's life. Every task is drawn from anonymised, permissioned activity and rewritten by hand, then graded blind by two reviewers with a third resolving disagreements. Models see a fresh held out set each round, so nothing can be memorised between releases.

  • Same time limit for models and the human control group.
  • Adversarial ledgers with duplicated, missing, and reversed entries.
  • A refusal on an unanswerable task scores higher than a confident guess.
Cash flow forecastingProjecting uneven, seasonal, multi client income.
Ledger reconciliationMatching payouts, fees, refunds, and chargebacks.
Expense classificationSorting real receipts into defensible categories.
Contract and fee mathRates, retainers, platform cuts, and net position.
Tax treatmentDeductibility judgments with a cited rationale.
Advisory judgmentKnowing when the honest answer is to refuse.

The business that sits on the other side

Each domain we grade maps to something a person currently pays for, waits for, or simply goes without. These are the openings we track.

95%Top score, reconciliation

Bookkeeping that finishes itself

Reconciliation is the largest recurring time cost for independent earners. It is also the domain where models already clear 95%.

90%Top score, advisory judgment

Advice priced for people without an accountant

Roughly nine in ten independent earners never pay for financial advice. Competent models change what that service can cost.

96%Top score, forecasting

Underwriting for thin file income

Lenders reject variable income because they cannot read it. A model that forecasts it well turns a rejection into a risk price.

91%Top score, tax treatment

Tax intelligence in the moment

Deductions are lost at the point of purchase, not at filing. This is the domain with the widest spread between models.

The goal

Make financial competence measurable, then make it useful.

Benchmarks are only worth running if something changes because of them. Ours exists to set the bar that decides which capabilities reach the people using Goaly, and in which order. Nothing ships into the product on vibes.

Give an independent earner the quality of financial thinking that used to require a firm on retainer, and be able to prove the thinking is sound.

Phase 01Shipped

Build the measurement

GoalyBench v1, six domains, blind grading, and a public methodology note.

Phase 02In progress

Grade the frontier continuously

Every major release is scored on a held out set within two weeks of launch.

Phase 032027

Publish the reliability bars

Per domain thresholds that state plainly what a model must reach before it touches a ledger.

Phase 042028

Open the evaluation set

Release the adversarial ledger corpus so other teams can reproduce and challenge the results.

GoalyBench results

Round six, scored and published.

Two frontier systems were evaluated at full reasoning budget against the same held out set, alongside a human control group and a median of eleven open weight models.

Rank 01

Claude Opus 5

93%

Strongest across every domain, and the only system to clear 90% on advisory judgment, where the correct answer is often a refusal.

Best domain
Forecasting, 96%
Weakest domain
Advisory, 90%
Rank 02

GPT 5.6 Sol Ultra

87%

Close on mechanical work and reconciliation, with the gap opening on tax treatment and on judgment calls that reward saying nothing.

Best domain
Forecasting, 91%
Weakest domain
Tax, 84%

GoalyBench v1.2 composite score

Percentage of 12,480 graded financial tasks answered correctly, blind scored.

Higher is better
Claude Opus 5Frontier model, full reasoning budget
93%
GPT 5.6 Sol UltraFrontier model, full reasoning budget
87%
Human control groupLicensed bookkeepers, same time limit
79%
Open weight medianMedian of eleven public models
64%
0255075100

Score by capability domain

The composite splits into six domains. Every score sits on the same 0 to 100 scale.

Claude Opus 5GPT 5.6 Sol Ultra
DomainGraded tasksClaude Opus 5GPT 5.6 Sol Ultra
Cash flow forecasting2,140
96%
91%
Ledger reconciliation2,380
95%
89%
Expense classification2,610
94%
88%
Contract and fee math1,720
92%
86%
Tax treatment2,090
91%
84%
Advisory judgment1,540
90%
84%

The widest gap sits in tax treatment and advisory judgment, the two domains where a wrong answer costs the most money.

Composite score across evaluation rounds

Six rounds of GoalyBench, each run on a freshly held out task set.

Claude Opus 5GPT 5.6 Sol Ultra
10090807060R1R2R3R4R5R693%87%

Both systems improved steadily, and the six point gap at the top of the board has held across the last three rounds. Values are also listed in the domain table above.

Scores are the mean of three independent runs per model. Grading is blind, the task set is held out and rotated every round, and the full methodology note is available on request from the research team.

Contact

Talk to the research team.

We publish what we find, and we would rather be corrected than quoted. If you work on financial reasoning, evaluation, or the products that depend on both, get in touch.

Research enquiriesresearch@goaly.app

Methodology questions, replication requests, and dataset access.

Partnershipspartnerships@goaly.app

Model providers and institutions who want a system evaluated.

Presspress@goaly.app

Interviews, figures, and citation of GoalyBench results.

Goaly ResearchRemote first, with the evaluation cohort spread across nine countries.
Response timeResearch and partnership mail is answered within three working days.