67 / 84retained slots triggered a score cap. A ceiling limits credit; it does not award it.
Toneme · Financial-agent training & evaluation data
Financial Agent RL Environments
A representative sample of ready-to-run financial-agent tasks built for post-training and evaluation. Source artifacts, executable workflows and reward evaluation are supplied together; recorded runs below show the environment in action.
Contact us to discuss a private evaluation pack or a broader task set.
- Tasks
- 14
- Categories
- 6
- Models
- 2
- Retained slots
- 84
Task-level final score distribution
Score / 100- DCF & Operations
- 20.00
- LBO
- 13.50
- Cash Flow
- 28.50
- Loan Tape
- 10.00
- M&A & Funding
- 9.34
- IRR & Comparables
- 19.68
Lower whisker / Q1 / Median / Q3 / Upper whisker
6.23 / 10.61 / 15.00 / 20.50 / 29.08
- DCF & Operations
- 46.59
- LBO
- 24.21
- Cash Flow
- 27.22
- Loan Tape
- 62.50
- M&A & Funding
- 41.00
- IRR & Comparables
- 38.85
Lower whisker / Q1 / Median / Q3 / Upper whisker
12.57 / 25.00 / 25.00 / 42.37 / 68.19
Each observation is a task mean across retained slots. Open box: 25th–75th percentiles; middle line: median. Whiskers reach the furthest observed task scores within 1.5×IQR; open circles show all outlying tasks. Numeric label: overall task-weighted mean.
01 / Inspect a real task
Look past the aggregate.
LNG valuation
Build a linked three-statement forecast, discounted cash-flow valuation and sensitivity analysis for an LNG receiving terminal.
Required deliverable
Formula-driven financial workbook + structured outputs
This model submission earned the score shown here. It is an inspectable example of the requested deliverable, not a reference answer.
- Final score
- 75.52
GPT-6 Astra · Slot 1
Inspect the recorded execution →Inspect the redacted income statement
| Line item | Year 3A | Year 4E | Year 5E | Year 6E | Year 7E | Year 8E | Year 9E |
|---|---|---|---|---|---|---|---|
| core | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| storage | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| ancillary | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| revenue | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| operating cost | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| surcharges | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| dispatch | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| gna | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| other operating | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| ebitda | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| depreciation | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| amortization | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| ebit | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| cash interest | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
| deposit income | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld | Withheld |
02 / Workflow coverage
Six workflows. Distinct challenges.
Mean final score by workflow
Equal task weights within each categoryNot yet evaluated.
Category sizes differ. Overall scores weight tasks equally, not categories equally. Missing results are not zero scores.
3retained slots earned full credit. Successful cases belong beside constrained outcomes.
14task environments pair executable workflows with structured reward evaluation.
03 / More samples
Where outputs fall short.
02 / Operating valuation
A forecast must respond to its inputs.
Build an integrated operating forecast and valuation for a shopping mall.
Required deliverable
Linked operating forecast and valuation workbook
Earned credit constrained by a causality check
The selected Astra submission earned 36.46 points before the cap. The recorded check did not detect real numeric causality, triggering a 25-point ceiling and reducing the final score to 25.00.
- Natural score
- 36.46
- Final score
- 25.00
- Score ceiling
- 25.00
GPT-6 Astra · Slot 1 · Scoring receipt excerpt
This selected submission illustrates a specific scoring constraint. It does not establish a general capability ranking or replace the complete results below.
03 / Capital budgeting
A ceiling does not award the points beneath it.
Construct a fresh, auditable capital-budgeting workbook from supplied tabular inputs and an output contract.
Required deliverable
Auditable capital-budgeting workbook + structured outputs
Low scores remain low after a cap is applied
The retained Astra submissions stayed below the 25-point ceiling. The recorded checks did not detect the required numeric causality. Applying a ceiling preserved the lower earned scores instead of automatically awarding 25 points.
- Natural score
- 10.95
- Final score
- 10.95
- Score ceiling
- 25.00
GPT-6 Astra · Slot 1 · Scoring receipt excerpt
This case illustrates the distinction between a score ceiling and earned credit. It is not a new model run.
04 / Inside a recorded run
Execution trajectories
Follow six selected executions from recorded actions to retained outcomes. These examples illustrate process; they are not category-average or best-case trajectories.
This trajectory could not be loaded. The evaluation results remain available.
Duration shows intervals between recorded events; it does not separate model inference from tool execution. One turn is one recorded agent step. Colors group actions by their reviewed meaning, not by an automated clustering model. Check progress reads an existing process status; Pause is a recorded pause. Action summaries are redacted; reported validation is not the private evaluation score.
05 / Inspect the evidence
Retained results
01DCF & OperationsGPT-5.6 Sol20.002/2 tasks · 6/6 capped slotsGPT-6 Astra46.592/2 tasks · 3/6 capped slots
| Task | Model | Mean score | Cap triggered | Slot evidence |
|---|---|---|---|---|
| Build a linked LNG forecast, DCF valuation and sensitivity analysis. | GPT-5.6 Sol | 15.00 | 3/3 | View slotsSlot 1 · 10.00 Natural score: 21.26 Valid retained submission Slot 2 · 25.00 Natural score: 43.79 Valid retained submission Slot 3 · 10.00 Natural score: 28.15 Valid retained submission |
| Build an integrated operating forecast and valuation for a shopping mall. | GPT-5.6 Sol | 25.00 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 31.87 Valid retained submission Slot 2 · 25.00 Natural score: 32.01 Valid retained submission Slot 3 · 25.00 Natural score: 34.69 Valid retained submission |
| Build a linked LNG forecast, DCF valuation and sensitivity analysis. | GPT-6 Astra | 68.19 | 0/3 | View slotsSlot 1 · 75.52 Natural score: 75.52 Valid retained submission Slot 2 · 72.00 Natural score: 72.00 Valid retained submission · Authorized replacement Slot 3 · 57.05 Natural score: 57.05 Valid retained submission |
| Build an integrated operating forecast and valuation for a shopping mall. | GPT-6 Astra | 25.00 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 36.46 Valid retained submission Slot 2 · 25.00 Natural score: 33.79 Valid retained submission Slot 3 · 25.00 Natural score: 35.72 Valid retained submission |
02LBOGPT-5.6 Sol13.504/4 tasks · 12/12 capped slotsGPT-6 Astra24.214/4 tasks · 12/12 capped slots
| Task | Model | Mean score | Cap triggered | Slot evidence |
|---|---|---|---|---|
| Model acquisition financing, operating cases and sponsor returns for a leveraged buyout. | GPT-5.6 Sol | 12.65 | 3/3 | View slotsSlot 1 · 10.00 Natural score: 21.27 Valid retained submission Slot 2 · 10.00 Natural score: 14.78 Valid retained submission Slot 3 · 17.95 Natural score: 17.95 Valid retained submission |
| Retrieve the required financial evidence and use it in a linked buyout model. | GPT-5.6 Sol | 10.00 | 3/3 | View slotsSlot 1 · 10.00 Natural score: 18.38 Valid retained submission Slot 2 · 10.00 Natural score: 32.94 Valid retained submission Slot 3 · 10.00 Natural score: 19.68 Valid retained submission |
| Build an acquisition model with linked financing, operating and return calculations. | GPT-5.6 Sol | 16.37 | 3/3 | View slotsSlot 1 · 10.00 Natural score: 26.46 Valid retained submission Slot 2 · 19.74 Natural score: 19.74 Valid retained submission Slot 3 · 19.36 Natural score: 19.36 Valid retained submission |
| Retrieve source financial information and connect it to the required buyout outputs. | GPT-5.6 Sol | 15.00 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 30.88 Valid retained submission Slot 2 · 10.00 Natural score: 13.82 Valid retained submission Slot 3 · 10.00 Natural score: 10.53 Valid retained submission |
| Model acquisition financing, operating cases and sponsor returns for a leveraged buyout. | GPT-6 Astra | 25.00 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 33.21 Valid retained submission Slot 2 · 25.00 Natural score: 27.86 Valid retained submission Slot 3 · 25.00 Natural score: 25.96 Valid retained submission |
| Retrieve the required financial evidence and use it in a linked buyout model. | GPT-6 Astra | 24.14 | 3/3 | View slotsSlot 1 · 23.58 Natural score: 23.58 Valid retained submission Slot 2 · 24.31 Natural score: 24.31 Valid retained submission · Authorized replacement Slot 3 · 24.53 Natural score: 24.53 Valid retained submission |
| Build an acquisition model with linked financing, operating and return calculations. | GPT-6 Astra | 25.00 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 31.85 Valid retained submission Slot 2 · 25.00 Natural score: 42.71 Valid retained submission Slot 3 · 25.00 Natural score: 42.55 Valid retained submission |
| Retrieve source financial information and connect it to the required buyout outputs. | GPT-6 Astra | 22.70 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 25.37 Valid retained submission Slot 2 · 19.44 Natural score: 19.44 Valid retained submission · Authorized replacement Slot 3 · 23.66 Natural score: 23.66 Valid retained submission |
03Cash FlowGPT-5.6 Sol28.502/2 tasks · 4/6 capped slotsGPT-6 Astra27.222/2 tasks · 5/6 capped slots
| Task | Model | Mean score | Cap triggered | Slot evidence |
|---|---|---|---|---|
| Reconcile financial disclosures into a formula-driven cash-flow statement. | GPT-5.6 Sol | 29.08 | 2/3 | View slotsSlot 1 · 25.00 Natural score: 37.14 Valid retained submission Slot 2 · 25.00 Natural score: 37.84 Valid retained submission Slot 3 · 37.23 Natural score: 37.23 Valid retained submission |
| Translate transaction assumptions into linked cash-flow effects and financial controls. | GPT-5.6 Sol | 27.91 | 2/3 | View slotsSlot 1 · 21.69 Natural score: 21.69 Valid retained submission Slot 2 · 25.00 Natural score: 38.22 Valid retained submission Slot 3 · 37.05 Natural score: 37.05 Valid retained submission |
| Reconcile financial disclosures into a formula-driven cash-flow statement. | GPT-6 Astra | 29.43 | 2/3 | View slotsSlot 1 · 25.00 Natural score: 37.86 Valid retained submission Slot 2 · 38.30 Natural score: 38.30 Valid retained submission Slot 3 · 25.00 Natural score: 35.76 Valid retained submission |
| Translate transaction assumptions into linked cash-flow effects and financial controls. | GPT-6 Astra | 25.00 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 34.59 Valid retained submission Slot 2 · 25.00 Natural score: 34.35 Valid retained submission Slot 3 · 25.00 Natural score: 34.59 Valid retained submission |
04Loan TapeGPT-5.6 Sol10.002/2 tasks · 6/6 capped slotsGPT-6 Astra62.502/2 tasks · 3/6 capped slots
| Task | Model | Mean score | Cap triggered | Slot evidence |
|---|---|---|---|---|
| Apply loan eligibility rules and reconcile the resulting portfolio metrics. | GPT-5.6 Sol | 10.00 | 3/3 | View slotsSlot 1 · 10.00 Natural score: 31.90 Valid retained submission Slot 2 · 10.00 Natural score: 34.43 Valid retained submission Slot 3 · 10.00 Natural score: 34.43 Valid retained submission |
| Calculate a borrowing base from loan-level eligibility and financing constraints. | GPT-5.6 Sol | 10.00 | 3/3 | View slotsSlot 1 · 10.00 Natural score: 17.14 Valid retained submission Slot 2 · 10.00 Natural score: 25.00 Valid retained submission Slot 3 · 10.00 Natural score: 24.29 Valid retained submission |
| Apply loan eligibility rules and reconcile the resulting portfolio metrics. | GPT-6 Astra | 100.00 | 0/3 | View slotsSlot 1 · 100.00 Natural score: 100.00 Valid retained submission Slot 2 · 100.00 Natural score: 100.00 Valid retained submission Slot 3 · 100.00 Natural score: 100.00 Valid retained submission |
| Calculate a borrowing base from loan-level eligibility and financing constraints. | GPT-6 Astra | 25.00 | 3/3 | View slotsSlot 1 · 25.00 Natural score: 25.00 Valid retained submission Slot 2 · 25.00 Natural score: 25.00 Valid retained submission Slot 3 · 25.00 Natural score: 25.00 Valid retained submission |
05M&A & FundingGPT-5.6 Sol9.342/2 tasks · 1/6 capped slotsGPT-6 Astra41.002/2 tasks · 6/6 capped slots
| Task | Model | Mean score | Cap triggered | Slot evidence |
|---|---|---|---|---|
| Compare financing structures through linked FX, tax and cash-flow scenarios. | GPT-5.6 Sol | 12.44 | 1/3 | View slotsSlot 1 · 5.42 Natural score: 5.42 Valid retained submission Slot 2 · 6.91 Natural score: 6.91 Valid retained submission Slot 3 · 25.00 Natural score: 33.12 Valid retained submission |
| Model alternative funding paths and their foreign-exchange and tax effects. | GPT-5.6 Sol | 6.23 | 0/3 | View slotsSlot 1 · 6.47 Natural score: 6.47 Valid retained submission Slot 2 · 7.95 Natural score: 7.95 Valid retained submission Slot 3 · 4.28 Natural score: 4.28 Valid retained submission |
| Compare financing structures through linked FX, tax and cash-flow scenarios. | GPT-6 Astra | 38.25 | 3/3 | View slotsSlot 1 · 34.49 Natural score: 34.49 Valid retained submission Slot 2 · 38.64 Natural score: 38.64 Valid retained submission Slot 3 · 41.63 Natural score: 41.63 Valid retained submission |
| Model alternative funding paths and their foreign-exchange and tax effects. | GPT-6 Astra | 43.74 | 3/3 | View slotsSlot 1 · 42.25 Natural score: 42.25 Valid retained submission Slot 2 · 44.49 Natural score: 44.49 Valid retained submission Slot 3 · 44.49 Natural score: 44.49 Valid retained submission |
06IRR & ComparablesGPT-5.6 Sol19.682/2 tasks · 6/6 capped slotsGPT-6 Astra38.852/2 tasks · 3/6 capped slots
| Task | Model | Mean score | Cap triggered | Slot evidence |
|---|---|---|---|---|
| Compare project cash flows and returns under funding and liquidity constraints. | GPT-5.6 Sol | 21.31 | 3/3 | View slotsSlot 1 · 24.32 Natural score: 24.32 Valid retained submission Slot 2 · 14.59 Natural score: 14.59 Valid retained submission Slot 3 · 25.00 Natural score: 25.54 Valid retained submission |
| Build an auditable capital-budgeting model from supplied inputs and output requirements. | GPT-5.6 Sol | 18.06 | 3/3 | View slotsSlot 1 · 20.68 Natural score: 20.68 Valid retained submission Slot 2 · 25.00 Natural score: 31.62 Valid retained submission Slot 3 · 8.51 Natural score: 8.51 Valid retained submission |
| Compare project cash flows and returns under funding and liquidity constraints. | GPT-6 Astra | 65.14 | 0/3 | View slotsSlot 1 · 64.73 Natural score: 64.73 Valid retained submission Slot 2 · 64.73 Natural score: 64.73 Valid retained submission Slot 3 · 65.95 Natural score: 65.95 Valid retained submission |
| Build an auditable capital-budgeting model from supplied inputs and output requirements. | GPT-6 Astra | 12.57 | 3/3 | View slotsSlot 1 · 10.95 Natural score: 10.95 Valid retained submission Slot 2 · 14.59 Natural score: 14.59 Valid retained submission Slot 3 · 12.16 Natural score: 12.16 Valid retained submission · Authorized replacement |
No tasks match these filters. Clear the search or reset the filters.
Search locates tasks without recalculating category means. Sorting describes this selected sample; it is not a fair capability ranking.
06 / Credit and constraints
Before and after the cap.
Not yet evaluated.
View all category values
| Workflow | Model | Natural score | Final score |
|---|---|---|---|
| DCF & Operations | GPT-5.6 Sol | 31.96 | 20.00 |
| DCF & Operations | GPT-6 Astra | 51.76 | 46.59 |
| LBO | GPT-5.6 Sol | 20.48 | 13.50 |
| LBO | GPT-6 Astra | 28.75 | 24.21 |
| Cash Flow | GPT-5.6 Sol | 34.86 | 28.50 |
| Cash Flow | GPT-6 Astra | 35.91 | 27.22 |
| Loan Tape | GPT-5.6 Sol | 27.86 | 10.00 |
| Loan Tape | GPT-6 Astra | 62.50 | 62.50 |
| M&A & Funding | GPT-5.6 Sol | 10.69 | 9.34 |
| M&A & Funding | GPT-6 Astra | 41.00 | 41.00 |
| IRR & Comparables | GPT-5.6 Sol | 20.88 | 19.68 |
| IRR & Comparables | GPT-6 Astra | 38.85 | 38.85 |
Natural score is the recorded pre-cap score, not accuracy. Other scoring gates may already affect it. A hard cap limits the final score; it never adds credit. Hover or tap a bar for task and slot scores.
07 / Methodology
Evidence before credit.
Our ready-to-run RL environment connects financial tasks to measurable reward. Natural scores capture earned quality; hard-cap gates constrain credit when critical requirements are unmet. Recorded execution and frozen submissions keep the result inspectable.
Evidence before credit
Toneme · Evaluation methodology- 01Task & artifacts
- 02Recorded execution
- 03Frozen submission
- 04EvaluationNatural scoreQuality earnedANDHard-cap gatesCredit constrainedFinal scoreA ceiling, never a bonus.
- 05Auditable result
Checks vary by task family. Infrastructure failures remain unscored.
A complete RL environment. Ready to run.
We provide the task environment, source artifacts, Docker configuration, execution contract and reward evaluation together. With the required runtime and images installed, researchers can run agents, collect trajectories and score their submissions through the supplied workflow. The evaluation pack includes the environment and grading components; this gallery shows public evidence.
Reward quality. Constrain shortcuts.
Natural score records earned task credit; hard-cap gates constrain credit when critical requirements are unmet. Applicable checks examine formulas, recalculation and input responsiveness, reducing reliance on plausible static answers. These mechanisms address specific failure modes; they do not establish complete resistance to reward hacking.
Control the run. Preserve the submission.
Docker configurations and recorded runtime profiles support inspection and replay. The Type06 replay uses a fresh container, read-only task inputs and a separate output directory, with submission checksums checked before and after execution. Recalculation and formula-cache observations help assess stale workbook values in that family. These controls reduce residual-state risks; replay still depends on the required images and calculation environment.
Keep the evidence attached.
Selected Astra runs include recorded steps and timestamps. Submission checksums, scoring receipts and the acceptance ledger bind retained results to their evidence. Under the recorded recovery profiles, infrastructure failures remain unscored; they are not converted into model failures or used to reroll valid low scores.
One calculation, throughout the page.
Scores use a 0–100 scale. We average retained slots within a task, then weight covered tasks equally. Category means apply the same rule within each category. Coverage is shown explicitly. Different task counts and selection protocols must be considered before comparing models.
Evaluate training efficacy in a scoped pilot.
This gallery is an illustrative sample, not a statistical claim about training lift or model ranking. Training efficacy should be evaluated on a scoped pilot against the target model and objective.
Continue the technical conversation
Discuss a private evaluation pack.
Interested in the private evaluation pack or a broader task set? Contact us by email to discuss your model, workflow and evaluation objective.
Request private evaluation pack