PROPRIETARYPAGE UPDATED

Toneme · Financial-agent training & evaluation data

Financial Agent RL Environments

A representative sample of ready-to-run financial-agent tasks built for post-training and evaluation. Source artifacts, executable workflows and reward evaluation are supplied together; recorded runs below show the environment in action.

Contact us to discuss a private evaluation pack or a broader task set.

Tasks
14
Categories
6
Models
2
Retained slots
84

Task-level final score distribution

Score / 100
OpenAIGPT-5.6 Sol14 / 14 tasks · 42 slots
Full-credit tasks: 0/14 (0.00%)Full-credit slots: 0/42 (0.00%)
OpenAIGPT-6 Astra14 / 14 tasks · 42 slots
Full-credit tasks: 1/14 (7.14%)Full-credit slots: 3/42 (7.14%)

Each observation is a task mean across retained slots. Open box: 25th–75th percentiles; middle line: median. Whiskers reach the furthest observed task scores within 1.5×IQR; open circles show all outlying tasks. Numeric label: overall task-weighted mean.

01 / Inspect a real task

Look past the aggregate.

Redacted model-submitted artifact

LNG valuation

Build a linked three-statement forecast, discounted cash-flow valuation and sensitivity analysis for an LNG receiving terminal.

Required deliverable
Formula-driven financial workbook + structured outputs

This model submission earned the score shown here. It is an inspectable example of the requested deliverable, not a reference answer.

Final score
75.52

GPT-6 Astra · Slot 1

Inspect the recorded execution →
Inspect the redacted income statement
IS | LNG Receiving TerminalRMB million unless labelled otherwise
Line itemYear 3AYear 4EYear 5EYear 6EYear 7EYear 8EYear 9E
coreWithheldWithheldWithheldWithheldWithheldWithheldWithheld
storageWithheldWithheldWithheldWithheldWithheldWithheldWithheld
ancillaryWithheldWithheldWithheldWithheldWithheldWithheldWithheld
revenueWithheldWithheldWithheldWithheldWithheldWithheldWithheld
operating costWithheldWithheldWithheldWithheldWithheldWithheldWithheld
surchargesWithheldWithheldWithheldWithheldWithheldWithheldWithheld
dispatchWithheldWithheldWithheldWithheldWithheldWithheldWithheld
gnaWithheldWithheldWithheldWithheldWithheldWithheldWithheld
other operatingWithheldWithheldWithheldWithheldWithheldWithheldWithheld
ebitdaWithheldWithheldWithheldWithheldWithheldWithheldWithheld
depreciationWithheldWithheldWithheldWithheldWithheldWithheldWithheld
amortizationWithheldWithheldWithheldWithheldWithheldWithheldWithheld
ebitWithheldWithheldWithheldWithheldWithheldWithheldWithheld
cash interestWithheldWithheldWithheldWithheldWithheldWithheldWithheld
deposit incomeWithheldWithheldWithheldWithheldWithheldWithheldWithheld
HTML reconstruction of the model-submitted income statement excerpt. Row labels and column headers are preserved; financial values and formulas are withheld. Typography and colors are adapted for this gallery.

02 / Workflow coverage

Six workflows. Distinct challenges.

Fixed 0–100 scale

Mean final score by workflow

Equal task weights within each category

Category sizes differ. Overall scores weight tasks equally, not categories equally. Missing results are not zero scores.

Credit constrained

67 / 84retained slots triggered a score cap. A ceiling limits credit; it does not award it.

Success retained

3retained slots earned full credit. Successful cases belong beside constrained outcomes.

Environment included

14task environments pair executable workflows with structured reward evaluation.

03 / More samples

Where outputs fall short.

Redacted task summaries

02 / Operating valuation

A forecast must respond to its inputs.

Build an integrated operating forecast and valuation for a shopping mall.

Required deliverable
Linked operating forecast and valuation workbook

Earned credit constrained by a causality check

The selected Astra submission earned 36.46 points before the cap. The recorded check did not detect real numeric causality, triggering a 25-point ceiling and reducing the final score to 25.00.

Natural score
36.46
Final score
25.00
Score ceiling
25.00

GPT-6 Astra · Slot 1 · Scoring receipt excerpt

This selected submission illustrates a specific scoring constraint. It does not establish a general capability ranking or replace the complete results below.

04 / Inside a recorded run

Execution trajectories

GPT-6 Astra · one retained slot per workflow

Follow six selected executions from recorded actions to retained outcomes. These examples illustrate process; they are not category-average or best-case trajectories.

      Duration shows intervals between recorded events; it does not separate model inference from tool execution. One turn is one recorded agent step. Colors group actions by their reviewed meaning, not by an automated clustering model. Check progress reads an existing process status; Pause is a recorded pause. Action summaries are redacted; reported validation is not the private evaluation score.

      05 / Inspect the evidence

      Retained results

      6 categories
      01DCF & OperationsView tasksGPT-5.6 Sol20.002/2 tasks · 6/6 capped slotsGPT-6 Astra46.592/2 tasks · 3/6 capped slots
      TaskModelMean scoreCap triggeredSlot evidence
      Build a linked LNG forecast, DCF valuation and sensitivity analysis.GPT-5.6 Sol15.003/3
      View slots
      Slot 1 · 10.00

      Natural score: 21.26
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 43.79
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 10.00

      Natural score: 28.15
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Build an integrated operating forecast and valuation for a shopping mall.GPT-5.6 Sol25.003/3
      View slots
      Slot 1 · 25.00

      Natural score: 31.87
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 32.01
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 34.69
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Build a linked LNG forecast, DCF valuation and sensitivity analysis.GPT-6 Astra68.190/3
      View slots
      Slot 1 · 75.52

      Natural score: 75.52
      Final score: 75.52
      Score ceiling: Not triggered

      Valid retained submission

      Slot 2 · 72.00

      Natural score: 72.00
      Final score: 72.00
      Score ceiling: Not triggered

      Valid retained submission · Authorized replacement

      Slot 3 · 57.05

      Natural score: 57.05
      Final score: 57.05
      Score ceiling: Not triggered

      Valid retained submission

      Build an integrated operating forecast and valuation for a shopping mall.GPT-6 Astra25.003/3
      View slots
      Slot 1 · 25.00

      Natural score: 36.46
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 33.79
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 35.72
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      02LBOView tasksGPT-5.6 Sol13.504/4 tasks · 12/12 capped slotsGPT-6 Astra24.214/4 tasks · 12/12 capped slots
      TaskModelMean scoreCap triggeredSlot evidence
      Model acquisition financing, operating cases and sponsor returns for a leveraged buyout.GPT-5.6 Sol12.653/3
      View slots
      Slot 1 · 10.00

      Natural score: 21.27
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 2 · 10.00

      Natural score: 14.78
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 3 · 17.95

      Natural score: 17.95
      Final score: 17.95
      Score ceiling: 25.00

      Valid retained submission

      Retrieve the required financial evidence and use it in a linked buyout model.GPT-5.6 Sol10.003/3
      View slots
      Slot 1 · 10.00

      Natural score: 18.38
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 2 · 10.00

      Natural score: 32.94
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 3 · 10.00

      Natural score: 19.68
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Build an acquisition model with linked financing, operating and return calculations.GPT-5.6 Sol16.373/3
      View slots
      Slot 1 · 10.00

      Natural score: 26.46
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 2 · 19.74

      Natural score: 19.74
      Final score: 19.74
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 19.36

      Natural score: 19.36
      Final score: 19.36
      Score ceiling: 25.00

      Valid retained submission

      Retrieve source financial information and connect it to the required buyout outputs.GPT-5.6 Sol15.003/3
      View slots
      Slot 1 · 25.00

      Natural score: 30.88
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 10.00

      Natural score: 13.82
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 3 · 10.00

      Natural score: 10.53
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Model acquisition financing, operating cases and sponsor returns for a leveraged buyout.GPT-6 Astra25.003/3
      View slots
      Slot 1 · 25.00

      Natural score: 33.21
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 27.86
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 25.96
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Retrieve the required financial evidence and use it in a linked buyout model.GPT-6 Astra24.143/3
      View slots
      Slot 1 · 23.58

      Natural score: 23.58
      Final score: 23.58
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 24.31

      Natural score: 24.31
      Final score: 24.31
      Score ceiling: 25.00

      Valid retained submission · Authorized replacement

      Slot 3 · 24.53

      Natural score: 24.53
      Final score: 24.53
      Score ceiling: 25.00

      Valid retained submission

      Build an acquisition model with linked financing, operating and return calculations.GPT-6 Astra25.003/3
      View slots
      Slot 1 · 25.00

      Natural score: 31.85
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 42.71
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 42.55
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Retrieve source financial information and connect it to the required buyout outputs.GPT-6 Astra22.703/3
      View slots
      Slot 1 · 25.00

      Natural score: 25.37
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 19.44

      Natural score: 19.44
      Final score: 19.44
      Score ceiling: 25.00

      Valid retained submission · Authorized replacement

      Slot 3 · 23.66

      Natural score: 23.66
      Final score: 23.66
      Score ceiling: 25.00

      Valid retained submission

      03Cash FlowView tasksGPT-5.6 Sol28.502/2 tasks · 4/6 capped slotsGPT-6 Astra27.222/2 tasks · 5/6 capped slots
      TaskModelMean scoreCap triggeredSlot evidence
      Reconcile financial disclosures into a formula-driven cash-flow statement.GPT-5.6 Sol29.082/3
      View slots
      Slot 1 · 25.00

      Natural score: 37.14
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 37.84
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 37.23

      Natural score: 37.23
      Final score: 37.23
      Score ceiling: Not triggered

      Valid retained submission

      Translate transaction assumptions into linked cash-flow effects and financial controls.GPT-5.6 Sol27.912/3
      View slots
      Slot 1 · 21.69

      Natural score: 21.69
      Final score: 21.69
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 38.22
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 37.05

      Natural score: 37.05
      Final score: 37.05
      Score ceiling: Not triggered

      Valid retained submission

      Reconcile financial disclosures into a formula-driven cash-flow statement.GPT-6 Astra29.432/3
      View slots
      Slot 1 · 25.00

      Natural score: 37.86
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 38.30

      Natural score: 38.30
      Final score: 38.30
      Score ceiling: Not triggered

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 35.76
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Translate transaction assumptions into linked cash-flow effects and financial controls.GPT-6 Astra25.003/3
      View slots
      Slot 1 · 25.00

      Natural score: 34.59
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 34.35
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 34.59
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      04Loan TapeView tasksGPT-5.6 Sol10.002/2 tasks · 6/6 capped slotsGPT-6 Astra62.502/2 tasks · 3/6 capped slots
      TaskModelMean scoreCap triggeredSlot evidence
      Apply loan eligibility rules and reconcile the resulting portfolio metrics.GPT-5.6 Sol10.003/3
      View slots
      Slot 1 · 10.00

      Natural score: 31.90
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 2 · 10.00

      Natural score: 34.43
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 3 · 10.00

      Natural score: 34.43
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Calculate a borrowing base from loan-level eligibility and financing constraints.GPT-5.6 Sol10.003/3
      View slots
      Slot 1 · 10.00

      Natural score: 17.14
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 2 · 10.00

      Natural score: 25.00
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Slot 3 · 10.00

      Natural score: 24.29
      Final score: 10.00
      Score ceiling: 10.00

      Valid retained submission

      Apply loan eligibility rules and reconcile the resulting portfolio metrics.GPT-6 Astra100.000/3
      View slots
      Slot 1 · 100.00

      Natural score: 100.00
      Final score: 100.00
      Score ceiling: Not triggered

      Valid retained submission

      Slot 2 · 100.00

      Natural score: 100.00
      Final score: 100.00
      Score ceiling: Not triggered

      Valid retained submission

      Slot 3 · 100.00

      Natural score: 100.00
      Final score: 100.00
      Score ceiling: Not triggered

      Valid retained submission

      Calculate a borrowing base from loan-level eligibility and financing constraints.GPT-6 Astra25.003/3
      View slots
      Slot 1 · 25.00

      Natural score: 25.00
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 25.00
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 25.00
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      05M&A & FundingView tasksGPT-5.6 Sol9.342/2 tasks · 1/6 capped slotsGPT-6 Astra41.002/2 tasks · 6/6 capped slots
      TaskModelMean scoreCap triggeredSlot evidence
      Compare financing structures through linked FX, tax and cash-flow scenarios.GPT-5.6 Sol12.441/3
      View slots
      Slot 1 · 5.42

      Natural score: 5.42
      Final score: 5.42
      Score ceiling: Not triggered

      Valid retained submission

      Slot 2 · 6.91

      Natural score: 6.91
      Final score: 6.91
      Score ceiling: Not triggered

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 33.12
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Model alternative funding paths and their foreign-exchange and tax effects.GPT-5.6 Sol6.230/3
      View slots
      Slot 1 · 6.47

      Natural score: 6.47
      Final score: 6.47
      Score ceiling: Not triggered

      Valid retained submission

      Slot 2 · 7.95

      Natural score: 7.95
      Final score: 7.95
      Score ceiling: Not triggered

      Valid retained submission

      Slot 3 · 4.28

      Natural score: 4.28
      Final score: 4.28
      Score ceiling: Not triggered

      Valid retained submission

      Compare financing structures through linked FX, tax and cash-flow scenarios.GPT-6 Astra38.253/3
      View slots
      Slot 1 · 34.49

      Natural score: 34.49
      Final score: 34.49
      Score ceiling: 49.00

      Valid retained submission

      Slot 2 · 38.64

      Natural score: 38.64
      Final score: 38.64
      Score ceiling: 49.00

      Valid retained submission

      Slot 3 · 41.63

      Natural score: 41.63
      Final score: 41.63
      Score ceiling: 49.00

      Valid retained submission

      Model alternative funding paths and their foreign-exchange and tax effects.GPT-6 Astra43.743/3
      View slots
      Slot 1 · 42.25

      Natural score: 42.25
      Final score: 42.25
      Score ceiling: 49.00

      Valid retained submission

      Slot 2 · 44.49

      Natural score: 44.49
      Final score: 44.49
      Score ceiling: 49.00

      Valid retained submission

      Slot 3 · 44.49

      Natural score: 44.49
      Final score: 44.49
      Score ceiling: 49.00

      Valid retained submission

      06IRR & ComparablesView tasksGPT-5.6 Sol19.682/2 tasks · 6/6 capped slotsGPT-6 Astra38.852/2 tasks · 3/6 capped slots
      TaskModelMean scoreCap triggeredSlot evidence
      Compare project cash flows and returns under funding and liquidity constraints.GPT-5.6 Sol21.313/3
      View slots
      Slot 1 · 24.32

      Natural score: 24.32
      Final score: 24.32
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 14.59

      Natural score: 14.59
      Final score: 14.59
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 25.00

      Natural score: 25.54
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Build an auditable capital-budgeting model from supplied inputs and output requirements.GPT-5.6 Sol18.063/3
      View slots
      Slot 1 · 20.68

      Natural score: 20.68
      Final score: 20.68
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 25.00

      Natural score: 31.62
      Final score: 25.00
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 8.51

      Natural score: 8.51
      Final score: 8.51
      Score ceiling: 25.00

      Valid retained submission

      Compare project cash flows and returns under funding and liquidity constraints.GPT-6 Astra65.140/3
      View slots
      Slot 1 · 64.73

      Natural score: 64.73
      Final score: 64.73
      Score ceiling: Not triggered

      Valid retained submission

      Slot 2 · 64.73

      Natural score: 64.73
      Final score: 64.73
      Score ceiling: Not triggered

      Valid retained submission

      Slot 3 · 65.95

      Natural score: 65.95
      Final score: 65.95
      Score ceiling: Not triggered

      Valid retained submission

      Build an auditable capital-budgeting model from supplied inputs and output requirements.GPT-6 Astra12.573/3
      View slots
      Slot 1 · 10.95

      Natural score: 10.95
      Final score: 10.95
      Score ceiling: 25.00

      Valid retained submission

      Slot 2 · 14.59

      Natural score: 14.59
      Final score: 14.59
      Score ceiling: 25.00

      Valid retained submission

      Slot 3 · 12.16

      Natural score: 12.16
      Final score: 12.16
      Score ceiling: 25.00

      Valid retained submission · Authorized replacement

      Search locates tasks without recalculating category means. Sorting describes this selected sample; it is not a fair capability ranking.

      06 / Credit and constraints

      Before and after the cap.

      View all category values
      WorkflowModelNatural scoreFinal score
      DCF & OperationsGPT-5.6 Sol31.9620.00
      DCF & OperationsGPT-6 Astra51.7646.59
      LBOGPT-5.6 Sol20.4813.50
      LBOGPT-6 Astra28.7524.21
      Cash FlowGPT-5.6 Sol34.8628.50
      Cash FlowGPT-6 Astra35.9127.22
      Loan TapeGPT-5.6 Sol27.8610.00
      Loan TapeGPT-6 Astra62.5062.50
      M&A & FundingGPT-5.6 Sol10.699.34
      M&A & FundingGPT-6 Astra41.0041.00
      IRR & ComparablesGPT-5.6 Sol20.8819.68
      IRR & ComparablesGPT-6 Astra38.8538.85

      Natural score is the recorded pre-cap score, not accuracy. Other scoring gates may already affect it. A hard cap limits the final score; it never adds credit. Hover or tap a bar for task and slot scores.

      07 / Methodology

      Evidence before credit.

      Our ready-to-run RL environment connects financial tasks to measurable reward. Natural scores capture earned quality; hard-cap gates constrain credit when critical requirements are unmet. Recorded execution and frozen submissions keep the result inspectable.

      Evidence before credit

      Toneme · Evaluation methodology
      1. 01Task & artifacts
      2. 02Recorded execution
      3. 03Frozen submission
      4. 04Evaluation
        Natural scoreQuality earned
        AND
        Hard-cap gatesCredit constrained
        Final score
        A ceiling, never a bonus.
      5. 05Auditable result
      Execution traceSteps, tools, recorded time
      Controlled replayDocker configuration · isolated runs
      Evidence bindingSubmission checksums · scoring receipts

      Checks vary by task family. Infrastructure failures remain unscored.

      Evaluation design overview. Checks and replay mechanisms vary by task family.

      A complete RL environment. Ready to run.

      We provide the task environment, source artifacts, Docker configuration, execution contract and reward evaluation together. With the required runtime and images installed, researchers can run agents, collect trajectories and score their submissions through the supplied workflow. The evaluation pack includes the environment and grading components; this gallery shows public evidence.

      Reward quality. Constrain shortcuts.

      Natural score records earned task credit; hard-cap gates constrain credit when critical requirements are unmet. Applicable checks examine formulas, recalculation and input responsiveness, reducing reliance on plausible static answers. These mechanisms address specific failure modes; they do not establish complete resistance to reward hacking.

      Control the run. Preserve the submission.

      Docker configurations and recorded runtime profiles support inspection and replay. The Type06 replay uses a fresh container, read-only task inputs and a separate output directory, with submission checksums checked before and after execution. Recalculation and formula-cache observations help assess stale workbook values in that family. These controls reduce residual-state risks; replay still depends on the required images and calculation environment.

      Keep the evidence attached.

      Selected Astra runs include recorded steps and timestamps. Submission checksums, scoring receipts and the acceptance ledger bind retained results to their evidence. Under the recorded recovery profiles, infrastructure failures remain unscored; they are not converted into model failures or used to reroll valid low scores.

      One calculation, throughout the page.

      Scores use a 0–100 scale. We average retained slots within a task, then weight covered tasks equally. Category means apply the same rule within each category. Coverage is shown explicitly. Different task counts and selection protocols must be considered before comparing models.

      Evaluate training efficacy in a scoped pilot.

      This gallery is an illustrative sample, not a statistical claim about training lift or model ranking. Training efficacy should be evaluated on a scoped pilot against the target model and objective.

      Continue the technical conversation

      Discuss a private evaluation pack.

      Interested in the private evaluation pack or a broader task set? Contact us by email to discuss your model, workflow and evaluation objective.

      Request private evaluation pack

      jaredleung@toneme.ai