AI News

AI math comparison tests and results

Measured results

Compare AI models on applied math tasks

This bank measures numerical correctness on twelve applied problems. It does not score the quality of explanations.

Last tested
2026-10-07T23:26:26.061111+00:00

Last successful source fetch
2026-10-06T09:04:42.241615+00:00

Last published
2026-10-08T00:41:47.842434+00:00

Dated benchmark leaderboard

Rows use the same task bank and show their actual test dates. Separate runs do not establish a current matched-cohort winner. The initial launch expansion includes Standard configurations; Flex is deferred.

Completed configurations and their original run dates. Scored usage costs are estimates; admission canaries and fee holds are tracked separately in the programme budget.
ModelTask passesBankTestedProcessingEstimated scored costEvidence
GPT-6 Luna
gpt-6-luna
36 / 36tested-best-bank-v2-reviewed2026-10-07T05:51:59.343583+00:00OpenAI direct API, Standard global
medium reasoning; 8192-token cap
$0.0061 (estimated)Original JSON
GPT-6.1 Sol
gpt-6.1-sol
36 / 36tested-best-bank-v2-reviewed2026-10-05T05:54:39.327428+00:00OpenAI direct API, Standard global
medium reasoning; 8192-token cap
$0.0500 (estimated)Original JSON
Claude Opus 5.5
claude-opus-5-5
33 / 36tested-best-bank-v2-reviewed2026-10-07T06:00:52.001925+00:00Claude direct API, Standard global
medium reasoning; 8192-token cap
$0.3078 (estimated)Original JSON
Gemini 3.8 Flash
gemini-3.8-flash
32 / 36tested-best-bank-v2-reviewed2026-10-05T05:54:39.327428+00:00Gemini Developer API, paid Standard
medium reasoning; 8192-token cap
$0.1359 (estimated)Original JSON
Claude Haiku 5.5
claude-haiku-5-5
25 / 36tested-best-bank-v2-reviewed2026-10-07T23:26:26.061111+00:00Claude direct API, Standard global
medium reasoning; 8192-token cap
$0.0118 (estimated)Original JSON
Claude Sonnet 5.5
claude-sonnet-5-5
2 / 36tested-best-bank-v2-reviewed2026-10-05T05:54:39.327428+00:00Claude direct API, Standard global
medium reasoning; 8192-token cap
$0.1355 (estimated)Original JSON
Run 2026-10-05T05:54:39.327428+00:00 · tested-best-bank-v2-reviewed · complete outputs
claude-sonnet-5-5 · math-allocation · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

$10.00 is 1000 cents. Each department's exact share is 1000/3 = 333.333... cents, so the floor is 333 each, which totals 999 cents. One cent remains.

All three fractional remainders are equal (about 0.333), so the tie goes to the lowest index, which is department A. A gets the extra cent.

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0020 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-allocation · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

$10.00 is 1000 cents. Each department's exact share is 1000/3 = 333.33... cents, so the floor is 333 each.

The floors sum to 999, leaving 1 cent. All three fractional remainders are equal, so the tie goes to the lowest index, which is A. A gets 334 and B and C get 333.

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0020 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-allocation · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

$10.00 is 1000 cents. Each department's exact share is 1000/3 = 333.333... cents, so each floor is 333 and the three floors total 999. One cent remains.

All three fractional remainders are equal (0.333...), so the tie goes to the lowest index. Department A gets the extra cent.

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0019 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-allocation · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0019 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-allocation · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0018 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-allocation · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0018 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-allocation · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0009 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-allocation · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0008 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-allocation · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0008 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-bayes · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

P(D|F) = 0.02·0.95 / (0.02·0.95 + 0.98·0.04) = 0.019 / (0.019 + 0.0392) = 0.019/0.0582 = 190/582 = 95/291.

Compute 95/291: 291·0.326 = 94.866, remainder 0.134. 0.134/291 = 0.00046048...
So 0.32646048...

More digits: 95/291 long division.
950/291 = 3 (873), r77
770/291 = 2 (582), r188
1880/291 = 6 (1746), r134
1340/291 = 4 (1164), r176
1760/291 = 6 (1746), r14
140 → 0, r140
1400/291 = 4 (1164), r236
2360/291 = 8 (2328), r32
320/291 = 1, r29
290 → 0, r290
2900/291 = 9 (2619), r281
2810/291 = 9 (2619), r191
1910/291 = 6 (1746), r164

Digits: 0.3264604810996...

{"probability": 0.3264604810996}

$0.0041 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-bayes · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

P(D|F) = 0.02*0.95 / (0.02*0.95 + 0.98*0.04) = 0.019 / (0.019 + 0.0392) = 0.019/0.0582 = 190/582 = 95/291.

Compute 95/291: 291*0.3 = 87.3, remainder 7.7 → 0.02646... Let's do long division.
95/291 = 0.326460481...
291*0.3 = 87.3; remainder 7.7
7.7/291 = 0.026460...
291*0.026 = 7.566; remainder 0.134
0.134/291 = 0.00046048...
291*0.00046 = 0.13386; remainder 0.00014
0.00014/291 = 4.8110e-7
Total: 0.3 + 0.026 + 0.00046 + 0.00000048110 = 0.3264604811

More digits: 0.00014/291 = 4.81099656e-7. So 0.326460481099656...

{"probability": 0.3264604810997}

$0.0038 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-bayes · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

P(defective | flag) = (0.02 × 0.95) / (0.02 × 0.95 + 0.98 × 0.04)

The numerator is 0.019 and the second term is 0.0392, so the denominator is 0.0582.

0.019 / 0.0582 = 190/582 = 95/291

Long division of 95/291:
- 950/291 = 3 remainder 77, giving digit 3
- 770/291 = 2 remainder 188, giving digit 2
- 1880/291 = 6 remainder 134 (1746), giving digit 6
- 1340/291 = 4 remainder 176 (1164), giving digit 4
- 1760/291 = 6 remainder 14 (1746), giving digit 6
- 140/291 = 0, giving digit 0
- 1400/291 = 4 remainder 236 (1164), giving digit 4
- 2360/291 = 8 remainder 32 (2328), giving digit 8
- 320/291 = 1 remainder 29, giving digit 1
- 290/291 = 0, giving digit 0
- 2900/291 = 9 remainder 281 (2619), giving digit 9
- 2810/291 = 9 remainder 191 (2619), giving digit 9

So 95/291 = 0.326460481099...

{"probability": 0.326460481099}

$0.0049 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-bayes · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"probability": 0.32646048109965636}

$0.0042 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-bayes · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"probability": 0.326460481099656}

$0.0039 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-bayes · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"probability": 0.32646048109965636}

$0.0039 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-bayes · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"probability":0.32646048109965636}

$0.0008 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-bayes · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"probability":0.32646048109965636}

$0.0009 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-bayes · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"probability":0.32646048109965636}

$0.0008 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-break-even · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

Price 40. Variable costs: 17 + 3 + 0.025·40 + 0.30 = 17 + 3 + 1 + 0.3 = 21.3.

Contribution per unit = 40 − 21.3 = 18.7.

Break-even units = 1870 / 18.7 = 100 exactly.

At 100 units: profit = 100 · 18.7 − 1870 = 1870 − 1870 = 0.

{"profit": 0, "units": 100}

$0.0022 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-break-even · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

Contribution per unit = 40 - 17 - 3 - (0.025×40 + 0.30) = 40 - 20 - 1.00 - 0.30 = 18.70.

Break-even units = 1870 / 18.7 = 100 exactly.

Profit at 100 units = 100 × 18.7 - 1870 = 0.

{"profit": 0, "units": 100}

$0.0019 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-break-even · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

Contribution per unit = 40 − 17 − 3 − (0.025×40 + 0.30) = 40 − 17 − 3 − 1.00 − 0.30 = 18.70.

Break-even units = 1870 / 18.70 = 100 exactly.

Profit at 100 units = 100 × 18.70 − 1870 = 0.

{"profit": 0, "units": 100}

$0.0019 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-break-even · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"profit": 0, "units": 100}

$0.0023 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-break-even · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"profit": 0, "units": 100}

$0.0024 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-break-even · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"profit": 0, "units": 100}

$0.0021 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-break-even · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"profit":0,"units":100}

$0.0008 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-break-even · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"profit":0,"units":100}

$0.0008 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-break-even · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"profit":0,"units":100}

$0.0008 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-capacity · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

Maximize 40A + 50B subject to 3A + 2B ≤ 120, 2A + 4B ≤ 160, with A and B whole numbers.

LP vertex: from 3A + 2B = 120 and A + 2B = 80, subtract to get 2A = 40, so A = 20 and B = 30. Contribution = 800 + 1500 = 2300.

Other vertices: A = 40, B = 0 gives 1600. A = 0, B = 40 gives 2000. So the LP optimum is at the integer point (20, 30) with value 2300.

Tie check: the objective ratio 40:50 differs from both constraint ratios (3:2 and 2:4), so the optimum is unique.

{"a": 20, "b": 30, "contribution": 2300}

$0.0032 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-capacity · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

Maximize 40A + 50B subject to 3A + 2B ≤ 120, 2A + 4B ≤ 160, with A and B whole numbers.

Where the two constraints meet: from the second, A + 2B ≤ 80. Subtracting that from the first gives 2A ≤ 40 (since 3A + 2B − (A + 2B) = 2A ≤ 120 − 80), so A = 20 and B = 30. Contribution there is 800 + 1500 = 2300.

Check the other vertices:
- A = 40, B = 0 gives 1600.
- A = 0, B = 40 gives 2000.

The vertex (20, 30) is integer and best among the vertices, so it is optimal for the LP and therefore for the integer problem.

Check for ties: the objective ratio 40:50 is not parallel to either constraint (3:2 or 2:4 = 1:2). The objective 4:5 differs from both, so the optimum is unique.

{"a": 20, "b": 30, "contribution": 2300}

$0.0041 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-capacity · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

Maximize 40A + 50B subject to 3A + 2B ≤ 120, 2A + 4B ≤ 160, with A and B whole numbers.

The LP vertex is where both constraints bind. From 3A + 2B = 120 and A + 2B = 80, subtracting gives 2A = 40, so A = 20 and B = 30. The contribution there is 800 + 1500 = 2300.

The other vertices are worse:
- (40, 0) gives 1600.
- (0, 40) gives 2000.

The vertex (20, 30) is integer, so it is the integer optimum. Check feasibility: labour is 60 + 60 = 120 and machine is 40 + 120 = 160, both within limits.

Ties: the objective 40A + 50B has slope ratio 4:5, while the constraint slopes are 3:2 and 1:2, so there are no ties along an edge.

{"a":20,"b":30,"contribution":2300}

$0.0038 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-capacity · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"a": 20, "b": 30, "contribution": 2300}

$0.0033 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-capacity · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"a": 20, "b": 30, "contribution": 2300}

$0.0028 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-capacity · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"a": 20, "b": 30, "contribution": 2300}

$0.0022 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-capacity · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0015 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-capacity · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0010 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-capacity · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0015 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-discount · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

Price after 20% off: 250 × 0.8 = 200.
After a further 15% off: 200 × 0.85 = 170.
After 8% tax: 170 × 1.08 = 183.6.
Effective pretax discount: (250 − 170)/250 = 32%.

{"discount_percent": 32, "final_price": 183.6}

$0.0016 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-discount · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

250 × 0.8 = 200; × 0.85 = 170; × 1.08 = 183.6.

Effective pretax discount: (250 − 170)/250 = 32%.

{"discount_percent": 32, "final_price": 183.6}

$0.0013 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-discount · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

Price after the first discount: 250 × 0.8 = 200.
Price after the second discount: 200 × 0.85 = 170.
Price after tax: 170 × 1.08 = 183.6.
Effective pretax discount: (250 − 170)/250 = 32%.

{"discount_percent": 32, "final_price": 183.6}

$0.0016 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-discount · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"discount_percent": 32, "final_price": 183.6}

$0.0016 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-discount · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"discount_percent": 32, "final_price": 183.6}

$0.0023 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-discount · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"discount_percent": 32, "final_price": 183.6}

$0.0018 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-discount · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"discount_percent":32,"final_price":183.6}

$0.0008 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-discount · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"discount_percent":32,"final_price":183.6}

$0.0008 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-discount · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"discount_percent":32,"final_price":183.6}

$0.0007 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-loan · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"interest":1557.1600003,"payment":564.88166668}

$0.0107 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-loan · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

{"interest":1557.1597885,"payment":564.88165786}

$0.0114 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-loan · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"interest":1557.1600004,"payment":564.88166668}

$0.0144 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-loan · attempt 1 · Failed

Gemini Developer API, paid Standard · medium reasoning

{"interest": 1557.159879961185, "payment": 564.8816616650494}

$0.0090 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-loan · attempt 2 · Failed

Gemini Developer API, paid Standard · medium reasoning

{"interest": 1557.1598794149, "payment": 564.8816616423}

$0.0247 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-loan · attempt 3 · Failed

Gemini Developer API, paid Standard · medium reasoning

{"interest": 1557.15988296969, "payment": 564.8816617904038}

$0.0072 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-loan · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"interest":1557.160000300248,"payment":564.881666679177}

$0.0050 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-loan · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"interest":1557.160000300236,"payment":564.8816666791765}

$0.0042 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-loan · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"interest":1557.1600003002552,"payment":564.8816666791773}

$0.0051 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-margin · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

Revenue: 210×30 = 6300; 30×24 = 720; total 7020.

Cost of goods: 240×18 = 4320.

Fees: 3% of 7020 = 210.6; per item fee 0.20×240 sold items = 48 (all 240 are sold). Total fees = 258.6.

Shipping: 2.50×240 = 600.

Profit = 7020 − 4320 − 258.6 − 600 = 1841.4.

Margin = 1841.4/7020 = 0.262307692...×100 = 26.2307692307...%

Check: 7020×0.2623 = 1841.346; remainder 0.054/7020 = 0.00000769, so 0.26230769. Margin is 26.230769230769...% (repeating 230769).

{"margin_percent": 26.230769230769, "profit": 1841.4, "revenue": 7020}

$0.0035 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-margin · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

Revenue: 210×30 = 6300; 30×24 = 720; total 7020.

Cost of goods: 240×18 = 4320.

Platform fees: 3% of 7020 = 210.60; plus 0.20×240 sold items = 48 (all 240 are sold). Total 258.60.

Shipping: 2.50×240 = 600.

Profit: 7020 − 4320 − 258.60 − 600 = 1841.40.

Margin: 1841.40/7020 = 0.262307692307... → 26.2307692307...%

Check: 7020×0.2623 = 1841.346; the remainder is 0.054/7020 = 0.0000076923, so 0.2623076923 and the margin is 26.23076923076923%.

{"margin_percent": 26.230769230769, "profit": 1841.4, "revenue": 7020}

$0.0034 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-margin · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

Revenue: 210×30 = 6300; 30×24 = 720; total 7020.

Cost of goods: 240×18 = 4320.

Fees: 3% × 7020 = 210.6; per item 0.20×240 = 48 (sold items: 210+30 = 240, so all items are sold).

Shipping: 2.50×240 = 600.

Profit = 7020 − 4320 − 210.6 − 48 − 600 = 1841.4.

Margin = 1841.4/7020 × 100 = 26.230769230...% (1841.4/7020 = 0.26230769...; check: 7020×0.2623 = 1841.346, remainder 0.054/7020 = 0.0000076923, so 0.2623076923). Margin percent = 26.23076923076923.

{"margin_percent": 26.23076923076923, "profit": 1841.4, "revenue": 7020}

$0.0035 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-margin · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"margin_percent": 26.23076923076923, "profit": 1841.4, "revenue": 7020}

$0.0045 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-margin · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"margin_percent": 26.23076923076923, "profit": 1841.4, "revenue": 7020}

$0.0038 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-margin · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"margin_percent": 26.23076923076923, "profit": 1841.4, "revenue": 7020}

$0.0055 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-margin · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"margin_percent":26.23076923076923,"profit":1841.4,"revenue":7020}

$0.0014 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-margin · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"margin_percent":26.23076923076923,"profit":1841.4,"revenue":7020}

$0.0018 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-margin · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"margin_percent":26.23076923076923,"profit":1841.4,"revenue":7020}

$0.0015 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-mix · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

Let x be litres of the 20% stock and y litres of the 60% stock.

x + y = 80 and 0.20x + 0.60y = 0.35 × 80 = 28.

Substituting x = 80 − y: 16 + 0.40y = 28, so y = 30 and x = 50.

{"litres_20": 50, "litres_60": 30}

$0.0016 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-mix · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

Let x be the litres of the 20% stock and y the litres of the 60% stock.

Volume: x + y = 80.
Concentrate: 0.20x + 0.60y = 0.35 × 80 = 28.

Substitute x = 80 − y: 0.20(80 − y) + 0.60y = 28, so 16 + 0.40y = 28, giving y = 30 and x = 50.

{"litres_20": 50, "litres_60": 30}

$0.0020 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-mix · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

Let x be the litres of the 20% stock and y the litres of the 60% stock.

Volume: x + y = 80.
Concentrate: 0.20x + 0.60y = 0.35 × 80 = 28.

Substitute x = 80 − y:
0.20(80 − y) + 0.60y = 28
16 + 0.40y = 28
y = 30, so x = 50.

{"litres_20": 50, "litres_60": 30}

$0.0019 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-mix · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0019 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-mix · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0014 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-mix · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0017 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-mix · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"litres_20":50,"litres_60":30}

$0.0006 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-mix · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"litres_20":50,"litres_60":30}

$0.0007 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-mix · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"litres_20":50,"litres_60":30}

$0.0009 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-npv · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

NPV = -1000 + 400 × (1/1.1 + 1/1.21 + 1/1.331)

The annuity factor is 1/1.1 + 1/1.21 + 1/1.331 = (1.21×1.1 + 1.1×1.1... let me compute over the common denominator 1.331:
1.21 + 1.1 + 1 = 3.31, so the factor is 3.31/1.331 = 2.486851990984...

Then 400 × 3.31/1.331 = 1324/1.331 = 994.7407963936...

Check the division: 1.331 × 994.74 = 1323.98..., since 1.331×994 = 1323.014 and 1.331×0.74 = 0.98494, the total is 1323.99894. The remainder is 1324 − 1323.99894 = 0.00106, and 0.00106/1.331 = 0.000796394. So the value is 994.740796394...

More digits: 0.00106/1.331 = 0.00079639368...
1.331 × 0.000796 = 0.001059476; the remainder is 0.000000524, and 0.000000524/1.331 = 3.9369e-7. So the value is 0.00079639369, and 1324/1.331 = 994.74079639369.

NPV = -5.25920360631

Check: 1000 − 994.74079639369 = 5.25920360631.

```json
{"npv": -5.25920360631}
```

$0.0051 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-npv · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

NPV = -1000 + 400 × (1/1.1 + 1/1.21 + 1/1.331)

The annuity factor is (1 − 1.1^-3)/0.1. Since 1.1^3 = 1.331, 1/1.331 = 0.7513148009015778.

The annuity factor is (1 − 0.7513148009015778)/0.1 = 2.486851990984222.

Multiplying by 400 gives 994.7407963936888.

So NPV = -5.2592036063112.

Exact check: 1/1.331 = 1000/1331. The annuity factor is 10 × (1 − 1000/1331) = 10 × 331/1331 = 3310/1331. Then 400 × 3310/1331 = 1324000/1331 = 994.74079639369...

Division check: 1331 × 994 = 1,323,014. The remainder is 986, and 986/1331 = 0.74079639...
- 1331 × 0.74 = 984.94, remainder 1.06
- 1.06/1331 = 0.000796393688...

So 994.740796393688, and NPV = -5.259203606312.

{"npv": -5.259203606312}

$0.0043 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-npv · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

NPV = -1000 + 400 × (1/1.1 + 1/1.21 + 1/1.331)

The annuity factor is 1/1.1 + 1/1.21 + 1/1.331 = (1.21×1.1 + 1.1×1.1 ... ) computed over the common denominator 1.331:
- 1/1.1 = 1.21/1.331
- 1/1.21 = 1.1/1.331
- 1/1.331 = 1/1.331

The sum is 3.31/1.331 = 2.486851990984222...

Check: 1.331 × 2.486851991 = 3.31000000. Good.

Multiplying by 400 gives 994.7407963936889, since 3.31 × 400 = 1324 and 1324/1.331 = 994.74079639368...

Long division of 1324/1.331 (that is, 1324000/1331):
- 1331 × 994 = 1,323,014, remainder 986.
- 9860/1331 = 7 (9317), remainder 543.
- 5430/1331 = 4 (5324), remainder 106.
- 1060/1331 = 0, remainder 1060.
- 10600/1331 = 7 (9317), remainder 1283.
- 12830/1331 = 9 (11979), remainder 851.
- 8510/1331 = 6 (7986), remainder 524.
- 5240/1331 = 3 (3993), remainder 1247.
- 12470/1331 = 9 (11979), remainder 491.
- 4910/1331 = 3 (3993), remainder 917.
- 9170/1331 = 6 (7986), remainder 1184.
- 11840/1331 = 8 (10648), remainder 1192.

This gives 994.74079639368..., so NPV = -5.259203606311...

Subtracting 1000 from 994.74079639368 gives -5.259203606320, using digits 994.740796393680. The digits so far are 994.74079639368, so NPV = -5.25920360632 (approximately).

```json
{"npv": -5.25920360632}
```

$0.0073 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-npv · attempt 1 · Failed

Gemini Developer API, paid Standard · medium reasoning

{"npv": -5.259196093163035}

$0.0045 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-npv · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"npv": -5.259203606311044}

$0.0042 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-npv · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"npv": -5.259203606311044}

$0.0040 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-npv · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"npv":-5.259203606311044}

$0.0018 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-npv · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"npv":-5.259203606311045}

$0.0016 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-npv · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"npv":-5.259203606311045}

$0.0018 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-queue · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

p = 0.98^3 = 0.941192
q = 1 - p = 0.058808

Success probability with one retry: 1 - q^2.
q^2 = 0.058808^2 = 0.003458620864 (check: 0.058808^2 = 0.058808 × 0.058808. 58808^2 = 58808×58808. 58808×58000 = 3,410,864,000; 58808×808 = 47,516,864; total 3,458,380,864. Hmm, recompute: 58808×58000 = 58808×58×1000 = 3,410,864×1000 = 3,410,864,000. 58808×808 = 58808×800 = 47,046,400 plus 58808×8 = 470,464, giving 47,516,864. Sum = 3,458,380,864.)
So q^2 = 3,458,380,864 × 10^-12 = 0.003458380864.

Success probability = 1 - 0.003458380864 = 0.996541619136.

Expected number of attempts = 1 + q = 1.058808 (the second attempt happens only if the first fails).

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0043 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-queue · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

p = 0.98^3 = 0.941192
q = 1 - p = 0.058808

Success probability with one retry: 1 - q^2 = p + q·p = 0.941192 + 0.058808·0.941192.
q^2 = 0.058808^2 = 0.003458380864 (check: 0.058808^2 = 0.0034583809 approx; exact: 58808^2 = 3,458,380,864, so q^2 = 3,458,380,864e-12 = 0.003458380864).
Success = 0.996541619136.

Expected number of complete attempts = 1 + q = 1.058808 (the second attempt happens only if the first fails).

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0030 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-queue · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

p = 0.98^3 = 0.941192
q = 1 - p = 0.058808

Success probability with one retry: 1 - q^2.
q^2 = 0.058808^2 = 0.003458620864 (check: 0.058808^2 = 58808^2 × 10^-12; 58808^2 = 58808×58808 = 58808×58000 = 3,410,864,000; 58808×808 = 47,516,864; total 3,458,380,864. So q^2 = 0.003458380864.)

Recheck: 58808×58000 = 58808×58 ×1000 = 3,410,864×1000 = 3,410,864,000. Correct. 58808×808 = 58808×800 = 47,046,400, plus 58808×8 = 470,464, total 47,516,864. Sum = 3,458,380,864. So q^2 = 0.003458380864.

Success probability = 1 - 0.003458380864 = 0.996541619136.

Expected number of attempts = 1 + q = 1.058808 (the second attempt happens only if the first fails).

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0043 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-queue · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0044 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-queue · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0037 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-queue · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0039 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-queue · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"expected_attempts":1.058808,"success_probability":0.996541619136}

$0.0013 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-queue · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"expected_attempts":1.058808,"success_probability":0.996541619136}

$0.0013 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-queue · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"expected_attempts":1.058808,"success_probability":0.996541619136}

$0.0013 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-variance · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

Mean = (4+7+7+10+12)/5 = 40/5 = 8.

Deviations: -4, -1, -1, 2, 4. Squares: 16, 1, 1, 4, 16, sum = 38.

Sample variance = 38/4 = 9.5.

Standard error = sqrt(9.5/5) = sqrt(1.9) ≈ 1.378404875209022.

Check: 1.3784^2 = 1.89998... and 1.378404875^2 ≈ 1.9. Good.

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

$0.0026 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-variance · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

Mean = (4+7+7+10+12)/5 = 40/5 = 8.

The squared deviations from the mean are 16, 1, 1, 4, 16, which sum to 38.

Sample variance = 38/4 = 9.5.

Standard error = sqrt(9.5/5) = sqrt(1.9) ≈ 1.378404875209022.

Check: 1.3784² = 1.89998..., which is consistent with sqrt(1.9). The more precise value is 1.37840487520902.

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

$0.0025 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-variance · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

Mean = (4+7+7+10+12)/5 = 40/5 = 8.

Deviations: -4, -1, -1, 2, 4. Squares: 16, 1, 1, 4, 16. Sum = 38.

Sample variance = 38/4 = 9.5.

Standard error = sqrt(9.5/5) = sqrt(1.9) ≈ 1.378404875209022.

Check: 1.3784² = 1.89998656; the remainder is 0.00001344, and dividing by 2·1.3784 gives 0.00000487, so the value is about 1.37840487. This agrees with the figure above.

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

$0.0029 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-variance · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.3784048752090222}

$0.0021 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-variance · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

$0.0019 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-variance · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

$0.0019 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-variance · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209022}

$0.0011 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-variance · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209022}

$0.0010 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-variance · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209022}

$0.0012 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-weighted · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

Week 1: total visits are 900 + 100 = 1000 and total purchases are 45 + 20 = 65, so conversion is 6.5%.

Week 2: total visits are 100 + 900 = 1000 and total purchases are 8 + 162 = 170, so conversion is 17%.

Change: 17 − 6.5 = 10.5 percentage points.

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0021 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-weighted · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

Week 1: total visits = 900 + 100 = 1000 and total purchases = 45 + 20 = 65, so the conversion rate is 6.5%.

Week 2: total visits = 100 + 900 = 1000 and total purchases = 8 + 162 = 170, so the conversion rate is 17%.

Change = 17 − 6.5 = 10.5 percentage points.

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0022 estimated; billed reconciliation remains separate.

claude-sonnet-5-5 · math-weighted · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

Week 1: total visits = 900 + 100 = 1000, total purchases = 45 + 20 = 65, so the conversion rate is 6.5%.

Week 2: total visits = 100 + 900 = 1000, total purchases = 8 + 162 = 170, so the conversion rate is 17%.

Change: 17 − 6.5 = 10.5 percentage points.

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0022 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-weighted · attempt 1 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17.0}

$0.0024 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-weighted · attempt 2 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17.0}

$0.0025 estimated; billed reconciliation remains separate.

gemini-3.8-flash · math-weighted · attempt 3 · Passed

Gemini Developer API, paid Standard · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17.0}

$0.0026 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-weighted · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"change_points":10.5,"week1_percent":6.5,"week2_percent":17}

$0.0009 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-weighted · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"change_points":10.5,"week1_percent":6.5,"week2_percent":17}

$0.0009 estimated; billed reconciliation remains separate.

gpt-6.1-sol · math-weighted · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"change_points":10.5,"week1_percent":6.5,"week2_percent":17}

$0.0009 estimated; billed reconciliation remains separate.

Run 2026-10-07T05:51:59.343583+00:00 · tested-best-bank-v2-reviewed · complete outputs
gpt-6-luna · math-allocation · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-allocation · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-allocation · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-bayes · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"probability":0.32646048109965636}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-bayes · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"probability":0.32646048109965636}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-bayes · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"probability":0.32646048109965636}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-break-even · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"profit":0,"units":100}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-break-even · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"profit":0,"units":100}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-break-even · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"profit":0,"units":100}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-capacity · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0002 estimated; billed reconciliation remains separate.

gpt-6-luna · math-capacity · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-capacity · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-discount · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"discount_percent":32,"final_price":183.6}

$0.0000 estimated; billed reconciliation remains separate.

gpt-6-luna · math-discount · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"discount_percent":32,"final_price":183.6}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-discount · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"discount_percent":32,"final_price":183.6}

$0.0000 estimated; billed reconciliation remains separate.

gpt-6-luna · math-loan · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"interest":1557.1600003,"payment":564.8816666792}

$0.0011 estimated; billed reconciliation remains separate.

gpt-6-luna · math-loan · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"interest":1557.16000065,"payment":564.881666694}

$0.0008 estimated; billed reconciliation remains separate.

gpt-6-luna · math-loan · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"interest":1557.16000030,"payment":564.881666679}

$0.0009 estimated; billed reconciliation remains separate.

gpt-6-luna · math-margin · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"margin_percent":26.23076923076923,"profit":1841.4,"revenue":7020}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-margin · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"margin_percent":26.2307692308,"profit":1841.4,"revenue":7020}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-margin · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"margin_percent":26.23076923076923,"profit":1841.4,"revenue":7020}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-mix · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"litres_20":50,"litres_60":30}

$0.0000 estimated; billed reconciliation remains separate.

gpt-6-luna · math-mix · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"litres_20":50,"litres_60":30}

$0.0000 estimated; billed reconciliation remains separate.

gpt-6-luna · math-mix · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"litres_20":50,"litres_60":30}

$0.0000 estimated; billed reconciliation remains separate.

gpt-6-luna · math-npv · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"npv":-5.259203606311044}

$0.0002 estimated; billed reconciliation remains separate.

gpt-6-luna · math-npv · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"npv":-5.25920360631}

$0.0003 estimated; billed reconciliation remains separate.

gpt-6-luna · math-npv · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"npv":-5.25920360631104432}

$0.0004 estimated; billed reconciliation remains separate.

gpt-6-luna · math-queue · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"expected_attempts":1.058808,"success_probability":0.996541619136}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-queue · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"expected_attempts":1.058808,"success_probability":0.996541619136}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-queue · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"expected_attempts":1.058808,"success_probability":0.996541619136}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-variance · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-variance · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209022}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-variance · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-weighted · attempt 1 · Passed

OpenAI direct API, Standard global · medium reasoning

{"change_points":10.5,"week1_percent":6.5,"week2_percent":17}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-weighted · attempt 2 · Passed

OpenAI direct API, Standard global · medium reasoning

{"change_points":10.5,"week1_percent":6.5,"week2_percent":17}

$0.0001 estimated; billed reconciliation remains separate.

gpt-6-luna · math-weighted · attempt 3 · Passed

OpenAI direct API, Standard global · medium reasoning

{"change_points":10.5,"week1_percent":6.5,"week2_percent":17}

$0.0001 estimated; billed reconciliation remains separate.

Run 2026-10-07T06:00:52.001925+00:00 · tested-best-bank-v2-reviewed · complete outputs
claude-opus-5-5 · math-allocation · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0029 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-allocation · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

1000 cents divided by 3 gives 333.333 each. Flooring gives 333 to each department, a total of 999, so 1 cent remains. All three fractional remainders are equal, so the tie goes to the lowest index, department A.

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0033 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-allocation · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"a_cents":334,"b_cents":333,"c_cents":333}

$0.0029 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-bayes · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"probability": 0.3264604810996564}

$0.0065 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-bayes · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"probability":0.3264604810996564}

$0.0069 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-bayes · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"probability":0.3264604810996564}

$0.0051 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-break-even · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"profit": 0, "units": 100}

$0.0033 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-break-even · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"profit": 0, "units": 100}

$0.0027 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-break-even · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"profit": 0, "units": 100}

$0.0033 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-capacity · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0059 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-capacity · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0072 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-capacity · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"a":20,"b":30,"contribution":2300}

$0.0056 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-discount · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"discount_percent": 32, "final_price": 183.6}

$0.0023 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-discount · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

250*0.8=200, *0.85=170, *1.08=183.6. Discount: 1-0.68=0.32 → 32%.

{"discount_percent":32,"final_price":183.6}

$0.0023 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-discount · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

250*0.8=200, *0.85=170, *1.08=183.6. Discount: 1-170/250=0.32 → 32%.

```json
{"discount_percent": 32, "final_price": 183.6}
```

Wait—no Markdown allowed.

{"discount_percent": 32, "final_price": 183.6}

$0.0031 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-loan · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"payment": 564.8816666792, "interest": 1557.1600003008}

$0.0338 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-loan · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"payment": 564.8816666795, "interest": 1557.160000308}

$0.0490 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-loan · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"payment": 564.8816666791765, "interest": 1557.160000300236}

$0.0532 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-margin · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"margin_percent": 26.230769230769, "profit": 1841.4, "revenue": 7020}

$0.0066 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-margin · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"margin_percent": 26.23076923076923, "profit": 1841.4, "revenue": 7020}

$0.0058 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-margin · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"margin_percent": 26.230769230769231, "profit": 1841.4, "revenue": 7020}

$0.0060 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-mix · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0023 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-mix · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0023 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-mix · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0025 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-npv · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"npv": -5.259203606311044}

$0.0099 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-npv · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"npv": -5.259203606311045}

$0.0103 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-npv · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"npv": -5.2592036063111}

$0.0116 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-queue · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0050 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-queue · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0052 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-queue · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"expected_attempts":1.058808,"success_probability":0.996541619136}

$0.0050 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-variance · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

$0.0068 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-variance · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.3784048752090222}

$0.0067 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-variance · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

$0.0135 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-weighted · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0029 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-weighted · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0029 estimated; billed reconciliation remains separate.

claude-opus-5-5 · math-weighted · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0029 estimated; billed reconciliation remains separate.

Run 2026-10-07T23:26:26.061111+00:00 · tested-best-bank-v2-reviewed · complete outputs
claude-haiku-5-5 · math-allocation · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-allocation · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-allocation · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"a_cents": 334, "b_cents": 333, "c_cents": 333}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-bayes · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

{"probability": 0.32758620689655172}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-bayes · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

{"probability": 0.3275862068965517}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-bayes · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"probability": 0.3264604811}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-break-even · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"profit": 0, "units": 100}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-break-even · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"profit": 0, "units": 100}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-break-even · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"profit": 0, "units": 100}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-capacity · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"a": 20, "b": 30, "contribution": 2300}

$0.0003 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-capacity · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"a": 20, "b": 30, "contribution": 2300}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-capacity · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"a": 20, "b": 30, "contribution": 2300}

$0.0003 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-discount · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"discount_percent": 32, "final_price": 183.6}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-discount · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"discount_percent": 32, "final_price": 183.6}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-discount · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"discount_percent": 32, "final_price": 183.6}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-loan · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"interest": 1557.16, "payment": 564.8816667}
```

$0.0015 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-loan · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"interest": 1557.1600018, "payment": 564.88166674}
```

$0.0021 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-loan · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"interest": 1557.16000003, "payment": 564.88166667919}
```

$0.0017 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-margin · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"margin_percent": 26.23076923076923, "profit": 1841.4, "revenue": 7020}

$0.0003 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-margin · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"margin_percent": 26.23076923076923, "profit": 1841.4, "revenue": 7020}

$0.0003 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-margin · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"margin_percent": 26.230769230769, "profit": 1841.4, "revenue": 7020}

$0.0003 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-mix · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-mix · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-mix · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"litres_20": 50, "litres_60": 30}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-npv · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"npv": -5.259203606311045}
```

$0.0003 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-npv · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"npv": -5.259203606311044}
```

$0.0004 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-npv · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"npv": -5.259203606311044}
```

$0.0006 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-queue · attempt 1 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"expected_attempts": 1.058808, "success_probability": 0.996541619136}
```

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-queue · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"expected_attempts": 1.058808, "success_probability": 0.996541619136}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-queue · attempt 3 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"expected_attempts": 1.058808, "success_probability": 0.996541619136}
```

$0.0003 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-variance · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.3784048752090222}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-variance · attempt 2 · Failed

Claude direct API, Standard global · medium reasoning

```json
{"mean": 8, "sample_variance": 9.5, "standard_error": 1.3784048752090222}
```

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-variance · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.3784048752090222}

$0.0002 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-weighted · attempt 1 · Passed

Claude direct API, Standard global · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-weighted · attempt 2 · Passed

Claude direct API, Standard global · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0001 estimated; billed reconciliation remains separate.

claude-haiku-5-5 · math-weighted · attempt 3 · Passed

Claude direct API, Standard global · medium reasoning

{"change_points": 10.5, "week1_percent": 6.5, "week2_percent": 17}

$0.0001 estimated; billed reconciliation remains separate.

Run history

  • 2026-10-04T08:20:14.082456+00:00 · Complete archived study (tested-best-bank-v1) · 108 / 108 required trials initiated · Archived evidence
  • 2026-10-05T05:41:13.124828+00:00 · complete · 108 / 108 required trials initiated · Archived evidence
  • 2026-10-07T05:45:33.170758+00:00 · complete · 36 / 36 required trials initiated · Archived evidence
  • 2026-10-07T05:52:07.718463+00:00 · complete · 36 / 36 required trials initiated · Archived evidence
  • 2026-10-07T19:33:27.474548+00:00 · complete · 36 / 36 required trials initiated · Archived evidence

Maintenance needs recovery: 3 checks failed on the latest attempt. Published test dates remain unchanged.

Source-check status: A source fetch has completed; factual changes require review. Latest attempt: 2026-10-08T00:41:23.352024+00:00.

Download the public task bank · Read the methodology