AI News

AI math comparison tests and results

Measured results

Compare AI models on applied math tasks

This bank measures numerical correctness on twelve applied problems. It does not score the quality of explanations.

Last tested
2026-10-04T17:01:24.616847+00:00

Last successful source fetch
2026-10-04T09:01:57.926856+00:00

Last published
2026-10-04T18:18:57.482142+00:00

Observed leaders

gpt-6.1-sol, gemini-3.8-flash

Descriptive pass counts after offline runner correction of the original outputs. Task ambiguities described in the audit limit broader conclusions.

Tested API configurations

ModelTask passesCost per passProcessing
GPT-6.1 Sol
gpt-6.1-sol
30 / 36$0.0013 (estimated)OpenAI direct API, Standard global
Medium reasoning; 8,192-token cap
Claude Sonnet 5.5
claude-sonnet-5-5
3 / 36$0.0270 (estimated)Claude direct API, Standard global
Medium reasoning; 8,192-token cap
Gemini 3.8 Flash
gemini-3.8-flash
30 / 36$0.0033 (estimated)Gemini Developer API, paid Standard
Medium reasoning; 8,192-token cap

Three documented current direct-API models from different providers, selected for coding/math use and practical cost. Not a claim that every leading tool is included. Account access and publication terms require verification before dispatch. No web, tools, prompt caching, batch or regional processing. Medium reasoning labels are provider-specific, not equal compute.

Scores by task group

ModelGroupPassesMedian latency
gpt-6.1-solFinance12 / 182.96 seconds (whole model)
gpt-6.1-solOptimization6 / 62.96 seconds (whole model)
gpt-6.1-solProbability6 / 62.96 seconds (whole model)
gpt-6.1-solStatistics6 / 62.96 seconds (whole model)
claude-sonnet-5-5Finance3 / 181.99 seconds (whole model)
claude-sonnet-5-5Optimization0 / 61.99 seconds (whole model)
claude-sonnet-5-5Probability0 / 61.99 seconds (whole model)
claude-sonnet-5-5Statistics0 / 61.99 seconds (whole model)
gemini-3.8-flashFinance12 / 182.56 seconds (whole model)
gemini-3.8-flashOptimization6 / 62.56 seconds (whole model)
gemini-3.8-flashProbability6 / 62.56 seconds (whole model)
gemini-3.8-flashStatistics6 / 62.56 seconds (whole model)
Historical and current-price estimates

gpt-6.1-sol: original run $0.0394; same recorded usage at rates reviewed 2026-10-04T06:52:40.966997+00:00: $0.0394. Both are estimates; billing reconciliation is separate.

claude-sonnet-5-5: original run $0.0811; same recorded usage at rates reviewed 2026-10-04T06:52:40.966997+00:00: $0.0811. Both are estimates; billing reconciliation is separate.

gemini-3.8-flash: original run $0.0995; same recorded usage at rates reviewed 2026-10-04T06:52:40.966997+00:00: $0.0995. Both are estimates; billing reconciliation is separate.

The 12 tasks

Open a task to inspect its exact prompt and, after testing, every model output. Hidden coding fixtures and answer keys are not served to the models or this page.

Outputs are generated by the named AI models and may contain errors. Scores cover the declared task checks; they do not verify every statement in an output.

Separate markup from margin Finance

A shop buys 240 items at $18 each. It sells 210 at $30 and 30 at $24. Platform fees are 3% of revenue plus $0.20 per sold item. Shipping costs $2.50 per sold item. Calculate revenue, profit after these costs, and profit margin as a percentage of revenue. Return only a JSON object with these keys: revenue, profit, margin_percent. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Failed
{"revenue":7020,"profit":1841.4,"margin_percent":26.23}

4.44 seconds · $0.0014 estimated

Attempt 2 · Failed
{"revenue":7020,"profit":1841.4,"margin_percent":26.23}

4.83 seconds · $0.0015 estimated

Attempt 3 · Failed
{"revenue":7020,"profit":1841.4,"margin_percent":26.23}

4.06 seconds · $0.0011 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
Revenue: 210×30 = 6300; 30×24 = 720; total 7020.
Cost of goods: 240×18 = 4320.
Platform fees: 0.03×7020 = 210.6; per-item fee: 240×0.20 = 48; total 258.6.
Shipping: 240×2.50 = 600.
Profit: 7020 − 4320 − 258.6 − 600 = 1841.4.
Margin: 1841.4/7020 = 0.262308 → 26.23%.

{"revenue": 7020, "profit": 1841.4, "margin_percent": 26.23}

2.24 seconds · $0.0023 estimated

Attempt 2 · Failed
Revenue: 210×30 = 6300; 30×24 = 720; total 7020.
Cost of goods: 240×18 = 4320.
Platform fees: 0.03×7020 = 210.60; per item 0.20×240 = 48; total 258.60.
Shipping: 2.50×240 = 600.
Profit: 7020 − 4320 − 258.60 − 600 = 1841.40.
Margin: 1841.40/7020 = 0.262308 → 26.23%.

{"revenue": 7020, "profit": 1841.4, "margin_percent": 26.23}

2.37 seconds · $0.0023 estimated

Attempt 3 · Failed
Revenue: 210×30 = 6300; 30×24 = 720; total 7020.
Cost of goods: 240×18 = 4320.
Platform fees: 0.03×7020 = 210.6; per item: 240×0.20 = 48; total fees 258.6.
Shipping: 240×2.50 = 600.
Profit: 7020 − 4320 − 258.6 − 600 = 1841.4.
Margin: 1841.4/7020 = 0.262308 → 26.23%.

{"revenue": 7020, "profit": 1841.4, "margin_percent": 26.23}

2.21 seconds · $0.0023 estimated

Gemini 3.8 Flash

Attempt 1 · Failed
{"revenue": 7020, "profit": 1841.4, "margin_percent": 26.23}

3.63 seconds · $0.0035 estimated

Attempt 2 · Failed
{"revenue": 7020, "profit": 1841.4, "margin_percent": 26.23}

3.42 seconds · $0.0031 estimated

Attempt 3 · Failed
{"revenue": 7020, "profit": 1841.4, "margin_percent": 26.23}

4.40 seconds · $0.0042 estimated

Solve a production mix Optimization

Make exactly 80 litres of a 35% concentrate from 20% and 60% stocks, with additive volumes. How many litres of each stock are needed? Return only a JSON object with these keys: litres_20, litres_60. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"litres_20":50,"litres_60":30}

2.92 seconds · $0.0006 estimated

Attempt 2 · Passed
{"litres_20":50,"litres_60":30}

2.50 seconds · $0.0007 estimated

Attempt 3 · Passed
{"litres_20":50,"litres_60":30}

2.52 seconds · $0.0006 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
Let x be litres of the 20% stock and 80 − x litres of the 60% stock.

0.20x + 0.60(80 − x) = 0.35 × 80 = 28

0.20x + 48 − 0.60x = 28, so 0.40x = 20 and x = 50.

That gives 50 litres of the 20% stock and 30 litres of the 60% stock.

{"litres_20": 50, "litres_60": 30}

2.07 seconds · $0.0017 estimated

Attempt 2 · Failed
Let x be the litres of the 20% stock. Then x + y = 80 and 0.2x + 0.6y = 0.35·80 = 28.

Substituting y = 80 − x: 0.2x + 48 − 0.6x = 28, so −0.4x = −20 and x = 50. Then y = 30.

{"litres_20": 50, "litres_60": 30}

1.83 seconds · $0.0015 estimated

Attempt 3 · Failed
Let x be litres of the 20% stock and 80 − x litres of the 60% stock.

0.20x + 0.60(80 − x) = 0.35 × 80
0.20x + 48 − 0.60x = 28
−0.40x = −20
x = 50

So 50 litres of the 20% stock and 30 litres of the 60% stock.

{"litres_20": 50, "litres_60": 30}

1.92 seconds · $0.0016 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"litres_20": 50, "litres_60": 30}

1.95 seconds · $0.0014 estimated

Attempt 2 · Passed
{"litres_20": 50, "litres_60": 30}

1.93 seconds · $0.0013 estimated

Attempt 3 · Passed
{"litres_20": 50, "litres_60": 30}

1.88 seconds · $0.0014 estimated

Aggregate conversion rates correctly Statistics

Campaign A has 900 visits and 45 purchases. Campaign B has 100 visits and 20 purchases. Next week A has 100 visits and 8 purchases; B has 900 visits and 162 purchases. Return each week’s total conversion percentage and next-minus-first percentage-point change. Return only a JSON object with these keys: week1_percent, week2_percent, change_points. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"week1_percent":6.5,"week2_percent":17,"change_points":10.5}

2.38 seconds · $0.0008 estimated

Attempt 2 · Passed
{"week1_percent":6.5,"week2_percent":17,"change_points":10.5}

2.92 seconds · $0.0008 estimated

Attempt 3 · Passed
{"week1_percent":6.5,"week2_percent":17,"change_points":10.5}

3.62 seconds · $0.0008 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
Week 1: total visits 1000, purchases 65, so 6.5%.
Week 2: total visits 1000, purchases 170, so 17%.
Change: 17 - 6.5 = 10.5 percentage points.

{"week1_percent": 6.5, "week2_percent": 17, "change_points": 10.5}

1.74 seconds · $0.0014 estimated

Attempt 2 · Failed
Week 1: total visits 1000, purchases 65, so 6.5%.
Week 2: total visits 1000, purchases 170, so 17%.
Change: 10.5 percentage points.

{"week1_percent": 6.5, "week2_percent": 17, "change_points": 10.5}

1.70 seconds · $0.0013 estimated

Attempt 3 · Failed
Week 1: total visits = 1000, purchases = 65, so 6.5%.
Week 2: total visits = 1000, purchases = 170, so 17%.
Change: 17 − 6.5 = 10.5 percentage points.

{"week1_percent": 6.5, "week2_percent": 17, "change_points": 10.5}

1.81 seconds · $0.0014 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"week1_percent": 6.5, "week2_percent": 17.0, "change_points": 10.5}

3.09 seconds · $0.0013 estimated

Attempt 2 · Passed
{"week1_percent": 6.5, "week2_percent": 17.0, "change_points": 10.5}

1.61 seconds · $0.0012 estimated

Attempt 3 · Passed
{"week1_percent": 6.5, "week2_percent": 17, "change_points": 10.5}

1.64 seconds · $0.0014 estimated

Interpret a screening result Probability

A defect affects 2% of units. A detector flags 95% of defective units and 4% of nondefective units. Given a flag, calculate the probability the unit is defective, as a decimal from 0 to 1. Return only a JSON object with these keys: probability. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"probability":0.32646048109965636}

2.71 seconds · $0.0007 estimated

Attempt 2 · Passed
{"probability":0.32646048109965636}

2.81 seconds · $0.0007 estimated

Attempt 3 · Passed
{"probability":0.32646048109965636}

2.44 seconds · $0.0006 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
P(defective | flag) = (0.02 × 0.95) / (0.02 × 0.95 + 0.98 × 0.04)

The numerator is 0.019 and the second term in the denominator is 0.0392, so the denominator is 0.0582.

0.019 / 0.0582 = 0.326460...

{"probability": 0.3265}

1.77 seconds · $0.0014 estimated

Attempt 2 · Failed
P(D)=0.02, P(F|D)=0.95, P(F|not D)=0.04.

Numerator: 0.02*0.95=0.019
Denominator: 0.019+0.98*0.04=0.019+0.0392=0.0582
Probability = 0.019/0.0582 = 0.326460...

{"probability": 0.32646}

1.84 seconds · $0.0014 estimated

Attempt 3 · Failed
P(D)=0.02, P(F|D)=0.95, P(F|not D)=0.04.
Numerator: 0.02*0.95=0.019
Denominator: 0.019+0.98*0.04=0.019+0.0392=0.0582
Probability=0.019/0.0582=0.326460...

0.0582*0.3265=0.01900; so ≈0.32646.

{"probability": 0.32646}

1.96 seconds · $0.0016 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"probability": 0.32646048109965636}

2.47 seconds · $0.0018 estimated

Attempt 2 · Passed
{"probability": 0.32646048109965635}

2.47 seconds · $0.0016 estimated

Attempt 3 · Passed
{"probability": 0.32646048109965636}

2.23 seconds · $0.0016 estimated

Calculate an amortizing payment Finance

A $12000 loan charges a nominal annual rate of 12%, compounded monthly, repaid with 24 equal end-of-month payments. Calculate the monthly payment and total interest, without intermediate rounding. Return only a JSON object with these keys: payment, interest. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Failed
{"payment":564.88,"interest":1557.16}

6.37 seconds · $0.0019 estimated

Attempt 2 · Failed
{"payment":564.88,"interest":1557.16}

12.10 seconds · $0.0042 estimated

Attempt 3 · Failed
{"payment":564.88,"interest":1557.16}

8.04 seconds · $0.0021 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
Monthly rate i = 0.01, n = 24.

Payment = 12000 · i / (1 − (1+i)^−24).

1.01^24 = 1.2697346485 (known value: 1.269734649).
1/1.269734649 = 0.787566 (more precisely 0.78756606).
1 − 0.78756606 = 0.21243394.
Payment = 120 / 0.21243394 = 564.8672...

Check: 0.21243394 × 565 = 120.02512, so the excess is 0.02512 / 0.21243394 = 0.11825. Payment = 565 − 0.11825 = 564.88175.

More precisely, the known payment factor for 1%, 24 months is 0.0470734722, so 12000 × 0.0470734722 = 564.8816664.

Total paid = 564.8816664 × 24 = 13557.16 (564.8816664 × 24 = 13557.1599936).
Interest = 1557.16 (about 1557.1599936).

{"payment": 564.8816664, "interest": 1557.15999}

4.07 seconds · $0.0038 estimated

Attempt 2 · Failed
Monthly rate i = 0.01, n = 24.

Payment = 12000 · 0.01 / (1 − 1.01^-24).

1.01^24 = 1.269734648531914, so 1.01^-24 = 0.787566178...

Check: 1/1.269734648531914 ≈ 0.78756618. Then 1 − 0.78756618 = 0.21243382.

Payment = 120 / 0.21243382 ≈ 564.8707.

More precisely: 0.21243382 × 565 = 120.0241, which is 0.0241 too high; 0.0241/0.21243 = 0.1134, so payment ≈ 564.8866. Recompute: 0.21243382 × 564.8866 = 0.21243382×564 = 119.81 (0.21243382×500 = 106.21691, ×64 = 13.59576, sum 119.81267), plus 0.21243382×0.8866 = 0.18835, total 120.00102. Slightly high by 0.00102, which is 0.0048 in payment. So payment ≈ 564.8818.

Refine 1.01^-24 more precisely: 1/1.269734648531914. 1.269734648531914 × 0.7875 = 0.99991 (1.269734648×0.7875 = 0.999915...). Compute: 1.269734648×0.78 = 0.990393025; ×0.0075 = 0.009523010; sum 0.999916035. Remainder 0.000083965/1.269734648 = 0.00006613. So the inverse is 0.78756613. Then 1 − that = 0.21243387.

Payment = 120/0.21243387. 0.21243387×564.88 = 0.21243387×564 = 119.812 (0.21243387×564: ×500 = 106.216935, ×64 = 13.595768, sum 119.812703), plus ×0.88 = 0.186941, total 119.999644. Remainder 0.000356/0.21243 = 0.001676. Payment ≈ 564.881676.

The known value is 564.8734? Check against the standard: the payment on 12000 at 1% for 24 months is 564.87 (the factor 0.0470735 × 12000 = 564.88). The factor 0.04707347 × 12000 = 564.8816. Good.

Total interest = 24 × 564.8817 − 12000 = 13557.16 − 12000 = 1557.16.

24 × 564.881676 = 13557.160. Interest ≈ 1557.16.

{"payment": 564.88, "interest": 1557.16}

6.79 seconds · $0.0087 estimated

Attempt 3 · Failed
Monthly rate i = 0.01, n = 24.

Payment = 12000 · 0.01 / (1 − 1.01^−24).

1.01^24 = 1.269734648531914...
1.01^−24 = 1/1.269734648531914 = 0.787566...

Compute: 1/1.2697346485 ≈ 0.78756691. Check: 1.2697346485 × 0.78756691 = 1.0000000 (1.2697346485×0.78 = 0.99039302; ×0.00756691 = 0.00960800; total 1.00000102, slightly high). Adjust: 0.00000102/1.26973 ≈ 0.0000008, so 0.7875661.

Then 1 − 0.7875661 = 0.2124339.
Payment = 120 / 0.2124339 = 564.8...

Compute: 0.2124339 × 565 = 120.0242. Excess 0.0242/0.2124339 = 0.1139. So the payment ≈ 564.886.

More precisely, the known value is 564.8671 for a 24-month 1% loan per 12000 (the factor 0.0470735 × 12000 = 564.882). The standard payment factor for 1%, 24 months is 0.04707347. Times 12000 = 564.88164.

Check: 0.04707347 × 12000 = 564.88164. This agrees with the estimate of about 564.886 within the rounding error of the hand calculation.

Total paid = 564.88164 × 24 = 13557.159; interest ≈ 1557.159.

More digits: the factor 0.0470734722 gives 564.88167 and 13557.1601, so the interest is 1557.16.

{"payment": 564.8816, "interest": 1557.16}

5.71 seconds · $0.0060 estimated

Gemini 3.8 Flash

Attempt 1 · Failed
{"payment": 564.88, "interest": 1557.16}

13.20 seconds · $0.0120 estimated

Attempt 2 · Failed
{"payment": 564.88, "interest": 1557.16}

12.40 seconds · $0.0102 estimated

Attempt 3 · Failed
{"payment": 564.88, "interest": 1557.16}

10.19 seconds · $0.0091 estimated

Calculate break-even with fees Finance

A product sells for $40. Variable costs are $17 production, $3 shipping, and a payment fee of 2.5% of sale price plus $0.30. Fixed monthly costs are $1870. Return the minimum whole units needed to break even and profit at that unit count. Return only a JSON object with these keys: units, profit. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"units":100,"profit":0}

2.34 seconds · $0.0007 estimated

Attempt 2 · Passed
{"units":100,"profit":0}

2.76 seconds · $0.0010 estimated

Attempt 3 · Passed
{"units":100,"profit":0}

2.57 seconds · $0.0007 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
Payment fee per unit: 0.025*40 + 0.30 = 1.00 + 0.30 = 1.30.
Variable cost per unit: 17 + 3 + 1.30 = 21.30.
Contribution margin: 40 - 21.30 = 18.70.
Break-even units: 1870 / 18.70 = 100 exactly.
Profit at 100 units: 100*18.70 - 1870 = 0.

{"units": 100, "profit": 0}

1.94 seconds · $0.0019 estimated

Attempt 2 · Failed
Payment fee per unit = 0.025*40 + 0.30 = 1.00 + 0.30 = 1.30.
Variable cost per unit = 17 + 3 + 1.30 = 21.30.
Contribution margin = 40 - 21.30 = 18.70.
Break-even units = 1870 / 18.70 = 100 exactly.
Profit at 100 units = 100*18.70 - 1870 = 0.

{"units": 100, "profit": 0}

1.84 seconds · $0.0019 estimated

Attempt 3 · Failed
Payment fee per unit = 0.025*40 + 0.30 = 1.00 + 0.30 = 1.30.
Variable cost per unit = 17 + 3 + 1.30 = 21.30.
Contribution margin = 40 − 21.30 = 18.70.
Break-even units = 1870 / 18.70 = 100 exactly.
Profit at 100 units = 100 × 18.70 − 1870 = 0.

{"units": 100, "profit": 0}

1.87 seconds · $0.0019 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"units": 100, "profit": 0}

2.89 seconds · $0.0021 estimated

Attempt 2 · Passed
{"units": 100, "profit": 0}

2.26 seconds · $0.0019 estimated

Attempt 3 · Passed
{"units": 100, "profit": 0}

1.77 seconds · $0.0013 estimated

Calculate independent service reliability Probability

Each of three independent services succeeds with probability 0.98. A workflow requires all three. It retries the complete workflow once if the first attempt fails; attempts are independent. Calculate eventual success probability and expected number of complete attempts. Return only a JSON object with these keys: success_probability, expected_attempts. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"success_probability":0.996541619136,"expected_attempts":1.058808}

3.99 seconds · $0.0012 estimated

Attempt 2 · Passed
{"success_probability":0.996541619136,"expected_attempts":1.058808}

2.70 seconds · $0.0011 estimated

Attempt 3 · Passed
{"success_probability":0.996541619136,"expected_attempts":1.058808}

3.04 seconds · $0.0010 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
p = 0.98^3 = 0.941192
Failure = 0.058808
Success probability = 1 - 0.058808^2 = 1 - 0.003458... compute: 0.058808^2 = 0.003458380864 (0.058808*0.058808: 0.058808*0.05=0.0029404; *0.008808=0.000517981...; sum: 0.0029404+0.00051798=0.00345838). So success = 0.996541619136.

Expected attempts = 1 + 0.058808 = 1.058808.

{"success_probability": 0.996541619136, "expected_attempts": 1.058808}

2.68 seconds · $0.0023 estimated

Attempt 2 · Failed
p = 0.98^3 = 0.941192
Failure = 0.058808
Success probability = 1 - 0.058808^2 = 1 - 0.003458700864 = 0.996541299136

Check 0.058808^2: 0.058808*0.058808 = 0.0034587... 0.058808^2 = (0.0588)^2 + 2*0.0588*0.000008 + 6.4e-11 = 0.00345744 + 0.0000009408 + 0.000000000064 = 0.003458380864. Recompute: 0.0588^2 = 0.00345744. Add 0.0000009408 → 0.0034583808, plus 6.4e-11 → 0.003458380864.
Success = 1 - 0.003458380864 = 0.996541619136

Expected attempts = 1 + 0.058808 = 1.058808

{"success_probability": 0.996541619136, "expected_attempts": 1.058808}

2.98 seconds · $0.0034 estimated

Attempt 3 · Failed
p = 0.98^3 = 0.941192
Failure = 0.058808
Success probability = 1 - 0.058808^2 = 1 - 0.003458540864 = 0.996541459136

0.058808^2: 0.058808*0.058808 = 0.003458 (0.058808*0.05=0.0029404; *0.008808=0.00051798...). Compute precisely: 0.058808*0.008808 = 0.000517980... 0.058808*0.008=0.000470464; *0.000808=0.0000475168 (0.058808*0.0008=0.0000470464; *0.000008=0.000000470464; sum 0.000047516864). Total 0.000517980864. Plus 0.0029404 = 0.003458380864.

So success = 1 - 0.003458380864 = 0.996541619136.

Expected attempts = 1 + 0.058808 = 1.058808

{"success_probability": 0.996541619136, "expected_attempts": 1.058808}

3.34 seconds · $0.0035 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"success_probability": 0.996541619136, "expected_attempts": 1.058808}

2.64 seconds · $0.0022 estimated

Attempt 2 · Passed
{"success_probability": 0.996541619136, "expected_attempts": 1.058808}

3.91 seconds · $0.0035 estimated

Attempt 3 · Passed
{"success_probability": 0.996541619136, "expected_attempts": 1.058808}

3.42 seconds · $0.0035 estimated

Compare sequential and additive discounts Finance

An item is $250 before discounts and tax. Apply 20% off, then a further 15% off the discounted price, then 8% tax. Return the final price and effective pretax discount percentage. Return only a JSON object with these keys: final_price, discount_percent. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"final_price":183.6,"discount_percent":32}

2.13 seconds · $0.0006 estimated

Attempt 2 · Passed
{"final_price":183.6,"discount_percent":32}

3.50 seconds · $0.0006 estimated

Attempt 3 · Passed
{"final_price":183.6,"discount_percent":32}

1.93 seconds · $0.0006 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
250 × 0.8 = 200; × 0.85 = 170; × 1.08 = 183.6.
Effective pretax discount: (250 − 170)/250 = 32%.

{"final_price": 183.6, "discount_percent": 32}

1.61 seconds · $0.0011 estimated

Attempt 2 · Failed
250 × 0.8 = 200; × 0.85 = 170; × 1.08 = 183.6. Effective pretax discount: (250 − 170)/250 = 32%.

{"final_price": 183.6, "discount_percent": 32}

1.60 seconds · $0.0010 estimated

Attempt 3 · Failed
250 × 0.8 = 200; × 0.85 = 170; × 1.08 = 183.6. The effective pretax discount is (250 − 170) / 250 = 32%.

{"final_price": 183.6, "discount_percent": 32}

1.77 seconds · $0.0011 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"final_price": 183.6, "discount_percent": 32}

1.70 seconds · $0.0012 estimated

Attempt 2 · Passed
{"final_price": 183.6, "discount_percent": 32}

2.05 seconds · $0.0014 estimated

Attempt 3 · Passed
{"final_price": 183.6, "discount_percent": 32}

2.18 seconds · $0.0015 estimated

Optimize a constrained production plan Optimization

Product A uses 3 labour hours and 2 machine hours and earns $40 contribution. B uses 2 labour and 4 machine hours and earns $50. Weekly limits are 120 labour and 160 machine hours. Products must be whole units. Maximize contribution; if tied choose more A. Return A, B and contribution. Return only a JSON object with these keys: a, b, contribution. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"a":20,"b":30,"contribution":2300}

4.77 seconds · $0.0012 estimated

Attempt 2 · Passed
{"a":20,"b":30,"contribution":2300}

4.13 seconds · $0.0011 estimated

Attempt 3 · Passed
{"a":20,"b":30,"contribution":2300}

4.35 seconds · $0.0011 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
Maximize 40A + 50B subject to 3A + 2B ≤ 120 and 2A + 4B ≤ 160, with A and B whole numbers.

The two constraints intersect where 3A + 2B = 120 and A + 2B = 80. Subtracting gives 2A = 40, so A = 20 and B = 30. The contribution there is 800 + 1500 = 2300.

Checking the other corner points:
- A = 0, B = 40 gives 2000.
- A = 40, B = 0 gives 1600.

The intersection is integer-valued and gives the highest contribution found. The objective slope (40:50) lies between the constraint slopes, so the optimum is at the vertex. Ties are not a concern, since the objective is not parallel to either constraint: 40/50 = 0.8, versus 3/2 = 1.5 and 2/4 = 0.5.

{"a":20,"b":30,"contribution":2300}

3.58 seconds · $0.0034 estimated

Attempt 2 · Failed
Maximize 40A + 50B subject to 3A + 2B ≤ 120 and 2A + 4B ≤ 160, with A and B whole numbers.

The two constraints intersect where 3A + 2B = 120 and A + 2B = 80. Subtracting gives 2A = 40, so A = 20 and B = 30. The contribution there is 800 + 1500 = 2300.

Checking the other corner points:
- A = 40, B = 0: 1600
- A = 0, B = 40: 2000

The intersection is integer-valued and gives the highest contribution, 2300. To check for ties, the objective ratio 40:50 differs from both constraint ratios (3:2 and 2:4), so no other point matches 2300.

{"a":20,"b":30,"contribution":2300}

3.01 seconds · $0.0030 estimated

Attempt 3 · Failed
Maximize 40A + 50B subject to 3A + 2B ≤ 120 and 2A + 4B ≤ 160, with A and B whole numbers.

The two constraints intersect where 3A + 2B = 120 and A + 2B = 80. Subtracting gives 2A = 40, so A = 20 and B = 30. The contribution there is 800 + 1500 = 2300.

Checking the other corners:
- A = 0, B = 40 gives 2000.
- A = 40, B = 0 gives 1600.

Since 2300 is the largest and the intersection is already whole numbers, it is optimal. The ratios 40/50 and 3/2 do not match, so there are no ties along an edge.

{"a":20,"b":30,"contribution":2300}

3.14 seconds · $0.0029 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"a": 20, "b": 30, "contribution": 2300}

3.02 seconds · $0.0022 estimated

Attempt 2 · Passed
{"a": 20, "b": 30, "contribution": 2300}

3.26 seconds · $0.0028 estimated

Attempt 3 · Passed
{"a": 20, "b": 30, "contribution": 2300}

3.58 seconds · $0.0030 estimated

Discount a project cash flow Finance

Pay $1000 now, then receive $400 at each year end for three years. Use a 10% annual discount rate. Return net present value without rounding intermediate values. Return only a JSON object with these keys: npv. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"npv":-5.259203606311044}

4.86 seconds · $0.0016 estimated

Attempt 2 · Passed
{"npv":-5.259203606311044}

8.02 seconds · $0.0027 estimated

Attempt 3 · Passed
{"npv":-5.259203606311044}

4.58 seconds · $0.0017 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
NPV = -1000 + 400*(1/1.1 + 1/1.21 + 1/1.331)

Annuity factor = 2.486851991...
Exactly: 1/1.1 = 0.909090909091, 1/1.21 = 0.826446280992, 1/1.331 = 0.751314800902. Sum = 2.486851990985.

Times 400 = 994.740796394.

NPV = -5.259203606

{"npv": -5.259203606}

2.49 seconds · $0.0018 estimated

Attempt 2 · Failed
NPV = -1000 + 400 × (1/1.1 + 1/1.21 + 1/1.331)

Annuity factor = (1 - 1.1^-3)/0.1 = 2.486851991...

Exact: 1/1.1 = 0.909090909..., 1/1.21 = 0.826446281, 1/1.331 = 0.751314801. Sum = 2.486851991.

Times 400 = 994.7407964

NPV = -5.2592036

More precisely, the sum is 2.4868519909842 (since 1.1^3 = 1.331, the exact sum is (1.21 + 1.1 + 1)/1.331 = 3.31/1.331 = 2.48685199098...). Times 400 = 1324/1.331 = 994.740796...; 1324/1.331 = 994.7407964...; NPV = -5.2592036.

Check: 1.331 × 994.7408 = 1324.0000 (1.331 × 994.7408 = 994.7408 + 0.331×994.7408 = 994.7408 + 329.2593 = 1324.0001). Good.

{"npv": -5.259203606311}

3.87 seconds · $0.0040 estimated

Attempt 3 · Failed
NPV = -1000 + 400 × (1/1.1 + 1/1.21 + 1/1.331)

Annuity factor = (1 - 1.1^-3)/0.1. 1.1^3 = 1.331, so 1/1.331 = 0.7513148009...

Annuity factor = (1 - 0.7513148009)/0.1 = 2.486851991

400 × 2.486851991 = 994.7407963

NPV = -5.2592037

More precisely: 1/1.331 = 0.75131480090157..., 1 - that = 0.24868519909843, factor = 2.4868519909843, times 400 = 994.74079639372, NPV = -5.25920360628.

{"npv": -5.25920360628}

2.97 seconds · $0.0026 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"npv": -5.259203606311044}

3.07 seconds · $0.0027 estimated

Attempt 2 · Passed
{"npv": -5.25920360631106}

3.32 seconds · $0.0032 estimated

Attempt 3 · Passed
{"npv": -5.259203606311044}

2.97 seconds · $0.0027 estimated

Allocate a shared expense Finance

Allocate exactly $10.00 to three departments with weights 1, 1, 1. Work in cents: floor shares, then allocate remaining cents by descending fractional remainder, ties to lower index. Return integer cents for departments A, B and C. Return only a JSON object with these keys: a_cents, b_cents, c_cents. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"a_cents":334,"b_cents":333,"c_cents":333}

1.61 seconds · $0.0004 estimated

Attempt 2 · Passed
{"a_cents":334,"b_cents":333,"c_cents":333}

1.65 seconds · $0.0007 estimated

Attempt 3 · Passed
{"a_cents":334,"b_cents":333,"c_cents":333}

2.99 seconds · $0.0006 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
{"a_cents": 334, "b_cents": 333, "c_cents": 333}

1.33 seconds · $0.0006 estimated

Attempt 2 · Passed
{"a_cents": 334, "b_cents": 333, "c_cents": 333}

1.23 seconds · $0.0006 estimated

Attempt 3 · Passed
{"a_cents": 334, "b_cents": 333, "c_cents": 333}

1.21 seconds · $0.0006 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"a_cents": 334, "b_cents": 333, "c_cents": 333}

2.24 seconds · $0.0015 estimated

Attempt 2 · Passed
{"a_cents": 334, "b_cents": 333, "c_cents": 333}

2.08 seconds · $0.0016 estimated

Attempt 3 · Passed
{"a_cents": 334, "b_cents": 333, "c_cents": 333}

2.75 seconds · $0.0015 estimated

Calculate sample uncertainty Statistics

For observations 4, 7, 7, 10, 12 calculate the arithmetic mean, unbiased sample variance (denominator n-1), and standard error of the mean. Return only a JSON object with these keys: mean, sample_variance, standard_error. Numeric fields must be numbers. No Markdown or explanation.

GPT-6.1 Sol

Attempt 1 · Passed
{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209022}

2.57 seconds · $0.0009 estimated

Attempt 2 · Passed
{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209022}

4.16 seconds · $0.0009 estimated

Attempt 3 · Passed
{"mean":8,"sample_variance":9.5,"standard_error":1.378404875209022}

2.66 seconds · $0.0009 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
Sum = 4+7+7+10+12 = 40, so the mean is 8.

Squared deviations: 16, 1, 1, 4, 16, which sum to 38. Sample variance = 38/4 = 9.5.

Standard error = sqrt(9.5/5) = sqrt(1.9) ≈ 1.378405.

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.3784048752090222}

2.16 seconds · $0.0017 estimated

Attempt 2 · Failed
Mean = (4+7+7+10+12)/5 = 40/5 = 8.

Squared deviations: 16, 1, 1, 4, 16, which sum to 38.

Sample variance = 38/4 = 9.5.

Standard error = sqrt(9.5/5) = sqrt(1.9) ≈ 1.378405.

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.3784048752090222}

2.03 seconds · $0.0017 estimated

Attempt 3 · Failed
Mean = (4+7+7+10+12)/5 = 40/5 = 8.

Squared deviations: 16, 1, 1, 4, 16, which sum to 38. Sample variance = 38/4 = 9.5.

Standard error = sqrt(9.5/5) = sqrt(1.9) ≈ 1.378405.

{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378405}

1.90 seconds · $0.0017 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

1.97 seconds · $0.0014 estimated

Attempt 2 · Passed
{"mean": 8.0, "sample_variance": 9.5, "standard_error": 1.378404875209022}

2.05 seconds · $0.0015 estimated

Attempt 3 · Passed
{"mean": 8, "sample_variance": 9.5, "standard_error": 1.378404875209022}

2.22 seconds · $0.0015 estimated

Reproduce this run

Download original JSON evidence · Download original CSV results

Scoring audit: 2026-10-04T18:18:02.762014+00:00 · 0 verdicts changed · zero new provider requests. Download appended correction.

Run history

  • 2026-10-04T08:20:14.082456+00:00 · complete · 108 / 108 required trials initiated · Archived evidence

Maintenance needs recovery: The daily source check is overdue. Published test dates remain unchanged.

Source-check status: A source fetch has completed; factual changes require review. Latest attempt: 2026-10-04T09:01:57.926856+00:00.

Download the public task bank · Read the methodology