AI News

AI coding comparison tests and results

Measured results

Compare AI models on Python coding tasks

This bank measures function implementations under a stated Python subset, not IDEs or whole-repository agents.

Last tested
2026-10-04T17:01:11.203235+00:00

Last successful source fetch
2026-10-04T09:01:57.926856+00:00

Last published
2026-10-04T18:18:57.482142+00:00

Observed leaders

gpt-6.1-sol

Descriptive pass counts after offline runner correction of the original outputs. Task ambiguities described in the audit limit broader conclusions.

Tested API configurations

ModelTask passesCost per passProcessing
GPT-6.1 Sol
gpt-6.1-sol
34 / 36$0.0021 (estimated)OpenAI direct API, Standard global
Medium reasoning; 8,192-token cap
Claude Sonnet 5.5
claude-sonnet-5-5
30 / 36
Originally 27 / 36; corrected offline
$0.0032 (estimated)Claude direct API, Standard global
Medium reasoning; 8,192-token cap
Gemini 3.8 Flash
gemini-3.8-flash
32 / 36$0.0075 (estimated)Gemini Developer API, paid Standard
Medium reasoning; 8,192-token cap

Three documented current direct-API models from different providers, selected for coding/math use and practical cost. Not a claim that every leading tool is included. Account access and publication terms require verification before dispatch. No web, tools, prompt caching, batch or regional processing. Medium reasoning labels are provider-specific, not equal compute.

Scores by task group

ModelGroupPassesMedian latency
gpt-6.1-solAlgorithms12 / 124.01 seconds (whole model)
gpt-6.1-solData correctness13 / 154.01 seconds (whole model)
gpt-6.1-solSafety and edge cases9 / 94.01 seconds (whole model)
claude-sonnet-5-5Algorithms6 / 122.17 seconds (whole model)
claude-sonnet-5-5Data correctness15 / 152.17 seconds (whole model)
claude-sonnet-5-5Safety and edge cases9 / 92.17 seconds (whole model)
gemini-3.8-flashAlgorithms11 / 126.39 seconds (whole model)
gemini-3.8-flashData correctness12 / 156.39 seconds (whole model)
gemini-3.8-flashSafety and edge cases9 / 96.39 seconds (whole model)
Historical and current-price estimates

gpt-6.1-sol: original run $0.0723; same recorded usage at rates reviewed 2026-10-04T06:52:40.966997+00:00: $0.0723. Both are estimates; billing reconciliation is separate.

claude-sonnet-5-5: original run $0.0951; same recorded usage at rates reviewed 2026-10-04T06:52:40.966997+00:00: $0.0951. Both are estimates; billing reconciliation is separate.

gemini-3.8-flash: original run $0.2410; same recorded usage at rates reviewed 2026-10-04T06:52:40.966997+00:00: $0.2410. Both are estimates; billing reconciliation is separate.

The 12 tasks

Open a task to inspect its exact prompt and, after testing, every model output. Hidden coding fixtures and answer keys are not served to the models or this page.

Outputs are generated by the named AI models and may contain errors. Scores cover the declared task checks; they do not verify every statement in an output.

Reconcile invoices and payments Data correctness

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input has invoices and payments lists of {id, cents}. Sum amounts by id; payments subtract from invoices. Return a dict of nonzero outstanding amounts sorted by id, including negative overpayments. IDs are strings. Empty lists are valid.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    totals = {}
    for invoice in data["invoices"]:
        key = invoice["id"]
        totals[key] = totals.get(key, 0) + invoice["cents"]
    for payment in data["payments"]:
        key = payment["id"]
        totals[key] = totals.get(key, 0) - payment["cents"]
    return {key: totals[key] for key in sorted(totals) if totals[key] != 0}

2.79 seconds · $0.0013 estimated

Attempt 2 · Passed
def solve(data):
    totals = {}
    for invoice in data["invoices"]:
        identifier = invoice["id"]
        totals[identifier] = totals.get(identifier, 0) + invoice["cents"]
    for payment in data["payments"]:
        identifier = payment["id"]
        totals[identifier] = totals.get(identifier, 0) - payment["cents"]
    return {
        identifier: totals[identifier]
        for identifier in sorted(totals)
        if totals[identifier] != 0
    }

4.17 seconds · $0.0014 estimated

Attempt 3 · Passed
def solve(data):
    totals = {}
    for invoice in data["invoices"]:
        identifier = invoice["id"]
        totals[identifier] = totals.get(identifier, 0) + invoice["cents"]
    for payment in data["payments"]:
        identifier = payment["id"]
        totals[identifier] = totals.get(identifier, 0) - payment["cents"]
    return {
        identifier: totals[identifier]
        for identifier in sorted(totals)
        if totals[identifier] != 0
    }

3.07 seconds · $0.0014 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
def solve(data):
    if isinstance(data, str):
        import json
        data = json.loads(data)
    totals = {}
    for item in (data.get("invoices") or []):
        k = str(item["id"])
        totals[k] = totals.get(k, 0) + int(item["cents"])
    for item in (data.get("payments") or []):
        k = str(item["id"])
        totals[k] = totals.get(k, 0) - int(item["cents"])
    return {k: totals[k] for k in sorted(totals) if totals[k] != 0}

2.13 seconds · $0.0022 estimated

Attempt 2 · Passed
def solve(data):
    if isinstance(data, str):
        import json
        data = json.loads(data)
    totals = {}
    for item in (data.get("invoices") or []):
        k = str(item["id"])
        totals[k] = totals.get(k, 0) + int(item["cents"])
    for item in (data.get("payments") or []):
        k = str(item["id"])
        totals[k] = totals.get(k, 0) - int(item["cents"])
    return {k: totals[k] for k in sorted(totals) if totals[k] != 0}

1.96 seconds · $0.0022 estimated

Attempt 3 · Passed
def solve(data):
    if isinstance(data, str):
        import json
        data = json.loads(data)
    totals = {}
    for item in data.get("invoices", []) or []:
        k = str(item["id"])
        totals[k] = totals.get(k, 0) + int(item["cents"])
    for item in data.get("payments", []) or []:
        k = str(item["id"])
        totals[k] = totals.get(k, 0) - int(item["cents"])
    return {k: totals[k] for k in sorted(totals) if totals[k] != 0}

2.16 seconds · $0.0023 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
def solve(data):
    balances = {}
    for inv in data.get("invoices", []):
        balances[inv["id"]] = balances.get(inv["id"], 0) + inv["cents"]
    for pmt in data.get("payments", []):
        balances[pmt["id"]] = balances.get(pmt["id"], 0) - pmt["cents"]
    return {k: balances[k] for k in sorted(balances) if balances[k] != 0}

3.85 seconds · $0.0034 estimated

Attempt 2 · Passed
import collections

def solve(data):
    balances = collections.defaultdict(int)
    for item in data.get('invoices', []):
        balances[item['id']] += item['cents']
    for item in data.get('payments', []):
        balances[item['id']] -= item['cents']
    return {k: balances[k] for k in sorted(balances) if balances[k] != 0}

4.39 seconds · $0.0037 estimated

Attempt 3 · Passed
def solve(data):
    balances = {}
    for inv in data.get('invoices', []):
        balances[inv['id']] = balances.get(inv['id'], 0) + inv['cents']
    for pmt in data.get('payments', []):
        balances[pmt['id']] = balances.get(pmt['id'], 0) - pmt['cents']
    return {k: balances[k] for k in sorted(balances) if balances[k] != 0}

5.08 seconds · $0.0040 estimated

Resolve duplicate event deliveries Data correctness

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is a list of {id, version, value}. Keep the highest version for each id; on equal version keep the last occurrence. Return the retained objects sorted lexicographically by id. Do not mutate the input.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    retained = {}
    for item in data:
        item_id = item["id"]
        if item_id not in retained or item["version"] >= retained[item_id]["version"]:
            retained[item_id] = item
    return [retained[item_id] for item_id in sorted(retained)]

3.02 seconds · $0.0010 estimated

Attempt 2 · Passed
def solve(data):
    latest = {}
    for item in data:
        identifier = item["id"]
        if identifier not in latest or item["version"] >= latest[identifier]["version"]:
            latest[identifier] = item
    return [latest[identifier] for identifier in sorted(latest)]

4.08 seconds · $0.0017 estimated

Attempt 3 · Passed
def solve(data):
    retained = {}
    for item in data:
        item_id = item["id"]
        if item_id not in retained or item["version"] >= retained[item_id]["version"]:
            retained[item_id] = item
    return [retained[item_id] for item_id in sorted(retained)]

3.15 seconds · $0.0010 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
def solve(data):
    best = {}
    for item in data:
        key = item["id"]
        cur = best.get(key)
        if cur is None or item["version"] >= cur["version"]:
            best[key] = item
    result = []
    for key in sorted(best, key=lambda k: str(k)):
        result.append(dict(best[key]))
    return result

1.83 seconds · $0.0015 estimated

Attempt 2 · Passed
def solve(data):
    best = {}
    for item in data:
        key = item["id"]
        cur = best.get(key)
        if cur is None or item["version"] >= cur["version"]:
            best[key] = item
    result = []
    for key in sorted(best, key=lambda k: str(k)):
        item = best[key]
        result.append(dict(item))
    return result

1.78 seconds · $0.0016 estimated

Attempt 3 · Passed
def solve(data):
    best = {}
    for item in data:
        key = item["id"]
        cur = best.get(key)
        if cur is None or item["version"] >= cur["version"]:
            best[key] = item
    result = []
    for key in sorted(best, key=lambda k: str(k)):
        item = best[key]
        result.append(dict(item))
    return result

1.88 seconds · $0.0016 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
def solve(data):
    best = {}
    for item in data:
        item_id = item['id']
        if item_id not in best or item['version'] >= best[item_id]['version']:
            best[item_id] = item
    return sorted(best.values(), key=lambda x: str(x['id']))

9.10 seconds · $0.0073 estimated

Attempt 2 · Passed
def solve(data):
    best = {}
    for item in data:
        item_id = item['id']
        if item_id not in best or item['version'] >= best[item_id]['version']:
            best[item_id] = item
    return sorted(best.values(), key=lambda x: str(x['id']))

6.79 seconds · $0.0058 estimated

Attempt 3 · Passed
def solve(data):
    best = {}
    for item in data:
        item_id = item['id']
        if item_id not in best or item['version'] >= best[item_id]['version']:
            best[item_id] = item
    return sorted(best.values(), key=lambda x: str(x['id']))

7.29 seconds · $0.0055 estimated

Merge maintenance windows Algorithms

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is a list of [start,end] integer half-open intervals, with start<=end. Drop empty intervals; merge overlapping or touching intervals. Return sorted merged intervals without mutating input.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    intervals = sorted((start, end) for start, end in data if start < end)
    merged = []
    for start, end in intervals:
        if merged and start <= merged[-1][1]:
            merged[-1][1] = max(merged[-1][1], end)
        else:
            merged.append([start, end])
    return merged

2.42 seconds · $0.0011 estimated

Attempt 2 · Passed
def solve(data):
    intervals = sorted((start, end) for start, end in data if start < end)
    merged = []
    for start, end in intervals:
        if merged and start <= merged[-1][1]:
            merged[-1][1] = max(merged[-1][1], end)
        else:
            merged.append([start, end])
    return merged

2.92 seconds · $0.0011 estimated

Attempt 3 · Passed
def solve(data):
    intervals = sorted((start, end) for start, end in data if start < end)
    merged = []
    for start, end in intervals:
        if merged and start <= merged[-1][1]:
            if end > merged[-1][1]:
                merged[-1][1] = end
        else:
            merged.append([start, end])
    return merged

2.65 seconds · $0.0011 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
def solve(data):
    # Keep only non-empty intervals, as fresh (start, end) pairs, so the input is never modified.
    items = []
    for iv in data:
        s, e = iv[0], iv[1]
        if s < e:
            items.append((s, e))
    items.sort()
    result = []
    for s, e in items:
        # Merge when the new interval overlaps or touches the previous one.
        if result and s <= result[-1][1]:
            if e > result[-1][1]:
                result[-1][1] = e
        else:
            result.append([s, e])
    return result

2.21 seconds · $0.0023 estimated

Attempt 2 · Passed
def solve(data):
    intervals = []
    for item in data:
        s, e = item[0], item[1]
        if s < e:
            intervals.append((s, e))
    intervals.sort()
    result = []
    for s, e in intervals:
        if result and s <= result[-1][1]:
            if e > result[-1][1]:
                result[-1][1] = e
        else:
            result.append([s, e])
    return result

1.86 seconds · $0.0018 estimated

Attempt 3 · Passed
def solve(data):
    intervals = []
    for item in data:
        s, e = item[0], item[1]
        if s < e:
            intervals.append((s, e))
    intervals.sort()
    result = []
    for s, e in intervals:
        if result and s <= result[-1][1]:
            if e > result[-1][1]:
                result[-1][1] = e
        else:
            result.append([s, e])
    return result

1.76 seconds · $0.0018 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
def solve(data):
    valid_intervals = [list(interval) for interval in data if interval[0] < interval[1]]
    if not valid_intervals:
        return []
    valid_intervals.sort(key=lambda x: (x[0], x[1]))
    merged = [valid_intervals[0]]
    for current in valid_intervals[1:]:
        last = merged[-1]
        if current[0] <= last[1]:
            if current[1] > last[1]:
                last[1] = current[1]
        else:
            merged.append(current)
    return merged

4.37 seconds · $0.0036 estimated

Attempt 2 · Passed
def solve(data):
    valid = [[iv[0], iv[1]] for iv in data if iv[0] < iv[1]]
    if not valid:
        return []
    valid.sort(key=lambda x: (x[0], x[1]))
    merged = [valid[0]]
    for iv in valid[1:]:
        if iv[0] <= merged[-1][1]:
            if iv[1] > merged[-1][1]:
                merged[-1][1] = iv[1]
        else:
            merged.append(iv)
    return merged

4.96 seconds · $0.0047 estimated

Attempt 3 · Passed
def solve(data):
    intervals = [list(interval) for interval in data if interval[0] < interval[1]]
    if not intervals:
        return []
    intervals.sort()
    merged = [intervals[0]]
    for current in intervals[1:]:
        if current[0] <= merged[-1][1]:
            if current[1] > merged[-1][1]:
                merged[-1][1] = current[1]
        else:
            merged.append(current)
    return merged

5.55 seconds · $0.0046 estimated

Redact nested customer records Safety and edge cases

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is arbitrary JSON. Recursively remove dictionary keys whose lowercase spelling equals password, token or secret. Preserve other values and list order; never mutate the original.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    def clean(value):
        if isinstance(value, dict):
            return {
                key: clean(item)
                for key, item in value.items()
                if key.lower() not in ("password", "token", "secret")
            }
        if isinstance(value, list):
            return [clean(item) for item in value]
        return value

    return clean(data)

5.82 seconds · $0.0023 estimated

Attempt 2 · Passed
def solve(data):
    if isinstance(data, dict):
        return {
            key: solve(value)
            for key, value in data.items()
            if key.lower() not in ("password", "token", "secret")
        }
    if isinstance(data, list):
        return [solve(value) for value in data]
    return data

6.12 seconds · $0.0022 estimated

Attempt 3 · Passed
def solve(data):
    if isinstance(data, dict):
        return {
            key: solve(value)
            for key, value in data.items()
            if key.lower() not in ("password", "token", "secret")
        }
    if isinstance(data, list):
        return [solve(value) for value in data]
    return data

6.99 seconds · $0.0024 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
import json

def solve(data):
    banned = {"password", "token", "secret"}

    def clean(x):
        if isinstance(x, dict):
            result = {}
            for k, v in x.items():
                if isinstance(k, str) and k.lower() in banned:
                    continue
                result[k] = clean(v)
            return result
        if isinstance(x, list):
            return [clean(i) for i in x]
        if isinstance(x, tuple):
            return [clean(i) for i in x]
        return x

    return clean(data)

2.01 seconds · $0.0021 estimated

Attempt 2 · Passed
import json

def solve(data):
    banned = {"password", "token", "secret"}

    def clean(v):
        if isinstance(v, dict):
            result = {}
            for k, val in v.items():
                if isinstance(k, str) and k.lower() in banned:
                    continue
                result[k] = clean(val)
            return result
        if isinstance(v, list):
            return [clean(x) for x in v]
        if isinstance(v, tuple):
            return tuple(clean(x) for x in v)
        return v

    if isinstance(data, str):
        try:
            parsed = json.loads(data)
        except (ValueError, TypeError):
            return data
        if isinstance(parsed, (dict, list)):
            return clean(parsed)
        return data
    return clean(data)

2.46 seconds · $0.0029 estimated

Attempt 3 · Passed
import math, json, csv, io, re, collections, heapq, unicodedata

def solve(data):
    banned = ("password", "token", "secret")

    def clean(x):
        if isinstance(x, dict):
            out = {}
            for k, v in x.items():
                if isinstance(k, str) and k.lower() in banned:
                    continue
                out[k] = clean(v)
            return out
        if isinstance(x, list):
            return [clean(i) for i in x]
        if isinstance(x, tuple):
            return [clean(i) for i in x]
        return x

    return clean(data)

2.10 seconds · $0.0023 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
def solve(data):
    forbidden = {"password", "token", "secret"}
    if isinstance(data, dict):
        return {
            k: solve(v)
            for k, v in data.items()
            if not (isinstance(k, str) and k.lower() in forbidden)
        }
    if isinstance(data, list):
        return [solve(item) for item in data]
    return data

6.49 seconds · $0.0055 estimated

Attempt 2 · Passed
def solve(data):
    if isinstance(data, dict):
        return {
            k: solve(v)
            for k, v in data.items()
            if not (isinstance(k, str) and k.lower() in {"password", "token", "secret"})
        }
    if isinstance(data, list):
        return [solve(item) for item in data]
    return data

5.11 seconds · $0.0039 estimated

Attempt 3 · Passed
def solve(data):
    if isinstance(data, dict):
        return {
            k: solve(v)
            for k, v in data.items()
            if not (isinstance(k, str) and k.lower() in ("password", "token", "secret"))
        }
    if isinstance(data, list):
        return [solve(item) for item in data]
    return data

6.98 seconds · $0.0043 estimated

Calculate seven-day activity totals Algorithms

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is {events:[[day,count],...], query_days:[integer,...]}. Return totals for each query day covering inclusive days q-6 through q. Events can be unsorted, repeated or negative-day; count is nonnegative. Query order must be preserved.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    events = sorted(data["events"], key=lambda event: event[0])
    queries = sorted(enumerate(data["query_days"]), key=lambda item: item[1])
    results = [0] * len(queries)

    left = 0
    right = 0
    total = 0

    for index, day in queries:
        while right < len(events) and events[right][0] <= day:
            total += events[right][1]
            right += 1

        while left < right and events[left][0] < day - 6:
            total -= events[left][1]
            left += 1

        results[index] = total

    return results

7.17 seconds · $0.0035 estimated

Attempt 2 · Passed
def solve(data):
    events = sorted(data["events"], key=lambda event: event[0])
    queries = sorted(enumerate(data["query_days"]), key=lambda item: item[1])
    totals = [0] * len(queries)

    left = 0
    right = 0
    total = 0
    n = len(events)

    for index, day in queries:
        while right < n and events[right][0] <= day:
            total += events[right][1]
            right += 1

        start = day - 6
        while left < right and events[left][0] < start:
            total -= events[left][1]
            left += 1

        totals[index] = total

    return totals

7.63 seconds · $0.0035 estimated

Attempt 3 · Passed
def solve(data):
    events = sorted(data["events"])
    queries = sorted(enumerate(data["query_days"]), key=lambda item: item[1])
    totals = [0] * len(queries)
    left = right = 0
    running_total = 0

    for index, day in queries:
        while right < len(events) and events[right][0] <= day:
            running_total += events[right][1]
            right += 1

        start = day - 6
        while left < right and events[left][0] < start:
            running_total -= events[left][1]
            left += 1

        totals[index] = running_total

    return totals

6.51 seconds · $0.0032 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
import bisect is not allowed, so I'll implement the binary search by hand.

The approach is to aggregate counts by day, sort the distinct days, build prefix sums, and answer each query as the prefix sum over days up to q minus the prefix sum over days up to q-7.

```python
def solve(data):
    events = data.get("events", []) or []
    queries = data.get("query_days", []) or []

    agg = {}
    for e in events:
        d = e[0]
        c = e[1]
        agg[d] = agg.get(d, 0) + c

    days = sorted(agg.keys())
    prefix = [0]
    for d in days:
        prefix.append(prefix[-1] + agg[d])

    def upto(x):
        # number of days <= x
        lo, hi = 0, len(days)
        while lo < hi:
            mid = (lo + hi) // 2
            if days[mid] <= x:
                lo = mid + 1
            else:
                hi = mid
        return prefix[lo]

    res = []
    for q in queries:
        res.append(upto(q) - upto(q - 7))
    return res
```

2.66 seconds · $0.0041 estimated

Attempt 2 · Failed
def solve(data):
    import bisect
    events = data.get("events", []) or []
    queries = data.get("query_days", []) or []
    agg = {}
    for e in events:
        d, c = e[0], e[1]
        agg[d] = agg.get(d, 0) + c
    days = sorted(agg)
    prefix = [0]
    for d in days:
        prefix.append(prefix[-1] + agg[d])
    res = []
    for q in queries:
        lo = bisect.bisect_left(days, q - 6)
        hi = bisect.bisect_right(days, q)
        res.append(prefix[hi] - prefix[lo])
    return res

1.81 seconds · $0.0026 estimated

Attempt 3 · Failed
def solve(data):
    import bisect
    events = data.get("events", []) or []
    queries = data.get("query_days", []) or []
    agg = {}
    for e in events:
        d, c = e[0], e[1]
        agg[d] = agg.get(d, 0) + c
    days = sorted(agg)
    prefix = [0]
    for d in days:
        prefix.append(prefix[-1] + agg[d])
    res = []
    for q in queries:
        lo = bisect.bisect_left(days, q - 6)
        hi = bisect.bisect_right(days, q)
        res.append(prefix[hi] - prefix[lo])
    return res

5.85 seconds · $0.0026 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
import collections

def solve(data):
    events = data.get("events", [])
    query_days = data.get("query_days", [])
    
    day_counts = collections.defaultdict(int)
    for day, count in events:
        day_counts[day] += count
        
    unique_days = sorted(day_counts.keys())
    n = len(unique_days)
    
    pref = [0] * (n + 1)
    for i in range(n):
        pref[i + 1] = pref[i] + day_counts[unique_days[i]]
        
    def bisect_left(target):
        lo, hi = 0, n
        while lo < hi:
            mid = (lo + hi) // 2
            if unique_days[mid] < target:
                lo = mid + 1
            else:
                hi = mid
        return lo

    def bisect_right(target):
        lo, hi = 0, n
        while lo < hi:
            mid = (lo + hi) // 2
            if target < unique_days[mid]:
                hi = mid
            else:
                lo = mid + 1
        return lo

    results = []
    for q in query_days:
        l = bisect_left(q - 6)
        r = bisect_right(q)
        results.append(pref[r] - pref[l])
        
    return results

6.84 seconds · $0.0059 estimated

Attempt 2 · Passed
import collections

def solve(data):
    events = data.get('events', [])
    query_days = data.get('query_days', [])

    counts = collections.defaultdict(int)
    for day, count in events:
        counts[day] += count

    results = []
    for q in query_days:
        total = sum(counts.get(q - i, 0) for i in range(7))
        results.append(total)

    return results

6.25 seconds · $0.0049 estimated

Attempt 3 · Failed
def solve(data):
    events = data.get("events", [])
    query_days = data.get("query_days", [])
    
    day_counts = collections.defaultdict(int)
    for day, count in events:
        day_counts[day] += count
        
    unique_days = sorted(day_counts.keys())
    prefix = [0] * (len(unique_days) + 1)
    for idx, d in enumerate(unique_days):
        prefix[idx + 1] = prefix[idx] + day_counts[d]
        
    def bisect_left(arr, x):
        lo, hi = 0, len(arr)
        while lo < hi:
            mid = (lo + hi) // 2
            if arr[mid] < x:
                lo = mid + 1
            else:
                hi = mid
        return lo

    def bisect_right(arr, x):
        lo, hi = 0, len(arr)
        while lo < hi:
            mid = (lo + hi) // 2
            if arr[mid] <= x:
                lo = mid + 1
            else:
                hi = mid
        return lo

    results = []
    for q in query_days:
        left = bisect_left(unique_days, q - 6)
        right = bisect_right(unique_days, q)
        results.append(prefix[right] - prefix[left])
        
    return results

6.42 seconds · $0.0057 estimated

Allocate cents without rounding loss Data correctness

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is {total:nonnegative integer,weights:[nonnegative integer,...]}. Allocate proportionally using largest remainders: floor each share, then give leftover cents to largest fractional remainders, ties to lower index. If weights empty return []; if all zero return zeros.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    total = data["total"]
    weights = data["weights"]
    weight_sum = sum(weights)

    if weight_sum == 0:
        return [0] * len(weights)

    shares = []
    remainders = []
    for weight in weights:
        share, remainder = divmod(total * weight, weight_sum)
        shares.append(share)
        remainders.append(remainder)

    leftover = total - sum(shares)
    order = sorted(range(len(weights)), key=lambda i: (-remainders[i], i))
    for i in order[:leftover]:
        shares[i] += 1

    return shares

2.73 seconds · $0.0017 estimated

Attempt 2 · Passed
def solve(data):
    weights = data["weights"]
    total = data["total"]
    weight_sum = sum(weights)

    if not weights:
        return []
    if weight_sum == 0:
        return [0] * len(weights)

    allocations = []
    remainders = []
    for weight in weights:
        allocation, remainder = divmod(total * weight, weight_sum)
        allocations.append(allocation)
        remainders.append(remainder)

    leftover = total - sum(allocations)
    order = sorted(
        range(len(weights)),
        key=lambda i: (-remainders[i], i)
    )
    for i in order[:leftover]:
        allocations[i] += 1

    return allocations

5.47 seconds · $0.0024 estimated

Attempt 3 · Passed
def solve(data):
    total = data["total"]
    weights = data["weights"]
    weight_sum = sum(weights)
    if not weight_sum:
        return [0] * len(weights)

    shares = [divmod(total * weight, weight_sum) for weight in weights]
    result = [share[0] for share in shares]
    leftover = total - sum(result)
    order = sorted(range(len(weights)), key=lambda i: (-shares[i][1], i))
    for i in order[:leftover]:
        result[i] += 1
    return result

3.03 seconds · $0.0015 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
def solve(data):
    total = data.get("total", 0)
    weights = data.get("weights", [])
    if not weights:
        return []
    s = sum(weights)
    n = len(weights)
    if s == 0:
        return [0] * n
    base = []
    rems = []
    for i, w in enumerate(weights):
        q, r = divmod(total * w, s)
        base.append(q)
        rems.append(r)
    left = total - sum(base)
    order = sorted(range(n), key=lambda i: (-rems[i], i))
    for k in range(left):
        base[order[k]] += 1
    return base

2.11 seconds · $0.0025 estimated

Attempt 2 · Passed
def solve(data):
    total = data.get("total", 0)
    weights = data.get("weights", [])
    if not weights:
        return []
    s = sum(weights)
    n = len(weights)
    if s == 0:
        return [0] * n
    base = []
    rems = []
    for i, w in enumerate(weights):
        num = total * w
        q, r = divmod(num, s)
        base.append(q)
        rems.append((r, i))
    left = total - sum(base)
    order = sorted(range(n), key=lambda i: (-rems[i][0], i))
    for i in order[:left]:
        base[i] += 1
    return base

2.17 seconds · $0.0026 estimated

Attempt 3 · Passed
def solve(data):
    total = data.get("total", 0)
    weights = data.get("weights", [])
    if not weights:
        return []
    s = sum(weights)
    n = len(weights)
    if s == 0:
        return [0] * n
    base = []
    rems = []
    for i, w in enumerate(weights):
        q, r = divmod(total * w, s)
        base.append(q)
        rems.append((r, i))
    left = total - sum(base)
    order = sorted(range(n), key=lambda i: (-rems[i][0], i))
    for k in range(left):
        base[order[k]] += 1
    return base

2.16 seconds · $0.0025 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
def solve(data):
    total = data['total']
    weights = data['weights']
    if not weights:
        return []
    total_weight = sum(weights)
    if total_weight == 0:
        return [0] * len(weights)

    shares = []
    remainders = []
    for i, w in enumerate(weights):
        prod = total * w
        shares.append(prod // total_weight)
        remainders.append((prod % total_weight, i))

    leftover = total - sum(shares)
    remainders.sort(key=lambda item: (-item[0], item[1]))

    for k in range(leftover):
        idx = remainders[k][1]
        shares[idx] += 1

    return shares

6.31 seconds · $0.0058 estimated

Attempt 2 · Passed
def solve(data):
    total = data['total']
    weights = data['weights']
    if not weights:
        return []
    total_weight = sum(weights)
    if total_weight == 0:
        return [0] * len(weights)
    
    shares = []
    remainders = []
    for i, w in enumerate(weights):
        q, r = divmod(w * total, total_weight)
        shares.append(q)
        remainders.append((-r, i))
    
    leftover = total - sum(shares)
    remainders.sort()
    for k in range(leftover):
        shares[remainders[k][1]] += 1
        
    return shares

7.91 seconds · $0.0059 estimated

Attempt 3 · Passed
def solve(data):
    total = data["total"]
    weights = data["weights"]
    if not weights:
        return []
    total_weight = sum(weights)
    if total_weight == 0:
        return [0] * len(weights)

    n = len(weights)
    floors = [(total * w) // total_weight for w in weights]
    leftover = total - sum(floors)

    sorted_indices = sorted(
        range(n),
        key=lambda i: (-((total * weights[i]) % total_weight), i)
    )

    result = list(floors)
    for i in range(leftover):
        result[sorted_indices[i]] += 1

    return result

9.03 seconds · $0.0084 estimated

Parse and aggregate CSV purchases Data correctness

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is CSV text with header customer,cents, including quoted fields. Ignore rows with a blank customer or cents that is not a signed base-10 integer. Sum by customer and return a dictionary, including zero totals. Preserve customer spelling.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    import csv
    import io
    import re

    totals = {}
    for row in csv.DictReader(io.StringIO(data)):
        customer = row.get("customer")
        cents = row.get("cents")

        if customer is None or not customer.strip() or cents is None:
            continue

        cents = cents.strip()
        if re.fullmatch(r"[+-]?[0-9]+", cents) is None:
            continue

        totals[customer] = totals.get(customer, 0) + int(cents)

    return totals

7.38 seconds · $0.0031 estimated

Attempt 2 · Failed
import csv
import io
import re


def solve(data):
    totals = {}
    for row in csv.DictReader(io.StringIO(data)):
        customer = row.get("customer")
        cents = row.get("cents")
        if customer is None or not customer.strip():
            continue
        if cents is None or re.fullmatch(r"[+-]?[0-9]+", cents) is None:
            continue
        totals[customer] = totals.get(customer, 0) + int(cents, 10)
    return totals

8.88 seconds · $0.0038 estimated

Attempt 3 · Failed
def solve(data):
    import csv
    import io
    import re

    totals = {}
    for row in csv.DictReader(io.StringIO(data)):
        customer = row.get("customer")
        cents = row.get("cents")
        if customer is None or not customer.strip():
            continue
        if cents is None or re.fullmatch(r"[+-]?[0-9]+", cents) is None:
            continue
        totals[customer] = totals.get(customer, 0) + int(cents, 10)

    return totals

10.15 seconds · $0.0043 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
import csv
import io
import re


def solve(data):
    result = {}
    reader = csv.reader(io.StringIO(data, newline=""))
    header = next(reader, None)
    if header is None:
        return result
    names = [h.strip().lower() for h in header]
    try:
        ci = names.index("customer")
        ai = names.index("cents")
    except ValueError:
        return result
    pattern = re.compile(r"[+-]?[0-9]+")
    for row in reader:
        if not row:
            continue
        if ci >= len(row) or ai >= len(row):
            continue
        customer = row[ci]
        if customer.strip() == "":
            continue
        raw = row[ai].strip()
        if not pattern.fullmatch(raw):
            continue
        result[customer] = result.get(customer, 0) + int(raw)
    return result

2.77 seconds · $0.0032 estimated

Attempt 2 · Passed
import csv
import io
import re


def solve(data):
    totals = {}
    reader = csv.reader(io.StringIO(data, newline=""))
    header = None
    ci = 0
    vi = 1
    pattern = re.compile(r"[+-]?[0-9]+")
    for row in reader:
        if header is None:
            if not row:
                continue
            header = [h.strip().lstrip("\ufeff").lower() for h in row]
            if "customer" in header and "cents" in header:
                ci = header.index("customer")
                vi = header.index("cents")
            else:
                ci = 0
                vi = 1
            continue
        if not row:
            continue
        if len(row) <= max(ci, vi):
            continue
        name = row[ci]
        val = row[vi].strip()
        if name.strip() == "":
            continue
        if not pattern.fullmatch(val):
            continue
        n = int(val)
        totals[name] = totals.get(name, 0) + n
    return totals

2.94 seconds · $0.0037 estimated

Attempt 3 · Passed
import csv
import io
import re

def solve(data):
    result = {}
    reader = csv.reader(io.StringIO(data, newline=""))
    header = None
    for row in reader:
        if not row or all(not c.strip() for c in row):
            continue
        header = [c.strip().lower() for c in row]
        break
    if header is None:
        return result
    try:
        ci = header.index("customer")
        ai = header.index("cents")
    except ValueError:
        return result
    pat = re.compile(r"[+-]?[0-9]+")
    for row in reader:
        if len(row) <= max(ci, ai):
            continue
        name = row[ci]
        if not name.strip():
            continue
        s = row[ai].strip()
        if not pat.fullmatch(s):
            continue
        result[name] = result.get(name, 0) + int(s)
    return result

2.81 seconds · $0.0033 estimated

Gemini 3.8 Flash

Attempt 1 · Failed
import csv
import io
import re

def solve(data):
    totals = {}
    reader = csv.DictReader(io.StringIO(data))
    for row in reader:
        customer = row.get('customer')
        cents = row.get('cents')
        if customer is None or customer.strip() == '':
            continue
        if cents is None or not re.fullmatch(r'[+-]?[0-9]+', cents):
            continue
        totals[customer] = totals.get(customer, 0) + int(cents)
    return totals

20.55 seconds · $0.0171 estimated

Attempt 2 · Failed · Output cap reached
" - no markdown, only code.
- "no custom classes, private/dunder attributes, globals/nonlocals, or async."
- "Preserve customer spelling."

Let's write down the function clearly:

```python
import csv
import io
import re

def solve(data):
    reader = csv.reader(io.StringIO(data))
    try:
        header = next(reader)
    except StopIteration:
        return {}
    
    header_map = {col.strip().lower(): i for i, col in enumerate(header)}
    cust_idx = header_map.get('customer')
    cents_idx = header_map.get('cents')
    if cust_idx is None or cents_idx is None:
        return {}
    
    pattern = re.compile(r'^[+-]?\d+$')
    totals = {}
    
    for row in reader:
        if len(row) <= max(cust_idx, cents_idx):
            continue
        cust = row[cust_idx]
        cents = row[cents_idx]
        if not cust or not cust.strip():
            continue
        if not pattern.match(cents):
            continue
        totals[cust] = totals.get(cust, 0) + int(cents)
        
    return totals
```

Everything looks solid and directly adheres to all instructions.import csv
import io
import re

def solve

39.25 seconds · $0.0308 estimated

Attempt 3 · Failed
import csv
import io
import re

def solve(data):
    if not data:
        return {}
    data = data.lstrip("\ufeff")
    reader = csv.DictReader(io.StringIO(data))
    totals = {}
    pattern = re.compile(r"[+-]?[0-9]+")
    for row in reader:
        customer = row.get("customer")
        cents = row.get("cents")
        if customer is None or cents is None:
            continue
        if not customer.strip():
            continue
        if not pattern.fullmatch(cents):
            continue
        totals[customer] = totals.get(customer, 0) + int(cents)
    return totals

21.48 seconds · $0.0177 estimated

Order deployment dependencies Algorithms

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is a dict mapping each named job to prerequisite names. Include names referenced only as prerequisites. Return the lexicographically smallest valid topological order, choosing the smallest ready name at every step. Return null for any cycle.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    import heapq

    graph = {}
    indegree = {}

    for job, prerequisites in data.items():
        graph.setdefault(job, set())
        indegree.setdefault(job, 0)

        for prerequisite in set(prerequisites):
            graph.setdefault(prerequisite, set()).add(job)
            indegree.setdefault(prerequisite, 0)
            indegree[job] += 1

    ready = [name for name, count in indegree.items() if count == 0]
    heapq.heapify(ready)
    order = []

    while ready:
        name = heapq.heappop(ready)
        order.append(name)

        for dependent in graph[name]:
            indegree[dependent] -= 1
            if indegree[dependent] == 0:
                heapq.heappush(ready, dependent)

    return order if len(order) == len(indegree) else None

9.50 seconds · $0.0044 estimated

Attempt 2 · Passed
import heapq


def solve(data):
    nodes = set(data)
    for prerequisites in data.values():
        nodes.update(prerequisites)

    indegree = {name: 0 for name in nodes}
    dependents = {name: [] for name in nodes}

    for job, prerequisites in data.items():
        for prerequisite in set(prerequisites):
            dependents[prerequisite].append(job)
            indegree[job] += 1

    ready = [name for name in nodes if indegree[name] == 0]
    heapq.heapify(ready)
    order = []

    while ready:
        name = heapq.heappop(ready)
        order.append(name)
        for job in dependents[name]:
            indegree[job] -= 1
            if indegree[job] == 0:
                heapq.heappush(ready, job)

    return order if len(order) == len(nodes) else None

4.25 seconds · $0.0023 estimated

Attempt 3 · Passed
def solve(data):
    import heapq

    names = set(data)
    for prerequisites in data.values():
        names.update(prerequisites)

    dependents = {name: [] for name in names}
    indegree = {name: 0 for name in names}

    for job, prerequisites in data.items():
        for prerequisite in set(prerequisites):
            dependents[prerequisite].append(job)
            indegree[job] += 1

    ready = [name for name in names if indegree[name] == 0]
    heapq.heapify(ready)
    order = []

    while ready:
        name = heapq.heappop(ready)
        order.append(name)
        for job in dependents[name]:
            indegree[job] -= 1
            if indegree[job] == 0:
                heapq.heappush(ready, job)

    return order if len(order) == len(names) else None

6.92 seconds · $0.0033 estimated

Claude Sonnet 5.5

Attempt 1 · Passed · Scoring corrected

Original verdict: failed. Corrected by offline replay after fixing the runner; the output is unchanged.

def solve(data):
    import heapq
    nodes = set()
    prereqs = {}
    for job, deps in data.items():
        nodes.add(job)
        if deps is None:
            deps = []
        elif isinstance(deps, (str, bytes)):
            deps = [deps]
        else:
            deps = list(deps)
        prereqs[job] = deps
        for d in deps:
            nodes.add(d)

    indeg = {n: 0 for n in nodes}
    children = {n: [] for n in nodes}
    for job, deps in prereqs.items():
        for d in set(deps):
            children[d].append(job)
            indeg[job] += 1

    heap = [n for n in nodes if indeg[n] == 0]
    heapq.heapify(heap)
    result = []
    while heap:
        n = heapq.heappop(heap)
        result.append(n)
        for c in children[n]:
            indeg[c] -= 1
            if indeg[c] == 0:
                heapq.heappush(heap, c)

    if len(result) != len(nodes):
        return None
    return result

3.16 seconds · $0.0040 estimated

Attempt 2 · Passed · Scoring corrected

Original verdict: failed. Corrected by offline replay after fixing the runner; the output is unchanged.

def solve(data):
    import heapq
    nodes = set()
    prereqs = {}
    for job, pre in data.items():
        nodes.add(job)
        if pre is None:
            pre = []
        elif isinstance(pre, (str, bytes)):
            pre = [pre]
        else:
            pre = list(pre)
        prereqs[job] = set(pre)
        for p in pre:
            nodes.add(p)
    indeg = {n: 0 for n in nodes}
    adj = {n: [] for n in nodes}
    for job, pres in prereqs.items():
        for p in pres:
            adj[p].append(job)
            indeg[job] += 1
    heap = [n for n in nodes if indeg[n] == 0]
    heapq.heapify(heap)
    result = []
    while heap:
        n = heapq.heappop(heap)
        result.append(n)
        for m in adj[n]:
            indeg[m] -= 1
            if indeg[m] == 0:
                heapq.heappush(heap, m)
    if len(result) != len(nodes):
        return None
    return result

2.85 seconds · $0.0041 estimated

Attempt 3 · Passed · Scoring corrected

Original verdict: failed. Corrected by offline replay after fixing the runner; the output is unchanged.

def solve(data):
    import heapq
    nodes = set()
    prereqs = {}
    for job, deps in data.items():
        nodes.add(job)
        if deps is None:
            deps = []
        elif isinstance(deps, (str, bytes)):
            deps = [deps]
        dl = list(deps)
        prereqs[job] = dl
        for d in dl:
            nodes.add(d)
    adj = {n: set() for n in nodes}
    indeg = {n: 0 for n in nodes}
    for job, dl in prereqs.items():
        for d in set(dl):
            if job in adj[d]:
                continue
            adj[d].add(job)
            indeg[job] += 1
    heap = [n for n in nodes if indeg[n] == 0]
    heapq.heapify(heap)
    result = []
    while heap:
        n = heapq.heappop(heap)
        result.append(n)
        for m in adj[n]:
            indeg[m] -= 1
            if indeg[m] == 0:
                heapq.heappush(heap, m)
    if len(result) != len(nodes):
        return None
    return result

3.24 seconds · $0.0042 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
import collections
import heapq

def solve(data):
    nodes = set(data.keys())
    for prereqs in data.values():
        nodes.update(prereqs)

    adj = collections.defaultdict(list)
    in_degree = {node: 0 for node in nodes}

    for job, prereqs in data.items():
        unique_prereqs = set(prereqs)
        in_degree[job] += len(unique_prereqs)
        for p in unique_prereqs:
            adj[p].append(job)

    heap = [node for node, deg in in_degree.items() if deg == 0]
    heapq.heapify(heap)

    result = []
    while heap:
        curr = heapq.heappop(heap)
        result.append(curr)
        for neighbor in adj[curr]:
            in_degree[neighbor] -= 1
            if in_degree[neighbor] == 0:
                heapq.heappush(heap, neighbor)

    if len(result) < len(nodes):
        return None

    return result

7.13 seconds · $0.0056 estimated

Attempt 2 · Passed
import collections
import heapq

def solve(data):
    nodes = set(data.keys())
    adj = collections.defaultdict(list)
    in_degree = {k: 0 for k in data}
    
    for job, prereqs in data.items():
        unique_prereqs = set(prereqs)
        in_degree[job] = len(unique_prereqs)
        for p in unique_prereqs:
            nodes.add(p)
            adj[p].append(job)
            if p not in in_degree:
                in_degree[p] = 0

    heap = [node for node in nodes if in_degree[node] == 0]
    heapq.heapify(heap)
    
    order = []
    while heap:
        u = heapq.heappop(heap)
        order.append(u)
        for v in adj[u]:
            in_degree[v] -= 1
            if in_degree[v] == 0:
                heapq.heappush(heap, v)
                
    if len(order) == len(nodes):
        return order
    return None

10.13 seconds · $0.0089 estimated

Attempt 3 · Passed
import heapq

def solve(data):
    nodes = set(data.keys())
    for prereqs in data.values():
        nodes.update(prereqs)

    adj = {node: [] for node in nodes}
    in_degree = {node: 0 for node in nodes}

    for job, prereqs in data.items():
        for p in set(prereqs):
            adj[p].append(job)
            in_degree[job] += 1

    heap = [node for node in nodes if in_degree[node] == 0]
    heapq.heapify(heap)

    order = []
    while heap:
        u = heapq.heappop(heap)
        order.append(u)
        for v in adj[u]:
            in_degree[v] -= 1
            if in_degree[v] == 0:
                heapq.heappush(heap, v)

    if len(order) < len(nodes):
        return None
    return order

8.62 seconds · $0.0073 estimated

Value remaining FIFO inventory Data correctness

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is {lots:[[quantity,unit_cents],...], sold:nonnegative integer}. Sell in lot order. Return {remaining_units,remaining_value_cents,unfilled_units}. All quantities/prices nonnegative; unfilled is sales beyond available inventory.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    to_sell = data["sold"]
    remaining_units = 0
    remaining_value_cents = 0

    for quantity, unit_cents in data["lots"]:
        sold_from_lot = min(quantity, to_sell)
        to_sell -= sold_from_lot
        remaining = quantity - sold_from_lot
        remaining_units += remaining
        remaining_value_cents += remaining * unit_cents

    return {
        "remaining_units": remaining_units,
        "remaining_value_cents": remaining_value_cents,
        "unfilled_units": to_sell,
    }

3.48 seconds · $0.0016 estimated

Attempt 2 · Passed
def solve(data):
    unfilled = data["sold"]
    remaining_units = 0
    remaining_value_cents = 0

    for quantity, unit_cents in data["lots"]:
        sold_from_lot = min(quantity, unfilled)
        unfilled -= sold_from_lot
        remaining = quantity - sold_from_lot
        remaining_units += remaining
        remaining_value_cents += remaining * unit_cents

    return {
        "remaining_units": remaining_units,
        "remaining_value_cents": remaining_value_cents,
        "unfilled_units": unfilled,
    }

3.10 seconds · $0.0016 estimated

Attempt 3 · Passed
def solve(data):
    unsold_demand = data["sold"]
    remaining_units = 0
    remaining_value_cents = 0

    for quantity, unit_cents in data["lots"]:
        sold_from_lot = min(quantity, unsold_demand)
        unsold_demand -= sold_from_lot
        remaining = quantity - sold_from_lot
        remaining_units += remaining
        remaining_value_cents += remaining * unit_cents

    return {
        "remaining_units": remaining_units,
        "remaining_value_cents": remaining_value_cents,
        "unfilled_units": unsold_demand,
    }

3.09 seconds · $0.0016 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
def solve(data):
    lots = data.get("lots", []) or []
    sold = data.get("sold", 0) or 0
    to_sell = sold
    remaining_units = 0
    remaining_value = 0
    for lot in lots:
        q = lot[0]
        p = lot[1]
        take = q if q < to_sell else to_sell
        to_sell -= take
        left = q - take
        remaining_units += left
        remaining_value += left * p
    return {
        "remaining_units": remaining_units,
        "remaining_value_cents": remaining_value,
        "unfilled_units": to_sell,
    }

2.14 seconds · $0.0025 estimated

Attempt 2 · Passed
def solve(data):
    lots = data.get("lots", []) or []
    sold = data.get("sold", 0) or 0
    to_sell = sold
    remaining_units = 0
    remaining_value = 0
    for lot in lots:
        q = lot[0]
        p = lot[1]
        take = q if q < to_sell else to_sell
        to_sell -= take
        left = q - take
        remaining_units += left
        remaining_value += left * p
    return {
        "remaining_units": remaining_units,
        "remaining_value_cents": remaining_value,
        "unfilled_units": to_sell,
    }

2.17 seconds · $0.0025 estimated

Attempt 3 · Passed
def solve(data):
    lots = data.get("lots", []) or []
    sold = data.get("sold", 0) or 0
    to_sell = sold
    remaining_units = 0
    remaining_value = 0
    for lot in lots:
        q = lot[0]
        p = lot[1]
        take = q if q < to_sell else to_sell
        to_sell -= take
        left = q - take
        remaining_units += left
        remaining_value += left * p
    return {
        "remaining_units": remaining_units,
        "remaining_value_cents": remaining_value,
        "unfilled_units": to_sell,
    }

2.15 seconds · $0.0025 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
def solve(data):
    lots = data.get('lots', [])
    to_sell = data.get('sold', 0)
    rem_units = 0
    rem_value = 0
    for q, p in lots:
        if to_sell >= q:
            to_sell -= q
        elif to_sell > 0:
            left_q = q - to_sell
            to_sell = 0
            rem_units += left_q
            rem_value += left_q * p
        else:
            rem_units += q
            rem_value += q * p
    return {
        'remaining_units': rem_units,
        'remaining_value_cents': rem_value,
        'unfilled_units': to_sell
    }

3.60 seconds · $0.0040 estimated

Attempt 2 · Passed
def solve(data):
    lots = data.get('lots', [])
    sold = data.get('sold', 0)
    
    to_sell = sold
    remaining_units = 0
    remaining_value_cents = 0
    
    for qty, unit_cents in lots:
        if to_sell > 0:
            take = min(qty, to_sell)
            qty_left = qty - take
            to_sell -= take
        else:
            qty_left = qty
        remaining_units += qty_left
        remaining_value_cents += qty_left * unit_cents
        
    return {
        'remaining_units': remaining_units,
        'remaining_value_cents': remaining_value_cents,
        'unfilled_units': to_sell
    }

4.85 seconds · $0.0044 estimated

Attempt 3 · Passed
def solve(data):
    lots = data.get("lots", [])
    remaining_to_sell = data.get("sold", 0)
    
    remaining_units = 0
    remaining_value_cents = 0
    
    for lot in lots:
        qty, price = lot[0], lot[1]
        if remaining_to_sell > 0:
            take = min(qty, remaining_to_sell)
            qty -= take
            remaining_to_sell -= take
        remaining_units += qty
        remaining_value_cents += qty * price
        
    return {
        "remaining_units": remaining_units,
        "remaining_value_cents": remaining_value_cents,
        "unfilled_units": remaining_to_sell
    }

4.58 seconds · $0.0039 estimated

Calculate service latency percentiles Algorithms

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is {values:[finite numbers,...],p:number in [0,100]}. Return nearest-rank percentile: sorted values at 1-based rank max(1,ceil(p/100*n)). Return null for no values. Do not interpolate or mutate input.

GPT-6.1 Sol

Attempt 1 · Passed
import math

def solve(data):
    values = sorted(data["values"])
    if not values:
        return None
    rank = max(1, math.ceil(data["p"] / 100 * len(values)))
    return values[rank - 1]

5.13 seconds · $0.0014 estimated

Attempt 2 · Passed
import math

def solve(data):
    values = sorted(data["values"])
    if not values:
        return None
    rank = max(1, math.ceil(data["p"] / 100 * len(values)))
    return values[rank - 1]

4.56 seconds · $0.0015 estimated

Attempt 3 · Passed
import math

def solve(data):
    values = sorted(data["values"])
    if not values:
        return None
    rank = max(1, math.ceil(data["p"] / 100 * len(values)))
    return values[rank - 1]

4.10 seconds · $0.0017 estimated

Claude Sonnet 5.5

Attempt 1 · Failed
def solve(data):
    values = data.get("values") if isinstance(data, dict) else None
    if not values:
        return None
    p = data.get("p", 0)
    s = sorted(values)
    n = len(s)
    rank = math.ceil(p / 100 * n)
    if rank < 1:
        rank = 1
    if rank > n:
        rank = n
    return s[rank - 1]

Wait: math must be imported. Here is the corrected final code:

import math

def solve(data):
    values = data.get("values") if isinstance(data, dict) else None
    if not values:
        return None
    p = data.get("p", 0)
    s = sorted(values)
    n = len(s)
    rank = math.ceil(p / 100 * n)
    if rank < 1:
        rank = 1
    if rank > n:
        rank = n
    return s[rank - 1]

2.76 seconds · $0.0033 estimated

Attempt 2 · Failed
def solve(data):
    values = data.get("values") if isinstance(data, dict) else None
    if not values:
        return None
    p = data.get("p", 0)
    s = sorted(values)
    n = len(s)
    rank = math.ceil(p / 100 * n)
    if rank < 1:
        rank = 1
    if rank > n:
        rank = n
    return s[rank - 1]

Wait, the math module must be imported. Here is the complete code:

import math

def solve(data):
    values = data.get("values") if isinstance(data, dict) else None
    if not values:
        return None
    p = data.get("p", 0)
    s = sorted(values)
    n = len(s)
    rank = math.ceil(p / 100 * n)
    if rank < 1:
        rank = 1
    if rank > n:
        rank = n
    return s[rank - 1]

2.57 seconds · $0.0033 estimated

Attempt 3 · Failed
def solve(data):
    values = data.get("values") if isinstance(data, dict) else None
    if not values:
        return None
    p = data.get("p", 0)
    s = sorted(values)
    n = len(s)
    rank = math.ceil(p / 100 * n)
    if rank < 1:
        rank = 1
    if rank > n:
        rank = n
    return s[rank - 1]

Wait, math must be imported; here is the complete version:

import math

def solve(data):
    values = data.get("values") if isinstance(data, dict) else None
    if not values:
        return None
    p = data.get("p", 0)
    s = sorted(values)
    n = len(s)
    rank = math.ceil(p / 100 * n)
    if rank < 1:
        rank = 1
    if rank > n:
        rank = n
    return s[rank - 1]

2.62 seconds · $0.0033 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
import math

def solve(data):
    values = data.get("values")
    if not values:
        return None
    p = data.get("p")
    n = len(values)
    rank = max(1, math.ceil(p / 100 * n))
    sorted_values = sorted(values)
    return sorted_values[rank - 1]

6.36 seconds · $0.0057 estimated

Attempt 2 · Passed
import math

def solve(data):
    values = data.get("values")
    if not values:
        return None
    p = data.get("p", 0)
    n = len(values)
    rank = max(1, math.ceil(p / 100 * n))
    return sorted(values)[rank - 1]

6.51 seconds · $0.0060 estimated

Attempt 3 · Passed
import math

def solve(data):
    values = data.get('values')
    if not values:
        return None
    p = data.get('p')
    n = len(values)
    rank = max(1, math.ceil((p / 100) * n))
    sorted_values = sorted(values)
    return sorted_values[rank - 1]

3.82 seconds · $0.0030 estimated

Normalize Unicode article slugs Safety and edge cases

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is a string. Apply Unicode NFKD normalization, drop combining marks, lowercase, retain only ASCII a-z and digits, replace each run of other characters with one hyphen, strip edge hyphens. Return the slug.

GPT-6.1 Sol

Attempt 1 · Passed
import unicodedata
import re

def solve(data):
    normalized = unicodedata.normalize("NFKD", data)
    text = "".join(char for char in normalized if not unicodedata.combining(char))
    return re.sub(r"[^a-z0-9]+", "-", text.lower()).strip("-")

2.09 seconds · $0.0010 estimated

Attempt 2 · Passed
import re
import unicodedata

def solve(data):
    normalized = unicodedata.normalize("NFKD", data)
    normalized = "".join(c for c in normalized if not unicodedata.combining(c))
    return re.sub(r"[^a-z0-9]+", "-", normalized.lower()).strip("-")

1.97 seconds · $0.0010 estimated

Attempt 3 · Passed
import unicodedata

def solve(data):
    normalized = unicodedata.normalize("NFKD", data)
    normalized = "".join(
        char for char in normalized if not unicodedata.combining(char)
    ).lower()

    result = []
    pending_hyphen = False
    for char in normalized:
        if "a" <= char <= "z" or "0" <= char <= "9":
            if pending_hyphen and result:
                result.append("-")
            result.append(char)
            pending_hyphen = False
        else:
            pending_hyphen = True

    return "".join(result)

3.17 seconds · $0.0016 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
import unicodedata

def solve(data):
    s = unicodedata.normalize("NFKD", data)
    s = "".join(ch for ch in s if not unicodedata.combining(ch))
    s = s.lower()
    out = []
    pending = False
    for ch in s:
        if ("a" <= ch <= "z") or ("0" <= ch <= "9"):
            if pending and out:
                out.append("-")
            pending = False
            out.append(ch)
        else:
            pending = True
    return "".join(out)

2.20 seconds · $0.0021 estimated

Attempt 2 · Passed
import unicodedata
import re

def solve(data):
    s = unicodedata.normalize("NFKD", data)
    s = "".join(c for c in s if not unicodedata.combining(c))
    s = s.lower()
    s = re.sub(r"[^a-z0-9]+", "-", s)
    return s.strip("-")

1.77 seconds · $0.0015 estimated

Attempt 3 · Passed
import unicodedata

def solve(data):
    s = unicodedata.normalize("NFKD", data)
    s = "".join(ch for ch in s if not unicodedata.combining(ch))
    s = s.lower()
    out = []
    pending = False
    for ch in s:
        if ("a" <= ch <= "z") or ("0" <= ch <= "9"):
            if pending and out:
                out.append("-")
            pending = False
            out.append(ch)
        else:
            pending = True
    return "".join(out)

2.19 seconds · $0.0021 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
import re
import unicodedata

def solve(data):
    normalized = unicodedata.normalize('NFKD', data)
    without_marks = ''.join(c for c in normalized if not unicodedata.combining(c))
    lowercased = without_marks.lower()
    slug = re.sub(r'[^a-z0-9]+', '-', lowercased)
    return slug.strip('-')

5.01 seconds · $0.0030 estimated

Attempt 2 · Passed
import re
import unicodedata

def solve(data):
    normalized = unicodedata.normalize('NFKD', data)
    without_marks = ''.join(
        c for c in normalized
        if not unicodedata.combining(c) and not unicodedata.category(c).startswith('M')
    )
    lowered = without_marks.lower()
    return re.sub(r'[^a-z0-9]+', '-', lowered).strip('-')

6.51 seconds · $0.0059 estimated

Attempt 3 · Passed
import re
import unicodedata

def solve(data: str) -> str:
    normalized = unicodedata.normalize('NFKD', data)
    without_marks = ''.join(
        c for c in normalized
        if not unicodedata.combining(c) and not unicodedata.category(c).startswith('M')
    )
    lowered = without_marks.lower()
    slug = re.sub(r'[^a-z0-9]+', '-', lowered)
    return slug.strip('-')

11.14 seconds · $0.0094 estimated

Apply a per-user rolling rate limit Safety and edge cases

Implement solve(data) in Python 3.9. Allowed imports: math, json, csv, io, re, collections, heapq, unicodedata. Use functions and built-in types; no custom classes, private/dunder attributes, globals/nonlocals, or async. Return only Python code, with no Markdown. Do not read files, use the network, run processes or access system state. Input is {limit:positive integer,window:positive integer,events:[{user,time},...]}. Times are nondecreasing integer seconds. Process in input order. Accept only when fewer than limit previously accepted events for that user have times strictly greater than time-window. Rejections do not consume allowance. Return a list of booleans.

GPT-6.1 Sol

Attempt 1 · Passed
def solve(data):
    from collections import deque

    limit = data["limit"]
    window = data["window"]
    accepted = {}
    result = []

    for event in data["events"]:
        user = event["user"]
        time = event["time"]
        cutoff = time - window

        if user not in accepted:
            accepted[user] = deque()
        times = accepted[user]

        while times and times[0] <= cutoff:
            times.popleft()

        allowed = len(times) < limit
        result.append(allowed)
        if allowed:
            times.append(time)

    return result

3.03 seconds · $0.0016 estimated

Attempt 2 · Passed
from collections import deque

def solve(data):
    limit = data["limit"]
    window = data["window"]
    histories = {}
    results = []

    for event in data["events"]:
        user = event["user"]
        time = event["time"]
        history = histories.setdefault(user, deque())
        cutoff = time - window

        while history and history[0] <= cutoff:
            history.popleft()

        accepted = len(history) < limit
        results.append(accepted)
        if accepted:
            history.append(time)

    return results

2.95 seconds · $0.0015 estimated

Attempt 3 · Passed
from collections import deque

def solve(data):
    limit = data["limit"]
    window = data["window"]
    accepted_by_user = {}
    result = []

    for event in data["events"]:
        user = event["user"]
        time = event["time"]
        queue = accepted_by_user.setdefault(user, deque())
        cutoff = time - window

        while queue and queue[0] <= cutoff:
            queue.popleft()

        accepted = len(queue) < limit
        result.append(accepted)
        if accepted:
            queue.append(time)

    return result

3.94 seconds · $0.0015 estimated

Claude Sonnet 5.5

Attempt 1 · Passed
def solve(data):
    from collections import deque

    limit = data["limit"]
    window = data["window"]
    events = data["events"]
    accepted = {}
    result = []
    for ev in events:
        user = ev["user"]
        t = ev["time"]
        dq = accepted.get(user)
        if dq is None:
            dq = deque()
            accepted[user] = dq
        # drop accepted events with time <= t - window
        cutoff = t - window
        while dq and dq[0] <= cutoff:
            dq.popleft()
        if len(dq) < limit:
            dq.append(t)
            result.append(True)
        else:
            result.append(False)
    return result

2.33 seconds · $0.0027 estimated

Attempt 2 · Passed
def solve(data):
    from collections import deque
    limit = data["limit"]
    window = data["window"]
    events = data["events"]
    accepted = {}
    result = []
    for ev in events:
        user = ev["user"]
        t = ev["time"]
        dq = accepted.get(user)
        if dq is None:
            dq = deque()
            accepted[user] = dq
        # drop accepted times that are no longer strictly greater than t - window
        while dq and dq[0] <= t - window:
            dq.popleft()
        if len(dq) < limit:
            dq.append(t)
            result.append(True)
        else:
            result.append(False)
    return result

2.19 seconds · $0.0027 estimated

Attempt 3 · Passed
def solve(data):
    from collections import deque
    limit = data["limit"]
    window = data["window"]
    events = data["events"]
    accepted = {}
    result = []
    for ev in events:
        user = ev["user"]
        t = ev["time"]
        dq = accepted.get(user)
        if dq is None:
            dq = deque()
            accepted[user] = dq
        cutoff = t - window
        while dq and dq[0] <= cutoff:
            dq.popleft()
        if len(dq) < limit:
            dq.append(t)
            result.append(True)
        else:
            result.append(False)
    return result

2.16 seconds · $0.0025 estimated

Gemini 3.8 Flash

Attempt 1 · Passed
import collections

def solve(data):
    limit = data['limit']
    window = data['window']
    events = data['events']
    
    user_history = collections.defaultdict(collections.deque)
    result = []
    
    for event in events:
        user = event['user']
        current_time = event['time']
        dq = user_history[user]
        
        threshold = current_time - window
        while dq and dq[0] <= threshold:
            dq.popleft()
            
        if len(dq) < limit:
            dq.append(current_time)
            result.append(True)
        else:
            result.append(False)
            
    return result

5.81 seconds · $0.0046 estimated

Attempt 2 · Passed
import collections

def solve(data):
    limit = data["limit"]
    window = data["window"]
    events = data["events"]

    history = collections.defaultdict(collections.deque)
    result = []

    for event in events:
        user = event["user"]
        t = event["time"]
        dq = history[user]
        cutoff = t - window
        while dq and dq[0] <= cutoff:
            dq.popleft()
        if len(dq) < limit:
            dq.append(t)
            result.append(True)
        else:
            result.append(False)

    return result

5.90 seconds · $0.0054 estimated

Attempt 3 · Passed
import collections


def solve(data):
    limit = data["limit"]
    window = data["window"]
    events = data["events"]

    user_history = collections.defaultdict(collections.deque)
    result = []

    for event in events:
        user = event["user"]
        t = event["time"]
        dq = user_history[user]
        cutoff = t - window

        while dq and dq[0] <= cutoff:
            dq.popleft()

        if len(dq) < limit:
            dq.append(t)
            result.append(True)
        else:
            result.append(False)

    return result

5.73 seconds · $0.0054 estimated

Reproduce this run

Download original JSON evidence · Download original CSV results

Scoring audit: 2026-10-04T18:18:02.762014+00:00 · 3 verdicts changed · zero new provider requests. Download appended correction.

Run history

  • 2026-10-04T08:20:14.082456+00:00 · complete · 108 / 108 required trials initiated · Archived evidence

Maintenance needs recovery: The daily source check is overdue. Published test dates remain unchanged.

Source-check status: A source fetch has completed; factual changes require review. Latest attempt: 2026-10-04T09:01:57.926856+00:00.

Download the public task bank · Read the methodology