
- Selected model attempts
- 3
- Selected model charges
- $0.0000 USD
- Human editing
- Not recorded
kingy tested
Original accepted output
{
"owner": "Jonah",
"launch_date": "2026-09-19",
"budget_usd": 900
}Qwen3-4B-Q4_K_M.gguf · seed 4101. Retained output for the original input. Changing the recipe does not generate or verify a new result.
All retained attempts for this task
| Exact model file | Accepted / attempts |
|---|---|
| Qwen3-4B-Q4_K_M.gguf | 3 / 3 |
| Qwen3-0.6B-Q8_0.gguf | 1 / 3 |
| Qwen3-1.7B-Q8_0.gguf | 0 / 3 |
Inspect the measured M4 Pro configuration →
Reproduce the outcome
- Download and extract the corrected-owner/date recipe package.
- Run python3 check_recipe.py. All nine historical scoring decisions must reproduce without inference, network access or installs.
- Inspect original-prompt.txt, expected.json and selected-output.txt; the selected output is Qwen3-4B attempt 1. The raw folder retains all nine original requests and responses.
- Change the original and corrected values in Make my version. Its exported instructions include the updated expected-answer JSON. New inputs have not been tested.
- If you choose to run a new local inference, keep the entire response, including any Markdown fences, in output.txt. Save the exported expected-answer JSON as expected-custom.json.
- Run python3 check_recipe.py --output output.txt --expected expected-custom.json. Retain failures and observed charges/time; never copy the historical success rate to a new input.
Original input and exact settings
- Original synthetic Monday note: Priya, September 12, 2026, budget USD 900.
- Synthetic Tuesday correction: Jonah, September 19, 2026; budget unchanged.
- temperature
0.6- top_p
0.95- max_tokens
256- context
4096- seeds
[4101,4102,4103]
Original prompt
Return JSON only with owner, launch_date and budget_usd. Monday note: owner was Priya, launch Sep 12 2026, budget 900. Correction Tuesday: owner is Jonah; launch moves to Sep 19 2026. Budget stays unchanged. Use final values; date format YYYY-MM-DD. /no_think
Attribution and rights: Kingy synthetic task and actual redacted output evidence; model weights remain under their separate license and are not redistributed.
Sources & limits
Small Qwen3 models: retained JSON pilot
Record checked 2026-09-07.
The selected Qwen3-4B file passed 3/3; 0.6B passed 1/3 and 1.7B passed 0/3 on this task. Three repetitions are not a reliability guarantee. Quantizations differ.
Only the original synthetic notes were tested. Adapted inputs and their newly constructed expected answers are untested.
Provider charges were $0; hardware, electricity and active human repair costs remain unrecorded.
Original formatting failures remain failed. Checking a corrected output later does not rewrite the retained scores.