kingy tested
What happened
Symptom
cudaMalloc cannot allocate a 1,200 MiB KV-cache buffer.
Conditions
Gemma 4 31B-it Q4_K_M
Ryzen 9 7950X / RTX 4090 24 GB / 128 GB RAM
65,536-token allocated context
Correction
Change both KV caches to Q8_0 for the same 64K workload, or reduce context to the retained passing 32K FP16 configuration.
Result
The same 64K workload passed with Q8_0 KV; 32K FP16 KV also passed.
What did not work
64K FP16 KV allocation failed in the retained run.
- Inspect the retained result and receipt before choosing runtime options.
- Keep model, runtime, hardware and other settings fixed when testing the change.
- Select Q8_0 KV cache for the 64K context.
- Record allocation success and task quality separately; retain the failed attempt.
Sources & limits
Gemma 4 31B-it owned RTX 4090 measured result
Record checked 2026-08-23.
Allocation success is not a task-quality assessment.
The exact reviewed hardware configuration matters; available memory differs across systems.