Skip to content
Kingy AI

Kingy / AI Fix Finder

← All fixes

Gemma 4 fails to allocate a 64K context on a 24 GB RTX 4090

The retained run passed after changing KV cache precision from FP16 to Q8_0.

kingy tested

What happened

Symptom

cudaMalloc cannot allocate a 1,200 MiB KV-cache buffer.

Conditions

Gemma 4 31B-it Q4_K_M

Ryzen 9 7950X / RTX 4090 24 GB / 128 GB RAM

65,536-token allocated context

Correction

Change both KV caches to Q8_0 for the same 64K workload, or reduce context to the retained passing 32K FP16 configuration.

Result

The same 64K workload passed with Q8_0 KV; 32K FP16 KV also passed.

What did not work

64K FP16 KV allocation failed in the retained run.

  1. Inspect the retained result and receipt before choosing runtime options.
  2. Keep model, runtime, hardware and other settings fixed when testing the change.
  3. Select Q8_0 KV cache for the 64K context.
  4. Record allocation success and task quality separately; retain the failed attempt.
Sources & limits

Gemma 4 31B-it owned RTX 4090 measured result

Record checked 2026-08-23.

Allocation success is not a task-quality assessment.

The exact reviewed hardware configuration matters; available memory differs across systems.

This still happens?

Send the smallest example that reproduces it. Use a public, non-sensitive example; leave out account details, keys and private files.

Help improve this tool

Did this help you finish?

About AI Fix Finder. Your feedback goes to Kingy’s private review queue.

Your experience with AI Fix Finder
What happened?