A 35-billion-parameter model on an 8 GB graphics card sounds like a typo. The 4-bit file alone is 22 GB. But Qwen3.6-35B-A3B, which Alibaba’s Qwen team released under Apache 2.0 in April 2026, is built in a way that makes it one of the best models you can run on ordinary hardware, if you put the right parts in the right memory.
Why it fits: most of the model sits idle
Qwen3.6-35B-A3B is a mixture-of-experts model. Each of its 40 layers holds 256 small “expert” networks, and for every token a router picks 8 of them, plus one shared expert that always runs. So although the model has about 36 billion parameters, only about 3.5 billion do any work on a given token.
The split is lopsided. Reading the parameter counts from the published weights, 92.9% of the text model is routed experts. Only about 2.45 billion parameters (the attention layers, the router, the shared expert and the embeddings) are needed for every token. In the Unsloth UD-Q4_K_XL file, that always-on part is about 2 GiB. The experts are the other 18 GiB.
That is the whole idea. Keep the small always-on part on the GPU, where it runs fast. Park the experts in system RAM, and let the CPU run the few that each token needs. Move as many expert layers back onto the GPU as your card has room for.
Why long context is cheap: most layers don’t keep a KV cache
The second reason is the attention design. Only one layer in four uses full attention and keeps a key-value cache. The other thirty are Gated DeltaNet layers, which keep a small state of fixed size no matter how long the conversation gets.
With 10 full-attention layers, 2 KV heads and a head size of 256, the cache costs 20 KiB per token at 16-bit precision, or about 10.6 KiB at 8-bit (q8_0). A full 131,072-token context fits in about 1.3 GiB. The model card advises “maintaining a context length of at least 128K tokens to preserve thinking capabilities”, and on this model that’s affordable even on a small card.
| KV cache type | 8K tokens | 32K | 128K | 256K |
|---|---|---|---|---|
| f16 | 160 MiB | 640 MiB | 2,560 MiB | 5,120 MiB |
| q8_0 | 85 MiB | 340 MiB | 1,360 MiB | 2,720 MiB |
| q4_0 | 45 MiB | 180 MiB | 720 MiB | 1,440 MiB |
The budget, card by card
Put those together for UD-Q4_K_XL with a 131K context, an 8-bit KV cache and at least 1 GiB left free on the card. The fixed cost on the GPU is about 3.9 GiB: 2.0 GiB of always-on weights, 1.3 GiB of KV cache, and about 0.6 GiB of runtime buffers. Everything else on the card is expert layers.
| Card | Always-on weights | Expert layers on GPU | KV cache | Runtime buffers | Free |
|---|---|---|---|---|---|
| 8 GB card | 2.0 GiB | 2.8 GiB | 1.3 GiB | 0.6 GiB | 1.3 GiB |
| 12 GB card | 2.0 GiB | 6.9 GiB | 1.3 GiB | 0.6 GiB | 1.2 GiB |
| 16 GB card | 2.0 GiB | 11.0 GiB | 1.3 GiB | 0.6 GiB | 1.1 GiB |
Source: Calculated from the model's config and the GGUF tensor sizes; the method reproduces a real llama.cpp memory log to the MiB (llama.cpp issue 25721). Not measured on these cards.
The system RAM figures are for llama.cpp alone. The server also keeps a prompt cache in RAM (up to 8 GiB by default), and your operating system needs room too, so plan for a 32 GB machine with an 8 GB card.
The command
With llama.cpp’s llama-server, one line does it. This is the 8 GB version; for 12 GB use --n-cpu-moe 25, and for 16 GB use --n-cpu-moe 16.
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_XL --no-mmproj \
-ngl 99 --n-cpu-moe 34 -fa on -c 131072 -ctk q8_0 -ctv q8_0 -np 1
What each part does:
-ngl 99puts every layer on the GPU to start with.--n-cpu-moe 34then moves the expert weights of the first 34 layers back to the CPU. This is the one number you tune: lower it until the GPU is nearly full.-c 131072sets the context, and-ctk q8_0 -ctv q8_0stores the KV cache at 8 bits.-fa onturns on flash attention.--no-mmprojskips the vision encoder, which-hfwould otherwise download and load onto the GPU.-np 1serves one conversation at a time, so memory isn’t split across parallel slots.
If you’d rather not do the arithmetic, current llama.cpp can place things itself. Set only the context with -c, and its automatic fitting (on by default) keeps the dense weights on the GPU and spills experts to RAM. Setting -ngl or --n-cpu-moe yourself turns that off.
How fast is it?
These are other people’s measurements, all on Qwen3.6-35B-A3B with mainline llama.cpp unless marked:
| GPU | Machine | Setup | Prompt (tok/s) | Output (tok/s) |
|---|---|---|---|---|
| RTX 3060 12 GB | i7-7700, 32 GB DDR4-2133 | UD-Q4_K_M, -ncmoe 24 |
413 | 38.9 |
| RTX 3060 12 GB | Ryzen 5600X, 32 GB DDR4 | Q4_K_M, -ncmoe 26, larger batch |
about 1,143 | n/a |
| RTX 4060 Ti 8 GB | Ryzen 9 7900X, DDR5 | Q4_K_S, q8_0 KV | n/a | 41.6 at 16K context, 24 at 200K |
| RTX 4070 12 GB | 64 GB DDR5 | Qwen3.5 (the previous model), Q4_K_M | 1,190 | 41.2 |
The pattern is consistent: around 40 tokens a second on a mid-range card, which is faster than most people read. On the same RTX 3060, moving more expert layers onto the GPU took generation from 28.1 tokens a second (all 40 on the CPU) to 42.6 (20 on the CPU).
With a mixture-of-experts model, the question isn’t whether it fits on your GPU. It’s how fast your system RAM is.
What to watch
- Generation speed is set by your RAM. Every token reads the active experts from system memory, so faster RAM means faster text. At the 8 GB setting, that’s roughly 490 MiB read from RAM per token.
- Long prompts are set by PCIe. Processing a big prompt copies the CPU-side experts to the GPU in batches. A larger batch helps prompt speed but uses more VRAM.
- Don’t squeeze the KV cache too hard. 8-bit is effectively lossless in current llama.cpp. 4-bit has hurt long reasoning in the one test I could find (on a different model, gpt-oss-20b: 37.9% on AIME25 at 16-bit, 21.7% at 4-bit). If you need memory, shorten the context before dropping to
q4_0. - It thinks by default. Qwen3.6 reasons before answering unless you pass
--reasoning off. The model card gives separate sampling settings for thinking and non-thinking modes; use those, and note that Qwen’s card and Unsloth’s guide disagree on the presence penalty for thinking mode. - Ollama and LM Studio. Ollama has no documented way to put only the experts on the CPU. LM Studio has one: a setting to force expert weights onto the CPU, with a per-layer slider in recent versions.
References
- Qwen3.6-35B-A3B model card and config Qwen, Hugging Face, April 2026.
- Qwen3.6-35B-A3B GGUF files Unsloth, Hugging Face.
- Qwen3.6: how to run locally Unsloth docs.
- llama-server README llama.cpp, September 2026.
- PR 15077: --n-cpu-moe llama.cpp, 4 August 2025.
- PR 16653: automatic fitting (--fit) llama.cpp, 15 December 2025.
- PR 28334: remove --mmap and --mlock in favour of --load-mode llama.cpp, 9 September 2026.
- PR 21038: KV cache rotation (and the AIME25 comment) llama.cpp, April 2026.
- Issue 25721: memory log used to check the budget method llama.cpp, July 2026.
- Issue 25859: RTX 3060 prompt-processing results llama.cpp, July 2026.
- Best way to run Qwen 3.6 35B MoE locally InsiderLLM, updated 22 September 2026.
- RTX 4060 Ti 8 GB results thread @above_spec (mirror), 7 May 2026.
- Qwen3.5 on an RTX 4070 with three flags Ken Imoto, 19 July 2026.
- Ollama FAQ Ollama.
- LM Studio 0.3.23 changelog LM Studio, 12 August 2025.