A close-up of a circuit board's dense circuitry in shallow focus, traces and surface-mounted components filling the frame in cool blue-grey tones.
LLM

Quantization roulette on consumer Blackwell: when a good model outputs nothing but exclamation marks

Bob OdellSeptember 14, 202618 min read

Same book, different printing presses.

The text is identical. One press turns out a clean, readable copy. The next turns out page after page of the same character repeated until the paper runs out. A third prints a perfect copy, slowly, while the dial on the front insists it is running at full speed.

You would not conclude from any of that that the book was badly written. On a graphics card you will conclude exactly that, and you will spend a weekend doing it.

That is the whole post. On consumer Blackwell, the quantization format is not a quality dial. It is a compatibility gate, and it fails in ways that look precisely like the model being broken.

I run a coding agent on my own hardware: two RTX 5090s, 32 GB of VRAM each, on consumer Blackwell silicon that reports itself as sm_120. I have written before about building and tuning the box, about making its inference server event-triggered instead of burning two CPU cores while idle, about the engine bake-off that prefix caching decided, about the attention architecture that made every agent turn re-read 162,000 tokens, and about what actually fits on 64 GB of VRAM.

This one is about the formats. A widely used, well-regarded 32B coding model produced nothing on this box but exclamation marks, on two separate kernel paths. A different model in FP8 loaded without a word of complaint and returned garbage. Neither was a bad model. Both were a format meeting silicon that had no working kernel for it.

Everything I measured below was measured on one box, between 2 and 5 June 2026. Read those dates as part of the content. This is the fastest-moving layer in the stack, and at the end of this post I come back to what changed underneath my own findings while I was writing it up.

The three failure signatures of a quantization mismatch

Before any format names, the shape of the problem. A format that does not suit your silicon fails in one of three ways, and they are not equally expensive.

SignatureWhat you actually seeWhat it costs you
It refuses to loadThe server dies during startup. An unsupported-configuration error, a missing kernel, a missing backend.An hour. This is the honest failure and the cheap one.
It loads and produces garbageHealth checks pass. Requests return 200. The output is nonsense, or one character forever.A weekend, because every instinct points at the model or the weights.
It loads, works, and quietly falls backCorrect output. A server that looks completely healthy. Throughput far below what the hardware can do.Months, because nothing is ever wrong enough to investigate.

The ranking is the useful part. Most people treat a refusal to load as the worst outcome and a clean load as success. It is the other way round. A refusal tells you the truth immediately and for free. A clean load tells you nothing at all, and the third signature can sit in a production lane indefinitely while you conclude that this is simply how fast the hardware is.

Three panels side by side: a solid closed barrier on the left, a dense field of identical repeated strokes in the centre, and a smooth arrow on the right that thins and fades out before it reaches the edge.

So the first diagnostic question is not "is this model any good." It is "which of those three am I looking at." They have different causes and completely different fixes.

When a well-regarded 32B coder outputs only !!!! on an RTX 5090

Here is signature two, in full.

I wanted a dense 32B coding model to test reasoning quality against the mixture-of-experts models that fit this box. I pulled a four-bit AWQ build of a very widely used 32B coder, published by the model's own authors. About 19 GB. It fitted with enormous room to spare. It loaded cleanly. The server came up healthy.

Every response was exclamation marks. Not truncated output, not a bad answer, not a loop of plausible-looking text. Just !!!! until it hit the token limit.

The obvious suspicion is a corrupt download, so I checked the files. The next suspicion is the kernel, so I switched kernel paths: AWQ has a Marlin path and a classic path, and vLLM will pick one for you. I ran it on both. Identical garbage on both.

At that point the arithmetic on blame changes. This is one of the most downloaded coding models in existence, in a quantization published by the people who trained it. If it were broken in general, it would be the loudest bug on the hub. It is not broken in general. It is broken here, and "here" means the AutoAWQ four-bit general matrix multiply format running on sm_120. The format miscomputes on this specific silicon and hands back numerically meaningless activations, and a model fed meaningless activations emits whatever token sits at the top of a flattened distribution. On this run that token was an exclamation mark.

The general rule I took from it: when a hugely popular model misbehaves only for you, you are not the first person to run the model. You are one of the first to run that format on that silicon. Popularity is evidence, and it points at your configuration rather than at the weights.

This class of fault is not rare and it is not always the kernel. vLLM's own tracker carries a closely related case where NVFP4 quantizations of a particular model family "silently produce garbage" because a set of layers was missing from the quantization config's ignore list (vLLM issue #40252). Different mechanism, same signature: clean load, healthy server, meaningless output, and the metadata rather than the model at fault.

NVFP4 vs AWQ vs FP8 on sm_120: the map as I found it

This is the table I would have wanted before I started. It is also the table with the shortest shelf life in this post, so it carries its date in the heading.

Measured on two RTX 5090s (sm_120), vLLM 0.22.0 and SGLang 0.5.12, between 2 and 5 June 2026.

FormatWhat happened on this box
NVFP4 (compressed-tensors)Worked, on real FlashInfer Cutlass kernels with no fallback, for the 30B mixture-of-experts I ran. The one format that gave me no trouble. See the caveat below, which is important.
AWQHit and miss, and the miss is silent. A 30B mixture-of-experts worked, with a patch. The dense 32B coder above produced pure garbage on both the Marlin and the classic kernel.
FP8Loaded, then produced garbage. The recorded cause was an untuned Blackwell fused mixture-of-experts kernel with no tuning config for this card.
MXFP4Not a quality question on this box. The 120B-class model that ships in it does not fit in 64 GB at all, which is the previous post in this series.

A grid of twelve square cells in three states: solid filled cells, diagonally hatched cells, and empty outlined cells crossed by a slash, scattered irregularly across the grid.

On the FP8 failure, the fragment I recorded at the time was a missing kernel configuration, along these lines:

Config file not found ... E=64,N=768,RTX_5090

I noted the shape of that message rather than preserving the full traceback, so treat it as the class of error and not as a quotation. The mechanism behind it is well documented upstream, and it is worth understanding because it generalizes past FP8. Fused mixture-of-experts kernels ship with per-GPU, per-shape tuning configs. vLLM's tracker records that the config directory contains no RTX 5090 entries for any data type, the only consumer-adjacent Blackwell configs covering a different professional card (vLLM issue #36095). With no matching config, the engine falls back to a default, and on this silicon that default path was not merely slow. It was wrong.

The same week, the alternate engine refused the AWQ build of the model that worked on vLLM, dying in its own loader on a zero-dimension tensor concatenation. Two engines, one model, one format, two unrelated faults. Which brings me to the caveat that matters more than the table.

Do not read "NVFP4 reliable" off that table as a property of NVFP4. It is a property of one quantization of one model on one engine version in one week of June. I went back to vLLM's issue tracker on 14 September 2026 specifically to check whether my own summary had aged well, and it had not. NVFP4 on sm_120 has a substantial failure history of its own: a device capability check that missed sm_120 entirely and broke NVFP4 mixture-of-experts kernels (#33416), a missing cutlass_scaled_mm kernel on a 30B NVFP4 mixture-of-experts (#24921), and a model that will not start at all because no NVFP4 mixture-of-experts backend supports the deployment configuration (#35065). Support is also actively landing in the other direction, including NVFP4 KV cache work aimed squarely at consumer Blackwell (PR #50288).

So the honest version of the table is not a ranking of formats. It is a demonstration that the unit of compatibility is the whole tuple: format, kernel, model architecture, engine version, and card. Change any one of the five and you are testing something new. My four rows are one sample of that space, dated, and the value in them is the shape rather than the verdict.

How to check you are on the fast kernel path and not a fallback

Signature three is the one nobody checks for, so here is the check.

The premise is simple and slightly uncomfortable: the kernels you believe you are running are not necessarily the kernels you are running. Engines are built to degrade gracefully. Faced with a shape they have no tuned kernel for, they pick something that works rather than stopping, and they tell you in a warning during a startup log that scrolls several hundred lines.

The canonical example is on the tracker, and it describes consumer Blackwell exactly. An NVFP4 checkpoint on an RTX 5090 falls back to a four-bit weight-only Marlin path while warning that there is no native FP4 support, and the model serves correctly the whole time, just far slower than the hardware should manage (vLLM issue #47749). Nothing crashes. Nothing looks wrong. You have bought a card for its native low-precision arithmetic and you are not using it.

Four checks, cheapest first.

1. Read the startup log for the words "fallback", "default config",
   "not supported", "emulated", and the name of a kernel you did not
   choose. Do not skim it. This is where the engine tells you.

2. Confirm the device is what you think it is. On Blackwell, verify
   the compute capability reads 12.0 rather than assuming the wheel
   detected it correctly.

3. Check that the enabling flags for your format are actually set in
   the running process, not in the command you believe you ran. Read
   them back off the live container, not out of your notes.

4. Measure prefill throughput and compare it with the figure the
   format is supposed to deliver. An order-of-magnitude gap is a
   fallback. This is the only check that catches a silent one.

Step four needs a number to compare against, which is the other reason to measure the healthy case while it is healthy. On this box, with prefix caching working on a roughly 30,000 token shared prefix, a cold turn took 19.7 seconds and subsequent warm turns took 0.10 seconds, around 190 times faster, at roughly 80 percent cache hit. After a later configuration fix the hit rate on an identical prefix read 99.8 percent. Those numbers are the baseline that makes a regression visible. Without them, "it feels slow today" is not actionable and nobody acts on it.

Point three deserves its own sentence, because it caught me. Read the flags off the running process. I have edited a launch template, restarted what I thought was the new configuration, and been served by a container that never picked up the change. The live process is the source of truth. Your notes are a hypothesis about it.

When the fix is a patch and not a config flag

Sometimes there is no flag, and the honest answer is that the engine has a bug.

The model that eventually went into my engineering lane is a 30B mixture-of-experts with multi-head latent attention, in AWQ. It crashed on load, on both engines, with an attribute error: the attention code asked a linear layer for its weight, and the layer did not have one. The reason is mundane once you see it. Quantized layers store packed tensors, qweight and scales and zero points, and no float weight attribute exists to read. This is a general property of AWQ and GPTQ layers rather than anything specific to my setup.

What made it a bug rather than a limitation was the shape of the surrounding code. A few lines earlier, the same function already computed a safe dtype with a guard that checked whether a weight attribute existed. Then it went and dereferenced .weight directly anyway. The guard was written and then not used. The fix was one line, swapping the direct dereference for the value the guard had already produced, and I applied it without rebuilding anything by mounting the patched file over the container's copy at launch.

That patch is the reason the lane works. It is also the part of this story I am least comfortable recommending, so here is the test I now use before running on a patched engine.

Is the failure a bug or a limitation? A bug is a mistake inside a code path that was clearly intended to support your case, and the tell is code written to handle your situation that does not quite run. A limitation is the absence of an implementation. Patching a bug restores intended behaviour. "Patching" a limitation means writing a kernel, and you will get it subtly wrong.

Does the patch change behaviour or restore it? One line that routes to an already-computed safe value is a different risk from one that disables a check or loosens a tolerance. If a patch makes an error go away without explaining why the error was wrong, that is a warning to walk away.

Does the output still get verified? A patch that lets a model load has proved nothing about what the model then produces. The exclamation-mark model loaded perfectly.

Can you carry it? Mine was a bind-mounted file and a one-line regeneration command, which is a maintainable position. A local fork of an engine that moves weekly is not.

And if it is a genuine bug, report it. I did not do that promptly, and the postscript makes the case better than I could: when I checked in September, upstream had reworked the quantized kv_b_proj path through a proper abstraction rather than the patched line I was carrying (vLLM PR #55741). The bug I sat on was real, other people hit it, and the fix that landed was better than mine. My patch is now most likely obsolete, which is the best outcome available and also an argument for filing the issue on the day you find it. A patch you carry quietly is technical debt with a countdown on it.

The counterpoint, and the reason not to reach for a patch first: a fault that looked like a model defect on this same box turned out to be two flag names. Tool calls were arriving malformed, and a zero-argument tool call would not parse into an empty object. That reads as a model too weak to emit clean structured output, which is exactly the conclusion I was drafting. It was the wrong parser. The engine had a newer parser for that model generation, plus a separate reasoning parser to keep the model's thinking out of the answer content. Changing those two flags and nothing else fixed the malformed calls, stopped thinking tags leaking into responses, and left the cache hit rate at 99.8 percent. One afternoon of "the model cannot do tool calls" was one line of configuration.

So the order is: configuration, then parser, then patch. Most things that look like a broken model are above the model.

A blue Ethernet connector plugged into a handheld network tester on a green cutting mat, with coiled cable beside it, mid-diagnostic.

The operational tax: a Hugging Face download that stalls at 2.3 GB and will not resume

None of the above is what actually cost me the most time. This did, and it is the section I nearly cut as too trivial to publish, which is a good sign that somebody is searching for it at one in the morning.

A 49.6 GB download stopped dead at 2.3 GB. Not slow. Zero bytes per second, indefinitely, with no error. The official client, a valid token, plenty of disk.

The box is on Wi-Fi, so the obvious conclusion is the link. I very nearly spent the evening on the network. Instead I measured it: a plain HTTPS transfer from the same host to the same service ran at 11 MB/s, steadily. The link was healthy. The transfer client was stuck.

Killing it and switching to a resumable plain-HTTPS transfer finished the file. For the model store I settled on a parallel downloader pointed at the resolve URLs, with a modest connection count, because this box's link develops serious bufferbloat under heavy parallel load. Gateway ping went from 1.7 ms to 159 ms with too many streams open, which turns a fast link into an unusable one for everything else on it.

Two things I learned the hard way with that downloader:

- A shard can sit there showing an unfinished control file while
  already being byte-complete. Compare the on-disk size against the
  size the service reports in its response headers before you decide
  it is stuck.

- It only fetches the files you list. List the weight shards and
  forget the index and tokenizer JSON, and the model will not load,
  with an error that says nothing about a missing download.

Here is the correction, and it is the reason this section is worth the words. My fix worked, but it was not the right fix. Checking upstream in September, the stall is a well-known and widely reported bug in the hub's newer content-addressed transfer backend, not a peculiarity of my machine. It has been reported stalling dead (xet-core #789), stalling whenever that backend is enabled (#850), and stalling specifically on cloud hosts while succeeding with the backend switched off (#800), with the same failure surfacing through higher-level libraries (transformers #45797). The described mechanism matches my symptom closely: chunks arriving out of order at scale, TCP spending its time reordering, a congestion window that never opens, and one connection eventually timing out and taking the whole transfer down with it. A 2.3 GB wall on a healthy 11 MB/s link is exactly what that looks like from outside.

And the documented workaround is a single environment variable that disables that backend and falls back to plain HTTPS, rather than the separate download tool I reached for. One variable would have replaced the entire detour.

I am leaving my longer route in this post rather than quietly replacing it with the one-liner, because the useful lesson is not the variable. It is the order of operations. Measure the link before you blame the link. Once I knew the link was fine at 11 MB/s, the fault was inside the client, and every fix from there, mine or the official one, was aimed at the right thing. The generalizable version is the same move as checking where an out-of-memory error fires, or checking which kernel is loaded: establish which layer is actually broken before you start fixing layers.

The decision rule

Three rules, in the order they save you time.

Prefer the format your silicon has first-class kernels for, and find out which one that is before you download 50 GB. Not the format with the best benchmark, and not the newest. The one whose kernels are compiled, tuned and tested for your exact compute capability. On a consumer card that is a smaller set than the marketing suggests, and the fastest way to establish it is to search the engine's issue tracker for your card's name. Ten minutes there is worth more than any blog post, this one included.

Verify output quality before you benchmark speed. Every performance number I collected on a garbage-producing configuration was wasted work, and there is something particularly bleak about a throughput graph for meaningless tokens. Send three prompts, read the answers with your own eyes, and only then start timing.

Treat "it loaded" as no evidence at all. A clean startup tells you the weights fit and the file format parsed. It tells you nothing about numerical correctness and nothing about which kernel you landed on. Both remaining failure signatures live entirely on the far side of a successful load.

What I would check first, next time

Search the engine's issue tracker for your card before choosing a format. Everything in this post that took me days is on that tracker in some form, often with a workaround in the thread. My own sample of four formats was less accurate, three months later, than an afternoon of reading would have been.

Record the healthy numbers while things are healthy. Prefill throughput, decode throughput, cache hit rate, startup time. A silent fallback is invisible without a baseline, and you cannot collect a baseline retrospectively.

Read the error's layer, not its wording. A stalled download that is not the network. An out-of-memory that is a shared-memory limit inside a kernel rather than VRAM. Exclamation marks that are a matrix multiply and not a model. In each case the message points at the wrong layer and the fix begins by finding the right one.

Same book, different printing presses. The one that prints gibberish is the expensive one, because you will spend the weekend rereading the book.

Date-stamp your own findings while you are at it. Mine were three months old when I came back to write them up, and one of the four rows had already stopped being true.

Image credits

Images from Unsplash.

Share

Related Posts