LLMThe hard part of moving a coding agent to a local model was never inference. It was keeping the safety discipline wrapped around it: isolation, a review gate that sends bad work back, and a kill switch that works without a deploy. The architecture I used, and the smoke test that passed while every real request failed.
LLMI thought my agent's base prompt was 4,000 tokens, because that is how long the prose is. Measured on the wire it was closer to 99,000, and the difference was tool schemas: about 55,000 tokens of tool definitions paid on every turn, before the agent had read a single file. What that costs you in window, and what it costs you in correctness.
LLMA widely used 32B coder produced only '!!!!' on two different kernel paths. Another model in FP8 loaded cleanly and returned garbage. Neither was a bad model. Both were quantization formats meeting silicon that had no working kernel for them. The three failure signatures, ranked by what they cost you, and how to tell them apart before you waste a weekend.
LLMThe arithmetic says a 4-bit 120B is about 60 GB and you have 64, so it should load. It does not, and not with tensor parallelism, an FP8 KV cache, CPU offload, or the memory fraction cranked to 0.97. Here is the memory budget that actually governs, the ceiling I derived, and a formula you can run in two minutes before downloading 60 GB.
LLMGeneration was a healthy 140 tokens per second. Prompt processing was 52 seconds, every turn, forever. The cause was not the GPU, the context length, or the system prompt: it was an attention architecture whose state cannot be cached across requests. How to spot it before you download 50 GB of weights, and why the same box now runs that architecture again.
LLMA coding agent is a 30-60 turn loop that re-sends a huge, mostly-unchanged context every turn, so the benchmark that decides your engine is cross-request prefix caching, not tokens per second. What I measured on two RTX 5090s, and the second axis nobody charts: tool-call reliability over a long loop.
LLMAn inference server that pins one CPU core per tensor-parallel rank at 100% while completely idle is not broken hardware. It is a busy-poll loop in the scheduler. How I found it, why a cpuset would only have hidden it, and the measurement three weeks later that mattered more than the two cores.
LLMA practical account of tuning a shared Ryzen 9950X / dual RTX 5090 box that runs CI, local LLM inference, and production web at once - which CPU and GPU settings measurably changed throughput and thermals, which ones did nothing, and the drift bug that made a frequency cap silently protect nothing.