Bob O'Dell
HomeAboutExperienceProjectsAI ToolsBlogContact
Bob O'Dell blog
Bob O'Dell/Blog

Blog

The open side of a desktop workstation: a large graphics card mounted low in the case, a liquid cooler's radiator and three fans across the top, and coolant tubes running between them.LLM

Putting a self-hosted model on real production work without betting the codebase on it

The hard part of moving a coding agent to a local model was never inference. It was keeping the safety discipline wrapped around it: isolation, a review gate that sends bad work back, and a kill switch that works without a deploy. The architecture I used, and the smoke test that passed while every real request failed.

September 28, 202616 min read
A pegboard wall densely covered with dozens of different hand tools hanging in rows, so many and so similar that no single one stands out.LLM

185 tools to 24: the agent context budget nobody measures

I thought my agent's base prompt was 4,000 tokens, because that is how long the prose is. Measured on the wire it was closer to 99,000, and the difference was tool schemas: about 55,000 tokens of tool definitions paid on every turn, before the agent had read a single file. What that costs you in window, and what it costs you in correctness.

September 24, 202617 min read
A close-up of a circuit board's dense circuitry in shallow focus, traces and surface-mounted components filling the frame in cool blue-grey tones.LLM

Quantization roulette on consumer Blackwell: when a good model outputs nothing but exclamation marks

A widely used 32B coder produced only '!!!!' on two different kernel paths. Another model in FP8 loaded cleanly and returned garbage. Neither was a bad model. Both were quantization formats meeting silicon that had no working kernel for them. The three failure signatures, ranked by what they cost you, and how to tell them apart before you waste a weekend.

September 14, 202618 min read
A GeForce RTX graphics card photographed close up in shallow focus, its fan shroud and heatsink fins filling the frame.LLM

What actually fits on 64 GB of VRAM (it is not a 120B model)

The arithmetic says a 4-bit 120B is about 60 GB and you have 64, so it should load. It does not, and not with tensor parallelism, an FP8 KV cache, CPU offload, or the memory fraction cranked to 0.97. Here is the memory budget that actually governs, the ceiling I derived, and a formula you can run in two minutes before downloading 60 GB.

September 8, 202617 min read
Looking straight up the centre void of a spiral staircase, its identical flights repeating away from the camera turn after turn.LLM

Why my coding agent re-read 162,000 tokens on every single turn

Generation was a healthy 140 tokens per second. Prompt processing was 52 seconds, every turn, forever. The cause was not the GPU, the context length, or the system prompt: it was an attention architecture whose state cannot be cached across requests. How to spot it before you download 50 GB of weights, and why the same box now runs that architecture again.

September 2, 202617 min read
A GeForce RTX graphics card photographed in selective focus, its fans and shroud filling the frame against a dark background.LLM

vLLM vs SGLang for agent workloads: the bake-off was decided by prefix caching, not throughput

A coding agent is a 30-60 turn loop that re-sends a huge, mostly-unchanged context every turn, so the benchmark that decides your engine is cross-request prefix caching, not tokens per second. What I measured on two RTX 5090s, and the second axis nobody charts: tool-call reliability over a long loop.

August 24, 202613 min read
Close-up of a computer processor and its surrounding cabling inside an open workstation case, lit by diffused indoor light.LLM

Two CPU cores, 100% busy, zero work: making a local LLM server event-triggered

An inference server that pins one CPU core per tensor-parallel rank at 100% while completely idle is not broken hardware. It is a busy-poll loop in the scheduler. How I found it, why a cpuset would only have hidden it, and the measurement three weeks later that mattered more than the two cores.

August 17, 202611 min read
A row of rack-mounted servers in a dimly lit server room, front panels and cabling visibleLLM

Tuning an LLM workstation: what actually moved the needle on CPU and GPU

A practical account of tuning a shared Ryzen 9950X / dual RTX 5090 box that runs CI, local LLM inference, and production web at once - which CPU and GPU settings measurably changed throughput and thermals, which ones did nothing, and the drift bug that made a frequency cap silently protect nothing.

August 15, 202612 min read