Series

The Model Isn't Everything

Three years of production LLM deployments moved the leverage downstream of model choice. Eleven parts on where it actually went, with the numbers.

Model upgrades produce roughly 3% quality gains. Engineering investment around the model produces 28–47%. Teams still treating model selection as the primary quality lever are optimising the cheapest, fastest-changing component in their stack.

Eleven parts on where the leverage moved: hallucination as a retrieval problem rather than a generation one, evals as the real product spec, why agent failures are almost always integration failures, why a domain expert beats a prompt engineer, and why the human baseline everyone measures against does not exist.

The last two are the ones regulated buyers should read first — calibration is the most under-invested feature in production AI, and pilots systematically lie about what production will do.

← All posts