Inference costs fall another order of magnitude
Serving costs for capable mid-tier models have dropped sharply again. The bottleneck for AI products is no longer compute price — it is data access and workflow integration.
Cost per million tokens for models in the capable mid-tier has continued its steep decline, driven by better serving stacks, speculative decoding, and hardware that finally matches transformer memory patterns.
When we modelled the platform's unit economics eighteen months ago, inference was the dominant line item. On current pricing it is a minority of run-rate cost, behind data acquisition, human review, and integration engineering.
This reallocation favours products with durable data positions and deep workflow embedding. A thin wrapper over a public model captures none of the surplus as costs fall; a system embedded in a customer's operating process captures most of it.
We have revised the platform's long-run gross margin assumptions upward accordingly, while holding the total capital requirement unchanged.
Investor takeaway
Falling inference prices accrue to products with data and workflow depth, not to thin model wrappers.
