Caching Cheaters on OpenRouter
Tarun Chitra — August 2026
The full essay — with the market-structure framing, the economics of open source, and the complete analysis — is here: The Souq.
OpenRouter is the largest and most successful public marketplace for open-weight models. As a marketplace, it matches AI users with open weight models that maximize quality at the lowest price for their task. The introduction of this marketplace, however, has fundamentally unbundled the product sold by closed weight frontier labs such as Anthropic and OpenAI. To see why, we need to first consider the agentic product delivered by a frontier lab. When a user uses an AI agent via a closed source harness such as Claude Code or OpenAI’s Codex, the default setup only allows users to utilize first-party models and compute that are provided by the lab. The user is effectively forced to buy compute and access to a model from a single provider.
However, open source harnesses, such as Nous Research (@NousResearch)’s Hermes — the largest application on OpenRouter, with 15–20% market share — allow for users to choose which models they want to run for particular tasks and which compute providers will utilize their hardware for inference. These models, when provided as open weight models, allow for a user to select a particular model made by a builder such as DeepSeek (@deepseek_ai) and then choose an inference provider such as Together AI (@togethercompute) to run their model. Together AI’s CEO Vipul Ved Prakash (@vipulved) recently wrote an excellent article, “The Economy of Tokens” that describes how this ecosystem evolved in a manner that allowed for the exponential growth needed for open weight models to be competitive with their closed source competitors.
Initially, open source harnesses were statically configured, choosing a single model for multiple tasks. However, as people realized that adapting the choice of model to a particular task would provide a better quality versus cost trade-off, it became clear that routers serve the economic role of a matching algorithm.[1] In particular, routers like OpenRouter perform two types of matching:
Capability Matching: Taking a user’s prompt or context and figuring out which model should be used for particular portions of the user’s workflow based on the quality and cost of the model. For instance, a frontier model might provide the best performance but cost 50 times more than a model that is only 20% less capable.[2]
Execution Matching: Given a model, a context, and a set of bidding inference providers, a user’s request is assigned to a provider that offers the best price and service-level agreements (SLAs), such as uptime and latency
Since Prakash’s article launched in June 2026, there have been multiple explosions in both open weight model quality and in the market dynamics that this unbundled market structure has evolved to.
The Game Theory of Routing
Given the explosion in usage and model quality on OpenRouter, a natural question is to ask what strategic behaviors have evolved within this marketplace. Matching markets in finance and compute often have deep strategic behaviors from market participants who try to optimize for market share or profit at the expense of other participants or users. To this end, we briefly formalize the OpenRouter market as a Tullock contest. From this lens, we analyze how OpenRouter has performed in terms of user welfare (e.g. how much has the quality and realized cost of users improved when they switch from closed source models to open weight models). Classical results on Tullock contests suggest that the main way a user’s welfare can be reduced in this contest is via strategic behavior from the inference providers that leads to higher costs or worsened quality. While there has been much written about worsening quality from third-party APIs in AI, there has been little written about the pricing dynamics in marketplaces like OpenRouter.
We ran paid experiments on OpenRouter and found explicit examples of strategic behavior from inference providers leading to worse prices to users. While there are a large set of potential attack vectors, we focused on the simplest attack: an inference provider quoting a lower price for a query while never serving cached tokens. As cached tokens are generally priced cheaper than input or output tokens on marketplaces like OpenRouter, it is theoretically possible for an inference provider to quote a lower price for input and/or output tokens than honest caching competitors by reporting that they never cached tokens. These attacks, however, rely on inference providers serving repeated queries to the same customer and reporting no cache usage. We find explicit empirical evidence for these strategies existing on OpenRouter’s market for model builder Z.ai (@Zai_org)’s GLM-5.2 model, with at least two providers performing actions that reduce user welfare. These strategies currently extract around 2% of OpenRouter’s revenue across the open-weight markets we measure, but if every provider used this strategy then the amount extracted could be up to 10% — on the order of $14 million a year.[3] Succinctly, our result can be viewed as saying the following: the cheapest quoted token is often not the cheapest conversation.
Right after a provider cuts its price, we send the same repeated prompt to it and to Z.ai at once. The cutting providers treat 71 percentage points less of the repeated text as cached, and pass on 52 points less of the savings a warm cache would give — so the identical follow-up query costs the user far more than it would at Z.ai. This measures how much more users pay on a repeated query, not why: the provider could be gaming the price, or simply incompetent (e.g.* restarting its machine, and wiping its cache, on every price update).*
In virtually all multi-agent marketplaces, there is often a question of whether particular participants can unfairly extract value from other participants or harm consumers. For instance, spoofing on traditional equities exchanges[4] or Maximal Extractable Value on blockchains are viewed as a means to distort prices against consumers and other market participants. Similarly, online ad auctions have long suffered a quality problem, where an ad buyer receives lower quality placement than what they were promised by a platform.[5] This problem is also known to be found in open weights markets — e.g. an inference provider running a cheaper model or quantizing it.[6]
How does routing actually work?
In order to measure market integrity on a platform like OpenRouter, we first have to understand how the router allocates tokens to different providers.[7] While the term “model routing” has been used to mean many things in academic and business contexts, routers like OpenRouter, Ramp (@tryramp), and Databricks (@databricks) share the following stylized structure:[8]
An application or user specifies a particular workflow that they want to execute (initial prompt, context, tool calls)
A capability routing algorithm analyzes the type of task being requested, maps a task to a set of evaluations (“evals”) that are most similar to the task, and then ranks the different models for suitability for the task based on these evals, community spend data, etc.
An execution routing algorithm takes in a desired model and the prices quoted by inference providers and allocates the tokens to an inference provider.
In the case of centralized/non-competitive routers (like Ramp and Databricks), their execution routing is not a competitive bidding process. Databricks routes across fixed, administrator-configured traffic splits to model endpoints the customer already pays for directly, with automatic fallbacks when an endpoint errors or rate-limits; Ramp selects among providers dynamically, scoring each model on its live latency and failure rates against a deadline, and passes through each provider’s list price. In both cases, the SLA guarantee comes from the router rather than from a contest: endpoints that violate latency or uptime targets are simply routed around. The same pattern runs inside large companies: Uber, Grab, LinkedIn, Expedia, and Instacart each route their internal AI traffic through a single gateway of this kind, and Palantir sells one to its customers.[9] We’ll briefly describe how these two routing algorithms work, but note that as the algorithms themselves are not open source software, we are describing their functionality based on stated documentation versus precise implementation.
The two stages of routing: capability routing maps a task to a model by trading quality against cost; execution routing then assigns the request to one of the inference providers serving that model. The second stage is where the pricing game is played
Capability Routing
A capability routing algorithm’s goal is to maximize the expected quality of an output while minimizing the cost. This quality score is constructed by statically analyzing the task at hand, mapping it to a set of evals for similar tasks, and then using the publicly posted performance of different models on that eval (e.g. on Arena.ai (@lmarena_ai))[10] to construct the score. This eval-mapped construction is implemented almost literally by LMSYS’s RouteLLM; most commercial routers instead learn a quality predictor from preference or usage data.
Execution Routing
Execution routing takes place after capability routing selects a model for a task. The goal, given a selected model, is to assign the token generation to an inference provider who supports that model. Centralized routers, like Ramp or Databricks, allocate requests without any bidding: Databricks by statically configured traffic percentages, Ramp by a dynamic rule tuned on live latency and failures. On the other hand, the model that OpenRouter uses is a Tullock Contest — a contest where participants expend costly effort to try to win a resource. The classic Tullock contest has a participant’s probability of winning the resource proportional to the effort (e.g. cost) that they spend. However, numerous studies have shown that non-proportional Tullock contests can be more expensive to manipulate.[11] OpenRouter’s documentation first filters out providers with recent outages and then defines the probability of winning as a function of the quoted prices alone; latency and throughput enter only if the user explicitly sorts on them — which turns the contest off entirely. We formalize this later in the post.
Competitive Equilibria
OpenRouter’s marketplace involves applications — entities that generate tokens, such as harnesses or specialized applications like Framer — and inference providers who provide compute. Applications send their context (prompts and auxiliary data) to OpenRouter’s API, which selects a model via capability routing for the user. Then the router uses a Tullock contest to allocate these input tokens to a particular inference provider.
Tullock Contests. Tullock Contests provide a theoretical framework for analyzing how a fixed rent — such as a payment for tokens — is distributed amongst a set of competing providers. They are used to describe everything from government procurement auctions to advertising and military war games.[12] Tullock contests have \(n\) participants who each contribute an “effort” \(x_i\) in order to realize a reward \(R\). Effort represents the relative quantity of resource that a provider expends to try to realize the reward versus another. The contest is fully specified by its allocation rule\(A\), which maps the efforts \(x_i\) to a set of probabilities \(\pi_i\) such that the expected reward of provider \(i\) is \(\pi_i R\).
The first contest studied in Tullock’s Efficient Rent Seeking (1980) is the allocation rule
\[\pi_i = A(x_1, \ldots, x_n)_i = \frac{x_i^r}{\sum_{j=1}^{n} x_j^r}\]
The exponent \(r\) represents decisiveness — a higher value of \(r\) implies that smaller relative changes in effort will cause larger disparities in expected reward. It can also be viewed as controlling the log odds for participants. For instance, if \(r = 2\), if participant \(i\) puts in two times more effort than participant \(j\), then \(i\) increases their probability of winning relative to that of j by four times. Put another way, when \(r > 1\), the impact in terms of expected reward of increasing effort is super-linear in effort.
OpenRouter is a multi-dimensional Tullock Contest. OpenRouter’s documentation[13] states that the allocation rule does the following:
Each of \(m\) providers posts a pricing vector \((p_i^I, p_i^O, p_i^C)\), where \(p_i^I\) is the price per million input tokens, \(p_i^O\) the price per million output tokens, and \(p_i^C\) the price per million cached tokens
First filters the set of \(m\) participants to the maximum subset of size \(n\) that have high quality score (where the quality score is not published but relies on a provider’s latency, uptime, and historical TPS performance)
OpenRouter computes a composite price \(P_i = P_i(p_i^I, p_i^O, p_i^C)\) and then allocates to provider \(i\) with probability \(\pi_i \propto 1/P_i^2\)
If we define the effort \(x_i = 1/P_i\) then, since \(\pi_i \propto 1/P_i^2 = x_i^2\), this can be viewed as a standard Tullock allocation rule with decisiveness 2. Higher decisiveness means that smaller increases in effort by a single participant can super-linearly increase their expected allocation and reward. In particular, this means that it is possible for an inference provider to make a small price decrease (which would increase their effort) to gain a much higher reward than their price increase saves in returns to the end user. The same super-linearity has a cruder cousin: because the rule scores each listed endpoint on its own, a decisiveness above one is not Sybil-resistant — an owner can split into economically identical duplicate endpoints and raise its combined share without touching its price.[14] We note, furthermore, that the function \(P_i\) that maps the user quoted prices to a scalar composite price is not explicitly defined in the OpenRouter documentation and we measure it from the router’s observed choices.[15] In short, \(P_i \sim p_i^I\).
While the documentation is underspecified, we can measure which price the router actually uses. OpenRouter lets a user ask for the cheapest provider directly, and when asked, the router has to reveal who it thinks is cheapest. We sent thousands of these requests. The informative moments are when one provider quotes the lowest input price while a different provider quotes the lowest output price — the router then has to pick a side. In every case we observed — 64 of 64 such choices across two markets — it picked the provider with the cheapest input tokens. If output price carried more than a few cents of weight per dollar of input price, the router would have made different choices. The cached-token quote has no effect we can detect: adding it to a model of the router’s choices improves prediction by exactly nothing (see plot below). For ranking providers, the contest price is effectively the input token price. Our exact fitting methodology is described in the methodology appendix.
Which price enters the contest? Left: every cheapest-provider pick the router revealed is explained by input price alone — give output price more than a few cents of weight per dollar of input price and the picks stop making sense. Right: adding the cached-token quote to a model of the router’s choices improves prediction by exactly nothing. The router ranks on input price and ignores the price where the money moves on repeated context
OpenRouter’s Dark Forest
We are now ready to look at the real behavior of these actors in the largest public routing environment, OpenRouter. We first describe the data captured from OpenRouter to analyze strategic behavior amongst inference providers, leaving the full methodological details to the appendix. This will allow us to analyze the precise Tullock contest used and what the optimal strategic behavior looks like in this market. We then look at high level statistical trends, especially in the near-frontier markets of Z.ai’s GLM-5.2 and Moonshot (@Kimi_Moonshot)’s Kimi-K3. We identify four archetypes of inference providers in these markets — suggesting there are some strategic and potentially manipulative behaviors at play.
Provider Pricing
Each provider on OpenRouter quotes three prices on the marketplace:[16]
Input Token Price, \(p_i^I\): Price per million tokens of input tokens (e.g. user prompt, other context/metadata)
Output Token Price, \(p_i^O\): Price per million tokens of output tokens (e.g. number of tokens in the response from the model)
Cached (read-only) Token Price, \(p_i^C\): Price per million tokens of cached tokens
The first two prices are relatively self-explanatory and give providers the ability to discriminate on pricing based on their expected cost for running different models on the hardware they have. The latter price, which is included by default in the OpenAI API, allows providers to offer a cheaper price for tokens that are repeated and resident in a KV-cache. This can lead to large cost savings in agentic workflows, where cached tokens are common, as prefixes of the prompt and/or task are often repeated in subsequent prompts made by the agent.
Cached Token Price is not Verifiable. One of the main reasons the service that OpenRouter offers is possible is because open weight models have standardized using OpenAI-compatible APIs. However, the components of the output of the API are not tamper-resistant as a user cannot independently verify the output of an API request. According to both the OpenAI and OpenRouter API documentations, the API operator provides both the OpenAI-compatible API response usage.prompt_tokens_details.cached_tokens and the OpenRouter API response for cached_tokens. In particular, this means that an inference provider can maliciously send a malformed OpenAI-API compatible response that claims there was no caching even if there was.
The Four Horsemen of Inference Provision. There are a number of different types of inference providers on OpenRouter. Virtually all model builders (e.g. the entity that trained the model, such as Z.ai or Minimax) run their own inference provider. The inference provider business is the main source of revenue for these companies and inference funds the training of future models. We bucket all other inference providers who participate on OpenRouter into four categories and provide examples from GLM-5.2 market:
Well-Funded Copy Pricers: These are inference providers whose compute is mainly used to service non-public, enterprise contracts, but participate in OpenRouter for marketing and/or to exist on the leaderboard. These entities range from venture funded (Together AI, Modal, Baseten, Fireworks, Venice, Crusoe) to public companies (Cloudflare, DigitalOcean). These entities seem to have the least amount of strategic pricing, precisely copying the price posted by the model builder (see below).
Well-funded copy pricers: venture-backed and public-company providers quoting exactly the model builder’s price — on the leaderboard for distribution, not competing on price
Specialist Providers: Inference providers who either use custom hardware (e.g. Cerebras) or improve kernel performance to offer faster inference. Usually, these are “fast” modes provided by the same well-funded providers; however, some, such as Wafer, are nascent and building a competitive edge via optimized provisioning. The pricing of these providers will be higher than that of the model builder.
Specialist providers: custom hardware and optimized kernels sold at a premium to the model builder’s price — faster tokens, not cheaper ones
Rebaters: These are providers who provide static prices below the model builder. They don’t dynamically adjust their prices but rather, effectively provide a rebate by pricing below what the model builder quotes.
Rebaters: static quotes pinned below the model builder — a standing discount, not a strategy that responds to the market
Repricers: These inference providers are frequently updating their pricing and dramatically undercutting the model builder — the sort of dynamic pricing OpenRouter’s own Alex Atallah (@alexatallah) has written about. In this screenshot, you can see that StreamLake (which is a datacenter business owned by Chinese social media application Kuaishou) and Novita are pricing at 5-10% of the price that the model builder is offering.
Repricers: StreamLake and Novita quoting 5–10% of the model builder’s price and updating constantly — the only archetype that plays the contest dynamically
Cache me if you can
One strategy that an inference provider can use to increase their revenue while reducing consumer welfare is to charge for cached tokens as if they were uncached. However, if this negatively influences the router’s probability of allocating to such a provider in the future, then they might lose long-term revenue. A natural question to ask is if OpenRouter’s Tullock Contest, which has a partially public scoring rule, allows for an execution of such a strategy without reducing long-term revenue. This is analogous, in many ways, to Maximal Extractable Value in blockchains, where strategic users are able to extract a profit from less sophisticated users by strategically updating pricing. A stylized version of such a strategy is:
A provider reduces their quoted price on input and output tokens dramatically below other providers
This increases their market share super-linearly, so their volume of tokens received from the router goes up relative to competing providers
They then report, by manipulating responses to the cached token API, that they didn’t use any cached tokens in order to charge the higher rate
Indirectly, one can think of the strategic provider as manipulating the caching mechanism to provide a form of stateful “memory” to the router, where if a provider is chosen at time t, they are more likely to be chosen at time t+1 despite charging the user a higher price.[17]
Example. As a simple example of how this is profitable relative to honest token reporting behavior, suppose that an honest / non-strategic provider offers $1 per million input and/or output tokens and $0.01 per million cached tokens, while a strategic provider offers $0.6 per million input and/or output tokens and $0.02 per million cached tokens. For a million token workload with half of the tokens cached, the honest provider charges $1/million tokens * 0.5 * 1 million tokens + $0.01/million tokens * 0.5 * 1 million tokens = $0.505 whereas the strategic provider charges $0.6/million tokens * 1 million tokens = $0.6. Despite charging less to increase their volume, they increase their profit relative to the honest provider (who gets less token volume).
Is there evidence that inference providers are executing this type of strategy in live OpenRouter data? Yes!
Results. We designed a sequence of experiments that we ran on OpenRouter where we repeatedly asked a sequence of prompts with varying levels of caching. Repeated prompts and/or shared context often occur in agentic workflows, so this setup is relatively realistic. We provide the full details of the experiments in the methodology section below. We looked, in particular, at how caching changed before and after an inference provider changed their price. From these experiments we find three results:
Repricers (especially Novita) appear to have significantly fewer cache hits than the model provider upon repricing / adjusting their quote
The lack of cache hits leads to decreased user welfare — the more a workflow uses repeated context, the more a user pays a repricer. We find cases where even though the repricer charges 50% less than the model builder, the user can end up paying more to use the “lower price” repricer than using the model builder
In the worst case, we find that on average it only takes a roughly 5 repeated queries for the loss that users face from a repricer to be worse than simply using the model builder’s API endpoint
Repricers have worse cache performance.
Right after a provider cuts its price, we send the same repeated prompt to it and to Z.ai at once. The cutting providers treat 71 percentage points less of the repeated text as cached, and pass on 52 points less of the savings a warm cache would give — so the identical follow-up query costs the user far more than it would at Z.ai. This measures how much more users pay on a repeated query, not why: the provider could be gaming the price, or simply incompetent (e.g.* restarting its machine, and wiping its cache, on every price update).*
We see different cache profiles for GLM-5.2. The model builder, Z.ai, is used as a benchmark and we compare how much of the response was cached (e.g. measuring cached token vs. output tokens) and the difference in cost. The three most aggressive repricing providers, StreamLake (a unit of Chinese social media application Kuaishou), Novita, and Baidu, have both significantly worse cache performance and higher prices on repeated queries.
Lack of caching means users pay more.
What happens around a price cut: the top panel shows the repricer’s cut relative to the model builder; the middle panel shows cache recognition collapsing at the cut — some providers report even fewer cached tokens on the second identical prompt, consistent with a full eviction; the bottom panel shows the user’s savings on repeated prompts decaying after the cut
Next, we look at how the caching behavior of repricing inference providers changes over time. In the top panel, we can see the price cut (for input tokens) that repricers make relative to the model builder. We look at how cache performance changes when there is a repricing event. The second panel shows how much of the input was cached at the time of a price cut. Curiously, some providers reported even fewer cached tokens on the second repeated prompt — suggesting a full cache eviction took place after the price cut. This leads, as the third panel shows, to a sharp decay in user savings on repeated prompts.
Are they doing it on purpose? We can’t subpoena anyone’s cache policy, and our experiments cannot distinguish a provider that evicts state strategically from one that rebuilds its serving stack every time it touches its price. But the distinction barely matters economically. The pattern — cut the visible price, drop the cache, bill repeated tokens at the full rate — is exactly what a strategic provider would choose, it recurs at hundreds of repricing events (Novita alone repriced roughly five hundred times in July), and the provider keeps the proceeds either way. A pricing model that advertises a discount the buyer predictably does not receive is misleading in effect, whatever its intent.[18]
Users can pay easily more than if they simply just used the model builder’s endpoint.
The cheapest first request is not always the cheapest conversation: projecting from observed request costs, roughly five repeated queries is where the worst repricer becomes more expensive than simply using the model builder’s endpoint
Given the large number of well-funded pricing copy pricers, a natural question to ask is, “how many repeated prompts does one need to ask a repricer before the excess that you pay for non-cached tokens is greater than using a higher quality provider endpoint?” We find that one needs roughly 5 repeated queries to the most offensive cache losing provider (Novita) before you start paying more than using the model builder’s API.
How much money could be extracted this way? We estimate the size of the prize directly from OpenRouter’s public data: for every open-weight model, the tokens a provider fails to recognize as cached, billed at the full input price instead of the cache price, summed across the market. Today the behavior extracts around $3.5 million a year, roughly 2% of OpenRouter’s revenue; if every provider used this strategy, the ceiling is on the order of $14 million a year, or 10%. The full derivation and its assumptions are in the cost appendix.
Annualized value a caching strategy could extract, over time, for GLM-5.2 alone versus every open-weight market
The maximal amount extractable has grown quickly over the weeks we can measure cleanly, tracking the growth of the open-weight market itself.
Annualized caching-attack value by model, across all open-weight markets
The amount varies across models, but a strategic provider does not have to pick one — it could aggregate the same behavior across every model it serves and realize a large profit.[19][20]
There is one more cost the contest hides, and it grows as the market matures. A warm cache is memory that only pays off where you earned it: reuse it at the provider that built it and repeated tokens are cheap; move, and the next provider charges full price to re-read everything. OpenRouter’s documented router shops the quoted price on every request and keeps no memory of where the last one went,[21] so it minimizes the price of the next token rather than the cost of the conversation. When one provider is clearly cheapest, the price-weighted draw keeps landing there and the cache stays warm; but when several providers sit close on price — exactly what a maturing market produces — the conversation is handed to whichever is a hair cheaper this instant, and every move lands cold. On GLM-5.2 we replay the public menus under this rule: a conversation that could reuse most of its context pays 2.33 times the session-aware cost with two providers, and 3.63 times with thirty-two. Competition makes it worse, not better — more providers means more moments when a rival is a fraction of a cent cheaper, more scattering, and more warm cache thrown away. It is competition on the wrong margin.
GLM-5.2 at 90% reusable context: what OpenRouter’s documented per-request routing costs a conversation, relative to an idealized router that keeps it on one warm provider, as the field grows from two providers to thirty-two. A counterfactual replay of the public menus, not a measurement of realized routing
How can OpenRouter improve their Tullock contest?
These experiments demonstrate that there are two key flaws with the current OpenRouter Tullock contest:
The contest is implicitly state-dependent (e.g. caching acts as “state” or “memory” that the provider manipulates to increase their profit) despite the fact that the contest selection function only depends on the current pricing
The lack of ability for the purchasing user to independently verify the state used by a provider to produce tokens means they can never be completely confident that they aren’t being overcharged
The simplest fix addresses the first issue at its root: rank providers by the price a user will actually pay, not the price a provider quotes. On a workload with repeated context, the expected cost of an input token at provider \(i\) is
\[\bar{p}_i = h_i\, p_i^{C} + (1 - h_i)\, p_i^{I},\]
where \(h_i\) is the probability that a repeated token is served from cache and \(p_i^{C}, p_i^{I}\) are the provider’s cached and input prices. If the contest runs on \(\bar{p}_i\) rather than the quoted \(p_i^{I}\) — with \(h_i\) measured empirically by OpenRouter from realized billing, conditioned on session continuity and workload type — then dropping the cache is no longer free: a provider that stops recognizing repeated tokens raises its own \(\bar{p}_i\) and loses allocation. The strategy we documented becomes self-defeating, because the price the provider was manipulating is now the price the contest scores.
Short of re-pricing the contest, the same signal can be used to police it. The first issue can also be addressed by adaptively / dynamically updating the exponent and quality scores used. To do this, the router can perform repeated experiments like the ones we ran above. More specifically, one can do repeated online updates to the decisiveness of the following form:
For an observation window of duration T, send a small number of randomized canary sessions to each provider — repeated prompts with a stable session identifier, mirroring the workloads users generate organically[^klnote]
For each provider, compare what the canaries were actually billed on repeated tokens to what the quoted cached-token price says they should have been billed; the gap is the provider’s overcharge on repeated context
Identify the providers whose overcharge is persistently larger than what ordinary cache churn (evictions, TTL expiry, redeployments) can explain
Adjust the exponent and/or lower the quality score of those providers until the gap closes
This is not exotic. Databricks’ own task router commits to a model at the start of a session precisely to preserve cache efficiency, which it calls “a critical cost driver”[22] — an explicit acknowledgment (tweet) that where a session goes, and whether its cache survives, is where the money is.
The second issue, however, is more pernicious and requires deeper changes to how the router operates. For instance, if the router mandated that all KV caches had to provide cryptographic attestations (e.g. in an NVIDIA (@nvidia) enclave) of their cached token counts, then this type of manipulation could be fully prevented. However, that would slow model performance down and/or make it harder for certain providers to compete with older hardware. As hardware for inference becomes more specialized, however, one could imagine such attestations becoming a standard to reduce marketplace manipulation and improve integrity guarantees for users.
The market measures the wrong thing
Step back from the individual providers, and the problem is not that a few of them are dishonest. It is that the contest scores every provider on the price they quote — a number that is not what a buyer actually pays. On any real conversation the money moves on repeated, cached tokens, and the quoted price barely touches them: a provider can advertise a discount on the tokens the router ranks while billing the repeated tokens the router ignores. Any market that ranks sellers on a metric they control, and that buyers cannot verify, will be gamed; the only open questions are by how much and who pays. Today it is about 2% of the market’s revenue, with a ceiling closer to 10%. The fix is not to police providers one at a time but to change what the contest measures — rank them by the effective, session-aware price a user will actually pay, with cache recognition read from realized billing, and dropping the cache stops being free: a provider that stops serving repeated tokens raises its own score and loses the traffic it was chasing.
None of this is new in kind. It is maximal extractable value, ported to inference: a public, partly-specified scoring rule that lets sophisticated players skim value from ordinary users by manipulating exactly the quantity the rule fails to measure. Crypto did not make MEV disappear by policing bots — it made the mechanism transparent and changed what earned the reward. OpenRouter can do the same. The cheapest quoted token is not the cheapest conversation, and until the contest scores the price the buyer actually pays, it never will be.
Disclosures
The author is an investor in Together AI and Nous Research via Robot Ventures, and an angel investor in Wafer.
Appendix
GLM-5.2. In this post, we spend most of our analysis on GLM-5.2. This is because we have the best historical data for this model and because it is the most competitive market on OpenRouter. We briefly describe that here. OpenRouter had some of its largest growth when Z.ai’s near frontier GLM-5.2 was released on June 13, 2026. The large, latent demand for the model led to a large uptick in the number of inference providers and the competitiveness of inference providers, as you can see below. The spread in how widely different providers quote also grew much faster for GLM-5.2. Given this large increase in the number of providers, we found this market the most likely to have caching-based manipulation.
Provider entry after GLM-5.2’s release: the market deepened rapidly, which is what makes it the best venue for studying strategic pricing
Price dispersion across GLM generations: quotes spread out much faster for GLM-5.2 than for prior generations of the model
Depth and dispersion at the final snapshot (Aug 8): many providers quoting a wide range of prices — a genuinely competitive book
The price dependence of the Tullock contest
The contest in the main text scores each provider on a single price \(P_i\), but a provider quotes three — input, output, and cached tokens — and OpenRouter’s documentation never says how they combine. Everything downstream — the decisiveness, the incentive to drop cache — depends on which price the contest actually reads, so we pin it down two ways, from the cleanest identification to the most complete.
The revealed ranking. A request sent with sort: price and fallbacks disabled forces the router to name the top of its own price ranking: whichever provider it returns is cheapest under whatever formula it runs. On most menus this reveals nothing, because one provider is cheapest on every price at once. The informative menus are those where the input-cheapest and output-cheapest providers differ, so we scan several markets every five minutes and spend only there. In 64 of 64 such splits across two markets, the router picked the lowest input price — including menus where that same provider quoted the most expensive output price on the board. On Xiaomi’s Mimo v2.5 Pro it chose DigitalOcean at $0.40 per million input tokens over three rivals whose output price was $0.87 against DigitalOcean’s $1.50; any formula placing more than about six cents of weight on a dollar of output price would have chosen a rival. The ranking is input price, and the output quote barely enters.
The held-out prediction. The deterministic test names the top of the ranking; to see the whole rule we treat each default-routed request as one draw from the live allocation and ask which quoted prices predict the provider that served it. We sent several thousand such requests, each with a fresh session identifier so the draws are independent, fit the choice on part of the data, and scored it on the rest. Input and output price carry all the predictive content; adding the cached-token quote improves out-of-sample prediction by nothing we can measure. The economic reading is the one the main text leans on: the price a provider advertises for cached tokens does not move the traffic it receives, so quoting a low headline price while billing repeated tokens at the full rate costs it nothing in the contest.
We are measuring behavior, not reading OpenRouter’s source: input price is a transparent stand-in for the production formula, and we label it as such wherever we use it. And “the cached quote adds nothing” is the sharp form of an identification limit, not mind-reading — providers set their cache price as a near-fixed fraction of input price, so the two move together, and a rule that quietly used cache price would be indistinguishable in our data from one that ignored it. Either way the conclusion the argument needs holds: cutting the input price is what wins traffic.
The caching experiments
The central claim is that a repricer bills repeated context as if it were fresh. Testing it means separating a real cache miss from the many mundane things that look like one, so the design is built around controls as much as measurements.
The market we measure against. We reconstruct the public OpenRouter market every five minutes — each provider’s input, output, and cached-token price, with reported traffic and cache use — and record the menu immediately before and after every experiment. Any block in which the menu moved mid-measurement is discarded; a price update in the middle of a run is otherwise easy to misread as provider behavior. When a provider changes price we freeze the event and compare its endpoint against the model builder’s — Z.ai for GLM-5.2 — with paid requests pinned to one provider and fallbacks disabled, so we measure what a buyer actually receives: the bill, the cached-token count, latency, output quality, and failures, rather than inferring any of them from the public quote.
The prompt. Each session opens with a long prompt that establishes state, then repeats related requests over several rounds. Every prompt carries text unique to its session, so nothing can be answered from OpenRouter’s own response cache — we are testing how a provider handles repeated context, not the router replaying an old answer. We run four workloads, from a deliberately transparent canary to realistic traffic:
a canary — “Here is a reference passage:
velvet velvet velvet …” (several thousand repetitions of a freshly drawn word), “reply with the single word ok” — then sent again, verbatim;a record set — a few thousand tokens of randomly generated ledger rows no two sessions share (“
id=7f3a2c amount=$412.18 …”), followed by a question about one row;a conversation — a question, the model’s answer appended, the next question on top of the growing transcript, so each turn is the last plus a little more;
an agent trace — a fixed log of tool calls (request, tool output, request) with only the final instruction changing.
In every case the opening context is unique to the session, the repeated portion is byte-identical across turns, and only the short tail is new — a provider with a working cache should recognize everything but the tail.
The treatments. Around that prompt we vary one thing at a time, and each variation doubles as a test of a mundane explanation. We preserve or discard the session identifier: if misses vanish when it is held fixed, the cause was reassignment, not the cache. We mutate the prefix after 25%, 75%, or all of its content: if misses track how much we changed, the cache is keying correctly. We stretch the wait between requests: if misses appear only after long gaps, that is ordinary expiry. And because a provider sometimes cuts price after losing traffic, we check that cache behavior worsens after a price change, not before. Each treatment gets the identical prompt, timing, and output schedule at the repricer and at the model-builder endpoint, so the only difference is who served it.
What the design isolates. A missed cache has two explanations that point at different actors, and the experiment is built to tell them apart. Pinned to a single provider under a stable session, the aggressive repricers’ caches mostly work — roughly ninety percent of repeated tokens recognized, against ninety-nine at Z.ai — and a ten-turn conversation stays cheaper at the repricer; that design isolates the provider. The losses in the main text concentrate where the session is not stable: at repricing events, and when the router itself moves a conversation to a new provider and lands it on a cold cache. That second cost belongs to the routing layer, not the provider serving the request, and we say which of the two each comparison identifies. The distinction is first-order: in the GLM-5.2 market input and cached tokens are about 98% of volume, so the billing of repeated context is most of what a conversation costs.
Two boundaries close the section. The yardstick is not Z.ai alone — unaffiliated providers who do not build the model recognize repeated tokens at ninety-nine percent as well, so a large cache gap is a property of specific providers, not of being third-party. And the experiment identifies an economic effect, not an intent: it can show that a buyer persistently overpays and that the pattern is inconsistent with random misses, but it cannot tell a deliberate eviction from a poorly built cache or an operational fault. Because repeated requests within a session are not independent, we compute uncertainty at the level of the price event, not the request.
The decisiveness measurement
The contest’s decisiveness \(r\) — the exponent on price — decides whether shaving a quoted price wins a proportional slice of traffic or a super-linear one, so we estimate it from the choices the router actually made rather than from the documentation’s word. On each menu the router faced, with its providers and their live prices, we ask which exponent best explains the provider it chose, and pool the answer across every choice we observed: 1,648 of them, six models, thirty-five providers. The fit is a decisiveness of 1.968, with a 95% interval of [1.856, 2.098] computed at the price-event level — the inverse-square rule the documentation implies, recovered from behavior.
The router’s live decisiveness, fit from the providers it actually chose: 1.968 (95% interval 1.86–2.10) across 1,648 choices — the inverse-square rule, as documented
This is not a reassuring number. A decisiveness of two means the return to undercutting is super-linear: shave your input price and you win far more than a proportional share of the order flow. That is precisely the market in which a misleading headline price pays — a provider that lowers its input price by dropping cache, then bills repeated tokens at the full rate, is rewarded well beyond the discount it only appears to give.
The claim boundary is about time. Decisiveness two is the standing rule, fit across every menu we see; it is not a promise about the hour after any single cut. The response to one provider’s price move is noisier and can fade within a day as it hits its own capacity and the router throttles it back. So the exponent is the shape of the long-run reward for being cheap — which is what a provider setting its price actually faces — not the impulse response to a single repricing, which is smaller and shorter-lived.
How we estimate the cost to users
We want one number: across OpenRouter’s open-weight markets, how much do users overpay because providers bill repeated tokens as if they were fresh? We build it from the daily public data, one model at a time, and then add it up.
For each provider on each model, we compare the fraction of tokens it reports as cached to the fraction the model builder itself reports on the same model. The model builder is the natural yardstick: it runs the reference implementation and has no reason to hide its own cache. Where a provider recognizes fewer repeated tokens than the builder, the difference is billed at the full input price instead of the cheaper cache price, and that gap — the shortfall in recognized tokens, times the price difference — is what the user overpays. We sum this over every provider and every open-weight model.
Two numbers come out. The first is what is being extracted today, given the cache gaps we actually observe: on the order of $3.5 million a year, about 2% of OpenRouter’s roughly $140 million in annual revenue. The second is the ceiling — what would be extracted if every third-party provider stopped recognizing cache entirely. That is about $14 million a year, or 10% of revenue. On GLM-5.2 alone, the market we study most closely, the ceiling is about $5.0 million. We exclude the model builders’ own endpoints from the ceiling: the builder is the honest party, so its volume is not something an attacker can take.
Three judgment calls set the scale, and the honest thing is to state them. We assume a little under half of a typical agentic prompt is repeated, cache-eligible context; we assume honest caching prices repeated tokens at roughly a tenth of the fresh rate; and we annualize from the window where the data is internally consistent. Moving each assumption across a reasonable range moves the ceiling between about $9 million and $18 million a year — the story does not turn on the exact figure. The benchmark is well anchored: the model-builder yardstick is available for the models that make up 92% of open-weight volume, and a conservative default covers the small remainder.
Two honest caveats. Our capture sees only about a tenth of OpenRouter’s total token flow, so if anything these figures understate the market. And the daily cache gap for the aggressive repricers is a persistent one — roughly eleven percent, whether or not they changed price that day — rather than a spike at each repricing. The five-minute experiments earlier in this post show when a provider drops its cache; the daily total simply integrates the steady shortfall that behavior leaves behind. Finally, this is money overpaid by users, not provider profit: we do not observe providers’ costs, so we make no claim about what they keep.
[^klnote]: We originally considered comparing the distribution of payments to the Tullock distribution implied by quoted prices (e.g. via KL divergence). We dropped it: payment shares differ from request shares whenever expected bills differ across providers, so a positive divergence arises under fully honest, heterogeneous pricing — it cannot identify shading.
This is the short, results-first version. The full essay — with a more philosophical market-structure framing, the economics of open source, and the complete analysis and the appendices referenced — can be found here.
Acknowledgments
Thanks to Danning Sui (@sui414), Michael Cieri (@mdcieri37), Yi Sun (@theyisun), Anirudh Pai (@ani_pai), Emilio Andere (@gpuemi), Vipul Ved Prakash (@vipulved), Georgios Konstantopoulos (@gakonst), Mallesh Pai (@malleshpai), Max Resnick (@maxresnick), Nagu Thogiti (@0xnagu), Karan Malhotra (@karan4d), Theo Diamandis (@theo_diamandis), Guillermo Angeris (@guilleangeris), and Haseeb Qureshi (@hosseeb).
Notes
[1] What the industry calls routing is, economically, a two-sided matching problem: capability matching pairs tasks with models, and execution matching pairs requests with inference providers. For the classical theory of matching markets, see David Gale and Lloyd S. Shapley, “College Admissions and the Stability of Marriage,” American Mathematical Monthly 69, no. 1 (1962): 9–15, https://doi.org/10.2307/2312726; and Alvin E. Roth and Marilda A. Oliveira Sotomayor, Two-Sided Matching (Cambridge University Press, 1990). On OpenRouter, capability matching is opt-in — by default the application names its model and the router performs only execution matching.
[2] “20% less capable” is deliberately loose: model evaluations mix absolute metrics (e.g. pass rates on coding or math benchmarks) with relative ones (e.g. Arena’s ELO-style preference ratings), so cross-model capability gaps are directional, not points on a single universal scale.
[3] On GLM-5.2 alone the ceiling is ~$5.0M/yr; across all open-weight markets, ~$14M/yr — about 10% of OpenRouter’s ~$140M annualized revenue as reported in Stripe-acquisition coverage (https://dealroom.co/news/141929-stripe-eyes-openrouter-at-10b-70x-the-startups-revenue/, July 2026). Realized extraction today is far smaller (~$3.5M/yr, ~2%). This is money overpaid by users, not provider profit. The full derivation, assumptions, and robustness checks are in the cost appendix.
[4] Commodity Exchange Act § 4c(a)(5)(C), 7 U.S.C. § 6c(a)(5)(C) (prohibiting “spoofing (bidding or offering with the intent to cancel the bid or offer before execution)”), https://www.law.cornell.edu/uscode/text/7/6c; see also CFTC, “Antidisruptive Practices Authority,” Interpretive Guidance and Policy Statement, 78 Fed. Reg. 31890 (May 28, 2013), https://www.govinfo.gov/content/pkg/FR-2013-05-28/pdf/2013-12365.pdf.
[5] Benjamin Edelman, Michael Ostrovsky, and Michael Schwarz, “Internet Advertising and the Generalized Second-Price Auction,” American Economic Review 97, no. 1 (2007): 242–259, https://www.aeaweb.org/articles?id=10.1257/aer.97.1.242; on buyers receiving counterfeit or lower-quality impressions than promised, Muhammad Ahmad Bashir et al., “A Longitudinal Analysis of the ads.txt Standard,” Proc. ACM IMC ’19 (2019), https://doi.org/10.1145/3355369.3355603. The auction operator itself can be the manipulator: in United States v. Google LLC (E.D. Va., April 17, 2025), the court found Google monopolized the open-web display ad-exchange and publisher ad-server markets, with trial evidence that Google preferenced its own exchange in the auctions it ran, https://en.wikipedia.org/wiki/United_States_v._Google_LLC_(2023).
[6] Irena Gao (@irena_gao), Percy Liang (@percyliang), and Carlos Guestrin (@guestrin), “Model Equality Testing: Which Model Is This API Serving?” (ICLR 2025), https://arxiv.org/abs/2410.20247, which found 11 of 31 commercial endpoints serving distributions that differ from the reference Llama weights; see also Moonshot AI’s K2 Vendor Verifier, https://github.com/MoonshotAI/K2-Vendor-Verifier, which benchmarks third-party Kimi API vendors against the official endpoint.
[7] What the industry calls routing is, economically, a two-sided matching problem: capability matching pairs tasks with models, and execution matching pairs requests with inference providers. For the classical theory of matching markets, see David Gale and Lloyd S. Shapley, “College Admissions and the Stability of Marriage,” American Mathematical Monthly 69, no. 1 (1962): 9–15, https://doi.org/10.2307/2312726; and Alvin E. Roth and Marilda A. Oliveira Sotomayor, Two-Sided Matching (Cambridge University Press, 1990). On OpenRouter, capability matching is opt-in — by default the application names its model and the router performs only execution matching.
[8] Individual routers implement, combine, or skip these stages — on OpenRouter, capability routing is opt-in (openrouter/auto); by default an application names its model and only execution routing runs. See https://openrouter.ai/docs/guides/routing/routers/auto-router.
[9] Uber, “Scaling GenAI at Uber with the GenAI Gateway,” https://www.uber.com/us/en/blog/genai-gateway/; Grab, “Grab AI Gateway,” https://engineering.grab.com/grab-ai-gateway; LinkedIn, “Behind the Platform: The Journey to Create the LinkedIn GenAI Application Tech Stack,” https://www.linkedin.com/blog/engineering/generative-ai/behind-the-platform-the-journey-to-create-the-linkedin-genai-application-tech-stack; Expedia Group, “Gateways, Guardrails and GenAI Models,” https://medium.com/expedia-group-tech/gateways-guardrails-and-genai-models-aa606379164d; Instacart, “Simplifying Large-Scale LLM Processing across Instacart with Maple,” https://tech.instacart.com/simplifying-large-scale-llm-processing-across-instacart-with-maple-63df4508d5be; Palantir, “LLM Capacity Management,” https://www.palantir.com/docs/foundry/aip/llm-capacity-management.
[10] Arena (formerly LMArena/Chatbot Arena), https://arena.ai. Methodology: Wei-Lin Chiang et al., “Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference” (ICML 2024), https://arxiv.org/abs/2403.04132.
[11] “More expensive to manipulate” in the precise sense that raising the decisiveness exponent raises the equilibrium expenditure required to move the allocation: rent dissipation increases in the exponent, reaching full dissipation of the contested rent for exponents above two. Gordon Tullock, “Efficient Rent Seeking,” in Toward a Theory of the Rent-Seeking Society (Texas A&M University Press, 1980), 97–112; Michael R. Baye, Dan Kovenock, and Casper G. de Vries, “The Solution to the Tullock Rent-Seeking Game When R > 2,” Public Choice 81 (1994): 363–380, https://personal.eur.nl/cdevries/Articles/publicchoicesolutiontullock.pdf.
[12] Gordon Tullock, “Efficient Rent Seeking,” in Toward a Theory of the Rent-Seeking Society, ed. James M. Buchanan, Robert D. Tollison, and Gordon Tullock (College Station: Texas A&M University Press, 1980), 97–112; Kai A. Konrad, Strategy and Dynamics in Contests (Oxford: Oxford University Press, 2009), https://global.oup.com/academic/product/strategy-and-dynamics-in-contests-9780199549603; Luis C. Corchón and Marco Serena, “Contest Theory,” in Handbook of Game Theory and Industrial Organization, Volume II (Cheltenham: Edward Elgar, 2018), 125–146, https://www.elgaronline.com/edcollchap/edcoll/9781788112772/9781788112772.00013.xml.
[13] OpenRouter, “Provider Selection,” https://openrouter.ai/docs/guides/routing/provider-selection. The default load balancer selects among providers with weight proportional to the inverse square of price, deprioritizing providers with recent outages.
[14] This is the false-name manipulation familiar from mechanism design: any allocation that is additive across identities and super-linear in effort rewards splitting one identity into several. In a mechanical replay of GLM-5.2’s public menus, ten economically identical duplicate endpoints raise an owner’s share by roughly four times. We see no evidence a provider does this today; the natural fix is to aggregate by owner — or by delivery grade — rather than by endpoint.
[15] The composite price is not explicitly defined in OpenRouter’s documentation; we measure it with deterministic price-sorted requests and by predicting default-routed choices from the quoted prices. See the methodology appendix.
[16] Technically there is a fourth quoted price for per-request/per-image pricing.
[17] Winning today to be the incumbent tomorrow is the classic switching-costs logic: Paul Klemperer, “Markets with Consumer Switching Costs,” Quarterly Journal of Economics 102, no. 2 (1987): 375–394, https://academic.oup.com/qje/article-abstract/102/2/375/1922549.
[18] The claim boundary: when we pin requests to a repricer and hold a session identifier fixed, its cache mostly works — roughly ninety percent of repeated tokens are recognized. The failures concentrate at repricing events and provider switches. Distinguishing a deliberate eviction policy from an operational one would require provider internals we do not have; the billing consequences do not depend on the distinction.
[19] This projection holds later-turn costs at the observed second-request level; directly measured multi-turn pinned sessions are in progress and will replace this extrapolation.
[20] A visible headline price with an economically important, less-salient continuation cost is the shrouded-attributes pattern: Xavier Gabaix (@xgabaix) and David Laibson, “Shrouded Attributes, Consumer Myopia, and Information Suppression in Competitive Markets,” Quarterly Journal of Economics 121, no. 2 (2006): 505–540, https://academic.oup.com/qje/article/121/2/505/1884013.
[21] OpenRouter, “Provider Selection,” https://openrouter.ai/docs/guides/routing/provider-selection. The default load balancer selects among providers with weight proportional to the inverse square of price, deprioritizing providers with recent outages.
[22] Databricks, “Smart Routing in Unity AI Gateway: Match Frontier Quality at 30% Lower Cost Per Task,” 2026, https://www.databricks.com/blog/smart-routing-unity-ai-gateway-match-frontier-quality-30-lower-cost-task. Its router assesses task complexity at the start of a session and commits to a model for the session’s duration, explicitly to preserve provider-side cache efficiency.

