Requesty, a London AI gateway, says the models carrying its production traffic are changing almost every month — and that open-weight systems now handle about half the tokens.

On August 24 the company reported that the most-used model on its router changed 11 times between January 1 and August 14. Eight different models held the top spot, each for about three weeks on average. On February 13, Anthropic's Opus 4.6 accounted for 78 percent of tokens routed through the service. By August 14, DeepSeek V4 Flash led with 32.8 percent. A separate company note this month said open-weight models carried 49.9 percent of tokens but only 13.2 percent of customer spend, or about $0.18 per million tokens versus $1.17 for closed models. Those figures come from Requesty's own logs.

The company argues that this churn is why applications should not hard-code a single provider. Developers point the OpenAI SDK at router.requesty.ai and get a catalog the firm now lists at more than 600 models, with routing by cost, latency, or availability, automatic failover, caching, and per-key budget caps. Pay-as-you-go customers pay a 5 percent markup on model prices, or they can attach their own provider keys. A July data post said the gateway routed across 136 providers and 502 distinct models in June, while 2,084 client applications reached the endpoint that month.

Requesty was founded in 2023 by Daniel Trugman and raised a $3 million seed round in September 2025, led by 20VC, with Tapestry VC, Insiders Ventures, and Tiny Supercomputer also participating. Its site currently claims more than 70,000 developers and more than 90 billion tokens processed each day. Enterprise features include single sign-on, role-based access, EU processing in Frankfurt, and a data-processing agreement.

The gateway sits in a field that also includes OpenRouter, Portkey, and self-hosted LiteLLM. Requesty's latest mix is a concrete sign that, on at least one busy router, cheap open-weight inference is no longer a side channel — even if most of the revenue still accrues to closed models.