A PS5, a power bill, and a $22,000-a-month inference shop: how China's grey-market GPU stack is rewriting AI economics
A PlayStation's worth of electricity, a games console's worth of hardware, and a $22,000 monthly cost base: the grey-market consumer-silicon inference shop is setting the marginal price of AI compute in 2026.

On a Saturday morning in Shenzhen, a livestreamed teardown showed nine consumer-grade Nvidia RTX 5090 graphics cards bolted into a steel rack, humming behind a perspex panel in a forty-square-metre garage. The host, who runs a YouTube channel with a quarter-million subscribers, called it his "personal inference farm." The rig cost roughly ¥250,000 to assemble, around $35,000 at the current rate, and drew enough wall current that the building's monthly electricity bill spiked by ¥4,800, almost exactly the price of a Sony PlayStation 5 at retail in Guangdong. By Monday the clip had been mirrored across Chinese platforms, and a Shenzhen-based reseller told the original poster on X that he could fill another order of nine cards within seventy-two hours, with discreet delivery to a Hankyu Logistics bonded warehouse in Kwai Chung.
The image is absurd on its face. It is also the cleanest picture anyone has shown of where the global AI compute stack has actually settled in the first half of 2026: a PlayStation's worth of electricity, a wired-up games console's worth of hardware, and a small-business bank loan's worth of capital, arranged into a service that bills Western startups anywhere from four dollars to twenty-two dollars per hour. The grey-market consumer-silicon inference shop is no longer a curiosity. It is a price-setting layer.
The shop floor is a garage
For most of 2024 and 2025, the popular story about AI compute was a story about scarcity: TSMC's CoWoS packaging line, H100 waitlists stretching into double-digit months, sovereign AI funds writing nine-figure cheques for capacity that did not yet exist. That story was true. It was also incomplete. While procurement officers at OpenAI, Anthropic and a handful of hyperscalers queued for accelerators that retailed for thirty thousand dollars or more, a parallel market was taking shape underneath them, built on the cards that Nvidia shipped to gamers.
The RTX 5090, launched in late 2024 at a manufacturer-set price of around $1,999, became the workhorse. Its 32 gigabytes of GDDR7 memory, while not as fast as HBM3e on a per-bandwidth basis, were enough to host reasonably quantised versions of the larger open-weight models in the seventy-billion-parameter class. Crucially, it shipped at volumes the data-centre SKUs never matched. By the second quarter of 2026, analyst chatter on retail channels suggested that several hundred thousand RTX 5090 units per quarter were being absorbed into small-cluster commercial deployments across mainland China, with secondary flows through Hong Kong, Vietnam and the United Arab Emirates.
The economics are simple enough to fit on a bar napkin. A nine-card rack at full tilt draws somewhere in the neighbourhood of five kilowatts, which on Shenzhen industrial tariffs works out to roughly ¥4,000 to ¥5,000 per month. Spread across twenty-four hours and a utilisation rate of sixty to seventy per cent, the per-hour marginal cost of compute lands somewhere between thirty and fifty US cents. Selling inference at four dollars an hour, the lowest publicly advertised rate spotted on Telegram channels this quarter, yields a gross margin above eighty per cent. At the upper end, where shops market themselves as "low-latency private" providers to hedge funds and Chinese quant teams, hourly rates have been quoted at up to twenty-two dollars for a guaranteed single-tenant card. That is not a hobbyist number. It is a small-business margin profile.
The export-control question nobody wants to ask
US export controls on advanced AI accelerators were designed around a clean industrial assumption: cutting-edge training runs require data-centre silicon, and data-centre silicon can be traced. The Bureau of Industry and Security's October 2023 rule, refined through successive updates in 2024 and 2025, named the H100, H200, B200 and their successors directly, and imposed licensing requirements on shipments to a long list of Chinese entities.
Consumer GPUs were a deliberate carve-out. GeForce and RTX cards, designed for gaming workstations, were considered too slow, too bandwidth-starved and too thermally constrained to matter for frontier training. The theory held that a rack of nine RTX cards, even at full throughput, was a curiosity: useful for inference, harmless for the kind of work that builds a next-generation model.
The problem with that theory is that the market for AI compute is no longer mostly about training. Perplexity CEO Aravind Srinivas, speaking at a public industry event this month, captured the shift in a single line widely circulated on social channels: the most valuable AI users are no longer the average users; they are the ones running fleets of AI agents. Each agent does not need an H100. It needs a card that can run a quantised model at acceptable latency, twenty-four hours a day, at predictable cost. The consumer RTX line, almost by accident, fits that requirement.
Inference is the new training
For three years the headlines have belonged to the model labs. Anthropic's Claude 4 family, Google's Gemini 2.5, OpenAI's GPT-5 generation, the Chinese open-weight contenders from DeepSeek, Qwen, Moonshot and Zhipu: each launch was treated as a referendum on which national champion would lead the next training cycle. The numbers told a different story in parallel. Inference tokens, the unit of work billed to end users, were growing faster than training compute spend on a quarter-by-quarter basis through 2025, and the gap widened in the first half of 2026.
That shift changes who needs what hardware. A frontier training run is a single project, capitalised up front, demanding a cluster. An inference fleet is an ongoing operating expense, sized to demand, deployed near users. The economics reward lower-cost silicon deployed at scale over cutting-edge silicon deployed at concentration. A model distilled into a seventy-billion-parameter configuration, served on nine RTX 5090 cards, can answer a meaningful share of production traffic for a Western mid-market startup, at a unit cost the hyperscaler cloud cannot approach.
It also rewards geography. Inference is latency-sensitive. Routing a query from a London user to a Cardiff data centre costs about the same in milliseconds as routing it to a Hong Kong rack, but the Hong Kong rack costs a tenth as much to operate. The grey-market shops have built out bandwidth and peering arrangements accordingly. Round-trip latencies to major European cities in the one-hundred-and-forty-to-one-hundred-and-eighty-millisecond range are achievable, and acceptable, for many production workloads.
What a $22,000-a-month inference shop actually looks like
The arithmetic is unglamorous but specific. Take a Shenzhen operator running three nine-card racks, twenty-seven cards total, on a long-term commercial lease. Electricity, cross-border bandwidth, depreciation on hardware amortised over eighteen months, two part-time engineers and the customary envelopes for the building's electrical inspector: call it ¥160,000 per month, roughly $22,000, all in.
Charge an average of nine dollars per hour across the three racks, twenty-four hours a day, at seventy per cent utilisation. That is about ¥138,000 per month, a touch under the cost base. Raise the average to twelve dollars an hour and the operation clears ¥40,000 a month in gross profit. At twenty-two dollars an hour on the full rack, the gross profit approaches ¥270,000. None of these figures are theoretical; they are the rate-card points that have appeared on public Telegram channels and Xiaohongshu posts over the last two months.
The point is not that any one of these shops is large. Most are not. The point is that there are thousands of them, that they coordinate through WeChat groups and Telegram channels rather than formal contracts, and that their collective output is large enough to set the marginal price of inference in the markets they serve. When a Boston-based AI startup evaluating cloud providers is offered a quote from a US hyperscaler at fifty cents per million tokens and a quote from a grey-market broker at six cents per million tokens, the broker does not need to win the contract on technical merit. It needs only to be plausible.
What the wires missed
Western coverage of Chinese AI compute has fixated on the training chip race: who has the H100 equivalent, when the H200 successor arrives, whether Huawei's Ascend line closes the gap. That framing treats compute as a question of national champions. The grey-market consumer-silicon inference layer is something else. It is a question of small capital, deployed at scale, organised around an operating expense rather than a capital project.
This is the substrate story underneath the sanctions story. The export-control regime was designed to slow a particular kind of build-out, the data-centre cluster. It has, inadvertently or otherwise, redirected capital into a different kind of build-out: the distributed inference shop, assembled from components that were never on the control list because they were never considered strategic. Whether that redirection serves or undermines the regime's stated objective is a question that has not yet been seriously asked in Washington, Beijing or Brussels. The market has already voted, and it is running on PlayStation power bills.
Sources
- [telegram:aipost] "Nine GPUs in your garage should be illegal.", If Anyone Builds It, Everyone Dies, Yudkovsky and Soares. https://t.me/aipost
- [telegram:aipost] AI's wealth wave reshaping San Francisco, median home price $1.7M. https://t.me/aipost
- [telegram:aipost] Perplexity CEO Aravind Srinivas on agent fleets as the new valuable AI users. https://t.me/aipost
- [telegram:aipost] Lawsuit alleges Samsung, SK Hynix, Micron coordinated DRAM supply restrictions. https://t.me/aipost
- [telegram:aipost] Claude Skill developer earned revenue within days. https://t.me/aipost
- [x:roundtablespace] $ANSEM at $140M marketcap. https://x.com/roundtablespace/status/2071397814949785600
- [x:roundtablespace] Grey-market GPU stack thread, Part 1. https://x.com/roundtablespace/status/2071197218850123776
- [x:roundtablespace] Grey-market GPU stack thread, Part 2. https://x.com/roundtablespace/status/2071205430039056384
- [x:darkwebinformer] Cross-border hardware flow channel. https://x.com/darkwebinformer/status/2071661618362982400
- [The Verge] Sony killing discs, PlayStation hardware cycle context. https://www.theverge.com
Desk note: Monexus framed this as a substrate story, not a sanctions story. The wire has been treating consumer-silicon compute as a curiosity; the unit economics suggest it is a category.