Skip to content
Hussh
Connect MCP

The Garage Supercomputer

Public essay: what it actually costs, in July 2026, to build a sovereign AI factory on a residential power drop — and why tokens per watt is the wrong metric.

Content

TL;DR: We priced a private, bare-metal AI supercomputer for a garage against July-2026 market data: a Blackwell-class pod is buyable off the shelf from ~$90k, energy turns out to be only 3–15% of cost per token (utilization — tokens per depreciated dollar — is the real metric), and the binding constraints are a nighttime noise ordinance and an insurance rider, not FLOPS.

Status as of 2026-07-27: see body.

By the founder of Hussh 🤫 — from the garage where Hussh started.

Everyone tells you frontier AI belongs to people with gigawatt campuses and sovereign wealth funds. We wanted to know what the floor actually is. Not in vibes — in purchase orders. So we designed a private, bare-metal AI factory sized for a two-car garage on an ordinary residential power drop, priced every component against July-2026 street data, and checked the arithmetic twice.

The results surprised us three times. Once about the metric everyone uses. Once about what you can actually buy. And once about what the real constraints turn out to be.

The metric everyone quotes is the wrong one

The brief we started with — the brief everyone starts with — was "maximum tokens per watt." It sounds rigorous. It is the wrong number, and here is why.

In the Pacific Northwest, residential power runs about 19¢ per kilowatt-hour, all-in. Take the best measured efficiency numbers in the open literature — roughly 2.8 tokens per joule for NVIDIA's B200, about 2.55 for AMD's MI355X, about 0.9 for the H100 (SemiAnalysis InferenceMAX, provisioned-power basis) — and the electricity cost of a million tokens works out to two to eight cents.

Now amortize the hardware. A serving node depreciates roughly 30–35% a year — we watched used H100s go from ~$50k at the 2024 peak to $18–22k refurbished by mid-2026. Spread a node's price over three years of output and capex comes to $0.15 to $1.70 per million tokens, depending almost entirely on one variable: how busy you keep the machine.

Energy is 3–15% of your cost per token. Depreciating silicon is the rest. The governing metric for anyone below hyperscale is not tokens per watt — it is tokens per depreciated dollar, which is a fancy way of saying utilization. A mediocre GPU running at 70% beats a marvel of efficiency running at 15%, and it isn't close. Our design reflects that: the single KPI on the wall is trailing-7-day tokens per depreciated dollar, and the rule is you don't buy the next node until the current one is measurably hungry.

What you can actually buy in July 2026

The second surprise: the best garage-scale silicon isn't the one in the headlines, and the one in the headlines isn't buyable anyway.

A GB300 NVL72 rack draws 120–132 kilowatts. A fully upgraded 400-amp residential service delivers 76.8 kW continuous — the flagship rack needs almost double the entire house. HGX B200 systems run $450k+ with waitlists. AMD's MI355X benches beautifully but effectively cannot be purchased at small scale — it's an OEM quote, not a cart button.

And the used 8×H100 chassis that seems like the obvious value play at $150–180k? At 0.9 tokens per joule it is three times less efficient than current silicon, with two more generations of depreciation still to absorb. The obvious buy is the wrong buy.

What's left is quietly remarkable: the RTX PRO 6000 Blackwell Server Edition — 96 GB per card, native FP4, explicitly a data-center product, in stock at retail for $12,781 a card. Eight of them in a standard PCIe server gives you 768 GB of VRAM at ~6.5 kW — two ordinary 240V circuits — for $130–155k from a tier-one OEM, or $90k from tinycorp's tinybox pro v2 if you like your vendors scrappy. That is enough memory to serve GLM-5.2 at FP4, Qwen3.5-397B at FP8, or a fleet of gpt-oss-120b replicas, through the same open stack (SGLang, vLLM, an OpenAI-compatible endpoint) the big labs use.

No proprietary interconnect. No rack-scale liquid loop. Standard chassis with real resale liquidity. The frontier of ownable AI is an air-cooled box you can lift with a friend.

The real constraints are wonderfully mundane

Physics was supposed to be the hard part. It isn't. A 400A service upgrade ($15–30k), six tons of cooling at full build (~$25k of mini-splits, mostly bypassed by Pacific Northwest air — this climate gives you free cooling more than 95% of hours), and the power bill for a 10 kW pod lands under $1,000 a month.

The actual binding constraints, in order: a nighttime noise ordinance (Washington caps residential property-line noise at 45 dBA after 10pm — that single number drives fan curves, acoustic treatment, and condenser placement more than any benchmark); an insurance rider (a standard homeowners policy covers about $2,500 of business equipment — you schedule the fleet on an inland-marine floater or you are self-insured and don't know it); and storage timing (NAND contract prices ran +70–75% quarter-over-quarter into 2026 and enterprise hard drives are effectively sold out for the year — the old "wait, it gets cheaper" rule is currently inverted; buy your drives with the first purchase order).

Nobody writes papers about noise ordinances. They should. That is where sovereign compute actually gets won or lost at this scale.

Fault tolerance by pairs, not clusters

The design rule that kept us honest: two of everything, replicated — not a distributed system that needs a team to babysit. Two used 100G switches at $1,600 each (no NOS license). Every node dual-homed. Two WAN providers plus a wireless third. Two ZFS storage nodes replicating hourly, with an open S3 layer (Garage) spanning both plus an encrypted replica in a rented half-rack downtown — your hardware, your keys; only the building is rented.

Just as instructive is what we rejected, because 2026's open-infrastructure landscape has sharp edges: three-node Ceph (a documented small-cluster trap — lose one node and it sits degraded, unable to heal), MinIO (community edition gutted in 2025, repository archived in 2026 — treat it as gone), and kdb+ (opaque six-figure licensing in a world where ClickHouse and ArcticDB are free and excellent). Boring, open, paired technology beats clever, licensed, singular technology at every layer of this build.

The honest economics

Three-year, straight-line, with energy included:

Build15% utilization40%70%
8× RTX PRO 6000 node (~$140k), big-MoE FP4$1.70/Mtok$0.67$0.41
Same node, 120B-class replicas$0.73$0.29$0.18
HGX B200 (~$450k)$0.80$0.31$0.19

Frontier open-weight APIs bill roughly $0.30–$2.00 per million tokens. So the crossover is real and reachable: above roughly 30–40% sustained utilization, owning beats renting — and every token stays yours. Below that, you are paying a sovereignty premium, and you should say so out loud rather than hide it in a spreadsheet. Sometimes the premium is worth paying: if the corpus is the edge, the corpus does not leave the building.

Why bother

Because the alternative is a world where intelligence is something you can only rent, from three landlords, on their terms, with your data as the rent.

A garage pod proves the other path closes. For the price of a nice car — $90k entry, ~$340k done properly with N+1 power, fire suppression, and an offsite replica — an individual, a family office, a clinic, a small fund, a warehouse owner can run frontier-class open models on hardware they own, behind locks they control, producing tokens that cost less than the API above modest utilization and answer to no one. That is the thesis behind 🤫 The AI Factory: every spare bay is a potential token factory; own the compute for your own work, and let idle hours earn.

Twenty-five years of building AI platforms inside the biggest companies in the world taught me where the leverage lives. It lives with whoever owns the machine. This time, that can be you.

Your data, your business. 🤫

The full engineering blueprint — bill of materials, electrical and thermal math, fault-domain design, and the 90-day build plan — stays with our GPs and partners. Building your own factory and want to compare notes? Reach us at hushh.ai.

Sources