← Back to assembly

PRIVATE AI · DECISION NOTE · 06 SEPTEMBER 2026

Rent capacity.
Own hardware.
Or use both.

The right choice starts with the workload: model, context length, response time, concurrency, data restrictions and uptime. This is a U.S.-dollar planning reference, not a benchmark, approved architecture or service quote.

DecisionCloud GPU rentalHybridOwned workstation / rack / Macs
Best fit to investigateVariable demand, experiments and short bursts.Predictable local work with deliberately approved cloud overflow.Sustained utilization, local control and workloads that fit the chosen hardware.
Upfront costNo GPU purchase. Setup, security and data transfer still take work.Local equipment plus cloud integration and testing.Parts, shared tools, assembly, site work and commissioning. See the selected build estimate.
Recurring costAllocated compute time, storage, backups, networking, support and administration.Local operation plus cloud usage, connection and two-environment administration.Power, cooling, maintenance labor, backups, software, spares and replacement cycles.
Time to startNo shipping, but region capacity, account/quota approval and setup can delay access.Depends on local readiness, data policy and a tested overflow workflow.Supplier delivery + site readiness + the step allowances + diagnostics. Out-of-stock items can dominate the schedule.
What you maintainProvider handles physical hardware. Your responsibility still includes the selected service's data, access and application obligations.Local hardware and both software environments; validate routing and failure behavior.Hardware, OS, applications, access, backups, monitoring and a real recovery plan.
Main trade-offLow hardware commitment; ongoing spend and external dependencies.Flexibility; more integration work and no automatic privacy guarantee.Control; capital risk, repair responsibility and limited installed capacity.

Cloud responsibility depends on the service. For example, AWS makes EC2 customers responsible for guest-OS patching, their applications and security-group configuration. AWS shared-responsibility source.

A dated cloud price example

Runpod's public pricing page listed an RTX Pro 6000 offering with 96GB VRAM, 188GB RAM and 16 vCPUs at $2.09/hour. Its standard network storage under 1TB was $0.07/GB/month. These are listing observations, not reserved capacity or an equivalent-performance claim. Runpod pricing, checked 2026-09-06.

ExampleCompute500GB network storageIllustrative subtotal / month
50 allocated hours$104.50$35.00$139.50
200 allocated hours$418.00$35.00$453.00
730 allocated hours$1,525.70$35.00$1,560.70

These subtotals exclude additional disks, backups, taxes, paid support, software and your labor. Stop/release policy, storage retention, region, actual capacity and checkout terms must be confirmed. Allocated time is not necessarily productive GPU time. Cloud services have different billing rules; do not apply this price to AWS or an API token service.

Make local operation visible

Use measured wall power after installation. An illustrative—not measured—average of 0.6kW × 730 hours × $0.22/kWh × a 1.2 cooling/overhead factor is $115.63/month. Replace every input with site evidence. Do not use a PSU's nameplate watts as average consumption.

At the guide's illustrative $125/person-hour, two maintenance hours would add $250/month. That makes $365.63/month before spares, backups, licenses, support and repairs. This is a worked arithmetic example, not a maintenance contract or expected power draw for any selected build.

3-year ownership cost = delivered hardware + tools + installation + site work + 36 × monthly operation + replacements − realized resale

Hybrid cost = local ownership cost + cloud usage + integration and ongoing connection costs

No break-even date is asserted. Compare matched successful jobs, latency, memory capacity and uptime first; equal GPU memory does not establish equal performance.

Time and maintenance allowances

For a small proof-of-concept only, allow 2–8 trained-person hours for a cloud environment, or 4–12 additional hours for hybrid routing and access tests, plus downloads and approvals. These are author planning ranges; client production scope may be much larger. Owned hardware assembly ranges are itemized by lesson in the guide.

Mac Studio is a different architecture

The guide includes three independent higher-memory Mac Studios and a lower-cost alternative. Each has its own operating system and memory. A network cable does not pool their RAM. Distributed inference needs a specifically supported runtime, networking design and measured workload validation. Mac unified memory and NVIDIA VRAM are not interchangeable performance measures.

The Mac alternative changes the compute-node purchase lines; it does not claim the same model capacity, training support or throughput. Check each variant's availability, enclosure fit, software support and actual benchmark before selecting it.

Before a client commitment

Record country/ZIP, budget, hardware already owned, models, concurrent users, retention, backup target, response-time target and required availability. Then obtain written component/rail/power approval, refreshed delivered prices and labor quotes. The rack servers remain configurable platforms requiring a supplier's exact internal BOM.