PRIVATE AI · DECISION NOTE · 06 SEPTEMBER 2026
Rent capacity.
Own hardware.
Or use both.
The right choice starts with the workload: model, context length, response time, concurrency, data restrictions and uptime. This is a U.S.-dollar planning reference, not a benchmark, approved architecture or service quote.
| Decision | Cloud GPU rental | Hybrid | Owned workstation / rack / Macs |
|---|---|---|---|
| Best fit to investigate | Variable demand, experiments and short bursts. | Predictable local work with deliberately approved cloud overflow. | Sustained utilization, local control and workloads that fit the chosen hardware. |
| Upfront cost | No GPU purchase. Setup, security and data transfer still take work. | Local equipment plus cloud integration and testing. | Parts, shared tools, assembly, site work and commissioning. See the selected build estimate. |
| Recurring cost | Allocated compute time, storage, backups, networking, support and administration. | Local operation plus cloud usage, connection and two-environment administration. | Power, cooling, maintenance labor, backups, software, spares and replacement cycles. |
| Time to start | No shipping, but region capacity, account/quota approval and setup can delay access. | Depends on local readiness, data policy and a tested overflow workflow. | Supplier delivery + site readiness + the step allowances + diagnostics. Out-of-stock items can dominate the schedule. |
| What you maintain | Provider handles physical hardware. Your responsibility still includes the selected service's data, access and application obligations. | Local hardware and both software environments; validate routing and failure behavior. | Hardware, OS, applications, access, backups, monitoring and a real recovery plan. |
| Main trade-off | Low hardware commitment; ongoing spend and external dependencies. | Flexibility; more integration work and no automatic privacy guarantee. | Control; capital risk, repair responsibility and limited installed capacity. |
Cloud responsibility depends on the service. For example, AWS makes EC2 customers responsible for guest-OS patching, their applications and security-group configuration. AWS shared-responsibility source.
A dated cloud price example
Runpod's public pricing page listed an RTX Pro 6000 offering with 96GB VRAM, 188GB RAM and 16 vCPUs at $2.09/hour. Its standard network storage under 1TB was $0.07/GB/month. These are listing observations, not reserved capacity or an equivalent-performance claim. Runpod pricing, checked 2026-09-06.
| Example | Compute | 500GB network storage | Illustrative subtotal / month |
|---|---|---|---|
| 50 allocated hours | $104.50 | $35.00 | $139.50 |
| 200 allocated hours | $418.00 | $35.00 | $453.00 |
| 730 allocated hours | $1,525.70 | $35.00 | $1,560.70 |
These subtotals exclude additional disks, backups, taxes, paid support, software and your labor. Stop/release policy, storage retention, region, actual capacity and checkout terms must be confirmed. Allocated time is not necessarily productive GPU time. Cloud services have different billing rules; do not apply this price to AWS or an API token service.
Make local operation visible
Use measured wall power after installation. An illustrative—not measured—average of 0.6kW × 730 hours × $0.22/kWh × a 1.2 cooling/overhead factor is $115.63/month. Replace every input with site evidence. Do not use a PSU's nameplate watts as average consumption.
At the guide's illustrative $125/person-hour, two maintenance hours would add $250/month. That makes $365.63/month before spares, backups, licenses, support and repairs. This is a worked arithmetic example, not a maintenance contract or expected power draw for any selected build.
3-year ownership cost = delivered hardware + tools + installation + site work + 36 × monthly operation + replacements − realized resale
Hybrid cost = local ownership cost + cloud usage + integration and ongoing connection costs
No break-even date is asserted. Compare matched successful jobs, latency, memory capacity and uptime first; equal GPU memory does not establish equal performance.
Time and maintenance allowances
For a small proof-of-concept only, allow 2–8 trained-person hours for a cloud environment, or 4–12 additional hours for hybrid routing and access tests, plus downloads and approvals. These are author planning ranges; client production scope may be much larger. Owned hardware assembly ranges are itemized by lesson in the guide.
- Daily automated checks: service health, storage capacity, backup job results and alerts with a named owner.
- Monthly planned review: updates in a test window, access changes, logs, storage growth, temperatures, UPS state and selected workload benchmarks. Inspection does not authorize opening electrical equipment.
- Quarterly planned exercise: restore a real backup and record recovery time. Test failure behavior without risking production data.
- Hardware service: follow the delivered unit's manufacturer schedule for filters, fans, batteries and warranty service. Quote incidents, replacements and call-outs separately.
Mac Studio is a different architecture
The guide includes three independent higher-memory Mac Studios and a lower-cost alternative. Each has its own operating system and memory. A network cable does not pool their RAM. Distributed inference needs a specifically supported runtime, networking design and measured workload validation. Mac unified memory and NVIDIA VRAM are not interchangeable performance measures.
The Mac alternative changes the compute-node purchase lines; it does not claim the same model capacity, training support or throughput. Check each variant's availability, enclosure fit, software support and actual benchmark before selecting it.
Before a client commitment
Record country/ZIP, budget, hardware already owned, models, concurrent users, retention, backup target, response-time target and required availability. Then obtain written component/rail/power approval, refreshed delivered prices and labor quotes. The rack servers remain configurable platforms requiring a supplier's exact internal BOM.