NVIDIA Repair-Flow Economic Model

A model of the income NVIDIA loses each year to its GPU repair/replacement process, and what a faster reverse-logistics system is worth. Built for the Lonny Orona / Greg DeLoccio meeting prep, grounded in the Alex Zhu reverse-logistics interview and the shiny-object process walk-through. A live animated prototype lives at lab/reverse-logistics-sim/index.html.

Status: strawman, updated 2026-06-26. $r$, $f$, $\tau$, $e$ and the repair-category structure are now interview-anchored (Greg 2026-06-26); $P$ and $C_r$ remain guesses to be corrected by the customer. The model is a conversation instrument, not a conclusion — ends with open questions, per RDI. Companion sim + math: 2026-06-24-reverse-logistics-sim-math; source at lab/reverse-logistics-sim/index.html. Live inputs use $e=0.99$ and the $C_r$ decomposition below ($C_r(60\text{d})\approx$10.8\text{k}$), not the old flat $$3\text{k}$.


Notation

SymbolMeaningSeedSource
$N$installed base (GPUs in service)100k🔴 guess (Meta alone ≈100k)
$f$annual failure rate → returns/yr $=Nf$9%🟡 our estimate
$e$entitlement share (returns NVIDIA owns)99%🟢 interview — <1% of RMAs ever rejected [G-0626]
$r$repair rate = fraction of failures fixed, not replaced60%🟢 interview (60/100)
$\bar r$technical repair ceiling90%🟢 interview — “chip rarely the problem”; ~90% known-fix [G-0626]
$\kappa_c$reman-line capacity cap on $r$75%🟡 interview-implied — repair competes with new production, revenue wins ~9/10 [G-0626]
$P$market price = value of any unit (sold-out)$30k🔴 guess
$C_r$total cost of one repair $= \gamma + \rho$~$10.8k @60d🔴 guess (decomposition below; $C_r$ grows with $\tau$)
$\tau$serial turnaround: warehouse-receipt → back in stock~60d e2e🟢 interview — ~60d end-to-end, 27–40d warehouse-to-stock [G-0626]
$\tau_a$approval-gauntlet latency (pre-warehouse)~10d🟢 interview — case mgmt → quality → finance, untracked [G-0626]
$\tau_c$customer replacement clock (from stock)~1–2 wk🟢 interview — decoupled from $\tau$ by the spares pool [G-0626]
$\rho_{ref}$refurb tail share (diagnose-from-scratch)10%🟢 interview — reman+repair are the other ~90% [G-0626]
$\delta$cost of capital12%assumption

The model

Premise — sold out. NVIDIA’s AI GPUs are capacity-allocated: demand exceeds supply and every unit built is sold. So a failed unit is either repaired (cost $C_r$) or replaced from new, and each replacement is a lost sale at the full price $P$ (the build cost $C_m$ is sunk either way and cancels). The static annual loss is:

$$\boxed{;\text{Loss} = Nf\big[(1-r),P + r,C_r\big];}$$

The lever: a repaired unit costs $C_r$; a replaced unit costs $P$. Since $P \gg C_r$, the marginal value of repairing one more unit instead of replacing it is $P - C_r \approx P$. That gap is the entire wedge.

Repair cost, decomposed

$C_r$ splits by when cost is incurred — intake (paid on every return, including the ~40/100 that can’t be fixed) plus fix (only on successes):

$$C_r = \underbrace{\gamma}{\text{intake}} + \underbrace{\rho}{\text{fix}}, \quad \gamma = \underbrace{L_{\text{in}}}{\text{freight}} + \underbrace{D}{\text{diagnosis}} + \underbrace{A}{\text{RMA/IT}}, \quad \rho = \underbrace{W}{\text{CM labor}} + \underbrace{M_c}{\text{consignment parts}} + \underbrace{M_t}{\text{turnkey parts}} + \underbrace{L_{\text{out}}}_{\text{return freight}}$$

Charging intake on all returns, the loss is: $$\text{Loss} = Nf\big[\gamma + r\rho + (1-r)(P+\sigma)\big]$$ (the extra $Nf(1-r)\gamma$ over the boxed form is the wasted diagnosis + freight on units that can’t be saved; $\sigma$ = scrap).

Two properties of $C_r$: (1) $M_c$ (NVIDIA-owned chips/boards) approaches $C_m$ when the fault is in expensive silicon, so $r$ and $C_r$ are joint outputs of the fault mix, not independent knobs; (2) $C_r$ is lumpy across three repair categories (Greg 2026-06-26): reman (current-revision, known fix, same live production line) and repair (N-minus-one, dedicated line) together are ~90% of volume — high-success, recall-shaped, cheaper; refurbishment is the ~10% tail — diagnose-from-scratch, may scrap, and is the pricier onesie-twosie work. The sim blends this as $C_r^{\text{eff}} = C_r(\tau),[1 + \rho_{ref}(\mu-1)]$ with a refurb cost multiple $\mu$ (strawman — we have the ~90/10 mix but no per-category cost yet). The recall-shaped reman+repair tail is what the $2.81B FY26 warranty reserve (filed 10-K; not the mis-aggregated $8.22B) prices, and the shape parametric insurance targets.

Repair rate is capacity-capped, not silicon-limited

The interview reframes what holds $r$ down. Greg: the chip is rarely the problem and ~90% of returns are known-defect/known-fix — so the ~40 of 100 that get replaced from new instead of repaired are not dead silicon. They are lost to (a) reman-line capacity — repair shares the assembly line with new production and, when constrained, revenue wins ~9/10 (a joint, negotiated call, not the CM’s); and (b) speed — units that miss the SLA window get filled from new stock. So the effective repair rate is a technical ceiling capped by capacity:

$$r_{\text{eff}}(\tau) = \min\big(\bar r(\tau),; \kappa_c\big), \qquad \bar r(\tau) = r_0 + k,(\tau_0 - \tau)$$

with $\bar r \le 0.90$ and $\kappa_c \approx 0.75$. This resolves the model’s old open question #1 (“is the 40% technical or speed-limited?”) toward capacity + speed — which makes the wedge larger, because both are things a faster, better-instrumented process can move, whereas dead silicon is not. Note the cap only binds below $\tau\approx27$ days: speed alone lifts $r$ to the mid-70s, but pushing past $\kappa_c$ needs an allocation decision, not just faster logistics.

Time

Turnaround $\tau$ enters three ways. By Little’s Law, $Nf,r,\tau$ units sit trapped in the pipeline, each worth $P$. And $r$ itself depends on speed: $r(\tau) = r_{\text{tech}},[1 - G(\tau)]$ — a technical ceiling $r_{\text{tech}}$ times the fraction fixed before the SLA window expires. Slow turnaround forces technically-repairable units onto the $P$ path. The full time-dependent loss:

$$\text{Loss}(\tau) = \underbrace{\delta P,Nf,r(\tau),\tau}_{\text{trapped capital}}

  • Nf\Big[\underbrace{\gamma_0 + a\tau}_{\text{admin dwell}}
  • r(\tau)\big(\rho_0 + w\tau\big)
  • \big(1-r(\tau)\big)(P+\sigma)\Big]$$

where $\rho_0 + w\tau$ captures re-work labor (re-diagnosis, handling) that grows in a fragmented process.

Two clocks — the correction the interview forces

The single $\tau$ above conflates two clocks Greg explicitly separated:

  • Serial repair clock $\tau$ — warehouse-receipt → back-in-stock, ~27–40 days (the only stretch NVIDIA tracks weekly). This is what squeezes $C_r$ and traps capital, and what drives pool replenishment.
  • Customer replacement clock $\tau_c$ — case-submit → replacement received. Because the spares pool buffers the customer, this is ~1–2 weeks from stock and is largely decoupled from the 30+-day serial clock — until the pool goes dry, when it balloons back toward $\tau$: $$\tau_c \approx \tau_a + s,\text{(shipFromStock)} + (1-s)\cdot 0 ;;\text{healthy pool}, \qquad \tau_c \to \tau_a + \tau ;;\text{on stockout}$$ In the sim, the stockout fraction $s(\tau)$ grows as the serial clock lengthens (the pool can’t keep up), so faster $\tau$ buys customer SLA indirectly — by keeping the pool full — not by making any single repair faster.

Why it matters for the wedge. NVIDIA measures the serial leg and not the customer leg; the ~1–2 weeks of approval latency $\tau_a$ (the three-gate gauntlet) is invisible because the clock only starts at warehouse arrival. So the customer-experienced SLA is understated, and its cheapest lever is software, not logistics.

The approval-latency lever (new, and nearly free)

$\tau_a\approx10$ days of case-management → quality → finance gates run before the unit ships, yet <1% of RMAs are ever rejected — pure latency, not a filter. Auto-fast-tracking the known-defect share $\kappa$ (~90% of volume — “6,000 H100s from this date range”) collapses it: $$\tau_a^{\text{eff}} = (1-\kappa),\tau_a ;\approx; 1\text{ day}.$$ This moves $\tau_c$ (and the customer SLA) directly, at ~zero physical cost, and is independent of the physical $\tau$ lever. It’s modeled in the sim as a toggle, parallel to the owed-back reconciliation toggle.

What a faster system is worth

Reducing $\tau$ by $\Delta\tau$ recovers: $$\text{Value} \approx \underbrace{Nf,\Delta r,(P - \rho)}_{\text{replace}\to\text{repair (dominant)}}

  • \underbrace{\delta P,Nf,r,\Delta\tau}_{\text{freed capital}}
  • \underbrace{Nf,r,w,\Delta\tau}_{\text{less re-work}}$$

Speed doesn’t make a single repair cheaper ($\rho$ barely moves). It re-routes failures from the $P$ path to the cheap $\rho$ path and unfreezes trapped capital. The leverage is large because NVIDIA is sold out — every trapped or given-away unit is worth the full $P$.


Illustrative numbers (strawman, interview-anchored τ)

$N=100k$, $f=9%$ → 9,000 returns/yr ($e=99%$ → ~8,900 entitled); $P=$30k$; $C_r^{\text{eff}}(\tau)=[1500+300+150\tau]\times1.10$. Today is ~60d end-to-end and best-case ~30d (Greg 2026-06-26) — not the 30d/5d the prior version used:

Serial τ$r_{\text{eff}}$ReplacementsRepairsTotal lossRecovered
60d (today)60%3,600 × $30k = $108M5,400 × $11.9k = $64M$172M
30d (best case)73.5%2,385 × $30k = $72M6,615 × $6.9k = $46M$118M+$54M

Loss scales linearly with $N$ — 10× at Meta’s 1M-GPU target. A replacement still costs several× a repair, so the replace→repair re-routing dominates.

Two caveats the interview adds: (1) below $\tau\approx27$d the reman-line capacity cap ($\kappa_c\approx75%$) binds — further gains need an allocation decision, not faster logistics; (2) this table is the serial clock. The customer clock is separate: ~$\tau_a+$ ship-from-stock ≈ 1–2 weeks when the pool is healthy, and auto-fast-tracking approval ($\tau_a: 10\to1$ day) improves the customer SLA independently of τ.


Assumptions / where it breaks

  • Sold out is load-bearing. Valid only while GPUs are supply-allocated. If demand softens or NVIDIA holds finished inventory, the replacement cost slides from $P$ down toward $C_m$ (build-extra instead of lost-sale). Greg 2026-06-26 reinforces this: warranty is filled from new units when repaired stock is short, and new buys are “what Jensen cares about.”
  • Single SKU, single price. DGX vs board vs cable have very different $P$ and $C_r$; the model blends them. The refurb-tail cost multiple $\mu$ is a strawman ($\rho_{ref}=10%$ known, per-category cost unknown).
  • Capacity cap $\kappa_c$ is a policy variable, not a constant. It’s the reman-line allocation between repair and revenue — a negotiated call that can move. Modeled as a hard cap; really a bargaining outcome.
  • The two clocks are coupled through one pool. The customer clock $\tau_c$ is modeled via a stockout fraction $s(\tau)$; the real coupling is a service-parts (base-stock) inventory problem — pool size $B$, demand variance, and replenishment lead time. Baxter Planning is literally the tool NVIDIA uses for this; our model is a reduced form of it.
  • Flat failure rate on the installed base (no early-life/bathtub curve).

Open questions (for Greg / Lonny)

  1. Is the ~40% un-repaired rate technical or speed-limited? Largely answered 2026-06-26: capacity + speed, not dead silicon (chip rarely the problem; ~90% known-fix). Remaining: the actual $\kappa_c$ split between repair and revenue on the reman line, and how it’s decided.
  2. Real end-to-end $\tau$? Answered: ~60d end-to-end, 27–40d warehouse-to-stock, +1–2 wk untracked approval. Remaining: the true mean/distribution and the customer clock $\tau_c$ they don’t yet measure.
  3. Real $C_r$ breakdown (freight / diagnosis / consignment parts) — and the per-category cost (reman vs repair vs refurb) that sizes $\mu$. Still open — the single biggest guess left in the model.
  4. On a long dwell, does the CM re-diagnose from scratch? Sizes $w$ — and is the refurb tail where most of that re-work lives?
  5. Real $N$ and loaded $P$? The two biggest multipliers, both guesses.
  6. Pool size $B$ and the from-stock service level — the parameters that set how decoupled $\tau_c$ actually is from $\tau$.

Sources: Greg en-route call, Jun 26; Alex Zhu, May 27; Lonny Orona, May 12; process walk-through, Jun 22. Companion: 2026-06-24-reverse-logistics-sim-math. Prototype: lab/reverse-logistics-sim/.