NVIDIA reverse-logistics repair flow — clean-room operational + economic model

Built from primary interview transcripts only. No prior TBD model, formula, unit price, or the two-clock structure was opened or reused; every equation, split, and price below is derived here or labelled [Speculation]. Purpose: quantify how revenue, cost, and bottom line move as repair turnaround time τ falls, and expose which lever actually moves the money.

Revision 2026-07-04 (post-comparison with 2026-06-20-nvidia-rma-process-map): added a customer-side collection layer (two-stock SD) after the process map made clear the $2B+ “bone pile” is non-collected material at customers, not a repair-capacity backlog. This corrects the earlier claim that the ρ>1 repair queue “regenerated the $2B bone pile” — it reproduced a same-magnitude but different pool. The collection layer, not repair, is the largest trapped-capital pool and the one TBD’s visibility portal actually moves.


Headline — the finding is which lever, not one number

Turnaround time is two different economic levers wearing one name, and they are worth ~20× different amounts:

Lever (what “faster τ” means)What it movesΔ per year, base caseRange (cons → opt)
Logistics/approval speed (cut dwell, capacity unchanged)Working capital in the pipeline; one-time sellable-inventory release+$59M/yr carrying saved + $236M one-time units freed to sell$8M → $452M/yr; $54M → $1.13B one-time
Repair throughput/coverage (faster line = more of the 60/40 gets repaired)Recurring cannibalized new-unit sales (sold-out market)+$1.3B/yr in units freed to sell$0.12B → $7.5B/yr
Collection speed + return rate (the portal: escalation triggers, ASN, “what you owe us” visibility, direct pickup)Working capital stranded at customers (the bone pile) + recurring non-return write-offs+$97M/yr carrying + $357M one-time freed + $364M/yr fewer write-offssee collection sweep below

Scenario = halve turnaround (≈60→30 days full clock / ≈33→16 days on the leg NVIDIA tracks); collection row = τ_c 60→30 d and non-return φ 10%→3%. Read this as: the slider the founders want to push is real, but the big prize is behind the capacity door, not the logistics door — and the door TBD’s own product opens is the collection door. All three are coupled by one identity (Little’s Law: trapped units = arrival rate × loop time), which is why the demo should show them together.

Trapped capital today (base case — three distinct pools, not one):

  • Repair pipeline (NVIDIA-owned units in repair): ~$471M (cons $106M → opt $2.26B).
  • Collection pool (defective units stranded at customer sites awaiting return — this is the “bone pile”): ~$774M — the largest working-capital pool, and the one NVIDIA has least visibility into.
  • Accumulating write-offs (units that never come back): growing ~$520M/yr.
  • The two-stock SD reaches ~$2.16B total trapped by year 2, reconciling to the observed $2B+ bone pile [Interview: In-person debrief, 2026-06-25] through the correct mechanism — slow/failed customer-side collection, not a repair-capacity backlog. (ρ>1 in repair is real and piles up its own WIP, but that is a separate, smaller pool — see §method 3.)

No verdict here — the load-bearing assumptions (unit value, return volume, and the repair-coverage shift) are [Speculation]/M-confidence and are what decide whether the base case or a number 10× smaller/larger is real. See “What would make this wrong.”


How we built it — three methods, one answer

1. Little’s Law (analytic sanity gate)

Backbone identity L = λ · W: units trapped in the pipeline = return rate × dwell time; trapped capital = trapped units × unit value.

  • Utilization ρ = λ / (c·μ) is the load-bearing check. Alex Zhu’s “100 returned, only ~60 can be repaired” [Interview: Alex Zhu, 2026-05-27] means repair throughput covers ~60% of returns → ρ ≈ 1/0.6 ≈ 1.7 > 1. A site over capacity has an unbounded queue — dwell time grows without limit. So the 27–40 day figure NVIDIA “tracks” [Interview: Greg, 2026-06-26] is survivorship: the dwell of units that get through, not of the system. The ones that don’t get through become the bone pile. Any model reporting a finite system-wide dwell today is broken; ours reports two regimes instead.
  • Stable regime (capacity brought to sufficiency, ρ<1): L = λ·τ is exact. At base λ = 400 units/wk and τ = 33 d (4.7 wk): L = 1,886 units, trapped capital $471M.

2. Text2Sim DES (per-node queueing, simulate_des)

Pooled M/M/c model of the compute-repair network (Dallas + Houston + Guadalajara pooled — see boundary note; no per-site rates exist in the transcripts to justify splitting). Service time = full node dwell (transit + customs + queue-at-CM + hands-on), normal(33, 9) days. Run at a scaled arrival rate (dwell and utilization are scale-invariant; absolute WIP scales linearly with λ via Little’s Law and is applied analytically). 12 replications, 95% CIs.

  • Sufficient capacity (ρ≈0.9): queue wait 0.0–0.8 days (mean ≈ 0.1), throughput = arrivals (~100%). → dwell ≈ service ≈ 33 d, queue negligible. Finding: once capacity is adequate, turnaround is a logistics problem baked into service time, not a queue-for-capacity problem. Scaled WIP = 1,881 (matches analytic 1,886).
  • Insufficient capacity (ρ≈1.8, ~today): only 55–56% of arrivals ever served — which equals capacity/offered-load = 1/ρ and independently reproduces Alex’s ~60% repair rate (not tuned to it). Queue wait grows 103 → 193 days as the run doubles 400→800 d — the unbounded-dwell signature of ρ>1. The unserved 44% is the bone pile forming.

3. Text2Sim SD (aggregate trapped inventory, simulate_sd)

Stock = units in pipeline; inflow = returns (400/wk); outflow = pipeline/τ (stable) or a capacity cap of 240/wk = 0.6·λ (constrained).

  • Stable: Pipeline converges to exactly 1,880 units, trapped capital $470M — matches Little’s Law to the dollar, equilibrium in ~20–25 weeks.
  • Constrained repair (ρ>1): the repair-side WIP grows linearly at (λ − capacity) = 160/wk → $2.08B at week 52. Correction: an earlier pass read this as “regenerating the $2B bone pile.” It does not — this is a repair backlog at the CM, a different (and smaller) pool than the bone pile, which is non-collected material at customers. The magnitude match was a coincidence. The bone pile is modeled properly in method 3b.

3b. Text2Sim SD — collection + repair + scrap loop (added 2026-07-04, scrap branch 2026-07-04b)

The process map is explicit that the $2B+ bone pile = non-collected / unreturned units at customers [Interview: In-person debrief, 2026-06-25], driven by “let it fail” deferral, BU-approval friction, ODM touch-inflation, and no day-31 escalation [Interview: Lonny, 2026-05-12 & 2026-06-17]. Under advance-replacement, NVIDIA ships a good unit before the defective one comes back, so it floats a spare for the whole collection lag — in a sold-out market, another unit not sold. Modeled as stocks feeding the repair loop, now including a scrap branch so returned-but-unrepairable units divert out of the pipeline (matching the s term in Channel B rather than assuming everything collected is fixable):

AtCustomer' = λ − Collected − Stranded      Collected  = AtCustomer / τ_c
Pipeline'   = RepairIn − Completions         Stranded   = AtCustomer · strand   → BonePile   (never returns; fraction φ)
BonePile'   = Stranded                       ScrapFlow  = Collected · s          → Scrapped   (returns, technically dead)
Scrapped'   = ScrapFlow                      RepairIn   = Collected − ScrapFlow ;  Completions = Pipeline / τ_r

Two hazards, deliberately distinct symbols: strand (per-week non-return hazard, tuned to give non-return fraction φ) and s (technical scrap fraction of returned units). Closed forms (verified against the SD run): collection pool = λ·τ_c·(1−φ)·V; write-off leak = λ·φ·V/yr; technical-scrap loss = λ·(1−φ)·s·V/yr — three separate levers.

  • Baseline (τ_c = 60 d, φ = 10%, s = 5%): AtCustomer settles at 3,096 units = $774M (the bone pile — larger than the now-$402M repair pipeline, which drops to 1,608 units as 5% diverts to scrap); total trapped crosses ~$2.16B at year 2. The permanent new-unit draw = 58 units/wk = non-return (40) + technical scrap (18)$753M/yr of lost sales before any capacity shortfall, split $520M/yr (Channel D) + $234M/yr (Channel B technical floor, untouchable by speed, capacity, or the portal).
  • Portal-improved (τ_c = 30 d, φ = 3%): collection pool → 1,668 units ($417M, releasing $357M); write-off leak → 12/wk ($364M/yr recovered); total trapped at year 2 → $1.17B~$1B less trapped than baseline. Note the technical-scrap $234M/yr does not move — it’s the floor the portal cannot reach.

3c. Segmented three-stream SD — reman / N-1 / refurb (added 2026-07-04c)

Greg’s three repair categories are not economically interchangeable: they differ on whether repair capacity is shared with revenue-generating new production. Modeled as three parallel repair pipelines (shares of base λ = 400/wk: reman 60%, N-1 30%, refurb 10% — the reman:N-1 split is [Speculation]; Greg gave only “~90% in the first two combined”), each with its own capacity type, turnaround, and unit value:

StreamCapacity typeVSteady-state SD result
Reman (current-gen, shared live production line)revenue-subordinate — the residual after new production takes its cut (here 130/wk vs 216/wk demand)$250Kdiverges — pipe grows 86/wk → $2.24B trapped by yr 2, bleeds $1.12B/yr into new-draw
N-1 (last-gen, dedicated repair line)dedicated, sized to demand$150Kstable at ~508 units = $76M (= λ·τ)
Refurb (one-off, separate shop)shop, sized$100Kstable at ~288 units = $29M

Two things the single-SKU model masked:

  1. The reman stream is the problem. It alone diverges and carries almost all the trapped capital and new-draw — because its capacity is borrowed from, and preempted by, new production (“revenue wins 9 out of 10,” [Interview: Greg, 2026-06-26]). N-1 and refurb, on independent capacity, stay bounded.
  2. Reman new-draw is the least recoverable dollar in the whole model. “Fixing” it means handing shared line-time back to repair — but that line-time otherwise sells at V, so reallocation just moves the loss from the warranty column to the new-sales column. Only net-new shared-line capacity (capex) truly recovers it. So the ~$1.3B Channel B “whale” is mostly reman, and mostly capex-gated — unreachable by scheduling, speed, or the portal.

And a cost the flat Channel C ($12K/unit) missed: each reman repair performed (130/wk) burns line-time that could have built a sellable unit — opportunity cost ≈ (t_repair/t_build)·V. At a [Speculation] t_repair/t_build = 0.3 → ~$75K/unit → ~$507M/yr of foregone new sales just to do the reman repairs, vs. the ~$81M/yr the flat opex assumed. Reman repair is not cheap; N-1 and refurb repair (dedicated/separate capacity that couldn’t make revenue anyway) genuinely are. The general rule the segmentation reveals: repair is cheap exactly when its capacity isn’t fungible with new sales.

Configuration question (surface, don’t conclude): moving reman onto dedicated lines (like N-1) would make reman repair genuinely cheap — but Greg keeps it on the build line deliberately (known-fix + fresh line knowledge). That tradeoff, and the shared-line allocation decision Greg said “shouldn’t be made by the CM on their own,” is a scheduling/coordination surface — plausibly software-addressable, and reman volume is NVIDIA-triggered (recall-like), so schedulable into line troughs. A question for Greg, not a claim.

4. Convergence gate — PASSED

QuantityLittle’s LawSDDES
Stable WIP (base)1,8861,8801,881
Stable trapped capital$471M$470M$470M
ρ verdict today>1 (≈1.7)>1 (linear growth)>1 (55% served, dwell↑)
Analytic = SD exactly; DES within ~1%; all three agree the system is over capacity today.
Within the skill’s “~2× and same ρ verdict” tolerance — the model is sound.

The economic model (my own structure)

Warranty repairs are free to the customer [Interview: Alex Zhu, 2026-05-27], and NVIDIA GPUs are supply-allocated / sold out — “those GPUs everybody’s fighting for, there’s never enough” [Interview: Greg, 2026-06-26]; imec’s Cedric: “NVIDIA’s narrow incentive is to replace chips as fast as possible… seller-side loses revenue” [Interview: imec/Cedric, 2026-06]. So the cost of the repair operation is not the repair — it’s the new units it consumes, each of which is a lost sale at price V. Three channels, each tagged with the lever that drives it:

Channel A — Working capital carried in the pipeline (LOGISTICS-τ lever). Trapped units = λ·τ. Carrying cost C_carry(τ) = k · λ · τ · V, falls linearly with τ. Sensitivity: dC_carry/dτ = k·λ·V = $3.6M/yr per day of turnaround (base). Cutting τ also releases λ·Δτ units one-time — in a sold-out market those sell at V: ΔRevenue_onetime = V·λ·Δτ = $14.3M per day of τ removed (base).

Channel B — Cannibalized new-unit sales / repair-vs-new split (CAPACITY lever + a fixed floor). Returns are served from repaired stock or from new stock pulled off the sold-out line [Interview: Alex Zhu, 2026-05-27; Greg, 2026-06-26]. The new-unit consumption rate has two distinct causes, and only one is a lever: R_lost = V·λ·(1−φ)·(s + shortfall), where

  • s = technical scrap — units that come back but physically cannot be repaired (cracked die, CoWoS/HBM bond damage). A hard floor: not moved by speed, capacity, or the portal — only by product/failure-mode engineering or pushing die failures back to TSMC (which the process map notes happens). Base ≈ $234M/yr (s = 5%).
  • shortfall = max(0, 1 − 1/ρ) — units that could be repaired but capacity can’t keep up. Base ≈ $1.3B/yr and moved by adding repair throughput.

Alex’s observed “60% repaired / 40% new” is the sum of these two — the transcripts don’t separate them, so the split (how much of the 40% is dead vs. merely capacity-starved) is a modeling structure, not a measured fact. Lifting capacity coverage (e.g. 60%→85%) is the whale ($1.3B/yr); the technical floor sits underneath it and no operational lever reaches it.

Channel C — Repair opex + expedite (mostly flat / countervailing). C_repair = c_r · λ · coverage ≈ $212M/yr (base, $12K/unit [Speculation]) — cheap next to a lost V-priced sale, which is the whole economic case for repairing rather than replacing. Expedite freight rises as τ falls (Greg: “we’d fly them by private jet” [Interview: In-person debrief, 2026-06-25]) — a countervailing cost, but the interviewees’ stated willingness to pay for speed implies they believe expedite < the A+B savings. Left as [Speculation]; flagged, not sized.

Channel D — Customer-side collection & non-return (the PORTAL lever). Advance-replacement floats a spare for the entire collection lag τ_c, so the collection pool is trapped capital exactly like the repair pipeline: C_carry_collection = k · λ · τ_c·(1−φ) · V (and a one-time sellable release of V·λ·Δ[τ_c·(1−φ)] when τ_c or φ falls). Separately, the fraction φ that never returns must be replaced from new, sold-out stock → recurring lost sales R_strand = V·λ·φ per year — a Channel-B-type leak that lives entirely on the customer side. Crucially, τ_c and φ are moved not by capex but by coordination/visibility (escalation triggers, ASN, “what you owe us,” direct pickup) — i.e. the dashboard Lonny and Greg sketched. This is the leg TBD’s product actually touches. Base case: τ_c 60→30 d frees $357M working capital ($97M/yr carrying), φ 10%→3% recovers $364M/yr of lost sales, and total trapped capital at year 2 falls ~$1B (see method 3b).

Bottom line(τ) = −[C_carry(τ_r) + C_carry_collection(τ_c,φ) + C_repair + C_expedite] − R_lost(coverage) − R_strand(φ) + one-time releases. Faster repair τ_r helps through A always and B only if speed brings capacity; faster/surer collection (τ_c↓, φ↓) helps through D — and D is the pool NVIDIA can’t currently see and the founders’ software can.

τ-sweep — Δ vs a halving of turnaround (≈60→30 d full / 33→16 d tracked leg)

MetricChannel / leverConservativeBaseOptimistic
Working capital released (one-time, sellable)A / logistics$54M$236M$1.13B
Carrying cost saved (recurring)A / logistics$8M/yr$59M/yr$452M/yr
Cannibalized-sales recovered (recurring)B / capacity portion$117M/yr$1.30B/yr$7.49B/yr
Technical-scrap loss (recurring floor — no lever recovers it)B / technical$33M/yr$234M/yr$1.70B/yr
Collection capital released (one-time)D / portal$130M$357M$1.55B
Non-return write-offs recovered (recurring)D / portal$63M/yr$364M/yr$2.18B/yr
Repair-pipeline capital today (level)$106M$471M$2.26B
Collection-pool capital today (the bone pile, level)$131M$774M$4.13B

Conservative: λ=150/wk, V=$150K, k=15%, coverage 60→70%, s=3%, τ_c 45→30 d, φ 5→2%. Base: λ=400/wk, V=$250K, k=25%, coverage 60→85%, s=5%, τ_c 60→30 d, φ 10→3%. Optimistic: λ=1,200/wk, V=$400K, k=40%, coverage 60→90%, s=8%, τ_c 90→30 d, φ 15→3%.

Sensitivity ranking (tornado — what moves the answer most)

  1. Repair-coverage shift (capacity) — $0.12–7.5B/yr. Dominates by 10–20×. Validate with Lonny/Alex.
  2. Non-return fraction φ / collection (Channel D, the portal lever) — $63M–2.2B/yr recurring
    • $0.1–1.6B one-time. Second-biggest recurring lever, and the one TBD’s product moves. Validate with Lonny (ARMA idle) + back-out φ from the $2B.
  3. Unit value V [Speculation] — linear on every channel; 2.7× swing. Validate with Greg/Lonny.
  4. Return volume λ — linear on every channel; 8× swing. Validate with Greg (weekly metrics).
  5. Collection lag τ_c — sets the bone-pile stock ($0.1–4.1B level); portal compresses it.
  6. Carrying rate k, scrap s, per-repair cost c_r — second-order.

Input ledger

AssumptionValue (cons → base → opt)SourceConf.Impact if wrong
Flow unit1 compute baseboard/tray ≈ 8 GPUs + 1 CPU, blended SKU[Interview: In-person debrief, 2026-06-25] (“1 CPU + 8 GPUs”)MHigh (defines V)
Return arrival rate λ150 → 400 → 1,200 units/wk[Interview: Lonny, 2026-05-12] “hundreds… becoming thousands”; [Greg, 2026-06-26] reman batches in thousands, “6,000 H100s”LHigh — linear on all channels
Repair turnaround τ (leg NVIDIA tracks: warehouse-receipt → back-in-stock)40 → 33 → 21 d[Interview: Greg, 2026-06-26] “27–40 days… best case 30, probably 60 now”MHigh (drives Channel A)
Pre-warehouse approval delay (untracked)7 → 10 → 14 d[Interview: Greg, 2026-06-26]; [Lonny, 2026-06-17] “day 31” SLAMMed (customer clock, not NVIDIA capital)
Repair coverage today (repair vs replace-from-new) — blends technical scrap + capacity gap~60% repaired / 40% new[Interview: Alex Zhu, 2026-05-27] “100 back, ~60 repaired”MHigh — Channel B
— of which capacity shortfall (movable by throughput)assumed to dominate the 40%[Speculation]LHigh — the whale
— of which technical scrap (fixed floor, see s)assumed minority of the 40%[Speculation]; “chip rarely the problem” [Lonny]LMed — the floor
Capacity coverage after fix70% → 85% → 90%[Speculation] (target state)LHigh — the whale
Utilization ρ today≈1.7 (>1)[Synthesis] from 60% coverageMHigh (regime selector)
Collection lag τ_c (customer hold → return)45 → 60 → 90 d[Interview: Lonny, 2026-05-12 & 2026-06-17] “weeks to months” ARMA idleLHigh — largest capital pool
Non-return fraction φ (permanent bone pile)2% → 10% → 15%[Speculation] — back-out from $2B non-collected [In-person debrief, 2026-06-25]LHigh — Channel D leak
# compute repair sites c3 (Dallas, Houston, Guadalajara; Foxconn/Wistron)[Interview: Greg, 2026-06-26]HLow (pooled)
Category mix (reman + repair vs refurb)~90% / ~10%[Interview: Greg, 2026-06-26]HMed
— reman : N-1 split within the 90%60% : 30% of total[Speculation] — Greg gave only the 90% combinedLHigh — opposite cannibalization behavior
Reman capacity typeshared w/ new production, revenue-subordinate[Interview: Greg, 2026-06-26] “same line… revenue takes precedent 9/10”HHigh — makes reman new-draw capex-gated
Per-stream unit value V (reman / N-1 / refurb)$250K / $150K / $100K[Speculation] — current-gen / discounted last-gen / salvageLMed
Reman line-time ratio t_repair / t_build0.3[Speculation]LMed — sets reman repair opportunity cost
RMA rejection rate<1%[Interview: Greg, 2026-06-26]HLow
Technical scrap / irreparable fraction s (of returned units)3% → 5% → 8%[Speculation]; die failures “go back to TSMC” & CoWoS/HBM bond damage [In-person debrief, 2026-06-25; process-map research]LMed — fixed floor of Channel B, ~$234M/yr base
Unit value V (sellable compute unit)$150K → $250K → $400K[Speculation] — 8× H100/H200 GPU ASP; “DGX costs millions” [Alex, 2026-05-27]LVery high — linear on everything
NVIDIA margin m (lost-sale profit)0.75[Public: NVIDIA data-center gross margin ~75%]MMed
Working-capital carrying rate k15% → 25% → 40%[Speculation] — cost of capital + fast GPU depreciation (Pluto sells H200 depreciation cover, [Ronit, 2026-05-22])LMed (Channel A)
Per-repair cost c_r$5K → $12K → $25K[Speculation] — component + skilled labor + cross-border logisticsLLow (cheap vs V)

Load-bearing (flagged): (1) repair-coverage shift, (2) non-return fraction φ / collection lag τ_c, (3) unit value V, (4) return volume λ. The headline is a [Speculation]-heavy estimate on exactly these — stated loudly per the ledger rule. φ and τ_c are the newest and least grounded (pure [Speculation]), yet they govern the largest capital pool; everything else is second-order.


Boundaries (confirmed / revised from the brief)

  • In scope: compute units (tray / baseboard / SXM, blended as one SKU v1) through the three compute repair nodes. Out: networking/switches (Vietnam / Israel / India), die failures (return to TSMC) [Interview: In-person debrief, 2026-06-25].
  • Two clocks, explicit: (a) NVIDIA-tracked leg warehouse-receipt → back-in-good-stock (27–40 d) — this is the leg that ties up NVIDIA’s working capital, so Channel A uses it; (b) customer-experienced clock (~60 d) = tracked leg + the untracked pre-warehouse approval delay (1–2 wk) + front/back transit. A third clock, τ_c, governs the bone pile ($2B+ non-collected) — the time a defective unit sits at the customer before it ships back (or never does). This is the largest trapped-capital pool and is now modeled explicitly (method 3b); it is driven by collection/coordination, not repair capacity — correcting an earlier reading of this model.
  • Pooled vs per-site: modeled as one pooled M/M/c network. The transcripts give no per-site capacity or routing split, so a 3-way split would add structure without information. Site-to-site transit variance (Dallas↔Guadalajara customs, the named bottleneck) is folded into the service-time spread. Flagged as a revision from the skill’s “DES per-site” default.

What would make this wrong (the 2–3 tests that matter)

  1. The repair-coverage shift is assumed, not observed. Channel B — the $1.3B/yr whale — rests on repair coverage rising from ~60% to ~85% once capacity/speed improves. If the real ceiling is 70%, the number roughly halves; if faster logistics doesn’t add throughput, Channel B is ~$0 and only the ~$59M/yr working-capital prize remains. Test: ask Lonny/Alex what fraction of returns can be repaired at full capacity, and whether the 60% is a capacity limit or a speed limit.
  2. Unit value V is a pure guess spanning $150K–$400K (2.7×), linear on every number here. Test: get Greg/Lonny to confirm the blended value of a returned compute unit (tray vs baseboard vs bare SXM — the mix matters).
  3. Return volume λ is L-confidence (150–1,200/wk, 8× range) and scales everything. Test: Greg sees weekly throughput/cycle-time metrics — ask for units-received/week by category. This single number tightens the whole model.
  4. Collection lag τ_c and non-return fraction φ are pure [Speculation] yet govern the largest pool (the bone pile) and Channel D — the value of TBD’s own product. If φ is 2% not 10%, the write-off leak is ~$70M/yr not $364M/yr; if the bone pile is mostly customer-owned (not core-return), it isn’t NVIDIA’s capital at all and Channel D collapses. Test: ask Lonny for mean ARMA idle time and the return-compliance rate, and confirm the defective unit is NVIDIA’s asset (core-return) — this decides whether the portal’s headline number is real.

Surprises / contradictions (per RDI)

  • The tracked turnaround time is survivorship, and it hides the real problem. The 27–40 day figure only describes units that clear a queue that (at ρ≈1.7) most units don’t clear on schedule. The honest system metric is bimodal: fast for the served ~60%, unbounded for the rest. NVIDIA is measuring the wrong clock — and the untracked approval leg means even the customer-experienced number is not being captured [Greg, 2026-06-26].
  • “Speed beats cost” may point at the smaller prize. The interviewees’ loudest signal is turnaround time [In-person debrief, 2026-06-25]. But the model says pure speed moves ~$59M/yr (working capital); the ~$1.3B/yr is behind capacity/coverage. Faster cycle time helps free capacity (less WIP per unit of throughput), so the instinct isn’t wrong — but a turnaround-time slider that ignores the repair-vs-new split undersells the economics by ~20×. This is a flag for the founders, not a conclusion.
  • The bone pile lives on the customer side, and it’s the biggest pool — a correction we owe ourselves. An earlier pass matched the $2B against a repair-capacity backlog; that was a magnitude coincidence between two different stocks. Modeled properly (method 3b), the $2B is non-collected material at customers — collection pool ($774M base) + accumulating write-offs — reaching ~$2.16B at year 2. It exceeds the repair pipeline, and unlike repair it’s fixed by coordination, not capex — which is why it’s the pool the founders’ portal can actually move.
  • The founders’ product and the biggest movable number are the same thing. Repair coverage (Channel B) is larger but capacity/capex-bound; the collection layer (Channel D, ~$0.7B stock + $364M/yr leak base) is software-bound. The demo’s turnaround slider should have a second slider: collection/return-rate.
  • “Repair is cheap” is only true off the revenue line. Segmenting Greg’s three categories showed the reman stream (repaired on the live production line) is the one that diverges and carries almost all the trapped capital + new-draw — and its dollars are the least recoverable in the model, because handing the shared line back to repair just moves the loss to new-sales. The cheap, genuinely-fixable repair is N-1 and refurb, which run on dedicated/separate capacity. So the biggest channel (B) is not one lever but two very different ones bolted together: a small recoverable piece (N-1/refurb capacity) and a large capex-gated piece (reman line).
  • The DES reproduced Alex’s 60% with no tuning — served fraction = 1/ρ fell out of the capacity constraint. Two independent interviewees (Alex’s 60%, Greg’s dwell) constrain the same ρ.
  • No contradiction found between sources on the flow structure — Greg (inside), Alex (inside, earlier), and Jason (outside observer, [2026-05-22]) describe the same chain. That agreement is itself worth flagging: either the picture is solid, or all three share a blind spot (e.g. the grey-market Shenzhen channel Jason raised, which none of the NVIDIA sources mention and which our model excludes).

Open questions → named contacts

  • Of the ~40% filled from new, how much is technically dead vs. merely capacity-starved?Lonny Orona (customer service / entitlement) and Alex Zhu (reverse-logistics ops, gave the 60%). This decides Channel B: the capacity-starved share is the ~$1.3B/yr whale a repair-capacity fix recovers; the technically-dead share (s) is a hard floor no operational lever — speed, capacity, or the portal — can touch (base ~$234M/yr). Also ask: do die failures get recovered under NVIDIA’s own warranty from TSMC, or eaten by NVIDIA?
  • Blended unit value V and the tray/baseboard/SXM mixGreg / Lonny. Linear on everything.
  • Weekly return volume λ by categoryGreg (“we can get into the metrics… 27–40 days on the metrics I see weekly”). Tightens the whole model.
  • Does NVIDIA carry the pipeline as its own working capital, and at what cost of capital / depreciation rate (k)?Greg / NVIDIA finance (the finance approval step in the RMA flow). Sets Channel A.
  • Mean collection lag τ_c (ARMA idle) and the return-compliance / non-return rate φ; and is the defective unit NVIDIA’s asset (core-return) or the customer’s?Lonny (owns ARMA / customer service) + Alex (reverse-logistics ops). Sets Channel D — the largest pool and the value of TBD’s portal. If it’s customer-owned, Channel D collapses.
  • The reman : N-1 split within the “90%”, and the reman line-time ratio (t_repair / t_build) — plus, could reman move to dedicated lines, or must it stay on the build line?Greg (owns the CM relationship + line allocation). Decides how much of the ~$1.3B Channel B whale is reman (capex-gated, unrecoverable by ops) vs N-1/refurb (genuinely recoverable), and the true cost of every reman repair. Also probes the shared-line scheduling surface Greg flagged as a “joint decision.”
  • Expedite cost curve — what does air-vs-ground actually cost per unit, and is the “fly it by private jet” instinct backed by a number? → Greg. The one countervailing cost.

Sources

Primary interviews (clean-room inputs): Greg DeLoccio, 2026-06-26 (fullest process walk-through); NVIDIA in-person, Greg+Lonny, 2026-06-25; Greg+Lonny, 2026-06-17; Alex Zhu, 2026-05-27; Lonny Orona, 2026-05-12; Jason (corroborating), 2026-05-22. Corroborating: Cedric (seller-side loses revenue on fast replacement). Methods: Little’s Law analytic pass; Text2Sim DES (12 reps, 95% CI) + SD (PySD) for the repair loop, converged; two-stock SD collection loop added 2026-07-04 (closed forms verified against the run). Post-hoc comparison against 2026-06-20-nvidia-rma-process-map. Clean-room: the original build opened no prior TBD economic/simulation model; all structure, equations, prices, and the two-clock treatment were derived independently. The collection layer was added afterward, once the clean-room artifact was saved, on comparison with the process map — and is flagged as such.