The NVIDIA repair-flow model, explained from scratch

This is a plain-language companion to the technical brief the clean-room operational model. It contains all the same reasoning, equations, and numbers — but it stops to explain each idea, defines the jargon the first time it appears, and leans on a single running analogy so you never have to hold an abstract symbol in your head. If you’ve never built a model like this, start here; if you know the material, the technical brief is faster.

This document is self-contained: Parts 1–12 build the intuition, and Parts 13–16 at the end collect every assumption, input value, formula, and result in reference tables, so you never have to leave this page to find a number.

Nothing here is a conclusion. A model surfaces what the evidence implies and which assumptions the answer hangs on — the founders decide what to do about it.


Part 0 — What is this kind of model, and why build one?

We had a business question we couldn’t answer by asking anyone directly:

If NVIDIA repairs broken GPUs faster, how much money does it make or save — and which kind of “faster” actually matters?

Nobody at NVIDIA can just tell you, because the answer depends on a dozen interacting moving parts: how many units break, how long repairs take, how many units never come back, what a GPU is worth, and so on. When a question has that many entangled parts, guessing is dangerous and “just ask an expert” doesn’t work — the expert only sees their slice.

So instead we build a model: a small, explicit, honest map of how units move through the process over time, put a number on each part, and let arithmetic carry the consequences. That’s all a model is — organized bookkeeping for stuff flowing through a pipeline, plus the discipline to be honest about which numbers we actually know versus which we’re guessing.

That honesty is the whole game. Every input in this model carries a tag:

  • [Interview: Name, Date] — someone who lives this problem told us.
  • [Public: Source] — from a filing or public document.
  • [Synthesis] — inferred by combining several sources.
  • [Speculation] — an honest guess, not yet grounded.

The reason for the tags: a precise-sounding number built on a guess is more dangerous than an admitted range, because it launders a guess into a fact. When you see “$1.3 billion” below, the tags let you check whether that rests on evidence or on a placeholder. And the real output of the exercise isn’t the headline figure — it’s a defensible range plus a short list of which unknowns most move the answer, and who could pin them down.


Part 1 — The thing we’re modeling, in plain English

When a GPU fails at a customer like Meta or Microsoft, NVIDIA doesn’t just mail back a fixed unit. A whole loop turns:

  1. The customer notices a failure and files a return request (an “RMA” — Return Merchandise Authorization, the industry term for “I’d like to send this broken thing back”).
  2. NVIDIA approves it (a multi-step paperwork gauntlet).
  3. NVIDIA ships the customer a working replacement right away so their data center keeps running — this is called advance replacement.
  4. The customer eventually ships the broken unit back.
  5. It gets repaired.
  6. The repaired unit goes onto a shelf of good spares (the “service pool”) to fill the next customer’s claim.

The analogy: a propane-tank exchange

Picture the propane-tank cage outside a hardware store. This maps onto NVIDIA’s operation almost perfectly, so hold onto it for the rest of the document:

  • The store keeps a shelf of full tanks ready to hand out → NVIDIA’s service pool of repaired GPUs.
  • Your tank fails; the clerk hands you a full one off the shelf and takes your empty → advance replacement.
  • Your empty goes to a refill shop, gets reconditioned, and comes back to the shelf → the repair loop.
  • The catch that makes this interesting: the store has a fixed number of tanks, and every one is spoken for. A tank sitting in the back being reconditioned is a tank not on the shelf earning money.

That last point is the crux, and it’s literally true for NVIDIA: GPUs are sold out. Greg told us they’re “the ones everybody’s fighting for — there’s never enough.” So every chip tied up anywhere in the repair loop is a chip NVIDIA could have sold instead. That single fact — a trapped unit is a lost sale — is the engine underneath every dollar figure in this model.

What we’re counting

Because trapped units are the cost, the whole model reduces to two questions: (a) how many units are stuck, and where? and (b) what does it cost to have them stuck there? Everything below is a careful answer to those two questions.


Part 2 — The one equation you need: Little’s Law

There is exactly one piece of math at the heart of this, and it’s genuinely intuitive.

Picture a coffee shop. 20 customers walk in per hour, and each spends about 15 minutes (¼ hour) inside. How many people are in the shop at any given moment? Twenty per hour times a quarter-hour each = 5 people. That’s the entire idea, and it has a name — Little’s Law:

Amount stuck in a process = how fast things arrive × how long each one stays.

In shorthand: L = λ × W, where L is the amount stuck (called “work-in-process” or WIP), λ (“lambda”) is the arrival rate, and W is the average time each item spends inside (called the dwell time).

Don’t let the Greek letter or “WIP” intimidate — λ is just a nickname for “arrivals per week,” and WIP is just “how many are currently inside.”

Worked example, repair version: suppose 400 broken units arrive per week and each spends about 5 weeks being repaired. Then at any moment, roughly:

L = 400 per week × 5 weeks = 2,000 units sitting in repair.

If each unit is worth $250,000, that’s $500 million of capital frozen in the repair pipe at all times. And notice what the equation instantly tells us: if you make repairs faster (shrink W), fewer units are stuck, and money is freed. Speeding up turnaround is, mathematically, the same thing as un-freezing capital. That is the lever the founders are excited about — and Little’s Law is what lets us price it.

We use Little’s Law two ways: as a quick pencil-and-paper estimate (the “analytic” pass), and as the sanity check that every fancier simulation has to agree with. If a simulation disagrees with Little’s Law by a lot, the simulation is wrong until proven otherwise.


Part 3 — Utilization, and why “over capacity” quietly breaks everything

Before trusting any dwell-time number, you have to ask one question that turns out to be decisive: is the repair operation keeping up with the incoming flow, or falling behind?

The measure for this is utilization, written ρ (“rho”). It’s just:

ρ = (work arriving) ÷ (work the shop can handle) = λ ÷ (number of repair stations × how fast each works).

Think of a supermarket checkout. If shoppers arrive slightly slower than the cashiers can ring them up, ρ is a bit below 1, lines stay short, and everything’s fine. But the instant shoppers arrive faster than the cashiers can handle — ρ above 1 — the line doesn’t just get long, it grows without limit. It never stabilizes. Every minute, the backlog is bigger than the minute before.

This matters enormously here, and it produces the model’s first surprising result. Alex Zhu told us that of every 100 units that come back, only about 60 get repaired; the other 40 are filled from new stock. Read as a capacity statement, that means repair throughput covers only ~60% of the incoming flow — so:

ρ ≈ 100 ÷ 60 ≈ 1.7, which is well above 1.

The repair system is over capacity. By the checkout-line logic, its backlog grows without bound. So the comfortable-sounding “27–40 days” that NVIDIA tracks for a repair cannot be the whole story — it’s the dwell time of the units that do get through, the survivors. It quietly ignores the ones piling up behind them. (This is called survivorship: measuring only the winners and mistaking that for the average.) A model that reports a tidy finite dwell time while ρ is above 1 is simply broken, and we say so rather than papering over it. The units that don’t get through are what NVIDIA calls the “bone pile” — and we’ll return to where it really lives.

The takeaway: always check ρ first. If the shop can’t keep up, “how do we speed up each repair?” is the wrong question — the right one is “how do we add capacity so the system is even stable?”


Part 4 — Two ways to compute the same thing (and why we do both)

Good modeling never trusts a single method. We compute every key number two independent ways and insist they agree — the way you’d check a big arithmetic result by redoing it a different way. If the two disagree, one is wrong, and we don’t report a number until we know which.

Method 1 — the analytic pass (Little’s Law by hand). Transparent, fast, no computer. It gives the first-order answer and the ρ check. Its weakness: it smooths over real-world messiness like random variation and queues bunching up.

Method 2 — simulation. We use a tool called Text2Sim, which runs two flavors of simulation:

  • DES — Discrete-Event Simulation. “Discrete events” means the computer tracks individual units one at a time: this unit arrives now, waits in line, gets served, leaves. It’s like watching every shopper move through the store. DES is the right tool when queues and randomness drive the answer — it naturally captures units bunching up and waiting.
  • SD — System Dynamics. Instead of tracking individuals, SD tracks the levels in the tanks over time — how the total pile rises and falls as flows run in and out. Think water levels in connected bathtubs. SD is the right tool for the aggregate picture: how the total trapped inventory grows week over week.

Two more terms you’ll meet:

  • Steady state — the point where a system stops changing and settles at a constant level. In Little’s-Law terms, the pile settles at exactly L = λ × W. Some of our systems reach steady state (the healthy ones); some never do (the over-capacity ones — their piles just keep growing).
  • Confidence interval (CI) — because a simulation has randomness built in, a single run is just one sample, like flipping a coin once. So we run it many times (here, 12 independent “replications”) and report a range: “the answer is X, give or take Y, 95% of the time.” A single run reported as if it were the truth is a rookie mistake; we report the range.

What the two methods said, and whether they agreed

For the healthy (adequate-capacity) case:

QuantityLittle’s Law (by hand)System DynamicsDiscrete-Event Sim
Units stuck in repair1,8861,8801,881
Capital frozen$471M$470M$470M

They agree to within a fraction of a percent — the analytic and SD numbers are essentially identical, and the DES lands right on top of them. When methods that work completely differently converge like that, you can trust the number. This is the convergence gate, and this model passed it.

The DES also taught us two things Little’s Law can’t:

  • In the healthy case, units barely wait in line at all (the queue delay was ~0 days) — meaning when there’s enough capacity, the ~33-day repair time is almost entirely transit and handling (shipping, customs at the Dallas–Guadalajara border, sitting at the contract manufacturer), not waiting-for-a-repair-slot. So turnaround, once capacity is adequate, is a logistics problem.
  • In the over-capacity case (ρ ≈ 1.7), the simulation served only 55–56% of arrivals — which is almost exactly the 1/ρ you’d predict, and it independently reproduces Alex’s “60% get repaired” from scratch, without being told to. And the average wait grew from 103 days to 193 days as we doubled the simulated time — the tell-tale “grows without limit” signature of a system over capacity. Two interviewees’ numbers (Alex’s 60%, Greg’s dwell time) turn out to describe the same overloaded system.

Part 5 — The three places units get stuck (the “stocks”)

A stock is modeling jargon for “a pile of units sitting somewhere.” A flow is the rate units move from one pile to the next (units per week). Our model has three stocks, each a real place in NVIDIA’s process. Back to the propane analogy:

Stock 1 — AtCustomer. Broken units the customer is holding but hasn’t shipped back yet. In the analogy: you got your full tank, but your empty is still rolling around in the garage. At NVIDIA these pile up because the customer’s team would rather “let it fail” than schedule the downtime to pack and ship the dead unit back — a real pain point the interviews surfaced repeatedly. This pile turns out to be the famous “$2 billion bone pile.”

Stock 2 — Pipeline. Units actually in the repair shop being fixed. The tanks in the back being reconditioned.

Stock 3 — BonePile (write-offs). Units that never come back at all — some customers simply never return the empty. These are gone for good, so NVIDIA must replace each with a brand-new unit.

The flows connecting them:

  • Collected = broken units finally shipped back (garage → repair shop).
  • Stranded = the trickle that never returns (garage → gone forever).
  • Completions = repaired units returning to the shelf (repair shop → service pool).

Written as bookkeeping (each line just says “this pile changes by what flows in minus what flows out”):

AtCustomer  changes by:  arrivals − Collected − Stranded
Pipeline    changes by:  RepairIn − Completions
BonePile    changes by:  Stranded

The new variables here are:

  • τ_c (“tau-C”) — the collection lag: how long, on average, a broken unit sits in the customer’s garage before shipping back. Interviews suggest “weeks to months.”
  • φ (“phi”) — the non-return fraction: the share that never comes back at all.

From these, three clean formulas fall out (and the simulation confirmed them exactly):

  • Units stuck at customers = λ × τ_c × (1 − φ) → at base numbers, 3,096 units = $774 million.
  • Yearly write-off leak (units gone forever, each replaced by a new sale lost) = λ × φ × V → ~$520 million per year.

An important subtlety worth pausing on: τ_c and φ are two different levers. Speeding up collection (smaller τ_c) shrinks the pile sitting in garages. Reducing the never-returned fraction (smaller φ) shrinks the permanent leak. They move different numbers, and a good fix addresses both.


Part 6 — The four ways this costs money (the “channels”)

A channel is just one distinct route by which the process drains money. We found four, and each one maps to a specific complaint NVIDIA voiced. For each, I’ll give the plain-English idea, then the formula (demystified), then the base-case number.

Channel A — Money frozen in the repair shop

Units in repair are capital you can’t spend. The cost of frozen capital per year is the carrying cost — interest you’re forgoing, plus the fact that GPUs lose value fast as newer chips arrive (a rate we call k, roughly 25% a year here).

Cost = k × λ × τ_r × V — that’s (yearly cost of frozen money) × (units in repair) × (their value).

Here τ_r (“tau-R”) is the repair leg (~5 weeks). How to shrink it: faster logistics and customs. Worth: about $59M/yr in carrying cost, plus a one-time $236M of units freed to sell when you halve the repair time.

Channel B — Giving away brand-new units for free

When the shelf of repaired spares runs empty, NVIDIA fills a warranty claim from a new unit — one it could have sold. In a sold-out market that’s a straight lost sale. This is Alex’s “60 repaired, 40 from new.” But — and this is a refinement worth understanding — that 40% happens for two completely different reasons, and only one is fixable:

Lost sales = V × λ × (1 − φ) × (s + shortfall)

  • s = technical scrap. Units that come back but physically cannot be repaired — a cracked silicon die, or the newer failure mode where the chip and its memory are bonded together so tightly (the “CoWoS/HBM” packaging) that you can’t separate them to fix one. These get scrapped. This is a hard floor: no amount of speed, capacity, or software recovers it. The only remedies are designing more-repairable hardware, or getting the chipmaker (TSMC) to eat die failures under NVIDIA’s own supplier warranty (which the interviews say happens for pure die failures). Base cost: ~$234M/yr (at s = 5%).
  • shortfall = the capacity gap. Units that could be repaired but there isn’t enough repair throughput to get to them in time. Base cost: ~$1.3 billion/yr — the single biggest number in the model — and it is movable, by adding repair capacity.

The honest caveat: Alex’s observed 60% is the sum of these two causes; the interviews don’t tell us the split. So “how much of the 40% is dead-on-arrival versus merely capacity-starved” is a real open question, and it decides whether the biggest channel has a large recoverable part or a big untouchable floor.

Channel C — The cost of actually doing the repairs

Parts, labor, and shipping for each repair. Relatively small (~$12K per unit, ~$212M/yr) — and that’s the whole economic case for repairing at all: fixing a unit costs a fraction of giving away a new one worth $250K. (Part 7 adds a crucial exception to “repairs are cheap.“)

Channel D — The bone pile (the collection layer)

This is the one the earlier version of the model got wrong, and it’s important. Two sub-costs: (1) money frozen in units sitting in customers’ garages, and (2) units that never come back, each replaced with a free new one.

Frozen: k × λ × τ_c × (1 − φ) × V. Leak: V × λ × φ per year.

Crucially, these are moved not by building factories but by coordination and visibility — return reminders, “day-31” escalation alerts, a shared view of “here’s what you still owe us,” arranging direct pickup. That is exactly the dashboard product Lonny and Greg sketched. Base case: speeding collection from ~60 to ~30 days frees $357M of working capital, cutting the never-returned rate from 10% to 3% recovers $364M/yr, and together they leave ~$1 billion less trapped after two years. This is the pool NVIDIA can’t currently even see, and it’s the one the founders’ software directly moves.

Why one unit shows up in two channels

A subtle but important point: a single $250,000 GPU can be counted as “trapped capital” (a frozen asset — Channels A and D) and as “a lost sale” (Channel B) — because it’s the same physical unit seen two ways. On the balance sheet it’s frozen money; on the income statement it’s a sale that didn’t happen. The model keeps these straight so we never double-count, but they’re two faces of one coin: a unit stuck in the loop.


Part 7 — Why one repaired GPU isn’t like another: reman vs. repair vs. refurbish

Here is the subtlety that most changed the picture, and it comes straight from Greg. Not all repairs are the same kind of thing. There are three, and they differ on one axis that turns out to decide everything: does repairing the unit compete for the same capacity that makes new units to sell?

  • Reman (remanufacturing) — current-generation units with a known defect and a known fix (think of a car recall: “everything we shipped between these dates has this problem”). High volume. The catch: these are repaired on the same live production line that builds new units.
  • Repair (N-1) — last year’s model (“N minus one”). Also high volume, but repaired on a separate, dedicated line that isn’t building anything new.
  • Refurbish — the ~10% oddball tail: one-off units of unknown condition, diagnosed on arrival, sometimes stripped for parts. Done in a separate repair shop. This is where most of the technical scrap (the s from Channel B) lives.

Why the shared-vs-dedicated distinction is decisive: when you repair a reman unit on the production line, you’re using a slot that could have built a brand-new unit to sell. So even a successful reman repair has a hidden price — a forgone sale. And when the line gets busy, Greg told us plainly, “revenue wins 9 times out of 10”: reman repairs get bumped in favor of new production.

We simulated the three streams as three separate pipelines. The result was stark:

StreamCapacity typeWhat happened in the simulation
Reman (shared production line)preempted by new productionDiverges — pile grows without limit to $2.24B frozen by year 2, and bleeds $1.12B/yr into giving away new units
N-1 (dedicated line)independent, sized to demandStable at ~$76M — perfectly healthy
Refurbish (separate shop)independentStable at ~$29M

Two lessons the earlier one-size-fits-all model completely hid:

  1. The reman stream is the problem. It’s the only one that spirals, and it carries nearly all the frozen capital and nearly all the give-away-a-new-unit cost — precisely because its capacity is shared with, and loses to, new sales.

  2. The reman losses are the least recoverable dollars in the entire model. You might think “just repair more reman units to stop giving away new ones.” But repairing a reman unit uses a line-slot that would otherwise sell for $250K — so shifting the slot from new-sales to repair doesn’t create value, it just moves the loss from one column to another. The only true fix is building more line capacity (expensive, slow, capital-heavy). So the giant $1.3B “capacity” number from Channel B is mostly reman, and mostly unreachable by scheduling, by speed, or by the founders’ software — only by capital investment.

There’s even a correction to “repairs are cheap” (Channel C): each reman repair performed burns a sellable production slot. Valuing that forgone sale, doing the reman repairs costs on the order of $507M/yr — not the ~$81M a flat “$12K per repair” assumed. The general rule the segmentation reveals, worth remembering: a repair is only cheap when the capacity it uses couldn’t have made revenue anyway. N-1 and refurbish (dedicated/separate lines) are genuinely cheap to repair; reman is not.

(A configuration question this raises — flagged, not answered: NVIDIA could move reman onto dedicated lines to make it cheap, but keeps it on the build line on purpose, because the team that just built the unit knows the fix best. That tradeoff, and the week-to-week “who gets the shared line” decision Greg said “shouldn’t be made by the contract manufacturer alone,” is itself a scheduling problem — possibly something software could help with. A question for Greg, not a claim.)


Part 8 — The three clocks (why “turnaround time” is ambiguous)

People say “turnaround time” as if it’s one number. It’s really three separate stopwatches, and NVIDIA only tracks one:

ClockWhat it coversWhose money it ties up
Approval (1–2 weeks, untracked)The paperwork gauntlet before anything shipsCustomer downtime; not NVIDIA’s frozen capital
Repair (27–40 days, the only tracked one)Warehouse receipt → fixed and back on the shelfNVIDIA’s repair-pipeline capital (Channel A)
Collection (weeks to months, untracked)How long the broken unit sits in the customer’s garageThe bone pile — the biggest pool (Channel D)

The uncomfortable insight: the clock NVIDIA doesn’t measure (collection) is probably the longest and ties up the most money. And the three are linked through the shelf of spares — if the total loop (collection + repair) runs long, or too many units never return, the shelf empties and NVIDIA has to give away new units (that’s Channel B firing). This is why the same $250K unit can be “frozen capital” in one place and “a lost sale” in another: it’s one circulating unit, seen from two ledgers.


Part 9 — Putting numbers on it: the assumption ledger

Every number that feeds the model lives in an assumption ledger — a table where each input carries its value, its source, our confidence in it (High/Medium/Low), and how much it would hurt if we’re wrong. This is the backbone of the whole exercise; without it, the outputs aren’t defensible.

Because several inputs are honest guesses, we never report a single confident figure. We report a range — conservative → base → optimistic — built from plausible low/base/high values of each input. The inputs that matter most (the ones where a wrong guess moves the answer the furthest) are flagged as load-bearing, and each names a person who could tighten it:

Load-bearing inputWhy it dominatesWho could pin it down
Repair-coverage / capacity shiftThe ~$1.3B Channel B number rests on lifting the repair rate; and how much of the 40%-from-new is capacity vs. technical scrapLonny, Alex
Non-return fraction φ & collection lag τ_cGovern the biggest single pool (the bone pile) and the value of the founders’ productLonny (also: is the dead unit even NVIDIA’s asset? If not, Channel D vanishes)
Unit value V ($150K–$400K)Multiplies every dollar figure; a 2.7× swingGreg, Lonny
Return volume λ (150–1,200/wk)Multiplies every figure; an 8× swingGreg (he sees weekly metrics)
Reman : N-1 splitDecides how much of the whale is capex-gated (reman) vs. recoverable (N-1)Greg

The point of this table is not modesty for its own sake. It’s a to-do list: these five questions, answered, would collapse most of the uncertainty in the model. That’s the real deliverable — not the headline number, but knowing exactly which five things to go learn next.


Part 10 — What the model actually says

Putting it all together, in plain language:

There isn’t one “turnaround” lever — there are three, worth wildly different amounts:

  1. Faster logistics (Channel A) — real but modest: ~$59M/yr in freed carrying cost, plus a one-time few-hundred-million of units freed to sell. This is the slider everyone’s excited about, and it’s the smallest of the three.
  2. More repair capacity (Channel B, capacity part) — the biggest number, ~$1.3B/yr — but it needs capital investment (new repair lines), and the largest slice of it (reman) can’t be recovered by cleverness at all, only by building net-new line.
  3. Better collection (Channel D) — ~$357M of freed capital plus ~$364M/yr recovered, and about $1 billion less trapped over two years. It’s moved by software and coordination — which is exactly the product the founders propose to build.

How big are the trapped pools today (base case)? Three distinct piles:

  • Repair pipeline: ~$471M
  • Collection pool (the bone pile itself): ~$774M — the largest, and the one nobody is measuring
  • Plus accumulating write-offs growing at ~$520M/yr

Add them up and the model reaches ~$2.16 billion trapped after two years — which lands right on the “$2 billion+ bone pile” NVIDIA quoted, and reaches it through the correct mechanism (units sitting uncollected at customers), not the mistaken one an earlier version guessed.

The punchline that matters most for the founders: the biggest number they could move with software (the collection layer, Channel D) sits on top of the second-largest pool of money in the system — and it’s the pool nobody is currently even measuring. The truly gigantic number (reman capacity) is bigger, but it’s a factory-building problem, not a software problem. So the model doesn’t say “go” — it says the thing you can build happens to be pointed at real, large, currently-invisible money, and here’s exactly how much and what would prove it.


Part 11 — What would make this wrong (and what to ask next)

Being honest about how this could be wrong is part of the rigor, not a disclaimer. The biggest ways:

  1. The repair-coverage improvement is assumed, not observed. If the real repairable ceiling is 70% rather than 85%, the whale roughly halves. Ask Lonny/Alex: at full capacity, what fraction of returns can actually be repaired — and is the 60% a capacity limit or a speed limit?
  2. Unit value V is a pure guess spanning $150K–$400K, and it multiplies everything. Ask Greg/Lonny to confirm the blended value of a returned compute unit.
  3. Collection lag and non-return rate (τ_c, φ) are guesses yet govern the largest pool — and the whole thing only counts as NVIDIA’s money if NVIDIA legally owns the un-returned unit. Ask Lonny: how long do units really sit at customers, how many never come back, and is the broken unit NVIDIA’s asset or the customer’s? (If it’s the customer’s, Channel D shrinks dramatically.)
  4. The reman : N-1 split and reman line-time cost are guesses that decide how much of the biggest channel is recoverable. Ask Greg.

Part 12 — The surprises (the most valuable part)

Good research tracks what contradicts expectations, because that’s where the learning is. The model’s surprises:

  • NVIDIA is measuring the wrong clock. The 27–40-day repair time it tracks is a survivor’s average — it ignores everything piling up behind an over-capacity queue.
  • The “$2B bone pile” lives on the customer side, not in the repair shop — a correction we owe ourselves. It’s units sitting uncollected in customers’ garages, fixed by coordination, not by factories.
  • The biggest software-movable number and the founders’ product are the same thing. Not by design — it just falls out of the arithmetic.
  • “Repairs are cheap” is only true off the revenue line. The reman stream, repaired on the live sales line, is the one that spirals and the one whose losses are least recoverable by anything short of building new capacity.


Reference — everything in one place

Parts 13–16 reproduce the full rigor of the technical brief: every input and where it came from, every formula, and every result. Nothing here is new — it’s the same model, collected for reference.

Part 13 — The complete input ledger

Every number that feeds the model, with its value across the three scenarios, where it came from, our confidence (High / Medium / Low), and — in plain English — what it actually is. Reading the “Source” column: [Interview] = someone told us; [Public] = a filing/document; [Synthesis] = inferred across sources; [Speculation] = an honest guess, not yet grounded.

Input (symbol)Plain meaningConservative → Base → OptimisticSourceConf.
Flow unitWhat “one unit” is1 compute baseboard/tray ≈ 8 GPUs + 1 CPU (blended)[Interview: In-person debrief 2026-06-25]M
λ (lambda)Units returned per week150 → 400 → 1,200 /wk[Interview: Lonny 2026-05-12; Greg 2026-06-26]L
τ_r (tau-R)Repair time (warehouse→back-in-stock)40 → 33 → 21 days[Interview: Greg 2026-06-26]M
Approval delayPaperwork gauntlet before shipping7 → 10 → 14 days[Interview: Greg; Lonny]M
Repair coverage today% of returns actually repaired (blends 2 causes)~60% repaired / 40% new[Interview: Alex 2026-05-27]M
— capacity partof the 40%, the share that’s just throughput-limitedassumed to dominate[Speculation]L
— technical partof the 40%, the share that’s physically deadassumed minority[Speculation]L
Capacity coverage after fixrepair % achievable with more capacity70 → 85 → 90%[Speculation]L
ρ (rho)Utilization = load ÷ capacity1.7 (over capacity)[Synthesis] from the 60%M
τ_c (tau-C)Collection lag (unit sits at customer)45 → 60 → 90 days[Interview: Lonny 2026-05-12 & 2026-06-17]L
φ (phi)Fraction that never comes back2% → 10% → 15%[Speculation], back-out from $2BL
c# compute repair sites3 (Dallas, Houston, Guadalajara)[Interview: Greg 2026-06-26]H
Category mixreman+repair vs. refurbish~90% / ~10%[Interview: Greg 2026-06-26]H
— reman : N-1 splitwithin the 90% (opposite cannibalization)60% : 30% of total[Speculation]L
Reman capacity typeshared w/ new production?shared, revenue-subordinate[Interview: Greg 2026-06-26]H
Per-stream Vvalue: reman / N-1 / refurb$250K / $150K / $100K[Speculation]L
t_repair / t_buildreman repair time vs. new-build time (sets its opportunity cost)0.3[Speculation]L
RMA rejection rate% of returns ever denied<1%[Interview: Greg 2026-06-26]H
sTechnical scrap fraction (of returned units)3 → 5 → 8%[Speculation]; die→TSMC, CoWoS/HBML
VBlended unit value$150K → $250K → $400K[Speculation]L
mNVIDIA gross margin (profit on a lost sale)0.75[Public: ~75%]M
kCarrying rate (cost of frozen capital + depreciation) /yr15 → 25 → 40%[Speculation]L
c_rCost to perform one repair$5K → $12K → $25K[Speculation]L

Load-bearing inputs (the handful that move the answer most, each with who could pin it down): (1) the repair-coverage / capacity shift — Lonny, Alex; (2) non-return fraction φ and collection lag τ_c — Lonny; (3) unit value V — Greg, Lonny; (4) return volume λ — Greg; (5) the reman:N-1 split — Greg. These five, answered, would collapse most of the model’s uncertainty.

Part 14 — Every formula in one place

The symbols are the nicknames from Parts 2–8: λ arrivals/wk, τ_r repair time, τ_c collection lag, W generic dwell time, L work-in-process (units stuck), V unit value, m margin, k carrying rate/yr, s technical-scrap fraction, φ non-return fraction, ρ utilization, c sites, μ per-site repair rate, coverage = repaired share, shortfall = the capacity gap.

#FormulaIn words
1L = λ · WLittle’s Law: units stuck = arrival rate × time each stays
2Trapped capital = L · V…times unit value = dollars frozen
3ρ = λ / (c · μ)Utilization. If ρ ≥ 1, the queue grows without limit
4Repair-pipeline WIP = λ · (1 − φ) · τ_runits that actually reach repair (the returned ones) × repair time
5Channel A carrying = k · λ · τ_r · Vyearly cost of capital frozen in repair; falls as τ_r falls
5bper-day A sensitivity = k · λ · V ÷ 7 ≈ $3.6M/yr per day (base); one-time release = V · λ · Δτ_r ≈ $14.3M per daywhat one day of faster repair is worth
6Channel B lost sales = V · λ · (1 − φ) · (s + shortfall)new units given away, from two causes…
6atechnical floor = V · λ · (1 − φ) · s…physically-dead units (no lever fixes this)
6bcapacity part = V · λ · (1 − φ) · shortfall, with shortfall = max(0, 1 − 1/ρ)…repairable-but-not-reached units (fixed by capacity)
7Channel C repair opex = c_r · λ · coverage (+ reman opportunity, below)cost of doing the repairs
8aChannel D collection pool = λ · τ_c · (1 − φ) unitshow many sit at customers
8bChannel D carrying = k · λ · τ_c · (1 − φ) · Vyearly cost of that frozen pool
8cChannel D write-off leak = V · λ · φ per yearnever-returned units, each a lost sale (independent of τ_c)
9aper-stream WIP = λ_stream · (1 − φ) · τ_streamreman / N-1 / refurb each get their own Little’s Law
9breman capacity = max(0, LineCapacity − NewDemand)reman gets whatever line-time new production leaves
9creman repair opportunity = (t_repair / t_build) · V per repaireach reman fix burns a slot that could’ve sold
10Bottom line(τ) = −[ C_carry(τ_r) + C_carry_collection(τ_c,φ) + C_repair + C_expedite ] − R_lost − R_strand + one-time releasesthe whole P&L picture as a function of speed

Two facts worth restating because they’re easy to miss: in formula 8c, the write-off leak depends on φ but not on τ_c — so faster collection shrinks the pile (8a), while a lower never-return rate shrinks the leak (8c). Two separate levers. And in 6a, the technical floor is the one term in the whole model that no operational lever — speed, capacity, or the portal — can move.

Part 15 — All the results

15.1 The convergence gate (do our two methods agree?)

For the healthy, adequate-capacity case, computed three independent ways:

QuantityLittle’s Law (hand)System DynamicsDiscrete-Event Sim
Units stuck in repair (base)1,8861,8801,881
Capital frozen$471M$470M$470M
Verdict on capacity todayover capacity (ρ≈1.7)over capacity (grows w/o limit)over capacity (55% served, wait rising)

They agree to a fraction of a percent — the model passes its own cross-check.

Simulation specifics (the rigor behind “the sim said”): the DES was run 12 independent times (replications) and reported with 95% confidence intervals. Healthy case: queue wait 0.0–0.8 days (mean ≈ 0.1), throughput ≈ 100% of arrivals, utilization ~72% → the ~33-day repair time is almost all transit/handling, not queue. Over-capacity case: 55–56% of arrivals served (≈ 1/ρ, independently reproducing Alex’s 60%), and average wait grew 103 → 193 days as the run doubled — the unbounded-queue signature. The System Dynamics model settled at exactly 1,880 units / $470M in the healthy case, matching Little’s Law to the dollar.

15.2 The three scenarios (what “conservative / base / optimistic” mean)

ConservativeBaseOptimistic
λ (returns/wk)1504001,200
V (unit value)$150K$250K$400K
k (carrying/yr)15%25%40%
capacity coverage60→70%60→85%60→90%
s (technical scrap)3%5%8%
τ_c (collection lag)45→30 d60→30 d90→30 d
φ (non-return)5→2%10→3%15→3%

15.3 The headline sweep — what each lever is worth

Scenario is a halving of turnaround (≈60→30 days end-to-end); the collection row is τ_c 60→30 d and φ 10→3%.

What it movesChannel / leverConservativeBaseOptimistic
Working capital freed (one-time)A / logistics$54M$236M$1.13B
Carrying cost saved (recurring)A / logistics$8M/yr$59M/yr$452M/yr
Cannibalized sales recovered (recurring)B / capacity$117M/yr$1.30B/yr$7.49B/yr
Technical-scrap loss (floor — no lever recovers it)B / technical$33M/yr$234M/yr$1.70B/yr
Collection capital freed (one-time)D / portal$130M$357M$1.55B
Non-return write-offs recovered (recurring)D / portal$63M/yr$364M/yr$2.18B/yr
Repair-pipeline capital today (level)$106M$471M$2.26B
Collection-pool capital today — the bone pile (level)$131M$774M$4.13B

15.4 The trapped-capital pools today (base case)

Three distinct piles, not one:

  • Repair pipeline (units in the shop): ~$471M
  • Collection pool (the bone pile — units sitting at customers): ~$774M, the largest
  • Write-offs (never-returned units): accumulating at ~$520M/yr
  • Total from the two-stock simulation: ~$2.16B trapped by year 2 — matching NVIDIA’s quoted “$2B+ bone pile,” via the correct mechanism (uncollected units), not a repair backlog.
  • With a scrap branch added, the permanent give-away-a-new-unit rate = 58 units/wk = 40 never-returned + 18 technically-dead → ~$753M/yr of lost sales before any capacity shortfall ($520M Channel D + $234M Channel B floor).

15.5 The collection-portal result (Channel D, base)

Improving collection from τ_c = 60→30 days and φ = 10→3%:

  • Collection pool: $774M → $417M (frees $357M of working capital)
  • Write-off leak: $520M/yr → $156M/yr (recovers $364M/yr)
  • Total trapped at year 2: $2.16B → $1.17B (~$1B less trapped)
  • The $234M/yr technical-scrap floor does not move — the portal can’t reach it.

15.6 The three repair streams (why reman is the whale)

Base split (reman 60% / N-1 30% / refurb 10% of λ), simulated as three separate pipelines:

StreamCapacitySimulation result
Reman (shared production line, revenue-subordinate)preempted by new salesDiverges$2.24B frozen by year 2, $1.12B/yr bled into new-unit give-aways
N-1 (dedicated repair line)independentStable at ~508 units = $76M
Refurbish (separate shop)independentStable at ~288 units = $29M

Plus the hidden cost the flat opex missed: performing the reman repairs (130/wk on the shared line) forgoes ~$507M/yr of new sales (at t_repair/t_build = 0.3), vs. the ~$81M a flat $12K/unit assumed. Reman losses are the least recoverable in the model — only net-new line capacity (capex) touches them; reallocating just moves the loss to the new-sales column.

15.7 Sensitivity ranking (what moves the answer most → validate first)

  1. Repair-coverage / capacity shift — $0.12–7.5B/yr. Dominates by 10–20×. → Lonny/Alex.
  2. Non-return φ / collection (Channel D — the portal’s lever) — $63M–2.2B/yr recurring + $0.1–1.6B one-time. Second-biggest recurring lever, and the one the founders’ product moves. → Lonny.
  3. Unit value V — linear on every channel; 2.7× swing. → Greg/Lonny.
  4. Return volume λ — linear on every channel; 8× swing. → Greg.
  5. Collection lag τ_c — sets the bone-pile level ($0.1–4.1B). → Lonny.
  6. Carrying rate k, scrap s, per-repair cost c_r — second-order.

Part 16 — Open questions, with named contacts

The model’s real output is this to-do list — the questions that would collapse its uncertainty:

  • Of the ~40% filled from new, how much is technically dead vs. merely capacity-starved? And do die failures get recovered from TSMC?Lonny, Alex. Decides whether the biggest channel has a large recoverable part or a hard floor.
  • Mean collection lag τ_c, the never-return rate φ, and — critically — is the broken unit NVIDIA’s asset (core-return) or the customer’s?Lonny, Alex. Sets the largest pool and the value of the portal. If it’s customer-owned, Channel D collapses.
  • Blended unit value V and the tray/baseboard/SXM mix.Greg, Lonny. Linear on everything.
  • Weekly return volume λ by category.Greg (he sees the weekly metrics). Tightens the whole model.
  • The reman : N-1 split, the reman line-time ratio, and whether reman could move to dedicated lines.Greg. Decides how much of the whale is capex-gated vs. recoverable.
  • Does NVIDIA carry the pipeline as its own working capital, and at what cost of capital (k)?Greg / NVIDIA finance. Sets Channel A.
  • The expedite cost curve — what does air-vs-ground actually cost per unit? → Greg. The one cost that rises as speed rises.

How to read the technical brief

You now have everything in this document — Parts 13–16 hold the full ledger, formulas, and results. The technical brief nvidia-reverse-logistics-operational-model-cleanroom-2026-07-04 is the source of record: the same content in dense form, plus the boundary/scope notes, the method write-ups, and the running commentary on what each result means. It’s the version to cite. When you read it:

  • λ, τ, φ, ρ, k, s, V — those are the nicknames from this document (arrivals, times, non-return fraction, utilization, carrying rate, scrap fraction, unit value).
  • Channels A/B/C/D — the four money-drains from Part 6.
  • “Little’s Law,” “DES,” “SD,” “ρ ≥ 1,” “steady state,” “CI” — all defined in Parts 2–4 above.
  • “reman / N-1 / refurb” — the three repair types from Part 7.

This is a companion for understanding, not a source of record — every number here traces to the technical brief and, through it, to the primary interviews. No conclusions are drawn; the founders decide.