NVIDIA RMA Process Map

A process-first read of NVIDIA’s reverse-logistics pathway. First: the workflow itself, from things we heard directly in meetings. Then: pain points and gaps. Then: research-filled overlay with confidence labels. Goal: walk in able to ask the questions that turn this into a deployable solution.


Don’t forget (the pocket card)

If you only re-read three things before walking in, re-read these.

  1. Do not call this an RFP. Greg + Lonny never used that word. They said “target rich environment” + “put some bullets together” + “let me find out.” The Palantir-Accenture-build-internal triangle is our hypothesis from the debrief, not their statement. Pressure-test it; don’t assert it.
  2. The single highest-leverage moment is the exec-sponsor question. “To make sure we scope to the right approval path — who would own sign-off on a Phase 1 engagement of this size?” If they name Brian Feller or anyone we haven’t found in public sources, the entire engagement plan updates from that one sentence.
  3. The Wistron Mobility Way question is the most concretely answerable open question we have. Fort Worth is two buildings — Heritage Parkway (324K sq ft, production) and Mobility Way (767K sq ft, publicly described only as “renovations”). Ask which building, and whether the repair line is the same SOW as production. Answer changes how we think about geography of the work.

§0 — Reading guide: what’s reliable vs. what’s our guess

A pre-flight calibration so we don’t mistake our own framings for theirs.

Source typeExamplesTreat as
In-meeting first-handThings Greg, Lonny, or Alex said in our callsReliable — primary evidence. Cited as [L] (Lonny 2026-05-12), [A] (Alex 2026-05-27), [G+L] (Greg + Lonny 2026-06-17), [G-0626] (Greg solo, 2026-06-26 en-route call)
Our debrief inferenceThings Bliss + Dustin said about the meeting after the meetingSpeculative — our interpretation, not their statement. Cited as [debrief]
External researchPublic sources, NVIDIA filings, trade press, research-agent outputInference — useful, never primary. Always confidence-labeled

A specific correction we owe ourselves

The “RFP” framing is ours, not theirs. Greg + Lonny did not say “we are running an RFP.” They said: “we’re such a target-rich environment”; “we just did a workshop with E&Y two weeks ago”; “how would you envision… what type of stuff would you do?”; “let me find out” (Lonny, on whether we could build it); “I’m definitely open for you guys to meet in person”; “we’d have to do an NDA”; “put some bullets together.” [G+L]

It was Bliss in the debrief who said “they’re either going to go to Palantir or they’re going to go internal or they’re going to go to Accenture” [debrief]. That is our hypothesis about the competitive shape — useful to plan against, but not something Greg or Lonny told us. Worth pressure-testing in Thursday’s meeting: Is there a formal procurement process running, or is this an executive-sponsored direct engagement?

Same calibration applies to “design partnership,” “RFP finalists,” and “Bare knuckle three-to-ten-day sprint.” Those are our framings, not their language.


§1 — The RMA workflow, as Greg + Lonny + Alex described it

This is the core process. Master this before extending to upstream / downstream.

The 12 steps, in order

Numbered by sequence. Each step has an owner, the tool(s) involved, and what crosses the seam to the next step.

Side legend: 🅝 = step NVIDIA controls. 🅒 = step customer controls. 🅒🅜 = step contract manufacturer controls. Customer- and CM-side steps are where NVIDIA’s visibility is structurally hardest.

#SideStepWho owns itTool / systemWhat flows out of this step
1🅒Customer detects a failed unit in their data centerCustomer’s data-center ops teamCustomer’s own fleet toolsA decision to open an RMA ticket [L] [A]
2🅒Customer opens RMA via NVIDIA portalCustomer (via web portal)Salesforce-backed customer portalA new case in NVIDIA’s SFDC [A] [G+L]
3🅝NVIDIA runs the three-gate approval gauntlet — case management (Lonny) → quality → finance (see Greg-0626 walk-through below)Lonny’s team + quality + financeSalesforce; failure-analysis lab for unknown codesValidated serial #, warranty entitlement, advance-replacement-or-standard decision, and an approval notice. <1% of RMAs ever rejected [L] [A] [G-0626]
4🅝NVIDIA dispositions: standard RMA or Advance Replacement (ARMA)Frontline + Lonny’s teamSalesforceA ship-replacement instruction [L]
5🅝NVIDIA ships replacement to customerWarehouse + 3PLExpeditors (3PL, replaced Omni)Delivery notice, tracking #, serial s, ship-to info — all communicated via email + spreadsheet, with customer reply-all chain forming here [G+L] [A]
6🅒Customer receives replacement; coordinates rack downtimeCustomer DC team + customer’s business unitCustomer-side ticketingA swap window — or “let it fail” deferral [L]
7🅒Customer physically swaps the unitCustomer DC technicianCustomer-side procedureA defective unit ready to return; 10-day window for return per NVIDIA terms [L, contractual]
8🅒 → 🅒🅜Customer wipes data + ships defective unit back to CM, ASN should fireCustomer + (sometimes) ODM integrator (Quanta etc.)3PL pickup (Expeditors); piloting direct pickup for one hyperscaler. ASN (Advance Shipping Notification) should fire to CM via EDI — today often email/paper [A]A defective unit in transit + an ASN the CM may or may not get cleanly [L] [G+L]
9🅒🅜CM receives unitWistron or FoxconnReceiving system; receipt acknowledgment EDI being turned on [A]A defective unit in CM intake
10🅒🅜 + 🅝CM dispositions into one of three categories: reman / repair / refurbish (see Greg-0626 below)Wistron / FoxconnShould be SAP-issued production order — today often not issued before CM starts [A]A repair job in process; compute routes Dallas/Houston/Guadalajara, networking routes Vietnam/Israel/India [G-0626]
11🅒🅜CM repairs using consignment + turnkey materialsWistron / Foxconn”Suspect sheets” (early ’90s spreadsheets) tracking component consumption [A]A repaired unit; consumption record for reconciliation
12🅝NVIDIA material manager reconciles consignment usageNVIDIA material mgmtManual reconciliation todayCleared inventory variance [A]
13🅒🅜 → 🅝Repaired unit reports back to Planning teamWistron / Foxconn → NVIDIA PlanningManual / spreadsheet today; SAP $2M+ planning automation deal in flight [A]Inventory signal back into planning
14🅝 → 🅒🅜Planning gives CM the next repair commitNVIDIA Planning teamBack-and-forth, partly manualA pull signal for CM to procure turnkey materials and start the next batch [A]
15🅝Repaired unit enters NVIDIA’s services poolNVIDIA warehouseNVIDIA inventory mgmtLike-for-like spare ready for next customer (same feature/function, different serial #) [G+L]
16🅝 → 🅒Dispatch from services pool to next customerNVIDIA warehouse + Expeditors3PL handoffCloses the loop for the next RMA cycle [G+L]

Read this table by where the side-marker changes color: every transition between 🅝 and 🅒 or 🅒🅜 is a seam where data has to cross an organizational boundary. Five of NVIDIA’s biggest pain points (§2) sit on those seams.

The five things Alex Zhu added on top of the steps

  1. Consignment vs. turnkey: NVIDIA owns the high-value components (chips, boards) and parks them at CM facilities. CMs procure their own low-cost parts (cables, third-party items) — that’s “turnkey.” Two material streams, one repair line. [A]
  2. Repair throughput is structurally insufficient: of every 100 returns, only ~60 actually get repaired. The other 40 are filled from new inventory — “we have to skew this 40 from new buy from resin, and new buy is all what Jensen cares about because new buy is basically revenue.” [A]
  3. ~90% of all repairs are “remanufacturing” — ECO/known-issue batches, like a car recall. The remaining ~10% are ad-hoc. All repairs are free to the customer. [A]
  4. The CM↔Planning loop is the unmeasured one. “Unless I receive a signal from a planning team, how am I supposed to know what to buy?” [A]
  5. SAP is being paid $2M+ specifically to automate the planning capability — “eradicate and get rid of the manual Excel-based planning model.” [A]

What Greg added on 2026-06-26 — the deepest process walk-through yet

A solo call with Greg (led by Dustin) on his drive back, walking the whole chain end to end. [G-0626] throughout. Six things that sharpen or correct the map above:

  1. The “customer” is two roles, not one — and the RMA hinges on telling them apart. Who initiates the RMA and where the unit ships to are the two identifiers that matter. Google is both customer and “sold-to” (buys direct from NVIDIA). xAI is the customer but not the sold-to — Dell or Quanta buys the chips, assembles, and stands up the data center, sitting in the middle as integrator. NVIDIA has no reliable directory of who to talk to on the customer side (“emails going to people that have left the company… it’s not a systemic, robust process”). Read steps 1–2 as customer / sold-to / integrator, not one node.

  2. A three-gate approval gauntlet runs before the unit ever ships to a warehouse — this is what the table compresses into step 3. Case filed → case management (Lonny’s team): warranty check, failure-log review, serial-number lookup in Salesforce → quality team: known failure code = approve, unknown = route to a failure-analysis labfinance: approves → approval notice to customer → customer ships to a warehouse. Less than 1% of RMAs have ever been rejected — the gauntlet is almost pure latency, not a filter. Greg’s own opportunity framing: high-volume known-defect batches (“we go to Microsoft and say everything we shipped you between these dates has this problem — 6,000 units, 6,000 serials”) should be auto-fast-tracked straight through, not walked gate by gate. Greg doesn’t own the approval criteria — Lonny does — so the detail behind each gate is a direct question for Lonny.

  3. Three repair categories, not two (this corrects the “~90% reman / ~10% ad-hoc” split in the Alex list above):

    • Remanufacturing (reman) — current-revision product, known defect + known fix; high volume; repaired on the same live mass-production line that built the original unit.
    • Repair (sort & repair) — N-minus-one (last year’s revision); also high volume; same site, separate dedicated repair line (keeps the current-revision production line undisturbed while the build team’s knowledge is still fresh).
    • Refurbishment — one-off, unknown condition, diagnosed on arrival; may be dismantled and salvaged for parts; the only true “repair-shop” path.
    • ~90% is reman + repair combined (the first two); refurb is the tail. Greg: reman and repair “are really recalls — known flaws”; refurb is the onesie-twosie diagnostic path. Routing rule: whoever built it usually repairs it.
  4. Repair geography is (at least) two separate networks. Compute (GPUs, trays, baseboard assemblies, SXM modules) routes to Dallas, Houston, or Guadalajara — all Foxconn/Wistron. Networking (switches — no GPUs on them) routes entirely separately: Vietnam, Israel (post-Mellanox), India. European compute still ships back to the US — no proximate compute repair capability there. A little reman runs in Taiwan, but most reman + refurb is Dallas / Houston / Guadalajara. This flips our earlier inference that the Mexico endpoint was Juarez (Wistron/Wiwynn) — Greg names Guadalajara; see §5.

  5. Cycle time, with real numbers for the first time. ~60 days end-to-end today (Greg’s estimate); ~30 days best case. Of that, 27–40 days is warehouse-receipt → repaired-unit-back-in-stock — the only stretch they track weekly. The approval gauntlet adds another 1–2 weeks that isn’t tracked at all, because NVIDIA only starts the clock at warehouse arrival, not at case submission. Two clocks, decoupled: the customer gets a replacement from stock in ~1–2 weeks (if inventory exists), while their own serial number takes 30+ days to repair and return to stock. Greg wants to start measuring both the customer-experienced time (case-submit → replacement received) and the serial-level throughput time (receipt → good inventory). The two metrics he cares about — throughput and cycle time — and he wants them translated to financial impact. (This interview-confirms the ~60-day baseline our economic model uses.)

  6. The two revenue-vs-repair tensions, confirmed and sharpened. Greg named exactly two (matching the model): (a) reman-line capacity — repair competes with new production on the same assembly line; when constrained it’s a joint, negotiated decision between NVIDIA’s service org and the revenue side (revenue wins ~9/10, “but they shouldn’t be making that decision on their own”). NVIDIA’s repair-side volume is “minuscule” next to revenue, so it combines forces with the revenue org to hold the CMs commercially accountable for both. (b) New-buy draw — when repaired stock is short, warranty gets filled from new units; “we try to limit the number of new buys in our service depots,” and GPUs especially are fought over (“there’s never enough”). This is the operational face of Alex’s 60-of-100 number and the model’s “replacement = lost sale at $P$.”

Logistics footnote: the Dallas ↔ Guadalajara cross-border move + customs clearance is the named queue-time bottleneck. Omni (bought by Forward Air) is being exited in the US — leaving Omni Union City and Omni Dallas, consolidating to Expeditors Dallas for US returns + fulfillment; Omni may be retained in Asia (Hong Kong / Taiwan), still debated. Speed is the dominant lever — Greg: “if we could fly these all by private jet, we would.”

Where Greg framed the problem

The frame Greg + Lonny used most in the meeting:

  • “We don’t even really know the score of the game right now.” They cannot measure their own SLA performance against the 30-day standard. [G+L]
  • “We come to the conversation, and we’re not even armed.” When a customer escalates, NVIDIA has no shared view of the case to argue from. [G+L]
  • “What’s the source of truth? Which dashboard do I believe?” Lonny’s frame: many incremental dashboards, no end-to-end system architecture, no agreed-on data. [G+L]
  • “We’re such a target-rich environment.” Greg’s framing — many places to start. [G+L]

§2 — Where the workflow breaks (from interviews only)

Each item below is a failure point Greg / Lonny / Alex named directly. Not inferred.

Workflow-level pain points

#Pain pointWhere it livesSource
P1Post-ticket spreadsheet cascade — DNs, tracking numbers, serial numbers, ship-to info all updated manually after the SFDC ticket is opened. Customer reply-all threads balloon to 16+ people; NVIDIA’s version and customer’s version diverge.Step 5 → step 9[G+L]
P2No automated ASN to CM — warehouse expects shipment notification; today it’s email/paper.Step 8 → step 9[A]
P3No production order issued before CM begins repair — CMs work without a system-of-record SAP PO.Step 9 → step 10[A]
P4”Suspect sheets” at CMs — component consumption tracked on ‘90s-style spreadsheets.Step 11[A]
P5Manual material reconciliation — material manager has to check CM consumption claims; risk of over-reporting.Step 12[A]
P6CM ↔ Planning back-and-forth is partly manual — repair commits not flowing as system signals.Step 13 → step 14[A]
P730-day SLA is not measured — “we don’t know the score of the game.”Step 4 → step 16 (the whole arc)[G+L]
P8No escalation triggers — “day 31 should fire a notification; today nothing does.”Spans the whole arc[G+L]
P9Customer business-unit approval bottleneck — DC teams can’t unilaterally take a rack down; BUs (Instagram, Facebook) prefer “let it fail” → ARMA units sit idle for weeks/months → inventory planning distortion.Step 6[L]
P10ODM integrators (Quanta etc.) add touches without value on returns — equipment gets handled multiple times waiting for carriers; NVIDIA piloting direct pickup with one hyperscaler.Step 8[L]
P11Repair throughput insufficient (~60/100) — gap is filled from new inventory, which is what Jensen “cares about.”Step 10 → step 13[A]
P12Telemetry walled off at customer fence line — chips wiped before return; first signal of product perf is ticket volume.Step 1 → step 2[G+L]
P13Stranded shipments — Phoenix five-truck incident: NVIDIA sent five trucks of “hundreds of millions of dollars of equipment” to customer site without coordination; trucks turned around to a security yard in the Phoenix sun.Step 5 / step 8 — coordination layer[G+L]
P14Hyperscale form-filling is operationally hostile — “you don’t want xAI or OpenAI to fill the form.”Step 2[A]
P15Mellanox / Israel team integration gap — “still operate like they’re separate companies.”Org-level — touches step 3[L] mentioned in May, never repeated in June — may be stale; treat as worth confirming, not as a current pain
P23Approval gauntlet is pure latency — case mgmt → quality → finance adds ~1–2 weeks before the unit even ships, yet <1% of RMAs are ever rejected. Known-defect batches walk every gate instead of being auto-fast-tracked. And the clock isn’t started until warehouse arrival, so this delay is invisible in the metrics.Step 3 (pre-warehouse)[G-0626]

How the pain points cluster

Three clusters, not 15 independent things:

  1. No system-of-record after the SFDC ticket. P1, P2, P3, P4, P5, P6 — every handoff from step 5 onward leaks into email + spreadsheets.
  2. Nobody measures the loop. P7, P8, P13 — there’s no end-to-end SLA dashboard, no escalation logic, no early-warning on stranded shipments.
  3. The customer side is opaque. P9, P10, P12, P14 — NVIDIA can’t see what customers are doing with units, can’t pull telemetry, can’t make customers fill forms cleanly, and can’t make customer BUs authorize swaps.

The dashboard Lonny + Greg sketched at the end of the meeting maps to cluster 2 first. Cluster 1 is what makes the dashboard possible (data has to be in a system to be dashboarded). Cluster 3 is where the bigger ambition lives.


§3 — What we don’t know about the systems and the process (interview gaps)

Strictly: things we cannot answer from what Greg / Lonny / Alex said.

Systems / tooling we don’t have a clear picture of

TopicWhat’s unclear
SalesforceEdition (Enterprise, Unlimited, Industries Cloud)? Service Cloud or custom? What custom objects exist? Lightning?
SAPECC or S/4HANA? On-prem, RISE, or HEC? What modules — SD, MM, PP, EWM? What is the $2M+ deal actually covering?
Baxter PlanningWhich modules? How does it sync with SAP? Does it consume Salesforce data? Does it consume CM repair data?
Data lakeWhat platform — Databricks, Snowflake, NVIDIA-internal? What data is already there? Who has access?
WMS / TMS being rolled outVendor names? In-house or partner systems being extended into 3PL?
EDIWhat platform? Who’s the integration partner (the EDI-with-couple-of-customers piece)?
Identity / SSOHow would a third-party developer authenticate? Okta? Azure AD?
API surfaceWhat APIs already exist that we could call? What would we need to build?
The EY workshop outputBeyond “two initiatives,” what did EY architecturally produce — a process map, a tooling proposal, a vendor short-list?
CM-side systemsWhat do Wistron and Foxconn actually use for repair tracking today, beyond “suspect sheets”? Is there any system at all?
3PL integrationHow does Expeditors plug into SAP / Salesforce today?
The customer portalIs it on Salesforce Experience Cloud, or a custom front-end on top of SFDC APIs?

Process facts we don’t have

TopicWhat’s unclear
RMA volumeHow many RMAs / month? / quarter? / year? Lonny said “hundreds today, thousands soon” but not a number.
Active RMAs at any momentThe thing the dashboard would put on its front page — never quantified.
Cycle time todayWe know the SLA spec (30 days). We don’t know actual mean or distribution. Partly answered 2026-06-26: ~60 days end-to-end (Greg’s estimate), ~30 best case; 27–40 days warehouse-receipt → back-in-stock (tracked weekly); approval adds ~1–2 weeks (untracked); clock starts only at warehouse arrival. Still open: true mean/distribution, and the customer-experienced case-submit → replacement time. [G-0626]
Stranded shipment frequencyOne Phoenix event. Monthly? Quarterly? One-off?
Spreadsheet thread countHow many parallel email/spreadsheet threads are running right now across the org?
CM consignment valueHow many dollars of NVIDIA-owned chips sit at Wistron + Foxconn at any given time?
ARMA idle timeLonny said “weeks or months.” What’s the mean? What’s the cost?
Per-customer SLA termsDo hyperscalers have custom contracts that tighten the 30 days?
Custom escalation pathsWhat contractual escalation rights does e.g. Meta have?
Repair cost per unitPer-SKU economics — never quantified.
Stranded inventory costWorking capital tied up in stranded equipment — never quantified.
Failure-mode ParetoGreg cited bent connector pins, heat sinks, backplane, network switches. Which dominates by frequency? By cost?
Geography of the workMexico — which CM, which city? Answered 2026-06-26: compute repairs route to Dallas, Houston, and Guadalajara (Foxconn/Wistron); networking to Vietnam, Israel, India; European compute returns to the US. Still open: Dallas July 2026 production-vs-repair and which Wistron building; the specific HK/Taiwan facilities. [G-0626]
Who else inside NVIDIA touches thisIT, Sales account teams, field engineers, Mellanox team, quality / FA team — how do they intersect with Lonny + Greg’s pillar?
The second EY workshop initiativeOne disclosed (the portal). The other — not yet.

Org / authority gaps

  • Who is the exec sponsor above Greg + Lonny? Named 2026-06-26: “Manu.” Greg + Lonny are peers; both report to Manu, who heads the service-supply-chain org (five functions: customer service, reverse logistics + warehousing, repair operations, planning, and Greg as the horizontal “fifth wheel”). Corroborated by the 2026-06-25 in-person debrief (Manu sits under Deb Shoquist). Spelling / LinkedIn identity still being confirmed via Ali. [G-0626]
  • What’s Greg’s actual budget authority? What size deal can he sign vs. needing the exec sponsor’s approval?
  • What’s the timeline they’re operating on?
  • Does Alex Zhu’s reverse-supply-chain transformation sit inside the same program as Greg + Lonny’s, or parallel to it?

§4 — The extended pathway (upstream + downstream), interview-anchored

The full chain widens beyond the RMA workflow. Greg + Lonny + Alex touched on the wings — less detail than the core, but it matters for understanding where their problem ends and where ours might start.

Upstream — failure detection at the customer (step 0, before step 1)

What they told us directly:

  • “You’re not gonna get telemetry off of their devices.” Customer data is highly confidential. Chips are wiped before return. [G+L]
  • First signal of product performance is ticket volume vs. install base. Any two data points start a trend. [G+L]
  • Logs come reactively after fault — never proactively. [G+L]
  • Greg’s prior gig (likely Infinera / photonics) ran a crawl program that scanned customer networks, predicted failures within ~90 days. Customers usually picked “let it fail + 4-hour SLA” over coordinating maintenance. But NVIDIA still got to pre-position spares. [G+L]
  • Lonny in the night-before chat softened this: “it’s not that the data doesn’t exist; it’s just inaccessible.” Some customers will share metrics selectively for triage. [LonnyChat]

What we don’t know: which customers will share what, through what mechanism. Whether NVIDIA has a triage-data program in flight. How customer-side fleet tools (Meta’s, Microsoft’s) map to NVIDIA’s case-open trigger.

Downstream — planning, services pool, and demand forecasting (steps 13–16+)

What they told us directly:

  • Baxter Planning does demand planning and forecasts both sales and failure rates. [L]
  • Planning ↔ CM signaling is partly manual today. SAP $2M+ deal is automating it. [A]
  • Services pool feeds the next ARMA cycle. Like-for-like (same feature/function, different serial #). [G+L]
  • Repair throughput insufficiency spills back into new-inventory consumption — economic conflict with new-buy revenue. [A]

What we don’t know: how Baxter actually models failure rates; what its inputs are; whether the SAP $2M deal touches Baxter or is upstream of it; the actual policy on when planning pulls from new inventory vs. waits for repair.

The four pillars — was this our impression or a real org chart?

Lonny in May described four pillars: dedicated repair lines, reverse logistics, demand planning (forecasts failure rates), systems group (automation/tooling). [L]

Greg in June described two EY workshop initiatives — the portal is one. [G+L]

We don’t actually know if these are the same program or different programs. This is a worth-asking question.


§5 — The RMA workflow with research overlay

Same 12-step process. Each gap from §3 backfilled where research can; confidence labels on every fill. Read this as: if the meeting confirms or denies these inferences, what do we update?

Geography of the work

StepInterview saidResearch adds (confidence)
5 / 8 / 16Expeditors (3PL) replaced OmniExpeditors is global, 340+ locations, non-asset-based; specialty in semiconductor reverse logistics. High [Expeditors corporate]
9–11Compute repair = Dallas, Houston, Guadalajara (Wistron + Foxconn); back-office HK; warehouse Taiwan [G-0626]Wistron Fort Worth is two buildings: 15200 Heritage Parkway (324K sq ft, $580M, primary, production); 14601 Mobility Way (767K sq ft, $181M, secondary, publicly described only as “renovations”). Mobility Way is 3x larger than the production primary site — strongest candidate for Lonny’s “repair line.” Medium confidence on the inference (no public source confirms repair). Greg-0626 now confirms Houston as a compute repair node — promotes the earlier “Foxconn Houston” speculation. [Fort Worth Report 2025-08-21; Hillwood; Dallas Innovates]
9–11Mexico repair routes = Guadalajara (Greg-0626, explicit)Updated 2026-06-26: Greg names Guadalajara, not Juarez. This flips our prior Medium-confidence inference (which favored Wistron/Wiwynn Juarez on the $23M warehouse-lease profile). Foxconn Guadalajara (GB200 megaplant, $500M–$900M, 240K servers/yr) is now the stated compute repair endpoint; the Juarez warehouse may still coexist as a logistics node. Now interview-anchored, High on Guadalajara as the Mexico repair site. [G-0626; Taipei Times 2025-05-09; DigiTimes; Mexico News Daily]
9–11Networking repair = Vietnam, Israel, India (separate from compute) [G-0626]Switches carry no GPUs and route through their own network: Vietnam, Israel (post-Mellanox), and India (likely Cumulus/Mellanox legacy). European compute still ships back to the US — no proximate compute repair in Europe. Interview-anchored, Medium-High (Greg flagged he’d confirm the networking specifics).
9–11Asia footprintNo public source names a specific NVIDIA repair facility in HK or Taiwan. Wistron Hsinchu + Hukou (production); Foxconn-NVIDIA Taiwan supercomputing cluster $1.4B H1 2026 (production). Greg-0626: a little reman runs in Taiwan; most reman + refurb is Dallas/Houston/Guadalajara. Reverse-flow specifics still externally invisible. Low confidence on identifying Asia repair facilities specifically.

Volumes that backfill the “hundreds → thousands” framing

StepInterview saidResearch adds (confidence)
1Meta has 100K GPUs; wants 1M in 5 yrsMeta Llama 3 published data: 16,384 H100 cluster, 54 days, 466 interruptions, ~78% hardware. ~9% annualized failure rate. High [Meta Llama 3 paper, 2024]
1”Hundreds today, thousands soon”At 9% × 100K = ~9K failures/yr ≈ 750/month. At 1M Meta GPUs ≈ 7.5K/month. Lonny’s math holds. High (arithmetic).
2 / 8Customer-side telemetry walled offMeta runs three detection systems: Fleetscanner (45–60 day cycles), Ripple (alongside live workloads), Hardware Sentinel (kernel-space exception analysis, outperforms test-based by 41%). >66% of training interruptions are SRAMs, HBMs, network switches. High [Meta Engineering Blog Jul 2025; ASPLOS 2025]. Meta has the data Greg + Lonny said they can’t see.

Failure modes that backfill Greg’s anecdotes

StepInterview saidResearch adds (confidence)
10”Chip is rarely the problem” — connector pins, heat sinks, backplane, network switchesGPU faults 30.1% of Meta interruptions; HBM3 17.2% [Meta Llama 3]. CoWoS-L thermal / CTE mismatch at 1400W Blackwell TDP is a confirmed driver — warping, HBM PHY microbump failures can render the entire chip inoperable. Once CoWoS-bonded, individual chiplets and HBM stacks can’t be replaced. Medium-High [SemiEngineering; Chiplet Summit 2025; proteanTecs]
1Hyperscalers wrote OCP RAS spec — Lonny didn’t name itOCP GPU & Accelerator RAS Requirements v1.7, published 2025-10-23, standardizes error reporting, crash dumps, RCA, error containment, Redfish/IPMI SEL/APEI formats. Customers are setting the serviceability bar NVIDIA’s hardware has to meet. High [OCP RAS v1.7]

Systems that backfill the stack

TopicResearch adds (confidence)
Baxter PlanningFounded 1993 by an ex-Texas Instruments service-parts planner; customers manage $11B+ inventory across 35K locations, 120 countries; AI-powered “BaxterPredict.” Marlin took majority 2024. High [Baxter Planning corporate; PRNewswire/Marlin 2024]. Strong fit with Lonny’s description.
NVIDIA Enterprise Support tiersBusiness Standard (4-hr Sev 1, 8x5 live) and Business Critical (1-hr Sev 1, 24/7) are the two named tiers. Premium TAM is mandatory for every DGX SuperPOD. High [NVIDIA Enterprise Support Policy 2025-05-05 PDF]
NVIDIA AI Enterprise (the subscription that bundles support)~$4,500/GPU/yr list, 1-year; ~$22,500/GPU 5-year; $2.00/GPU/hour PAYG via cloud marketplace. High [Dell APD ac566091; Insight; NVIDIA marketplace]
Mellanox Global Expedite RMASilver / Gold / Platinum, with 4-Hour Expedite RMA as an explicit upgrade option. High [network.nvidia.com Mellanox RMA PDF]

Who runs this org — the org gap research can partly fill

TopicResearch adds (confidence)
Likely exec sponsor todayBrian Feller, VP Global Planning, Logistics & Services — joined NVIDIA May 2021 from Dell (19 yr). His own stated scope: “his team includes a growing Services organization to support reverse logistics, repair, and RMA fulfillment.” Round Rock, TX. Medium-High [ON Partners placement announcement; theorg.com; LinkedIn]
Critical wrinkleNVIDIA has an OPEN public Workday req — VP, Global Service Operations (JR1999315) — explicitly names Baxter and SAP and reports to “EVP of Global Operations” (Shoquist). $352K–$558K base. The literal seat above Greg + Lonny may not be filled. Either Feller is being elevated or Services is being split out under a new VP. High that the req exists. Unknown if filled. [NVIDIA Workday JR1999315; theladders.com]
Above that layerDebora Shoquist (EVP Operations) — scope includes supplier mgmt, CM mgmt, supply planning, logistics, quality. Greg + Lonny’s org rolls up here. High [NVIDIA Newsroom bio]
Alex’s “VP”Most plausibly Brian Feller (scope match). Medium-High
Trivedi successor (EVP Enterprise Sales)Not publicly named. Trivedi retired April 2026; now on Enphase board (June 2026). Possibly absorbed by Jay Puri or still unfilled. Unknown [SDxCentral; GlobeNewswire 2026-06-15]

The economics that interview data alone doesn’t surface

TopicResearch adds (confidence)
NVIDIA warranty reserve$2.81B FY26not the widely-quoted $8.22B (WarrantyWeek mis-aggregation). Accrual rate 0.46% → 0.92% → 1.15% in 2 years. Claims paid jumped to $957M (+337% YoY). Attribution per 10-K footnote: “primarily Compute & Networking.” High [NVIDIA FY26 10-K]
The “all repairs free” + “~90% remanufacturing” puzzleUnder ASC 460/450-20, recall reserves typically get separately disclosed when probable + estimable. NVIDIA’s 10-K does not break out recall-related charges. Two possible reads, both speculative: (a) ECO repairs treated as routine warranty (softens income-statement signal), or (b) the rising accrual rate IS the recall disclosure, dressed as warranty. No sell-side analyst has flagged this. It’s the strongest balance-sheet hook for a Phase 2 financialization angle. Medium confidence in the framing; High confidence that this is publicly invisible.
AMD comparisonAMD warranty reserve $308M FY25 (vs. NVIDIA’s $2.81B) — mirror curve, ~10x smaller. High [AMD FY25 10-K]. AMD does not segment data-center vs. client warranty.
Intel / Broadcom / MarvellDisclose nothing material in warranty reserves. NVIDIA + AMD ≈ 80% of US semi-industry warranty reserve; NVIDIA alone ~74%. High [respective 10-Ks; WarrantyWeek 23rd Annual Report 2026-04-16]

NVIDIA’s existing forward-fleet products (where our wedge does not go)

NVIDIA already operates a forward fleet-management layer. We sit downstream of where it stops:

  • Mission Control (GA on Blackwell) — AI factory infra mgmt; provisioning; monitoring; error diagnosis.
  • Run:ai (acquired Dec 30 2024 for ~$700M) — GPU workload orchestration; folded into Mission Control.
  • Base Command Manager (ex-Bright Computing) — cluster provisioning.
  • Fleet Command, Fleet Intelligence — edge / fleet telemetry.

The boundary: Mission Control + Run:ai is what runs while a unit is alive in a customer’s data center. Lonny’s portal is what NVIDIA needs after a unit fails. They don’t overlap because customers strip telemetry — Mission Control’s signals stop at the customer’s fence line. Medium-High [NVIDIA docs; cross-ref [[nvidia-primer-2026-06-18]]]


§6 — Where the workflow breaks, with research-overlaid pain points

Same three clusters from §2; research surfaces new pain points the interviews didn’t.

Cluster 1: No system-of-record after the SFDC ticket

The interview pain points (P1–P6) stand. New from research:

  • P16. The CM “suspect sheet” failure mode is industry-typical. Public reporting on Wistron and Foxconn reverse-logistics IT maturity is sparse, suggesting this is not a NVIDIA-specific gap but a common floor across high-volume CM ops. Medium confidence based on absence of public Wistron/Foxconn repair-mgmt-platform disclosures.
  • P17. SAP $2M+ planning automation may sit alongside, not under, the dashboard. Greg explicitly described his Cisco-era job as moving away from siloed ERP — a tension with the active SAP procurement. Medium speculation; worth pressure-testing.

Cluster 2: Nobody measures the loop

The interview pain points (P7, P8, P13) stand. New from research:

  • P18. OCP RAS v1.7 (Oct 2025) is hyperscaler-defined, not NVIDIA-defined. The customers wrote the spec for what NVIDIA’s serviceability should look like. Any dashboard NVIDIA buys probably has to meet hyperscaler-side data expectations to be acceptable. High that the spec exists.
  • P19. The 30-day SLA is publicly visible only in DGX support docs; whether hyperscalers have tighter custom SLAs is opaque externally. Unknown.

Cluster 3: The customer side is opaque

The interview pain points (P9, P10, P12, P14) stand. New from research:

  • P20. Meta has the telemetry data NVIDIA says it can’t get. Hardware Sentinel runs entirely customer-side, beats test-based methods 41%. The data exists; whether Meta would share it (selectively, for triage) is a relationship question, not a technology question. High that Meta has it; Unknown on sharing.
  • P21. The China RMA collapse has produced no publicly named replacement geography. Gray-market Shenzhen repair shops fill the gap (~500 units/month per shop, $1,400–$2,800/GPU). High that the problem exists; Unknown if NVIDIA is doing anything about it inside the customer-service ops scope. Per CLAUDE.md, surface only if Lonny raises it.

New organizational pain point research alone surfaces

  • P22. The org is mid-restructure. The open VP, Global Service Operations req (JR1999315) sits between Brian Feller (whose scope already says “growing Services organization”) and the EVP layer. This is a layer in flight. Engagement timing has to account for the possibility that the exec sponsor changes during the engagement.

§7 — Questions for the meeting

The whole point of this doc. Ordered by the answer’s leverage on whether we can deploy a solution.

Tier 1 — Authority, coordination, and timeline (must answer)

  1. “Who would own sign-off on a Phase 1 engagement of this size? Is that you, or does it route through Brian’s org?” The single highest-leverage question. If they name someone we haven’t found in public sources, that’s the person who matters.
  2. “How does Alex Zhu’s reverse-supply-chain transformation work intersect with what you’re doing? Same program, parallel program, or different problem?” Alex was our connector to Greg + Lonny but we don’t know if his project overlaps. If it overlaps, we walk into a coordination problem blind. If it’s parallel, the buying centers may be different.
  3. “How does this engagement get contracted — MSA + SOW, vendor procurement, or something else?” Reveals whether this is structured procurement (the RFP we inferred) or direct exec engagement.
  4. “What’s the timeline you’re operating on? When does a decision need to happen?” Tighter than 4 weeks → accelerate prototype. Longer than 8 weeks → embed and shape the spec.

Tier 2 — Scope and success (must answer)

  1. “Of the two EY workshop initiatives, what’s the second one? Is it a sibling portal effort or something internal?” Tells us what we integrate with vs. compete against.
  2. “What systems do you want this to talk to on day 1 vs. day 90?” SFDC, SAP (which modules / version), Baxter, data lake, Expeditors, CM systems, customer-side EDI.
  3. “What’s the success metric at 6 months? At 12 months?” SLA attainment measurable for the first time? Cycle-time reduction? Stranded-shipment events to zero?
  4. “What’s the data-lake situation today? What’s in it that’s reliable, and who owns the pipelines?” Affects whether we build on top or alongside.

Tier 3 — The process and the failure modes (worth asking)

  1. “Walk me through a real recent RMA — say, the Phoenix five-truck incident. What happened in each system at each step?” Surfaces the actual workflow vs. the conceptual one.
  2. “On the CM side, what’s the current playbook — do Wistron and Foxconn use any system at all today beyond the suspect sheets?” Tells us whether B2B integration is greenfield or a brownfield upgrade.
  3. “How do escalations actually happen today? Who at the customer calls who at NVIDIA?” Maps the human network we replace or augment.
  4. “Of the failure modes you mentioned — connector pins, heat sinks, backplane, network switches — is there a Pareto? Which dominates by frequency? By dollar?” Calibrates where the wedge sits and what failure-mode-aware tooling would actually pay back.
  5. “~60 of 100 repair rate — what’s the bottleneck? Capacity? Component availability? Diagnosis time?” Different bottlenecks imply different solutions.
  6. “What’s the EDI integration like with the couple of large customers who already have it? What data flows?” A template we could extend.

Tier 4 — Geography and physical footprint (worth asking)

  1. “Dallas July 2026 — production, repair, or both? Which Wistron building? Heritage Parkway or Mobility Way?” The most concretely answerable open question we have.
  2. “When you say Mexico, is that Juarez (Wistron/Wiwynn) or Guadalajara (Foxconn)?”
  3. “Asia footprint — which Hong Kong office, which Taiwan warehouse, specifically?”

Tier 5 — Customer-side (worth asking)

  1. “Which hyperscalers do you most want involved in the discovery sprint? Whose pain do we want this to reflect first?”
  2. “Are any customers willing to share triage data today, selectively, for fault diagnosis?” Lonny softened on this in the night-before chat — worth probing.
  3. “Are you aligned with OCP RAS v1.7? Is that shaping what you’re scoping?”

Tier 6 — Economics (listen, don’t press)

  1. “~90% remanufacturing — how is that accounted for? Standard warranty or treated as something different?” Listen for signal; don’t push. The answer’s value is the Phase 2 financialization signal, not the Phase 1 dashboard.

§8 — What we’d need to know in order to actually deploy a solution

The bridge from “understanding their problem” to “shipping the dashboard.” Each of these is something we don’t know yet and would need to know before writing a real SOW.

The minimum-viable deployment requires answers to

  1. Identity + access: How does a third-party developer authenticate into NVIDIA’s environment? SSO provider, account provisioning, security review timeline.
  2. Data access: Read-only API or pipeline access to (a) SFDC case data, (b) SAP relevant modules, (c) data-lake slice, (d) ASN feed, (e) Expeditors shipment events.
  3. Scope contract: Is what we build owned by NVIDIA, jointly owned, or owned by our entity and licensed to NVIDIA? (Joe Malchow’s frame argues for the third — independent operator.)
  4. Hosting + infrastructure: Where does the dashboard live — NVIDIA cloud (which one), our cloud, on-prem?
  5. Compliance posture: NVIDIA has government / defense customers and SOC 2 / FedRAMP-relevant data flows. What’s the security baseline our deployment has to meet?
  6. Customer integration: For any customer-facing feature, we need the customer to opt in. Who at the customer side signs that?
  7. CM / 3PL integration: For Wistron / Foxconn / Expeditors data flows, do they sign data-sharing agreements with NVIDIA, with us, or with both?
  8. Change management: Lonny + Alex both flagged change management as a real-world constraint. Who at NVIDIA owns the rollout, and what’s the deployment cadence — pilot team, then expansion, or company-wide?
  9. Operating-model decision: Are we vendor (SaaS subscription), services partner (FDE-style), or operator (independent data co)? The contracting vehicle follows this answer.
  10. Exit / continuity: If the engagement ends, what happens to the data, the dashboard, the operational handoff?

The questions that determine whether Phase 2 (financialization) is real

  • Does NVIDIA’s CFO org have anyone working on warranty risk transfer today? (Don’t press; listen.)
  • Does the ~90% remanufacturing accounting story have anyone curious about it?
  • Would NVIDIA share aggregated, anonymized warranty data with a third-party underwriter (Munich Re / Swiss Re / Assurant) if we operated the data layer?

§9 — Confidence summary

ClaimConfidenceBasis
The 12-step RMA workflow above is the workflow Greg/Lonny/Alex describedHighDirect evidence, three independent interviews
Pain points P1–P15 are real and namedHighDirect interview quotes
Brian Feller is currently the most plausible exec sponsorMedium-HighPublic placement announcement + scope language match
The open VP Global Service Operations req (JR1999315) reports to Shoquist and names Baxter + SAPHighNVIDIA Workday primary
Wistron Mobility Way (767K sq ft) is the likely Dallas repair lineMediumInference from size + “renovations” framing + Lonny’s July timeline; no public confirmation
Mexico repair = Guadalajara (was: Juarez more likely)HighUpdated 2026-06-26 — Greg names Guadalajara directly, flipping the prior Juarez inference [G-0626]
Compute repair = Dallas/Houston/Guadalajara; networking = Vietnam/Israel/IndiaHigh / Medium-HighDirect interview [G-0626]; Greg flagged he’d confirm networking specifics
~60-day end-to-end cycle time; 27–40 days warehouse-to-stock; +1–2 wk approval (untracked)HighDirect interview [G-0626] — upgrades the model’s prior 60-day guess
Three repair categories (reman / repair / refurb), ~90% in the first twoHighDirect interview [G-0626] — corrects the earlier 2-category split
Exec sponsor above Greg + Lonny = “Manu” (under Shoquist)Medium-HighNamed in [G-0626]; corroborated by 2026-06-25 in-person debrief; spelling/identity TBC
Warranty reserve $2.81B FY26 (NOT $8.22B)HighNVIDIA FY26 10-K primary
9% Meta failure rate scales linearly to 7.5K failures/month at 1M GPUsHighArithmetic on Meta Llama 3 primary data
OCP RAS v1.7 is hyperscaler-defined, Oct 2025HighOCP primary
NVIDIA + AMD ≈ 80% of US semi warranty reserveHigh10-K primary
The “RFP” with Palantir / Accenture / build-internal finalistsSpeculation, not statedBliss said this in debrief; Greg/Lonny did not
”Design partnership” framingSpeculation, not statedLonny in night-before chat said “design partnership” once internally to Dustin — never to Greg in the meeting
Per-RMA cost economicsUnknownNot disclosed in any source
Whether Greg/Lonny have a procurement timeline alreadyUnknownNot addressed in any conversation

Anchor transcripts: 2026-05-12-lonny-orona; 2026-05-27-nvidia-reverse-logistics-supply-chain-discussion; 2026-06-17-logistics-catch-up-wgreg-lonny-nvidia; Greg en-route call, 2026-06-26. Internal context: 2026-06-17-lonny-greg-debrief; 2026-06-17-lonny-chat. Full evidence + research: 2026-06-20-nvidia-pre-meeting-knowledge-map; nvidia-primer-2026-06-18. Earlier 1-pager (preserved): 2026-06-20-nvidia-meeting-1pager-briefing.