NVIDIA RMA Process Map
A process-first read of NVIDIA’s reverse-logistics pathway. First: the workflow itself, from things we heard directly in meetings. Then: pain points and gaps. Then: research-filled overlay with confidence labels. Goal: walk in able to ask the questions that turn this into a deployable solution.
Don’t forget (the pocket card)
If you only re-read three things before walking in, re-read these.
- Do not call this an RFP. Greg + Lonny never used that word. They said “target rich environment” + “put some bullets together” + “let me find out.” The Palantir-Accenture-build-internal triangle is our hypothesis from the debrief, not their statement. Pressure-test it; don’t assert it.
- The single highest-leverage moment is the exec-sponsor question. “To make sure we scope to the right approval path — who would own sign-off on a Phase 1 engagement of this size?” If they name Brian Feller or anyone we haven’t found in public sources, the entire engagement plan updates from that one sentence.
- The Wistron Mobility Way question is the most concretely answerable open question we have. Fort Worth is two buildings — Heritage Parkway (324K sq ft, production) and Mobility Way (767K sq ft, publicly described only as “renovations”). Ask which building, and whether the repair line is the same SOW as production. Answer changes how we think about geography of the work.
§0 — Reading guide: what’s reliable vs. what’s our guess
A pre-flight calibration so we don’t mistake our own framings for theirs.
| Source type | Examples | Treat as |
|---|---|---|
| In-meeting first-hand | Things Greg, Lonny, or Alex said in our calls | Reliable — primary evidence. Cited as [L] (Lonny 2026-05-12), [A] (Alex 2026-05-27), [G+L] (Greg + Lonny 2026-06-17), [G-0626] (Greg solo, 2026-06-26 en-route call) |
| Our debrief inference | Things Bliss + Dustin said about the meeting after the meeting | Speculative — our interpretation, not their statement. Cited as [debrief] |
| External research | Public sources, NVIDIA filings, trade press, research-agent output | Inference — useful, never primary. Always confidence-labeled |
A specific correction we owe ourselves
The “RFP” framing is ours, not theirs. Greg + Lonny did not say “we are running an RFP.” They said: “we’re such a target-rich environment”; “we just did a workshop with E&Y two weeks ago”; “how would you envision… what type of stuff would you do?”; “let me find out” (Lonny, on whether we could build it); “I’m definitely open for you guys to meet in person”; “we’d have to do an NDA”; “put some bullets together.” [G+L]
It was Bliss in the debrief who said “they’re either going to go to Palantir or they’re going to go internal or they’re going to go to Accenture” [debrief]. That is our hypothesis about the competitive shape — useful to plan against, but not something Greg or Lonny told us. Worth pressure-testing in Thursday’s meeting: Is there a formal procurement process running, or is this an executive-sponsored direct engagement?
Same calibration applies to “design partnership,” “RFP finalists,” and “Bare knuckle three-to-ten-day sprint.” Those are our framings, not their language.
§1 — The RMA workflow, as Greg + Lonny + Alex described it
This is the core process. Master this before extending to upstream / downstream.
The 12 steps, in order
Numbered by sequence. Each step has an owner, the tool(s) involved, and what crosses the seam to the next step.
Side legend: 🅝 = step NVIDIA controls. 🅒 = step customer controls. 🅒🅜 = step contract manufacturer controls. Customer- and CM-side steps are where NVIDIA’s visibility is structurally hardest.
| # | Side | Step | Who owns it | Tool / system | What flows out of this step |
|---|---|---|---|---|---|
| 1 | 🅒 | Customer detects a failed unit in their data center | Customer’s data-center ops team | Customer’s own fleet tools | A decision to open an RMA ticket [L] [A] |
| 2 | 🅒 | Customer opens RMA via NVIDIA portal | Customer (via web portal) | Salesforce-backed customer portal | A new case in NVIDIA’s SFDC [A] [G+L] |
| 3 | 🅝 | NVIDIA runs the three-gate approval gauntlet — case management (Lonny) → quality → finance (see Greg-0626 walk-through below) | Lonny’s team + quality + finance | Salesforce; failure-analysis lab for unknown codes | Validated serial #, warranty entitlement, advance-replacement-or-standard decision, and an approval notice. <1% of RMAs ever rejected [L] [A] [G-0626] |
| 4 | 🅝 | NVIDIA dispositions: standard RMA or Advance Replacement (ARMA) | Frontline + Lonny’s team | Salesforce | A ship-replacement instruction [L] |
| 5 | 🅝 | NVIDIA ships replacement to customer | Warehouse + 3PL | Expeditors (3PL, replaced Omni) | Delivery notice, tracking #, serial s, ship-to info — all communicated via email + spreadsheet, with customer reply-all chain forming here [G+L] [A] |
| 6 | 🅒 | Customer receives replacement; coordinates rack downtime | Customer DC team + customer’s business unit | Customer-side ticketing | A swap window — or “let it fail” deferral [L] |
| 7 | 🅒 | Customer physically swaps the unit | Customer DC technician | Customer-side procedure | A defective unit ready to return; 10-day window for return per NVIDIA terms [L, contractual] |
| 8 | 🅒 → 🅒🅜 | Customer wipes data + ships defective unit back to CM, ASN should fire | Customer + (sometimes) ODM integrator (Quanta etc.) | 3PL pickup (Expeditors); piloting direct pickup for one hyperscaler. ASN (Advance Shipping Notification) should fire to CM via EDI — today often email/paper [A] | A defective unit in transit + an ASN the CM may or may not get cleanly [L] [G+L] |
| 9 | 🅒🅜 | CM receives unit | Wistron or Foxconn | Receiving system; receipt acknowledgment EDI being turned on [A] | A defective unit in CM intake |
| 10 | 🅒🅜 + 🅝 | CM dispositions into one of three categories: reman / repair / refurbish (see Greg-0626 below) | Wistron / Foxconn | Should be SAP-issued production order — today often not issued before CM starts [A] | A repair job in process; compute routes Dallas/Houston/Guadalajara, networking routes Vietnam/Israel/India [G-0626] |
| 11 | 🅒🅜 | CM repairs using consignment + turnkey materials | Wistron / Foxconn | ”Suspect sheets” (early ’90s spreadsheets) tracking component consumption [A] | A repaired unit; consumption record for reconciliation |
| 12 | 🅝 | NVIDIA material manager reconciles consignment usage | NVIDIA material mgmt | Manual reconciliation today | Cleared inventory variance [A] |
| 13 | 🅒🅜 → 🅝 | Repaired unit reports back to Planning team | Wistron / Foxconn → NVIDIA Planning | Manual / spreadsheet today; SAP $2M+ planning automation deal in flight [A] | Inventory signal back into planning |
| 14 | 🅝 → 🅒🅜 | Planning gives CM the next repair commit | NVIDIA Planning team | Back-and-forth, partly manual | A pull signal for CM to procure turnkey materials and start the next batch [A] |
| 15 | 🅝 | Repaired unit enters NVIDIA’s services pool | NVIDIA warehouse | NVIDIA inventory mgmt | Like-for-like spare ready for next customer (same feature/function, different serial #) [G+L] |
| 16 | 🅝 → 🅒 | Dispatch from services pool to next customer | NVIDIA warehouse + Expeditors | 3PL handoff | Closes the loop for the next RMA cycle [G+L] |
Read this table by where the side-marker changes color: every transition between 🅝 and 🅒 or 🅒🅜 is a seam where data has to cross an organizational boundary. Five of NVIDIA’s biggest pain points (§2) sit on those seams.
The five things Alex Zhu added on top of the steps
- Consignment vs. turnkey: NVIDIA owns the high-value components (chips, boards) and parks them at CM facilities. CMs procure their own low-cost parts (cables, third-party items) — that’s “turnkey.” Two material streams, one repair line.
[A] - Repair throughput is structurally insufficient: of every 100 returns, only ~60 actually get repaired. The other 40 are filled from new inventory — “we have to skew this 40 from new buy from resin, and new buy is all what Jensen cares about because new buy is basically revenue.”
[A] - ~90% of all repairs are “remanufacturing” — ECO/known-issue batches, like a car recall. The remaining ~10% are ad-hoc. All repairs are free to the customer.
[A] - The CM↔Planning loop is the unmeasured one. “Unless I receive a signal from a planning team, how am I supposed to know what to buy?”
[A] - SAP is being paid $2M+ specifically to automate the planning capability — “eradicate and get rid of the manual Excel-based planning model.”
[A]
What Greg added on 2026-06-26 — the deepest process walk-through yet
A solo call with Greg (led by Dustin) on his drive back, walking the whole chain end to end. [G-0626] throughout. Six things that sharpen or correct the map above:
-
The “customer” is two roles, not one — and the RMA hinges on telling them apart. Who initiates the RMA and where the unit ships to are the two identifiers that matter. Google is both customer and “sold-to” (buys direct from NVIDIA). xAI is the customer but not the sold-to — Dell or Quanta buys the chips, assembles, and stands up the data center, sitting in the middle as integrator. NVIDIA has no reliable directory of who to talk to on the customer side (“emails going to people that have left the company… it’s not a systemic, robust process”). Read steps 1–2 as customer / sold-to / integrator, not one node.
-
A three-gate approval gauntlet runs before the unit ever ships to a warehouse — this is what the table compresses into step 3. Case filed → case management (Lonny’s team): warranty check, failure-log review, serial-number lookup in Salesforce → quality team: known failure code = approve, unknown = route to a failure-analysis lab → finance: approves → approval notice to customer → customer ships to a warehouse. Less than 1% of RMAs have ever been rejected — the gauntlet is almost pure latency, not a filter. Greg’s own opportunity framing: high-volume known-defect batches (“we go to Microsoft and say everything we shipped you between these dates has this problem — 6,000 units, 6,000 serials”) should be auto-fast-tracked straight through, not walked gate by gate. Greg doesn’t own the approval criteria — Lonny does — so the detail behind each gate is a direct question for Lonny.
-
Three repair categories, not two (this corrects the “~90% reman / ~10% ad-hoc” split in the Alex list above):
- Remanufacturing (reman) — current-revision product, known defect + known fix; high volume; repaired on the same live mass-production line that built the original unit.
- Repair (sort & repair) — N-minus-one (last year’s revision); also high volume; same site, separate dedicated repair line (keeps the current-revision production line undisturbed while the build team’s knowledge is still fresh).
- Refurbishment — one-off, unknown condition, diagnosed on arrival; may be dismantled and salvaged for parts; the only true “repair-shop” path.
- ~90% is reman + repair combined (the first two); refurb is the tail. Greg: reman and repair “are really recalls — known flaws”; refurb is the onesie-twosie diagnostic path. Routing rule: whoever built it usually repairs it.
-
Repair geography is (at least) two separate networks. Compute (GPUs, trays, baseboard assemblies, SXM modules) routes to Dallas, Houston, or Guadalajara — all Foxconn/Wistron. Networking (switches — no GPUs on them) routes entirely separately: Vietnam, Israel (post-Mellanox), India. European compute still ships back to the US — no proximate compute repair capability there. A little reman runs in Taiwan, but most reman + refurb is Dallas / Houston / Guadalajara. This flips our earlier inference that the Mexico endpoint was Juarez (Wistron/Wiwynn) — Greg names Guadalajara; see §5.
-
Cycle time, with real numbers for the first time. ~60 days end-to-end today (Greg’s estimate); ~30 days best case. Of that, 27–40 days is warehouse-receipt → repaired-unit-back-in-stock — the only stretch they track weekly. The approval gauntlet adds another 1–2 weeks that isn’t tracked at all, because NVIDIA only starts the clock at warehouse arrival, not at case submission. Two clocks, decoupled: the customer gets a replacement from stock in ~1–2 weeks (if inventory exists), while their own serial number takes 30+ days to repair and return to stock. Greg wants to start measuring both the customer-experienced time (case-submit → replacement received) and the serial-level throughput time (receipt → good inventory). The two metrics he cares about — throughput and cycle time — and he wants them translated to financial impact. (This interview-confirms the ~60-day baseline our economic model uses.)
-
The two revenue-vs-repair tensions, confirmed and sharpened. Greg named exactly two (matching the model): (a) reman-line capacity — repair competes with new production on the same assembly line; when constrained it’s a joint, negotiated decision between NVIDIA’s service org and the revenue side (revenue wins ~9/10, “but they shouldn’t be making that decision on their own”). NVIDIA’s repair-side volume is “minuscule” next to revenue, so it combines forces with the revenue org to hold the CMs commercially accountable for both. (b) New-buy draw — when repaired stock is short, warranty gets filled from new units; “we try to limit the number of new buys in our service depots,” and GPUs especially are fought over (“there’s never enough”). This is the operational face of Alex’s 60-of-100 number and the model’s “replacement = lost sale at $P$.”
Logistics footnote: the Dallas ↔ Guadalajara cross-border move + customs clearance is the named queue-time bottleneck. Omni (bought by Forward Air) is being exited in the US — leaving Omni Union City and Omni Dallas, consolidating to Expeditors Dallas for US returns + fulfillment; Omni may be retained in Asia (Hong Kong / Taiwan), still debated. Speed is the dominant lever — Greg: “if we could fly these all by private jet, we would.”
Where Greg framed the problem
The frame Greg + Lonny used most in the meeting:
- “We don’t even really know the score of the game right now.” They cannot measure their own SLA performance against the 30-day standard.
[G+L] - “We come to the conversation, and we’re not even armed.” When a customer escalates, NVIDIA has no shared view of the case to argue from.
[G+L] - “What’s the source of truth? Which dashboard do I believe?” Lonny’s frame: many incremental dashboards, no end-to-end system architecture, no agreed-on data.
[G+L] - “We’re such a target-rich environment.” Greg’s framing — many places to start.
[G+L]
§2 — Where the workflow breaks (from interviews only)
Each item below is a failure point Greg / Lonny / Alex named directly. Not inferred.
Workflow-level pain points
| # | Pain point | Where it lives | Source |
|---|---|---|---|
| P1 | Post-ticket spreadsheet cascade — DNs, tracking numbers, serial numbers, ship-to info all updated manually after the SFDC ticket is opened. Customer reply-all threads balloon to 16+ people; NVIDIA’s version and customer’s version diverge. | Step 5 → step 9 | [G+L] |
| P2 | No automated ASN to CM — warehouse expects shipment notification; today it’s email/paper. | Step 8 → step 9 | [A] |
| P3 | No production order issued before CM begins repair — CMs work without a system-of-record SAP PO. | Step 9 → step 10 | [A] |
| P4 | ”Suspect sheets” at CMs — component consumption tracked on ‘90s-style spreadsheets. | Step 11 | [A] |
| P5 | Manual material reconciliation — material manager has to check CM consumption claims; risk of over-reporting. | Step 12 | [A] |
| P6 | CM ↔ Planning back-and-forth is partly manual — repair commits not flowing as system signals. | Step 13 → step 14 | [A] |
| P7 | 30-day SLA is not measured — “we don’t know the score of the game.” | Step 4 → step 16 (the whole arc) | [G+L] |
| P8 | No escalation triggers — “day 31 should fire a notification; today nothing does.” | Spans the whole arc | [G+L] |
| P9 | Customer business-unit approval bottleneck — DC teams can’t unilaterally take a rack down; BUs (Instagram, Facebook) prefer “let it fail” → ARMA units sit idle for weeks/months → inventory planning distortion. | Step 6 | [L] |
| P10 | ODM integrators (Quanta etc.) add touches without value on returns — equipment gets handled multiple times waiting for carriers; NVIDIA piloting direct pickup with one hyperscaler. | Step 8 | [L] |
| P11 | Repair throughput insufficient (~60/100) — gap is filled from new inventory, which is what Jensen “cares about.” | Step 10 → step 13 | [A] |
| P12 | Telemetry walled off at customer fence line — chips wiped before return; first signal of product perf is ticket volume. | Step 1 → step 2 | [G+L] |
| P13 | Stranded shipments — Phoenix five-truck incident: NVIDIA sent five trucks of “hundreds of millions of dollars of equipment” to customer site without coordination; trucks turned around to a security yard in the Phoenix sun. | Step 5 / step 8 — coordination layer | [G+L] |
| P14 | Hyperscale form-filling is operationally hostile — “you don’t want xAI or OpenAI to fill the form.” | Step 2 | [A] |
| P15 | Mellanox / Israel team integration gap — “still operate like they’re separate companies.” | Org-level — touches step 3 | [L] mentioned in May, never repeated in June — may be stale; treat as worth confirming, not as a current pain |
| P23 | Approval gauntlet is pure latency — case mgmt → quality → finance adds ~1–2 weeks before the unit even ships, yet <1% of RMAs are ever rejected. Known-defect batches walk every gate instead of being auto-fast-tracked. And the clock isn’t started until warehouse arrival, so this delay is invisible in the metrics. | Step 3 (pre-warehouse) | [G-0626] |
How the pain points cluster
Three clusters, not 15 independent things:
- No system-of-record after the SFDC ticket. P1, P2, P3, P4, P5, P6 — every handoff from step 5 onward leaks into email + spreadsheets.
- Nobody measures the loop. P7, P8, P13 — there’s no end-to-end SLA dashboard, no escalation logic, no early-warning on stranded shipments.
- The customer side is opaque. P9, P10, P12, P14 — NVIDIA can’t see what customers are doing with units, can’t pull telemetry, can’t make customers fill forms cleanly, and can’t make customer BUs authorize swaps.
The dashboard Lonny + Greg sketched at the end of the meeting maps to cluster 2 first. Cluster 1 is what makes the dashboard possible (data has to be in a system to be dashboarded). Cluster 3 is where the bigger ambition lives.
§3 — What we don’t know about the systems and the process (interview gaps)
Strictly: things we cannot answer from what Greg / Lonny / Alex said.
Systems / tooling we don’t have a clear picture of
| Topic | What’s unclear |
|---|---|
| Salesforce | Edition (Enterprise, Unlimited, Industries Cloud)? Service Cloud or custom? What custom objects exist? Lightning? |
| SAP | ECC or S/4HANA? On-prem, RISE, or HEC? What modules — SD, MM, PP, EWM? What is the $2M+ deal actually covering? |
| Baxter Planning | Which modules? How does it sync with SAP? Does it consume Salesforce data? Does it consume CM repair data? |
| Data lake | What platform — Databricks, Snowflake, NVIDIA-internal? What data is already there? Who has access? |
| WMS / TMS being rolled out | Vendor names? In-house or partner systems being extended into 3PL? |
| EDI | What platform? Who’s the integration partner (the EDI-with-couple-of-customers piece)? |
| Identity / SSO | How would a third-party developer authenticate? Okta? Azure AD? |
| API surface | What APIs already exist that we could call? What would we need to build? |
| The EY workshop output | Beyond “two initiatives,” what did EY architecturally produce — a process map, a tooling proposal, a vendor short-list? |
| CM-side systems | What do Wistron and Foxconn actually use for repair tracking today, beyond “suspect sheets”? Is there any system at all? |
| 3PL integration | How does Expeditors plug into SAP / Salesforce today? |
| The customer portal | Is it on Salesforce Experience Cloud, or a custom front-end on top of SFDC APIs? |
Process facts we don’t have
| Topic | What’s unclear |
|---|---|
| RMA volume | How many RMAs / month? / quarter? / year? Lonny said “hundreds today, thousands soon” but not a number. |
| Active RMAs at any moment | The thing the dashboard would put on its front page — never quantified. |
| Cycle time today | [G-0626] |
| Stranded shipment frequency | One Phoenix event. Monthly? Quarterly? One-off? |
| Spreadsheet thread count | How many parallel email/spreadsheet threads are running right now across the org? |
| CM consignment value | How many dollars of NVIDIA-owned chips sit at Wistron + Foxconn at any given time? |
| ARMA idle time | Lonny said “weeks or months.” What’s the mean? What’s the cost? |
| Per-customer SLA terms | Do hyperscalers have custom contracts that tighten the 30 days? |
| Custom escalation paths | What contractual escalation rights does e.g. Meta have? |
| Repair cost per unit | Per-SKU economics — never quantified. |
| Stranded inventory cost | Working capital tied up in stranded equipment — never quantified. |
| Failure-mode Pareto | Greg cited bent connector pins, heat sinks, backplane, network switches. Which dominates by frequency? By cost? |
| Geography of the work | [G-0626] |
| Who else inside NVIDIA touches this | IT, Sales account teams, field engineers, Mellanox team, quality / FA team — how do they intersect with Lonny + Greg’s pillar? |
| The second EY workshop initiative | One disclosed (the portal). The other — not yet. |
Org / authority gaps
- Who is the exec sponsor above Greg + Lonny? Named 2026-06-26: “Manu.” Greg + Lonny are peers; both report to Manu, who heads the service-supply-chain org (five functions: customer service, reverse logistics + warehousing, repair operations, planning, and Greg as the horizontal “fifth wheel”). Corroborated by the 2026-06-25 in-person debrief (Manu sits under Deb Shoquist). Spelling / LinkedIn identity still being confirmed via Ali.
[G-0626] - What’s Greg’s actual budget authority? What size deal can he sign vs. needing the exec sponsor’s approval?
- What’s the timeline they’re operating on?
- Does Alex Zhu’s reverse-supply-chain transformation sit inside the same program as Greg + Lonny’s, or parallel to it?
§4 — The extended pathway (upstream + downstream), interview-anchored
The full chain widens beyond the RMA workflow. Greg + Lonny + Alex touched on the wings — less detail than the core, but it matters for understanding where their problem ends and where ours might start.
Upstream — failure detection at the customer (step 0, before step 1)
What they told us directly:
- “You’re not gonna get telemetry off of their devices.” Customer data is highly confidential. Chips are wiped before return.
[G+L] - First signal of product performance is ticket volume vs. install base. Any two data points start a trend.
[G+L] - Logs come reactively after fault — never proactively.
[G+L] - Greg’s prior gig (likely Infinera / photonics) ran a crawl program that scanned customer networks, predicted failures within ~90 days. Customers usually picked “let it fail + 4-hour SLA” over coordinating maintenance. But NVIDIA still got to pre-position spares.
[G+L] - Lonny in the night-before chat softened this: “it’s not that the data doesn’t exist; it’s just inaccessible.” Some customers will share metrics selectively for triage.
[LonnyChat]
What we don’t know: which customers will share what, through what mechanism. Whether NVIDIA has a triage-data program in flight. How customer-side fleet tools (Meta’s, Microsoft’s) map to NVIDIA’s case-open trigger.
Downstream — planning, services pool, and demand forecasting (steps 13–16+)
What they told us directly:
- Baxter Planning does demand planning and forecasts both sales and failure rates.
[L] - Planning ↔ CM signaling is partly manual today. SAP $2M+ deal is automating it.
[A] - Services pool feeds the next ARMA cycle. Like-for-like (same feature/function, different serial #).
[G+L] - Repair throughput insufficiency spills back into new-inventory consumption — economic conflict with new-buy revenue.
[A]
What we don’t know: how Baxter actually models failure rates; what its inputs are; whether the SAP $2M deal touches Baxter or is upstream of it; the actual policy on when planning pulls from new inventory vs. waits for repair.
The four pillars — was this our impression or a real org chart?
Lonny in May described four pillars: dedicated repair lines, reverse logistics, demand planning (forecasts failure rates), systems group (automation/tooling). [L]
Greg in June described two EY workshop initiatives — the portal is one. [G+L]
We don’t actually know if these are the same program or different programs. This is a worth-asking question.
§5 — The RMA workflow with research overlay
Same 12-step process. Each gap from §3 backfilled where research can; confidence labels on every fill. Read this as: if the meeting confirms or denies these inferences, what do we update?
Geography of the work
| Step | Interview said | Research adds (confidence) |
|---|---|---|
| 5 / 8 / 16 | Expeditors (3PL) replaced Omni | Expeditors is global, 340+ locations, non-asset-based; specialty in semiconductor reverse logistics. High [Expeditors corporate] |
| 9–11 | Compute repair = Dallas, Houston, Guadalajara (Wistron + Foxconn); back-office HK; warehouse Taiwan [G-0626] | Wistron Fort Worth is two buildings: 15200 Heritage Parkway (324K sq ft, $580M, primary, production); 14601 Mobility Way (767K sq ft, $181M, secondary, publicly described only as “renovations”). Mobility Way is 3x larger than the production primary site — strongest candidate for Lonny’s “repair line.” Medium confidence on the inference (no public source confirms repair). Greg-0626 now confirms Houston as a compute repair node — promotes the earlier “Foxconn Houston” speculation. [Fort Worth Report 2025-08-21; Hillwood; Dallas Innovates] |
| 9–11 | Mexico repair routes = Guadalajara (Greg-0626, explicit) | Updated 2026-06-26: Greg names Guadalajara, not Juarez. This flips our prior Medium-confidence inference (which favored Wistron/Wiwynn Juarez on the $23M warehouse-lease profile). Foxconn Guadalajara (GB200 megaplant, $500M–$900M, 240K servers/yr) is now the stated compute repair endpoint; the Juarez warehouse may still coexist as a logistics node. Now interview-anchored, High on Guadalajara as the Mexico repair site. [G-0626; Taipei Times 2025-05-09; DigiTimes; Mexico News Daily] |
| 9–11 | Networking repair = Vietnam, Israel, India (separate from compute) [G-0626] | Switches carry no GPUs and route through their own network: Vietnam, Israel (post-Mellanox), and India (likely Cumulus/Mellanox legacy). European compute still ships back to the US — no proximate compute repair in Europe. Interview-anchored, Medium-High (Greg flagged he’d confirm the networking specifics). |
| 9–11 | Asia footprint | No public source names a specific NVIDIA repair facility in HK or Taiwan. Wistron Hsinchu + Hukou (production); Foxconn-NVIDIA Taiwan supercomputing cluster $1.4B H1 2026 (production). Greg-0626: a little reman runs in Taiwan; most reman + refurb is Dallas/Houston/Guadalajara. Reverse-flow specifics still externally invisible. Low confidence on identifying Asia repair facilities specifically. |
Volumes that backfill the “hundreds → thousands” framing
| Step | Interview said | Research adds (confidence) |
|---|---|---|
| 1 | Meta has 100K GPUs; wants 1M in 5 yrs | Meta Llama 3 published data: 16,384 H100 cluster, 54 days, 466 interruptions, ~78% hardware. ~9% annualized failure rate. High [Meta Llama 3 paper, 2024] |
| 1 | ”Hundreds today, thousands soon” | At 9% × 100K = ~9K failures/yr ≈ 750/month. At 1M Meta GPUs ≈ 7.5K/month. Lonny’s math holds. High (arithmetic). |
| 2 / 8 | Customer-side telemetry walled off | Meta runs three detection systems: Fleetscanner (45–60 day cycles), Ripple (alongside live workloads), Hardware Sentinel (kernel-space exception analysis, outperforms test-based by 41%). >66% of training interruptions are SRAMs, HBMs, network switches. High [Meta Engineering Blog Jul 2025; ASPLOS 2025]. Meta has the data Greg + Lonny said they can’t see. |
Failure modes that backfill Greg’s anecdotes
| Step | Interview said | Research adds (confidence) |
|---|---|---|
| 10 | ”Chip is rarely the problem” — connector pins, heat sinks, backplane, network switches | GPU faults 30.1% of Meta interruptions; HBM3 17.2% [Meta Llama 3]. CoWoS-L thermal / CTE mismatch at 1400W Blackwell TDP is a confirmed driver — warping, HBM PHY microbump failures can render the entire chip inoperable. Once CoWoS-bonded, individual chiplets and HBM stacks can’t be replaced. Medium-High [SemiEngineering; Chiplet Summit 2025; proteanTecs] |
| 1 | Hyperscalers wrote OCP RAS spec — Lonny didn’t name it | OCP GPU & Accelerator RAS Requirements v1.7, published 2025-10-23, standardizes error reporting, crash dumps, RCA, error containment, Redfish/IPMI SEL/APEI formats. Customers are setting the serviceability bar NVIDIA’s hardware has to meet. High [OCP RAS v1.7] |
Systems that backfill the stack
| Topic | Research adds (confidence) |
|---|---|
| Baxter Planning | Founded 1993 by an ex-Texas Instruments service-parts planner; customers manage $11B+ inventory across 35K locations, 120 countries; AI-powered “BaxterPredict.” Marlin took majority 2024. High [Baxter Planning corporate; PRNewswire/Marlin 2024]. Strong fit with Lonny’s description. |
| NVIDIA Enterprise Support tiers | Business Standard (4-hr Sev 1, 8x5 live) and Business Critical (1-hr Sev 1, 24/7) are the two named tiers. Premium TAM is mandatory for every DGX SuperPOD. High [NVIDIA Enterprise Support Policy 2025-05-05 PDF] |
| NVIDIA AI Enterprise (the subscription that bundles support) | ~$4,500/GPU/yr list, 1-year; ~$22,500/GPU 5-year; $2.00/GPU/hour PAYG via cloud marketplace. High [Dell APD ac566091; Insight; NVIDIA marketplace] |
| Mellanox Global Expedite RMA | Silver / Gold / Platinum, with 4-Hour Expedite RMA as an explicit upgrade option. High [network.nvidia.com Mellanox RMA PDF] |
Who runs this org — the org gap research can partly fill
| Topic | Research adds (confidence) |
|---|---|
| Likely exec sponsor today | Brian Feller, VP Global Planning, Logistics & Services — joined NVIDIA May 2021 from Dell (19 yr). His own stated scope: “his team includes a growing Services organization to support reverse logistics, repair, and RMA fulfillment.” Round Rock, TX. Medium-High [ON Partners placement announcement; theorg.com; LinkedIn] |
| Critical wrinkle | NVIDIA has an OPEN public Workday req — VP, Global Service Operations (JR1999315) — explicitly names Baxter and SAP and reports to “EVP of Global Operations” (Shoquist). $352K–$558K base. The literal seat above Greg + Lonny may not be filled. Either Feller is being elevated or Services is being split out under a new VP. High that the req exists. Unknown if filled. [NVIDIA Workday JR1999315; theladders.com] |
| Above that layer | Debora Shoquist (EVP Operations) — scope includes supplier mgmt, CM mgmt, supply planning, logistics, quality. Greg + Lonny’s org rolls up here. High [NVIDIA Newsroom bio] |
| Alex’s “VP” | Most plausibly Brian Feller (scope match). Medium-High |
| Trivedi successor (EVP Enterprise Sales) | Not publicly named. Trivedi retired April 2026; now on Enphase board (June 2026). Possibly absorbed by Jay Puri or still unfilled. Unknown [SDxCentral; GlobeNewswire 2026-06-15] |
The economics that interview data alone doesn’t surface
| Topic | Research adds (confidence) |
|---|---|
| NVIDIA warranty reserve | $2.81B FY26 — not the widely-quoted $8.22B (WarrantyWeek mis-aggregation). Accrual rate 0.46% → 0.92% → 1.15% in 2 years. Claims paid jumped to $957M (+337% YoY). Attribution per 10-K footnote: “primarily Compute & Networking.” High [NVIDIA FY26 10-K] |
| The “all repairs free” + “~90% remanufacturing” puzzle | Under ASC 460/450-20, recall reserves typically get separately disclosed when probable + estimable. NVIDIA’s 10-K does not break out recall-related charges. Two possible reads, both speculative: (a) ECO repairs treated as routine warranty (softens income-statement signal), or (b) the rising accrual rate IS the recall disclosure, dressed as warranty. No sell-side analyst has flagged this. It’s the strongest balance-sheet hook for a Phase 2 financialization angle. Medium confidence in the framing; High confidence that this is publicly invisible. |
| AMD comparison | AMD warranty reserve $308M FY25 (vs. NVIDIA’s $2.81B) — mirror curve, ~10x smaller. High [AMD FY25 10-K]. AMD does not segment data-center vs. client warranty. |
| Intel / Broadcom / Marvell | Disclose nothing material in warranty reserves. NVIDIA + AMD ≈ 80% of US semi-industry warranty reserve; NVIDIA alone ~74%. High [respective 10-Ks; WarrantyWeek 23rd Annual Report 2026-04-16] |
NVIDIA’s existing forward-fleet products (where our wedge does not go)
NVIDIA already operates a forward fleet-management layer. We sit downstream of where it stops:
- Mission Control (GA on Blackwell) — AI factory infra mgmt; provisioning; monitoring; error diagnosis.
- Run:ai (acquired Dec 30 2024 for ~$700M) — GPU workload orchestration; folded into Mission Control.
- Base Command Manager (ex-Bright Computing) — cluster provisioning.
- Fleet Command, Fleet Intelligence — edge / fleet telemetry.
The boundary: Mission Control + Run:ai is what runs while a unit is alive in a customer’s data center. Lonny’s portal is what NVIDIA needs after a unit fails. They don’t overlap because customers strip telemetry — Mission Control’s signals stop at the customer’s fence line. Medium-High [NVIDIA docs; cross-ref [[nvidia-primer-2026-06-18]]]
§6 — Where the workflow breaks, with research-overlaid pain points
Same three clusters from §2; research surfaces new pain points the interviews didn’t.
Cluster 1: No system-of-record after the SFDC ticket
The interview pain points (P1–P6) stand. New from research:
- P16. The CM “suspect sheet” failure mode is industry-typical. Public reporting on Wistron and Foxconn reverse-logistics IT maturity is sparse, suggesting this is not a NVIDIA-specific gap but a common floor across high-volume CM ops. Medium confidence based on absence of public Wistron/Foxconn repair-mgmt-platform disclosures.
- P17. SAP $2M+ planning automation may sit alongside, not under, the dashboard. Greg explicitly described his Cisco-era job as moving away from siloed ERP — a tension with the active SAP procurement. Medium speculation; worth pressure-testing.
Cluster 2: Nobody measures the loop
The interview pain points (P7, P8, P13) stand. New from research:
- P18. OCP RAS v1.7 (Oct 2025) is hyperscaler-defined, not NVIDIA-defined. The customers wrote the spec for what NVIDIA’s serviceability should look like. Any dashboard NVIDIA buys probably has to meet hyperscaler-side data expectations to be acceptable. High that the spec exists.
- P19. The 30-day SLA is publicly visible only in DGX support docs; whether hyperscalers have tighter custom SLAs is opaque externally. Unknown.
Cluster 3: The customer side is opaque
The interview pain points (P9, P10, P12, P14) stand. New from research:
- P20. Meta has the telemetry data NVIDIA says it can’t get. Hardware Sentinel runs entirely customer-side, beats test-based methods 41%. The data exists; whether Meta would share it (selectively, for triage) is a relationship question, not a technology question. High that Meta has it; Unknown on sharing.
- P21. The China RMA collapse has produced no publicly named replacement geography. Gray-market Shenzhen repair shops fill the gap (~500 units/month per shop, $1,400–$2,800/GPU). High that the problem exists; Unknown if NVIDIA is doing anything about it inside the customer-service ops scope. Per CLAUDE.md, surface only if Lonny raises it.
New organizational pain point research alone surfaces
- P22. The org is mid-restructure. The open VP, Global Service Operations req (JR1999315) sits between Brian Feller (whose scope already says “growing Services organization”) and the EVP layer. This is a layer in flight. Engagement timing has to account for the possibility that the exec sponsor changes during the engagement.
§7 — Questions for the meeting
The whole point of this doc. Ordered by the answer’s leverage on whether we can deploy a solution.
Tier 1 — Authority, coordination, and timeline (must answer)
- “Who would own sign-off on a Phase 1 engagement of this size? Is that you, or does it route through Brian’s org?” The single highest-leverage question. If they name someone we haven’t found in public sources, that’s the person who matters.
- “How does Alex Zhu’s reverse-supply-chain transformation work intersect with what you’re doing? Same program, parallel program, or different problem?” Alex was our connector to Greg + Lonny but we don’t know if his project overlaps. If it overlaps, we walk into a coordination problem blind. If it’s parallel, the buying centers may be different.
- “How does this engagement get contracted — MSA + SOW, vendor procurement, or something else?” Reveals whether this is structured procurement (the RFP we inferred) or direct exec engagement.
- “What’s the timeline you’re operating on? When does a decision need to happen?” Tighter than 4 weeks → accelerate prototype. Longer than 8 weeks → embed and shape the spec.
Tier 2 — Scope and success (must answer)
- “Of the two EY workshop initiatives, what’s the second one? Is it a sibling portal effort or something internal?” Tells us what we integrate with vs. compete against.
- “What systems do you want this to talk to on day 1 vs. day 90?” SFDC, SAP (which modules / version), Baxter, data lake, Expeditors, CM systems, customer-side EDI.
- “What’s the success metric at 6 months? At 12 months?” SLA attainment measurable for the first time? Cycle-time reduction? Stranded-shipment events to zero?
- “What’s the data-lake situation today? What’s in it that’s reliable, and who owns the pipelines?” Affects whether we build on top or alongside.
Tier 3 — The process and the failure modes (worth asking)
- “Walk me through a real recent RMA — say, the Phoenix five-truck incident. What happened in each system at each step?” Surfaces the actual workflow vs. the conceptual one.
- “On the CM side, what’s the current playbook — do Wistron and Foxconn use any system at all today beyond the suspect sheets?” Tells us whether B2B integration is greenfield or a brownfield upgrade.
- “How do escalations actually happen today? Who at the customer calls who at NVIDIA?” Maps the human network we replace or augment.
- “Of the failure modes you mentioned — connector pins, heat sinks, backplane, network switches — is there a Pareto? Which dominates by frequency? By dollar?” Calibrates where the wedge sits and what failure-mode-aware tooling would actually pay back.
- “~60 of 100 repair rate — what’s the bottleneck? Capacity? Component availability? Diagnosis time?” Different bottlenecks imply different solutions.
- “What’s the EDI integration like with the couple of large customers who already have it? What data flows?” A template we could extend.
Tier 4 — Geography and physical footprint (worth asking)
- “Dallas July 2026 — production, repair, or both? Which Wistron building? Heritage Parkway or Mobility Way?” The most concretely answerable open question we have.
- “When you say Mexico, is that Juarez (Wistron/Wiwynn) or Guadalajara (Foxconn)?”
- “Asia footprint — which Hong Kong office, which Taiwan warehouse, specifically?”
Tier 5 — Customer-side (worth asking)
- “Which hyperscalers do you most want involved in the discovery sprint? Whose pain do we want this to reflect first?”
- “Are any customers willing to share triage data today, selectively, for fault diagnosis?” Lonny softened on this in the night-before chat — worth probing.
- “Are you aligned with OCP RAS v1.7? Is that shaping what you’re scoping?”
Tier 6 — Economics (listen, don’t press)
- “~90% remanufacturing — how is that accounted for? Standard warranty or treated as something different?” Listen for signal; don’t push. The answer’s value is the Phase 2 financialization signal, not the Phase 1 dashboard.
§8 — What we’d need to know in order to actually deploy a solution
The bridge from “understanding their problem” to “shipping the dashboard.” Each of these is something we don’t know yet and would need to know before writing a real SOW.
The minimum-viable deployment requires answers to
- Identity + access: How does a third-party developer authenticate into NVIDIA’s environment? SSO provider, account provisioning, security review timeline.
- Data access: Read-only API or pipeline access to (a) SFDC case data, (b) SAP relevant modules, (c) data-lake slice, (d) ASN feed, (e) Expeditors shipment events.
- Scope contract: Is what we build owned by NVIDIA, jointly owned, or owned by our entity and licensed to NVIDIA? (Joe Malchow’s frame argues for the third — independent operator.)
- Hosting + infrastructure: Where does the dashboard live — NVIDIA cloud (which one), our cloud, on-prem?
- Compliance posture: NVIDIA has government / defense customers and SOC 2 / FedRAMP-relevant data flows. What’s the security baseline our deployment has to meet?
- Customer integration: For any customer-facing feature, we need the customer to opt in. Who at the customer side signs that?
- CM / 3PL integration: For Wistron / Foxconn / Expeditors data flows, do they sign data-sharing agreements with NVIDIA, with us, or with both?
- Change management: Lonny + Alex both flagged change management as a real-world constraint. Who at NVIDIA owns the rollout, and what’s the deployment cadence — pilot team, then expansion, or company-wide?
- Operating-model decision: Are we vendor (SaaS subscription), services partner (FDE-style), or operator (independent data co)? The contracting vehicle follows this answer.
- Exit / continuity: If the engagement ends, what happens to the data, the dashboard, the operational handoff?
The questions that determine whether Phase 2 (financialization) is real
- Does NVIDIA’s CFO org have anyone working on warranty risk transfer today? (Don’t press; listen.)
- Does the ~90% remanufacturing accounting story have anyone curious about it?
- Would NVIDIA share aggregated, anonymized warranty data with a third-party underwriter (Munich Re / Swiss Re / Assurant) if we operated the data layer?
§9 — Confidence summary
| Claim | Confidence | Basis |
|---|---|---|
| The 12-step RMA workflow above is the workflow Greg/Lonny/Alex described | High | Direct evidence, three independent interviews |
| Pain points P1–P15 are real and named | High | Direct interview quotes |
| Brian Feller is currently the most plausible exec sponsor | Medium-High | Public placement announcement + scope language match |
| The open VP Global Service Operations req (JR1999315) reports to Shoquist and names Baxter + SAP | High | NVIDIA Workday primary |
| Wistron Mobility Way (767K sq ft) is the likely Dallas repair line | Medium | Inference from size + “renovations” framing + Lonny’s July timeline; no public confirmation |
| Mexico repair = Guadalajara (was: Juarez more likely) | High | Updated 2026-06-26 — Greg names Guadalajara directly, flipping the prior Juarez inference [G-0626] |
| Compute repair = Dallas/Houston/Guadalajara; networking = Vietnam/Israel/India | High / Medium-High | Direct interview [G-0626]; Greg flagged he’d confirm networking specifics |
| ~60-day end-to-end cycle time; 27–40 days warehouse-to-stock; +1–2 wk approval (untracked) | High | Direct interview [G-0626] — upgrades the model’s prior 60-day guess |
| Three repair categories (reman / repair / refurb), ~90% in the first two | High | Direct interview [G-0626] — corrects the earlier 2-category split |
| Exec sponsor above Greg + Lonny = “Manu” (under Shoquist) | Medium-High | Named in [G-0626]; corroborated by 2026-06-25 in-person debrief; spelling/identity TBC |
| Warranty reserve $2.81B FY26 (NOT $8.22B) | High | NVIDIA FY26 10-K primary |
| 9% Meta failure rate scales linearly to 7.5K failures/month at 1M GPUs | High | Arithmetic on Meta Llama 3 primary data |
| OCP RAS v1.7 is hyperscaler-defined, Oct 2025 | High | OCP primary |
| NVIDIA + AMD ≈ 80% of US semi warranty reserve | High | 10-K primary |
| The “RFP” with Palantir / Accenture / build-internal finalists | Speculation, not stated | Bliss said this in debrief; Greg/Lonny did not |
| ”Design partnership” framing | Speculation, not stated | Lonny in night-before chat said “design partnership” once internally to Dustin — never to Greg in the meeting |
| Per-RMA cost economics | Unknown | Not disclosed in any source |
| Whether Greg/Lonny have a procurement timeline already | Unknown | Not addressed in any conversation |
Anchor transcripts: 2026-05-12-lonny-orona; 2026-05-27-nvidia-reverse-logistics-supply-chain-discussion; 2026-06-17-logistics-catch-up-wgreg-lonny-nvidia; Greg en-route call, 2026-06-26. Internal context: 2026-06-17-lonny-greg-debrief; 2026-06-17-lonny-chat. Full evidence + research: 2026-06-20-nvidia-pre-meeting-knowledge-map; nvidia-primer-2026-06-18. Earlier 1-pager (preserved): 2026-06-20-nvidia-meeting-1pager-briefing.