NVIDIA Knowledge Map — Pre-Meeting Synthesis

Purpose: stitch together every claim, fact, name, number, and process detail captured across our three NVIDIA conversations into one map. First what we know (transcript-anchored). Then what we don’t know (explicit gaps). Then what’s likely true (research-filled with confidence labels). Built per RDI methodology — surfaces evidence, does not conclude.

How to read this doc

  • §1 What We Know — every claim comes from one of the three transcripts; cited inline.
  • §2 What We Don’t Know — gaps the three transcripts leave open, organized by topic.
  • §3 What’s Likely True (Research-Filled) — best-available answers to the §2 gaps, with confidence tags. Do not treat §3 claims as conversation-anchored facts.
  • §4 Top Questions for the Meeting — prioritized.
  • §5 Surprises & Tensions — per RDI methodology.

Source shorthand:


§1 — What We Know (transcript-anchored)

A. The people

Lonny Orona — Customer Service Operations lead at NVIDIA. Eight weeks in as of May 12 2026 [Lonny-0512]. Previously 8 years at Meta in network infrastructure, rising from managing in-fence fiber and connectivity → all capex/opex for global data center builds → leading a $20M efficiency program targeting $40M in year two [Lonny-0512]. Left Meta frustrated with repeated layoff cycles (“same tagline about efficiency, but didn’t we do that last time?”) and what he called organizational immaturity / cultural erosion [Lonny-0512]. Frames his value at NVIDIA as someone who can “sit on either side of the table” — having been both a hyperscale customer of NVIDIA and now a hyperscale supplier [Lonny-0512]. His team handles compute science frontline support — first call from customers, warranty and entitlement validation, advance replacement decisions [Lonny-0512].

Greg Dalcello — Senior Director, Service Team. 22 years at Cisco in high-tech supply chain (factory management, fulfillment, outsourced contract manufacturing); 2.5 years at MyNextPower (solar farm builder, Fremont) managing steel and dry-drop logistics; now at NVIDIA driving operations transformation [Greg+Lonny-0617]. Background business/management, not engineering — “got the only job I could” after business school in ‘92 [Greg+Lonny-0617]. Got into transformation when his Cisco boss put him on a global ERP rollout for 18 months — “I just learned a ton” [Greg+Lonny-0617]. ~15 years in transformation/enablement portfolio management [Greg+Lonny-0617]. Per Lonny in May: “one week ahead of Lonny in the role” [Lonny-0512] — so Greg also joined NVIDIA early-to-mid 2026.

Alex Zhu — 2 years at NVIDIA; previously ~2 years at Apple; before that 8-9 years in consulting [Alex-0527]. Cupertino-based. Graduated 2012. Works with a VP at NVIDIA on reverse supply chain operations — RMA, ASN, warehouse mgmt, planning optimization, repair operations, dashboards/KPIs [Alex-0527]. Confirmed via LinkedIn: Penn State Supply Chain & Information Systems background; prior SAP SD (Sales & Distribution) implementation experience [briefing-alex-zhu-2026-05-27-final]. Alex was the person who connected Greg and Lonny [Greg+Lonny-0617] — he is a cross-organizational node.

The unnamed VP that Alex works under. Alex referenced “a VP here” who oversees the reverse supply chain transformation he’s working on [Alex-0527]. Not named in any conversation.

The exec sponsor above Greg + Lonny — now named: “Manu.” Greg and Lonny are peers, both reporting to Manu, who heads the service-supply-chain org. Manu’s org has four functions plus Greg as the horizontal “fifth wheel” coordinator: customer service (Lonny), reverse logistics + warehousing, repair operations (the CMs), and planning [Greg-0626]. Corroborated by the 2026-06-25 in-person debrief, which places Manu under Deb Shoquist (COO/EVP Operations). Spelling and LinkedIn identity are still being confirmed (Ali is the mutual connection chasing this). This resolves the long-open “unidentified exec sponsor” gap — though Manu’s actual budget authority vs. needing Shoquist sign-off is still unknown. Joe Malchow advised the bullet doc needs to land at this level.

B. The “four pillar” customer service org

Lonny described a newly-forming organizational design at NVIDIA [Lonny-0512]:

  1. Dedicated repair lines — separated from new-production manufacturing
  2. Reverse logistics — distinct from forward delivery
  3. Demand planning — forecasts both sales and projected failure rates
  4. Systems group — tasked with automation and tooling implementation

Lonny’s own pillar is compute-science frontline support — first call, warranty validation, advance-replacement decisions [Lonny-0512].

This structure was Lonny’s framing in May. By June 17, Greg + Lonny were describing two big initiatives that came out of an Ernst & Young workshop two weeks earlier — the customer-facing portal/dashboard is one of them [Greg+Lonny-0617]. The second initiative was not disclosed (pending NDA).

The four pillars and the EY workshop initiatives are likely the same overall transformation program — but Lonny and Greg didn’t explicitly stitch them together for us.

Greg’s 2026-06-26 framing of the same org, from the top down: Manu’s service-supply-chain org has four functions — customer service (Lonny), reverse logistics + warehousing (one group), repair operations (the CMs), and planning — plus Greg as the horizontal “fifth wheel” providing IT/coordination across all four [Greg-0626]. This maps cleanly onto Lonny’s four pillars (Greg’s “systems group” role = the fifth wheel). Greg was explicit that the whole project is primarily solving Lonny’s problem — “he’s my customer, the customer of IT” — and that keeping Lonny (our champion) happy is the objective [Greg-0626].

Mellanox / Israel team — Lonny noted in May that the legacy Mellanox network team and NVIDIA’s frontline support team “still operate like they’re separate companies” [Lonny-0512]. Did not come up again in June.

C. Scale of the problem

  • Meta installed base: 100,000 GPUs today; target 1,000,000 over 5 years [Lonny-0512].
  • Current return volume: “We’re struggling to get hundreds of units back. So if that number becomes thousands of units, we’re just going to be calling over” [Lonny-0512].
  • Repair capacity: If 100 units returned, only ~60 can be repaired; the remaining 40 must be filled from new inventory [Alex-0527]. “This 40 we have to skew from new buy from resin… new buy is all what Jensen cares about because new buy is basically revenue” [Alex-0527].
  • Repair mix: ~90% of repairs are “remanufacturing” (ECO/recall-style known-issue batches — like an Audi/BMW/Mercedes recall) vs. ~10% ad-hoc legacy product issues [Alex-0527].
  • Phoenix incident: NVIDIA sent 5 trucks of “hundreds of millions of dollars of equipment” to a customer receipt location with no advance coordination. Customer couldn’t receive them. Trucks turned around to a security yard, equipment parked in the sun in Phoenix [Greg+Lonny-0617].

D. The customer set

Named by Alex Zhu (May 27): the major CSPs (Cloud Service Providers) NVIDIA’s reverse-logistics org services include Meta, Google, CoreWeave, xAI, OpenAI [Alex-0527]. Alex also called Microsoft NVIDIA’s “biggest customer” mid-conversation [Alex-0527].

Lonny framed hyperscalers as “different animals — they are not shy to pound for what they want; they spend a lot of money. I come from one” [Greg+Lonny-0617].

In the night-before chat, Lonny told Dustin he is hearing the same complaints from “neoclouds and hyperscalers” — i.e., the theoretical ICP for our solution stretches across both categories [LonnyChat-0617].

Government / defense market: Lonny said NVIDIA is “accelerating into the government entity market” — was on a call about it the week before June 17 — and expects government customers to be even more stringent about data stripping than hyperscalers [Greg+Lonny-0617]. Per CLAUDE.md, surface this for the human rather than re-frame around it.

E. What gets returned and what fails

Product surface: NVIDIA sells SXM modules (GPU-only assembly) and trays with 2–6 or 2–8 GPUs plus 1–2 CPUs in different configurations [Greg+Lonny-0617]. DGX units cost “millions of dollars” each [Alex-0527].

The chip itself is rarely the problem. Greg: “Most of our returns the chip is rarely the problem. It’s usually something else on that assembly that is the problem… We salvage a very, very high percentage of our chips. Either they remain on the board as we repair it, or we even find a way to take that GPU off and reuse it on a different assembly” [Greg+Lonny-0617]. NVIDIA does not repair GPUs themselves [Greg+Lonny-0617].

Specific failure modes named by Greg/Lonny:

  • Bent connector pins in cable cartridges — “you can start to get a signal degrade… you don’t need to do gorilla grip on these things; you can just slide it and seat it” — but operators sometimes slam them in, bending pins, propagating signal degradation through the system [Greg+Lonny-0617].
  • Backplane / slot damage — Greg’s example from a former life: a 6-foot-high rack transport system where backplane replacement in the field was hard, requiring panel removal and two-person extraction [Greg+Lonny-0617].
  • Heat sink issues — passing reference [Greg+Lonny-0617].
  • Network switches — repairable; one of the main board-level assemblies NVIDIA does repair (versus GPUs themselves) [Greg+Lonny-0617].

Repair vs. replace vs. scrap:

  • Repairs feed into a services pool for spares [Greg+Lonny-0617].
  • Customer under warranty gets a like-for-like replacement — same feature/functionality, different serial number [Greg+Lonny-0617].
  • Advance replacement of a defective unit can be a refurbishment (industry convention: a unit that has been repaired and tested) [Greg+Lonny-0617].
  • DOA window: if a brand-new module fails within ~30–45 days, customer gets a new one as DOA [Greg+Lonny-0617].

Who actually does the repair: Contract manufacturers. Specifically Wistron and Foxconn for NVIDIA [Lonny-0512]. NVIDIA “gives them the playbook” — diagnosis steps, repair steps, acceptance criteria [Lonny-0512]. Back office support in Hong Kong; warehouse operations in Taiwan; Dallas repair line going live July 2026 [Lonny-0512]. International transit Asia and Mexico adds “week-plus each way” [Lonny-0512].

F. The end-to-end RMA workflow

Reconstructed from Alex’s May 27 walkthrough cross-referenced with Greg/Lonny’s June 17 detail:

  1. Customer opens RMA via Salesforce portal — manual form completion is a “major pain point for enterprise customers.” Alex’s example: “If you’re working for xAI or OpenAI you have to fill the form — you don’t want to do that” [Alex-0527]. Even with future automation (QR-code style, Amazon-like), a human must diagnose because “DGX could cost millions of dollars” — too high-value to auto-trigger [Alex-0527].

  2. Triage and the three-gate approval gauntlet — Lonny’s team validates serial number, warranty tier (standard vs. extended), advance-replacement vs. standard [Lonny-0512]. Greg’s 2026-06-26 detail: this is a three-gate sequence — case management (Lonny: warranty/failure-log/serial checks in Salesforce) → quality (known failure code = approve; unknown → failure-analysis lab) → finance (approves) → approval notice → customer ships. <1% of RMAs have ever been rejected, so the gauntlet is latency, not a filter; known-defect batches should be auto-fast-tracked [Greg-0626].

  3. Warehouse expects ASN — Advance Shipping Notification. Currently not flowing properly. NVIDIA is turning on automated ASN via EDI (electronic data interface) — both the outbound ASN to warehouse and a receipt-acknowledgment EDI back [Alex-0527]. “Versus today these are all emails and papers — really bad” [Alex-0527].

  4. Warehouse receives, dispositions, sends to CM — Currently NVIDIA “doesn’t even create a production order before the CM starts operation” [Alex-0527]. Greg: 30-day SLA from RMA arrival in warehouse to replacement at customer [Greg+Lonny-0617]. 2026-06-26 cycle-time numbers: ~60 days end-to-end today, ~30 best case; 27–40 days is warehouse-receipt → repaired-unit-back-in-stock (the only stretch tracked weekly); the approval gauntlet adds ~1–2 untracked weeks because the clock starts only at warehouse arrival. Two decoupled clocks — customer gets a replacement from stock in ~1–2 weeks, while their own serial takes 30+ days to repair [Greg-0626].

  5. CM repair using consignment + turnkey materials:

    • Consignment — NVIDIA-owned high-value components (chips, boards) sitting at the CM physically; CM takes what it needs; material manager reconciles to catch over-use [Alex-0527].
    • Turnkey — CM-procured cables and third-party parts (Taiwan / China suppliers); CM procurement model [Alex-0527].
    • CMs track component consumption on “suspect sheets” — early ’90s style spreadsheets [Alex-0527]. Short-term plan: automated templates. Long-term: B2B API integration for system-to-system seamless tracking [Alex-0527].
  6. CM ↔ Planning team handoff — Planning team coordinates repair commitments. CM needs a signal from planning to know what to procure. Planning team is being automated by NVIDIA paying SAP $2M+ to replace manual Excel-based planning [Alex-0527].

  7. Repaired unit back to services pool, dispatched via Expeditors — 3PL Expeditors International replaced Omni [Lonny-0512]. Per Lonny, ODM integrators (Quanta etc.) add little value on returns and create handling inefficiencies — equipment gets touched multiple times waiting for carriers. NVIDIA is piloting direct pickup from hyperscalers with one customer, bypassing the ODM [Lonny-0512].

  8. Customer receives replacement — but business-unit approval bottleneck inside hyperscalers: data-center teams can’t unilaterally take a rack down; BUs (Instagram, Facebook) often prefer “let it fail” rather than authorize proactive replacement → advance replacement units sit idle for weeks or months, distorting NVIDIA’s inventory planning [Lonny-0512].

Where it breaks: Steps 3–6 essentially happen outside any system. Greg: “Everything downstream of the SFDC ticket happens via email and spreadsheets — DNs, tracking numbers, serial numbers, ship-to info — all updated manually. Customers reply-all, add more people, threads balloon to 16+ people, customer version and Nvidia version of the data diverge. We cannot effectively communicate with our customers re what is going on” [Greg+Lonny-0617]. Greg’s punchline: “We come to the conversation and we’re not even armed.”

Where the bottleneck is not (per Alex): people, in the abstract. “When you join the company you’re hired for a specific reason… your job is to do a best job at your job description. The way I used to do my job no longer scales — but the only complaint is that I’m now working nights and weekends.” Alex frames the root cause as the absence of a horizontal connector team and single source of truth [Alex-0527].

G. The current tooling stack (named by name across conversations)

SystemFunctionStatus as of June 2026Source
Salesforce (SFDC)Front-end customer portal + case managementIn production[Greg+Lonny-0617] [Alex-0527]
SAPMaterial planning, ERP back end$2M+ planning automation deal in flight with SAP[Alex-0527]
Baxter PlanningService-parts demand planning (forecasts failure rates)In use[Lonny-0512]
Expeditors3PL — replaced OmniReplaced Omni; non-asset-based 3PL[Lonny-0512]
WMS + TMSWarehouse + transportation managementRolling out now; also pushing WMS into 3PL partners[Greg+Lonny-0617]
Data lakeBackend data storeNVIDIA has one; “a lot of this stuff resides there” — MVP would build off it[Greg+Lonny-0617]
EDIElectronic Data Interchange — for ASN, receipt acknowledgmentBeing turned on — moving away from email[Alex-0527]
Customer web portalFront-end for RMA submission, sits on top of SFDCIn production[Greg+Lonny-0617] [Alex-0527]
EDI with select customersSome hyperscalers have established EDI integration sending NVIDIA dataEstablished with “a couple of large customers”[Greg+Lonny-0617]
”Suspect sheets”Manual spreadsheets used by CMs to track component consumption”Early ’90s style” — being replaced with templates → B2B APIs[Alex-0527]
Ernst & YoungWorkshop facilitator that defined the two big initiativesWorkshop ~2 weeks before June 17[Greg+Lonny-0617]

Greg’s framing: this stack mirrors Cisco circa 2007–09 before his ERP transformation — siloed instances, data latency, dual transactions [Greg+Lonny-0617]. He sees the right long-term answer as moving to B2B with partners using their systems — not having NVIDIA maintain dual transactions across NVIDIA + CM/3PL systems [Greg+Lonny-0617].

H. Telemetry — what NVIDIA can and cannot see

The clearest structural fact across both June 17 conversations:

  • “You’re not gonna get telemetry off of their devices.” Customer data is highly confidential. Chips are virtually wiped before being returned to NVIDIA [Greg+Lonny-0617].
  • Government customers will be even more stringent about stripping data [Greg+Lonny-0617].
  • First signal of product performance comes from ticket volume against install base. “Any two data points start a trend” [Greg+Lonny-0617].
  • NVIDIA does receive logs from customers — reactively (after a fault has occurred), never proactively [Greg+Lonny-0617].
  • Greg’s prior company (likely Infinera, photonics) had a crawl program that scanned customer networks, ran degradation models, predicted failures within ~90 days. Customers usually chose “let it fail” + 4-hour SLA replacement rather than coordinate maintenance windows — but the telemetry still allowed Greg’s company to pre-position spares [Greg+Lonny-0617].

Lonny softened this in the night-before chat: “It’s not that the data doesn’t exist. It’s just like, what the hell is pychant data or telemetry data? I don’t know. … Customers are having the same problems too.” He said some customers are willing to share metrics selectively for triage [LonnyChat-0617]. Per [Debrief-0617]: it’s unclear how much is structural vs. just a lack of integration.

I. The current vision — shared customer portal

Greg + Lonny’s articulated goal [Greg+Lonny-0617]:

  • A shared internal + external portal where customers see real-time state of their engagement with NVIDIA without calling or emailing.
  • Visible data: active RMAs, shipment status, tracking numbers, serial numbers, what the customer owes NVIDIA in return.
  • Automated escalation triggers when SLAs are breached (e.g., day 31 fires a notification) or when equipment sits unshipped (e.g., 49 hours after pickup notification).
  • Drill-down from macro scorecard (RMA volume, closure rate, average turnaround) to line-item detail.
  • Internal view = more milestones. Customer-facing view = curated but live subset.
  • MVP approach: quick-and-dirty using data already in the data lake. Decision pending whether to build on SFDC portal or as a separate solution.

Greg framed the underlying tension: customers are aggressive about pushing for what they want; NVIDIA needs to be equally assertive about what it needs back (returns, ASN data, serial numbers from customer side) [Greg+Lonny-0617]. The portal also serves an internal coverage purpose — when a customer calls anyone at NVIDIA, that person can look in one place instead of restarting the email chain [Greg+Lonny-0617].

J. The RFP context (from debrief and night-before chat)

Per [Debrief-0617]:

  • NVIDIA is actively running an RFP for a data integration dashboard for their reverse supply chain.
  • Finalists: Palantir, Accenture, or build internally.
  • NVIDIA explicitly said yes when Bliss + Dustin asked if they could build it.
  • Thursday 8am Zoom confirmed (turned into in-person rescheduling).

Per [LonnyChat-0617]:

  • Explicit leadership directive at NVIDIA: get off legacy tooling.
  • Framing Lonny gave for the engagement: “we’re treating this as like an internal call, getting the ball rolling, getting the dialogue going. This goes well, have an in-person meeting next.”
  • Lonny’s goal: a design partnership with NVIDIA.

K. Commercial / warranty model (what we have)

  • All repairs free to customer. No charge model [Alex-0527].
  • Two warranty packages: standard + extended (extended carries premium SLAs / faster lead time) [Alex-0527].
  • Term of standard warranty: “multiple years” — Alex said he wasn’t an expert on the specifics [Alex-0527].
  • 30-day SLA from RMA arrival in warehouse to replacement to customer (standard) [Greg+Lonny-0617].
  • DOA window: ~30–45 days for brand-new units [Greg+Lonny-0617].
  • Three repair categories (Greg-0626 corrects the earlier 2-category split): reman (current-revision, known defect + fix, runs on the same live mass-production line), repair / sort-and-repair (N-minus-one, separate dedicated line at the same site), and refurbishment (one-off, unknown condition, diagnosed on arrival, may be salvaged for parts). ~90% is reman + repair combined; refurb is the tail. Reman + repair are “really recalls”; all repairs free to customer [Greg-0626; Alex-0527].
  • $8B warranty reserve figure — Dustin quoted publicly-reported $8B; Alex did not confirm and said he was “not an expert” [Alex-0527]. Independently corrected to $2.81B per NVIDIA’s FY26 10-K — see §3 below.

L. What Lonny / Greg / Alex want from us

Across all three conversations:

  • Alex (May 27): offered to introduce us to Greg (“Senior Director, Service Team — he’s new, he’s super busy”). Open to follow-up after PTO. Open to a campus visit. [Alex-0527]
  • Lonny (May 12): extended olive branch — “let me introduce you to Greg, who’s one week ahead of me.” Acknowledged actively procuring external tooling. Said his team has no time for in-house build. [Lonny-0512]
  • Greg + Lonny (June 17): Asked for a short bullet doc outlining what an engagement could look like and how Bliss/Dustin would participate. Greg will explore the NDA process — NVIDIA will issue its own NDA (straightforward for two individuals). After NDA, they’ll share the EY workshop initiative breakdown. Latter part of next week for in-person. [Greg+Lonny-0617]
  • Lonny said explicitly when Dustin asked about engagement scope: “What type of time do you guys have. Infinite? It seems like.” And earlier: “We need to think out of the box here.” [Greg+Lonny-0617]
  • Lonny in night-before chat: “more important than solving the problem, we need to communicate to the customers that we’re on the journey” [LonnyChat-0617].

M. What the 2026-06-26 Greg call added (post-meeting walk-through)

A solo call with Greg on his drive back (led by Dustin) — the most complete end-to-end description of the flow we have. Key additions, with where each is folded into this map:

  • “Customer” ≠ “sold-to.” Who initiates the RMA and where it ships to are the two identifiers that matter. Google buys direct (customer = sold-to); xAI is the customer but Dell/Quanta is the sold-to and integrator. No reliable customer-POC directory inside NVIDIA — “emails going to people that have left the company” [Greg-0626]. (New; not previously in §D.)
  • Three-gate approval gauntlet before warehouse shipment, <1% ever rejected → see §F step 2 [Greg-0626].
  • Three repair categories (reman / repair / refurb), ~90% in the first two → see §E [Greg-0626].
  • Repair geography is two networks: compute → Dallas / Houston / Guadalajara; networking → Vietnam / Israel / India; European compute returns to the US → see §3.B [Greg-0626].
  • Cycle-time numbers (~60d end-to-end, 27–40d warehouse-to-stock, +1–2 wk untracked approval, two decoupled clocks) → see §F step 4 [Greg-0626].
  • Two revenue-vs-repair tensions, confirmed: (a) reman-line capacity competes with new production on the same line — a joint, negotiated decision (revenue wins ~9/10); (b) warranty filled from new units when repaired stock is short (“limit new buys in service depots,” GPUs fought over). NVIDIA combines forces with the revenue org to hold CMs accountable for both because its repair-side volume is “minuscule” next to revenue [Greg-0626].
  • Logistics: Dallas ↔ Guadalajara cross-border + customs is the named queue-time bottleneck. Omni (now owned by Forward Air) is being exited in the US (Union City + Dallas) → consolidating to Expeditors Dallas; Omni may be kept in Asia, still debated [Greg-0626].
  • Org: Manu named as the head above Greg + Lonny (peers) → see §A and §B [Greg-0626].
  • Quanta / integrators: Lonny questioning their RMA value-add; Greg’s VAR analogy (integrators act like distributors — buy material, build, install, sometimes run the data center) ties back to the customer/sold-to split [Greg-0626].

§2 — What We Don’t Know

The three conversations leave material gaps. Listed here without speculation; §3 attempts to fill what’s fillable.

A. Org & buying center

  1. Who is the executive sponsor above Greg + Lonny? Not named in any conversation. Named 2026-06-26: “Manu” (Greg + Lonny are peers under him; Manu sits under Shoquist per the 2026-06-25 debrief). Spelling/identity still being confirmed. Still open: Manu’s actual budget authority vs. needing Shoquist sign-off. [Greg-0626]
  2. Where does customer-service ops actually report in NVIDIA’s org chart? Under Shoquist (Operations)? Under a separate Service / Customer Success EVP? Under Sales?
  3. Who is Alex Zhu’s VP? Alex referenced “a VP here” but didn’t name them.
  4. Is the four-pillar structure formalized? Or aspirational org design that Lonny is articulating?
  5. What is Greg’s actual budget authority? Is he Senior Director, with director-level signoff? Or a VP-level position with deeper authority?
  6. What’s the relationship between EY workshop initiatives and the four pillars? Same program or distinct?
  7. What happened to Mellanox / Israel team integration? Came up in May, never again in June.
  8. Where do NVIDIA’s account managers / Field CTOs fit? Customer requests come in “from all directions”; who else is involved when a customer escalates?
  9. Who in NVIDIA’s exec org owns warranty reserve from an operational standpoint? CFO Kress holds the BS; who runs against it day-to-day?

B. Procurement / RFP

  1. RFP timeline. Decision date? Bid response deadline?
  2. Contract size in play. Is this six-figure, low-seven-figure, eight-figure?
  3. Technical requirements of the RFP — what exactly is being asked for? Discovery + prototype + production system?
  4. What is Palantir actually proposing? Bliss came from Palantir — would know the Foundry pitch.
  5. What is Accenture proposing? Standard SI playbook?
  6. Is “build internal” credible? Or political cover?
  7. What’s NVIDIA IT’s role? They are likely the technical evaluator and contracting party.
  8. Procurement process — who signs the master agreement? PO? SaaS contract? Statement of Work?
  9. What other vendors did NVIDIA consider (Servigistics, ServiceMax, Syncron, ReverseLogix) and rule out?

C. Customer detail

  1. Top 5 customers by RMA volume. Not the same as top 5 by revenue. Who has the loudest support tickets?
  2. Per-customer SLA terms. Do hyperscalers have custom contractual SLA tighter than the generic 30-day?
  3. Custom escalation paths or contractual obligations by customer (e.g., does Meta have escalation rights different from CoreWeave?).
  4. What % of warranty reserve is allocated against named hyperscaler returns?
  5. Number of active RMAs at any given moment. Hundreds? Thousands? Tens of thousands?
  6. What does the existing EDI integration with “a couple of large customers” actually pass? Which customers?
  7. What about NVIDIA’s customers in the secondary market (refurb resold to tier-2 DCs)?

D. Workflow / data quantities

  1. RMAs per month / quarter / year — not stated.
  2. Failure rates by SKU (H100, H200, B100, B200, GB200, network switches).
  3. Failure rates by batch / ECO.
  4. Average cycle time today vs. 30-day SLA. Partly answered 2026-06-26: ~60 days end-to-end, ~30 best case; 27–40 days warehouse-to-stock (tracked); +1–2 wk approval (untracked). Still open: true mean/distribution and the customer-experienced case-submit → replacement clock. [Greg-0626]
  5. Number of stranded-shipment incidents like Phoenix — Greg cited one; is it monthly? Quarterly?
  6. Volume of CM consignment inventory at any time — what’s the working capital tied up?
  7. Average time advance-replacement units sit idle — Lonny said “weeks or months” in May; what’s the mean?
  8. What’s NVIDIA’s repair throughput today at Hong Kong / Taiwan facilities?

E. Technical stack details

  1. Salesforce edition / configuration. Custom objects? Lightning experience?
  2. SAP module set — ECC or S/4HANA? On-prem, HEC, or RISE? What’s the $2M+ deal scope?
  3. Baxter Planning integration state — what does Baxter pull from SAP? What does it push?
  4. Data lake platform — Databricks? Snowflake? NVIDIA proprietary?
  5. Identity / SSO / security context for a third-party developer / vendor.
  6. API surface available for external integrators.
  7. What did the EY workshop actually produce? Process map? Tooling architecture? Prioritized initiatives?
  8. What systems are the WMS / TMS being rolled out — vendor-named?

F. Customer-side telemetry

  1. What logs / metrics WILL customers share, through what mechanism?
  2. Is there any program for telemetry sharing? (e.g., a triage portal?)
  3. Are hyperscalers driving OCP RAS adoption that would change the telemetry calculus?
  4. What’s the relationship between NVIDIA Mission Control / Run:ai and Lonny’s reverse-flow world? Different orgs, different data — but is there a bridge?
  5. What does Meta’s Hardware Sentinel / Fleetscanner / Ripple actually expose to NVIDIA’s RMA flow? Lonny is ex-Meta and would know — but didn’t volunteer it.

G. Failure modes

  1. NVIDIA’s actual failure-mode taxonomy and Pareto. What’s the top 10 failure modes by frequency / cost?
  2. The bent connector pin issue Greg cited — is it a Pareto driver? Or anecdotal?
  3. CoWoS-L thermal cycling on 1400W Blackwell — confirmed driver of failures?
  4. HBM3 vs. HBM3e failure rate differential — relevant since Blackwell uses HBM3e.
  5. Power conversion / VRM failures.
  6. Fleet-wide ECOs in flight right now that dominate repair volume — Alex said 90% is remanufacturing, but which specific ECOs?
  7. What’s the chip-level vs. board-level vs. system-level failure split?

H. Commercial / financial detail

  1. Extended warranty pricing. What does NVIDIA charge? What’s the take rate?
  2. Repair cost per unit, per SKU.
  3. Warranty reserve trajectory — what does FY27 look like at current scaling?
  4. Margin impact of repair-vs-replace decision — Alex said “new buy is what Jensen cares about because new buy is revenue.” What’s the implied margin difference?
  5. Cost of stranded inventory (the Phoenix five-truck event, idle ARMA units).
  6. What does NVIDIA pay Wistron / Foxconn for repair labor? Per-unit? Cost-plus? Fixed?
  7. Cost of the EDI / B2B integration work with CMs.
  8. What’s the financial impact of remanufacturing 90% of repairs being free to customer (i.e., effectively a recall obligation)?

I. Cross-border / repair geography

  1. The Mexico operation — which CM? Which facility? Answered 2026-06-26: Guadalajara (Foxconn). Compute repair = Dallas / Houston / Guadalajara; networking = Vietnam / Israel / India. [Greg-0626]
  2. Dallas — production AND repair, or just production? Lonny says repair line July 2026; public reporting says production. Greg-0626 confirms Dallas is a compute repair node, but which building (Heritage Parkway vs. Mobility Way) and the production-vs-repair SOW split is still open.
  3. Asia repair footprint — back-office Hong Kong, warehouse Taiwan — specifically which facilities? (Greg-0626: a little reman runs in Taiwan; specific buildings still unidentified.)
  4. The China RMA hole — restricted chip RMA impossibility in China; has this created any new authorized repair flows in adjacent geographies?
  5. Export controls in the repair flow — how do EAR rules apply when a chip crosses a border for repair? Where does Lonny/Greg own this vs. compliance org?

J. Strategy / sequencing

  1. Is the dashboard the wedge into something bigger (financialization, telemetry) or is it the whole opportunity from NVIDIA’s POV?
  2. How does this relate to NVIDIA Mission Control, Base Command, Run:ai, Fleet Command, Fleet Intelligence? These are forward-fleet products — is there a thesis to bridge them with reverse-flow?
  3. Why hasn’t NVIDIA itself acquired a reverse-logistics SaaS (Servigistics, ServiceMax)? They’ve acquired heavily in adjacent forward-fleet (Run:ai, Bright Computing). Strategic decision or oversight?
  4. Why hasn’t a Servigistics-class incumbent gone after NVIDIA with a vertical sleeve?
  5. What’s the success criterion at 6 months? 12 months? Are they expecting an ROI metric, or a process delivery, or both?

§3 — What’s Likely True (Research-Filled)

For each gap, the best available answer from external research with explicit confidence labels:

  • High — primary source (10-K, press release, named CM disclosure)
  • Medium-High — multiple trade-press sources triangulate
  • Medium — single source or inference from org logic
  • Low — speculative; flagging for completeness
  • Unknown — not determinable from public sources

A. Exec sponsor / org reporting line

The most consequential finding from external research: the seat above Greg + Lonny may not be filled.

Most plausible current exec sponsor: Brian Feller, VP Global Planning, Logistics & Services.

  • Joined NVIDIA May 2021 from Dell (19 years, most recently VP Global Server Supply Chain Planning). MIT LGO.
  • His own stated scope: “supply, production and capacity planning of NVIDIA’s full product portfolio as well as global logistics and fulfillment for over 20 manufacturing and distribution sites. Additionally, his team includes a growing Services organization to support reverse logistics, repair, and RMA fulfillment.”
  • That phrasing is a near-exact verbal overlap with Greg/Lonny’s scope (RMA, repair, reverse logistics, hyperscaler fulfillment), and the word growing matches Lonny being 8 weeks in and the EY workshop two weeks before our June 17 meeting.
  • Based in Round Rock, TX (not Santa Clara) — Dell-era legacy. Worth knowing in advance: he may not show up in person at a Santa Clara meeting.
  • Confidence: HIGH that Feller owns the org Greg/Lonny sit inside. MEDIUM that he is the direct budget owner on a $1M+ engagement vs. requiring Shoquist sign-off. [Sources: ON Partners placement announcement; theorg.com; ZoomInfo; LinkedIn]

Critical finding — open public req that may be the literal seat above Greg/Lonny: VP, Global Service Operations (JR1999315) posted on NVIDIA’s Workday portal, Santa Clara, $352K–$558K base.

  • Job description explicitly states: “lead NVIDIA’s worldwide service strategy and execution… coordinates all aspects of after-sales support, warranty and repair operations, refurbishment, customer experience, and global logistics… reports directly to the EVP of Global Operations… coordinate global service planning and execution, including demand forecasting, inventory optimization, reverse logistics, and systems integration (e.g., Baxter, SAP).”
  • The explicit mention of Baxter and SAP is the giveaway — those are the exact systems Lonny named. This is the same org.
  • This req being posted as open in mid-2026 means either Brian Feller is being elevated and the seat below him is being backfilled, or NVIDIA is splitting Services out from Planning/Logistics under a new VP. Either way, the named exec sponsor for our deal may be a person who is not yet in seat.
  • Confidence: HIGH that this req exists and reports to Shoquist. UNABLE TO DETERMINE whether it has been filled or who is in the running. [Source: nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/Vice-President--Global-Service-Operations_JR1999315; theladders.com; bebee.com]

EVP Operations: Debora Shoquist — scope explicitly includes “supplier management, contract-manufacturing management, supply planning, logistics, quality management.” Customer-service ops rolls up here through Feller (or the open VP req). On a $1M+ services engagement, she is the budget authority above the VP layer. Confidence: HIGH [Public: NVIDIA Newsroom bio].

Alternate roll-up: Jay Puri (EVP Worldwide Field Operations) — his bio explicitly includes “support services organizations” as part of his scope. There’s a real ambiguity here. At Cisco (Greg’s former employer), Customer Experience rolls up under Sales/Services post-2017 — that’s the Puri pattern. At Meta (Lonny’s former employer), DC Engineering rolls up under Ops — that’s the Shoquist pattern. Public signals suggest NVIDIA splits: forward customer support sits with Puri; backend RMA / repair / reverse logistics sits with Shoquist via Feller. Confidence: MEDIUM-HIGH that Puri does NOT own this specific reverse-logistics workflow. [Public: nvidia.com management team bio]

EVP Enterprise Sales: Shanker Trivedi RETIRED April 2026 after 17 years. Now on Enphase Energy board (announced June 15 2026). Successor not publicly named — may be Puri absorbed the function or the role is still unfilled. Confidence: HIGH on Trivedi retirement; Unknown on successor. [Sources: SDxCentral; GlobeNewswire 2026-06-15]

CFO: Colette Kress owns warranty reserve on the balance sheet. Confidence: HIGH [Public: NVIDIA Newsroom].

Alex Zhu’s “VP” — most plausibly Brian Feller, given Feller’s own stated scope (reverse logistics + RMA fulfillment) matches Alex’s reverse-supply-chain operations work. Confidence: MEDIUM-HIGH.

One-sentence ask to drop into the Thursday meeting: “To make sure we’re scoping for the right approval path — who would own sign-off on a Phase 1 engagement of this size? Is that you, or does it route through Brian’s org?” The way they answer is the single highest-information moment of the meeting. If they name a different person, that’s the person who matters and we have not found them in public sources.

B. Repair geography

Wistron Fort Worth — two buildings, not one. The full footprint NVIDIA’s $500B US AI infrastructure commitment includes:

  • 15200 Heritage Parkway: 324,600 sq ft, $580M (primary site). High confidence this is production. [Public: Fort Worth Report 2025-08-21; Hillwood newsroom 2025]
  • 14601 Mobility Way: 766,994 sq ft, $181M (secondary site — three times larger than primary). Publicly described only as “renovations.” [Public: Fort Worth Report 2025-08-21; Dallas Innovates]
  • Combined: >1M sq ft, $761M total, 888 jobs, operational early 2026.

Strongest inference: The Mobility Way 767K sq ft building is the most plausible candidate for the July 2026 repair line. It is three times larger than the primary production site, publicly described only as “renovations,” and Lonny’s July 2026 timeline lands inside the early-2026 operational window for the Fort Worth complex. Confidence: MEDIUM — this is inference, not confirmed. Highest-leverage question for Thursday: which Wistron building, and is the SOW the same as production or a separate contract?

Foxconn Houston: Fairbanks Logistics Park, 4-building 1M sq ft Class A (acquired from Dalfen for $142M), $450M build-out, GB300 NVL72 production, 600 jobs, production start Q1 2026 with humanoid robot trials. Confidence: HIGH on the production facts. No public source describes a repair line — whether “Dallas + Foxconn” in Lonny’s remark points to Houston is speculative. [Public: Dalfen Industrial 2025; Connect CRE; Tom's Hardware 2025; Assembly Mag]

Mexico — RESOLVED 2026-06-26: Greg names Guadalajara directly. This flips the prior inference (which favored Juarez on the warehouse-lease profile).

  • Foxconn Guadalajara (Tonalá, Jalisco) — the stated Mexico compute-repair endpoint [Greg-0626]: 450m plant on 40-hectare property; $500M–$900M investment (sources disagree); world’s largest GB200 facility; planned capacity 20K servers/month = 240K/yr; production underway; primarily for Project Stargate / OpenAI. Confidence: now HIGH as the Mexico repair node (interview-anchored), having previously been HIGH-on-production / UNKNOWN-on-repair. [Greg-0626; Public: TechSpot; TMTPost; Mexico News Daily; Mexico Business]
  • Wistron / Wiwynn Juarez (prior leading candidate — now demoted): factory upgrade up to $16.7M, warehouse lease through 2030 up to $23M; part of $1.1B–$1.2B US/Mexico subsidiary capital injection May 2025. We previously called the $23M warehouse lease the leading Mexico-endpoint candidate; Greg’s direct naming of Guadalajara supersedes that. Juarez may still serve as a logistics/warehouse node, but it is not the stated repair site. [Public: Taipei Times 2025-05-09; DigiTimes 2025-05-07; Wistron Q1 2025 release]
  • Houston, TX is now also confirmed as a compute-repair node alongside Dallas and Guadalajara (promotes the earlier “Foxconn Houston” speculation); networking products route entirely separately through Vietnam, Israel, and India [Greg-0626].

Asia footprint — Lonny’s description holds but is externally invisible.

  • Wistron Hsinchu Science Park + Hukou Township: AI server production; NTD 21.5B (~$680M) facility upgrades approved 2025 [Public: DigiTimes Nov 2024; Taipei Times 2025-06-20].
  • Foxconn-NVIDIA Taiwan supercomputing cluster: $1.4B, ready H1 2026 [Public: AOL/Reuters].
  • No public source identifies a specific NVIDIA repair facility in Hong Kong or Taiwan. Lonny’s internal description (back-office Hong Kong, warehouse Taiwan) does not appear in any external source. Confidence: LOW on specific Asia repair facility identification.
  • ASE / Amkor on the repair flow: confirmed as packaging partners (forward); silent on reverse-flow involvement. Confidence: UNKNOWN.

China RMA crisis — context only, do not reframe around it:

  • RTX 4090 RMA in China: simple repairs (fan) done in Hong Kong; GPU/memory replacement blocked because Taiwan won’t export repaired card back. Board partners offering full refunds instead. [Public: WCCFTech; VideoCardz; Tom's Hardware]
  • A100/H100 in China: NVIDIA does not provide warranty or repair; gray-market Shenzhen industry repairs hundreds-to-500/month per shop, $1,400–$2,800/GPU. [Public: Reuters via The Decoder; Techzine; Tom's Hardware July 2025]
  • H200 license restart confirmed. Jensen confirmed NVIDIA has received purchase orders from Chinese customers + export licenses; H200 manufacturing restarted; authorized agents reappearing for licensed Asian enterprise customers. [Public: Tom's Hardware; Wecent; Asia Times May 2026]
  • No publicly named NVIDIA adjacent-Asia authorized repair flow to capture customers who can’t service in-country. Confidence: UNABLE TO CONFIRM. Per CLAUDE.md, surface for the human only if Lonny brings it up — don’t reframe.

C. Extended warranty / DGX support tier economics

Standard hardware warranty:

  • Consumer GPUs: 2–3 years. Explicitly voids warranty for datacenter use or commercial GPU clusters. [Public: NVIDIA manufacturer warranty page]
  • DGX systems ship with “Standard DGX Hardware Warranty and a minimum 3 years of DGX Support services”; renewable annually after that. [Public: nvidia.com/en-us/data-center/dgx-support/]
  • PCIe data-center cards installed in OEM servers inherit the server’s base warranty plus any warranty upgrades. [Public: Lenovo Press LP1732 product guide]

Enterprise support tier structure — two named tiers confirmed:

TierCoverageSev 1 responseConfidence
Business Standard24x7 case filing; live support 8am–5pm local business hours4 hoursHIGH [NVIDIA Enterprise Support Policy 2025-05-05 PDF]
Business Critical24/7 live support, mission-critical deployments1 hourHIGH [NVIDIA Enterprise Support Policy 2025-05-05 PDF]

Premium TAM (Technical Account Manager) is mandatory for every DGX SuperPOD. [Public: docs.nvidia.com/dgx-superpod/faq]

Mellanox networking (separate tier structure): Silver / Gold / Platinum, with Silver as default end-to-end, Gold adding 24-hour advance hardware replacement, and 4-Hour Expedite RMA as an explicit optional upgrade. [Public: network.nvidia.com/pdf/support/Mellanox_Global_Expedite_RMA_Service.pdf]

NVIDIA AI Enterprise — the most legible piece of recurring-support economics:

SKUList PriceSource
NVAIE 1-year subscription, per GPU, 8x5 Standard support~$4,500/GPU/yrdell.com APD ac566091; insight.com 731-AI7003 [HIGH]
NVAIE 3-year subscription, per GPU~$13,500 listdell.com APD ac566092 [MEDIUM]
NVAIE 5-year subscription, per GPU~$22,500 listdell.com APD ac566093 [MEDIUM-HIGH]
NVAIE Inception / education tier$1,125/GPU/yr (~75% discount)NVIDIA licensing guide [MEDIUM]
NVAIE Essentials (cloud marketplace)$2.00/GPU/hr pay-as-you-goNVIDIA marketplace [HIGH]

Cross-check: $4,500/GPU/yr against an H100 list of ~$30K is ~15% annualized — comparable to enterprise software-attached hardware ratios.

What is NOT publicly disclosed:

  1. The price delta between Business Standard and Business Critical on a DGX system. No reseller catalog surfaces a DGX hardware support uplift SKU. This is a clean question to ask Lonny/Greg.
  2. SuperPOD Premium TAM annualized cost — “mandatory” but never priced publicly.
  3. Attach rate / take rate for extended warranty or NVAIE — neither 10-K nor any earnings call breaks it out. NVIDIA does not segment NVAIE / support revenue. The economics are deliberately opaque.
  4. Per-RMA cost (logistics + FA + repair labor + replacement unit COGS) — not in our vault, not in any public source.
  5. Hyperscaler bespoke terms (Microsoft, Meta, OpenAI, CoreWeave): size and duration disclosed (e.g., CoreWeave $6.3B through April 2032), but warranty term, indemnity caps, SLA credits NOT in any SEC filing. [Medium]

The economic puzzle that connects to our wedge:

  • ~90% of repairs are “remanufacturing” (ECO/recall-pattern, per Alex Zhu) and all repairs are free to customer.
  • Under ASC 460/450-20, recall reserves typically get specifically disclosed when probable and estimable. NVIDIA’s 10-K does not break out recall-related charges separately. [High; cross-ref [[reverse-logistics-warranty-tam-2026-05-29]]]
  • Two interpretations, both speculative pending Greg/Alex follow-up:
    1. NVIDIA’s accounting classifies ECO repairs as routine warranty under the accrual model rather than discrete recall — softening the income-statement signal.
    2. The distinction between “remanufacturing” and “recall” is real for engineering but immaterial for accounting — and the rising accrual rate IS the disclosure, dressed as warranty.
  • This is publicly invisible. No sell-side coverage, no WarrantyWeek piece (which focused on consumer 16-pin connector melt), no tech press flags it. A digital-twin partner who can distinguish recall-pattern from ad-hoc failure at the unit level has unique balance-sheet relevance — and this is the strongest Phase 2 financialization hook surfaced by the three conversations.

AMD comparison (one cycle behind, ~10x smaller):

YearAMD ReserveClaims paidAccrualsAccrual rate
FY23$85M$106M$126M~0.56%
FY24$188M$110M$213M~0.83%
FY25$308M$238M$358M~1.03%

AMD’s curve mirrors NVIDIA’s. Critical caveat: AMD does NOT segment warranty between MI300/data-center and client/gaming, so the directional similarity is consistent with — but doesn’t prove — an MI300 contribution. [Public: AMD FY25 10-K]

Intel files no product-warranty XBRL concept at all [Public: SEC company facts CIK 50863]. Broadcom and Marvell disclose nothing material. NVIDIA + AMD ≈ 80%+ of US semiconductor industry warranty reserve; NVIDIA alone ~74%. This concentrates the buyer pool for our long-term thesis to two names. [Synthesis; WarrantyWeek 23rd Annual Report 2026-04-16]

D. Customer concentration (corroborates Alex’s list)

  • FY26 10-K customer concentration: Two direct customers ≈ 36% of FY26 revenue; four direct customers each >10% = collectively 61%; top single customer ≈ 22%. Top 5 hyperscalers ≈ “a little over 50%” of revenue per CFO Colette Kress on Q4 FY26 earnings call. Confidence: High [Public: NVIDIA FY26 10-K customer-concentration footnote; Q4 FY26 earnings call].
  • The named customer set (from Alex Zhu + public triangulation): Microsoft, Google, Meta, Amazon (Tier 1 hyperscalers); CoreWeave, Oracle (neoclouds); xAI, OpenAI (frontier labs); sovereigns. Confidence: High [Alex-0527; Public: DCD / CIO Dive 2026].

E. Warranty reserve trajectory — corrected

The $8B figure Dustin cited and Alex didn’t confirm is wrong. Per NVIDIA FY26 10-K product-warranty footnote:

FYReserve ($M, end)Claims paid ($M)Accruals ($M)Accrual rate
FY2382109145
FY2430654278~0.46%
FY251,2902191,203~0.92%
FY262,8079572,474~1.15%

Confidence: High [Public: NVIDIA FY26 10-K, accession 0001045810-26-000021]. Additions attributed “primarily related to Compute & Networking segment” (data center, not consumer). The widely-circulated $8.22B WarrantyWeek figure does not reconcile to NVIDIA’s filing and contradicts WarrantyWeek’s own industry-aggregate report. Pitch deck must reflect $2.81B.

F. Failure modes — independent corroboration

  • Meta Llama 3 paper (2024): 16,384 H100 cluster; 466 interruptions over 54 days; ~78% hardware-related (~363 failures); GPU faults 30.1%; HBM3 17.2%; one failure every ~3 hours; ~9% annualized failure rate. Confidence: High [Public: Meta Llama 3 paper; Tom's Hardware].
  • Scaling math: at 100K Meta GPUs ≈ 9,000 failures/year ≈ 750/month. At 1M Meta GPUs (Lonny’s target) ≈ 90,000 failures/year ≈ 7,500/month. Confidence: High (arithmetic check).
  • Meta detection systems: Fleetscanner (micro-benchmarks every 45–60 days), Ripple (alongside active workloads), Hardware Sentinel (kernel-space exception analysis, outperforming testing-based methods by 41%). Over 66% of training interruptions stem from component failures in SRAMs, HBMs, and network switches. Confidence: High [Public: Meta Engineering Blog, July 2025; ACM ASPLOS 2025 paper].
  • CoWoS-L thermal / CTE mismatch: confirmed driver at 1400W Blackwell TDP — warping, microbump failures in HBM PHY can render the entire chip inoperable. Once CoWoS-bonded, individual chiplets and HBM stacks can’t be replaced. Confidence: Medium-High [Public: SemiEngineering; Chiplet Summit 2025; proteanTecs].
  • OCP GPU & Accelerator RAS Requirements v1.7 (2025-10-23): standardizes error reporting, crash dumps, RCA, error containment, Redfish/IPMI SEL/APEI formats. Hyperscalers are writing the serviceability spec, not NVIDIA. Confidence: High [Public: OCP RAS v1.7].

G. Competitive landscape

  • RFP finalists (Palantir / Accenture / build-internal) — confirmed [Debrief-0617].
  • Adjacent category, not in this RFP but credible if NVIDIA broadens: PTC Servigistics (service-parts planning, ~$144.5M revenue 2024 Syncron), Syncron (aftermarket service lifecycle), PTC ServiceMax (field service execution, acquired by PTC for $1.46B in 2023), ReverseLogix, Optoro (acquired by Blue Yonder Aug 2025). No vendor in this list specifically solves reverse-logistics + repair workflow + service-parts planning + 3PL/logistics integration for semiconductor / data-center hardware as a unified surface. Confidence: Medium-High [Public: PTC/SAP/Salesforce IR; trade press].

H. NVIDIA forward-fleet products (the boundary to our wedge)

NVIDIA owns the forward fleet-management layer; reverse layer is open.

  • NVIDIA Mission Control — GA on Blackwell; AI factory infra mgmt (provisioning, monitoring, error diagnosis).
  • NVIDIA Run:ai — acquired Dec 30 2024 for ~$700M; GPU workload orchestration, scheduling, utilization; folded into Mission Control.
  • NVIDIA Base Command Manager (BCM) — ex-Bright Computing acquisition; cluster provisioning/management.
  • NVIDIA Fleet Command — hybrid-cloud platform for edge AI deployment.
  • NVIDIA AI Enterprise — software platform (NIM, SDKs, drivers, k8s, support SLA).
  • NVIDIA Fleet Intelligence — host-based agent streaming telemetry to managed cloud service.

The boundary: Mission Control + Run:ai is what a hyperscaler runs while a unit is in their data center. Lonny’s portal is what NVIDIA itself needs after a unit fails. The reason these don’t overlap is structural — customers strip telemetry before returning chips, so Mission Control’s signals stop at the customer fence line. Confidence: Medium-High [Public: NVIDIA docs; cross-ref [[nvidia-primer-2026-06-18]] §1].

I. Repair vs. replace economics

  • System OEMs (Dell, HPE, Supermicro) carry flat-to-declining warranty reserves despite booming AI-server revenue:
    • Dell FY26 warranty reserve $450M (flat-ish $467M → $450M).
    • HPE FY25 $284M (declining $318M → $284M).
    • Supermicro FY25 $17.0M (flat, ~0.2–0.3% accrual on $20B+ revenue).
    • Confidence: High [Public: Dell/HPE/SMCI 10-Ks].
  • The cost concentrates at NVIDIA rather than distributing down the chain — under standard semiconductor supplier terms, liability caps at component purchase price and excludes consequential/indirect damages (UCC §2-719). The integrators don’t absorb it. Confidence: High [Public: Law Insider warranty-cap clause library; Stevens & Bolton; Lexology].

J. Financialization layer (Phase 2 horizon)

Per reverse-logistics-warranty-tam-2026-05-29 and financialization-primer-2026-05-29:

  • Closest existing analog: Munich Re + TWAICE for lithium-ion battery performance warranties (2019). Reinsurer assumes battery OEM’s warranty obligation; OEM transfers part of reserve as premium; specialist holds float and earns investment income. Confidence: High that the structural template exists [Public: Munich Re / TWAICE].
  • Munich Re aiSure covers AI output errors, not AI hardware warranty. Mosaic + Munich Re Feb 2026 partnership brings up to ~$15M initial coverage for AI vendor errors. Does not cover data-center hardware warranty. Confidence: High [Public: Munich Re aiSure; Reinsurance News 2026].
  • No public example of a specialist insurer writing warranty-liability risk transfer for data-center GPU/server hardware. Battery analog proves the structure; nobody has built it for compute. Confidence: Medium (absence of evidence is not absence of activity).

§4 — Top Questions for the Meeting

Prioritized for the in-person meeting. Don’t ask all of them — pick the 3–4 that the conversation hasn’t already answered.

Tier 1 — Buying authority & timeline (must answer)

  1. Who is the exec sponsor for this initiative above the two of you? Who owns the budget? Who signs the SOW?
  2. What’s the procurement timeline? When does the RFP have to land a decision?
  3. What’s the engagement shape you’re imagining? Discovery sprint? Embedded design partner? SaaS contract? Statement of work + milestone payments?
  4. What’s the contracting vehicle? NVIDIA Master Services Agreement? Vendor procurement? NDA → MSA → SOW?

Tier 2 — Scope & success (must answer)

  1. Of the two EY workshop initiatives, what’s the second one? Is it customer-facing dashboard sibling, or internal-ops?
  2. What’s success at 6 / 12 months? Are there hard metrics (SLA attainment, cycle time)? Or is it process delivery?
  3. What systems do you want this to integrate with on day 1 vs. day 60? (SFDC, SAP, Baxter, data lake, EDI feeds, Expeditors, CM systems.)
  4. Are you imagining we sit on top of the existing data lake, or do we need our own data store? (Joe Malchow’s structural argument — arm’s-length independent repository — depends on this.)

Tier 3 — Org context (worth asking)

  1. Where do you report in NVIDIA’s org chart? (Under Shoquist? Under whoever replaced Trivedi?)
  2. How does the Israel / Mellanox team integration play with this? Same systems? Different RMA flow?
  3. Does Alex Zhu’s reverse-supply-chain transformation work cross over with what you’re doing? Same data layer or different?
  4. What’s NVIDIA IT’s role in this engagement? They’d typically be a key technical evaluator on a data-platform deal.

Tier 4 — Customer side (worth asking but may not yet be knowable)

  1. Which hyperscaler customers would you most want to involve in the discovery sprint? Joe’s structural argument relies on hyperscalers participating.
  2. Which customers already have EDI integration with you? (“a couple of large customers” per Greg.)
  3. What’s the current escalation pattern — who at the customer side calls who at NVIDIA?

Tier 5 — Geography & failure context (lower priority but useful)

  1. Dallas facility going live July — is it production, repair, or both?
  2. What’s the Mexico operation?
  3. Of the failure modes you cited (bent connector pins, heat sink, backplane), is there a Pareto?

Tier 6 — Commercial / financialization (Phase 2 territory, listen for openings)

  1. Is there warranty-economics work happening anywhere in NVIDIA’s CFO org? (Don’t push; just listen for signals.)
  2. What does “we don’t measure the SLA” actually mean operationally? Is there a reporting infrastructure today that just doesn’t include it, or no instrumentation at all?

§5 — Surprises & Tensions (per RDI methodology)

These are the things from the three conversations that are surprising, contradictory, or otherwise high-signal.

  1. The four-pillar org Lonny described in May and the EY workshop initiatives Greg described in June may or may not be the same program. If they’re separate, NVIDIA is running multiple parallel transformations.

  2. NVIDIA is paying SAP $2M+ to automate planning at the exact moment Greg’s framing argues against siloed ERP investment. Either Greg’s transformation thesis hasn’t reached the SAP procurement decision, or the $2M SAP deal is for a narrower scope than it sounds.

  3. All repairs are free. This implies that 90% of repair volume — “remanufacturing” / known-issue ECOs — is functionally a continuous recall program with no revenue-recovery model. NVIDIA’s warranty reserve growth (0.46% → 1.15% accrual rate in 2 years) is a direct economic signal of this.

  4. The customer base is concentrated and includes named hyperscalers Lonny used to work for. This is asymmetric leverage: Lonny knows from inside how Meta’s BU-approval bottleneck works, because he lived it.

  5. The chip is rarely the problem. The thing being repaired is the board, network switch, connector, or heat-sink level — not the GPU die. This pushes the wedge away from “GPU failure prediction” and toward “system-level component lifecycle management.”

  6. NVIDIA doesn’t repair its own GPUs. Repair-line capacity is entirely with Wistron + Foxconn. NVIDIA owns consignment inventory at CM facilities and the playbook; the CM owns the labor and floor.

  7. The “RFP” framing emerged only in the debrief. In the actual meeting, Greg + Lonny described a “target rich environment” and asked for a bullet doc. Bliss + Dustin reframed this internally as an RFP with Palantir / Accenture / build-internal finalists — but Greg and Lonny did not use that language explicitly. Worth pressure-testing in the meeting whether this is a structured RFP or an executive-sponsored direct procurement.

  8. Lonny’s “we don’t know how much is structural vs. lack of integration” on telemetry (from the night-before chat) contradicts the harder line he and Greg took in the actual meeting (“you’re not gonna get telemetry off their devices”). The truth is probably in between — some customers will share, some won’t, and the boundary depends on relationships.

  9. Greg’s prior company’s “crawl program” precedent is a tell. He has lived the pattern of “instrument the customer network, predict failures within 90 days, pre-position spares.” This is a forward-looking design instinct that contradicts his harder line on telemetry. If we can build something that gives him that without crossing the customer-confidentiality line, we’re aligned with his intuition.

  10. Alex’s Apple framing is a useful anchor: Apple’s supply chain solves advanced optimization problems (90% → 99%). NVIDIA is solving foundational problems. This is consistent with the rest of the picture and tells us the bar for “this is good enough to be useful” is lower than it would be at an Apple-style customer. Phase 1 doesn’t have to be elegant; it has to give them the score of the game.

  11. The seat above Greg + Lonny may not be filled. NVIDIA has a public Workday req — VP, Global Service Operations (JR1999315) — that explicitly names Baxter and SAP and reports to EVP Global Operations. Brian Feller (VP Global Planning, Logistics & Services, ex-Dell, Round Rock TX) is the strongest current candidate for the day-to-day sponsor. The open req means a major reorg is in flight at the layer we are pitching into — likely either Feller elevation or a new VP being slotted in below him. This shapes everything about timing and authority.

  12. Wistron Fort Worth is two buildings, not one. The publicly known Heritage Parkway primary site (324K sq ft, $580M) is dwarfed by the secondary Mobility Way building (767K sq ft, $181M) that is described only as “renovations.” The Mobility Way building is the strongest candidate for Lonny’s “July 2026 repair line.” If true, NVIDIA is building a ~750K sq ft North American GPU repair operation without publicly announcing it.

  13. The economics of “all repairs free” + “~90% remanufacturing” is publicly invisible. Under ASC 460/450-20, recall reserves typically get separately disclosed. NVIDIA’s 10-K does not break out the recall-pattern category. Either NVIDIA’s accounting treats ECO repairs as routine warranty (softening the income-statement signal), or the rising 0.46% → 1.15% accrual rate IS the recall disclosure dressed as warranty. No sell-side analyst has flagged this. It is the strongest balance-sheet hook for a Phase 2 financialization wedge that surfaced from the three conversations.

  14. The approval gauntlet rejects <1% — it’s latency, not a filter (2026-06-26). Three sequential gates (case management → quality → finance) run before the unit ships, adding ~1–2 weeks, yet almost nothing is ever turned down. High-volume known-defect batches (“6,000 H100s from this date range”) walk every gate instead of being auto-fast-tracked. The strongest pure-software/time wedge in the whole flow, and it’s invisible in NVIDIA’s metrics because the clock only starts at warehouse arrival. [Greg-0626]

  15. Two decoupled clocks, neither measured end-to-end (2026-06-26). The customer gets a replacement from stock in ~1–2 weeks; their own serial number takes 30+ days to repair and return to stock. NVIDIA tracks the 27–40 day warehouse-to-stock leg but not the customer-experienced case-submit → replacement time. The metric they report understates the pain the customer actually feels by roughly half. [Greg-0626]

  16. Repair geography is two separate networks (2026-06-26). Compute (TX + Mexico) and networking (Vietnam/Israel/India) are distinct reverse flows with no shared routing — the reverse “supply chain” is at least two supply chains. European compute has no proximate repair capability and ships back to the US. [Greg-0626]


§6 — Confidence Summary

ClaimConfidenceBasis
Lonny’s pain description is accurateHighDirect interview corroborated by Alex Zhu independently
Greg’s “we don’t know the score” framing is accurateHighDirect interview
~90% of repairs are reman + repair (now 3 categories: reman / repair / refurb), ~60/100 repair rateHighCorroborated 2026-06-26 by Greg, who also corrected the split from 2 categories to 3 [Greg-0626]
~60-day end-to-end cycle time; 27–40 days warehouse-to-stock; +1–2 wk untracked approvalHighDirect interview [Greg-0626]
Approval gauntlet (case mgmt → quality → finance); <1% of RMAs ever rejectedHighDirect interview [Greg-0626]
Exec sponsor above Greg + Lonny = “Manu” (under Shoquist)Medium-High[Greg-0626] + 2026-06-25 in-person debrief; spelling/identity TBC
All repairs free to customerMedium-HighSingle source (Alex Zhu); consistent with public DGX support docs
30-day RMA SLA, currently not measuredHighDirect interview (Greg)
Wistron + Foxconn as repair CMs; Dallas line July 2026Medium-HighDirect interview; public sources confirm facility but frame as production
Tech stack: SFDC + SAP + Baxter + ExpeditorsHighDirect interview (Lonny + Alex independently)
Customer set: Meta + Google + CoreWeave + xAI + OpenAI + MicrosoftHighDirect interview (Alex) + 10-K customer concentration
Warranty reserve $2.81B FY26 (NOT $8.22B)HighNVIDIA FY26 10-K primary
9% annualized GPU failure rate at hyperscaleHighMeta Llama 3 paper primary
Tier 1 hyperscalers ≈ “a little over 50%” of FY26 revenueHighCFO Q4 FY26 earnings call
Shanker Trivedi retired April 2026HighSDxCentral primary
Mission Control / Run:ai is forward-layer onlyMedium-HighNVIDIA marketing materials + Lonny’s telemetry boundary
RFP finalists: Palantir / Accenture / build-internalHighDirect debrief; not pressure-tested in meeting
The “RFP” structure is formalMediumBliss + Dustin framing in debrief; not confirmed by Greg/Lonny
Exec sponsor: Brian Feller (VP Planning/Logistics/Services) is the strongest current candidateMedium-HighPublic placement announcement + scope language match
Open VP req (JR1999315) for “Global Service Operations” reporting to ShoquistHighNVIDIA Workday posting; theladders mirror
Wistron Mobility Way 767K sq ft building is the likely Dallas repair lineMediumInference from size + “renovations” framing + Lonny’s July timeline
Foxconn Guadalajara as Mexico repair destinationHighConfirmed 2026-06-26 — Greg names Guadalajara directly [Greg-0626]
Wistron Juarez warehouse lease as Mexico repair destinationDemotedSuperseded by Greg’s direct naming of Guadalajara; Juarez may be a logistics node only
Houston as a compute repair node; networking repaired in Vietnam/Israel/IndiaHigh / Medium-HighDirect interview [Greg-0626]; Greg flagged he’d confirm networking specifics
NVAIE list price ~$4,500/GPU/yrHighMultiple channel listings (Dell, Insight)
Business Standard 4-hr Sev1; Business Critical 1-hr Sev1HighNVIDIA Enterprise Support Policy 2025-05-05
AMD warranty trajectory mirrors NVIDIA’s, ~10x smallerHighAMD FY25 10-K primary
Per-RMA repair cost economicsUnknownNot disclosed in any source
Extended warranty attach / take rateUnknownNVIDIA does not segment this
Whether the EY initiative routes through Feller, Shoquist, or the open VP reqUnknownCannot determine from public info — single best question to ask Thursday

§7 — Where This Synthesis Sits in the Vault

This memo is a knowledge map anchored to the three transcripts the user named. For the broader pitch context, see:


Sources (transcripts): Lonny Orona — 2026-05-12; Alex Zhu — 2026-05-27; Greg + Lonny — 2026-06-17; Bliss + Dustin debrief — 2026-06-17; Lonny night-before chat — 2026-06-17; Greg en-route call — 2026-06-26

Note (2026-06-26): this pre-meeting map has been updated in place with detail from Greg’s post-meeting en-route call ([Greg-0626]). The header still reads “pre-meeting” by date, but the §1 facts now incorporate the in-person meeting’s follow-up. Sections folding in 06-26 detail: §1.A/§1.B (Manu, org), §1.E (3 repair categories), §1.F (approval gauntlet, cycle time), §1.M (digest), §3.B (Guadalajara), §5 (surprises 14–16), §6 (confidence).