Aureak
Insights·July 2026

What investors miss in AI infrastructure diligence

The four errors that recur in the memos I read.

Investors ask the right questions about AI infrastructure businesses. They stress-test the technology, size the market, probe management, model the unit economics. What they consistently mis-weight is what actually determines the return.

Four errors recur in the diligence memos I read, at frequencies that suggest they’re systemic rather than idiosyncratic. All four are correctable — but only if you know to look.

1. Over-weighting Nvidia dependency as a risk

Almost every diligence memo on a GPU-cloud or AI infrastructure business includes some version of: “The company is heavily dependent on Nvidia. If Nvidia is disrupted, or if AMD, Groq, or custom silicon takes share, the thesis is impaired.”

This is directionally true and analytically empty.

Nvidia dependency is not a risk factor. It’s the operating condition of the entire AI infrastructure industry for the next three-to-five years, minimum. Nvidia captures 87–92% of the AI training accelerator market by revenue.1 AMD’s Instinct line is genuinely competitive on FLOPS-per-dollar for some workloads and is scaling, but it remains a minority share in every quarter through 2027 on any credible forecast. Custom silicon (AWS Trainium, Google TPU, Microsoft Maia) matters within hyperscalers but doesn’t show up in the third-party market. Groq and other inference startups have compelling niches but move a tiny fraction of aggregate GPU-hours.

Being Nvidia-dependent is like being water-dependent if you’re in the beverage business. It’s not the differentiator. What matters is the price you pay for Nvidia hardware, and the services you can layer on top.

The real question the memo should be asking is: at what commercial terms does the target buy Nvidia? A hyperscaler paying 20–30% off list with priority allocation, an anchor tenant guarantee, and multi-year commitments is running a fundamentally different business than a Neocloud paying 5–10% off list with quarterly purchase order cycles. That gap — 15–20 percentage points on the largest capex line — determines whether the business earns 200 basis points of operating margin or 800 at the same rental price. Neither is disrupted by Nvidia. Only one earns its cost of capital.

Similarly, the question isn’t “what if Nvidia releases a new chip and strands your current fleet.” Nvidia will release a new chip. It has been shipping a new architecture every 12–18 months (Hopper in 2023, Blackwell in 2024, Rubin in 2026). The question is whether the target’s economic depreciation matches its accounting depreciation. A GB200 is depreciated over 4–5 years for tax and financial reporting purposes. If Rubin arrives in 2026 with roughly 2x throughput at similar rental prices, real economic life may collapse to 30 months. This is not a hypothetical scenario. It is the base case.

What to actually diligence, instead of Nvidia-dependency boilerplate:

  • The target’s Nvidia pricing tier. Direct-to-Nvidia allocation with a signed multi-year commitment versus buying through a systems integrator (Supermicro, Dell) is a 15–25% cost of goods difference on the largest line item.
  • Depreciation schedule versus Nvidia’s product cadence. If accounting life is 5 years and product cadence is 18 months, mid-life stranding is priced into every deal, whether the model shows it or not.
  • Hedges: multi-vendor buying (Instinct MI355X alongside GB200), or asset-lite structures (leases, joint ventures, capacity contracts) that offload architecture risk onto someone else’s balance sheet.

2. Under-weighting the cost of capital

The single biggest silent driver of AI infrastructure returns is the weighted average cost of capital. And it is systematically under-modeled.

Consider two companies buying identical GPU clusters at identical prices. Company A is a hyperscaler with an AA credit rating, financing at 5–6% WACC. Company B is a scaled Neocloud financing 60% of its GPU capex with asset-backed debt at 11–13%, mezzanine paper at 15–18%, and 40% equity at a 20% target IRR — blended WACC around 13–15%.2

Over the depreciable life of the asset — call it four years, though see the point above — that 700–900 basis point spread compounds. On a $500M GPU capex program, the finance-cost drag alone is $35–45M per year for Company B versus Company A. That is 8–10% of gross revenue on the same physical infrastructure. And it is before you touch operating cost differences.

Yet the standard diligence model treats WACC as a discount rate applied at the terminal step to get to NPV. It doesn’t flow through into the year-by-year P&L. Which means the model shows both companies earning “attractive” gross margins, and only in the return calculation does the gap appear — by which point management has already convinced everyone the assumptions are conservative.

Two specific errors follow.

Over-optimistic ROIC. Return on invested capital is often modeled at nominal operating margin without adjusting for the cost of debt service on the underlying capex. A business with 18% EBITDA margins and 8% asset-backed debt on 70% of its capex is earning meaningfully worse ROIC than one at the same margin with 30% debt at 5%. This distinction is invisible in EBITDA multiples.

Under-modeled refinancing risk. GPU-backed debt has typically been financed at 3–5 year tenors, matching the depreciable life of the asset. That means the debt refinances into whatever rate environment exists in 2028–2030. If rates rise, or the credit market for GPU-backed lending narrows — which it will if utilization or rental rates soften — the target refinances at a higher spread. The deal that penciled at 15% IRR pencils at 8%, or negative.

What to actually diligence:

  • Debt structure line by line. Tranche size, tenor, coupon, covenants, collateral. Blended WACC. Refinancing waterfall.
  • Sensitivity to rate scenarios. A 200 bp rise in refinance rates in 2028 should be a modeled case, not a footnote.
  • Interest coverage under stress. What does the ratio look like at 60% utilization instead of 80%? Many GPU-cloud businesses are one downturn away from covenant breach.
  • Off-balance-sheet structures. SPVs and joint ventures that carry debt not appearing in the parent’s consolidated financials — the Meta–Blue Owl “Beignet” template is becoming a common pattern3 — matter for enterprise value even when they don’t appear on the balance sheet.

3. Missing the platform economics

If you asked five investors what makes a Neocloud defensible, most would answer “capacity” — the number of GPUs installed, the pipeline of contracted GPUs, the total booked megawatts. This is a measurable, comparable answer. It’s also almost entirely wrong as a moat.

Capacity is not a moat. Capacity is a fungible input. Whoever installs it earns the return on that specific fleet for its depreciable life; nothing stops a competitor from installing an identical fleet next door. Anchor customers can, and do, switch Neoclouds when their initial term rolls off, if the switching cost is low enough.

The real moat is above the capacity layer. It’s the platform.

The hyperscalers understand this. AWS is not defensible because it has the most servers. It’s defensible because a customer running on AWS has integrated identity and access management, VPC networking, managed databases, event streaming, deployment pipelines, monitoring, compliance attestations, and a hundred adjacent services. Extracting from AWS to a competitor is a 12–24 month engineering project. That is the moat.

The good Neoclouds are quietly building the same thing at a smaller scale. Managed model serving. Managed fine-tuning. Data-plane integration with the customer’s data lake. MLOps tooling. Observability. Role-based access control that works across the customer’s tenants. These are expensive to build and slow to build. But once a customer’s ML pipeline is deeply integrated, switching costs rise from “a spreadsheet exercise” to “a re-architecture.”

The average diligence memo pays this almost no attention. It counts GPUs. It measures utilization. It projects rental rates. It doesn’t ask: what fraction of the target’s customers use two or more services beyond raw compute? What’s the average number of workloads per customer that touch shared infrastructure? What percentage of revenue comes from customers who have been on the platform for more than 24 months?

Those are the numbers that separate a business from a warehouse.

What to actually diligence:

  • Service breadth beyond IaaS. How many services does the average customer consume? What is the trajectory over time?
  • Net retention among customers past their initial term. This is the real signal for platform stickiness.
  • Engineering allocation. What percentage of engineering headcount is on infrastructure hardening versus new platform services? Companies stuck in raw GPU rental mode allocate very little to the latter.
  • Contractual structure. Volume commitments, minimum commitments, or true consumption pricing? True consumption pricing with high service breadth is the strongest signal that customers are staying because it’s easier, not because they’re contractually locked.

4. Treating utilization as an input, not a scenario

Utilization is where most diligence models quietly stop being useful.

Every discounted-cash-flow model I’ve seen for a GPU-cloud business — DCF being the standard way to value an asset by projecting the cash it throws off each year and discounting those flows back to present value — assumes 75–85% steady-state utilization once the fleet is booked. This is a defensible assumption in the current environment where a handful of large AI labs are absorbing every GPU that becomes available. It has been true for the past two years. It will not be true forever.

What actually happens to utilization when it declines? Three scenarios, all plausible over a 5-year hold:

Model efficiency step-change. A DeepSeek-style efficiency breakthrough — where an equivalent-quality model requires 5–10x less training compute — reduces aggregate demand overnight. The DeepSeek-V3 result in early 2025 was a preview. A repeat in the training or inference stack cuts industry-wide GPU demand for months while the market re-equilibrates. If your target is levered 60% against 80% utilization assumptions, that shock is existential.

Customer concentration. Many Neoclouds derive 40–70% of revenue from one to three anchor customers.4 If Anthropic finishes a training run and shifts to inference-heavy consumption on a different chip class, or if OpenAI reduces its Azure or CoreWeave commitment during a renegotiation, the anchor fleet sits at 40% utilization while depreciation continues at full rate. This is priced into the equity of every Neocloud whose top customer represents more than 30% of revenue, whether or not the memo mentions it.

Chip-generation obsolescence. When Rubin ships in volume in 2026–2027, a GB200 at year 3 of its 5-year life competes against a chip with materially better throughput per dollar. Customers do not renew their reserved capacity on GB200 at 2024 pricing. They buy Rubin from a competitor. GB200 utilization at year 4 might be 50% on debt that was underwritten to 80%.

None of these scenarios should be treated as tail risk. All three have happened, or been approximated, in the past 24 months.

What to actually diligence:

  • Utilization sensitivity. What does return look like at 80%, 70%, 60%, 50%? At what utilization does the business breach covenants? At what utilization does it fail to service debt?
  • Customer concentration. If the top three customers walked, what’s the take-rate on their capacity from the residual market?
  • Reserved versus on-demand mix in the revenue base. A reserved-heavy business has 12–24 months of runway to absorb an efficiency shock. An on-demand-heavy business has weeks.

What this means for your diligence memo

The good version of an AI infrastructure diligence memo does four things the average version doesn’t.

It prices Nvidia dependency correctly — not as a boilerplate risk factor, but as an operating condition whose terms (allocation tier, discount depth, refresh cadence) actually determine the returns.

It flows cost of capital through the P&L, not just the terminal step. Interest coverage, debt service, and refinancing risk should be first-page items, not appendix items.

It measures platform breadth as a proxy for durability. Number of services per customer, net retention past the initial term, engineering allocation to platform versus infrastructure. A business that lives entirely on raw compute is worth less than one that lives on compute plus services, even at identical revenue today.

It treats utilization as a scenario, not an input. Every model should have a 60% utilization case and an efficiency-shock case. The businesses that survive both are the businesses worth owning.

None of this is difficult analysis. All of it takes work. And in a category where hundreds of billions of dollars have moved in twenty-four months, the memos that skip these steps are the ones underwriting the losses.

Rahul Mittal

Founder, Aureak. Former AWS and Amazon infrastructure and supply chain leader.

Aureak advises institutional investors on AI infrastructure diligence and portfolio company oversight. Book a call.

References

  1. Nvidia’s share of AI training accelerator revenue is estimated at 87–92% for 2024–2025 across independent analyst frameworks (SemiAnalysis, Omdia, Yole Group). Share of installed AI training capacity is comparable.
  2. Neocloud WACC estimates drawn from the CoreWeave S-1 filing (March 2025), Nebius financial disclosures, and public reporting on Blackstone- and Magnetar-led GPU-backed lending facilities. Actual blended rates for private Neoclouds vary widely and depend heavily on anchor-tenant credit quality.
  3. Reference to the Meta–Blue Owl “Beignet” joint venture structure announced 2025, in which off-balance-sheet SPVs hold data center assets financed through infrastructure private credit.
  4. Customer concentration estimates from the CoreWeave S-1 (Microsoft as approximately 62% of 2024 revenue), Nebius disclosures, and industry reporting on anchor tenant contracts across the Neocloud category.