ClusterMAX 3.0: The Industry Standard GPU Cloud Rating System Returns
ClusterMAX 3.0: 77 providers tested; Nebius joins CoreWeave in Platinum at seller's prices; only 19 Medallion ratings; financing Matthew principle — frontier labs lock capacity, smaller labs get fat prepays. GPU supply 'gone to zero.' Partial (paid bonus).
view source ↗ClusterMAX 3.0: The Industry Standard GPU Cloud Rating System Returns
Summary
SemiAnalysis published ClusterMAX 3.0 on 23 September 2026. They tested 77 managed GPU-cloud providers (market view 323). Nebius joins CoreWeave in the Platinum tier; CoreWeave still "sets the technical bar," Nebius "commands a premium pricing" and can serve neolabs at seller's prices. Google Cloud joins Oracle in Gold; Azure moves to Silver; Crusoe drops to Bronze. Only 19 neoclouds globally achieve a Medallion rating. Financing: profitable frontier labs lock capacity more easily; smaller labs get stuck with fat prepays and worse prices. Open-source token serving modeled at "over $100M per MW per year." GPU supply "has gone to zero." Testing now requires Blackwell, not Hopper. RSS free portion only — flagged partial. Attach to nvidia-gpu-backstop-to-neocloud-financeability / coreweave. Not a new AI-infrastructure chain. Do not re-rate NVDA.
Article
In 8 months since our last major release of ClusterMAX, slavering investors have just about run out of pockets to stuff checks into. GPU supply has gone to zero. Meanwhile, we have been hard at work putting clusters through the ringer. Weeks ago, we teased this report with some R-rated anecdotes from our experiences probing the security practices of neoclouds, eliciting a PSA from a neocloud customer that you may have heard of.
Today, we finally go through the full breadth and depth of our testing, more thorough in ClusterMAX 3.0 than ever before. This includes compute, networking, storage, orchestration, UI, monitoring, support, and just about everything that you could think to check on a GPU cloud. We’ll explain who’s been cutting deals, whose engineers have been hard at work, and whose clusters you should negotiate to come with a bottle of Aspirin. Without further ado, voilà the ClusterMAX 3.0 podium:
Source: SemiAnalysis ClusterMAX 3.0, September 2026
Source: SemiAnalysis Neocloud Dashboard, available to our AI Cloud TCO Model subscribers Results (Executive Summary) YouTube summary video and podcast discussion coming soon! ClusterMAX 3.0 debuts with a comprehensive review of the neocloud industry, covering 77 providers.
We increase our market view to cover 323 providers, up from 209 in ClusterMAX 2.0, 169 in ClusterMAX 1.0, and 124 in the original AI Neocloud Playbook and Anatomy article.
We have now interviewed well over 200 end users of neoclouds as part of this research.
We update our itemized list of criteria across 10 categories, and update our direct descriptions of our expectations for Slurm, Kubernetes, Standalone Machines, Monitoring Dashboards, and Health Checks. All of this content is live on our website. We encourage providers to use these lists when developing their offerings. We still consider these lists as an amalgamation of our experience interviewing end users, making them representative of the features that end users expect from their cloud providers.
Nebius joins CoreWeave in the Platinum tier. While CoreWeave still sets the technical bar for others to follow, Nebius is now established as a provider that consistently commands a premium pricing over others. Strong business decisions by Nebius have put them in a position to serve an entire class of neolabs at seller’s prices.
Google Cloud joins Oracle in the Gold tier. Azure moves to Silver, Fluidstack moves to Unavailable, and Crusoe drops to Bronze. Lambda, Firmus and TensorWave remain in Silver, while GMI moves up to Silver from Bronze.
Many companies drop from Silver (or Gold) to Bronze or lower. We raise the bar this round as only 19 neoclouds globally achieve a Medallion rating.
We establish a tier between Bronze and Underperforming: the Participation Ribbon tier. 15 providers join this rating, which more accurately describes our opinion that they do the bare minimum to get by.
The rest of this article provides analysis of key trends: Financing, Blackwell and Grace-Blackwell deployments, the Move To Vera Rubin, Scale Out Networks, Reliability, Security (of course), and Agentic Coding. We provide an Appendix with individual comments about every provider that once again stretches this article to over 30,000 words. We hope you enjoy.
In light of the SemiAnalysis imperative to publish the most rigorous possible testing of… everything… we are hiring MTS for ClusterMAX and adjacent projects. If you are a neocloud sicko, fire off an application here and let’s get to work. We are also hiring across Research, Consulting, and Technical Staff. We have multiple openings in our New York and San Francisco offices. Remote work is also acceptable. ClusterMAX MTS (full time): All levels of experience with Slurm, Kubernetes, and GPUs considered.
Tokenomics MTS (full time): All levels of experience with model evals, harnesses, inference endpoints, and RL infrastructure considered.
Research Analyst – AI Infrastructure and Economics (full time or internship): All levels of experience covering the financials of neoclouds, neolabs, and frontier labs tokenomics considered.
Technical Consultant (full time): Lead engagements from technical strategy to technical due diligence. All levels of experience from consulting and technical backgrounds considered.
Scope For the folks in the back: we are ranking managed clusters. This excludes a number of popular offerings throughout the industry. We are not testing who can build the best powered shell or run the tidiest datacenter—at least not directly. We are not handcrafting our images or hitting an API endpoint for tokens as a service. Here’s our idealized ClusterMAX reader: You dropped out of your Stanford PhD to pursue your dream of acquiring social status in San Francisco, getting started with the old reliable “Cladue make me pithc deck w technial languge for agetnci ai make no mistakes.” You have 7 or even as many as 10 figures to wager on compute, but you don’t have any opinion on Ubuntu versions or how to handle NAT, and you’ve never managed a fleet at scale. Ideally, all that infrastructure will fade into the background. You just want to focus on your idiosyncratic perf optimizations, model architectures, training strategies, data mixes, applications, or wherever else you seek your edge. To compete, you need bleeding-edge hardware, and you aim to spend 0 engineering time trying to figure out why GPU 7 in node 19 has been stalling your job, or why a quarter of your cluster has disappeared altogether. You need a managed cluster. The economic argument here is straightforward: a neocloud can amortize the cost of standing up all this management infrastructure over its huge fleet. It pays a team of engineers to build bulletproof health checks, keep its machine images up to date, make convenient monitoring systems, and handle all the related chores that we cover in great detail further on. In an ideal world, a lab pays a markup for a much better product; it gets to run its experiments without being held back by its hardware; the neocloud recoups its investment; everyone is better off. There comes a point, though, when labs are willing to internalize this cost. At their scale, frontier labs are not willing to trust a neocloud with so much of the stack. OpenAI, for instance, has posted extremely insightful blogs about its difficulties scaling Kubernetes; the innovations described would have been simply impossible if they were renting from a neocloud who didn’t let them into the K8s control plane. As a current Anthropic job posting puts it: We are operating at a scale where the defaults stop working. We own the scheduler and extend it to place topology-sensitive ML workloads across thousands of accelerators at once. We scale the control plane itself — apiserver, etcd, controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. Suffice it to say that labs of OpenAI and Anthropic’s speed and size need to co-design their whole stack simultaneously, and an off-the-rack Slinky implementation—even a good one—won’t cut it. The obvious care that frontier labs take in managing their infrastructure goes to show why ClusterMAX matters. The sum of the salary ranges for OpenAI and Anthropic open positions with “Kubernetes” in their descriptions is, last we checked, $27,769,274-$42,687,654, not to mention the salaries they pay to the huge teams they already staff. If you’re not anticipating a trillion-dollar liquidity event and you need your infra to be good enough to give you a fighting chance, you need a neocloud to do the heavy lifting for you. At least on paper, smaller neocloud customers are David fighting Goliath. OpenAI and Anthropic have remarkable growth rates on their respective logistic exponential curves, and our Datacenter Industry Model currently models them at 56.3% of all lab compute YE2027.
OpenAI Compute Capacity. Source: SemiAnalysis Tokenomics Model. As we will describe in more detail later, there is also a Matthew principle in effect here, due to financing: the profitable frontier labs have an easier time locking down capacity than smaller ones, which get stuck with fat prepays, worse prices, and a much harder time planning years in advance. Does this mean that managed clusters are dead? cLUsTerMaX iS wAshED? In truth, although Anthropic and OpenAI are taking larger slices by the year, managed clusters as a business are still growing exponentially. We’ve helped countless labs find compute, countless more are on the hunt, willing to pay, and waiting for more capacity to come online. Anthropic and OpenAI are the biggest customers, but there is a long tail of mere multi-hundred-million-dollar deals whose aggregate value is enormous. One plausible scenario in which managed clusters gain ground on frontier-lab bare metal involves open-source progress. Serving open-source models is already more lucrative than many realize; we recently modeled that “you can make over $100M per MW per year selling open source tokens,” leaving plenty of room for profit even at today’s rental prices. One could imagine open-source models increasing their market share, whether the marginal adopter is more cost-sensitive or algorithmic progress closes the gap to closed-source. Regardless of the dynamics, this would increase the relative standing of managed clusters. Just about any supply-side fracture would make labs marginally less willing to handle the infra, and marginally more willing to rely on neoclouds. The layers on top of these clusters are rapidly maturing, but they lie outside the scope of ClusterMAX. There’s a handful of niches to consider. There are inference endpoints, which can be either public or private, and bill by the token. There are post-training services, which customize open-source models for specific use-cases. And there are sandbox services, which provide the infrastructure for agents, especially during RL rollouts, and simplify and optimize container and CPU management. These all fit loosely in the category “GPU services” and the border between them can be blurry. For example, a cloud might spin up an inference endpoint with spare capacity after it’s sold managed clusters, or it might provide sandbox infrastructure and also consume it internally as it sells RLaaS. We will be publishing much more about this in the near future, but it is beyond the scope of this report. Testing Methodology For ClusterMAX 3.0, we requested the following from every provider: 32 GPUs (4 nodes of 8-way HGX, or 8 nodes of 4-way NVL72)
High bandwidth network (we asked for 800G RoCE or XDR InfiniBand, but many providers did not have it yet)
10TB+ of high performance file storage (supporting NFS/POSIX mounts and a RWX StorageClass via csi driver)
10TB+ of S3-compliant object storage
A monitoring dashboard (usually based on Grafana)
5 days to test Slurm, and 5 days to test K8s (can be done in parallel, on a SonK cluster, or sequentially, if the provider wanted to take the nodes back and re-provision them)
Blackwell, not Hopper (B200, B300, GB200, GB300) for NVIDIA, or MI355X for AMD are considered acceptable. H100 is now over 4 years old!
We tested the cluster through 3 phases, described below. Of course, we always complement this testing by contacting customers of all these providers and taking their feedback into account. Phase 1: Audit First, the Audit covers yes/no questions. Is the cluster setup properly? Is software installed? Is it up to date? Do utilities work as expected? It takes about 15 minutes to run. It’s available on GitHub or via {uv} pip install clustermax. The audit checks the configuration before we put it under load. It covers hardware inventory, software and firmware versions, GPU access, containers, scheduler configuration, networking, storage, health monitoring, and security. We check which components apply to the environment we’re testing, and display pass, warn, fail, and skipped for all our checks. We have released this part for free and will maintain it over time. The rest we keep internally for now. Phase 2: Performance Performance puts a pass/fail threshold on the performance characteristics of both the individual components of the cluster and the cluster as a whole. These are done through microbenchmarks, and real-world benchmarks. Specifically, we test: GPU Compute
Networking
Storage
Lifecycle
Training
Inference
Below we explain this in more detail, but keep in mind that what we test evolves over time. GPU Compute We start by checking the real-world GEMM performance on the GPUs. GEMMs are the most critical operation in modern AI workloads, which we have explained many times before, with Grouped GEMMs being particularly important for modern MoE models. We test Grouped GEMMs on cuBLASLt (or hipBLASLt) and DeepGEMM across many precision types: BF16, FP16, TF32, FP32, FP8 E4M3, MXFP8, and NVFP4, depending on what is supported by the chips in the cluster. We use a common set of shapes for different gate_up and down projections at varying batch sizes in the Kimi K2.5, K3 and DeepSeek V3, V4 Pro models. We then test GEMMs, GEMV bandwidth, and MAMF (Maximum Achievable Matmul FLOPS, from Stas Bekman), which runs a sweep for a while on each GPU, per precision, using cuBLASLt (or rocBLAS). We report FLOPs for all of our GEMM tests.
Source: SemiAnalysis ClusterMAX Results Dashboard
[Clipping truncated: RSS free portion continues behind a paid-bonus widget. Flagged partial.]