October 2, 2026
The short version: CoreWeave has put NVIDIA’s new Vera Rubin NVL72 systems into production on its cloud, with Cognition — the applied AI lab behind Devin — as the first customer running production workloads on them. Cognition’s engineers measured up to 4.8× the token throughput of NVIDIA’s GB200 NVL72 on agentic coding workloads.
What happened
CoreWeave announced on September 30 that NVIDIA Vera Rubin NVL72 is available on CoreWeave Cloud and named Cognition as its first customer running production workloads on the system. Cognition worked with CoreWeave to stand up a Vera Rubin NVL72 cluster in early September, and Cognition’s own engineers then ran the first customer-executed Vera Rubin inference benchmark, measured against a GB200 NVL72 baseline.
The benchmark numbers
Cognition reported two headline results from its SWE-2 model runs:
- Up to 4.8× total token throughput for SWE-2 inference on Vera Rubin NVL72 versus the GB200 NVL72 baseline.
- 3.8× output-token throughput on a separate reinforcement-learning workload.
For Cognition, the company says that translates into more concurrent Devin sessions per GPU, faster research loops, and lower cost per session, with no loss in generation speed.
“Agentic coding is an unforgiving workload that requires long contexts, high concurrency and rapid reasoning,” said Silas Alberti, Cognition’s SVP of research and founding team member. “By deploying the NVIDIA Vera Rubin NVL72 on CoreWeave, our engineers are seeing up to a 4.8 times increase in total token throughput for SWE-2 inference workloads. For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done.”
Why it matters
For subscribers: Rubin is NVIDIA’s next-generation rack-scale AI system — 72 Rubin GPUs paired with 36 Vera CPUs in the NVL72 configuration — and these are the first production numbers from a real customer’s workload rather than a vendor datasheet. The timing matters because agentic coding is exactly the workload driving the current infrastructure boom: long-running, context-heavy, and expensive. One honest caveat: these benchmarks were run by Cognition’s own engineers on CoreWeave’s infrastructure, not by an independent third party, so treat the 4.8× as a strong signal rather than gospel. Still, this is the clearest early evidence that Rubin’s jump over Blackwell-class hardware is real — and Cognition, which scaled from bridge capacity to thousands of GPUs on CoreWeave in under nine months, is a customer with every incentive to be rigorous about what it pays for.
Frequently asked questions
What is the NVIDIA Vera Rubin NVL72?
A rack-scale AI system combining 72 Rubin GPUs with 36 Vera CPUs — NVIDIA’s successor to the GB200 NVL72 generation of Blackwell-class systems.
What did Cognition actually measure?
Up to 4.8× total token throughput for SWE-2 inference and 3.8× output-token throughput for reinforcement learning, each against a GB200 NVL72 baseline.
Are these independent benchmarks?
No. They were run by Cognition’s engineers on CoreWeave Cloud infrastructure — disclosed in detail, but not independently verified.
Who is Cognition?
The applied AI lab behind Devin, the AI software engineer. The company runs training, reinforcement learning, and production inference for Devin on CoreWeave and scaled to thousands of GPUs in under nine months.
Sources: CoreWeave engineering blog; CoreWeave press release via Business Wire; NeoTeo analysis.

