NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training


Think of a professional athlete. What separates elite performers is what happens between games: continuous refinement, adjusting to new opponents and sharpening skills based on what the last game exposed.

Agentic AI works the same way. A model is no longer asked for an answer. It’s given a goal and has to keep adapting as environments shift, edge cases emerge and tools change. Unlike a generative model responding to a prompt, an agentic model must plan, use different tools and recover from problems it encounters mid-run.

That’s why post-training, the phase that refines a model after initial training on raw data, is no longer a one-time finishing step. It’s continuous, because the environment that agentic models operate in shifts fast. The tools an agent uses can change week to week. Edge cases surface in production that no test set anticipated. Each deployment brings its own codebase, policies and environment.

Post-training runs loop back from production as new problems surface. The compute footprint grows not because any single run is larger, but because the runs never stop. Agentic AI introduces a new compute pattern for post-training, making it the central workload of the agentic era and the primary driver of intelligence per dollar.

The goal of post-training is to maximize intelligence per dollar by maximizing the yield of every forward and backward pass in the continuous learning cycle. The forward pass — inference — is measured in cost per token. That means that every improvement to cost per token flows directly into intelligence per dollar. 

Agentic Post-Training Demystified

Post-training is where intelligence is built. In pretraining, the model learns to predict the next token, which gives it fluency but not intelligence. Post-training is where it learns to write code, plan a multistep task, use a search tool and recover when something goes wrong. Inference is what comes after: the model working on the job, priced in cost per token.

Because there’s no answer key to memorize, only a reward, the model learns by reinforcement learning (RL) techniques. When given a task, it writes out an attempt — the forward pass — the same work it does on the job. The attempt is scored, and the lesson updates the model’s weights — the backward pass. Across millions of attempts, intelligence grows.

Each step is compute intensive, and running this loop at scale is an orchestration problem: thousands of environments generating rollouts in parallel, rewards being verified and updated weights flowing back into training with accelerators fully utilized. NVIDIA NeMo open libraries, such as NeMo Gym for training environments and NeMo RL for distributed post-training, turn post-training from bespoke research code into repeatable infrastructure. 

Why Intelligence per Dollar Extends Cost per Token 

If inference is the revenue engine, post-training is the multiplier: the more capable the model, the higher the value of every token served. 

Cost per token is the key metric for the inference factory: the all-in cost of delivering 1 million tokens. Intelligence per dollar sits one layer up, answering a different question: what does it cost to build a model worth serving, and keep it worth serving as its environment changes?

The two are nested, not competing. AI infrastructure that lowers cost per token also lowers the cost of every point of intelligence built into the model. And every point of intelligence built in raises the value of every token the inference factory serves. 

In other words, cost per token measures operating yield; intelligence per dollar measures whether the investment in model intelligence is paying off. 

Maximizing Intelligence per Dollar: Post-Training Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra — an open weight, 550-billion-parameter mixture-of-experts (MoE) model, offers verifiable benchmarks and a fully disclosed post-training recipe run on NeMo RL. It scored 71.7% on a standard real-world coding benchmark, SWE-bench verified, where it produced a working fix for roughly seven in 10 real software bugs from open source projects, each one checked against the project’s own tests.  

Illustrative 20 billion rollout tokens, based on prior-generation Nemotron 3 Super’s ~1.2 million rollouts at ~10,000 tokens each, scaled up for the larger Ultra model. Intelligence per dollar between platforms is independent of this assumption; the absolute values scale with the token count.

The NVIDIA Blackwell platform lowers cost per run and makes the frequent post-training the agentic era demands economically viable. That intelligence is reaped across every token served.

The NVIDIA Vera Rubin platform extends the trajectory further, training the largest models with one-fourth the GPUs of the Blackwell generation. It was codesigned from end to end to maximize intelligence per dollar for the agentic post-training load: more rollouts per run, more environments in play and post-training cycles that never stop.

Post-Training Workflows in Action

Prime Intellect’s Lab continuously post-trains frontier open models on NVIDIA Blackwell and uses NVIDIA Dynamo for inference orchestration. With Vera Rubin, Prime Intellect plans to scale reinforcement learning environments, generate more rollouts per run and accelerate training-to-inference iteration loops to maximize intelligence per dollar for businesses.

Prime Intellect has optimized its sandbox infrastructure to integrate with NVIDIA Vera CPUs, enabling low-latency, energy-efficient reinforcement learning. Open source tools and models such as NVIDIA Nemotron and NVIDIA NeMo Gym are also integrated into its software stack. When comparing realistic RL sandbox workloads against alternative x86 architectures, Prime Intellect found that Vera delivers, on average, 30% greater throughput per CPU.

Perplexity’s RL post-training stack runs asynchronously across hundreds of NVIDIA GPUs, with an RDMA-based weight transfer engine that syncs trillion-parameter models in under two seconds between training and inference compute nodes. The resulting post-trained Qwen3 235B models are then served on NVIDIA GB200 NVL72 systems.

Together AI provides post-training as a service, including supervised fine-tuning, RL and direct preference optimization. The service is delivered via a feature-rich application programming interface and software development kit that supports the full range of post-training on its AI Native Cloud platform. It has been running on NVIDIA’s platform and optimized kernel libraries, and is looking to harness the Vera Rubin platform next.

Learn more about NVIDIA Vera Rubin, the platform for AI factories to maximize intelligence per dollar across workloads. And explore NVIDIA’s full-stack platform for training frontier models.

Shift will tidy up your home for free, but will record the chores to train robots


Shift is offering to clean homes for free, but there is one important condition. The company will record those chores to build training data for future home robots.

The New York-based startup is currently offering free cleaning services, in which a vetted operator visits a home and wears a camera-equipped device while performing routine household work. The footage can then help AI systems understand how people clean homes outside controlled lab settings.

Your messy home is valuable AI training data

AI companies have already used text, images, and videos from the internet to train software models. But robots need a different kind of data. They need to understand physical spaces, household objects, and the messy logic of everyday chores.

A robot cannot learn home cleaning only from staged lab videos. Real homes have cluttered tables, dishes stacked in awkward ways, stains in corners, and objects placed where they should not be. That kind of chaos is what makes household footage useful.

Shift is not the only company chasing this kind of physical AI data. In India, startups and data vendors are already building businesses around this demand, paying workers to record first-person videos of everyday tasks and supplying that footage to AI companies. For robotics firms, ordinary human labour is becoming valuable training material.

This is where it starts to feel a bit dystopian

Cleaning may only be the start. In the announcement video, Shift says it eventually plans to move into other areas like plumbing, cooking, and building.

Today, we’re launching shift. We’re starting by cleaning your apartment in New York City, for free.

Here’s how it works. Book a shift cleaning. A vetted shift operator comes to your home wearing one of our devices. They clean. They leave. You pay nothing.

In exchange, we record… pic.twitter.com/oBrCXcEz5G

— shift (@joinshiftX) May 28, 2026

For years, the fear around AI has mostly been about office jobs. Writers, coders, designers, and customer support teams have already felt the pressure, and in some cases, that fear has started turning into job losses.

Trades have largely escaped that conversation because physical work is harder to automate. A chatbot can write an email, but it cannot fix a leaking pipe or clean a messy kitchen. Companies like Shift are trying to close that gap by collecting footage of people doing those exact tasks.

AI and robotics may still need time to match the efficiency and precision of a human worker. But watching companies collect this kind of data to train advanced robots feels like the opening scene of the kind of sci-fi movie that does not end well for humans.

Why Cohere’s ex-AI research lead is betting against the scaling race


AI labs are racing to build data centers as large as Manhattan, each costing billions of dollars and consuming as much energy as a small city. The effort is driven by a deep belief in “scaling” — the idea that adding more computing power to existing AI training methods will eventually yield superintelligent systems capable of performing all kinds of tasks.

But a growing chorus of AI researchers say the scaling of large language models may be reaching its limits, and that other breakthroughs may be needed to improve AI performance.

That’s the bet Sara Hooker, Cohere’s former VP of AI Research and a Google Brain alumna, is taking with her new startup, Adaption Labs. She co-founded the company with fellow Cohere and Google veteran Sudip Roy, and it’s built on the idea that scaling LLMs has become an inefficient way to squeeze more performance out of AI models. Hooker, who left Cohere in August, quietly announced the startup this month to start recruiting more broadly.

In an interview with TechCrunch, Hooker says Adaption Labs is building AI systems that can continuously adapt and learn from their real-world experiences, and do so extremely efficiently. She declined to share details about the methods behind this approach or whether the company relies on LLMs or another architecture.

“There is a turning point now where it’s very clear that the formula of just scaling these models — scaling-pilled approaches, which are attractive but extremely boring — hasn’t produced intelligence that is able to navigate or interact with the world,” said Hooker.

Adapting is the “heart of learning,” according to Hooker. For example, stub your toe when you walk past your dining room table, and you’ll learn to step more carefully around it next time. AI labs have tried to capture this idea through reinforcement learning (RL), which allows AI models to learn from their mistakes in controlled settings. However, today’s RL methods don’t help AI models in production — meaning systems already being used by customers — to learn from their mistakes in real time. They just keep stubbing their toe.

Some AI labs offer consulting services to help enterprises fine-tune their AI models to their custom needs, but it comes at a price. OpenAI reportedly requires customers to spend upwards of $10 million with the company to offer its consulting services on fine-tuning.

Techcrunch event

San Francisco
|
October 27-29, 2025

“We have a handful of frontier labs that determine this set of AI models that are served the same way to everyone, and they’re very expensive to adapt,” said Hooker. “And actually, I think that doesn’t need to be true anymore, and AI systems can very efficiently learn from an environment. Proving that will completely change the dynamics of who gets to control and shape AI, and really, who these models serve at the end of the day.”

Adaption Labs is the latest sign that the industry’s faith in scaling LLMs is wavering. A recent paper from MIT researchers found that the world’s largest AI models may soon show diminishing returns. The vibes in San Francisco seem to be shifting, too. The AI world’s favorite podcaster, Dwarkesh Patel, recently hosted some unusually skeptical conversations with famous AI researchers.

Richard Sutton, a Turing award winner regarded as “the father of RL,” told Patel in September that LLMs can’t truly scale because they don’t learn from real world experience. This month, early OpenAI employee Andrej Karpathy told Patel he had reservations about the longterm potential of RL to improve AI models.

These types of fears aren’t unprecedented. In late 2024, some AI researchers raised concerns that scaling AI models through pretraining — in which AI models learn patterns from heaps of datasets — was hitting diminishing returns. Until then, pretraining had been the secret sauce for OpenAI and Google to improve their models.

Those pretraining scaling concerns are now showing up in the data, but the AI industry has found other ways to improve models. In 2025, breakthroughs around AI reasoning models, which take additional time and computational resources to work through problems before answering, have pushed the capabilities of AI models even further.

AI labs seem convinced that scaling up RL and AI reasoning models are the new frontier. OpenAI researchers previously told TechCrunch that they developed their first AI reasoning model, o1, because they thought it would scale up well. Meta and Periodic Labs researchers recently released a paper exploring how RL could scale performance further — a study that reportedly cost more than $4 million, underscoring how expensive current approaches remain.

Adaption Labs, by contrast, aims to find the next breakthrough, and prove that learning from experience can be far cheaper. The startup was in talks to raise a $20 million to $40 million seed round earlier this fall, according to three investors who reviewed its pitch decks. They say the round has since closed, though the final amount is unclear. Hooker declined to comment.

“We’re set up to be very ambitious,” said Hooker, when asked about her investors.

Hooker previously led Cohere Labs, where she trained small AI models for enterprise use cases. Compact AI systems now routinely outperform their larger counterparts on coding, math, and reasoning benchmarks — a trend Hooker wants to continue pushing on.

She also built a reputation for broadening access to AI research globally, hiring research talent from underrepresented regions such as Africa. While Adaption Labs will open a San Francisco office soon, Hooker says she plans to hire worldwide.

If Hooker and Adaption Labs are right about the limitations of scaling, the implications could be huge. Billions have already been invested in scaling LLMs, with the assumption that bigger models will lead to general intelligence. But it’s possible that true adaptive learning could prove not only more powerful — but far more efficient.

Marina Temkin contributed reporting.



1,000 artists release ‘silent’ album to protest UK copyright sell-out to AI


The U.K. government is pushing forward with plans to attract more AI companies to the region through changes to copyright law that would allow developers to train AI models on artists’ content on the internet — without permission or payment — unless creators proactively “opt out.” Not everyone is marching to the same beat, though.

On Monday, a group of 1,000 musicians released a “silent album,” protesting the planned changes. The album — titled “Is This What We Want?” — features tracks from Kate Bush, Imogen Heap, and contemporary classical composers Max Richter and Thomas Hewitt Jones, among others. It also features co-writing credits from hundreds more, including big names like Annie Lennox, Damon Albarn, Billy Ocean, The Clash, Mystery Jets, Yusuf / Cat Stevens, Riz Ahmed, Tori Amos, and Hans Zimmer. 

But this is not Band Aid part 2. And it’s not a collection of music. Instead, the artists have put together recordings of empty studios and performance spaces — a symbolic representation of what they believe will be the impact of the planned copyright law changes. 

“You can hear my cats moving around,” is how Hewitt Jones described his contribution to the album. “I have two cats in my studio who bother me all day when I’m working.”

To put an even more blunt point on it, the titles of the 12 tracks that make up the album spell out a message: “The British government must not legalize music theft to benefit AI companies.”

The album is just the latest move in the U.K. to bring attention to the issue of how copyright is being handled in AI training. Similar protests are underway in other markets, like the U.S., highlighting a global concern among artists.

Ed Newton-Rex, who organized the project, has simultaneously been leading a bigger campaign against AI training without licensing. A petition he started has now been signed by more than 47,000 writers, visual artists, actors, and others in the creative industries, with nearly 10,000 of them signing up in just the last five weeks since the U.K. government announced its big AI strategy. 

Newton-Rex said he has also been “running a nonprofit in AI for the last year where we’ve been certifying companies that basically don’t scrape and train on great work without permission.” 

Newton-Rex arrived at advocating for artists after having batted for both sides. Classically trained as a composer, he later built an AI-based music composition platform called Jukedeck that let people bypass using copyrighted works by creating their own. Its catchy pitch, where he rapped and riffed on the virtues of using AI to write music, won the TechCrunch Startup Battlefield competition in 2015. Jukedeck was eventually acquired by TikTok, where he worked for some time on music services. 

After several years at other tech companies like Snap and Stability, Newton-Rex is back to considering how to build the future without burning the past. He’s contemplating that idea from a pretty interesting vantage point: He now lives in the Bay Area with wife Alice Newton-Rex, VP of product at WhatsApp. 

The album release comes just ahead of the planned changes to copyright law in the U.K, which would force artists who do not want their work used for AI training purposes to proactively “opt out.”

Newton-Rex thinks this effectively creates a lose-lose situation for artists since there is no opt-out method in place, or any clear way of being able to track what specific material has been fed into any AI system. 

“We know that opt-out schemes are just not taken up,” he said. “This is just going to give 90% [to] 95% of people’s work to AI companies. That’s without a doubt.”

The solution, say the artists, is to produce work in other markets where there might be better protections for it. Hewitt Jones — who threw a working keyboard into a harbor in Kent at an in-person protest not long ago (he fished it out, broken, afterwards) — said he’s considering markets like Switzerland for distributing his music in the future. 

But the rock and hard place of a harbor in Kent are nothing compared to the Wild West of the internet. 

“We’ve been told for decades to share our work online because it’s good for exposure. But now AI companies and, incredibly, governments are turning around and saying, ‘Well, you put that online for free …” Newton-Rex said. “So now artists are just stopping making and sharing their work. A number of artists have contacted me to say this is what they’re doing.”

The album will be posted widely on music platforms sometime Tuesday, the organizers said, and any donations or proceeds from playing it will go to the charity Help Musicians.