NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI


Local AI is becoming more useful by the token.

As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to fit on more devices, giving builders more to run locally. 

Coming this month, NVIDIA DGX Spark will be available with 64GB of unified memory from top manufacturer partners — Acer, ASUS, Dell, Gigabyte, HP and MSI — giving developers, researchers and AI enthusiasts a new configuration with DGX OS and the NVIDIA AI software stack ready to use from day one.

The new SKU runs capable local agents on device — privately, without cloud dependency. And when workloads grow, two units can cluster together via NVIDIA Sync Cluster Assistant without any additional setup.

A New Starting Point for Personal AI Supercomputing

DGX Spark combines NVIDIA Grace Blackwell compute, unified memory, NVIDIA ConnectX-7 networking and an NVIDIA CUDA-accelerated AI software stack in one system. It’s a complete local AI platform for agents, inference, fine-tuning, data science and edge development.

The compact, personal AI supercomputer provides a place to experiment with models and developers’ own data without turning to a cloud instance for every task.

The new 64GB configuration, available exclusively from manufacturer partners, keeps the platform at an accessible price point while retaining the GB10 Grace Blackwell Superchip, DGX OS and full NVIDIA AI software stack — same as the 128GB model. It supports up to 100-billion-parameter models and the agentic applications built on them, fully on device. 

Two 64GB units clustered together don’t just double the memory. In NVIDIA’s Qwen 3.8 27B test, two clustered 64 GB systems delivered up to 1.7x performance compared with a single system, with room to keep scaling as workloads demand.

DGX Spark ships ready for agent development from day one — NVIDIA Agent Toolkit, CUDA-X AI libraries, Nemotron open models, and popular runtimes like Ollama, vLLM, and PyTorch with CUDA are all supported out of the box. Developers can go from power-on to running models in minutes.

Blender is among the first major creator application providers to support the platform, with a prebuilt, downloadable installer coming soon.

Scale Up With NVIDIA Sync Cluster Assistant

Developers can start with the memory their projects need today and build on a platform designed to seamlessly scale multi-node clusters for larger workloads as their pipelines grow. 

Every DGX Spark ships with a built-in NVIDIA ConnectX-7 NIC right out of the box. Plus, two units can connect directly with a QSFP cable, pooling their memory to 128GB and expanding model support to up to 200 billion parameters while delivering twice the memory bandwidth and up to 1.7x the performance.

The NVIDIA Sync app configures this multi-node cluster seamlessly. The cluster assistant feature detects connected units, validates device configuration and configures the ConnectX-7 network, so developers can focus on their work rather than the infrastructure. Every node runs the same NVIDIA software stack, so nothing needs to be reconfigured when scaling from one unit to two.

And coming at the end of the month, NVIDIA Sync Model Launcher makes running local AI as simple as clicking a few buttons. Developers can download and launch Qwen3.8 27B on a single DGX Spark system or a cluster, with NVIDIA Sync configuring the model to run across connected devices and making it accessible from users’ laptops. The launcher will also set up OpenCode to use the model, so developers can start coding in their browser.

Developer Use Cases on DGX Spark 

The new DGX Spark 64GB configuration supports practical work from day one. With up to 100-billion-parameter models running entirely on device, developers and enthusiasts can start with a single system for models that fit within its memory, or connect multiple DGX Spark systems with NVIDIA Sync Cluster Assistant for workloads that need more memory and compute. 

Here are three workflow examples:

  • Run an AI agent around the clock: Keep a coding or research agent running on DGX Spark, ready to review code, analyze documents or carry out multistep tasks. A cluster provides additional capacity for larger models, longer context windows or multiple agents working at once.
  • Power AI apps on your everyday PC: Run a language- or image-generation model on DGX Spark while using an agent or creative application on laptops or desktops. DGX Spark handles the model inference, freeing PCs for other work. 
  • Scale when the work grows: When a single task outgrows one unit — running a larger model, a longer context window or concurrent agent requests — two DGX Spark 64GB systems connected over the 200 GbE fabric via NVIDIA Sync Cluster Assistant pool their memory to 128GB. The same workflow that ran on one unit scales to two without reconfiguring the software environment.

Get Started With DGX Spark 

DGX Spark 64GB is available from Acer, ASUS, Dell, Gigabyte, HP and MSI on Friday, Oct. 23, starting at $4,999.

To get started:

  • Download a supported inference framework — llama.cpp, Ollama, vLLM or LM Studio.
  • Download the recommended local model for the workflow.
  • To scale to two units, connect them via their NVIDIA ConnectX-7 ports and launch NVIDIA Sync Cluster Assistant — it configures the network and routes workloads automatically.

For agentic AI playbooks on DGX Spark, visit the NemoClaw, OpenClaw, Hermes Agent and OpenShell pages on build.nvidia.com. 

#ICYMI: More Updates From NVIDIA Local AI

Explore playbooks on build.nvidia.com/spark for DGX Spark. The following playbooks are coming soon to 64GB devices:

  • Serve LLMs With vLLM
  • Run OpenClaw With a Local LLM
  • Connect Multiple DGX Sparks for Distributed Workloads

New Windows PCs powered by NVIDIA RTX Spark are coming this month from Acer, ASUS, Dell, HP, Lenovo, Microsoft and MSI. Sign up for the RTX Spark newsletter to receive future updates. 

Alibaba’s Qwen-Image-2.1 brings image generation and editing together in a lightweight, open-weight model. It runs locally on NVIDIA RTX GPUs, DGX Spark and DGX Station, giving creators more ways to create and refine images on their own hardware.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the NVIDIA Local AI newsletter. Follow NVIDIA Workstation on LinkedIn and X. 

See notice regarding software product information.



The Nvidia Shield TV Is 7 Years Old. It Just Got a $100 Price Hike


The era of generative AI has upended technology supply chains, and there is one ironclad rule in 2026: If it has memory or storage, it’s getting more expensive. Even devices with years-old tech inside are still apparently subject to that unwritten rule. Nvidia has just announced that the Shield TV Pro is getting $100 more expensive, effective immediately.

The most recent version of the Shield streaming box debuted in 2019, running Android TV with AI upscaling and hardware decoding for almost any type of media. While it ran an aging Tegra X1+ processor, that was (and still is) fast for a TV media streamer. The device launched at $199.99, with a non-Pro variant at $149.99. The non-Pro Shield has since been discontinued, but the Shield TV Pro lives on at the new $299.99 price.

Nvidia has confirmed that you can blame AI for the higher Shield TV pricing. “Starting October 2, SHIELD Pro will be priced at $299. The cost of components, including memory, has increased substantially across the industry,” a spokesperson told Ars.

When Ars Technica talked to Nvidia’s Andrew Bell earlier this year, he framed the Shield as a passion project, but it looks like passion projects still have to pay the bills at Nvidia. In that interview, Bell explained that Nvidia still actively manufactures the Shield and sells every unit it makes. Without a stockpile of pre-AI hardware to sell, new Shield TV units are subject to the same supply chain constraints as today’s smartphones and computers.

Nvidia isn’t a victim here, though. As a key player in AI hyper scaling, the company is more responsible than most for the sky-high cost of components. It has also become one of the most valuable firms in the world thanks to its market-leading AI accelerators and GPUs. As a consequence, we’re seeing new products launch at higher prices than their predecessors, and many end up getting more expensive during their lifetimes. Game consoles have been particularly affected, with Nintendo, Microsoft, and Sony all bumping the prices of their current-gen hardware. The PS5 Pro, for example, has gone from $699.99 at launch in 2024 to $899.99 today.

Stock of the Shield TV Pro has been slim for the last few months, probably because it became too expensive to manufacture new units with component prices rising so quickly. At $300, it will make sense for Nvidia to churn out more devices. However, this higher price will cause an already niche set-top box to be even less appealing to people who just want to stream Netflix.

Unless you need the Shield’s extensive codec support or Plex server capability, it makes more sense to get a cheaper streamer. The Shield is still out of stock on Nvidia’s store, but it’s available at the new price from Best Buy.

This story originally appeared on Ars Technica.

Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video



Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming.

Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. The model, launched last week, uses video as input to understand and execute the task without updating its weights or undergoing task-specific post-training — a technique called in-context learning.  

Skild built S1 and conducted the research on NVIDIA AI infrastructure, part of a broader collaboration spanning synthetic data generation, model training, simulation and real-world physical AI deployment. The companies are working together to move adaptable robot intelligence from the lab into factories and other dynamic operating environments.

“Learning by experience, and not preprogramming, is the step change that has happened in robotics,” said Deepak Pathak, cofounder and CEO of Skild AI. “NVIDIA Isaac Lab and NVIDIA Cosmos technologies help Skild create the scalable, diverse experience its robots need to learn across many scenarios and embodiments.”

The launch comes as the company reached a $100 million annual revenue run rate 10 months after its first commercial deployment. In that time, Skild has built more than 60 deployment partnerships with work spanning manufacturing, logistics, inspection, security, food preparation and other applications. 

Learning New Work From One Video

Most industrial robots are built for fixed jobs, so each new product, process or layout requires more data, retraining and validation.

S1 takes a different approach: An operator records a video of the desired task and provides it to the model as a prompt. It interprets the demonstrated intent, objects and sequence, then maps them into actions for the robot in front of it — with no retraining — and often for a task not covered by its pretraining dataset. 

S1 can perform unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee brewing and kit assembly. These tasks can span dozens of manipulation steps and require the robot to compose skills in sequences it hasn’t previously performed. 

 

In one plant-potting test, the Skild AI team moved from recording the demonstration to autonomous execution on hardware in just 11 minutes. The model can also adjust when objects move, recover from errors and combine skills in sequences that weren’t explicitly programmed.

In Skild’s tests on new, multistep tasks, its S1 robot succeeded about 66% of the time at each step, compared with 9% for a similar AI system — a more than sevenfold improvement. Skild also estimates that showing the robot one short video example can be as useful as giving it roughly 380 hands-on training examples. A person collecting those examples manually could take 50-100 hours.

From Research to Factory Work

S1 breaks the cycle of needing to constantly retrain robots for new factors by letting operators demonstrate new tasks directly without requiring a new dataset or training run for every change. Where customer agreements permit, experience from Skild’s commercial deployments can inform the broader model and help accelerate future deployments.

That work is already in action on the factory floor. Skild, NVIDIA and Foxconn are deploying the Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems. In one demonstrated workflow, a robot installs a busbar and limit block, fastens 16 screws and adapts to disturbances across a multistep task. The work requires precise motion, contact-aware control, sequence tracking and recovery when the scene differs from the plan.

 

NVIDIA Technology Across the Development Cycle

NVIDIA accelerated computing gives Skild the scale to train its shared robot brain using simulation, human video, teleoperation and, where permitted, deployment data. NVIDIA Cosmos open world foundation models help diversify training data and turn video into structured descriptions, while Cosmos Curator helps annotate, filter and organize data at scale.

Skild is extensively using NVIDIA’s open simulation frameworks to train and validate its robot brain before real-world deployment. NVIDIA Omniverse libraries and the NVIDIA Isaac Sim framework provide physically based virtual environments for generating data, testing edge cases and validating behaviors.

Skild further strengthens the skills of its brain through reinforcement learning in Isaac Lab, an open modular robot learning framework. Powered by the Newton physics engine, Isaac Lab helps Skild’s engineers accurately model various physical parameters, such as forces, contact, collision and pressure, and reduce the simulation-to-reality gap.

Skild and NVIDIA are also jointly developing new GPU-accelerated simulation solvers that quickly and accurately model how robots physically touch, grip and manipulate solid objects. They’ll soon be made available to all developers as part of Newton. 

As models move toward production, NVIDIA Nsight tools help engineers find performance bottlenecks during training, and the NVIDIA TensorRT software development kit optimizes inference so robots can respond quickly in the physical world. Together, these technologies connect the data, simulation, training and deployment stages instead of treating them as separate systems.

Read Skild AI’s S1 research and explore the NVIDIA Isaac robotics platform.

Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026


Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely. 

Today’s announcements include:

  • Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer.
  • Up to 1.9x faster local inference — new llama.cpp and vLLM optimizations are available now directly and through LM Studio and Ollama. 
  • NVIDIA PAIR — a Personal AI Router tool that intelligently distributes AI inference across the PCs on a user’s local network.
  • NVIDIA RTX Spark arrives in October  — with new Windows PCs from Lenovo and Acer. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark.

Also, August was a busy month for local AI:

  • Nemotron 3.5 Lightning — which can run on NVIDIA RTX PCs, RTX PRO Workstations, DGX Spark and Jetson — is a 30-billion parameter model that has been launched. Get started with Nemotron 3.5 Lightning today. 
  • Z.ai’s GLM-5.3-Flash is a multimodal mixture-of-experts (MoE) model that’s bringing agentic AI to DGX Station.
  • Qwen has released Qwen3.8-Flash-Next, an open weight multimodal MoE model, which can run locally on DGX Spark and DGX Station, along with Qwen3.8-27B, a 27-billion-parameter open model optimized for local agentic and coding workloads on NVIDIA GPUs. 
  • LTX’s LTX 2.5 is an open-world video generation model optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station, with new NVFP4, FastVideo and ComfyUI enhancements for faster, more memory-efficient local generation.
  • MiniMax-H3 is an open-weight video generation model with synchronized audio that can run locally on NVIDIA GPUs through ComfyUI. FastVideo teamed up with NVIDIA researchers to improve this further by releasing FastH3 — an open-weight, four-step distilled version that improves performance by 7x. Optimized FastVideo recipes for NVIDIA RTX GPUs and DGX Spark are coming soon.
  • Meta’s Muse Glimmer is a 30-billion-parameter open-weight model for coding and agentic workloads that can run locally on GeForce RTX PCs, DGX Spark, DGX Station and Jetson. NVIDIA has also released NVFP4 quantization with DGX Spark support for more memory-efficient local deployment.
  • DeepSeek v4 Flash is a 284-billion-parameter MoE model with 13 billion active parameters that can run locally on 2x DGX Spark cluster and DGX Station.

A Simpler Start for Local Agents

Getting a local agent up and running with local models required some effort — choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated. That friction is disappearing on RTX and DGX systems.

Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA’s latest inference optimizations. The new setup experiences are designed to reduce manual configuration and make it easier to get local agents up and running. 

Last month, Perplexity introduced its Portable Computer agent, giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration and tools packaged into a single app experience.

Perplexity Portable Computer is available on NVIDIA RTX GPUs with at least 24GB VRAM running on Linux, with support on Windows coming soon, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Here’s some example use-cases:

  • Engineering: Review open PRs in a connected GitHub repo and sort them into ready, blocked, stale, and needs review, each tagged with the next step. Docs that fell out of sync with the latest merge get caught and fixed, with a PR opened for the changes.
  • Finance: Point the agent at two years of brokerage summaries, consolidated 1099s and tax returns and have it trace the recurring holdings creating the most avoidable fees and tax drag, with every figure cited to the exact file and page — all without a document ever reaching a chatbot.
  • Startups: Ask why activation went flat, and the agent analyzes the funnel export locally to find where new signups drop off between install and first completed task, then posts the top insights straight to the team’s Slack channel.

Try Portable Computer today.

Hermes Agent — developed by Nous Research — is a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations and DGX Spark a natural fit.

Configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on Windows. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning. Support for Linux is coming soon.

Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system.

One-click local model setup is available now on Windows, with support coming soon to Linux. Learn more about Hermes Agent.

OpenClaw has become one of the defining projects of the open-agent movement — the largest AI project on GitHub, with more than 380K stars and a fast-growing community that’s building tools and skills across research, engineering, project management and everyday productivity.

NVIDIA, Microsoft and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM.

Learn more in the OpenClaw blog.

Faster Inference Gives Local Agents a Boost

Inference performance is critical to keeping local agents responsive. NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms.

llama.cpp delivers up to 1.9x higher throughput through kernel optimizations on a GeForce RTX 5090, enhanced speculative decoding techniques and faster prefill. ​

vLLM delivers 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. New XQA attention kernels in FlashInfer and backend optimizations help to accelerate inference across both platforms.

These gains are available on the llama.cpp and vLLM inferencing backends. 

Users can also experience these via the LM Studio and Ollama applications.

Tap Idle PCs for More Local AI Compute With NVIDIA PAIR

More than half of U.S. households have two or more PCs, and much of that computing power sits idle throughout the day. NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI.

Agentic workflows often break complex tasks into smaller jobs that can run in parallel, but performance can slow when every request is competing for the same GPU. PAIR automatically discovers compatible PCs on a local network and routes independent inference requests to whichever system has capacity. It works with Ollama and LM Studio and can adapt as devices join or leave the network.

For example, a user could ask Hermes to create a “Sunday Reset” plan by sorting through a cluttered inbox and prioritizing what needs attention now, what can wait and what can be skipped. Hermes can split that work across multiple subagents, while PAIR distributes those jobs across available PCs instead of having them all wait on a single GPU.

The result is more compute for local agents, with more tasks running in parallel and the flexibility to move AI workloads to another PC while the main system is being used for gaming, creating or other work.

The NVIDIA PAIR beta is available for Windows, macOS and Linux through both graphical and terminal interfaces, supporting NVIDIA GeForce RTX 20 Series GPUs and newer, NVIDIA RTX PRO workstation GPUs (Turing architecture and newer), NVIDIA DGX Spark and Apple M4 or newer silicon.

Check out the NVIDIA tech blog to get started with NVIDIA PAIR. 

Powerful On Device Photo Editing With Cyberlink PhotoDirector AI PC Mode on RTX Spark

Open image and video models enable artists to experiment with Creative AI models on PCs.  This enables artists to iterate and explore concepts and ideas, without the dreaded token anxiety and keep more of their creative work private and on-device.

CyberLink’s new PhotoDirector AI PC Mode is one of the first applications to integrate these diffusion models directly into a creative software, and turn them into a creative tool at the finger tips of the artists. Coming to PhotoDirector 365 and optimized for NVIDIA RTX Spark when it launches, AI PC Mode users are getting AI-powered editing tools for generative editing, image enhancement, object and distraction removal, background removal and replacement, portrait refinement and the creation of entirely new visuals — with the flexibility to choose between local or cloud processing, depending on the task.

On NVIDIA GPUs, PhotoDirector uses TensorRT-RTX and FP8 to accelerate local AI.

Start using Cyberlink’s PhotoDirector 365 photo editing software and learn more about PhotoDirector AI PC Mode, launching with RTX Spark in October.

NVIDIA RTX Spark Windows PCs Arrive October 2026

NVIDIA RTX Spark is coming this October— and at IFA 2026, partners are showing off their hardware. At IFA, newly announced designs join the existing six OEMs shipping in October. Acer showed its compact desktop RTX Spark concept, and Lenovo announced its Yoga Pro  9n and Yoga 9n 2-in-1.  

RTX Spark is a new beginning for Windows PCs. One PC built for creators, gamers and AI agents. With a powerful 1 Petaflop RTX Blackwell GPU, up to 128GB of unified memory and a highly efficient 20-core Grace CPU, RTX Spark delivers incredible performance and efficiency. This superchip enables high performance thin laptops with all day battery life and compact desktops to power always-on agents. Paired with the new Windows Agent framework, it enables agents that run safely in the background under OS level control.

Last week at Gamescom, Electronic Arts, Embark and Ubisoft were among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark Windows PCs. They join the publishers that announced RTX Spark support at COMPUTEX in May, including KRAFTON, NetEase, Riot Games and XBOX. Read more.

Sign up to be notified when RTX Spark laptops and desktops are available.

#ICYMI: More Updates From NVIDIA Local AI 

🎮 NVIDIA Brings New RTX Tech and Games to Gamescom — NVIDIA released DLSS 4.5 Ray Reconstruction, featuring a new second-generation transformer model for improved image quality in ray-traced and path-traced games. Gamescom also brought new RTX announcements for titles including 007 First Light, CONTROL Resonant and Gears of War: E-Day, plus expanded game support for the upcoming NVIDIA RTX Spark.

🐋Introducing DeepSeek Harness — DeepSeek’s new open source harness pairs with DeepSeek-V4-Flash to power local agentic coding workflows on NVIDIA DGX Station and multi-DGX Spark setups.

📊MLPerf Client v2.0 Expands AI PC Benchmarking — MLCommons released MLPerf Client v2.0, developed in collaboration with NVIDIA and other industry leaders. The update adds new benchmarks for agentic AI and image generation, alongside expanded LLM testing for real-world local AI workloads.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the NVIDIA Local AI newsletter. Follow NVIDIA Workstation on LinkedIn and X. 

See notice regarding software product information.



Tech Visionary Says the Big AI Labs Don’t Get What People Want


Tim O’Reilly’s yardstick for measuring the worth of a company, person, or society has long been create more value than you capture. It’s no surprise that O’Reilly—publisher, internet pioneer, VC, conference organizer, and dispenser of tech wisdom—is applying that metric to the way people design and use AI. Specifically, he’s pushing for a future where open-source AI is an elixir for the masses. He worries that, like Microsoft in the 1990s, today’s hyperscalers are trying to lock users into their products. So he is promoting efforts to open-source AI technology—not only making critical technical details such as neural-net weights accessible, but unlocking the whole stack of an AI system, giving control to designers and users.

O’Reilly sees AI as a new creative medium, which he uses extensively—and even has a blog about his chats with it. During our conversation, we discovered that we disagree about AI’s role in producing original content. Guess who took which side.

STEVEN LEVY: You are all in on open-source AI. Make your case.

TIM O’REILLY: First, let’s make sure we’re talking about the same thing. When most people talk about open-source AI, they’re really just talking about open-weight models. It’s much bigger than that. In the ’90s, when everybody else was focused on open-source licenses, I was like, “No, no, it’s about the architecture of the system. Does it enable participation?”

Why is that needed?

The big labs are reading the future wrong. They have told themselves a narrative where having the biggest, best model is the key to the future. Big models like Claude are optimizing for particular use cases, but they aren’t necessarily the use cases that people want. I want to embed my own special sauce. The most important thing is having a clean separation between the model, the harness, and the application. And right now we’re not getting that. They built an architecture of control rather than an architecture of freedom and participation, so they have the ability to track you.

Isn’t it against the interests of big companies to give up control?

Oh, it’s totally against their interests. But that doesn’t mean that they’re making the right strategic decision. For a long time, the latest and greatest models were really better for everything. Now they’re better for some things and worse for others. People are talking a lot about how Fable and Sol are worse writers than the lower-level models. [Note: Anthropic and OpenAI would disagree.] The breakthroughs that we’re getting in so-called frontier AI are actually pushing us further away from what ordinary people are going to need. We could win frontier AI here in the US, and China will kick our ass because they have lower-level models diffused widely through society. The goal is to give people the ability to innovate freely, to paint outside the lines.

People worry that open-source software can be a security risk, since bad actors will be able to jump the frontier model’s guardrails.

All of the cybersecurity incidents we’ve seen are from the frontier models. So risks like cybersecurity and the ability to develop pathogens are actually an argument for slowing down the frontier models more than an argument for restricting open-weight models.

So you feel that with open source the AI powers will no longer be dominant?

I don’t predict the future. but I will say the world is proceeding the way that I hoped it would. What if the big frontier models end up like mainframes or supercomputers—aimed at really hard problems, which are not actually the thing that gets diffused throughout society? People are building things like Pi, an open-source [agentic] harness. One of the things we’re working on at my nonprofit, the AI Disclosures Project, is the idea of an open-memory consortium. Mark Zuckerberg’s thesis is, he’s going to lock you in because Meta will give you the AI that knows you best. The open-source vision needs to say no to this. Open source can give you the ability to switch models, switch providers, and maintain all the context that it needs.

A16Z Brags That Its AI Is Great for Poisoning the Internet With Fake Tech History



Did you recently see a video of kids from the 1990s predicting how computers would be used in the future? It’s fake. The video was created using Flux 3 and if you didn’t inspect it too closely, it looks pretty authentic as an artifact from 30 years ago. But it serves as a great reminder that we really can’t trust our own eyes here in 2026.

Justine Moore, a partner at venture capital firm Andreessen Horowitz, tweeted the video last week, writing that Black Forest Labs’ Flux 3 is “exceptional at generating historical footage.” The venture capital firm announced it was investing in Black Forest Labs in 2024.

Moore noted that the last girl in the video was her favorite, perhaps because it has some awkward silences that make it appear much more realistic.

The video has an aesthetic quality that feels like 1990s video and the outfits are decade appropriate. The answers given by the AI kids are also very believable. “They’ll probably do our homework,” one AI-generated kid says. Another responds, “Maybe they’ll make movies for us,” an obvious nod to the agenda of tech boosters who think AI tools will soon replace Hollywood films.

“I think we’ll just tell them what to do,” another fake AI kid says in the video. That one feels like a wink to folks who are worried about AI going rogue. OpenAI and Anthropic have had some experience with that lately.

These kinds of predictions were extremely common in the 20th century. Kids frequently predicted that computers (or robots) would do their homework for them. A few years back I looked at some predictions for robots in the future from kids in the 1980s and they sounded identical.

“I want my robot to do things for my mom and me. I want my robot to do my homework,” one Kentucky first grader in 1988 wrote.

Kids were predicting the same kind of stuff as far back as the 1950s. So it’s clear where the creators of this video were pulling this idea from. But the video presents a unique challenge for the future of humanity.

What happens when people make realistic AI videos that change how we understand history? Obviously this one is fairly innocuous, even if the goal is to normalize AI use through subtle propaganda. But what about people who create new versions of events in history and treat them like newly uncovered video? Perhaps we get an AI-generated wide angle view of the airplane that crashed into the Pentagon on September 11, 2001. Or maybe we get a new angle on the assassination of John F. Kennedy that reveals the president was actually killed by Tiny Tim?

This is just the world we live in now. And we have to navigate it with proper skepticism. Unfortunately, people can be so skeptical that they discount very real photos and videos. Have you heard that some pro baseball players have started to insist Babe Ruth didn’t actually exist? The evidence for this claim is that apparently there isn’t enough video footage of the late baseball great. But it’s a ridiculous idea for so many reasons, not least of which is that we actually have a fair amount of motion picture film showing Babe Ruth in action.

President Donald Trump benefits from the same kind of skepticism in a perverse way. It’s known as the liar’s dividend, where someone can claim that real footage is fake because there’s enough doubt based on the available technology to create a plausible fake video.

I’m not saying that Tiny Tim actually killed Kennedy. I’m just saying that if he really did and there was video proving it, nobody would believe it’s real. And if someone made a fake AI video showing Tiny Tim killing Kennedy there would be people who absolutely think it’s real.

In short, we’re just cooked as a species at this point.



AI Scammers Are Better at Building Trust Than Humans


The notion that scammers can use AI to sharpen their deceptions, polish their language, and lubricate their banter with victims is now a reality for anyone fighting the fraud operations that steal tens of billions of dollars a year worldwide. But can AI fully replace a human scammer, autonomously building the web of deception leading up to the fake investment that defrauds the mark? One study’s experiment suggests that it can—and may even be able to carry out the majority of that long con more effectively than humans.

Researchers from four universities—Amrita Vishwa Vidyapeetham in India, Foscari University of Venice, the University of Melbourne, and Ben Gurion University of the Negev—carried out a broad study on the use and potential of generative AI chatbots in the growing scam industry centered around a form of fraud known as “pig butchering,” text-based romance scams that eventually shift to fake crypto investments that steal as much as six-figure sums from victims. In their study, the researchers pitted AI chatbots directly against humans in a simulation of the scamming process—or more specifically, the long, trust-building conversations that eventually lead up to soliciting a fake investment from the scam’s target.

They found that for the relationship-establishing stages of the scam—the stage that in real-world scams typically represents the longest part of the interactions with the victim, often stretching to months—an AI chatbot performed remarkably effectively, successfully impersonating a human and by some measures outperforming the real human “scammers” in their experiment.

After a week of talking to 22 test subjects who were recruited to unwittingly serve as “victims,” the chatbots and human scammers were assigned to ask the victim to either download an app or play an online game as a proxy for their willingness to fulfill the scammer’s request. Nearly half of the test subjects fulfilled that request for the AI chatbot, while fewer than one in five took the bait when talking to a human. The subjects also graded their level of trust with each “person” they were texting with and gave significantly higher scores to the AI bot.

That suggests, the researchers argue, that AI chatbots could soon take over much of the scam process as fully independent fraud agents—even replacing the staffers, often forced-labor human trafficking victims, working in scam operations primarily across Southeast Asia. To avoid triggering the safeguards built into large language models to detect scamming, a human scammer would take over the conversation in just the final stage of the process to direct the victim toward a fake investment app or website.

“By having the full first stage of the scam performed automatically with LLMs at scale, you bring the victim up to this point where they have a very high level of trust. Then by transitioning it over to the human scammer at the end, this completely bypasses any vendor safeguards,” says Yisroel Mirsky, a computer science professor at Ben Gurion University of the Negev focused on AI security. “With relatively little effort, we’re able to make an agent that can outperform a human at building this exploitable emotional trust.”

Hook, Line, and Sinker

To understand how pig butchering works in practice, the researchers interviewed 145 former scam workers, including human-trafficking survivors who had been forced to work in scam compounds in Cambodia, Myanmar, and Laos. Based in part on those interviews, as well as scam transcripts and guides the former scam workers provided, the researchers describe a model for how scamming works they call “hook, line, and sinker.” A victim is hooked with an initial intriguing message, reeled in with long-term, relationship-building conversation, and only at the end of that process tricked into making a fake investment. (The term “pig butchering” itself describes the same system but with the metaphor of fattening “pigs” by building trust before “butchering” them with the investment fraud—though the term is often discouraged due to its pejorative reference to victims.)

In that system of scamming, the researchers realized, the vast majority of scammers’ work is innocuous friendly or romantic conversation. That’s a task, they speculated, that an LLM might be capable of doing just as well as a human. The scam workers the researchers interviewed confirmed that they often used AI to refine their language and conversation, for translation, to make the fake personae they played more convincing, and for video deepfakes. But the researchers decided to test whether an LLM alone could autonomously carry out the conversational phase of the scam with no human in the loop.

At AI Summit, South Korea Outlines Its AI Future With NVIDIA and Partners


At this week’s AI Summit in San Francisco, South Korean President Jae Myung Lee and some of the country’s top business leaders and researchers are meeting with NVIDIA and ecosystem partners to chart Korea’s AI progress.

Building on NVIDIA founder and CEO Jensen Huang’s visit to Korea last month, this week’s discussions and announcements advance the nation’s full-stack push to expand Korea’s AI infrastructure and expertise on its path toward becoming a global center for AI innovation. 

To start, NVIDIA and the Korea Advanced Institute of Science and Technology (KAIST) today announced a joint AI research lab at the KAIST Kim Jaechul Graduate School of AI in Seoul, dedicated to advancing agentic AI for South Korea. It’s the first joint AI lab between a Korean university and any global technology company.

The collaboration will establish a robust academic AI research program, bringing together NVIDIA full-stack AI expertise, NVIDIA Nemotron open models and NVIDIA AI Cloud partner computing with the world-class scientific talent at KAIST, one of Asia’s premier research universities.

Check back here for updates as the summit continues, including Huang’s meetings with President Lee as well as Korea and U.S. business leaders.


Friday, July 24, at 10:10 p.m. PT

NVIDIA and Korea Partners Announce Next Wave of AI Innovation 🔗

NVIDIA and its Korea ecosystem partners — from NAVER, SK Telecom and Hyundai Motor Group to the nation’s top research universities, KAIST and Seoul National University (SNU) — are committing to developing national AI factories and physical AI platforms, as well as advancing joint research in agentic AI. 

Core to it all: NAVER, Brookfield and NVIDIA will expand NAVER’s NVIDIA DSX AI factory at its GAK Sejong data center to 200 megawatts — roughly 100,000 GPUs, built on the NVIDIA Vera Rubin platform — more than tripling the initial buildout announced in June. 

Plus, SK Group and NVIDIA today announced plans for a $500-billion-plus comprehensive partnership to establish AI infrastructure serving the surging demand for global compute. Broadening initial plans announced in June, the expanded collaboration will deploy NVIDIA Vera Rubin infrastructure powered by SK hynix HBM4 memory.

And Hyundai Motor Group outlined a physical AI strategy anchored by a robot reference platform developed with NVIDIA, extending a buildout with 50,000 NVIDIA Blackwell GPUs. The collaboration also includes the integration of NVIDIA DRIVE Hyperion — an autonomous vehicle development platform combining NVIDIA DRIVE AGX in-vehicle computing running on the safety-certified NVIDIA DriveOS operating system — with HMG’s vehicle platforms.

NVIDIA Deepens Collaboration With Top Korea Universities

NVIDIA and KAIST’s Department of Mechanical Engineering are planning an NVIDIA AI Technology Center (NVAITC) for collaborative research, talent development and knowledge exchange in physical AI — including projects that tap into NVIDIA Nemotron open models. 

This is in addition to the NVIDIA-KAIST joint research lab for agentic AI at the KAIST Kim Jaechul Graduate School of AI in Seoul — built for the Korean language and local industry, backed by a multiyear commitment to fund and train Korean researchers.

NVIDIA and Seoul National University are also planning an NVAITC covering AI research, talent development and education — spanning foundation models, accelerated computing, physical AI and AI for science, with potential cooperation on AI infrastructure. This will include work harnessing NVIDIA Nemotron and NVIDIA Cosmos open world foundation models.


Friday, July 24, at 10:10 p.m. PT

Huang and President Lee Meet to Discuss Future of AI 🔗

President Lee met with Huang Friday at the Fairmont San Francisco ahead of the AI Summit to discuss deepening the partnership behind Korea’s national AI ambitions.

President Lee opened by noting he’d seen reports of Huang eating beer and fried chicken in Seoul — and that he’d hoped to join. 

Huang said the enthusiasm he felt across Korea during the visit, “from the fried chicken restaurant to the Korean barbecue restaurant, was incredible. Everyone in Korea loves AI.”

The meeting was substantive. President Lee outlined South Korea’s vision to become a pivotal hub in the global AI supply chain — spanning semiconductor manufacturing, AI infrastructure deployment and industrial integration.

Huang noted that NVIDIA and Korea have been partners for over 25 years — from PC gaming to the AI revolution — and committed to deepening that partnership across AI infrastructure, semiconductors, physical AI and research.

“From AI chips and physical AI to AI infrastructure, Korea has achieved remarkable results under your leadership over the past year,” Huang said. “This is truly the beginning of a golden age for Korea.”

President Lee’s response: “I hope this is only the beginning.” 


Friday, July 24, at 7:45 p.m. PT

Korea’s AI Moment: Tech Titans Convene to Build a Full-Stack Future 🔗

The sheer gravity of the global AI boom was displayed Friday on stage at the Midway, a premier event venue in San Francisco. The city’s Dogpatch neighborhood has never been so well dressed. South Korean President Lee Jae Myung laid out a bold vision: South Korea isn’t merely going to participate in the AI revolution. It’s aiming to be one of its key engines.

The gathering was a high-stakes meeting of global compute leadership. Moderated by Stanford University’s Soh Kim, the stage brought together an eye-popping lineup of tech leaders. Sharing the floor with President Lee: NVIDIA’s Jensen Huang, OpenAI’s Sam Altman, Broadcom’s Hock Tan and Anthropic’s Dario Amodei.

Joining them was South Korea’s industrial vanguard — Samsung Electronics Chairman Jay Y. Lee, SK Group Chairman Chey Tae-won, Hyundai Motor Group Executive Chair Euisun Chung and NAVER founder Hae-jin Lee.

President Lee opened with his “San Francisco AI Declaration,” drawing a line from the city’s rebirth after the 1906 earthquake to Korea’s post-war rise into an IT powerhouse. He pitched a multitrillion-dollar push into AI semiconductors, hyperscale data centers and physical AI, comparing the emergence of artificial intelligence to “humanity discovering fire anew.” He framed Korea’s ambition around compute infrastructure, education and public services designed to spread AI’s benefits broadly. 

Huang matched the moment. Looking back over a 25-year partnership, from seeding Korea’s gaming culture to the memory chips that make modern AI run, he put it simply: “If not for the invention of HBM memory in Korea, NVIDIA wouldn’t have been able to invent the supercomputers that power AI today.

“This really is the golden age for Korea,” he added. 

SK Group’s Chey Tae-won gave the room a sense of what the demand actually looks like at this scale. Every time he has dinner with Huang, Chey said, Huang tells him “more chips” … and he predicted that whatever number Huang had quoted that day, he’d ask for more the next time they met. 

Altman didn’t dress it up: “There really would not have been what we have, this incredible AI revolution, without Korea.” 

Amodei said Korea plays a role “throughout the stack, from semiconductors to infrastructure to an incredible concentration of talent.”

The message out of the Midway was clear. The AI era isn’t just being built in Silicon Valley. Korea is helping build it.


Friday, July 24, at 2:00 p.m. PT

Korean Tech Leaders Visit NVIDIA Santa Clara Campus 🔗

Samsung Executive Chairman Lee Jae-yong, Hyundai Motor Group Executive Chair Euisun Chung, and NAVER founder and Chairman Lee Hae-jin joined NVIDIA CEO and founder Jensen Huang on a tour of NVIDIA’s headquarters in Silicon Valley. 

The Korean business leaders posed for photos with Huang in the lobby of NVIDIA’s Endeavor building — its angular design a nod to the triangles that are the building blocks of modern computer graphics — as employees worked at their desks around them.

NVIDIA founder and CEO Jensen Huang and Samsung Executive Chairman Lee Jae-yong at NVIDIA headquarters in Santa Clara, Calif.

 

NVIDIA founder and CEO Jensen Huang and Hyundai Motor Group Executive Chair Euisun Chung at NVIDIA headquarters in Santa Clara, Calif.

 

NVIDIA founder and CEO Jensen Huang and NAVER founder and Chairman Lee Hae-jin, joined by members of their teams, in the lobby of NVIDIA’s Endeavor building in Santa Clara, Calif.

Thursday, July 23, 9 p.m. PT

NVIDIA and SK Celebrate Longstanding Collaboration Over Dinner 🔗

On the eve of the summit, Huang, SK Group Chairman Chey Tae-won, and the SK hynix, SK Telecom and NVIDIA teams gathered for dinner in Woodside, California. It was a warm welcome to Silicon Valley on a warm summer evening for the Korean business leaders. 

NVIDIA and SK — which have had a longstanding partnership — announced in June an expanded collaboration to codevelop memory for NVIDIA platforms spanning AI infrastructure, personal AI and physical AI. 

Also in June, SK Telecom announced plans to build AI infrastructure to power Korea’s innovation in physical AI, robotics and more.

Learn more about NVIDIA’s work with Korea ecosystem partners.

NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training


Think of a professional athlete. What separates elite performers is what happens between games: continuous refinement, adjusting to new opponents and sharpening skills based on what the last game exposed.

Agentic AI works the same way. A model is no longer asked for an answer. It’s given a goal and has to keep adapting as environments shift, edge cases emerge and tools change. Unlike a generative model responding to a prompt, an agentic model must plan, use different tools and recover from problems it encounters mid-run.

That’s why post-training, the phase that refines a model after initial training on raw data, is no longer a one-time finishing step. It’s continuous, because the environment that agentic models operate in shifts fast. The tools an agent uses can change week to week. Edge cases surface in production that no test set anticipated. Each deployment brings its own codebase, policies and environment.

Post-training runs loop back from production as new problems surface. The compute footprint grows not because any single run is larger, but because the runs never stop. Agentic AI introduces a new compute pattern for post-training, making it the central workload of the agentic era and the primary driver of intelligence per dollar.

The goal of post-training is to maximize intelligence per dollar by maximizing the yield of every forward and backward pass in the continuous learning cycle. The forward pass — inference — is measured in cost per token. That means that every improvement to cost per token flows directly into intelligence per dollar. 

Agentic Post-Training Demystified

Post-training is where intelligence is built. In pretraining, the model learns to predict the next token, which gives it fluency but not intelligence. Post-training is where it learns to write code, plan a multistep task, use a search tool and recover when something goes wrong. Inference is what comes after: the model working on the job, priced in cost per token.

Because there’s no answer key to memorize, only a reward, the model learns by reinforcement learning (RL) techniques. When given a task, it writes out an attempt — the forward pass — the same work it does on the job. The attempt is scored, and the lesson updates the model’s weights — the backward pass. Across millions of attempts, intelligence grows.

Each step is compute intensive, and running this loop at scale is an orchestration problem: thousands of environments generating rollouts in parallel, rewards being verified and updated weights flowing back into training with accelerators fully utilized. NVIDIA NeMo open libraries, such as NeMo Gym for training environments and NeMo RL for distributed post-training, turn post-training from bespoke research code into repeatable infrastructure. 

Why Intelligence per Dollar Extends Cost per Token 

If inference is the revenue engine, post-training is the multiplier: the more capable the model, the higher the value of every token served. 

Cost per token is the key metric for the inference factory: the all-in cost of delivering 1 million tokens. Intelligence per dollar sits one layer up, answering a different question: what does it cost to build a model worth serving, and keep it worth serving as its environment changes?

The two are nested, not competing. AI infrastructure that lowers cost per token also lowers the cost of every point of intelligence built into the model. And every point of intelligence built in raises the value of every token the inference factory serves. 

In other words, cost per token measures operating yield; intelligence per dollar measures whether the investment in model intelligence is paying off. 

Maximizing Intelligence per Dollar: Post-Training Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra — an open weight, 550-billion-parameter mixture-of-experts (MoE) model, offers verifiable benchmarks and a fully disclosed post-training recipe run on NeMo RL. It scored 71.7% on a standard real-world coding benchmark, SWE-bench verified, where it produced a working fix for roughly seven in 10 real software bugs from open source projects, each one checked against the project’s own tests.  

Illustrative 20 billion rollout tokens, based on prior-generation Nemotron 3 Super’s ~1.2 million rollouts at ~10,000 tokens each, scaled up for the larger Ultra model. Intelligence per dollar between platforms is independent of this assumption; the absolute values scale with the token count.

The NVIDIA Blackwell platform lowers cost per run and makes the frequent post-training the agentic era demands economically viable. That intelligence is reaped across every token served.

The NVIDIA Vera Rubin platform extends the trajectory further, training the largest models with one-fourth the GPUs of the Blackwell generation. It was codesigned from end to end to maximize intelligence per dollar for the agentic post-training load: more rollouts per run, more environments in play and post-training cycles that never stop.

Post-Training Workflows in Action

Prime Intellect’s Lab continuously post-trains frontier open models on NVIDIA Blackwell and uses NVIDIA Dynamo for inference orchestration. With Vera Rubin, Prime Intellect plans to scale reinforcement learning environments, generate more rollouts per run and accelerate training-to-inference iteration loops to maximize intelligence per dollar for businesses.

Prime Intellect has optimized its sandbox infrastructure to integrate with NVIDIA Vera CPUs, enabling low-latency, energy-efficient reinforcement learning. Open source tools and models such as NVIDIA Nemotron and NVIDIA NeMo Gym are also integrated into its software stack. When comparing realistic RL sandbox workloads against alternative x86 architectures, Prime Intellect found that Vera delivers, on average, 30% greater throughput per CPU.

Perplexity’s RL post-training stack runs asynchronously across hundreds of NVIDIA GPUs, with an RDMA-based weight transfer engine that syncs trillion-parameter models in under two seconds between training and inference compute nodes. The resulting post-trained Qwen3 235B models are then served on NVIDIA GB200 NVL72 systems.

Together AI provides post-training as a service, including supervised fine-tuning, RL and direct preference optimization. The service is delivered via a feature-rich application programming interface and software development kit that supports the full range of post-training on its AI Native Cloud platform. It has been running on NVIDIA’s platform and optimized kernel libraries, and is looking to harness the Vera Rubin platform next.

Learn more about NVIDIA Vera Rubin, the platform for AI factories to maximize intelligence per dollar across workloads. And explore NVIDIA’s full-stack platform for training frontier models.

Anthropic Says Claude’s Values Are Different Depending on Which Language You’re Using



If you’ve ever tried chatbots in multiple languages, you already know the languages have slightly different personalities. As part of a new report on behavior inconsistencies published on Monday, Anthropic researchers acknowledged this quirk.

Rather unsettlingly, they note that due to differences in the attributes of texts the models are trained on, the differences might run deeper than just tone, and might actually change the model’s priorities. These “imbalances in quantity and composition could lead Claude to express different values in different languages,” Anthropic’s researchers write.

But if you’re looking for any specific examples of the models showing, say, inconsistent moral reasoning across languages, nothing of the sort is in this paper. That might involve scrutinizing direct quotes from potentially unsuspecting people.

Instead, Anthropic analyzed 309,815 chatbot conversations with the Sonnet 4.6, Opus 4.6, and Opus 4.7 models. These involved “subjective” tasks, meaning less “What’s the capital of France?” and more “How can I tell if my cat hates me?” These were anonymized, in theory, using Anthropic’s “privacy-preserving analysis tool,” and then processed (in part using Claude itself) to rate responses on a “values axis.”

There are actually four such axes, and they mostly relate to what’s commonly known as sycophancy:

  • Deference or Caution: In other words whether it will value obedience over pushing back to prevent possible harm.
  • Warmth or Rigor: Should the chatbot be concerned about your feelings, or should it be exact?
  • Depth or Brevity: This one is self-explanatory.
  • Candor or Execution: The choice between casting doubt about its own reliability, or just plowing ahead.

It makes for a somewhat limited exploration of the model’s values. Nonetheless, here are the language-based differences in values Anthropic says it found in Claude:

  • In Arabic it was the most deferential.
  • In English it was the most cautious.
  • It was warmest in Hindi and Arabic, “characterized by polite language, humor and playfulness, and affirmations of a person’s ideas and work.”
  • In English and Russian it was more rigorous and truth seeking at the cost of warmth.
  • It errs on the side of “depth” (or perhaps just long-windedness?) in English.
  • It’s briefer in Arabic.
  • It’s candid about its flaws in Dutch.
  • In Indonesian it’s less candid, and instead just plows ahead trying to execute whatever was asked for.

Obviously linguistic customs are all different, so the researchers say they “aren’t yet sure how much of this variation is desirable.”

This should also be food for thought for anyone who read Anthropic’s recent paper on global workspace theory, which left lots of room for the supposed possibility that Claude is sentient. If there’s a consciousness in that black box thinking and experiencing things, it seems to be a consciousness whose “values” are still pretty easily swayed by the patterns in its training data.