I met a lot of weird robots at CES — here are the most memorable


CES has always been a robot extravaganza, and this year’s event saw the announcement of a number of important robotics developments, including the new, production-ready debut of Atlas, the humanoid from Boston Dynamics. Then there were all the robots on the showroom floor, where bots often serve as good marketing for the companies involved. If they don’t always give a totally accurate representation of where commercial deployment is at the moment, they do give visitors a peek at where it might be headed. And, of course, they sure are fun to look at. I spent a decent amount of time perusing the bots on display this week. Here are some of the most memorable ones I encountered.

The ping pong player

The movie Marty Supreme just came out a month ago, so I guess it’s only appropriate that there was a ping-pong-playing robot at this year’s convention. The Chinese robotics firm Sharpa had rigged up a full-bodied bot to play some competitive table tennis against one of the firm’s staff. When I stopped by the Sharpa booth, the robot was losing to its human competitor, 5-9, and I would not characterize the game that was occurring as particularly fast-paced. Still, the spectacle of seeing a robot play ping pong was impressive enough on its own, and I’m sure I have known some humans whose paddle skills were basically equivalent to (or slightly worse than) the bot’s. A Sharpa rep told me that the company’s main product is its robotic hand, and that the full-bodied bot had been debuted at CES to demonstrate the hand’s dexterity.

The boxer

One of the exhibits that drew the largest crowds involved robots from the Chinese company EngineAI, which is developing humanoid robots. The bots, dubbed the T800 (a nod to the Terminator franchise), were in a mock boxing ring and were styled as fighting machines. That said, I never saw any of the bots actually hit each other. Instead, they would sort of shadowbox near each other, never actually making contact. They were also a little unpredictable. One kept walking out of the ring and into the audience, which naturally got a rise out of onlookers. At another point, one of the bots tripped over its own feet and then face-planted on the floor, where it lay for awhile before it decided to get up again. So, not exactly a Mike Tyson situation, but the machines still managed to evoke a spooky kind of humanoid behavior that made for high-quality entertainment. I overheard an observer quip: “That’s too much like Robocop.”

The dancer

Dancing robots have long been a staple at CES, and this year was no different. This year, the dance-move torch was carried by bots from Unitree, a major Chinese robotics manufacturer that has been scrutinized for potential ties to the Chinese military. Unitree has made a number of impressive announcements about its product base, including a humanoid bot that can supposedly run at speeds of up to 11 mph. I didn’t see any evidence of anything nefarious at Unitree’s booth this week—just a lot of bots that were feeling the groove.

Techcrunch event

San Francisco
|
October 13-15, 2026

The convenience store clerk

I stopped by the booth for Galbot, another Chinese company that says it is focused on multi-modal large language models and general purpose robotics. Galbot’s booth had been styled to look like a convenience store, and its bot appeared to have been synched with a menu app. A customer would come to the booth, select an item from the menu, and then the bot would go and fetch the selected merch for them. After I chose Sour Patch Kids, the bot dutifully retrieved a box off the shelf for me. According to the company’s website, the robot has been deployed in a number of real-world settings, including as an assistant at Chinese pharmacies.

The housekeeper

Creating a machine that can fold laundry has long been one of the core ambitions of the commercial robotics community. The ability to pick up a T-shirt and fold it is considered a fundamental test of automated competence. For that reason, I was fairly impressed by the display over at Dyna Robotics, a firm that develops advanced manipulation models for automated tasks. There, a pair of robotic arms could be seen efficiently folding laundry and placing it in a pile. A Dyna representative told me that the firm had already established partnerships with a number of hotels, gyms, and factories.

One of those businesses, the rep told me, is Monster Laundry, based in Sacramento, California. Monster integrated Dyna’s shirt-folding robot into its operations late last year and now describes itself as the “first laundry center in North America to debut a state-of-the-art robotic folding system from Dyna.” 

Dyna also has some impressive backing. It concluded an $120 million Series A fundraising round in September that included funding from Nvidia’s NVentures, as well as from Amazon, LG, Salesforce, and Samsung.

The butler

I also stopped by LG’s section of CES to take a look at its new home robot, CLOid. It was cute but was not the fastest bot on the block. You can read my full review of that experience here.

Today’s NYT Mini Crossword Answers for Jan. 10


Looking for the most recent Mini Crossword answer? Click here for today’s Mini Crossword hints, as well as our daily answers and hints for The New York Times Wordle, Strands, Connections and Connections: Sports Edition puzzles.


Today’s Mini Crossword is not only the longest of this week, but I believe it’s the toughest. Read on for all the answers. And if you could use some hints and guidance for daily solving, check out our Mini Crossword tips.

If you’re looking for today’s Wordle, Connections, Connections: Sports Edition and Strands answers, you can visit CNET’s NYT puzzle hints page.

Read more: Tips and Tricks for Solving The New York Times Mini Crossword

Let’s get to those Mini Crossword clues and answers.

completed-nyt-mini-crossword-puzzle-for-jan-10-2026.png

The completed NYT Mini Crossword puzzle for Jan. 10, 2026.

NYT/Screenshot by CNET

Mini across clues and answers

1A clue: Pieces of legislation
Answer: ACTS

5A clue: Like blue whales and the Burj Khalifa
Answer: BIG

8A clue: Cuisine with tom yum soup
Answer: THAI

9A clue: This clue number ÷ 9
Answer: ONE

10A clue: Classic internet prank
Answer: RICKROLL

12A clue: Ranked above all others
Answer: ATTHETOP

13A clue: “What was ___ was saying?”
Answer: ITI

14A clue: This clue number – 9
Answer: FIVE

15A clue: Home of MoMA
Answer: NYC

16A clue: Read receipt below an Instagram message
Answer: SEEN

Mini down clues and answers

1D clue: Alphabetically first subway line in 15-Across
Answer: ATRAIN

2D clue: Either blank of “___ ___ Bang Bang”
Answer: CHITTY

3D clue: Bit of strategy
Answer: TACTIC

4D clue: Religious wearer of a turban known as a dastar
Answer: SIKH

5D clue: Baby’s knitted shoe
Answer: BOOTIE

6D clue: Smitten
Answer: INLOVE

7D clue: Writer’s alternative to a ballpoint
Answer: GELPEN

11D clue: Often-heckled sports figures
Answer: REFS


Don’t miss any of our unbiased tech content and lab-based reviews. Add CNET as a preferred Google source.




Top 27 Sites to Hire HVAC CAD Designers & 3D Modeling Drafters for Architectural Projects


Todays post features 27 top site to hire HVAC CAD designers and 3D modeling drafters for your architectural projects. It is not always simple to find the best 3D model drafter or HVAC shop drawing designer for construction work on buildings. It is similar to searching for that particular missing Lego block in the toy mess. You know it is there, but have no clue where it may be. That is where this top website list proves to be useful. They are all highly competent personnel prepared to dissect intricate blueprints into understandable and implementable designs. They include Cad Crowd, which is a promise due to its highly competent freelancers who will be capable of bringing your project ideas to life without the inconvenience of having to surf through what appears to be an infinite amount of surfing.

cadcrowd-logo

1. Cad Crowd

Cad Crowd is where to get the best 3D modeling drafters and CAD HVAC designers. Cad Crowd website offers the customer a pre-screened list of technical precision specialists and creative problem-solving skills specialists. From one HVAC drawing to 3D building models, to complete drafting solutions, Cad Crowd is the quickest solution to obtain the best HVAC design freelancer. Cad Crowd stands out among the rest of the firms in that it is a reliability and quality-based firm which makes every effort to ensure that the correct talent is allocated to the project. Such companies would be ideal for them in view of their fussiness towards creative problem-solving, precision, and timeliness.

Website: Cadcrowd.com

caddrafter logo

2. CADDrafter

CADDrafter provides a low-profile platform through which a person gets enabled to access CAD professionals who provide any drafting service, ranging from HVAC plans to 3D modeling. The company provides technical drawings with precision through exposure to industry standards. Members gain a head start by witnessing professionals with books of experience in the industry, where the HVAC system gets its place in the master plan in case of full building plans. The company is most suitable for companies that need assistance with complex schematics, proofreading, and customized drafting services. Although less exhaustive than Cad Crowd, CADDrafter is a good choice for organizations and companies looking to outsource expert drafting services for their construction and building company.

Website: Caddrafter.us

RELATED: HVAC Duct Shop Drawings: The Complete 2025 Guide for Freelancers and Construction Service Firms

Uppteam logo

3. Uppteam

Uppteam is a website that utilizes freelancing to hire an army of professionals on the payroll, such as CAD designers and draftsmen. Not CAD or HVAC design itself, the site does, however, provide the customer with a means of access to individuals who do have such capability. The job can be put out on bid, bids sought, and a freelancer employed who otherwise would have been sought as an employee. Uppteam can even be used for professional personnel for construction project work done on mechanical system drawings or 3D modeling software. As the site is global and multi-industry, customers prefer to spend time sorting through the choices in attempting to recognize professionals who truly have expertise in HVAC. 

Website: Uppteam.com

kwork logo

4. Kwork

Kwork is one of the freelancing sites where buyers browse rows of professionals’ services. There are indeed CAD and drafting services, but the site is not specialized in offering such services. Advance-booked customers make the process of hiring them simple and quick. If an air con service, Kwork can offer a transport service with the aim of getting CAD drawing drafting specialists or 3D building drawings specialists. The site has ginormous lists of irrelevant services, so there will not be many professional HVAC drafting experts. But one that you can use in case you want a hassle-free and easy recruitment process. 

Website: Kwork.com

Naukri logo

5. Naukri

Naukri is said to be among India’s finest job websites where the employment partners and the employment seekers interact with each other in sizable numbers from various sectors. It is freelance CAD drafting of HVAC or 3D modeling only but there also exists a method for advertising full-time or permanent employment vacancies. Indian companies can utilize it as an entry point for talent pool. To freelance drafter customers who buy on demand, they will never obtain Naukri so welcoming and accurate to use as sites that provide technical design engineering services independently. It suits those companies, which have requirements in-house, need to hire in house employees rather than project-based freelancers.

Website: Naukri.com

Uptalent

 6. Uptalent

Uptalent is a freelancer site where one is exposed to a talent pool of many professionals such as CAD designing and drafting. There is a site where clients list their projects and allow them to be bid upon by freelancers who list their experience and qualifications. Rather than purchasing technical assistance in 3D modeling and HVAC CAD designing, Uptalent can serve as an alternative where one can outsource. But because it is known that the site has sufficient work to keep me busy that renders it wastage of time to be and be too an assessor while seeking specialty HVAC experience, it is an open market for companies who would prefer purchasing skill and pay but is not one-to-one service like Cad Crowd. 

Website: Uptalent.com

RELATED: Relevance of MEP Drafting Services for Architectural Design Firms & Construction Companies

DesigningDraftingServices

7. DesigningDraftingServices

DesigningDraftingServices is an elite firm offering technical drawing and drafting services of the highest quality to mechanical engineering design, construction, and architecture clients. It offers tailored services such as 3D architecture models and HVAC design. It offers customers with customers who need the best schematics that only insert into completed projects. It accomplishes this in accuracy and according to specifications in the widest meaning of the term. 

DesignDraftingServices is for companies needing professional work in urgency mode rather than window shopping in a bazaar. Though less varied in freelance service providers than Cad Crowd, its past experience in providing drafts of the service would compel architecture firms to be compelled to use it for precise, precise, and customized technical reports. 

Website: Designingdrafting.com

CADhero

8. CADHero

CADHero produces a drafting and design firm whose service is actionable to use in engineering, construction, and architecture. It provides skilled professionals with the art of modeling HVAC systems and precise CAD drawing that can be integrated into master building plans. Clients require professional work from conceptual drawings to highly detailed specifications for projects. CADHero provides personal one-on-one service, and customers are able to talk to draftsmen directly in their efforts to communicate requirements. It is not as big in professional staff as Cad Crowd but is a great 3D architectural drafting firm. It is perfect for customers that require professional assistance and precise 2D and 3D modeling for architectural design and mechanical engineering design. 

Website: Bluebyte.biz

enginerio logo

9. Enginerio

Enginerio is an engineering outsourcing company with design and drafting capabilities to serve the needs of many different industries. As an electrically, mechanically, and building systems-capable firm, it also has draft support services for HVAC in 3D modeling. There are customers who are able to outsource the skill of veteran lead drafters and design engineering professionals in an effort to achieve form and function when producing drafts. Enginerio is appropriate for organizations that would never hire a lone freelancer but would outsource to specialized firms. Though smaller than Cad Crowd, Enginerio possesses the technical draft and design capability of firms seeking such capability. 

Website: Enginerio.com

Australian Design Drafting Services

10. Australian Design & Drafting Services

Australian Design & Drafting Services is an Australian company providing bespoke architectural, civil engineering, and mechanical designs. It is an Australian business company whose customers require 3D modeling and HVAC CAD drafting service adequate for regional legislation and law. It is best suited to provide technical drawings adequate to fulfill the needs of the construction industry. The technical drawings are guaranteed by clients to provide full documentation, one-to-one service, and on-time turnover. Though it does not have Cad Crowd’s global talent pool of freelancers, the service is more than valuable for Australian businesses seeking skilled local services to deliver design for HVAC, CAD drafting, and project assistance. 

Website: Astdcad.com.au

RELATED: Differences between 3D Modeling, CAD, and BIM Explained

StartCAD

11. StartCAD 

StartCAD is an organization that offers professional CAD and drafting services for building work, engineering, and architecture clients. It offers solutions for 3D modeling, mechanical drawing system design, and HVAC design. StartCAD is a service one can rely on to provide accurate technical drawings that will be part of larger building plans. Even though as good as the service is, its pool of talent to work with is not as large as Cad Crowd, whose larger talent pool of experienced experts are at your beck and call. StartCAD would suit best companies who would like to have direct access to professional engineers and drafters who will industry standard target-design and model support. 

Website: Startcad.com

PZ Drafting

12. PZ Drafting 

PZ Drafting offers CAD drafting, mechanical, and architectural design services. Technical drawing, 3D modeling, and HVAC system designing is offered. Professional schematics and project specification suggestions are offered to the customers. While augmented by reliable draft services, the site is not welcoming in the global freelance base and talent pool to Cad Crowd. It is best suited for clients that are being driven toward hiring an efficient small group to create satisfactory technical output. Companies that must be able to come back with elastic options will enjoy the convenience and predisposition of massive freelance sites. 

Website: Pzdrafting.com

Chemionix

13. Chemionix 

Chemionix offers technological accuracy in construction drawing and construction engineering. Chemionix can deliver HVAC CAD drafting, 3D building models, and mechanical drawings with minute details. The clients have direct access to the 2D and 3D drafting experts with deep experience of the finer aspects of construction and specification drafting. Chemionix does not enjoy access and freelancer numbers of network Cad Crowd. The service works best for firms that would prefer an in-house team of drafts to provide on a project-by-project basis rather than with heaps of freelancers. Chemionix is a great piece of software but never left the world. 

Website: Chemionix.com

CAD Outsourcing

14. CADOutsourcing 

CADOutsourcing provides drafting business solutions and CAD to architects, building contractors, and engineers. CADOutsourcing allows preparation of 3D modeling and HVAC design and technical drawing. The firm is most suitable where it is a matter of meeting industry and precision requirements, and on the most complex work requiring skilled hands. Even though it offers good blueprints, CADOutsourcing is behind Cad Crowd in alternatives as well as 3D design freelancers. Those looking for alternatives or need a right freelancer within a timeline can use other websites. CADOutsourcing is suitable for businesses requiring regular service from a professional vendor who will produce technical drawings and modeling according to project specs-and standards-of quality.

Website: Cadoutsourcing.com

RELATED: Fundamentals of BIM & Modeling Design Services at Building Information Modeling Companies

Gsource-Technologies

15. Gsource Technologies

Gsource Technologies is an organization offering CAD drafting and 3D modeling solutions for building construction and engineering projects. It offers mechanical blueprints, HVAC, and accurate technical drawings. Its customers require accurate professional work to be done by skilled drafters & detailers. While Gsource Technologies is of higher precision, it is disappointing in comparison to Cad Crowd in freelancers’ numbers and world presence. It is not perfect for clients who want more than one freelancer or faster uploading of a project. Gsource Technologies is ideal for organizations that need single specialization of an activity, especially organizations where its area of focus is technical precision and touch feel with a vision drafting department. 

Website: Gsourcedata.com

Archdraw Outsourcing

16. Archdraw Outsourcing

Archdraw Outsourcing offers draftings and engineering project drawings and CAD professional services. Archdraw offers HVACS plans, 3D modeling, and technically correct plans. Precise and accurate schematics to industry standard that can improve building work can be offered to customers. Though for a low price it will be challenging to compete with Archdraw Outsourcing quality work, not always skill to compete with Cad Crowd who like having ginormous pool of pre-screened solo experts from around the world. Best value in Archdraw Outsourcing is thus to organizations who would like to have a specific provider and not the freelancers themselves. It works best with some project draft work on accuracy and reliability. 

Website: Arcdrawoutsourcing.com

Tawoostechcom

17. Al Tawoos Al Abyadh Tech Cont 

Al Tawoos Al Abyadh Tech Cont is a technologically enhanced CAD and drafting in engineering and design solution provider company. The company can provide HVAC plan, 3D modeling, and schematics with total commitment. The clients have the luxury of talking to the experts themselves who are handier in the backend of standards and accuracy. The premium nature of service cannot beat Cad Crowd’s global talent pool and marketplaces. Companies seeking the other option of freelancers or in-direct job posting may be disappointed on the platform. Al Tawoos Al Abyadh Tech Cont is best utilized for compliant work involving fewer employees and non-scalable or non-diversifiable architectural drawing needs. 

Website: Tawoostech.com

Seekcom

18. SEEK 

SEEK is an open website for jobs and does not employ freelance HVAC CAD designers and 3D modeling drafters for their firm. It is designed with regard to full-time and permanent employment in every business. One-off or temporary rough draft project-type work requested by the applicants cannot create the capacity required within the time frame. The platform is better suited for standard job searching, and it is not an immediate tech contract roll-out. SEEK is not presenting a best-of-breed freelancer website, and its ads are not traditionally formatted in the technical skill set deserving of architectural HVAC drafting. For high-end CAD firms who are seeking services, the right websites like Cad Crowd are much better and safer. 

Website: Seek.com

RELATED: MEP Shop Drawing Services: The Secret Weapon for Freelancers and Firms in Successful Building Construction 

peopleperhour logo

19. PeoplePerHour 

PeoplePerHour is a generic freelance site for any firm but not for HVAC CAD and 3D building modeling. There may be best freelancers at hand, but pools of talent change and skill tests aren’t fixed. There is much hunting and digging required to find premium-level talent who can carry out premium-quality HVAC drafting or premium-quality 3D modeling. The site is tolerable for vanilla freelance work but not yet tolerable for high-stakes architecture work with accuracy and technical skills. Companies needing quality pre-screened draft service are best served on expert sites such as Cad Crowd, which involve design and engineering professionals

Website: Peopleperhour.com

kolabtree logo

20. Kolabtree 

Kolabtree would be the most relevant to client-specialist and research collaborations, but technical draft specialist collaborations. Much less likely, however, would be using advanced 3D modelers or HVAC CAD drafters on building projects. While the website can be used as a reference for work on data analysis or research, but not for architectural or engineering blueprints. Customers to be recruited with project files would barely be provided with such profiles. The website is not so valuable for CAD architecture, and businesses in need of detailed HVAC floor planning as well as 3D modeling will find more satisfactory results on websites like Cad Crowd. 

Website: Kolabtree.com

Guru logo

21. Guru 

Guru is an open freelancers website with a good number of different services but no particular experience in 3D HVAC CAD and architectural modeling under strict control. Freelancers who lose faith in the capability of freelancers fetch freelancers hunters. Freelance work on model designing or HVAC designing cannot be promised to get done. Freelance work of any kind but independent draft work of architecture would suit Guru. Companies necessitating precision and pre-shortlisted experts are better placed on platforms such as Cad Crowd where highly restricted experts are accorded the advantage of technical design and modeling skills. 

Website: Guru.com

indeedcom logo

22. Indeed 

Indeed is a very specialized full-time and contract workers job board but not for freelance job opening employment. The clients that are seeking HVAC CAD designers or drafters that have the capability to perform 3D modeling for specialty or temporary services will not find Indeed to be of much use. The website can only be used to look for full-time work and not freelance specialty work, thus taking eons to acquire drafting expertise. Companies that require accuracy, knowledge, and round-the-clock uptime for technical services cannot obtain the required staff. These platforms will never have to contend with draft-based project specifications, but standalone markets such as Cad Crowd offer direct access to specialist professionals with the required expertise and comprehensive portfolios. 

Website: Indeed.com

RELATED: Impacts of BIM Design on Reducing Carbon Footprint for Architectural Firms & CAD Companies

Truelancer logo

23. Truelancer 

Truelancer is an open freelance platform offering services in a broad category but not architectural or engineering CAD designers. Though one can locate some of the drafters providing drafting services, it is not possible to ensure they are skilled in 3D modeling or HVAC. It is not an environment that will offer full architectural services where accuracy matters. Those companies requiring professional and skilled CAD personnel might take time and have mediocre results. For top-level drafting, the best for it to happen is on websites such as Cad Crowd where customers truly know freelancers hired who have considerable work experience done in HVAC model creation and designing. 

Website: Truelancer.com

Freelancer

24. Freelancer 

Freelancer.com is an open website with some freelance categories of employment, i.e., a number of CAD services but no HVAC CAD or for one 3D architecture modeling. Freelancers may do their best to see that they have the ability but quality checking and warrantees are unthinkable. Technical accuracy, proper schematics, and trade specifications, dandy work cannot be assured at all times. Freelancer is fine for run-of-the-mill freelance services but mission-critical schematics. Firms searching for professional HVAC CAD or 3D architecture visualization service will do best at specialty stores like Cad Crowd with maximum competency and adequate experience. 

Website: Freelancer.com

ZipRecruiter Logo

25. ZipRecruiter

ZipRecruiter is longer-term job career job promotion and shorter freelance- or project-hire work. Firms looking to employ HVAC CAD drafters or 3D drafting modelers to complete construction projects will not be blessed with the luxury of having the ability to count on the option of being able to rely on skilled temporary labor. The site has no experienced or independently managed pool of freelancers and is far too unrepresentative for technical-level or short-term drafting work.

Although suitable for long-term investment, companies that require instant access to experienced freelancers will be at a loss. Specialized sites such as Cad Crowd are more appropriate due to the immediate access provided to experienced specialists with credentials available for 3D architectural modeling and HVAC design. 

Website: Ziprecruiter.ie

Upwork-logo

26. Upwork 

Upwork is a gigantic freelancing platform with numerous categories but with no 3D modeling or HVAC CAD to play with in architecture. There are a few good freelancers, but good professionals need to be hunted thoroughly to find them. Pretty big projects that need to be worked on exactly and to specifications may not necessarily be done flawlessly. Upwork is okay for ad hoc freelance jobs but by no means technical drafting where precision is needed. Businesses would do well to hire them on platforms such as Cad Crowd with membership restricted to pre-approved professionals who have some technical design, HVAC drafting, and 3D modeling experience on their résumés. 

Website: Upwork.com

RELATED: Architectural Plans, CAD Drawing Costs & Architect Service Pricing: Full Breakdown

fiverr-logo

27. Fiverr 

Fiverr is a high-volume freelancing platform with an overwhelmingly enormous list of services that have no or minimal relation to one another, typically in infinitesimally small quantities. Whereas other independent contractors may be qualified to do CAD or 3D modeling, the platform is not architectural drafting- or HVAC drafting-based. Possessing technical expertise, precision, and complying with regulations for work can be targeted by the cracks. Micro or technical work is best done on Fiverr than high-end HVAC drafting or high-end 3D modeling. Companies that need top-notch skilled technical engineers to perform architectural tasks will perform far better on specialty boards like Cad Crowd where they can have A-list pre-screened technical experts on their payroll. 

Website: Fiverr.com

The bottom line 

All architectural projects always have problems that they carry with them, but not when you utilize the services of the perfect HVAC CAD designer or 3D modeling drafter. With all of these sites at your fingertips, now you have your own way of being able to get to some of the most sought-out sites to get the information to the experts and allow time for the overall picture. 

How Cad Crowd can help

From developing hundreds of HVAC floor plans to precise 3D models, we have a lot of professionals waiting to get contracted and assist. In case you don’t want to test and re-do, then visit Cad Crowd to get HVAC CAD designers & 3D modeling drafters for architectural work and start your project with the right professional today. Contact us for a free quote.

author avatar

MacKenzie Brown is the founder and CEO of Cad Crowd. With over 18 years of experience in launching and scaling platforms specializing in CAD services, product design, manufacturing, hardware, and software development, MacKenzie is a recognized authority in the engineering industry. Under his leadership, Cad Crowd serves esteemed clients like NASA, JPL, the U.S. Navy, and Fortune 500 companies, empowering innovators with access to high-quality design and engineering talent.

Connect with me: LinkedInXCad Crowd

Introducing the VS Code Insiders Podcast


December 1, 2025 by James Montemagno, @jamesmontemagno.

VS Code Podcast Hero Banner - A modern podcast studio with VS Code logo, microphones, and code flowing in the background

Ever wonder what goes on behind the scenes at the world’s most popular code editor? The VS Code Insiders Podcast is here to pull back the curtain and give you an insider’s look at the features, decisions, and people shaping the future of Visual Studio Code.

What’s the podcast all about?

The official podcast from the Visual Studio Code team goes beyond the release notes. Each episode features conversations with the developers, product managers, and community contributors who are actively building VS Code’s next generation of features.

From deep dives into experimental tooling and new extensions to candid discussions about where coding is headed with AI, this podcast is your backstage pass to the evolution of modern development. Whether you’re a seasoned developer, a curious tinkerer, or just obsessed with clean commits and clever workflows, there’s something here for you.

Recent episode highlights

Here are a few recent episodes that showcase some of the exciting topics covered:

Subscribe today!

Visit the VS Code Insiders Podcast home page to browse past episodes and subscribe on your favorite podcast app today!

Download and install VS Code Insiders to try the newest features as soon as they ship.

Stay up to date with the latest changes in the continuously updating Insiders release notes.

Happy coding.

Vmake AI Agent Review: The Easy Way to Edit Videos in 2026


A person edits video on a large monitor featuring an AI assistant, surrounded by cameras and audio equipment in a modern workspace.

In the past, video content used to be a luxury. Now, it is the baseline. If you run a business or build a brand, you are in the video business. The problem is that traditional editing hasn’t really caught up with this reality. It is still slow, technical, and you even have to worry about frame rates and codecs. This is the friction point that Vmake AI targets. It isn’t trying to be a better version of Adobe Premiere. It is trying to be a production assistant that handles the tedious parts of the job for you.

The most common starting point for using these tools isn’t usually creation; it’s repair. You often have a video that works, but it’s trapped in the wrong format or covered in old branding. Maybe you downloaded a clip from TikTok, and it has that bouncing logo. To use it on Instagram, you need to clean it up. 

A reliable way to remove watermark from video files is essential here. It analyzes the pixels around the logo and fills in the gaps. It turns a “dead” file into a usable asset without forcing you to crop out half the frame.

What is Vmake AI Agent?

Many other AI Agents are complex and wait for you to click a button. While Vmake introduces a “Video Agent” that works differently. It feels less like using a tool and more like giving instructions to a freelancer. You don’t hunt for the “crop” tool or the “caption” button. Instead, just use a chat window.

You can paste a URL to a product page and tell the Agent to “make a promotional video.” It reads the page. It pulls the product images. It writes a script based on the selling points. This changes the workflow from “how do I do this?” to “what do I want?” It handles the technical translation. You provide the intent; the machine handles the execution. 

This is particularly useful for User Generated Content (UGC) styles. The system is able to generate speaking videos of people, using AI avatars to deliver the script. Plus it looks pretty natural. 

How To Use Vmake AI Agent For Workflow?

Talking about features is abstract. Seeing how the work actually flows makes it clearer. Here is what a typical session looks like when you move from a raw idea to a finished post.

  1. Briefing the System: You start in the chat interface. You don’t need raw footage yet. If you are selling a backpack, you drop the link to your store. You type a command: “Create a fast-paced video for TikTok focusing on durability.” The Agent scans the link. It identifies the key features—waterproof material, reinforced zippers, and hidden pockets. It generates a storyboard and a script.
  2. The Generation Phase: The system builds the video. It selects stock footage or uses the images from your URL. It adds a voiceover. This takes a minute or two. The result is a draft. It might be 80% there. The tool gets the timing right, picks the music, and structures the narrative according to your needs.
  3. Refining in Canvas: No AI is perfect. You will likely want to change something. This is where you switch to the Canvas editor. It looks more like a standard editor, but simplified. You might swap out an image that doesn’t fit. You might tweak the voiceover speed. This step is about polishing. You take the “good” draft and make it “great.”
  4. Captions and Optimization: Social video needs text. Most people watch on mute. The Auto Captions feature comes in handy here. You can use it to generate captions for your videos. Moreover, you can also pick the style you want. 
  5. Final Packaging The last step is the thumbnail. A bad thumbnail kills a video before it starts. The AI analyzes your video and generates a few options. It finds the most interesting frame. It adds a hook title like “Best Travel Hack?” You pick the winner and export.

Using Hooks to Improve Viewer Engagement

Getting a video made is only half the battle. Getting people to watch it is the other half. The platform includes specific tools for this, focusing on the “Hook.” This is the first three seconds of the video. It is the make-or-break moment.

The AI Hook tool generates specific opening clips designed to stop the scroll. It uses visual effects and motion to grab attention immediately. It isn’t guessing; it uses data on what actually works in performance marketing. If you are running ads, this is crucial. You can test three different hooks for the same video to see which one drives more views.

Using Auto Captions for the Audience That Watches Videos Without Audio

Transcribing audio manually is painful. It takes forever. You have to type, pause, rewind, type again. Vmake automates this for seven languages. It handles English, Spanish, Portuguese, and others.

This isn’t just about translation. It is about accessibility. If your video doesn’t have text, you are ignoring the commuter on the train who doesn’t have headphones. The tool ensures your message lands even when the volume is zero. It also allows for batch processing. If you have ten videos from a campaign, you don’t have to caption them one by one. You can upload the whole folder and let the system run.

Who Vmake Is Designed For?

This isn’t for filmmakers. If you are editing a documentary, you want manual control. You want to color grade every pixel. This tool is for marketers and for the solo entrepreneur who needs to post three times a day. It is for the e-commerce store owner who needs product videos for fifty items.

It democratizes the “agency” look. In the past, getting a polished video, having captions, and 4K quality required hiring freelancers. But now? It’s simple to use tools like Vmake.  Vmake is an efficient video editor to edit videos online, and it is available in multiple video editing features, including video enhancer, watermark remover, auto caption generator,  hook generator and more. 

Final Thoughts

The biggest change here isn’t any single feature. It is the logic of how people create. The world is moving away from manual assembly and towards a direction where you act less like an editor and more like a producer. You verify the work rather than doing it from scratch.

Vmake represents this shift clearly. It bundles the strategy (script generation), the production (video creation), and the post-production (editing and captioning) into one flow. It removes the friction. It lets you focus on the story you want to tell, rather than the software you need to learn.

ExpressVPN two-year plans are up to 78 percent off right now


ExpressVPN is back on sale again, and its two-year plans are up to 78 percent off right now. You can get the Advanced tier for $101 for 28 months. This is marked down from the $392 that this time frame normally costs. On a per-month basis, it works out to roughly $3.59 for the promo period.

Image for the large product module

ExpressVPN

We’ve consistently liked ExpressVPN because it’s fast, easy to use and widely available across a large global server network. In fact, it’s our current pick for best premium VPN. One of the biggest drawbacks has always been its high cost, and this deal temporarily solves that issue.

In our review we were able to get fast download and upload speeds, losing only 7 percent in the former and 2 percent in the latter worldwide. We found that it could unblock Netflix anywhere, and its mobile and desktop apps were simple to operate. We gave ExpressVPN an overall score of 85 out of 100.

The virtual private network service now has three tiers. Basic is cheaper with fewer features, while Pro costs more and adds extra perks like support for 14 simultaneous devices and a password manager. Advanced sits in the middle and includes the password manager but only supports 12 devices.

The Basic plan is $78 right now for 28 months, down from $363, and the Pro plan is $168, down from $560. That’s 78 percent and 70 percent off, respectively. All plans carry a 30-day money-back guarantee for new users, so you can try it without committing long term if you’re on the fence.

Follow @EngadgetDeals on X for the latest tech deals and buying advice.



Top 10 Open-source Reasoning Models in 2026


Introduction 

AI in 2026 is shifting from raw text generators to agents that act and reason. Experts predict a focus on sustained reasoning and multi-step planning in AI agents. In practice, this means LLMs must think before they speak, breaking tasks into steps and verifying logic before outputting answers. Indeed, recent analyses argue that 2026 will be defined by reasoning-first LLMs-models that intentionally use internal deliberation loops to improve correctness. These models will power autonomous agents, self-debugging code assistants, strategic planners, and more. 

At the same time, real-world AI deployment now demands rigor: “the question is no longer ‘Can AI do this?’ but ‘How well, at what cost, and for whom?’”. Thus, open models that deliver high-quality reasoning and practical efficiency are critical.

Reasoning-centric LLMs matter because many emerging applications- from advanced QA and coding to AI-driven research-require multi-turn logical chains. For example, agentic workflows rely on models that can plan and verify steps over long contexts. Benchmarks of 2025 show that specialized reasoning models now rival proprietary systems on math, logic, and tool-using tasks. In short, reasoning LLMs are the engines behind next-gen AI agents and decision-makers.

In this blog, we will explore the top 10 open-source reasoning LLMs of 2026, their benchmark performance, architectural innovations, and deployment strategies.

What Is a Reasoning LLM?

Reasoning LLMs are models tuned or designed to excel at multi-step, logic-driven tasks (puzzles, advanced math, iterative problem-solving) rather than one-shot Q&A. They typically generate intermediate steps or thoughts in their outputs.

For instance, answering “If a train goes 60 mph for 3 hours, how far?” requires computing distance = speed×time before answering-a simple reasoning task. A true reasoning model would explicitly include the computation step in its response. More complex tasks similarly demand chain-of-thought. In practice, reasoning LLMs often have thinking mode: either they output their chain-of-thought in text, or they run hidden iterations of inference internally.

Modern reasoning models are those refined to excel at complex tasks best solved with intermediate steps, such as puzzles, math proofs, and coding challenges. They typically include explicit reasoning content in the response. Importantly, not all LLMs need to be reasoning LLMs: simpler tasks like translation or trivia don’t require them. In fact, using a heavy reasoning model everywhere can be wasteful or even “overthinking.” The key is matching tools to tasks. But for advanced agentic and STEM applications, these reasoning-specialist LLMs are essential.

Architectural Patterns of Reasoning-First Models

Reasoning LLMs often employ specialized architectures and training:

  • Mixture of Experts (MoE): Many high-end reasoning models use MoE to pack trillions of parameters while activating only a fraction per token. For example, Qwen3-Next-80B activates only 3B parameters via 512 experts, and GLM-4.7 is 355B total with ~32B active. Moonshot’s Kimi K2 uses ~1T total parameters (32B active) across 384 experts. Nemotron 3 Nano (NVIDIA) uses ~31.6B total (3.2B active, via a hybrid MoE Transformer). MoE allows huge model capacity for complex reasoning with lower per-token compute.
  • Extended Context Windows: Reasoning tasks often span long dialogues or documents. Thus many models natively support huge context sizes (128K-1M tokens). Kimi K2 and Qwen-coder models support 256K (extensible to 1M) contexts. LLaMA 3.3 extends to 128K tokens. Nemotron-3 supports up to 1M context length. Long context is crucial for multi-step plan tracking, tool history, and document understanding.
  • Chain-of-Thought and Thinking Modes: Architecturally, reasoning LLMs often have explicit “thinking” modes. For example, Kimi K2 only outputs in a “thinking” format with <think>…</think> blocks, enforcing chain-of-thought. Qwen3-Next-80B-Thinking automatically includes a <think> tag in its prompt to force reasoning mode. DeepSeek-V3.2 exposes an endpoint that by default produces an internal chain of thought before final answers. These modes can be toggled or controlled at inference time, trading off latency vs. reasoning depth.
  • Training Techniques: Beyond architecture, many reasoning models undergo specialized training. OpenAI’s gpt-oss-120B and NVIDIA’s Nemotron all use RL from feedback (often with math/programming rewards) to boost problem-solving. For example, DeepSeek-R1 and R1-Zero were trained with large-scale RL to directly optimize reasoning capabilities. Nemotron-3 was fine-tuned with a mix of supervised fine-tuning (SFT) on reasoning data and multi-environment RL . Qwen3-Next and GPT-OSS both adopt “thinking” training where the model is explicitly trained to generate reasoning steps. Such targeted training yields markedly better performance on reasoning benchmarks.
  • Efficiency and Quantizations: To make these large models practical, many use aggressive quantization or distillation. Kimi K2 is natively INT4-quantized. Nemotron Nano was post-quantized to FP8 for faster throughput. GPT-OSS-20B/120B are optimized to run on commodity GPUs. Moonshot’s MiniMax also emphasizes an “efficient design”: only 10B activated parameters (with ~230B total) to fit complex agent tasks.

Collectively, these patterns – MoE scaling, huge contexts, chain-of-thought training, and careful tuning – define today’s reasoning LLM architectures.

1. GPT-OSS-120B

GPT-OSS-120B is a production-ready open-weight model released in 2025.  It uses a Mixture-of-Experts (MoE) design with 117B total / 5.1B active parameters. 

GPT-OSS-120B achieves near-parity with OpenAI’s o4-mini on core reasoning benchmarks, while running on a single 80GB GPU. It also outperforms other open models of similar size on reasoning and tool use.

 

It also comes in a 20B version optimized for efficiency: the 20B model matches o3-mini and can run on just 16GB of RAM, making it ideal for local or edge use. Both models support chain-of-thought with <think> tags and full tool integration via APIs. They support high instruction-following quality and are fully Apache-2.0 licensed.

Key specs: 

Variant

Total Params

Active Params

Min VRAM (quantized)

Target Hardware

Latency Profile

gpt-oss-120B

117B

5.1B

80GB

1x H100/A100 80GB

180-220 t/s ​

gpt-oss-20B

21B

3.6B

16GB

RTX 4070/4060 Ti

45-55 t/s ​

 

Strengths and Limits

  • Pros: Near-proprietary reasoning (AIME/GPQA parity), single-GPU viable, full CoT/tool APIs for agents.
  • Cons: 120B deploy still needs tensor-parallel for <80GB setups; community fine-tunes nascent; no native image/vision.

Optimized for latency

  •  GPT-OSS-120B can run on 1×A100/H100 (80GB), and OSS-20B on a 16GB GPU.
  •  Strong chain-of-thought & tool use support.

2. GLM-4.7

GLM-4.7  is a 355B-parameter open model with task-oriented reasoning enhancements. It was designed not just for Q&A but for end-to-end agentic coding and problem-solving. GLM-4.7 introduces “think-before-acting” and multi-turn reasoning controls to stabilize complex tasks. For example, it implements “Interleaved Reasoning”, meaning it performs a chain-of-thought before every tool call or response. It also has “Retention-Based” and “Round-Level” reasoning modes to keep or skip inner monologue as needed. These features let it adaptively trade latency for accuracy.

Performance‑wise,  GLM‑4.7 leads open-source models across reasoning, coding, and agent tasks. On the Humanity’s Last Exam (HLE) benchmark with tool use, it scores ~42.8 %, a significant improvement over GLM‑4.6 and competitive with other high-performing open models. In coding, GLM‑4.7 achieves ~84.9 % on LiveCodeBench v6 and ~73.8 % on SWE-Bench Verified, surpassing earlier GLM releases.

The model also demonstrates robust agent capability on benchmarks such as BrowseComp and τ²‑Bench, showcasing multi-step reasoning and tool integration. Together, these results reflect GLM-4.7’s broad capability across logic, coding, and agent workflows, in an open-weight model released under the MIT license.

Key Specs

  • Architecture: Sparse Mixture-of-Experts
  • Total parameters: ~355B (reported)
  • Active parameters: ~32B per token (reported)
  • Context length: Up to ~200K tokens
  • Primary use cases: Coding, math reasoning, agent workflows
  • Availability: Open-weight; commercial use permitted (license varies by release)

Strengths

  • Strong performance in multi-step reasoning and coding
  • Designed for agent-style execution loops
  • Long-context support for complex tasks
  • Competitive with leading open reasoning models

Weaknesses

  • High inference cost due to scale
  • Advanced reasoning increases latency
  • Limited English-first documentation

3. Kimi K2 Thinking

Kimi K2 Thinking is a trillion-parameter Mixture-of-Experts model designed specifically for deep reasoning and tool use. It features approximately 1 trillion total parameters but activates only 32 billion per token across 384 experts. The model supports a native context window of 256K tokens, which extends to 1 million tokens using Yarn. Kimi K2 was trained in INT4 precision, delivering up to 2x faster inference speeds.

The architecture is fully agentic and always thinks first. According to the model card, Kimi K2-Thinking only supports thinking mode, where the system prompt automatically inserts a <think> tag. Every output includes internal reasoning content by default.

Kimi K2 Thinking leads across the shown benchmarks, scoring 44.9% on Humanity’s Last Exam, 60.2% on BrowseComp, and 56.3% on Seal-0 for real-world information collection. It also performs strongly in agentic coding and multilingual tasks, achieving 61.1% on SWE-Multilingual, 71.3% on SWE-bench Verified, and 83.1% on LiveCodeBench V6.

Overall, these results show Kimi K2 Thinking outperforming GPT-5 and Claude Sonnet 4.5 across reasoning, agentic, and coding evaluations.

Key Specs

  • Architecture: Large-scale MoE
  • Total parameters: ~1T (reported)
  • Active parameters: ~32B per token
  • Experts: 384
  • Context length: 256K (up to ~1M with scaling)
  • Primary use cases: Deep reasoning, planning, long-context agents
  • Availability: Open-weight; commercial use permitted

Strengths

  • Excellent long-horizon reasoning
  • Very large context window
  • Strong tool-use and planning capability
  • Efficient inference relative to total size

Weaknesses: 

  • Truly enormous scale (1T) means daunting training/inference overhead. 
  • Still early (new release), so real-world adoption/tooling is nascent.

4. MiniMax-M2.1

MiniMax-M2.1 is another agentic LLM geared toward tool-interactive reasoning. It uses a 230B total param design with only 10B activated per token, implying a large MoE or similar sparsity. 

The model supports interleaved reasoning and action, allowing it to reason, call tools, and react to observations across extended agent loops. This makes it well-suited for tasks involving long sequences of actions, such as web navigation, multi-file coding, or structured research tasks.

MiniMax reports strong internal results on agent benchmarks such as SWE-Bench, BrowseComp, and xBench. In practice, M2.1 is often paired with inference engines like vLLM to support function calling and multi-turn agent execution.

Key Specs

  • Architecture: Sparse, agent-optimized LLM
  • Total parameters: ~230B (reported)
  • Active parameters: ~10B per token
  • Context length: Long context (exact size not publicly specified)
  • Primary use cases: Tool-based agents, long workflows
  • Availability: Open-weight (license details limited)

Strengths

  • Purpose-built for agent workflows
  • High reasoning efficiency per active parameter
  • Strong long-horizon task handling

Weaknesses

  • Limited public benchmarks and documentation
  • Smaller ecosystem than peers
  • Requires optimized inference setup

5. DeepSeek-R1-Distill-Qwen3-8B

DeepSeek-R1-Distill-Qwen3-8B represents one of the most impressive achievements in efficient reasoning models. Released in May 2025 as part of the DeepSeek-R1-0528 update, this 8-billion parameter model demonstrates that advanced reasoning capabilities can be successfully distilled from massive models into compact, accessible formats without significant performance degradation.

The model was created by distilling chain-of-thought reasoning patterns from the full 671B parameter DeepSeek-R1-0528 model and applying them to fine-tune Alibaba’s Qwen3-8B base model. This distillation process used approximately 800,000 high-quality reasoning samples generated by the full R1 model, focusing on mathematical problem-solving, logical inference, and structured reasoning tasks. The result is a model that achieves state-of-the-art performance among 8B-class models while requiring only a single GPU to run.

Performance-wise, DeepSeek-R1-Distill-Qwen3-8B delivers results that defy its compact size. It outperforms Google’s Gemini 2.5 Flash on AIME 2025 mathematical reasoning tasks and nearly matches Microsoft’s Phi 4 reasoning model on HMMT benchmarks. Perhaps most remarkably, this 8B model matches the performance of Qwen3-235B-Thinking on certain reasoning tasks—a 235B parameter model. The R1-0528 update significantly improved reasoning depth, with accuracy on AIME 2025 jumping from 70% to 87.5% compared to the original R1 release.

The model runs efficiently on a single GPU with 40-80GB VRAM (such as an NVIDIA H100 or A100), making it accessible to individual researchers, small teams, and organizations without massive compute infrastructure. It supports the same advanced features as the full R1-0528 model, including system prompts, JSON output, and function calling—capabilities that make it practical for production applications requiring structured reasoning and tool integration.

Key Specs

  • Model type: Distilled reasoning model
  • Base architecture: Qwen3-8B (dense transformer)
  • Total parameters: 8B
  • Training approach: Distillation from DeepSeek-R1-0528 (671B) using 800K reasoning samples
  • Hardware requirements: Single GPU with 40-80GB VRAM
  • License: MIT License (fully permissive for commercial use)
  • Primary use cases: Mathematical reasoning, logical inference, coding assistance, resource-constrained deployments

Strengths

  • Exceptional performance-to-size ratio: matches 235B models on specific reasoning tasks at 8B size
  • Runs efficiently on single consumer-grade GPU, dramatically lowering deployment barriers
  • Outperforms much larger models like Gemini 2.5 Flash on mathematical reasoning
  • Fully open-source with permissive MIT licensing enables unrestricted commercial use
  • Supports modern features: system prompts, JSON output, function calling for production integration
  • Demonstrates successful distillation of advanced reasoning from massive models to compact formats

Weaknesses

  • While impressive for its size, still trails the full 671B R1 model on the most complex reasoning tasks
  • 8B parameter limit constrains multilingual capabilities and broad domain knowledge
  • Requires specific inference configurations (temperature 0.6 recommended) for optimal performance
  • Still relatively new (May 2025 release) with limited production battle-testing compared to more established models

6. DeepSeek-V3.2 Terminus

DeepSeek’s V3 series (codename Terminus”) builds on the R1 models and is designed for agentic AI workloads. It uses a Mixture-of-Experts transformer with ~671B total parameters and ~37B active parameters per token.

DeepSeek-V3.2 introduces a Sparse Attention architecture for long-context scaling. It replaces full attention with an indexer-selector mechanism, reducing quadratic attention cost while maintaining accuracy close to dense attention.

As shown in the below figure, the attention layer combines Multi-Query Attention, a Lightning Indexer, and a Top-K Selector. The indexer identifies relevant tokens, and attention is computed only over the selected subset, with RoPE applied for positional encoding.

The model is trained with large-scale reinforcement learning on tasks such as math, coding, logic, and tool use. These skills are integrated into a shared model using Group Relative Policy Optimization.

                                  Fig- Attention-architecture of deepseek-v3.2

DeepSeek reports that V3.2 achieves reasoning performance comparable to leading proprietary models on public benchmarks. The V3.2-Speciale variant is further optimized for deep multi-step reasoning.

DeepSeek-V3.2 is MIT-licensed, available via production APIs, and outperforms V3.1 on mixed reasoning and agent tasks.

Key specs

  • Architecture: MoE transformer with DeepSeek Sparse Attention
  • Total parameters: ~671B (MoE capacity)
  • Active parameters: ~37B per token
  • Context length: Supports extended contexts up to ~1M tokens with sparse attention
  • License: MIT (open-weight)
  • Availability: Open weights + production API via DeepSeek.ai

Strengths

  • State-of-the-art open reasoning: DeepSeek-V3.2 consistently ranks at the top of open-source reasoning and agent tasks.
  • Efficient long-context inference: DeepSeek Sparse Attention (DSA) reduces cost growth on very long sequences relative to standard dense attention without significantly hurting accuracy.
  • Agent integration: Built-in support for thinking modes and combined tool/chain-of-thought workflows makes it well-suited for autonomous systems.
  • Open ecosystem: MIT license and API access via web/app ecosystem encourage adoption and experimentation. 

Weaknesses

  • Large compute footprint: Despite sparse inference savings, the overall model size and training cost remain significant for self-hosting.
  • Complex tooling: Advanced thinking modes and full agent workflows require expertise to integrate effectively.
  • New release: As a relatively recent generation, broader community benchmarks and tooling support continue to mature.

7. Qwen3-Next-80B-A3B

Qwen3-Next is Alibaba’s next-gen open model series emphasizing both scale and efficiency. The 80B-A3B-Thinking variant is specially designed for complex reasoning: it combines hybrid attention (linearized + sparse mechanisms) with a high-sparsity MoE. Its specs are striking: 80B total parameters, but only ~3B active (512 experts with 10 active). This yields very fast inference. Qwen3-Next also uses multi-token prediction (MTP) during training for speed.

Benchmarks show Qwen3-Next-80B performing excellently on multi-hop tasks. The model card highlights that it outperforms earlier Qwen-30B and Qwen-32B thinking models, and even outperforms the proprietary Gemini-2.5-Flash on several benchmarks. For example, it gets ~87.8% on AIME25 (math) and ~73.9% on HMMT25, better than Gemini-2.5-Flash’s 72.0% and 73.9% respectively. It also shows strong performance on MMLU and coding tests.

Key specs: 80B total, 3B active. 48 layers, hybrid layout with 262K native context. Fully Apache-2.0 licensed.

Strengths: Excellent reasoning & coding performance per compute (beats larger models on many tasks); huge context; extremely efficient (10× speed up for >32K context vs older Qwens).

Weaknesses: As a MoE model, it may require specific runtime support; “Thinking” mode adds complexity (always generates a <think> block and requires specific prompting).

8. Qwen3-235B-A22B

Qwen3-235B-A22B represents Alibaba’s most advanced open reasoning model to date. It uses a massive Mixture-of-Experts architecture with 235 billion total parameters but activates only 22 billion per token, achieving an optimal balance between capability and efficiency. The model employs the same hybrid attention mechanism as Qwen3-Next-80B (combining linearized and sparse attention) but scales it to handle even more complex reasoning chains.

The “A22B” designation refers to its 22B active parameters across a highly sparse expert system. This design allows the model to maintain reasoning quality comparable to much larger dense models while keeping inference costs manageable. Qwen3-235B-A22B supports dual-mode operation: it can run in standard mode for quick responses or switch to “thinking mode” with explicit chain-of-thought reasoning for complex tasks.

Performance-wise, Qwen3-235B-A22B excels across mathematical reasoning, coding, and multi-step logical tasks. On AIME 2025, it achieves approximately 89.2%, outperforming many proprietary models. It scores 76.8% on HMMT25 and maintains strong performance on MMLU-Pro (78.4%) and coding benchmarks like HumanEval (91.5%). The model’s long-context capability extends to 262K tokens natively, with optimized handling for extended reasoning chains.

The architecture incorporates multi-token prediction during training, which improves both training efficiency and the model’s ability to anticipate reasoning paths. This makes it particularly effective for tasks requiring forward planning, such as complex mathematical proofs or multi-file code refactoring.

Key Specs

  • Architecture: Hybrid MoE with dual-mode (standard/thinking) operation
  • Total parameters: ~235B
  • Active parameters: ~22B per token
  • Context length: 262K tokens native
  • License: Apache-2.0
  • Primary use cases: Advanced mathematical reasoning, complex coding tasks, multi-step problem solving, long-context analysis

Strengths

  • Exceptional mathematical and logical reasoning performance, surpassing many larger models
  • Dual-mode operation allows flexibility between speed and reasoning depth
  • Highly efficient inference relative to reasoning capability (22B active vs. 235B total)
  • Native long-context support without requiring extensions or special configurations
  • Comprehensive Apache-2.0 licensing enables commercial deployment

Weaknesses

  • Requires MoE-aware inference runtime (vLLM, DeepSpeed, or similar)
  • Thinking mode adds latency and token overhead for simple queries
  • Less mature ecosystem compared to LLaMA or GPT variants
  • Documentation primarily in Chinese, with English materials still developing

9. MiMo-V2-Flash

MiMo-V2-Flash represents an aggressive push toward ultra-efficient reasoning through a 309 billion parameter Mixture-of-Experts architecture that activates only 15 billion parameters per token. This 20:1 sparsity ratio is among the highest in production reasoning models, enabling inference speeds of approximately 150 tokens per second while maintaining competitive performance on mathematical and coding benchmarks.

The model uses a sparse gating mechanism that dynamically routes tokens to specialized expert networks. This architecture allows MiMo-V2-Flash to achieve remarkable cost efficiency, operating at just 2.5% of Claude’s inference cost while delivering comparable performance on specific reasoning tasks. The model was trained with a focus on mathematical reasoning, coding, and structured problem-solving.

MiMo-V2-Flash delivers impressive benchmark results, achieving 94.1% on AIME 2025, placing it among the top performers for mathematical reasoning. In coding tasks, it scores 73.4% on SWE-Bench Verified and demonstrates strong performance on standard programming benchmarks. The model supports a 128K token context window and is released under an open license permitting commercial use.

However, real-world performance reveals some limitations. Community testing indicates that while MiMo-V2-Flash excels on mathematical and coding benchmarks, it can struggle with instruction following and general-purpose tasks outside its core training distribution. The model performs best when tasks closely match mathematical competitions or coding challenges but shows inconsistent quality on open-ended reasoning tasks.

Key Specs

  • Architecture: Ultra-sparse MoE (309B total, 15B active)
  • Total parameters: ~309B
  • Active parameters: ~15B per token (20:1 sparsity)
  • Context length: 128K tokens
  • License: Open-weight, commercial use permitted
  • Inference speed: ~150 tokens/second
  • Primary use cases: Mathematical competitions, coding challenges, cost-sensitive deployments

Strengths

  • Exceptional efficiency with 15B active parameters delivering strong math and coding performance
  • Outstanding cost profile at 2.5% of Claude’s inference cost
  • Fast inference at 150 t/s enables real-time applications
  • Strong mathematical reasoning with 94.1% AIME 2025 score
  • Recent release represents cutting-edge MoE efficiency techniques

Weaknesses

  • Instruction-following can be inconsistent on general-purpose tasks
  • Performance is strongest within math and coding domains, less reliable on diverse workloads
  • Limited ecosystem maturity with sparse community tooling and documentation
  • Best suited for narrow, well-defined use cases rather than general reasoning agents

10. Ministral 14B Reasoning

Mistral AI’s Ministral 14B Reasoning represents a breakthrough in compact reasoning models. With only 14 billion parameters, it achieves reasoning performance that rivals models 5-10× its size, making it the most efficient model in this top-10 list. Ministral 14B is part of the broader Mistral 3 family and inherits architectural innovations from Mistral Large 3 while optimizing for deployment in resource-constrained environments.

The model employs a dense transformer architecture with specialized reasoning training. Unlike larger MoE models, Ministral achieves its efficiency through careful dataset curation and reinforcement learning focused specifically on mathematical and logical reasoning tasks. This targeted approach allows it to punch well above its weight class on reasoning benchmarks.

Remarkably, Ministral 14B achieves approximately 85% accuracy on AIME 2025, a leading result for any model under 30B parameters and competitive with models several times larger. It also scores 68.2% on GPQA Diamond and 82.7% on MATH-500, demonstrating broad reasoning capability across different problem types. On coding benchmarks, it achieves 78.5% on HumanEval, making it suitable for AI-assisted development workflows.

The model’s small size enables deployment scenarios impossible for larger models. It can run effectively on a single consumer GPU (RTX 4090, A6000) with 24GB VRAM, or even on high-end laptops with quantization. Inference speeds reach 40-60 tokens per second on consumer hardware, making it practical for real-time interactive applications. This accessibility opens reasoning-first AI to a much broader range of developers and use cases.

Key Specs

  • Architecture: Dense transformer with reasoning-optimized training
  • Total parameters: ~14B
  • Active parameters: ~14B (dense)
  • Context length: 128K tokens
  • License: Apache-2.0
  • Primary use cases: Edge reasoning, local development, resource-constrained environments, real-time interactive AI

Strengths

  • Exceptional reasoning performance relative to model size (~85% AIME 2025 at 14B)
  • Runs on consumer hardware (single RTX 4090 or similar) with strong performance
  • Fast inference speeds (40-60 t/s) enable real-time interactive applications
  • Lower operational costs make reasoning AI accessible to smaller teams and individual developers
  • Apache-2.0 license with minimal deployment barriers

Weaknesses

  • Lower absolute ceiling than 100B+ models on the most difficult reasoning tasks
  • Limited context window (128K) compared to million-token models
  • Dense architecture means no parameter efficiency gains from sparsity
  • May struggle with extremely long reasoning chains that require sustained computation
  • Smaller model capacity limits multilingual and multimodal capabilities

Model Comparison Summary

 

Model

Architecture

Params (Total / Active)

Context Length

License

Notable Strengths

GPT-OSS-120B 

Sparse / MoE-style

~117B / ~5.1B

~128K

Apache-2.0

Efficient GPT-level reasoning; single-GPU feasibility; agent-friendly

GLM-4.7 (Zhipu AI)

MoE Transformer

~355B / ~32B

~200K input / 128K output

MIT

Strong open coding + math reasoning; built-in tool & agent APIs

Kimi K2 Thinking (Moonshot AI)

MoE (≈384 experts)

~1T / ~32B

256K (up to 1M via Yarn)

Apache-2.0

Exceptional deep reasoning and long-horizon tool use; INT4 efficiency

MiniMax-M2.1

MoE (agent-optimized)

~230B / ~10B

Long (not publicly specified)

MIT

Engineered for agentic workflows; strong long-horizon reasoning

DeepSeek-R1 (distilled)

Dense Transformer (distilled)

8B / 8B

128K

MIT

Matches 235B models on reasoning; runs on single GPU; 87.5% AIME 2025

DeepSeek-V3.2 (Terminus)

MoE + Sparse Attention

~671B / ~37B

Up to ~1M (sparse)

MIT

State-of-the-art open agentic reasoning; long-context efficiency

Qwen3-Next-80B-Thinking

Hybrid MoE + hybrid attention

80B / ~3B

~262K native

Apache-2.0

Extremely compute-efficient reasoning; strong math & coding

Qwen3-235B-A22B

Hybrid MoE + dual-mode

~235B / ~22B

~262K native

Apache-2.0

Exceptional math reasoning (89.2% AIME); dual-mode flexibility

Ministral 14B Reasoning

Dense Transformer

~14B / ~14B

128K

Apache-2.0

Best-in-class efficiency; 85% AIME at 14B; runs on consumer GPUs

MiMo-V2-Flash

Ultra-sparse MoE

~309B / ~15B

128K

MIT

Ultra-efficient (2.5% Claude cost); 150 t/s; 94.1% AIME 2025

 

Conclusion

Open-source reasoning models have advanced quickly, but running them efficiently remains a real challenge. Agentic and reasoning workloads are fundamentally token-intensive. They involve long contexts, multi-step planning, repeated tool calls, and iterative execution. As a result, they burn through tokens rapidly and become expensive and slow when run on standard inference setups.

The Clarifai Reasoning Engine is built specifically to address this problem. It is optimized for agentic and reasoning workloads, using optimized kernels and adaptive techniques that improve throughput and latency over time without compromising accuracy. Combined with Compute Orchestration, Clarifai dynamically manages how these workloads run across GPUs, enabling high throughput, low latency, and predictable costs even as reasoning depth increases.

These optimizations are reflected in real benchmarks. In evaluations published by Artificial Analysis on GPT-OSS-120B, Clarifai achieved industry-leading results, exceeding 500 tokens per second with a time to first token of around 0.3 seconds. The results highlight how execution and orchestration choices directly impact the viability of large reasoning models in production.

In parallel, the platform continues to add and update support for top open-source reasoning models in the community. You can try these models directly in the Playground or access them through the API and integrate them into their own applications. The same infrastructure also supports deploying custom or self-hosted models, making it easy to evaluate, compare, and run reasoning workloads under consistent conditions.

As reasoning models continue to evolve in 2026, the ability to run them efficiently and affordably will be the real differentiator.



AI Copilot Keeps Berkeley’s X-Ray Particle Accelerator on Track


In the rolling hills of Berkeley, California, an AI agent is supporting high-stakes physics experiments at the Advanced Light Source (ALS) particle accelerator.

Researchers at the Lawrence Berkeley National Laboratory ALS facility recently deployed the Accelerator Assistant, a large language model (LLM)-driven system to keep X-ray research on track.

The Accelerator Assistant — powered by an NVIDIA H100 GPU harnessing CUDA for accelerated inference — taps into institutional knowledge data from the ALS support team and routes requests through Gemini, Claude or ChatGPT. It writes Python and solves problems, either autonomously or with a human in the loop.

This is no small task. The ALS particle accelerator sends electrons traveling near the speed of light in a 200-yard circular path, emitting ultraviolet and X-ray light, which is directed through 40 beamlines for 1,700 scientific experiments per year. Scientists worldwide use this process to study materials science, biology, chemistry, physics and environmental science.

At the ALS, beam interruptions can last minutes, hours or days, depending on the complexity, halting concurrent scientific experiments in process. And much can go wrong: the ALS control system has more than 230,000 process variables.

“It’s really important for such a machine to be up, and when we go down, there are 40 beamlines that do X-ray experiments, and they are waiting,” said Thorsten Hellert, staff scientist from the Accelerator Technology and Applied Physics Division at Berkeley Lab and lead author of a research paper on the groundbreaking work.

Until now, facility staff troubleshooting issues have had to quickly identify the areas, retrieve data and gather the right personnel for analysis under intense time pressure to get the system back up and running.

“The novel approach offers a blueprint for securely and transparently applying large language model-driven systems to particle accelerators, nuclear and fusion reactor facilities, and other complex scientific infrastructures,” said Hellert.

The research team demonstrated that the Accelerator Assistant can autonomously prepare and run a multistage physics experiment, cutting setup time and reducing efforts by 100x.

Applying Context Engineering Prompts to Accelerator Assistant

The ALS operators interact with the system through either a command line interface or Open WebUI, which enables interaction with various LLMs and is accessible from control room stations, as well as remotely. Under the hood, the system uses Osprey, a framework developed at Berkeley Lab to apply agent-based AI safely in complex control systems.

Each user is authenticated and the framework maintains personalized context and memory across sessions, and multiple sessions can be managed simultaneously. This allows users to organize distinct tasks or experiments into separate threads. These inputs are routed through the Accelerator Assistant, which makes connections to the database of more than 230,000 process variables, a historical database archive service and Jupyter Notebook-based execution environments.

“We try to engineer the context of every language model call with whatever prior knowledge we have from this execution up to this point,” said Hellert.

Inference is done either locally — using Ollama, which is an open-source tool for running LLMs with a personal computer, on an H100 GPU node located within the control room network — or externally with the CBorg gateway, which is a lab-managed interface that routes requests to external tools such as ChatGPT, Claude or Gemini.

The hybrid architecture balances secure, low-latency, on-premises inference with access to the latest foundation models. Integration with EPICS (Experimental Physics and Industrial Control System) enables operator-standard safety constraints for direct interaction with accelerator hardware. EPICS is a distributed control system used in large-scale scientific facilities such as particle accelerators. Engineers can write Python code in Jupyter Notebook that can communicate with it.

Basically, conversational input is turned into a clear natural language task description for objectives without redundancy. External knowledge such as personalized memory tied to users, documentation and accelerator databases are integrated to assist with terminology and context.

“It’s a large facility with a lot of specialized expertise,” said Hellert. “Much of that knowledge is scattered across teams, so even finding something simple — like the address of a temperature sensor in one part of the machine — can take time.”

Tapping Accelerator Assistant to Aid Engineers, Fusion Energy Development

Using the Accelerator Assistant, engineers can start with a simple prompt describing their goal. Behind the scenes, the system draws on carefully prepared examples and keywords from accelerator operations to guide the LLM’s reasoning.

“Each prompt is engineered with relevant context from our facility, so the model already knows what kind of task it’s dealing with,” said Hellert.

Each agent is an expert in that field, he said.

Once the task is defined, the agent brings together its specialized capabilities — such as finding process variables or navigating the control system — and can automatically generate and run Python scripts to analyze data, visualize results or interact safely with the accelerator itself.

“This is something that can save you serious time — in the paper, we say two orders of magnitude for such a prompt,” said Hellert.

Looking ahead, Hellert aims to have the ALS engineers put together a wiki that documents the many processes that go on to support the experiments. These documents could help the agents run the facilities autonomously — with a human in the loop to approve the course of action.

“On these high-stakes scientific experiments, even if it’s just a TEM microscope or something that might cost $1 million, a human in the loop can be very important,” said Hellert.

The work has already expanded beyond ALS as part of the DOE’s Genesys mission, with the framework being deployed across U.S. particle accelerator facilities. Next up, Hellert just began collaborating with engineers at the ITER fusion reactor — the world’s largest — in France for implementing the framework for use in the fusion reactor facility. He also has a collaboration in the works with the Extremely Large Telescope ELT, in northern Chile.

Benefiting Humanity: Scientific Impact of Experiments Supported by ALS

Beyond optimizing the accelerator and other industrial operations, the work at the ALS directly enables scientific breakthroughs with global impact. The facility’s stable X-ray beams underpin research in health, climate resilience and planetary science.

During the COVID-19 pandemic, ALS researchers helped characterize a rare antibody that could neutralize SARS-CoV-2. Structural biology experiments at Beamline 4.2.2 revealed how six molecular loops of the antibody latch onto and disable the viral spike protein. The findings supported the rapid development of a therapeutic that remained effective through multiple variants.

ALS science also contributes to climate-focused research. Metal-organic frameworks (MOFs) — a class of porous materials capable of capturing water or carbon dioxide from air — were extensively studied across several ALS beamlines. These experiments supported foundational work that ultimately led to the 2025 Nobel Prize in Chemistry, recognizing the transformative potential of MOFs for sustainable water harvesting and carbon management.

In planetary science, ALS measurements of samples returned from NASA’s OSIRIS-REx mission helped trace the chemical history of asteroid Bennu. X-ray analyses provided evidence that such asteroids carried water and molecular precursors of life to early Earth, deepening our understanding of the origins of the planet’s habitable conditions.

 

c# – How to make web app publishing play nice with build acceleration/FUTDC in VS build?


I have a setup where there is one web application and multiple libraries containing controllers, Razor views and static content under wwwroot. Something like this:

C:.
├───AspNetCoreTest
└───WebProcessorLibrary
    ├───Areas
    │   └───TestArea
    │       ├───Controllers
    │       └───Views
    │           └───AreaTest
    ├───Controllers
    ├───Views
    │   ├───Home
    │   └───Shared
    └───wwwroot
        ├───Areas
        │   └───TestArea
        │       ├───Content
        │       └───Scripts
        ├───Content
        └───Scripts

In this example there is only one library, but it could be more.

It is configured to publish on build into its own bin directory:

<PublishDir>$(OutputPath)</PublishDir>
<PublishUrl>$(PublishDir)</PublishUrl>
<DeployOnBuild>true</DeployOnBuild>

So after running dotnet build the web application’s bin directory looks like this:

C:\xyz [master ≡ +0 ~9 -0 !]> dir .\tests\apps\AspNetCoreTest\bin\Debug\net9.0\

    Directory: C:\xyz\tests\apps\AspNetCoreTest\bin\Debug\net9.0

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
d----            1/8/2026  8:49 PM                wwwroot
-a---            1/8/2026  8:49 PM           2030 AspNetCoreTest.deps.json
-a---            1/8/2026  8:49 PM           8704 AspNetCoreTest.dll
-a---            1/8/2026  8:49 PM         156672 AspNetCoreTest.exe
-a---            1/8/2026  8:49 PM          23444 AspNetCoreTest.pdb
-a---            1/8/2026  8:49 PM            416 AspNetCoreTest.runtimeconfig.json
-a---            1/8/2026  8:49 PM          42602 AspNetCoreTest.staticwebassets.endpoints.json
-a---            1/8/2026  8:49 PM           2310 AspNetCoreTest.staticwebassets.runtime.json
-a---            1/8/2026  8:49 PM           4608 Dayforce.Web.dll
-a---            1/8/2026  8:49 PM          22016 Dayforce.Web.NetCore.dll
-a---            1/8/2026  8:49 PM          28968 Dayforce.Web.NetCore.pdb
-a---            1/8/2026  8:49 PM          20936 Dayforce.Web.pdb
-a---            1/8/2026  8:49 PM          15872 TestModels.dll
-a---            1/8/2026  8:49 PM          35484 TestModels.pdb
-a---            1/8/2026  8:49 PM            558 web.config
-a---            1/8/2026  8:49 PM          84480 WebProcessorLibrary.dll
-a---            1/8/2026  8:49 PM          48424 WebProcessorLibrary.pdb
-a---            1/8/2026  8:49 PM          24416 WebProcessorLibrary.staticwebassets.endpoints.json
-a---            1/8/2026  8:49 PM           2154 WebProcessorLibrary.staticwebassets.runtime.json

C:\xyz [master ≡ +0 ~9 -0 !]>

And the cshtml views are compiled into the WebProcessorLibrary.dll assembly – verified.

All is in order.

Now I open the solution in VS IDE. Build acceleration and FUTDC are enabled – verified.

My scenario:

  1. Touch a static content, for example .\tests\apps\WebProcessorLibrary\wwwroot\Content\Site.css
  2. Build in VS IDE
  3. Check if the static content is published (it is not).

Observe:

C:\xyz [master ≡ +0 ~9 -0 !]> dir .\tests\apps\WebProcessorLibrary\wwwroot\Content\Site.css,.\tests\apps\AspNetCoreTest\bin\Debug\net9.0\wwwroot\Content\Site.css

    Directory: C:\xyz\tests\apps\WebProcessorLibrary\wwwroot\Content

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:54 PM           1792 Site.css

    Directory: C:\xyz\tests\apps\AspNetCoreTest\bin\Debug\net9.0\wwwroot\Content

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:54 PM           1792 Site.css

C:\xyz [master ≡ +0 ~9 -0 !]> touch.exe .\tests\apps\WebProcessorLibrary\wwwroot\Content\Site.css
C:\xyz [master ≡ +0 ~9 -0 !]> dir .\tests\apps\WebProcessorLibrary\wwwroot\Content\Site.css,.\tests\apps\AspNetCoreTest\bin\Debug\net9.0\wwwroot\Content\Site.css

    Directory: C:\xyz\tests\apps\WebProcessorLibrary\wwwroot\Content

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:56 PM           1792 Site.css

    Directory: C:\xyz\tests\apps\AspNetCoreTest\bin\Debug\net9.0\wwwroot\Content

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:54 PM           1792 Site.css

C:\xyz [master ≡ +0 ~9 -0 !]>

So the new timestamp is 8:56PM, but the published content is still at 8:54PM.

Now let us build the solution in VS IDE (the WebProcessorLibrary project targets both net472 and net9.0 and it builds into separate shared bin directory dedicated to library projects, i.e. web app builds into its own):

Build started at 8:57 PM...
1>------ Build started: Project: WebProcessorLibrary, Configuration: Debug Any CPU ------
1>WebProcessorLibrary -> C:\dayforce\DevOps\Dayforce.Web\tests\bin.core\Debug\net9.0\WebProcessorLibrary.dll
1>WebProcessorLibrary -> C:\dayforce\DevOps\Dayforce.Web\tests\bin\WebProcessorLibrary.dll
========== Build: 1 succeeded, 0 failed, 6 up-to-date, 0 skipped ==========
========== Build completed at 8:57 PM and took 02.340 seconds ==========

So FUTDC has handed off the build for WebProcessorLibrary to the msbuild, but the DLL was not updated, of course:

C:\xyz [master ≡ +0 ~9 -0 !]> dir .\tests\bin.core\Debug\net9.0\WebProcessorLibrary.dll, .\tests\apps\AspNetCoreTest\bin\Debug\net9.0\WebProcessorLibrary.dll

    Directory: C:\xyz\tests\bin.core\Debug\net9.0

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:49 PM          84480 WebProcessorLibrary.dll

    Directory: C:\xyz\tests\apps\AspNetCoreTest\bin\Debug\net9.0

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:49 PM          84480 WebProcessorLibrary.dll

C:\xyz [master ≡ +0 ~9 -0 !]>

The dll timestamp of 8:49PM predates the timestamp of the touched css file. And that makes sense, since the css file is not a dependency of the binary and we do not want it to be.

But it also means the published static content has not been updated either:

C:\xyz [master ≡ +0 ~9 -0 !]> dir .\tests\apps\WebProcessorLibrary\wwwroot\Content\Site.css,.\tests\apps\AspNetCoreTest\bin\Debug\net9.0\wwwroot\Content\Site.css

    Directory: C:\xyz\tests\apps\WebProcessorLibrary\wwwroot\Content

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:56 PM           1792 Site.css

    Directory: C:\xyz\tests\apps\AspNetCoreTest\bin\Debug\net9.0\wwwroot\Content

Mode                 LastWriteTime         Length Name
----                 -------------         ------ ----
-a---            1/8/2026  8:54 PM           1792 Site.css

C:\xyz [master ≡ +0 ~9 -0 !]>

Which is a problem.

What is the proper way to fix it without any of the following:

  • Disabling FUTDC for any project
  • Disabling build acceleration for any project (it is enabled – verified)
  • Causing C# code to recompile because of static content changes

AI Recruiting Tools for a Fairer Hiring Future


Hiring Future
ID 348055955 | Ai And Hiring ©
Wanan Yossingkum | Dreamstime.com

Key Takeaways

  • AI recruiting tools can streamline hiring processes and reduce human biases for the hiring future.
  • Ensuring fairness requires careful design, regular audits, and transparency in AI systems.
  • Legal frameworks are evolving to address the challenges posed by AI in the recruitment process.

Table of Contents

  • Introduction
  • The Promise of AI in Recruitment
  • Challenges in Ensuring Fairness
  • Legal and Ethical Considerations
  • Best Practices for Implementing AI Recruiting Tools
  • Conclusion

Artificial Intelligence (AI) has rapidly become a transformative force in the recruitment industry, enabling organizations to streamline their processes for identifying, attracting, and hiring top talent. As hiring volumes increase for the hiring future and competition intensifies, companies are turning to robust AI recruiting tools to automate labor-intensive tasks, enhance the candidate experience, and eliminate subjectivity from the decision-making process. But with this technological leap, tough questions are being raised regarding transparency, fairness, and the potential for bias in algorithm-driven hiring practices.

As recruiters leverage AI’s capabilities—from resume parsing to predictive assessments—the promise is to deliver more equitable outcomes, but only when these systems are built and managed with fairness as a core principle. The effectiveness of this promise hinges on recognizing common pitfalls in AI systems and proactively addressing them. Only then can we hope to create truly inclusive hiring processes where every candidate receives an equal opportunity to succeed.

For large organizations, AI is now essential for managing the complexity and sheer scale of applications, but ensuring these tools contribute to a level playing field is a continuous challenge. To help organizations harness the full benefits without unintended consequences, this article breaks down the opportunities and hurdles in achieving equity with AI-powered recruitment.

From emerging best practices to the rapidly evolving regulatory landscape, understanding how to utilize AI responsibly in recruitment is crucial for achieving long-term organizational and societal progress. In the following sections, you’ll learn about the advantages, risks, laws, and guidelines that shape this vital conversation for HR professionals and business leaders alike.

Introduction

The adoption of AI in recruitment has surged for hte hiring future. Corporations and agencies increasingly rely on algorithmic assistance to process vast candidate pools, accelerate hiring timelines, and deliver objective, data-driven recommendations. AI offers the promise of identifying ideal applicants more consistently and fairly by automating repetitive steps and applying precise criteria. Yet, realizing these fairness benefits is complex and fraught with the risk of replicating historic biases that are coded into the training data or created by opaque logic within the AI itself.

The Promise of AI in Recruitment

By automating the screening of resumes, scheduling interviews, and even assessing skills, AI tools offer unmatched efficiency. Some organizations report that carefully implemented AI systems have reduced screening costs by as much as 75% and cut the average time-to-hire from 44 days to just 11. More importantly, these tools can enforce consistent evaluation criteria, sidestepping individual recruiter biases and supporting efforts to build diverse workplaces.

  • Efficiency: For the hiring future, automating manual recruitment steps saves both time and operational costs.
  • Consistency: Every applicant is evaluated against uniform standards, ensuring a merit-based shortlist.
  • Data-Driven Insight: Advanced analytics provide recruitment teams with feedback and opportunities to refine their hiring strategies based on historical outcomes.

When properly calibrated, AI can act as a “bias interrupter,” flagging and removing job criteria that indirectly exclude underrepresented groups, according to several studies summarized by SHRM. Over time, these improvements can contribute to more positive candidate experiences and stronger business outcomes.

Challenges in Ensuring Fairness

Despite their advantages, AI recruiting tools can unintentionally reinforce or amplify discrimination unless carefully managed. One of the leading causes is bias in training data: if the algorithms are trained on historic hiring data that reflects past inequities, the resulting recommendations can mirror those injustices. Another concern is the lack of transparency in proprietary systems—often referred to as “black box” algorithms—which makes it difficult for employers and regulators to identify the root cause of errors or biases.

  • Biased Data: Historical discrimination or imbalanced datasets can perpetuate systemic exclusions.
  • Opaque Processes: Proprietary AI algorithms are often not fully explainable, making root-cause analysis of unfairness difficult.
  • Human Exclusion: Total reliance on AI overlooks the contextual judgment and ethical reasoning human recruiters provide, especially in the final decision stages.

Recent research published by Reuters demonstrates that even small, unintentional biases baked into algorithms can produce systemic disadvantages against marginalized candidates, often without clear warning signs.

Legal and Ethical Considerations

Regulators are catching up to the risks posed by AI-powered hiring. For example, New York City’s Local Law 144 mandates regular audits for bias and requires organizations to disclose their use of automated employment decision tools. Still, many of these regulations are in their infancy. Gaps remain regarding enforceable standards for acceptable levels of bias, coverage of partially automated tools, and mechanisms for candidates to challenge algorithmic decisions.

With the regulatory landscape evolving at a rapid pace—both in the U.S. and globally—organizations need to adopt a proactive approach. Legal compliance must be seen not only as a risk mitigation strategy but also as an ethical imperative for fostering inclusion in the workforce. The European Union’s AI Act and guidelines from the U.S. Equal Employment Opportunity Commission (EEOC) are expected to further define benchmarks for responsible use in the near future, as reported by The Wall Street Journal.

Best Practices for Implementing AI Recruiting Tools

Organizations seeking to maximize the benefits of AI recruiting tools while minimizing unintended harm should follow several best practices:

  1. Regular Bias Audits: Conduct frequent assessments of AI outputs for fairness, addressing any detected bias before it can impact candidates’ lives.
  2. Transparency: Offer candidates and recruiters insight into how AI-driven recommendations are made, ensuring accountability at every decision point.
  3. Human Oversight: Keep recruiters involved as the final decision-makers, using AI insights as support, rather than a sole arbiter.
  4. Diverse Training Data: Train algorithms on representative data sets to minimize the risk of encoding historical inequities.
  5. Compliance Monitoring: Monitor legal and regulatory developments to quickly adapt internal practices and policies accordingly.

Additionally, fostering a multidisciplinary team—including technologists, HR professionals, ethicists, and legal advisors—can further enhance the review and governance of AI in hiring, ensuring that a broad range of perspectives is considered.

Conclusion

AI recruiting tools present tremendous potential to foster fairness, speed, and objectivity in the hiring future processes. Achieving these positive outcomes requires deliberate attention to system design, frequent bias monitoring, human participation, and transparent practices. As the landscape of laws and best practices evolves, organizations must remain vigilant, continually refining their use of AI to advance workplace equity and inclusion for all candidates.

Find a Home-Based Business to Start-Up >>> Hundreds of Business Listings.