Yesterday, OpenAI officially released ‘gpt-oss-120b’ and ‘gpt-oss-20b’, which are 2 new open-weight language models that can be used by devs, researchers, and companies at a lower cost, assuming that deploying and operating the model locally are the most prioritized and uncompromisable factors.

OpenAI Logo

Both of these models are trained on NVIDIA H100 GPUs and are designed to run best on the massive network of GPUs powered by the NVIDIA CUDA platform. With optimizations for the NVIDIA Blackwell platform thanks to the introduction of NVFP4 4-bit precision, they can hit an impressive 1.5 million tokens per second on GB200 NVL72 systems, delivering huge efficiency gains for inference.

As for capabilities, OpenAI described them as having the ability to do advanced reasoning, utilizing tools, apply chain-of-thought processing, and are fundamentally designed to run anywhere from consumer hardware and embedded systems to cloud environments and data centers.

Capitalizing on this, Cloud providers like Amazon, Baseten, Microsoft, and of course, NVIDIA via NIM Microservices, are offering those models on the respective platforms as an “online alternative”, while the rest of the people can find it through places like Hugging Face and GitHub.

CEO Sam Altman has stated: “We’re excited to make this model, the result of billions of dollars of research, available to the world to get AI into the hands of the most people possible,”.

Open-weight language models are different from fully open-sourced models, though, as the former is about allowing users to “inspect and build upon the process of refining inputs and output” for a better outcome, while the latter will have all of its source code and documentation publicized for literally anyone to fork, modify, and deploy in their own discretion.

The release of ‘gpt-oss-120b’ and ‘gpt-oss-20b’ are highly anticipated as it has been delayed a couple of times before, and adding on the fact that GPT-4 “is the dumbest model that users will ever use” that subsequently led to the creation of multiple GPT-4.x.x (or whatever name it is) builds made for specific tasks with specialized properties like lighting fast processing or deep, slow, methodical thinking, this might be the best time for accelerating AI adoption.

Facebook
Twitter
LinkedIn
Pinterest

Related Posts

Subscribe via Email

Enter your email address to subscribe to Tech-Critter and receive notifications of new posts by email.