GPUs, Part 1: From Quake III Arena to the AI Factory
Do you need an AI Factory, or do you need some GPU-capable servers for a few workloads? How do you accurately measure the ROI? We discuss how answers to these and other questions will have a direct impact on tokenomics.
https://delivery-p155402-e1860468.adobeaemcloud.com/adobe/assets/urn:aaid:aem:7fdd2679-b849-45f1-80f6-04743c65c8a2/as/ModernizationMonday-GPU-AdobeStock_1862904371.avif
GPU
2026-10-12T00:00:00.000Z
6
Adam Stone
Field CTO
blog author portrait Adam Stone

Modernization Monday is a field guide to what it actually takes to modernize your infrastructure for AI, published every other Monday. Each installment takes one piece of the puzzle and breaks it down in plain terms. This is post two of nine… I mean ten. I changed my mind. Twelve. I may change it again. Treat “twelve” like a soft target.

Last time, we established the base metric of our AI unit economics, tokenomics. In this installment and the three that follow, we look at the biggest single lever on that number, which is where and how you buy the factory's machinery. This is where decisions get real, so let’s walk through them the way a principal architect would… by asking a whole slew of questions. Do you need an AI Factory, or do you need some GPU-capable servers for a few workloads? Are you sold on a cool idea without a practical need? How do you accurately measure the ROI? Is purchasing a requirement, or is cloud or neocloud a better fit? Which GPU server? What size? How many? From whom? Where is it going? Can you power it on? Do I bet the over if the wind is blowing out to left-center at Wrigley this weekend? Ginger or Mary Ann?

baseball field

Each answer lays the foundation for the next decision. There are enough decisions that compute gets four installments. The first two are about the core of the machine itself, the GPU. This one covers a brief overview of the GPU, what it does, and the vocabulary you need to follow a GPU conversation. The next one covers how to size the right one. After that, we zoom out to what the GPU goes in, and whether you should own it, rent it, or both.

Let's start at the beginning

The core physical technology that powers the most influential and disruptive software of our time was invented for the sole purpose of making video games look better. In 1999, NVIDIA released the GeForce 256 and marketed it as the world's first Graphics Processing Unit (GPU). It was brought into existence so I didn't have to replace the CPU in my ancient PC to make Quake III Arena gameplay smooth as glass while I sat at my desk back in college. As it turns out, the math behind making a dragon look more lifelike on a screen is remarkably similar to the math behind teaching a computer advanced pattern recognition. I am oversimplifying a bit here, but both come down to multiplying enormous grids of numbers.

After a while, the visionaries at NVIDIA recognized the potential of leveraging this technology beyond making gamers happier. In 2006, NVIDIA introduced the Compute Unified Device Architecture (CUDA), software that gave programmers an interface to leverage GPUs for workloads outside of those that lower grades and frustrate spouses. Most investors treated CUDA like an expensive side project.

Six years later (2012, for those who were about to pull out a calculator), a University of Toronto team trained an image-recognition model called AlexNet on two off-the-shelf gaming GPUs and beat the competition so decisively that the research world changed direction as fast as a hiker who spots someone on the trail in an 80s hockey mask.

Then a lot happened in the following ten years that I could go into, but our marketing team would never approve the length of this post. Let's just throw in one quick fact: in 2016, Jensen Huang delivered the first NVIDIA DGX-1 supercomputer to a small research lab called OpenAI.

When ChatGPT arrived in late 2022, an unexpected race began. Who would have thought parallel math processing would ever be so popular? Revenge of the nerds. NVIDIA had spent more than fifteen years building both the hardware and the software interface to rapidly expand GPU usage. Stock goes boom. No, I didn't have any. It took CUDA roughly sixteen years to pay off, and the lesson applies to more than stock prices. Infrastructure bets that matter most are measured in years, not quarters.

data science and AI applications

What a GPU actually does

Here is the explanation I use with executives who have never heard the term “CUDA core” and never want to. A CPU is like a handful of brilliant professors who can solve almost any problem, one problem at a time, in order. A GPU is like a stadium full of students who can each do simple arithmetic, all at the same time. Give the professors a complicated, branching task and they shine. Give the stadium a million identical multiplication problems and it finishes before the professors have even wiped off the chalkboard.

Rendering a video game frame is a million identical problems. What color is this pixel? What color is the one next to it? So on and so on. Running an AI model is the same type of problem with different inputs and outputs. The CPUs are the foremen. The GPUs are the machines on the assembly line. The foremen schedule the work and move the data. The GPUs work the math.

So, why is it still called a GPU if the “G” is for Graphics? For the same reason you still “dial” a phone number and “hang up” when you are done. Some old names stick. Move on.

A quick detour: dials, weights, and parameters

Speaking of dials, think of an AI model as a software machine with billions of dials that has learned patterns from an enormous amount of data. A model can be ready-to-go for users to start interacting with. We call this “inferencing.” For inferencing, think of all the dials being set to where they need to be. In AI language, the “weights” are set.

Before a model is ready for inferencing, it needs to be trained. Training is the process of ingesting an enormous amount of data and learning from it to turn every single dial (weight) to exactly the right position for the use case. When you come across the word “parameters” in a model’s name or description, you are seeing the count of the dials. A 32-billion-parameter model has 32 billion dials to turn. Training a model is an intensive (and very expensive) process.

If you take a model that someone else has already trained and continue its training on your own data, so it learns your vocabulary, your documents, or your task, that process is called “fine-tuning.” Fine-tuning is adjusting the weights and hence is also intensive and expensive.

I’d be remiss if I didn’t throw one more term out there, because it decides whether you can run a model on your own hardware at all or have to rent it. Some model makers publish their finalized weights for anyone to download. Think of it as a cousin of open-source software: you get the finished product, not necessarily the recipe. The term for these is “open-weights model.” Models like Llama (from Meta), Qwen (from Alibaba Cloud), Mistral (from Mistral AI, a French company), and Gemma (from Google DeepMind) are very popular and downloadable. Depending on the model’s license, you can run it on your own hardware, fine-tune it, and keep inference traffic entirely within your environment.

Other model makers keep the dials locked away, and you pay to use them, either by monthly subscription or by the token. The models behind ChatGPT, Claude, and Gemini work this way. You can rent time on their AI Factories, but you can’t build your own with their models. Some people call these “closed models.” I don’t. I call them by their names.

When you are “sizing” a GPU for a model (which is the next installment’s topic), you are sizing for an open-weights model.

The generations, attempted briefly

GPU generations now arrive roughly every year, and the naming conventions make sense if you know them and make absolutely no sense at first glance. Isn’t B before H? NVIDIA names all their GPUs after scientists. “Hopper” is named after Rear Admiral Grace Hopper, a U.S. Navy computer pioneer. She was instrumental in programming the Harvard Mark I, an IBM Automatic Sequence Controlled Calculator (ASCC), in the 1940s, was the author of one of the first compilers, and was one of the leaders that brought the world COBOL. She passed away in 1992 without ever publicly apologizing to a single college student for that last one.

Fun fact time: Grace Hopper’s team taped a moth found in a relay of the Harvard Mark II into the computer’s logbook in 1947, labeled “first actual case of bug being found.” The word “bug” was already around, but that moth is why every programmer on earth now says “debugging.”

handwritten notes

In the NVIDIA ecosystem, Hopper was the generation that carried the first wave of generative AI. First came the H100, followed by the memory-upgraded H200. Blackwell came next, starting with the B200, which… come on. Is there a B100 somewhere that everyone just forgot about? (There actually was.) Grace Blackwell GB200 and GB300 systems pair the GPU with NVIDIA's own Grace CPU. Vera Rubin has arrived and roughly doubles performance yet again, with facility consequences that are, frankly, impressive. A single NVL72 (72 GPUs) Vera Rubin rack requires 190 to 230 kW. A. Single. Rack. That one rack draws the average power of nearly 200 homes. Color me speechless.

A Vera Rubin rack houses 1.3 million individual components and nearly 1,300 chips, all conveniently packaged into a rack weighing roughly 4,000 lbs. That is fighting-weight territory for a hippopotamus. Hippos eat roughly 1 to 1.5 percent of their body weight per night, about 88 lbs. of grass. A Vera Rubin rack running flat out burns through the energy equivalent of about 160 gallons of gasoline a day, or six fill-ups of the pickup truck it weighs as much as. Rewind back to the point of tokenomics, though. Who cares what it weighs or eats, as long as the floor can hold it, the facility can keep it healthy over time, and the tokens coming out are more valuable than the electrons going in.

Fun fact number two: A hippo can surprisingly run at about 19 mph over short distances. Back in high school, I could potentially have outrun an average hippo, although I never came across one in Rhode Island. I could run more than 20 mph. I was very fast. A Vera Rubin rack runs for several million dollars. It is also very fast.

AI race hippo

Each generation does more work per watt than the last and raises the power and cooling bars along with it, both of which matter for tokenomics. Which generation to buy, and whether to wait for the next one, comes after learning how to choose GPUs and GPU counts based on workloads.

Now you can follow the conversation

You now know where the GPU came from, what it does, what a model is, and how to read the names on the roadmap. That is enough to follow most GPU conversations in the conference room without just nodding along and hoping the person speaking at you isn’t going to ask your opinion.

Next time, we open the spec sheet and define the four numbers that decide which GPU fits your model, and we’ll illustrate two examples that show how fast the arithmetic gets real.

If your team is working through this right now, it is exactly the kind of conversation we have every day. Get in touch and we can talk through where you are.

Next time: GPUs, part two. Sizing the machinery: memory, precision, interconnectivity, and power.

false
Blog
Artificial Intelligence,Networking
3
technology-area
true
related-cards