The Salad Bar Problem of Computing

We were promised that smaller, open-source models would save the polar bears. The logic was simple: if we stop sending every single 'write a haiku about a depressed toaster' request to a warehouse-sized data center in Iowa, we save electricity. It’s the digital equivalent of bringing your own tote bag to the grocery store. But then tools like Ollama arrived and made running these models so easy that I’m now running fourteen different versions of Llama 3 just to see which one is better at explaining why my cat looks at me with such profound judgment.

Enter the Jevons Paradox. Named after William Stanley Jevons—a 19th-century economist who was probably a real blast at parties—this principle states that when you make a resource more efficient, people don't use less of it. They use way, way more. It’s like when a restaurant opens an 'all-you-can-eat' salad bar. You don't go there to eat a sensible amount of kale; you go there to see how many chickpeas the human structural frame can support before it collapses.

By making AI 'cheap' and 'local,' we haven't reduced the carbon footprint. We’ve just distributed the environmental guilt across millions of glowing MacBooks. I used to feel bad about a server farm cooling itself with a river; now I’m just worried my desk is going to melt through the floor like the blood from a Xenomorph.

My Laptop Is Screaming and It’s My Fault

There is a specific sound a MacBook Pro makes when it realizes you are about to ask a 7-billion parameter model to summarize a 400-page PDF of 'The Silmarillion.' It’s a high-pitched whir that sounds like a miniature jet engine trying to escape its earthly tether. Before Ollama, I’d think twice. I’d use my own brain, which runs on roughly 20 watts and the occasional sourdough crust. Now? I’m clicking 'Run' because it’s free. Except it’s not free. It’s costing my local power grid its dignity.

a laptop sitting on an ice pack with steam rising
Photo by Dmitriy Ryndin on Pexels

The 'rebound effect' is a comedic tragedy. We optimized the code so it uses 50% less energy per token, so naturally, we all decided to generate 1,000% more tokens. It’s like buying a fuel-efficient Prius and then deciding that, because gas is cheap, you should spend your weekends driving in circles around the state of Nebraska just to feel the wind in your hair. We are currently driving the AI Prius through Nebraska at 120 miles per hour.

I recently saw a post on Hacker News where someone was proud of their 'Jev-style decision model' for local inference. They had automated their entire email inbox to be sorted by a local LLM. Every time a newsletter lands, their GPU kicks into gear, consuming enough power to briefly dim the streetlights in their neighborhood, all so an AI can tell them that a 20% off coupon for socks isn't 'high priority.' We are burning the rainforest to avoid reading about discount hosiery.

The Thermodynamics of Boredom

Let’s talk about the 'Ollaya' effect—that specific dopamine hit you get when you realize you can download a new model and have it running in thirty seconds. It’s addictive. It’s the digital version of hoarding craft supplies for a hobby you don't actually have. I have 400GB of quantized models on my drive. Do I use them? No. But I might need a specific uncensored version of Mistral if I ever decide to write a gritty reboot of The Magic School Bus.

Every time we download a new 4-bit quantization, we tell ourselves we’re being 'efficient.' We’re 'saving VRAM.' But VRAM is a gas that expands to fill its container. If you give a nerd 24GB of VRAM, they will find a way to use 23.9GB of it to generate high-resolution images of a steampunk squirrel. It is a fundamental law of the universe, right up there with gravity and the fact that you will always find an old AA battery in your kitchen drawer that is actually dead.

We’ve reached a point where 'local' doesn't mean 'sustainable.' It just means 'unregulated.' When OpenAI runs a model, they at least have an incentive to keep power costs down because they enjoy having billions of dollars. When I run a model, my only incentive is my own curiosity, which is a bottomless pit of energy consumption fueled by a desire to see if an AI can write a convincing apology letter from a dog to a rug.

What This Actually Means

Making things easier to use almost always results in us using them into the ground. The Jevons Paradox isn't a glitch in the system; it’s a feature of human psychology. We don't want 'enough.' We want 'more for less,' and once we get 'more for less,' we just take 'way too much.' The accessibility of open-source AI is a triumph of engineering, but it’s an absolute disaster for anyone who likes glaciers.

If we actually want to address the carbon footprint of AI, we have to stop pretending that 'efficiency' is a magic wand. Efficiency is just a bigger bucket. If you want to save water, you don't just get a more efficient bucket; you eventually have to turn off the tap. But nobody wants to turn off the tap when the tap is currently generating a 3D-render of Shrek in the style of Wes Anderson.

Ultimately, my laptop is going to keep screaming, and I’m going to keep downloading 'efficient' models until the heat from my CPU finally cooks the burrito I left next to the exhaust port. We’re not saving the world; we’re just making the end of it very, very interactive.

Quick Answers

Does running AI locally actually save energy?
Technically yes, per individual query, but you’ll end up running so many queries that your smart meter will start smoking. It’s like saying eating one grape is healthy, then eating 14,000 grapes.

What is the Jevons Paradox in simple terms?
The more efficient you make a lightbulb, the more likely people are to leave the lights on in the garage for three weeks straight just because 'it's cheap.'

Should I stop using Ollama?
No, it’s great software, just maybe stop asking it to rewrite the US Constitution in the voice of a surfer named 'Chad' every time you’re bored at 2 AM.