A Masterclass in Self-Sabotage

Silicon Valley has finally achieved the impossible: they have built a perpetual motion machine that generates nothing but noise. For the last decade, we were told the internet was a vast, shimmering ocean of human knowledge ripe for the taking. Now, thanks to the miracle of large language models, that ocean is roughly 40% plastic by volume and rising. It is truly impressive to watch billion-dollar companies spend a fortune to build tools that effectively poison their own water supply.

We are currently witnessing the "Synthetic Erosion" of the digital commons, which is a fancy way of saying we've let the robots rewrite the script and now the robots are bored with what they're reading. Since November 2022, the volume of AI-generated sludge—pardon me, "optimized content"—has reached a point where the next generation of models is being trained on the hallucinations of the previous one. It’s like trying to make a photocopy of a photocopy of a polaroid of a ghost. The resolution isn't exactly improving.

The Great Data Enclosure Act

If you enjoyed the era of the open internet, I hope you took screenshots. We are rapidly moving toward a world where anything worth reading is locked behind a $20-a-month velvet rope. Reddit, Twitter, and every major news outlet have realized that their archives are the only thing standing between AI labs and total structural collapse. Consequently, the "commons" is being fenced off faster than a suburban housing development in 2005.

  • Reddit signed a $60 million deal with Google to let them graze on human snark.
  • News Corp is charging millions for the privilege of letting an LLM summarize their paywalled articles into a three-sentence bulleted list.
  • The rest of us get to navigate an "open" web that consists primarily of 14,000-word blog posts about "The Top 10 Best Toasters of 2024" written by a bot that has never seen bread.

This is the "data-scarcity paradox." The more content we generate, the less information we actually have. We are drowning in text, yet starving for a single sentence that wasn't statistically predicted by a GPU in a shipping container in Iowa. The irony of paying a subscription fee to access human-written text so you can use it to train a bot to replace the human you just paid to read is a level of meta-capitalism that deserves its own circle of hell.

a heavy iron padlock on a rusty gate leading to an empty library
Photo by David McElwee on Pexels

Modeling the Heat Death of the Internet

Researchers at Oxford and Cambridge recently coined the term "Model Collapse." It’s a beautiful concept, really. It describes what happens when an AI is trained on too much AI-generated data: it forgets the edges of reality. It starts to believe that the average human has eight fingers and that the capital of France is probably a brand of artisanal cheese. By flooding the zone with synthetic garbage, the industry has ensured that the only way to build a "smart" model is to find data that hasn't been touched by an AI yet.

This has created a hilarious new gold rush for "pristine" data. Companies are now scouring the basements of university libraries and digitizing yellowing paper from 1954 because, back then, people actually had to think before they typed. We are effectively strip-mining the past to fuel a future that can’t even describe a sunset without mentioning "a symphony of colors" or "a testament to the beauty of nature."

What This Actually Means

The "Digital Commons" is dead; it’s just that the funeral is being live-streamed by a bot to an audience of other bots. We are transitioning into a two-tier information society. The wealthy and the powerful will have access to "Verified Human" data—private, curated enclaves of actual thought. The rest of us will be left to wander the synthetic wasteland, arguing with customer service chatbots that use 500 words to tell us they can't help.

This isn't a glitch in the system; it’s the inevitable result of treating human expression as a raw commodity to be harvested rather than a conversation to be joined. We’ve commodified thought to the point of worthlessness. If you want to know what the future looks like, imagine a bot reading a bot’s summary of a bot’s interpretation of a tweet—forever.

It’s a tragedy, sure, but at least the shareholders are happy for the next fiscal quarter. And isn't that the real meaning of progress? To build a machine so sophisticated that it eventually chokes on its own output while we watch from behind a paywall?

Quick Answers

Is the internet actually getting worse?
Yes, objectively. Search results are now a graveyard of AI-generated SEO bait designed to sell you affiliate links for products that don't exist.

Can't we just filter out the AI content?
We’re trying, but it’s an arms race where the person building the filters is also the person building the bots that bypass them. It’s a very lucrative circle.

Where can I find real human writing?
Mostly in physical books, old journals, or deep within private Discord servers where the scraping bots haven't found the invite link yet.

Will AI eventually run out of things to say?
Technically no, it will just keep saying the same three things in increasingly complex and confident ways until we all stop noticing.