The Price of 4.8 Million Articles
In 2011, a 24-year-old programmer named Aaron Swartz was arrested for downloading academic papers from JSTOR via an MIT server. The government did not see a researcher or an activist; they saw a felon. They leveraged the Computer Fraud and Abuse Act (CFAA) to threaten him with 35 years in prison and $1 million in fines. Swartz wasn't selling the data. He wasn't feeding it into a proprietary model to sell subscriptions. He was an advocate for Open Access who believed that the sum of human knowledge shouldn't be locked behind paywalls.
Today, companies like Meta, Google, and OpenAI ingest the entire public internet—billions of images, trillions of words, and countless copyrighted works—to build commercial products valued in the hundreds of billions of dollars. They do so without permission, without compensation, and without the fear of a single FBI knock at the door. The scale of the 'scraping' is magnitudes larger than anything Swartz was accused of, yet the legal response has shifted from handcuffs to 'fair use' debates in civil courts. This is the shadow precedent: a world where data extraction is a crime for the citizen but a business model for the corporation.
The Weaponization of the CFAA
The Computer Fraud and Abuse Act was designed in 1986 to catch hackers, but it has historically been used as a cudgel against anyone who uses a computer in a way a platform doesn't like. In the Swartz case, the prosecution argued that by circumventing basic IP blocks to download files, he had 'exceeded authorized access.' This interpretation turned a terms-of-service violation into a federal crime. It was a strategy of total destruction aimed at a private individual who lacked the lobbying power to redefine the law.
Compare this to the current defense strategies of Big Tech. When Meta or OpenAI are caught scraping personal blogs, news sites, or private repositories, their legal teams argue that the 'transformative' nature of AI makes the scraping legal under fair use. They aren't worried about the CFAA because they have the capital to argue that their 'access' is inherently authorized by the public nature of the web. The law hasn't changed significantly since 2011; only the status of the entity doing the scraping has.

Photo by Brett Sayles on Pexels
The Enclosure of the Digital Commons
What we are seeing is a modern version of the Enclosure Acts, where common land was fenced off for private profit. In the 2010s, the legal system was used to protect the 'fences'—the paywalls of JSTOR and Elsevier—by making examples of those who tried to open them. Now, the same legal system is being asked to look the other way while corporations tear down those same fences for their own benefit. If an individual scrapes a database to provide free information, it is theft. If a trillion-dollar company scrapes a database to train a chatbot it charges $20 a month for, it is innovation.
This hypocrisy creates a dangerous incentive structure. It tells developers and activists that the only safe way to interact with data is to be part of a massive corporate hierarchy. If you have a legal department and a lobbying budget, the data belongs to you. If you are a lone actor with a script and a vision for the public good, you are a target. The 'unauthorized access' that destroyed Aaron Swartz has become the foundational fuel for the most profitable industry of the 21st century.
What This Actually Means
The divergence between the Swartz prosecution and the AI era proves that 'justice' in the digital realm is often just a reflection of market cap. We have built a system that criminalizes the liberation of information while subsidizing its exploitation. By allowing corporations to claim fair use for massive-scale scraping while maintaining strict criminal penalties for individuals, we are effectively giving Big Tech a monopoly on the use of the public internet.
If we do not address this disparity, we will end up with a digital landscape where the only people allowed to 'read' the web at scale are machines owned by a handful of companies. The spirit of the open web—the one Aaron Swartz died trying to protect—is being suffocated not by a lack of technology, but by a legal system that treats data as a liability for the public and an asset for the powerful. We must demand a legal framework that applies the same standards of 'access' and 'fair use' to a kid in a basement as it does to a CEO in a boardroom.
Quick Answers
Wasn't what Aaron Swartz did actually illegal?
Technically, he violated a terms-of-service agreement and bypassed mechanical blocks, which the government used to trigger the CFAA. The controversy lies in the fact that the punishment—decades in prison—was obscenely disproportionate to a non-commercial act of downloading papers he already had legal access to through MIT.
How is AI scraping different from what Swartz did?
In terms of technical action, it isn't; both involve automated scripts fetching data from servers. The difference is purpose and power: AI labs use the data to create new commercial products, whereas Swartz intended to make the data freely available to the public.
Why aren't AI companies being prosecuted under the CFAA?
Recent Supreme Court rulings, like Van Buren v. United States, have narrowed the CFAA, making it harder to prosecute people for simply violating a website's rules. However, the biggest factor is political and economic; the government is currently more interested in winning the 'AI arms race' than in policing the methods of the companies winning it.



