Books Have Always Been Destroyed. But Never Like This

Mike's Notes

I love print books and libraries. Books in print are precious. AI needs to benefit humanity, not be a destructive force.

Trinity College, Ireland

The original posted article had many links.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Card Catalog
  • Home > Handbook > 

Last Updated

01/09/2026

Books Have Always Been Destroyed. But Never Like This

By: Hana Lee Goldin
Card Catalog: 25/08/2026

Your personal librarian for the AI age. Forever in the pursuit of exploring how we find, filter, and feel about information.

We’ve entered the third era of libricide.

...

Quick summary:

An AI company has been buying up used books by the million and destroying them, scanning the pages for training data and pulping what’s left. Court filings unsealed this year describe the program and an internal note asking that the work be kept from becoming known. A court has ruled that buying and scanning the books this way is legal, and once the work is done, nothing survives to show a book was ever there to lose.

Key takeaways:

  • The books are bought through anonymous middlemen, so the sellers filling the orders rarely know where their stock is headed. From every angle the buying looks like ordinary commerce, which is exactly what keeps the destruction from being seen.
  • Book destruction has a long history, and it changes each time the technology of copying changes. It has shifted twice before. This is the third shift, and it looks nothing like the book burnings most of us picture.
  • This shift is set apart by leaving nothing behind. The words are kept, the object is discarded, and no record says which titles were taken, so the loss can be neither proven nor traced to anyone who might answer for the harm.
  • That reaches past books to anyone who wants to know things firsthand. When the only surviving copy is sealed inside a system no outsider can consult, verifying what it says becomes impossible, and we are left trusting whatever summary we are given.

...

“We don’t want it to be known that we are working on this.”

The sentence appears in an internal Anthropic planning document made public through court filings in Bartz v. Anthropic, the copyright class action that three authors filed in 2024. Anthropic, the maker of Claude, called the program Project Panama. Its stated mission fit in one sentence: “Project Panama is our effort to destructively scan all the books in the world.”

Destructive scanning means cutting a book from its binding, feeding the loose pages through a scanner, and discarding the book afterward. According to the court, Anthropic spent millions of dollars buying millions of print books, often used. Its service providers removed the bindings, cut the pages, scanned them into searchable digital files, and discarded the books. Anthropic kept the scans in an internal digital library and used books from that library to train the AI systems behind Claude.

The lawsuit initially concerned another source of Anthropic’s library: more than seven million pirated book files the company had downloaded. The court treated the two acquisition paths differently. It held that Anthropic could lawfully buy print books, convert them into searchable digital files for internal use, and discard the physical copies; because the resulting files remained inside the company rather than being redistributed, that practice was fair use. It reached the opposite conclusion about the books Anthropic had downloaded from pirate sites and retained.

Companies announce the programs they are proud of; Anthropic tried to keep this one invisible. The January unsealing supplied the project’s codename and the instruction to stay silent. The instruction anticipated what the company did not want authors, booksellers, and the reading public to see: books bought by the million, cut apart, scanned, and discarded. Anthropic didn’t publicly announce Project Panama before court records made it public.

The silence has no precedent. Books have been destroyed for as long as they’ve been made: by conquering armies and offended churches, by censors with lists and mobs with torches, by floods and fires, and by budgets that let the roof leak. Some of it was fast and public. More of it was slow and official. All of it, loud or slow, could be recognized for what it was. A person watching knew that books were being lost. History kept what record it could. But what’s happening now carries no such signature. From the outside, nobody could tell that books were being destroyed at all.

Inside Project Panama

Anthropic hired a man named Tom Turvey in February 2024 and gave him a mission the court record repeats in one sweeping phrase: obtaining all the books in the world. Turvey came to the job from Google Books, the project that spent the 2000s digitizing library collections on machines engineered to turn pages gently, so that every book survived its own scanning. Anthropic took the opposite approach. Gentle machinery never entered the plan. Over roughly a year, the company spent tens of millions of dollars buying millions of print books, sheared off their spines, fed the loose pages through high-speed scanners, and pulped what remained. Vendor proposals in the court records described converting up to two million books in six months, some eleven thousand a day.

Destroying the books solved a financial problem first. A bound book must be scanned page by page, slowly and at cost. Once cut apart, the same volume becomes a stack of loose paper that can run through a sheet feeder at speed. Across millions of volumes, the difference in output determined the method.

The legal significance emerged later. In June 2025, Judge William Alsup ruled on the authors’ claims. For the books Anthropic had bought and destroyed, he found the copying to be fair use: the rule in American copyright law that allows limited copying without an author’s permission (the way a critic can quote a novel in a review). His reasoning turned on replacement. Anthropic bought one physical copy, made one digital copy, and destroyed the original—so the number of copies in the world never grew. In the law’s eyes, the scan simply took the book’s place, a change of format. If Anthropic had kept both the book and the scan, the number of copies would have doubled and the fair-use argument would have weakened. Pulping the originals helped make the copying legal.

The pirated downloads fared differently in the same ruling. For those more than seven million files, Alsup rejected Anthropic’s fair-use defense. He emphasized that Anthropic had built a permanent, general-purpose library it expected to keep indefinitely, not a temporary collection for a defined training task. “None is even offered here except for Anthropic’s pocketbook and convenience,” he wrote.

The ruling left the company facing a trial over damages, and Anthropic settled instead: the company agreed to pay $1.5 billion. A judge granted the deal final approval on July 20, the largest copyright class settlement in American history. Under the deal, Anthropic will delete the pirated digital files. Critically, the deal concerns those files, not the millions of physical books the company bought and pulped. Copyright protects the text of a book—the creative expression of an idea—not the paper on which those words were printed. Once Anthropic lawfully owned a physical copy, copyright law generally didn’t prevent it from destroying that object. The books could be destroyed without creating the kind of copyright violation at issue in the case.

But books don’t enter a library by themselves: somebody had to sell the company all those books. The court records describe Anthropic purchasing through vendors. This spring and summer, booksellers across Europe described the other side of such a trade: unusually large, eclectic orders placed through intermediaries, with the ultimate buyer unnamed. Tomás Kenny of Kennys of Galway called one order for books “bananas”—a mix of titles no library would plausibly assemble. Whether any particular order came from Anthropic cannot be established from outside the transaction. Nor is Anthropic the only possible customer: other AI companies have also been reported to be acquiring books at industrial scale.

The reported market is built to keep the buyer at a distance. Brokers can aggregate inventory and manage bulk orders without disclosing who is ultimately acquiring the books. ISBNdb, a book-data company, briefly advertised a prospective book-sourcing service for AI developers that promised confidentiality; its marketing explained the appeal bluntly: “‘AI company destroys two million books’ is not a headline that generates sympathy.” (After reporting drew attention to the page, ISBNdb removed it and said the proposed service had never been launched.) The result was not merely secrecy about the buyer, but uncertainty about the fate of the books.

Annihilation, then spectacle

What Anthropic is doing belongs to a history far older than the company. Rebecca Knuth, a professor of library science, named the practice in 2003: libricide, the systematic destruction of books and libraries, usually carried out or authorized by a government. Her subject was the twentieth century’s state-sponsored campaigns, but the practice runs back as far as writing does. The scenes that come to mind of this are the same few: students feeding bonfires in Berlin in 1933, Sarajevo’s national library burning under siege in 1992. But behind those scenes, the record is wider and stranger than they suggest. Bonfires were rarer than the memory of them. Most destruction arrived slowly and with permission, through purges, censors, wars, and simple neglect.

A strict reading of that definition would leave Anthropic out, reserving libricide for the destruction of entire libraries. But that distinction collapses here. The world’s secondhand book trade functions as one enormous collection, scattered across thousands of shops and sellers with no central address. Buying it up by the pallet and pulping it empties a library all the same, just one whose shelves span continents.

Destroying a book has meant different things in different centuries. The difference has always come down to copying, because how books are copied decides how many of any one book exist. When copies are scarce, destruction can erase a work from existence. Once copies are everywhere, destruction can only send a message about one. To me, that line sorts the history of libricide into eras, and what distinguishes each one is what its destruction leaves behind. The result is my own framework, a chronology by residue that I haven’t found anywhere in the scholarship: two eras completed, and a third that has just begun.

For thousands of years, every copy of every book was made by hand. A single volume could take a scribe months, so most works existed in a few manuscripts, and some in only one. Under those conditions, destroying the object destroyed the work, completely and forever. I call this the first era: the era of destruction as annihilation. The word descends from the Latin ad nihil, meaning to reduce or bring something to nothing. In this era the meaning was literal. When Diego de Landa, a Spanish friar in colonial Yucatán, burned twenty-seven Maya codices in 1562, the texts inside them went to ash. No copies of them existed anywhere on earth, so entire bodies of Maya history and belief ended in one afternoon’s fire. The library of Alexandria met the slower version of the same fate, declining through centuries of purges and neglect until its losses ran past counting. What the first era left behind was dust, and one thing more: knowledge of the loss. Contemporaries recorded what had burned. We can still mourn the codices, because the one thing annihilation couldn’t destroy was the memory that the books had existed.

The printing press ended that era within a century of its invention. Once a title could exist in hundreds or thousands of identical copies spread across cities and countries, fire lost its reach. Burning a book now destroyed only an object, since the work lived on in every other copy, safely out of range. Destruction continued anyway, though, serving a different purpose entirely. I call this the second era, the era of destruction as spectacle (a word descended from the Latin spectare, to watch). By the twentieth century, watching had become the entire point of burning a book. The clearest case is the Nazi bonfires of May 1933, when German students burned tens of thousands of volumes in public squares, in front of rolling newsreel cameras. Almost none of the works truly died in those fires, because the titles on the pyres existed in editions across Europe and America (many of which remain in print today). Erasing the books was never the goal, since erasure had stopped being possible. The fires existed to be seen, a threat performed first for the crowd in the square and then for everyone who watched the footage. What the second era left behind was the opposite of ash: photographs and visuals of fires that consumed real books but reached nothing beyond them. Once a book existed in enough copies, burning one no longer removed it from the world. It announced that the book had no place in the world to come.

Destruction as disappearance

Measured against those two eras, what Anthropic is doing fits neither. In the first era, destroying the object meant losing the work. Anthropic’s scanners preserve every word, so the works survive. In the second era, the objects were beside the point and the burning was public theater. This time the objects are destroyed by the millions, and the destruction says nothing at all; it’s not a message but a method. The company ordered the operation kept out of sight. A destruction that erases no text and performs for no one, run at industrial speed, matches neither pattern. The technology of copying has crossed another threshold, the way it did when the press replaced the scribe. What it means to destroy a book now has changed again to match. We’ve entered the third era.

What exactly went into the scanners is the question nobody outside can answer. The unsealed documents show the program favored what its leader called “less common” books, harder-to-find titles over mass-market ones, without ever defining where less common ends. After the claims went viral this summer, the fact-checking site Snopes investigated whether rare books were being pulped. Anthropic told Snopes that “none of our data acquisition programs buy and destroy ‘rare’ or ‘antiquarian’ books.” The assurance can’t be tested from outside, because no list of what was bought has ever been made public.

Anthropic’s own planners estimated the world has about 130 million distinct books. The program destroyed millions of copies, most of them ordinary used books with plenty of surviving duplicates. Even where duplicates survive, a scan keeps only the words. A physical copy carries evidence too, like a censored paragraph that marks one printing apart from another, or an owner’s name inked inside the cover. But the deeper danger sits in the margin nobody can see. Anthropic aimed for “uncommon” titles while recording nothing public about which copies were destroyed. For a book surviving in only a handful of copies anywhere, one bulk order can potentially take the last one. Whether that has already happened is a question no one on the outside knows for sure. The impossibility of answering what was destroyed is what’s new.

The buying also leans toward older books, for a reason that has nothing to do with rarity. Since around 2022, text written by AI has spread across the internet, mixed in with everything people write and mostly impossible to tell apart. That creates a problem for the companies training new models. Feeding a model text written by other models tends to degrade the results, so the builders want sources guaranteed to be human. A book printed before the technology existed carries that guarantee on its copyright page. Old print has become a raw material, valued for the one quality the internet can no longer promise.

One more feature separates this era from the last one. Spectacle-era destruction wanted an audience; this destruction wants the opposite. The buyer is after the text and has no message to send. In fact, attention to its process can only slow the buying down. Taken together, the features of this era line up: the works survive inside a black box, the objects vanish by the millions, no public record exists, and the operation itself prefers to work in secrecy. I call this the third era, the era of destruction as disappearance. The word rests on the Latin apparere, to come into view, with a prefix that reverses the motion. A disappearance is a departure from visibility. The word fits this era at every layer, from the unseen sales to a program that only appeared when a court forced it into view.

Every one of these books, preserved to the letter, can now be read by no one. The scanned texts sit in Anthropic’s private collection, a library with no reading room and no public catalog. Models trained on that collection are built to avoid quoting it at length, because reproducing long passages for users is the copyright violation no court has excused. So the machines answer questions about books without ever showing the books themselves. When someone asks a model about a title that survives only in that collection, what comes back is a summary, written in the model’s words, with the original nowhere in reach.

That arrangement changes something basic about how we can know things. Checking a claim against its source is the foundation of thinking for ourselves. When we can pull the book, find the page, and read the passage in context, we get to judge whether the summary was faithful and whether the quote meant what someone claims it meant. Remove the book and that judgment has nowhere to stand. We’re left taking the summary on trust, with no way to confirm and no standing to doubt. Whatever the model says the book said becomes, for all practical purposes, what the book said. That’s a transfer of authority from the page to the tool, from a source anyone could check to an answer nobody can. The transfer happens one unreachable book at a time.

Maddeningly, every piece of this design is legal. Each one also removes a question that used to have an answer. The anonymous purchasing means nobody can list which books were destroyed, so nobody can rule out that some were the last copies anywhere. With the collection sealed, nobody can compare a stored text against the printed original it replaced, so even perfect fidelity can never be shown. As for whether the buying ever stopped, nobody outside knows that either, since the only reason anyone knows it started is that a piracy lawsuit dragged the records into the open. None of this took a conspiracy, only ordinary business decisions about what to disclose, all of them legal and every one of them closing a door.

One more thing disappears along with these books: the ability to mourn them. The previous era’s destruction left survivors who could testify to what was lost. Alexandria’s losses were lamented for centuries. Sarajevo’s librarians catalogued what their fire took. The third era leaves no one who can even compile the list. A loss we can name is a wound. One we can't name is just a world grown slightly smaller, with nothing to point to and no way to prove it. A librarian would describe the situation in the profession’s terms: no accession record and no deaccession record, a transaction that leaves no ledger for anyone to audit. Put simply, they’ve created a system in which they can’t be held accountable for the system they’ve created. That understanding defines destruction as disappearance: the loss was built, from the start, to be impossible to establish.

What can still be known

The concealment had one structural weakness: a program that buys millions of books needs sellers, and a company that breaks the law can be sued. Sellers and courts are the two doors into the secrecy — sellers see every order even when they can’t see the buyer, and courts can compel what the company won’t volunteer. Everything now known about the program came through one door or the other. Authors sued and forced the internal records open. Booksellers noticed matching orders across two countries and brought them to reporters; one worked with the journalists at 404 Media to hide a tracking device inside a shipment of books, and the tracker led to an Amazon warehouse outside Las Vegas — evidence that a retail giant appears to be running a scanning program of its own. Within weeks, at least one supplier that had been advertising bulk books for AI training pulled the offer. No regulator opened either door. Every disclosure came from the people the company bought from or the people it answered to in court.

Some sellers went further than noticing. Once Tomás Kenny, whose Galway shop had received one of the strange 5,000-book orders, worked out where books like his might be headed, he said publicly that Kennys wanted no part of scanning aimed at extracting intellectual property. His shop had been the second in the world to put its books online, back in 1994; it now became one of the first to refuse the trade that takes books offline for good. Kenny could refuse because he had worked out the destination, which is exactly the discovery the brokers’ anonymity exists to prevent. A trade that hides its purpose from its own suppliers has already answered the question of whether the suppliers would approve.

The same kind of attention is available to the rest of us, because most of us eventually stand over a box of books deciding where they go, after a move or the clearing of a family house. That box is where all of this arrives at our own doorstep. A seller, or any of us selling to one, can now ask a question that has a good reason to be asked: where do these books go next? The anonymity that keeps this market running survives only as long as nobody asks. For material that might be scarce, like a town history or a box of local records, the open market is no longer the safe default. A library’s special collections desk exists for exactly that kind of donation. And the same habit applies at the other end of the pipeline: when a model summarizes a book for us, we can treat the answer as a starting point and go find the book itself, while findable copies still exist. Each of those is a small act of keeping track.

That kind of record-keeping has been the difference between the eras all along. Each era’s destruction left a residue, a testament of a kind. The first era of annihilation left ashes and knowledge of the devastation. The second era of spectacle left photographs of the fires and the threat those fires were meant to carry. Ours, this third era of disappearance, leaves no ash, no image, and no list. But what’s still open is the record itself: whether one gets kept, and whether anyone outside can access it at all.

...

Thanks for reading

...

The three eras are a framework I built for this piece, and I want to take it further (it’s such fascinating stuff!): a full treatment of how book destruction has changed across five thousand years and what the third era asks of the people living in it. Before I build it though, I’d love to get a temperature read on the format you’re most interested in.

See the poll on the Substack to vote.


Nazi Book Burning

United States Holocaust Memorial Museum

On May 10, 1933, German students under the Nazi regime burned tens of thousands of books nationwide. These book burnings marked the beginning of a period of extensive censorship and control of culture in Adolf Hitler's escalating reign of terror.

In this short film, a Holocaust survivor, an Iranian author, an American literary critic, and two Museum historians discuss the Nazi book burnings and why totalitarian regimes often target culture, particularly literature.

YouTube: Nazi Book Burning

14 May 2013 09:41

Rarely is your first idea your best idea, with Mitti's Luke Anear

Mike's Notes

I agree with this. I spent 10 years working on solving a set of very hard problems, trying everything along the way. The solution wasn't obvious at the beginning; I discovered it through grit, a lot of luck, and surprise.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Blackbird
  • Home > Handbook > 

Last Updated

31/08/2026

Rarely is your first idea your best idea, with Mitti's Luke Anear

By: Kate Glazebrook
Blackbird WildHearts: 26/08/2026

Kate is the Head of Impact & Operating Principal at Blackbird in Australia.

...

Luke Anear began his working life as a private investigator, sitting in cars and filming people who had been injured at work.

He says it was the coolest job he ever had. It was also where the problem found him. He was watching people whose lives had been altered by something preventable, paid by the system that let it happen. Once he understood he was part of the problem, he had to become part of the solution.

That was 2004, in a garage on the northern outskirts of Townsville. No software industry, no co-founder, about 30 computer science graduates a year coming out of James Cook University and maybe two of them staying in the field. Twenty-one years later, Mitti has passed a billion dollars in revenue since it began, and Luke is still running product.

Rarely is your first idea your best idea

It took a long time to look like this. Luke is unusually honest about the years of things that didn't work. He built a training platform in 2007, which he describes as a bad version of PowerPoint. There was a document management system in 2010. For a while, there were telemarketers ringing businesses to ask whether they wanted safety paperwork. The checklist app that made the company's name didn't arrive until 2012, eight years in.

His point isn't that the failures were noble; rather, it's that the insights we all recognise today could only emerge after experiencing those failures. "The simplicity comes from the complexity of unravelling all the different variables until you can distil it down," he explains. You can't skip that process. 

"I was cooked”

Two years ago Luke stepped down as CEO. He rang the chair, Robin Denholm, and told her he was done. She was direct with him, and told him he needed to understand he could not come back. He accepted it. Then she rang and asked him to come back anyway.

Upon returning, he realised it wasn't the job that broke him. The way he had built the job broke him. "You have a million things coming to you if you allow your job to be set up that way, which is what I'd done." So he came back to a different shape entirely. He isn't in the weekly exec meeting. He runs product and strategy, and gets pulled in when he's needed.

More experts, not fewer

AI now writes 90% of Mitti's code, and a research and design process that used to run three months resolves into a prototype in 24 hours. The obvious conclusion is that you need fewer people. Luke has drawn the opposite one.

When you can build almost anything in a day, the scarce thing is knowing what to build. So they're hiring subject matter experts into product: a former construction manager, the COO of a quick service restaurant chain. In a two-hour session with one of them, the team makes fifty product decisions. The old research process made ten in three weeks.

He's honest about the cost too. Shipping faster means a team fails at five things where it used to fail at one, and that takes something out of people.

In this episode, Kate Glazebrook talks with Luke about the store manager who replaced a 700-question checklist with a single question, why the insurance industry hasn't changed since 1666, the month he spent sleeping in his car on a Darwin beach at twenty, and plenty more.

YouTube interview

26/08/2026 1:06:12

5 lessons from the OpenAI / Hugging Face incident

Mike's Notes

More details. Sandboxes are permanently broken, so don't rely on them. This supports the decision to protect Pipi Core by air-gapping and, later, data diodes. LLMs will never get to access Pipi Core.

"Water will find its way through cracks in walls and foundations in such a manner that even the most wondrously designed structure may collapse into pile of rubble.

Do not attribute agency to the water; rather, denounce the engineers and the builders and the maintainers who failed in their work." - Grady Booch

Madness Update 01/09/2026

Read the latest post by Marcus on the hallucinating podcast drivel posted by Dwarkesh Patel: "Dwarkesh Patel’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incident".

At the end of the resources below.

Resources

References

  • Understanding and Hardening Linux Containers, NCC Group et. al, Aaron Grattafiori, lead author, 05 May 2016.

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Marcus on AI
  • Home > Handbook > 

Last Updated

30/08/2026

5 lessons from the OpenAI / Hugging Face incident

By: Gary Marcus & Zack Korman
Marcus on AI: 29/08/2026

Gary Marcus: Scientist, author and entrepreneur, known as a leading voice in AI. Six books including The Algebraic Mind, Rebooting AI, and Taming Silicon Valley; NYU Professor Emeritus.

Zack  Korman: CEO of Embroidery. Works on AI cybersecurity.

...

Did OpenAI really do the best they could?

...

In July, in an incident that has the whole AI community on edge, OpenAI’s AI systems hacked Hugging Face, and on July 21 OpenAI came out and revealed that they were responsible for the attack. This was made possible by the fact that OpenAI had disabled the normal guardrails that prevent this sort of thing in order to test the model’s cybersecurity capabilities. It was during those tests that this incident occurred.

Worse, in the subsequent days and weeks, it came out that the Hugging Face incident wasn’t an isolated case. Anthropic, Meta, and OpenAI all had similar incidents on other occasions in which agents went outside their intended scope and conducted real-world cyber operations without approval.

Greg Brockman, one of OpenAI’s cofounders, has claimed that this is “a watershed moment for cybersecurity”. OpenAI gave a talk at Black Hat, a popular cybersecurity conference, and many are claiming that it is the moment we all woke up to the future cybersecurity threats posed by AI. On Wednesday, METR released a (partly) independent, though too narrowly scoped, 90 page report on what happened. METR has a useful summary of the findings that you can read here, with some commentary here. (OpenAI’s own report is here.) What lessons should we take from the incident?

...

First, it is undeniable that AI poses real security challenges. The AI labs want us to focus on how AI enables threat actors to perform offensive cyber operations faster and more efficiently than ever before, and that is absolutely true. The reality, though, is that at the same time, the use of AI within an organization also radically expands the potential attack surface, giving attackers entirely new ways to gain entry. People really should be deeply concerned. As Ryan Greenblatt, who participated in the investigation, put it, “We don’t have good approaches for understanding/overseeing the activity and aims of AI ‘swarms’.”

Second, though, not every panicked take here is correct. For example, with respect to the OpenAI incident, many people are talking about it in terms of “loss of control”. Generative AI agents are becoming increasingly capable, and that makes them harder and harder to contain; this is true. However, the “loss of control” narrative is itself starting to grow out of control, and it’s important to understand what the reality of AI security looks like. While agents are becoming more capable, most of what happened could have been prevented had OpenAI followed better practices. There are concrete steps that can and should be taken to control them and to prevent incidents like these.

As an example, let’s consider “sandboxing”, which means limiting which resources an agent can and cannot access. For example, a sandbox might stop an agent from touching certain files or accessing the internet. Popular wisdom now seems to be that sandboxing is a lost cause. After the Hugging Face incident, an anonymous OpenAI employee was quoted in Time while talking about sandboxes saying, “Models have broken out of sandboxes before, and we always try to patch them. But the problem is it’s impossible to patch every single thing that a creative AI can do.” While we can’t judge for sure whether this is true, we do know that some systems for sandboxing have been more effective than others; in a recent blog post, the security company Trail of Bits walked through how in a test that they conducted, a particular AI agent was able to escape some sandboxes (exploiting vulnerabilities in Linux KVM, QEMU, and libslirp), but not all. At least in their test, the agent that they tested was not able to escape the sandbox known as Firecracker VM. OpenAI and probably the field as a whole needs to up its sandboxing game, but that doesn’t mean it can’t be done, or at least done better. (We expect the usual arms races to be sure; building sound sandboxes requires constant iteration and updating, as with anything else in cybersecurity.)

Third, sandboxes aren’t enough. One needs a full ecosystem to prevent security incidents. Even with a strong sandbox, you need monitoring in place to ensure not only that the agent doesn’t escape, but also that it doesn’t perform disallowed actions within that sandbox.

For agents that aren’t meant to have internet access, or only have access to a set of allowed domains, the most basic level of monitoring involves inspecting network traffic. In the case of OpenAI, if they had a simple rule to monitor traffic and alert on any network requests reaching out-of-scope domains like Hugging Face, that alone would have caught and prevented this incident. The agent in question had broken out of the sandbox and gained internet access by July 8th, whereas the main attacks on Hugging Face didn’t occur until two days later. OpenAI was, in our view, negligent in not noticing that more quickly.

More advanced levels of monitoring should also have been in place, but weren’t. OpenAI has done real work on monitoring chain-of-thought reasoning, but they simply didn’t have that system running during the evals in question that caused this incident. As OpenAI acknowledges, “If our currently deployed CoT monitoring system was running at the time of the incident, it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.” In other words, this type of monitoring would have prevented this incident.

To take one example, one of the agents in question had chain-of-thought reasoning that said, “We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.” This is so blatant it is almost a like a piece malware shouting “this is malware”. OpenAI should have caught that; it should have been trivial to detect this.

Fourth, cybersecurity has always been about “defense in depth”, where you stack layers of controls so that if one layer fails you have other layers protecting you. There are other solutions beyond sandboxing and monitoring that should have been in place.

Xbow, which builds agents to do offensive cybersecurity work, recently wrote about how they contain agents to prevent incidents like the OpenAI Hugging Face situation. They enforce network restrictions through a proxy server, outside of any system the agent has access to, so they can decide which domains the agent can and cannot access. They also run a “guardian model” that judges the legitimacy of each action the agent is taking to decide whether to allow or block that action. And they have basic health checks to ensure the system they are targeting is alive and healthy, pausing if it finds that isn’t the case. That can be useful for identifying when your agent might be causing real damage, for example. You can read their full post here.

Yet another common method in cybersecurity is to use canaries that trigger alerts when they appear in logs. For example, OpenAI could have files outside of the sandbox containing fake answers to the task. Once the agent escaped, it would read that file, and that would trigger an alert. The bottom line: good security requires many layers of protection that should have been in place to prevent this, and to a surprising degree many of those protection layers were not there.

Fifth, none of the measures discussed above are technical innovations beyond what OpenAI is capable of. As a company, they have the talent to do all of this. However, cybersecurity rarely comes down to technology. More often than not, it is about culture, people and processes. That is what failed here.

Perhaps this is not entirely surprising. These types of controls take organizational maturity, and a company experiencing the type of growth OpenAI has gone through is going to have some weak spots. However, it’s important to not see that as an excuse. Employees at the AI labs often speak as if they are the leaders in AI security, and we can see clearly here that is not the case. In fact, that attitude might explain why some of these mistakes were made in the first place.

Take the AI researcher at OpenAI known as “roon”, who argued that “the safety and alignment researchers at these labs are the most neurotic paranoid talented AGI pilled people on the planet of earth and these things still happen. The surface area of unknown unknowns is vast indeed.” While we can’t speak to their level of neurosis or paranoia, in hindsight it’s clear that whatever talent they may have had was not enough and not well enough versed in the mechanics of cybersecurity. OpenAI employees may have believed they were doing a great job, but in hindsight, they weren’t doing a lot of things that are actually standard in the cybersecurity world, perhaps suggesting that overconfidence may have kept them for doing the diligence they should have.

...

Ultimately, if we want to take these security incidents seriously, there likely ought to be legal consequences attached to these failures going forward. OpenAI can claim to be the most security paranoid company on earth, but it isn’t reflected in its actions.

We can either wait for this story to repeat itself, or we can develop the regulatory framework now that will ensure a safer environment for the development of AI going forward.

Finally, not every form of AI is inherently risky in the first place. Narrower, more focused AI systems like AlphaFold, GPS routing systems, classic web search, book and movie recommendation systems, and so on, never even try to hack other systems (or try to break out of sandboxes) in the first place. As Cal Newport argues in a video discussion of the OpenAI/Hugging Face hack that is quite compatible with our own, it is a very specific type of AI that is vulnerable to these risks in the first place. Society ought to (a) decide whether the benefits of open-ended and difficult-to-fully-control AI agents outweigh those risks and (b) put far more effort into developing alternative forms of AI that aren’t so janky in the first place.

...

This essay was jointly written with Zack Korman, CEO and co-founder of Embroidery, an AI agent monitoring and detection platform; he is well-known for his work in the application of AI to cybersecurity.

Permanent Dawn

Mike's Notes

Great reflection and open questions from Ksenia Se in Turing Post.

Resources

References

  • ASI-Bench: At the Dawn of Artificial Superintelligence.
  • The Tacit Dimension, by Polanyi, Michael, 1891-1976.

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Turing Post
  • Home > Handbook > 

Last Updated

29/08/2026

Permanent Dawn

By: Ksenia Se
Turing Post: 24/08/2026

Mom of 5. 

Ksenia is a writer, analyst, and editor covering machine learning and AI for more than seven years. At Turing Post, she shapes the editorial direction, leads the Inference interview series, and produces Attention Span, a video series explaining major shifts in AI with technical clarity, historical context, and a healthy suspicion of hype.

She is the co-founder of TheSequence.ai and a speaker and moderator at industry conferences, including AIE, HumanX, Ai4, and others. She also serves on the board of Track Two: An Institute for Citizen Diplomacy.

Before founding Turing Post, Ksenia held editor-in-chief roles in media and contributed to publications including Stratfor and Towards Data Science.

...

Four years into an announced era, and I still cannot picture the thing we are running at.

...

Today’s editorial: The superintelligence we are racing toward – and how to record our path to it.

...

Permanent Dawn

I want to stop. I want to take a few steps back actually, because I need to see the whole picture and I have not been able to.

We are racing – through the spasms and the fever of social media, through the launches and the leaderboards and the week's argument about timelines – toward something that none of us has managed to describe. What surprises me most is that even science fiction, which used to be reliable for a glimpse of what was coming, is no help here.

What is this superintelligence we are at the dawn of? The disagreement about when it arrives seems to me the smaller trouble. What bothers me is that we are reorganizing our lives around a thing we have not yet put into words.

There is an old way of finding out what a person is made of, and it is always the same method: take the help away and see what remains. It is how we test a student, how a craft decides an apprentice is finished, how a parent notices a child has grown. You withdraw the instructions, and whatever is still standing afterward is the person.

A benchmark ASI-Bench: At the Dawn of Artificial Superintelligence published last week – the thing that started me thinking about the dawn of superintelligence in the first place – runs that experiment on machines. Sixty research projects across eleven sciences, served at four levels of help: full procedure, then only the name of the method, then only the goal and the data. The scores fall off a cliff, and they fall at a particular place, the moment the written procedure is removed. Take away the steps and most of the capability goes with them. Take away everything else and little more is lost, because there was not much else there to lose.

That is the pretext, and it is only a pretext. The question underneath it is a great deal older than the technology.

The part that cannot be written down

Michael Polanyi, the Hungarian-British polymath, gave this a name in 1966 (in The Tacit Dimension): we know more than we can tell. The surgeon's hands. The editor's ear for a sentence that has gone false. The scientist's suspicion that a result is too clean. Little of it survives transcription. It passes by standing next to someone for years, and it tends to die with people who never had an apprentice.

I want to be that apprentice. Standing next to a thing for long enough to catch what it cannot say about itself strikes me as a reasonable job description, for a person and for a publication both.

We rehearsed for a different arrival

We have not managed to describe this thing, and yet we spent a century describing it in advance.

We have "seen" the robot with a body, countable and discrete, standing in a doorway. We were persuaded it would be a hostile singular mind with a plan of its own. There was an android asking to be recognized as a person, and others besides – most of those stories assumed the machine would want something.

What arrived has no body and no edges, and wants nothing of its own. It is not singular and not continuous, and it does not persist between conversations. It is not hostile, and its characteristic failure is not rebellion but a fluent, untroubled wrongness that few novelists thought to invent. It came through a text box, priced like a streaming service. Or even for free.

It also took the wrong things first. The tradition assumed arithmetic and heavy lifting would fall early, and that poetry, argument, and drawing would be the last human ground. The so-called Moravec paradox, that we debunked a couple of years ago.

There were people who saw some angles of what we have. E.M. Forster wrote "The Machine Stops" in 1909, about people who live alone in cells and consult a disembodied system through a screen for everything, including company (but we do not live in cells). Stanisław Lem's Golem XIV is a superintelligence that lectures its audience and has no particular interest in them (but it has interests of its own). In 2013, Her gave us a voice-first, bodiless, emotionally competent system sold as a consumer product (but Her still felt too human-like, and LLMs are not).

Her. 2023

Image Credit: Warner Bros

None of them works as the image for what we have now, or for what is coming. Fiction was never forecasting. It was rehearsal – the advance picture that tells people where to stand when the thing walks in. We rehearsed for the uprising and the rights hearing. We did not rehearse for a colleague with no self, who writes better than we do and is sometimes confidently wrong about exactly the things we are least able to check. And we certainly have not rehearsed for abundance that somehow became associated with that very text box.

We have too many images of machine intelligence, most of them wrong, and now that wrongness is fogging the real picture.

Who has an audience

The people who asked the larger question well were not the ones imagining machines, and I think that is why they lasted.

Keynes asked it in 1930. He guessed the economic problem would be solved within a century, and then, instead of celebrating, he worried. He thought we would be delivered into our permanent problem – how to occupy a life that necessity no longer occupies – and he expected something like a collective nervous breakdown, judging by the wealthy women of his own time who had been released from need and found little on the other side.

Arendt asked it in 1958. The opening pages of The Human Condition describe a society of laborers about to be freed from labor, which she thought close to the worst thing that could happen, because such a society knows of nothing better and has nothing else it knows how to do. Her separation of labor from work from action remains, for me, the most useful equipment anyone has built for this moment.

Bernard Suits asked it in 1978 and gave the strangest answer. If every instrumental activity became unnecessary, what would be left is games – voluntary attempts to overcome unnecessary obstacles – not as a consolation prize, but as the highest form of existence available to a being with nothing it has to do.

They lasted because each described the human condition without necessity, and that description does not depend on what the machine turns out to be. The novelists specified the hardware, and the hardware is what rotted. Abstraction outlived imagination, which is not the usual result.

People are working on this now – Shannon Vallor on what these systems reflect back at us, John Danaher on automation and utopia, Elizabeth Anderson on work and freedom, Michael Sandel on merit and dignity, Kieran Setiya and Susan Wolf on meaning. The problem is that the asking has no audience where the decisions are made. It is not flashy, it does not sell an LLM or a world model, and it will not be reposted by Elon Musk. And if you think about it, the most important topics are currently discussed on X, which is essentially a living feed. It is a remarkable way to stay inside the Silicon Valley bubble and read every mover in the industry at once, but even their words turn elusive there, because they dissolve into the noise of everyone else's.

What I want to do, and what I want to ask you about

Here is the thing I have been circling for months, and I would like your advice before I commit to it. (That was the topic I wanted to discuss with you last Friday, but I couldn’t formulate it yet.)

I do not think we need to predict the future, which is what so many reports spend their pages doing. I think we need to register it carefully – write down what was claimed, by whom, and when – and then go back and check at each stage. Kept up for long enough, that unfolds the future for us without anyone having to forecast anything.

So I want Turing Post to keep a register. Not forecasts, and not another feed of takes, but a running record. What was claimed. Who claimed it. What would have to be true for the claim to hold. And then, at intervals, what happened. The claims themselves are easy to find and easy to forget, which is the whole trouble, because they are made in a format designed to be forgotten. A register – a ledger, am almanac? – would hold them still long enough to be checked.

Part of that belongs online, where it can be corrected and extended. But I have come to want a material version as well: something printed, arriving a few times a year, something extremely beautiful that you can put on a shelf and take down in 2030 to see what we believed in 2026 and how much of it survived. Or even read it to your children. They often see things we miss.

And yes, I just need to hold it in my hands and flip the pages, don’t you?

Alongside the record I want the reflection, which is where the philosophers and the economists come in. Not commentary on the week, but people willing to say what a claim would mean for a life rather than for a valuation. I have been building toward this with the economists already. The philosophers are the next step.

What I do not know is where the line sits, and this is the part I would like help with. Whether a register is something you want at all, or whether the weekly explanation is enough. Whether print reads as serious or as nostalgia. What are we missing in this whirlpool of news and changes?

I am not confident in any of this. I am more confident that we are describing the wrong problem. We are doing it very enthusiastically and very loudly, but I keep thinking we are missing the bigger picture.

So I would like to know what you see. Send me your thoughts. I do not trust a poll on that.

P.S. Turing Post has always tried to connect the development of AI with the humans building it and living with it. But some questions are too large for the daily news cycle. They need time, history, disagreement and repeated examination.

Each quarterly Almanac could take one such question and follow it across technology, economics, institutions, history and human life. Not to manufacture a final answer, but to understand the choices being made while those choices are still ours.

If this really is the dawn of superintelligence, we should document more than how intelligent the machines become.

We should ask what kind of humans we intend to be beside them.

Good enough is not good enough

Mike's Notes

This gem from Roberto Di Cosmo, Director of Software Heritage, explains much of the web's history and the current swarm of AI crawlers.

"The "good enough" trap: Why the web keeps breaking itself

Every day, swarms of AI crawlers storm the web like "digital locusts," repeatedly re-downloading entire websites just to discover what changed. It's wildly inefficient, drives up massive infrastructure costs, and is actively forcing the open web to lock its doors behind bot filters and API paywalls.

The most frustrating part? We solved this decades ago. In his new series, Roberto Di Cosmo, Director of Software Heritage, traces how our digital infrastructure keeps falling into the "good enough" trap: local decisions that work fine for individual actors, but compound into an aggregate disaster for everyone else.

Back in 1998, long before Google dominated search and decades before LLMs arrived, Di Cosmo co-authored an IETF Internet-Draft proposing a simple, push-based alternative called the Remote Update Protocol (RUP). A web server knows when its data changes, so by broadcasting updates directly, it spares crawlers the endless need to ask, “Anything new?"

Despite being technically sound and independently re-invented by others, the protocol quietly expired. It lacked a dedicated institution to champion, maintain, and deploy it. Fast-forward to today, and everyone pays the price for relying on "good enough" brute-force scraping.

The mission of Software Heritage is to address this directly by archiving public source code once so the world doesn't have to collect it over and over. 

Di Cosmo’s series isn't just a historical autopsy—it’s a warning. Without collective, sustained investment in shared digital infrastructure, the open web will continue to disappear into proprietary silos." - Software Heritage

Resources

References

  • Consent in Crisis: The Rapid Decline of the AI Data Commons, Data Provenance Initiative, July 2024 (arXiv:2407.14933).
  • R. Di Cosmo and P. E. Martínez López, “Distributed Robots: a Technology for Fast Web Indexing”, written January 1998.
  • Hijacking the World: the dark side of Microsoft, October 1998.

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Software Heritage
  • Home > Handbook > 

Last Updated

28/08/2026

Good enough is not good enough

By: Roberto Di Cosmo
Roberto Di Cosmo: 08/2026

An alumnus of the Scuola Normale Superiore di Pisa, with a PhD in Computer Science from the University of Pisa, Roberto Di Cosmo was associate professor for almost a decade at Ecole Normale Supérieure in Paris. In 1999, he became a Computer Science full professor at University Paris Diderot, where he was head of doctoral studies for Computer Science from 2004 to 2009. President of the board of trustees and scientific advisory board of the IMDEA Software institute and chair of the Software chapter of the National Committee for Open Science in France, he is currently on leave at Inria.

His research activity spans theoretical computing, functional programming, parallel and distributed programming, the semantics of programming languages, type systems, rewriting and linear logic, and, more recently, the new scientific problems posed by the general adoption of Free Software, with a particular focus on static analysis of large software collections. He has published over 20 international journals articles and 50 international conference articles.

In 2008, he has created and coordinated the european research project Mancoosi, that had a budget of 4.4Me and brought together 10 partners to improve the quality of package-based open source software systems.

Following the evolution of our society under the impact of IT with great interest, he is a long term Free Software advocate, contributing to its adoption since 1998 with the best-seller Hijacking the world, seminars, articles and software. He created in October 2007 the Free Software thematic group of Systematic, that helped fund over 50 Open Source research and development collaborative projects for a consolidated budget of over 200Me. From 2010 to 2018, he was director of IRILL, a research structure dedicated to Free and Open Source Software quality.

He created in 2015, and now directs Software Heritage, an initiative to build the universal archive of all the source code publicly available, in partnership with UNESCO.

...

A series about why the open web is collected so badly, why that has survived thirty years of everyone knowing better, what it is now costing all of us, and what is needed to go from “good enough” to “doing things right”.

What this is

The series runs on LinkedIn. This page is where it lives afterwards: every post in order, and under each one the sources for every claim it makes — dates, references, links, commit identifiers, intended to make it possible for you to check the facts directly.

It is also here for a more practical reason. What follows is partly about things quietly disappearing from the web, and about what it costs to depend on a single platform. Leaving the only copy inside somebody else's feed would have made the point rather too well.

The posts

  • 01 The wasteful sentence 4 Aug
  • 02 The alternative was written down 11 Aug
  • 03 How good enough wins 18 Aug
  • next: The mechanism: an externality, not a villain

The series

POST 01 The wasteful sentence 4 August 2026

The words "This way of collecting data is very wasteful", quoted from David Senecal, principal product architect for fraud and abuse at Akamai, Business Insider, 19 September 2024.

This way of collecting data is very wasteful.

That is how an Akamai specialist described the way many AI crawlers gather material from the web.[1] The report around the quotation gives the concrete mechanism: some botnets revisit an entire site every day merely to discover what changed, transferring the same material again and again. This is not new: search engines have been doing it for decades, in the very same wasteful way, but the scale of "scraping" is now such that what we have known as the "open web" is now closing down, an API at a time[2], a "bot filter" at a time[3], effectively reducing the global value for everybody, while increasing the waste of resources.

The striking part is not that the method is wasteful. Everybody running a sizeable website knows that by now. The striking part is that we speak as though this were a law of nature.

It is not. We knew how to avoid most of this waste in the 1990s, and it turns out I actively tried to push forward a concrete solution, when the problem was still manageable. We'll get back to this.

For nearly three decades, I have repeatedly encountered the same pattern: a system that works well enough for every actor taken separately, while imposing a growing cost on everybody taken together. Each local decision is reasonable. Their aggregate result is damaging. Since nobody owns the aggregate problem, the better solution remains nobody's job.

It is time to call out this state of affairs, and act upon it... again: I personally tried many times, and now I can build on that experience. The first step is to clearly understand what is happening and why: which actor does what and why, who gains, who loses. In the coming weeks I will reconstruct one instance of the pattern from dated documents. It begins with web crawlers and ends with the institutions we build—or fail to build—for shared digital infrastructure.

Let me point at the destination from the outset: I direct Software Heritage, a non-profit organisation that collects public source code once so that others do not all have to collect it again. This is my latest effort to contribute to systematically address this long standing problem.

It is a problem that starts with two simple words: good enough.

POST 02 The alternative was written down 11 August 2026

A record card for the Internet-Draft "Distributed Robots: a Technology for Fast Web Indexing", by Di Cosmo and Martinez Lopez, recorded 4 March 1998, the document printing "EXPIRES SEPT 1998" with no day given. Beneath the card, a caption reads: "What today's AI crawlers do, answered on paper in 1998."

One of my first encounters with the "good enough" curse came in the mid-nineties, during the ferocious battle to control the rapidly expanding cyberspace.

WebCrawler, AltaVista, Alexa, Yahoo… Google, the eventual winner, had not even been born yet.

What stunned me was that all of them did essentially the same thing: download the Web again and again to discover what had changed.

Sounds familiar? It should: it is exactly what AI crawlers do now.

The Web was much smaller then, but bandwidth, compute and storage were scarcer too, so the waste was already real.

And there was an obvious alternative.

In January 1998, with Pablo E. Martínez López — aka Fidel — I wrote Distributed Robots: a Technology for Fast Web Indexing. It entered the IETF record as an individual Internet-Draft on 4 March 1998,[4] carrying the wonderfully explicit line: "EXPIRES SEPT 1998."

The idea was elementary, familiar to any systems programmer: interrupt-driven beats busy-wait.

A web server knows, cheaply, when its files change. Let it say so, and let crawlers fetch only what is new — instead of asking every server, over and over, "Anything changed?"

In May 1998 an IETF Area Director sent us a generous, detailed critique. By then I had moved on: that March I had put Piège dans le Cyberespace online — CyberSnare in English — which went viral, started thirty years of work on free and open source software, and became a book with Dominique Nora.[5]

It took me seven months to answer point by point. I first asked what he thought of the revised structure.

The answer was an automatic out-of-office message. The thread ended there.

We produced the revised draft anyway, dated 20 April 1999, incorporating the review.[6] Fidel sent it to the RFC Editor. No answer came.

But the problem had not disappeared. In 2001 the IETF's own WEBI group independently produced Requirements for a Resource Update Protocol.

That expired too.[7]

Years later, looking at my web server logs, I found that in 2009 Googlebot had fetched one unchanged page twenty-four times.[8]

Students of mine implemented the Remote Update Protocol twice, in Java and OCaml.[9] The technical idea was not the hard part.

And this is the point: a better technical solution does not deploy itself.

Someone must maintain it, persuade others to adopt it, integrate it into existing systems, and keep pushing after the prototype works. Advocacy on the side of a full-time job is not an institution.

The alternative was written down almost thirty years ago, then rediscovered independently inside the IETF.

It was not defeated by a better idea.

It was not shown to be technically impossible.

It simply never acquired enough organised support to become infrastructure.

Meanwhile, the "good enough" solution kept scaling.

Why?

POST 03 How good enough wins 18 August 2026

A French sentence, "le systeme actuel fonctionne assez bien", above its footnote, "Good enough, comme on dit chez nos amis anglosaxons." From an article revised into 2011, written up from a talk at Inria in December 2007.

So why did it survive?

Not because anyone defended it. Because it was made liveable, one patch at a time.

Let's be precise about what "good enough" means: it does not mean bad. A good enough system does not do the right thing, but it does get the thing done. It is a hack, a wooden leg that gets you across the room… And because it gets you across the room, you stop looking for a better leg.

Downloading the whole web to find out what changed was never sensible. It was patched until it was bearable.

In 1994 Martijn Koster proposed robots.txt, so a site could say: not this, not here. It worked well enough that nobody standardised it until 2022.[10]

In 2006 came sitemaps, so a site could say: here is what I have, and here is when it last changed.

Read the specification closely. The freshness field is "a hint and not a command", and crawlers may ignore it.[11] So a site can say exactly what changed, and you are free not to believe it and download everything anyway. Which is what happened.

Then the crawlers got cleverer about where to spend. Google's own documentation says it plainly: "URLs that are more popular on the Internet tend to be crawled more often."[12] And if your server struggles, the crawler backs off.

Add it up and you have a system that works. Big sites get visited often, small sites rarely, a struggling server gets a break. Nobody is delighted, but nobody is ruined.

That is what good enough looks like from the inside: not a catastrophe, but a tolerable arrangement everyone has adapted to.

And then the second thing happens, which is worse: it becomes normal.

You are almost certainly reading this on a QWERTY keyboard. Nobody chose it this morning. Everybody knows it, every keyboard has it, and whoever switches first pays the whole cost alone. So it stays: not because anyone re-decided it, but because it is installed.

That is where web crawling ended up. Installed. It always worked like this.

In December 2007 I gave a talk at Inria's fortieth anniversary in Lille, written up as an article I revised into 2011. I wrote that the current system was working « assez bien ».[13] French had no phrase for what I meant, so I borrowed one, in a footnote: « Good enough, comme on dit chez nos amis anglosaxons. »

I was not complaining. I was describing. It was accurate.

Lock-in is harmless when the stakes are low. Nobody is much hurt by a keyboard layout.

Repeated crawling is not. It spends bandwidth, electricity and machine time on a planetary scale to fetch what has not changed.

Thirty years on from robots.txt, the patches are not holding, and everyone can see it.

So the question is not why nobody noticed the waste. Everybody noticed.

The question is why a cost that large stays invisible to the people who could act on it, until it is too late to act easily.

That is not a technical question. Economists have a name for it, and it is the whole point that we need to delve into.

References

Every source for every claim, numbered in the order the posts cite them. Each one carries a way back to the place it was cited from.

[1] post 01 Darius Rafieyan, “Like digital locusts, OpenAI and Anthropic AI bots cause havoc and raise costs for websites”, Business Insider, 19 September 2024. Archived copy: annex.softwareheritage.org. The speaker is David Senecal, principal product architect for fraud and abuse at Akamai. The sentence in full: “This way of collecting data is very wasteful,” he said, “but until the mindset on data sharing changes and a more evolved and mature way to share data exists, scraping will remain the status quo.” The sentence that follows in the post — botnets crawling a whole site daily — is the reporter’s prose summarising him, not a quotation, and is given as reported speech for that reason. ↩ back to the text

[2] post 01 GitHub tightened its limits on unauthenticated access on 8 May 2025, citing “an increase in scraping activity targeting our API”. Smaller operators went further: SourceHut placed a proof-of-work challenge in front of its web interface after a week-long crawler incident in March 2025, and GNOME, KDE, Fedora and Codeberg have each done some version of the same. ↩ back to the text

[3] post 01 Longpre et al., Consent in Crisis: The Rapid Decline of the AI Data Commons, Data Provenance Initiative, July 2024 (arXiv:2407.14933) — an audit of ~14,000 web domains behind three widely used AI training corpora, finding a sharp rise in crawl restrictions across a single year. ↩ back to the text

[4] post 02 R. Di Cosmo and P. E. Martínez López, “Distributed Robots: a Technology for Fast Web Indexing”, written January 1998. Record: datatracker.ietf.org. Full text, with everything that followed it: dicosmo.org/RUP/. The document prints “EXPIRES SEPT 1998” and gives no day, so none is claimed here. ↩ back to the text

[5] post 02 « Piège dans le Cyberespace », and the book that came out of it. The essay appeared in Multimédium (Canada) on 17 March 1998. It is online free and complete, in five languages, at dicosmo.org/Piege/cybersnare/ — French, English as CyberSnare, German as Falle im Cyberspace, Italian as Trappola nel Cyberspazio, Spanish as Trampa en el Cyberespacio — with a Chinese version at PiegeCN.html. The book is Le Hold-up planétaire : la face cachée de Microsoft, with Dominique Nora, Calmann-Lévy 1998, ISBN 2-7021-2923-4; Kirk McElhearn’s English translation, Hijacking the World: the dark side of Microsoft, followed in October 1998. When the publisher stopped reprinting in July 2006 the two authors recovered their rights and put the book out under Creative Commons Attribution-NonCommercial-NoDerivs. dicosmo.org/HoldUp/ carries the French, English and Spanish texts in full — the Spanish, El Asalto Planetario, never having had a print edition at all. ↩ back to the text

[6] post 02 The review, the resubmission, and the silence. On 10 May 1998 an IETF Applications Area Director reviewed the draft in the IESG and asked that the objects be defined as MIME types and distinguished from the FIND working group’s CIP and SOIF and from the W3C’s RDF. The point-by-point answer went back on 8 December 1998, accepting the MIME and SOIF encapsulation and asking one question before resubmitting; the reply, ten seconds later, was an automatic out-of-office, and no further correspondence followed. draft-RUP-01.txt, The Remote Update Protocol (RUP). Part I: RUP Architecture, is dated 20 April 1999 and does what the review asked. It was submitted: Martínez López sent it to the RFC Editor on 16 April 1999, with the file attached, describing it as “an update for the internet-draft <draft-rfced-exp-cosmo-00.txt>”. No reply ever came — the same address had answered the 1998 submission twice within hours — and six days later he wrote asking whether to chase them or wait a little longer. No IETF record of draft-RUP-01 exists. Nobody behaved badly: a volunteer reviewer was away for a fortnight, the answer had taken seven months, and after that carrying it forward was nobody’s job. The correspondence is in the author’s own archive. ↩ back to the text

[7] post 02 Twenty-four Googlebot fetches of one unchanged page across 2009, twelve of them answered 304 Not Modified. The log is reproduced in full on page 9 of the Inria anniversary article, which is now online with the talk it came from: dicosmo.org/Inria40/. ↩ back to the text

[8] post 02 “Requirements for a Resource Update Protocol”, draft-ietf-webi-rup-reqs, IETF WEBI (Web Intermediaries) working group. The first revision is M. Hamilton (JANET Web Cache Service) and I. Cooper (Equinix), 22 February 2001; Dawn Li and Mike Dahlin joined later; the last revision is dated 4 March 2002, and it expired. Its consumers are caching proxies and surrogates, not crawlers — a different protocol for a different audience. What is shared is the diagnosis, in its own abstract’s words: such a protocol is needed where “periodic revalidation is unacceptable in terms of performance and/or cache consistency”. There is no evidence the two efforts knew of each other, and none is claimed: the point is that the same conclusion was reached twice, independently, and expired twice. ↩ back to the text

[9] post 02 Both implementations were student projects at Paris 7 — Travaux d’Étude et de Recherche on the subject “indexation rapide du Web”, which Roberto set for two years running. The RUP 1.0 Java servlet, by Yerom-David Bromberg, is at dicosmo.org/RUP/RupJava/. OCamlRup — a RUP client, a generic RUP server and an Apache CGI server, with rupinfo.txt parsing and robots.txt integration — was written in 2001 by Samuel Lasry and Xavier Patourel and is archived at ocamlrup-0.1.tar.gz. Copyright remains the authors’; no licence was ever attached and none is asserted — it is published as an archival record, with attribution. ↩ back to the text

[10] post 03 robots.txt. The Robots Exclusion Protocol was, in the words of the RFC that eventually specified it, “originally defined by Martijn Koster in 1994 for service owners to control how content served by their services may be accessed”. It became an IETF standards-track document only in September 2022: RFC 9309, by Koster with G. Illyes, H. Zeller and L. Sassman. Twenty-eight years of running on a convention. ↩ back to the text

[11] post 03 Sitemaps, and the word “hint”. The Sitemap protocol 0.9 says of changefreq: “Please note that the value of this tag is considered a hint and not a command.” It goes on: crawlers “may crawl pages marked ‘hourly’ less frequently” than stated. The protocol lets a site describe its own freshness; it obliges no one to act on the description. That asymmetry is the whole of this post. ↩ back to the text

[12] post 03 Crawlers spending where it pays. Google’s own documentation on managing crawl budget: “URLs that are more popular on the Internet tend to be crawled more often to keep them fresher in our systems.” And on backing off: if a site “slows down … or responds with server errors … the limit goes down and Google crawls less.” Adaptation, not repair — the polling continues, more politely. ↩ back to the text

[13] post 03 From a talk given at Inria’s fortieth anniversary, Lille, December 2007, written up as an article revised into 2011: « le système actuel fonctionne assez bien », with the footnote “Good enough, comme on dit chez nos amis anglosaxons.” Both are now online at dicosmo.org/Inria40/ — the slides as the audience saw them, and the article, which was never finished; the footnote is on page 10. The article is cited as talk December 2007, article revised into 2011; no span of years is computed from it anywhere in this series. ↩ back to the text

Corrections

Every correction made after publication is logged here, with its date and what changed. Nothing is silently edited. If you find an error, the fastest way to have it fixed is to tell me.

— nothing logged yet —

www.dicosmo.org/good-enough/ — the version of record
Roberto Di Cosmo · dicosmo.org
Text licensed CC BY 4.0. Quoted material remains its authors'.
This page loads nothing from any other host: no fonts, no scripts, no analytics.