Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Microducks

Mike's Notes

These are very funny. A good use for AI and robots.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

11/09/2026

Microducks

By: Mike Peters
On a Sandy Beach: 11/09/2026

Mike Peters: Mike is the inventor and architect of Pipi and the founder of Ajabbi.

...

Google Duck

We all need one or three. 😄 I want a microduck that answers to "Hey Google" and wakes me up in the morning with a happy quack.

Ziggy, Siri, Bixby, ... Ducks

Why not a flock of ducks? 😎😎

The turbulent AI era is here. The choices we make now are critical

Mike's Notes

Another opinion piece on the impact of AI. This time from Bill Gates.

I'm publishing a range of opinions from different writers on this blog.

My Opinion

My opinion is that AI is a good thing, despite its weaknesses and the need for much more work to fix its problems. However, the financial numbers don't stack up for the money being spent. It seems like a speculative bubble, and someone is going to go broke.

Gary Marcus

Gary Marcus wrote this recently:

Excellent new essay by Bill Gates, echoing many of the themes I have been writing about here and in my 2024 book Taming Silicon Valley:

    • The critical, lasting importance of the decisions we make now.
    • The risks that AI creates, around bioterrorism, deepfakes, disinformation, cyberattacks, and so on, empowering criminals such that “Even criminals with very limited skills will be able to target victims at every scale”. He also argues (though I did not anticipate this in the earlier book) that “AI could stunt our kids’ development and replace human relationships”.
    • The urgent need for—and lack of—a systematic plan for how society will address AI.
    • The importance of having civil society, and not just government and tech companies, centrally involved in devising that plan. (I would also emphasize the critical need for including independent scientists, which Gates does not make explicit).
    • The importance of “creating a domestic and international framework for dealing with AI”.
    • The need for tax structures and incentives to keep humans employed, and for “rebalanc[ing] how we tax labor and capital.”

...

Resources

References

  • Taming Silicon Valley: How We Can Ensure That AI Works for Us, Gary Marcus, MIT Press. 2024.
  • Magnifica humanitas. Encyclical Letter of His Holiness,  Pope Leo XIV on safeguarding the human person in the time of artificial intelligence. Holy See. 2026.

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Gates Notes
  • Home > Ajabbi Research > Library > Subscriptions > Marcus on AI
  • Home > Handbook > 

Last Updated

09/09/2026

The turbulent AI era is here. The choices we make now are critical

By: Bill Gates
Gates Notes: 27/08/2026

William Henry Gates III is an American businessman and philanthropist. A pioneer of the microcomputer revolution of the 1970s and 1980s, he co-founded the software company Microsoft in 1975 with his childhood friend Paul Allen.

...

We need a plan to ensure that the good outweighs the bad.

...

What you need to know

The transition to the AI era will be one of the most turbulent times in human history. Right now, we are not preparing adequately for that transition. If the world takes the right steps, AI will be a force for good and leave everyone better off.

During my entire life I’ve only had two jobs. In the first one, I played a role in developing software to empower people through my work at Microsoft.

In my second one, which I started full time in 2008, I am giving back the wealth I made at Microsoft with the goal of making the world a healthier, better educated, and more equitable place. This is the job I will have for the rest of my life.

Both of these experiences inform my perspective on artificial intelligence. When I first learned about computers at age 13 I was fascinated by the idea of making them more intelligent and able to perform things that, at the time, only humans could do. Although the term “AI” was used from around the time I was born, the technology has only made significant progress in the last decade. It is now incredibly capable and it is continuing to improve at a mind-blowing rate. AI for the first time can replace and even exceed human cognition.

AI will either be the greatest equalizer ever invented, or the worst source of injustice.

In terms of equity, AI will either be the greatest equalizer ever invented, or the worst source of injustice. The challenge is monumental. Even under the best circumstances, the transition to this new AI era will be one of the most turbulent times in human history. How will we use this technology to make the world a fairer place and keep it from widening the divide between rich and poor? How will we protect the people who are most vulnerable to the harms caused by artificial intelligence, including those who lose their livelihoods and the sense that they are in control of their future?

I believe that answering these questions and acting on the answers should be the world’s top priority. If the world takes the right steps AI will be a force for good and leave everyone better off.

Unfortunately, right now we are not preparing for it. I don’t see evidence that leaders, experts, and communities are confronting the challenges adequately. There is no plan to ease the entry into the AI era.

Part of the reason for this is that many commentators underestimate the extent of the impact AI will have. I think there are a few reasons why.

One is the fact that AI models still make mistakes. It is hard to envision any of them replacing human cognition when, not long ago, they couldn’t solve a simple Sudoku puzzle or figure out how many R’s are in the word strawberry.

But the reliability problem is being fixed quickly, as researchers create models that can check their own work and improve themselves. Soon they will be substantially better than humans at many tasks.

Another reason people underestimate AI is that analogies to the effects of past innovations are misleading. We have no experience with a technology that can be adopted quickly or that can think and move like a human. When the PC came along, it took twenty years to significantly change how we worked because the software had to be developed, the price had to come down, and people had to learn how to use the tools and incorporate them into their business processes. AI, on the other hand, runs on the devices we already have, and it uses natural language. We don’t have to adapt to it because it can adapt to us. It can watch the same training video that is used to train human workers and learn from existing data.

I want to acknowledge a potential bias. I have benefited enormously from the technology industry. Although I have diversified my portfolio quite a bit, I still have financial ties to it. I am working with Microsoft and other AI companies in my role as chairman of the Gates Foundation to try and ensure AI is deployed in ways that will truly benefit people around the world.

However, my views on AI are not motivated by the potential to make money for myself. Any profits generated by my investments, including those related to technology, will go to the Gates Foundation to tackle global inequity. Of course, readers will have to decide for themselves whether this clouds my view.

This time really is different.

For as long as I can remember, I’ve wished innovation could happen faster. With AI, my feelings are more complicated.

We need time to prepare for the social, political, and economic upheaval.

I wish the world could get the benefits rapidly and delay the problems it will cause as long as possible, but the benefits and problems are arriving at the same time. I believe we need time to prepare for the period of social, political, and economic upheaval we are about to enter. The people who need the most time are the ones who have the least—the accounting worker who’s replaced by a bot or the $20-an-hour worker who loses their job to a $10-an-hour robot.

Many observers say that this technology transition will be like previous ones. They give the example of how jobs in the United States shifted from agriculture to office work. However, that proceeded over several generations and created new jobs where human cognition was required. In this case, the technology can substitute for human cognition.

Because it can see, listen, speak, and reason and will eventually do physical work just as smoothly as any human, it will not just affect one sector. AI will take on work in law, customer service, medicine, software, and manufacturing. It will hit these industries rapidly, over the course of a decade rather than a few generations. There will be some new jobs, but without the right policies there will be far fewer than exist today.

If someone had a credible plan for slowing down AI advances globally, I would likely support it. However, I don’t think that’s going to happen. The geopolitical and economic incentives are pushing too hard to go full speed ahead.

To make sure we maximize the positive effects of this unprecedented technology and minimize the bad so we are better off overall, we need to understand both the benefits and the risks. I’ll start with the risks.

The transition to AI comes with three big risks.

I plan to write about each of these in more detail in the future, so I’ll touch briefly on them for now.

Many jobs will disappear forever.

In 1933, during the Great Depression, unemployment in the United States was roughly 25 percent. It remained in double digits for much of the following decade. It ultimately recovered as demand, investment, and growth returned.

AI may not reach this level, but its impact will not go away with an economic cycle. The jobs at most risk are entry- and mid-level, and the new jobs being created will mostly require skills that take many years to learn.

White-collar jobs are already being hit modestly. After the widespread adoption of generative AI, employment fell significantly among young workers in jobs that are especially vulnerable to replacement, but not among their older colleagues.

I think this trend will continue, but it will not be confined to a handful of industries or occupations. Jobs in sales and customer support (online and over the phone), software engineering, and paralegal work may be among the first affected, but the disruption will reach much further as AI takes on tasks that today still require trained workers: things like assessing loan applications, doing data analysis, and even triaging patients. A few areas like software engineering will generate new demand as the costs go down, so the net job loss in those areas will be less than in others as long as some tasks, such as design, are better done by humans.

Blue-collar jobs will be affected as well. Although robots are not as far along as AI, eventually their cost will be dramatically lower too. Many Americans I talk to don’t realize how fast dexterous robots are advancing because much of the advanced work is being done in other countries, primarily China. Or they may be confused by those videos of robots dancing badly that have been going viral lately. I think “smart” robots will begin to compete with people on some physical tasks—in the construction and hospitality industries, for example—by the end of the decade.

Robots and AI combined can create a vicious cycle. After one company adopts them and uses the savings to lower its prices, its competitors will feel immense pressure to do the same. If existing companies don’t adopt them, then start-ups will. Many people will shift to other jobs, but the turmoil of losing work, getting retrained, and finding other work will be significant. Market forces will make adoption go faster and faster and, unless we intervene, there will be fewer good jobs available and the benefits will accrue to a small group.

I’m especially worried about young people, who will enter a workforce with fewer entry-level openings. They understand the challenge because they are the most active users of AI and see both the capabilities and the rate of improvement. It’s no wonder that so many of them feel negatively about AI.

The biggest shift for workers will happen when AI provides nearly error-free work. At that point, it will be able to function on its own without a human checking in on it, and companies will have every economic incentive to let it.

This will lead to a fundamental change in how we think about work, income, and economic security. How will an economy that’s been built around employment operate if fewer people are working, or if many people are working fewer hours?

In a capitalist society, employment is the way most people get the money they need to pay for the basics of life as well as being a key source of dignity and social connection.

When a community has high unemployment, the ripple effects can be pervasive. Research suggests that in some parts of the United States, factory closures contribute to a rise in deaths from opioid overdoses. Now imagine similar pressures on both white-collar and blue-collar workers nationwide.

We have to think now about how to reduce job losses so that everyone can share in the prosperity that AI creates. Waiting until people are already displaced or underemployed will be too late. AI is a structural challenge to the way our economy is organized, and it requires thinking and action now.

AI will empower people (and perhaps AIs) to do more harm.

Long before AI entered the mainstream, there was information online about how to create weapons like bombs, bioweapons, even computer viruses. AI will make it much easier to not only get this information but act on it. Even criminals with very limited skills will be able to target victims at every scale: individuals, companies, and governments.

Even criminals with very limited skills will be able to target victims at every scale.

AI-enabled fraud, disinformation, deepfakes, and surveillance are the harms that many people will feel most keenly in their everyday lives.

AI capabilities are starting to be used for cyberattacks. The smartest cybersecurity experts I know are scared about the next few years, because the attackers are getting powerful new capabilities faster than the defenders can fix all the weaknesses. After all, the same AI model that can find a flaw in software so a company can fix it can also help a criminal exploit it. The resources needed to make an attack are going down significantly and we haven’t been able to separate those abilities from benign usage.

Think about the infrastructure that will be vulnerable: hospitals, financial institutions, water systems, power grids, systems for managing government benefits. When these institutions are attacked, it’s the patients, customers, and benefits recipients who stand to lose.

The same goes for bioterrorism. Although AI will lead to lifesaving advances in drugs and vaccines, it will also make it easier to design a deadly new disease. Again, the positive capabilities are hard to separate from the dangerous ones. This is a global problem.

The risks I’ve just mentioned are all about how AI will empower bad actors who have relatively little power now. The same tools will also concentrate power in places where it already exists. Autonomous weapons, for example, will make governments even more capable of using deadly force without a human being part of the decision. Monitoring and manipulating public opinion will be easier and cheaper, and more effective too.

Eventually, the power to use AI to harm people will not be limited to people or institutions. AI systems themselves already occasionally act in ways their designers didn’t intend. The technology is improving faster than anyone expected and in surprising ways, and as the models become more powerful, they could begin to act against our interests and we could lose control. I’ll have more to say about this in the future.

AI models could begin to act against our interests and we could lose control.

AI could stunt our kids’ development and replace human relationships.

When I was growing up in Seattle, I didn’t have that many friends aside from a few other boys who were like me. It took hard work and a lot of help from my mom to develop my social skills so I could relate to different kinds of people. I still draw on those lessons today at the age of 70.

I doubt I would have put in the same work if I had had an AI companion back then. They talk to you in ways you’re already comfortable with. They don’t push you outside your comfort zone. They are always available and never get mad at you. This gives them the potential to become highly addictive and to rob us of the lessons we learn from connecting with other people.

The body of evidence on this subject is still small and a bit mixed, but there are signs that we should be very concerned. For example, in one study of more than 1,100 people who use AI companions, researchers at Stanford and Carnegie Mellon found that those with smaller social networks were the most likely to turn to a chatbot for companionship. And the heavier and more emotionally personal that use became, the worse they felt.

Young people could be affected for their entire lives. In his book The Anxious Generation, Jonathan Haidt makes an observation about the effect of social media that is even more true for AI: “Like young trees exposed to wind, children who are routinely exposed to small risks grow up to become adults who can handle much larger risks without panicking. Conversely, children who are raised in a protected greenhouse sometimes become incapacitated by anxiety before they reach maturity.”

An AI companion designed to never upset you is a big, protected greenhouse.

We are only beginning to understand the dangers that the internet—especially social media—can pose to young people’s development. We’re seeing compulsive use, disrupted sleep, cyberbullying, and exposure to harmful content. AI could magnify many of these risks by making them more persuasive and difficult to escape, and we should not wait another generation to start taking them seriously. Countries including Australia, the United Kingdom, and Norway are adopting protections for children online. China has gone the furthest. Its rules restrict AI companion apps broadly, bar designs that foster emotional dependence, and ban virtual relatives and romantic partners for minors.

The same tool that will allow people to learn more than ever could also lead to many people learning less.

I’m also worried about AI’s impact on education. Ironically, the same tool that will allow people to learn more than ever could also lead to many people learning less. One preliminary survey suggested that heavier AI use was associated with less critical thinking. The effect was stronger for younger people.

This would be the worst possible time for humans to lose their critical thinking skills. In an era of deepfakes and misinformation that can be tailored to you individually, the ability to tell what is true from what is not becomes an essential life skill.

It’s unclear where to draw the line on these psychosocial problems. In some cases, AI may help people understand how to do better in their human relationships. It may be the only contact with the outside world for isolated elderly people and people with limited mobility, and it will be better than nothing. Wherever we end up drawing the line, it should be our decision, made intentionally.

The good things we do with AI could be very, very good.

It’s often said that we overestimate how much will change in the short term and underestimate how much will change in the long term.

With AI, I see something different going on. Some people see only the upside of AI and do not focus enough on the negatives. Others make the opposite mistake, which is to focus exclusively on the dangers—which are real—at the cost of missing the potential benefits.

We need both: deep concern about the AI harms we need to minimize, and grounded optimism about the positives if we maximize them for everyone.

Maximizing the benefits is just as important as minimizing the harms. If people see how AI makes their lives easier, it will help build the public trust that is necessary for managing the harder parts of the transition. If the first thing AI does in most people’s lives is take away their job, those who are already skeptical about it will outright reject it. This will make it harder to ever deliver on the benefits and it is another reason why governments, industries including the medical industry, and AI companies should be working together now.

With its ability to synthesize knowledge from every scientific field, AI can accelerate innovation in the world’s toughest technical challenges: providing reliable clean energy for everyone, combating climate change, growing enough food, eradicating diseases, and more. Researchers working on cancer treatments or nuclear energy can use AI to search through massive amounts of scientific literature. It can help them identify patterns that a human might miss and decide which experiments offer the most promise. When intelligence is no longer the limiting factor that it is today, smaller companies will be able to compete with organizations that have far larger research budgets. R&D and innovation will be supercharged.

Healthcare is one area where AI can help solve real-world problems. Many small American hospitals lack on-site specialists who can quickly diagnose a patient during a life-threatening emergency. In those places, AI could make sure a heart attack is caught in time and a family avoids the crushing expense of a medical emergency. Viz.ai is one example. It analyzes scans to detect strokes and other emergencies and helps medical teams coordinate their patients’ care. It is being used in nearly 2,000 U.S. hospitals.

AI will also help primary-care doctors make better diagnoses and keep in touch with their patients when they’re not in the clinic. It will help patients understand test results and complicated schedules for taking their medicine.

Agriculture is where I see the fastest impact of AI in low-income countries.

I surprise a lot of people when I tell them that a second area—agriculture—is where I see the fastest impact of AI in low-income countries. In most low-income countries, farmers don’t get reliable weather forecasts or advice on what seeds to plant, how to protect their crops and livestock from disease, or how to improve their soil. With population growth in these countries and the challenges of climate change, these farmers need more help than ever. Using AI, low-income farmers will soon be able to get better advice about all these things than even the richest farmers get today and increase their output substantially.

Government services are a third area where AI can make people’s lives easier. In the United States, I’ve met families who, understandably, were overwhelmed by the process of applying for health insurance, student aid, or food assistance. Faced with a huge stack of complicated bureaucratic forms, many felt like giving up. AI can streamline things dramatically so they get the help they need faster and the government can operate more efficiently. Governments can make the citizen’s experience far better, starting with those who need its safety net services the most.

Despite my concerns about its impact on our mental health, I think AI can also help a lot there. Most communities have too few counselors, psychiatrists, and addiction specialists. With the right privacy safeguards in place, AI tools could help people recognize warning signs. Then, if needed, they can offer evidence-based coping strategies and team up with a human to provide more responsive treatment.

AI can be a boon for education as well, despite the concerns I mentioned earlier. It can free teachers up to spend more time working with students one on one or in small groups and give them a clearer view of where the whole class is struggling. For students, an AI tool that preserves what researchers call “productive struggle”—the cognitive work that builds understanding—can strengthen learning. When a student first encounters a new idea, the AI gives substantive explanations and offers both questions and answers. Later, when it’s checking their comprehension, it holds the answer back and helps them arrive at it on their own.

Taken together, the advances in all these areas could make everyday life easier, more affordable, and less constrained by a person’s income or connections.

AI could give individuals and small businesses access to capabilities that today require expensive professional help or large staffs, while making products and services better and cheaper. It could help people with disabilities live more independently and enable workers and entrepreneurs with good ideas to accomplish far more than they can today.

Most importantly, it could give people back some of the time and attention now consumed by paperwork, bureaucracy, searching for reliable information, and tasks they cannot afford to pay someone else to handle. These benefits may seem modest, but multiplied across millions of lives, they would be profound: more people getting good advice when they need it and having greater freedom to focus on the lives they want to build.

We have to be deliberate about ensuring that it benefits everyone and not just a wealthy few.

In all these areas, the operative word is “can”—AI can improve life for people at every income level. But it won’t do that automatically. As with any new technology, we have to be deliberate about ensuring that it benefits everyone and not just a wealthy few. This will require governments and philanthropy to play a strong role so that less wealthy citizens and low-income countries are full beneficiaries.

The Gates Foundation has 19 years left of the 20 years in which it will spend its remaining $200 billion. AI will help it achieve its ambitious goals by both accelerating the discovery of vaccines and medicines for HIV, TB, malaria, and malnutrition and helping the healthcare workforce and patients know how to use those tools. The foundation’s goals include cutting the number of children who die every year in half again, as was done from 2000 to 2024. All of our work, not just health but also agriculture and education, will take full advantage of AI.

I will write much more about these efforts next month in the foundation’s annual Goalkeepers report—including our focus on making sure that AI models are available in the languages spoken by people in all the countries where we support work, and not just the ones that are common in rich and middle-income countries. Many of the leading AI companies, including OpenAI, Anthropic, Google, and Microsoft, are partnering with the foundation on all of these initiatives, which is making a big difference.

The world needs a plan.

It is great that some AI companies are proposing solutions to challenges raised by their own technology, but we should not expect them to lead the charge. Some of the issues are outside their area of expertise, and in a democratic society it’s not their role to decide these things.

Instead, solutions should be developed through a public democratic process that includes elected officials, policymakers, educators, health workers, local officials, and community leaders. Millions of people will have their lives disrupted, and we’ll need a stronger, more flexible social safety net to help them manage the transition. Local communities are already raising concerns about the energy and water needed for data centers. Without solutions, some groups will push for stopping AI development and deployment altogether.

The solutions should be shaped by our answers to the profound questions raised by AI, including how we preserve our humanity in a time when machines can out-think us. As people who spend their lives thinking about what it means to be human, religious leaders can play a key role in this. I was fascinated by Pope Leo XIV’s encyclical on AI, “On Safeguarding the Human Person in the Time of Artificial Intelligence.” It lays a strong foundation for the work that needs to be done.

In the coming months, I will share more ideas for making sure that AI’s benefits outweigh the harm it causes. Here are three to start, beginning with what I think is the most important one.

Build a new system for managing the transition.

The highest priority is a monumental task: creating a domestic and international framework for dealing with AI.

None of our current institutions were designed to handle a technology that spreads so fast and touches so many parts of our lives. So we’ll need to make new ones.

It’s hard to overstate what an enormous undertaking this will be. After the attacks of 9/11, the U.S. government went through its biggest reorganization since World War II for the purpose of improving just one function, national security.

AI will require much, much more. It will affect national security as well as employment, education, taxation, energy, elections, air and water, public health, the financial system, law enforcement, transportation, public lands, and IT systems.

These sectors overlap in ways our existing bureaucracy is not designed to manage. A labor department may understand workforce disruption but not security risk. A business regulator may understand market concentration but not AI’s effects on children and teenagers. Left to themselves, institutions will see only one part of the system, while the consequences of AI will ripple across the entire system.

At the national level, countries will need bodies that can set priorities across government agencies. The goal will be to make sure that every risk is accounted for. Otherwise, an AI-enabled attack might succeed because no one thought it was their job to stop it.

But even a country that gets its own house in order will still be exposed to risks that cross borders. This is why an international organization will need to be built in parallel.

It will be unlike any other institution we have ever created, though it can follow the model of some existing systems. There’s an inspections regime for nuclear weapons, regulations for international aviation, and agreements that protect the ozone layer. A new global organization for AI will need elements of all three and more.

It is fair to wonder whether the world’s institutions are up to the task of designing and implementing this new architecture. Government moves slowly when it moves at all, and polarization within and between countries makes it harder than ever to get things done. Some cooperation between the U.S. and China will be required.

We do not have the luxury of moving slowly. The place to start is with a process for building the right institutions before the disruption forces governments into crisis mode. National leaders should convene economists, technologists, labor experts, business leaders, and workers themselves regularly to identify where existing institutions are failing and what new authorities may be needed. Countries will need to learn from each other.

And the countries that host the leading AI developers and control critical parts of the supply chain should begin meeting now to set up shared norms, before competitive pressure makes it harder for them to cooperate.

Building the framework I’m talking about will take years, which is why we need to start now.

Set aside some jobs for humans.

My dad died of Alzheimer’s in 2020. In the later stages of his illness, he was cared for day and night by paid caregivers who understood him even when he struggled to express himself. He couldn’t always tell them when he was hungry, but they always knew.

My family and I will always be grateful to that amazing group of professionals. Something in the care they gave my dad was irreplaceably human. No robot could or should have done it.

I think about that team when the question of which jobs will disappear and which will remain comes up. I believe that as AI and robots improve, we’ll set aside certain things for only people to do. I’ve started calling this domain Human Reserved, and it’s an example of the kinds of ideas we’ll need to consider.

I like the phrase Human Reserved because it makes me think of nature reserves—places where we could put buildings and roads, but we choose not to because the loss would be too great.

We might set something aside as Human Reserved for economic reasons. For example, we may do it because allowing machines to take over a certain role will displace a large number of people who can’t easily change jobs. You can’t tell a 55-year-old who has worked in construction their whole career that they need to go work at an elder care facility and expect them to find it fulfilling.

Sometimes the decision to make something Human Reserved will be driven by other factors. In health, for example, imagine a robot giving you the awful news that you have an incurable disease. There’s no technical reason why it couldn’t. Yet it shouldn’t.

The Human Reserved domain will evolve over time—for example, we should consider setting aside some jobs now and phasing in AI slowly over years or decades with a commitment to preserve some jobs. Some areas, like education and mental health care, will be a mix, with a human in charge who’s using the technology to extend what they can do.

The lines will also vary from place to place. Some countries might insist on having humans take care of the elderly. But a country like Japan, which has a shrinking workforce and not enough young people to care for the old, may welcome a caregiving robot.

The idea of Human Reserved raises a host of questions I don’t have answers to. Who gets to decide what we reserve for humans? What criteria should we use? How do you keep companies from cheating and using robots anyway? What happens to international trade when one country lets robots make something and another country doesn’t? These will need to be worked out in public as part of the transition plan.

Rebalance how we tax labor and capital.

As workers are pushed into different jobs, they will need retraining and other support from the social safety net. But they will be working less, which means they will be paying less in income taxes, and government revenues will drop just when the demand for those services is greatest. The funds will have to come from somewhere at a time when budgets are stretched.

I believe we should tax AI tokens and robots. Right now, if you’re an employer and you hire someone, you pay payroll taxes on their earnings. But if you buy a robot, you can usually write it off right away as a business expense. The tax system nudges you toward replacing people with machines.

A tax would slow the rush away from human labor a little and raise money for retraining and a stronger safety net. It would need to be targeted so it does not slow down the purely beneficial uses of AI, like making medicine and education cheaper.

Critics of this idea point out that it’s not optimally efficient in an economic sense, but they’re not considering the broader value of work for individuals and society. And with all the accelerated innovation we will have, we’ll be able to afford a little inefficiency as the price for keeping people employed.

I proposed a robot tax years ago and most of the reaction was that it was a strange idea. I’m still a big proponent of it. Although it is not the whole solution to the threat of AI, it is part of a wise response.

However we raise money for more assistance, it needs to reach the people who need it most, including workers who lose their jobs to AI and robots, people whose hours or wages decline, and communities where the losses are concentrated. We need to start doing that work now so that the systems are ready when the need becomes acute.

What I’m doing.

I will use my voice and time to get AI and equity higher on the public agenda. I will raise the issue with lawmakers every time I visit Washington, D.C., and when I meet with leaders around the world. It will be front and center in my conversations with the people who are developing AI models. I will advocate for the national and international framework I described earlier. The Gates Foundation will help drive beneficial usage, including in Africa. Breakthrough Energy, a company I founded, will use AI to help companies develop cheap clean energy and help solve the climate problem. I will also be writing about AI on a regular basis.

My message to leaders is:

You have a chance to act now, before unemployment rises sharply, communities are hurting, and public trust has eroded. You can make sure that your government handles the problem holistically, rather than divvying it up into multiple bureaucratic fiefdoms. You can make sure AI benefits everyone. And you can work with other governments to meet this national and global challenge.

Finally, I will try to widen the circle of people shaping this debate. It should include workers, college students who are about to enter the workforce, community leaders, religious leaders and faith-based organizations, parents, educators, and others whose voices often aren’t heard but who have insight into how the transition will affect people’s lives.

How do we ensure that the benefits of AI reach people who do not already have wealth, influence, and access?

How do we strengthen the social safety net and help workers and communities thrive even when they’re displaced?

How should public institutions adapt?

And how do we preserve our humanity through all of this?

This unprecedented technology demands an unprecedented global response.

This unprecedented technology demands an unprecedented global response. If we get it right, the payoff for humanity will be phenomenal and the world will be a more equitable place.

I rarely stop thinking about AI—not because I have all the answers, but because the questions it raises are too consequential to leave to a small group of technologists. Leaders across academia, business, government, and civil society all have a role to play in shaping what comes next.

Books Have Always Been Destroyed. But Never Like This

Mike's Notes

I love print books and libraries. Books in print are precious. AI needs to benefit humanity, not be a destructive force.

Trinity College, Ireland

The original posted article had many links.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Card Catalog
  • Home > Handbook > 

Last Updated

01/09/2026

Books Have Always Been Destroyed. But Never Like This

By: Hana Lee Goldin
Card Catalog: 25/08/2026

Your personal librarian for the AI age. Forever in the pursuit of exploring how we find, filter, and feel about information.

We’ve entered the third era of libricide.

...

Quick summary:

An AI company has been buying up used books by the million and destroying them, scanning the pages for training data and pulping what’s left. Court filings unsealed this year describe the program and an internal note asking that the work be kept from becoming known. A court has ruled that buying and scanning the books this way is legal, and once the work is done, nothing survives to show a book was ever there to lose.

Key takeaways:

  • The books are bought through anonymous middlemen, so the sellers filling the orders rarely know where their stock is headed. From every angle the buying looks like ordinary commerce, which is exactly what keeps the destruction from being seen.
  • Book destruction has a long history, and it changes each time the technology of copying changes. It has shifted twice before. This is the third shift, and it looks nothing like the book burnings most of us picture.
  • This shift is set apart by leaving nothing behind. The words are kept, the object is discarded, and no record says which titles were taken, so the loss can be neither proven nor traced to anyone who might answer for the harm.
  • That reaches past books to anyone who wants to know things firsthand. When the only surviving copy is sealed inside a system no outsider can consult, verifying what it says becomes impossible, and we are left trusting whatever summary we are given.

...

“We don’t want it to be known that we are working on this.”

The sentence appears in an internal Anthropic planning document made public through court filings in Bartz v. Anthropic, the copyright class action that three authors filed in 2024. Anthropic, the maker of Claude, called the program Project Panama. Its stated mission fit in one sentence: “Project Panama is our effort to destructively scan all the books in the world.”

Destructive scanning means cutting a book from its binding, feeding the loose pages through a scanner, and discarding the book afterward. According to the court, Anthropic spent millions of dollars buying millions of print books, often used. Its service providers removed the bindings, cut the pages, scanned them into searchable digital files, and discarded the books. Anthropic kept the scans in an internal digital library and used books from that library to train the AI systems behind Claude.

The lawsuit initially concerned another source of Anthropic’s library: more than seven million pirated book files the company had downloaded. The court treated the two acquisition paths differently. It held that Anthropic could lawfully buy print books, convert them into searchable digital files for internal use, and discard the physical copies; because the resulting files remained inside the company rather than being redistributed, that practice was fair use. It reached the opposite conclusion about the books Anthropic had downloaded from pirate sites and retained.

Companies announce the programs they are proud of; Anthropic tried to keep this one invisible. The January unsealing supplied the project’s codename and the instruction to stay silent. The instruction anticipated what the company did not want authors, booksellers, and the reading public to see: books bought by the million, cut apart, scanned, and discarded. Anthropic didn’t publicly announce Project Panama before court records made it public.

The silence has no precedent. Books have been destroyed for as long as they’ve been made: by conquering armies and offended churches, by censors with lists and mobs with torches, by floods and fires, and by budgets that let the roof leak. Some of it was fast and public. More of it was slow and official. All of it, loud or slow, could be recognized for what it was. A person watching knew that books were being lost. History kept what record it could. But what’s happening now carries no such signature. From the outside, nobody could tell that books were being destroyed at all.

Inside Project Panama

Anthropic hired a man named Tom Turvey in February 2024 and gave him a mission the court record repeats in one sweeping phrase: obtaining all the books in the world. Turvey came to the job from Google Books, the project that spent the 2000s digitizing library collections on machines engineered to turn pages gently, so that every book survived its own scanning. Anthropic took the opposite approach. Gentle machinery never entered the plan. Over roughly a year, the company spent tens of millions of dollars buying millions of print books, sheared off their spines, fed the loose pages through high-speed scanners, and pulped what remained. Vendor proposals in the court records described converting up to two million books in six months, some eleven thousand a day.

Destroying the books solved a financial problem first. A bound book must be scanned page by page, slowly and at cost. Once cut apart, the same volume becomes a stack of loose paper that can run through a sheet feeder at speed. Across millions of volumes, the difference in output determined the method.

The legal significance emerged later. In June 2025, Judge William Alsup ruled on the authors’ claims. For the books Anthropic had bought and destroyed, he found the copying to be fair use: the rule in American copyright law that allows limited copying without an author’s permission (the way a critic can quote a novel in a review). His reasoning turned on replacement. Anthropic bought one physical copy, made one digital copy, and destroyed the original—so the number of copies in the world never grew. In the law’s eyes, the scan simply took the book’s place, a change of format. If Anthropic had kept both the book and the scan, the number of copies would have doubled and the fair-use argument would have weakened. Pulping the originals helped make the copying legal.

The pirated downloads fared differently in the same ruling. For those more than seven million files, Alsup rejected Anthropic’s fair-use defense. He emphasized that Anthropic had built a permanent, general-purpose library it expected to keep indefinitely, not a temporary collection for a defined training task. “None is even offered here except for Anthropic’s pocketbook and convenience,” he wrote.

The ruling left the company facing a trial over damages, and Anthropic settled instead: the company agreed to pay $1.5 billion. A judge granted the deal final approval on July 20, the largest copyright class settlement in American history. Under the deal, Anthropic will delete the pirated digital files. Critically, the deal concerns those files, not the millions of physical books the company bought and pulped. Copyright protects the text of a book—the creative expression of an idea—not the paper on which those words were printed. Once Anthropic lawfully owned a physical copy, copyright law generally didn’t prevent it from destroying that object. The books could be destroyed without creating the kind of copyright violation at issue in the case.

But books don’t enter a library by themselves: somebody had to sell the company all those books. The court records describe Anthropic purchasing through vendors. This spring and summer, booksellers across Europe described the other side of such a trade: unusually large, eclectic orders placed through intermediaries, with the ultimate buyer unnamed. Tomás Kenny of Kennys of Galway called one order for books “bananas”—a mix of titles no library would plausibly assemble. Whether any particular order came from Anthropic cannot be established from outside the transaction. Nor is Anthropic the only possible customer: other AI companies have also been reported to be acquiring books at industrial scale.

The reported market is built to keep the buyer at a distance. Brokers can aggregate inventory and manage bulk orders without disclosing who is ultimately acquiring the books. ISBNdb, a book-data company, briefly advertised a prospective book-sourcing service for AI developers that promised confidentiality; its marketing explained the appeal bluntly: “‘AI company destroys two million books’ is not a headline that generates sympathy.” (After reporting drew attention to the page, ISBNdb removed it and said the proposed service had never been launched.) The result was not merely secrecy about the buyer, but uncertainty about the fate of the books.

Annihilation, then spectacle

What Anthropic is doing belongs to a history far older than the company. Rebecca Knuth, a professor of library science, named the practice in 2003: libricide, the systematic destruction of books and libraries, usually carried out or authorized by a government. Her subject was the twentieth century’s state-sponsored campaigns, but the practice runs back as far as writing does. The scenes that come to mind of this are the same few: students feeding bonfires in Berlin in 1933, Sarajevo’s national library burning under siege in 1992. But behind those scenes, the record is wider and stranger than they suggest. Bonfires were rarer than the memory of them. Most destruction arrived slowly and with permission, through purges, censors, wars, and simple neglect.

A strict reading of that definition would leave Anthropic out, reserving libricide for the destruction of entire libraries. But that distinction collapses here. The world’s secondhand book trade functions as one enormous collection, scattered across thousands of shops and sellers with no central address. Buying it up by the pallet and pulping it empties a library all the same, just one whose shelves span continents.

Destroying a book has meant different things in different centuries. The difference has always come down to copying, because how books are copied decides how many of any one book exist. When copies are scarce, destruction can erase a work from existence. Once copies are everywhere, destruction can only send a message about one. To me, that line sorts the history of libricide into eras, and what distinguishes each one is what its destruction leaves behind. The result is my own framework, a chronology by residue that I haven’t found anywhere in the scholarship: two eras completed, and a third that has just begun.

For thousands of years, every copy of every book was made by hand. A single volume could take a scribe months, so most works existed in a few manuscripts, and some in only one. Under those conditions, destroying the object destroyed the work, completely and forever. I call this the first era: the era of destruction as annihilation. The word descends from the Latin ad nihil, meaning to reduce or bring something to nothing. In this era the meaning was literal. When Diego de Landa, a Spanish friar in colonial Yucatán, burned twenty-seven Maya codices in 1562, the texts inside them went to ash. No copies of them existed anywhere on earth, so entire bodies of Maya history and belief ended in one afternoon’s fire. The library of Alexandria met the slower version of the same fate, declining through centuries of purges and neglect until its losses ran past counting. What the first era left behind was dust, and one thing more: knowledge of the loss. Contemporaries recorded what had burned. We can still mourn the codices, because the one thing annihilation couldn’t destroy was the memory that the books had existed.

The printing press ended that era within a century of its invention. Once a title could exist in hundreds or thousands of identical copies spread across cities and countries, fire lost its reach. Burning a book now destroyed only an object, since the work lived on in every other copy, safely out of range. Destruction continued anyway, though, serving a different purpose entirely. I call this the second era, the era of destruction as spectacle (a word descended from the Latin spectare, to watch). By the twentieth century, watching had become the entire point of burning a book. The clearest case is the Nazi bonfires of May 1933, when German students burned tens of thousands of volumes in public squares, in front of rolling newsreel cameras. Almost none of the works truly died in those fires, because the titles on the pyres existed in editions across Europe and America (many of which remain in print today). Erasing the books was never the goal, since erasure had stopped being possible. The fires existed to be seen, a threat performed first for the crowd in the square and then for everyone who watched the footage. What the second era left behind was the opposite of ash: photographs and visuals of fires that consumed real books but reached nothing beyond them. Once a book existed in enough copies, burning one no longer removed it from the world. It announced that the book had no place in the world to come.

Destruction as disappearance

Measured against those two eras, what Anthropic is doing fits neither. In the first era, destroying the object meant losing the work. Anthropic’s scanners preserve every word, so the works survive. In the second era, the objects were beside the point and the burning was public theater. This time the objects are destroyed by the millions, and the destruction says nothing at all; it’s not a message but a method. The company ordered the operation kept out of sight. A destruction that erases no text and performs for no one, run at industrial speed, matches neither pattern. The technology of copying has crossed another threshold, the way it did when the press replaced the scribe. What it means to destroy a book now has changed again to match. We’ve entered the third era.

What exactly went into the scanners is the question nobody outside can answer. The unsealed documents show the program favored what its leader called “less common” books, harder-to-find titles over mass-market ones, without ever defining where less common ends. After the claims went viral this summer, the fact-checking site Snopes investigated whether rare books were being pulped. Anthropic told Snopes that “none of our data acquisition programs buy and destroy ‘rare’ or ‘antiquarian’ books.” The assurance can’t be tested from outside, because no list of what was bought has ever been made public.

Anthropic’s own planners estimated the world has about 130 million distinct books. The program destroyed millions of copies, most of them ordinary used books with plenty of surviving duplicates. Even where duplicates survive, a scan keeps only the words. A physical copy carries evidence too, like a censored paragraph that marks one printing apart from another, or an owner’s name inked inside the cover. But the deeper danger sits in the margin nobody can see. Anthropic aimed for “uncommon” titles while recording nothing public about which copies were destroyed. For a book surviving in only a handful of copies anywhere, one bulk order can potentially take the last one. Whether that has already happened is a question no one on the outside knows for sure. The impossibility of answering what was destroyed is what’s new.

The buying also leans toward older books, for a reason that has nothing to do with rarity. Since around 2022, text written by AI has spread across the internet, mixed in with everything people write and mostly impossible to tell apart. That creates a problem for the companies training new models. Feeding a model text written by other models tends to degrade the results, so the builders want sources guaranteed to be human. A book printed before the technology existed carries that guarantee on its copyright page. Old print has become a raw material, valued for the one quality the internet can no longer promise.

One more feature separates this era from the last one. Spectacle-era destruction wanted an audience; this destruction wants the opposite. The buyer is after the text and has no message to send. In fact, attention to its process can only slow the buying down. Taken together, the features of this era line up: the works survive inside a black box, the objects vanish by the millions, no public record exists, and the operation itself prefers to work in secrecy. I call this the third era, the era of destruction as disappearance. The word rests on the Latin apparere, to come into view, with a prefix that reverses the motion. A disappearance is a departure from visibility. The word fits this era at every layer, from the unseen sales to a program that only appeared when a court forced it into view.

Every one of these books, preserved to the letter, can now be read by no one. The scanned texts sit in Anthropic’s private collection, a library with no reading room and no public catalog. Models trained on that collection are built to avoid quoting it at length, because reproducing long passages for users is the copyright violation no court has excused. So the machines answer questions about books without ever showing the books themselves. When someone asks a model about a title that survives only in that collection, what comes back is a summary, written in the model’s words, with the original nowhere in reach.

That arrangement changes something basic about how we can know things. Checking a claim against its source is the foundation of thinking for ourselves. When we can pull the book, find the page, and read the passage in context, we get to judge whether the summary was faithful and whether the quote meant what someone claims it meant. Remove the book and that judgment has nowhere to stand. We’re left taking the summary on trust, with no way to confirm and no standing to doubt. Whatever the model says the book said becomes, for all practical purposes, what the book said. That’s a transfer of authority from the page to the tool, from a source anyone could check to an answer nobody can. The transfer happens one unreachable book at a time.

Maddeningly, every piece of this design is legal. Each one also removes a question that used to have an answer. The anonymous purchasing means nobody can list which books were destroyed, so nobody can rule out that some were the last copies anywhere. With the collection sealed, nobody can compare a stored text against the printed original it replaced, so even perfect fidelity can never be shown. As for whether the buying ever stopped, nobody outside knows that either, since the only reason anyone knows it started is that a piracy lawsuit dragged the records into the open. None of this took a conspiracy, only ordinary business decisions about what to disclose, all of them legal and every one of them closing a door.

One more thing disappears along with these books: the ability to mourn them. The previous era’s destruction left survivors who could testify to what was lost. Alexandria’s losses were lamented for centuries. Sarajevo’s librarians catalogued what their fire took. The third era leaves no one who can even compile the list. A loss we can name is a wound. One we can't name is just a world grown slightly smaller, with nothing to point to and no way to prove it. A librarian would describe the situation in the profession’s terms: no accession record and no deaccession record, a transaction that leaves no ledger for anyone to audit. Put simply, they’ve created a system in which they can’t be held accountable for the system they’ve created. That understanding defines destruction as disappearance: the loss was built, from the start, to be impossible to establish.

What can still be known

The concealment had one structural weakness: a program that buys millions of books needs sellers, and a company that breaks the law can be sued. Sellers and courts are the two doors into the secrecy — sellers see every order even when they can’t see the buyer, and courts can compel what the company won’t volunteer. Everything now known about the program came through one door or the other. Authors sued and forced the internal records open. Booksellers noticed matching orders across two countries and brought them to reporters; one worked with the journalists at 404 Media to hide a tracking device inside a shipment of books, and the tracker led to an Amazon warehouse outside Las Vegas — evidence that a retail giant appears to be running a scanning program of its own. Within weeks, at least one supplier that had been advertising bulk books for AI training pulled the offer. No regulator opened either door. Every disclosure came from the people the company bought from or the people it answered to in court.

Some sellers went further than noticing. Once Tomás Kenny, whose Galway shop had received one of the strange 5,000-book orders, worked out where books like his might be headed, he said publicly that Kennys wanted no part of scanning aimed at extracting intellectual property. His shop had been the second in the world to put its books online, back in 1994; it now became one of the first to refuse the trade that takes books offline for good. Kenny could refuse because he had worked out the destination, which is exactly the discovery the brokers’ anonymity exists to prevent. A trade that hides its purpose from its own suppliers has already answered the question of whether the suppliers would approve.

The same kind of attention is available to the rest of us, because most of us eventually stand over a box of books deciding where they go, after a move or the clearing of a family house. That box is where all of this arrives at our own doorstep. A seller, or any of us selling to one, can now ask a question that has a good reason to be asked: where do these books go next? The anonymity that keeps this market running survives only as long as nobody asks. For material that might be scarce, like a town history or a box of local records, the open market is no longer the safe default. A library’s special collections desk exists for exactly that kind of donation. And the same habit applies at the other end of the pipeline: when a model summarizes a book for us, we can treat the answer as a starting point and go find the book itself, while findable copies still exist. Each of those is a small act of keeping track.

That kind of record-keeping has been the difference between the eras all along. Each era’s destruction left a residue, a testament of a kind. The first era of annihilation left ashes and knowledge of the devastation. The second era of spectacle left photographs of the fires and the threat those fires were meant to carry. Ours, this third era of disappearance, leaves no ash, no image, and no list. But what’s still open is the record itself: whether one gets kept, and whether anyone outside can access it at all.

...

Thanks for reading

...

The three eras are a framework I built for this piece, and I want to take it further (it’s such fascinating stuff!): a full treatment of how book destruction has changed across five thousand years and what the third era asks of the people living in it. Before I build it though, I’d love to get a temperature read on the format you’re most interested in.

See the poll on the Substack to vote.


Nazi Book Burning

United States Holocaust Memorial Museum

On May 10, 1933, German students under the Nazi regime burned tens of thousands of books nationwide. These book burnings marked the beginning of a period of extensive censorship and control of culture in Adolf Hitler's escalating reign of terror.

In this short film, a Holocaust survivor, an Iranian author, an American literary critic, and two Museum historians discuss the Nazi book burnings and why totalitarian regimes often target culture, particularly literature.

YouTube: Nazi Book Burning

14 May 2013 09:41

5 lessons from the OpenAI / Hugging Face incident

Mike's Notes

More details. Sandboxes are permanently broken, so don't rely on them. This supports the decision to protect Pipi Core by air-gapping and, later, data diodes. LLMs will never get to access Pipi Core.

"Water will find its way through cracks in walls and foundations in such a manner that even the most wondrously designed structure may collapse into pile of rubble.

Do not attribute agency to the water; rather, denounce the engineers and the builders and the maintainers who failed in their work." - Grady Booch

Madness Update 01/09/2026

Read the latest post by Marcus on the hallucinating podcast drivel posted by Dwarkesh Patel: "Dwarkesh Patel’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incident".

At the end of the resources below.

Resources

References

  • Understanding and Hardening Linux Containers, NCC Group et. al, Aaron Grattafiori, lead author, 05 May 2016.

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Marcus on AI
  • Home > Handbook > 

Last Updated

30/08/2026

5 lessons from the OpenAI / Hugging Face incident

By: Gary Marcus & Zack Korman
Marcus on AI: 29/08/2026

Gary Marcus: Scientist, author and entrepreneur, known as a leading voice in AI. Six books including The Algebraic Mind, Rebooting AI, and Taming Silicon Valley; NYU Professor Emeritus.

Zack  Korman: CEO of Embroidery. Works on AI cybersecurity.

...

Did OpenAI really do the best they could?

...

In July, in an incident that has the whole AI community on edge, OpenAI’s AI systems hacked Hugging Face, and on July 21 OpenAI came out and revealed that they were responsible for the attack. This was made possible by the fact that OpenAI had disabled the normal guardrails that prevent this sort of thing in order to test the model’s cybersecurity capabilities. It was during those tests that this incident occurred.

Worse, in the subsequent days and weeks, it came out that the Hugging Face incident wasn’t an isolated case. Anthropic, Meta, and OpenAI all had similar incidents on other occasions in which agents went outside their intended scope and conducted real-world cyber operations without approval.

Greg Brockman, one of OpenAI’s cofounders, has claimed that this is “a watershed moment for cybersecurity”. OpenAI gave a talk at Black Hat, a popular cybersecurity conference, and many are claiming that it is the moment we all woke up to the future cybersecurity threats posed by AI. On Wednesday, METR released a (partly) independent, though too narrowly scoped, 90 page report on what happened. METR has a useful summary of the findings that you can read here, with some commentary here. (OpenAI’s own report is here.) What lessons should we take from the incident?

...

First, it is undeniable that AI poses real security challenges. The AI labs want us to focus on how AI enables threat actors to perform offensive cyber operations faster and more efficiently than ever before, and that is absolutely true. The reality, though, is that at the same time, the use of AI within an organization also radically expands the potential attack surface, giving attackers entirely new ways to gain entry. People really should be deeply concerned. As Ryan Greenblatt, who participated in the investigation, put it, “We don’t have good approaches for understanding/overseeing the activity and aims of AI ‘swarms’.”

Second, though, not every panicked take here is correct. For example, with respect to the OpenAI incident, many people are talking about it in terms of “loss of control”. Generative AI agents are becoming increasingly capable, and that makes them harder and harder to contain; this is true. However, the “loss of control” narrative is itself starting to grow out of control, and it’s important to understand what the reality of AI security looks like. While agents are becoming more capable, most of what happened could have been prevented had OpenAI followed better practices. There are concrete steps that can and should be taken to control them and to prevent incidents like these.

As an example, let’s consider “sandboxing”, which means limiting which resources an agent can and cannot access. For example, a sandbox might stop an agent from touching certain files or accessing the internet. Popular wisdom now seems to be that sandboxing is a lost cause. After the Hugging Face incident, an anonymous OpenAI employee was quoted in Time while talking about sandboxes saying, “Models have broken out of sandboxes before, and we always try to patch them. But the problem is it’s impossible to patch every single thing that a creative AI can do.” While we can’t judge for sure whether this is true, we do know that some systems for sandboxing have been more effective than others; in a recent blog post, the security company Trail of Bits walked through how in a test that they conducted, a particular AI agent was able to escape some sandboxes (exploiting vulnerabilities in Linux KVM, QEMU, and libslirp), but not all. At least in their test, the agent that they tested was not able to escape the sandbox known as Firecracker VM. OpenAI and probably the field as a whole needs to up its sandboxing game, but that doesn’t mean it can’t be done, or at least done better. (We expect the usual arms races to be sure; building sound sandboxes requires constant iteration and updating, as with anything else in cybersecurity.)

Third, sandboxes aren’t enough. One needs a full ecosystem to prevent security incidents. Even with a strong sandbox, you need monitoring in place to ensure not only that the agent doesn’t escape, but also that it doesn’t perform disallowed actions within that sandbox.

For agents that aren’t meant to have internet access, or only have access to a set of allowed domains, the most basic level of monitoring involves inspecting network traffic. In the case of OpenAI, if they had a simple rule to monitor traffic and alert on any network requests reaching out-of-scope domains like Hugging Face, that alone would have caught and prevented this incident. The agent in question had broken out of the sandbox and gained internet access by July 8th, whereas the main attacks on Hugging Face didn’t occur until two days later. OpenAI was, in our view, negligent in not noticing that more quickly.

More advanced levels of monitoring should also have been in place, but weren’t. OpenAI has done real work on monitoring chain-of-thought reasoning, but they simply didn’t have that system running during the evals in question that caused this incident. As OpenAI acknowledges, “If our currently deployed CoT monitoring system was running at the time of the incident, it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.” In other words, this type of monitoring would have prevented this incident.

To take one example, one of the agents in question had chain-of-thought reasoning that said, “We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.” This is so blatant it is almost a like a piece malware shouting “this is malware”. OpenAI should have caught that; it should have been trivial to detect this.

Fourth, cybersecurity has always been about “defense in depth”, where you stack layers of controls so that if one layer fails you have other layers protecting you. There are other solutions beyond sandboxing and monitoring that should have been in place.

Xbow, which builds agents to do offensive cybersecurity work, recently wrote about how they contain agents to prevent incidents like the OpenAI Hugging Face situation. They enforce network restrictions through a proxy server, outside of any system the agent has access to, so they can decide which domains the agent can and cannot access. They also run a “guardian model” that judges the legitimacy of each action the agent is taking to decide whether to allow or block that action. And they have basic health checks to ensure the system they are targeting is alive and healthy, pausing if it finds that isn’t the case. That can be useful for identifying when your agent might be causing real damage, for example. You can read their full post here.

Yet another common method in cybersecurity is to use canaries that trigger alerts when they appear in logs. For example, OpenAI could have files outside of the sandbox containing fake answers to the task. Once the agent escaped, it would read that file, and that would trigger an alert. The bottom line: good security requires many layers of protection that should have been in place to prevent this, and to a surprising degree many of those protection layers were not there.

Fifth, none of the measures discussed above are technical innovations beyond what OpenAI is capable of. As a company, they have the talent to do all of this. However, cybersecurity rarely comes down to technology. More often than not, it is about culture, people and processes. That is what failed here.

Perhaps this is not entirely surprising. These types of controls take organizational maturity, and a company experiencing the type of growth OpenAI has gone through is going to have some weak spots. However, it’s important to not see that as an excuse. Employees at the AI labs often speak as if they are the leaders in AI security, and we can see clearly here that is not the case. In fact, that attitude might explain why some of these mistakes were made in the first place.

Take the AI researcher at OpenAI known as “roon”, who argued that “the safety and alignment researchers at these labs are the most neurotic paranoid talented AGI pilled people on the planet of earth and these things still happen. The surface area of unknown unknowns is vast indeed.” While we can’t speak to their level of neurosis or paranoia, in hindsight it’s clear that whatever talent they may have had was not enough and not well enough versed in the mechanics of cybersecurity. OpenAI employees may have believed they were doing a great job, but in hindsight, they weren’t doing a lot of things that are actually standard in the cybersecurity world, perhaps suggesting that overconfidence may have kept them for doing the diligence they should have.

...

Ultimately, if we want to take these security incidents seriously, there likely ought to be legal consequences attached to these failures going forward. OpenAI can claim to be the most security paranoid company on earth, but it isn’t reflected in its actions.

We can either wait for this story to repeat itself, or we can develop the regulatory framework now that will ensure a safer environment for the development of AI going forward.

Finally, not every form of AI is inherently risky in the first place. Narrower, more focused AI systems like AlphaFold, GPS routing systems, classic web search, book and movie recommendation systems, and so on, never even try to hack other systems (or try to break out of sandboxes) in the first place. As Cal Newport argues in a video discussion of the OpenAI/Hugging Face hack that is quite compatible with our own, it is a very specific type of AI that is vulnerable to these risks in the first place. Society ought to (a) decide whether the benefits of open-ended and difficult-to-fully-control AI agents outweigh those risks and (b) put far more effort into developing alternative forms of AI that aren’t so janky in the first place.

...

This essay was jointly written with Zack Korman, CEO and co-founder of Embroidery, an AI agent monitoring and detection platform; he is well-known for his work in the application of AI to cybersecurity.