Ranked: Duolingo’s Most Popular Languages in Every Country in 2024

Mike's Notes

This could be a rough guide to what languages to initially target for translating Pipi UI and documentation. Pipi 9 has already been set up and tested to be multi-lingual and multi-script.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Visual Capitalist
  • Home > Learn > Reference > i18n

Last Updated

21/04/2025

Ranked: Duolingo’s Most Popular Languages in Every Country in 2024

By: Marcus Lu & Amy Kuo
Visual Capitalist: 6/04/2025

Key Takeaways

  • English is the number #1 ranked language on Duolingo in 134 different countries.
  • Other popular languages include Spanish (#1 in 33 countries) and French (#1 in 16 countries).

About half of the world’s population speaks at least two languages.

Learning a second language is a valuable endeavor, improving cognitive function, social awareness, and providing more career opportunities.

But what languages are people around the world trying to learn, and how does it differ from country to country?

We map the most popular languages studied on Duolingo from their 2024 language report.

The Most Popular Languages For Learning

English is the number #1 ranked language on Duolingo in 134 different countries. It is also the most prominent language on the planet, with approximately 1.5 billion speakers.

Browse the below table to see the top language and second-most popular language in every country Duolingo is available in.

Country Top Language on DuoLingo 2nd Language
๐Ÿ‡ฆ๐Ÿ‡ซ Afghanistan English German
๐Ÿ‡ฆ๐Ÿ‡ฑ Albania English German
๐Ÿ‡ฉ๐Ÿ‡ฟ Algeria English French
๐Ÿ‡ฆ๐Ÿ‡ฉ Andorra English Spanish
๐Ÿ‡ฆ๐Ÿ‡ด Angola English French
๐Ÿ‡ฆ๐Ÿ‡ท Argentina English Portuguese
๐Ÿ‡ฆ๐Ÿ‡ฒ Armenia English French
๐Ÿ‡ฆ๐Ÿ‡น Austria English German
๐Ÿ‡ฆ๐Ÿ‡ฟ Azerbaijan English Turkish
๐Ÿ‡ง๐Ÿ‡ญ Bahrain English French
๐Ÿ‡ง๐Ÿ‡ฉ Bangladesh English Korean
๐Ÿ‡ง๐Ÿ‡พ Belarus English German
๐Ÿ‡ง๐Ÿ‡ช Belgium English French
๐Ÿ‡ง๐Ÿ‡ฏ Benin English Spanish
๐Ÿ‡ง๐Ÿ‡ด Bolivia English Portuguese
๐Ÿ‡ง๐Ÿ‡ท Brazil English Spanish
๐Ÿ‡ง๐Ÿ‡ฌ Bulgaria English German
๐Ÿ‡ง๐Ÿ‡ซ Burkina Faso English French
๐Ÿ‡ง๐Ÿ‡ฎ Burundi English French
๐Ÿ‡จ๐Ÿ‡ป Cabo Verde English French
๐Ÿ‡ฐ๐Ÿ‡ญ Cambodia English Chinese
๐Ÿ‡จ๐Ÿ‡ฒ Cameroon English German
๐Ÿ‡จ๐Ÿ‡ซ Central African Republic English French
๐Ÿ‡น๐Ÿ‡ฉ Chad English French
๐Ÿ‡จ๐Ÿ‡ฑ Chile English Portuguese
๐Ÿ‡จ๐Ÿ‡ณ China English Japanese
๐Ÿ‡จ๐Ÿ‡ด Colombia English French
๐Ÿ‡ฐ๐Ÿ‡ฒ Comoros English French
๐Ÿ‡จ๐Ÿ‡ฌ Republic of the Congo English Spanish
๐Ÿ‡จ๐Ÿ‡ท Costa Rica English Spanish
๐Ÿ‡จ๐Ÿ‡ฎ Cรดte d'Ivoire English Spanish
๐Ÿ‡ญ๐Ÿ‡ท Croatia English Spanish
๐Ÿ‡จ๐Ÿ‡บ Cuba English French
๐Ÿ‡จ๐Ÿ‡พ Cyprus English Spanish
๐Ÿ‡จ๐Ÿ‡ฟ Czechia English German
๐Ÿ‡จ๐Ÿ‡ฉ Democratic Republic of the Congo English French
๐Ÿ‡ฉ๐Ÿ‡ฏ Djibouti English French
๐Ÿ‡ฉ๐Ÿ‡ด Dominican Republic English French
๐Ÿ‡ช๐Ÿ‡จ Ecuador English French
๐Ÿ‡ช๐Ÿ‡ฌ Egypt English French
๐Ÿ‡ธ๐Ÿ‡ป El Salvador English French
๐Ÿ‡ฌ๐Ÿ‡ถ Equatorial Guinea English French
๐Ÿ‡ช๐Ÿ‡ท Eritrea English French
๐Ÿ‡ช๐Ÿ‡ช Estonia English Spanish
๐Ÿ‡ช๐Ÿ‡น Ethiopia English French
๐Ÿ‡ซ๐Ÿ‡ท France English Spanish
๐Ÿ‡ฌ๐Ÿ‡ฆ Gabon English Spanish
๐Ÿ‡ฌ๐Ÿ‡ช Georgia English German
๐Ÿ‡ฉ๐Ÿ‡ช Germany English German
๐Ÿ‡ฌ๐Ÿ‡ท Greece English Spanish
๐Ÿ‡ฌ๐Ÿ‡น Guatemala English French
๐Ÿ‡ฌ๐Ÿ‡ณ Guinea English French
๐Ÿ‡ฌ๐Ÿ‡ผ Guinea-Bissau English French
๐Ÿ‡ญ๐Ÿ‡น Haiti English Spanish
๐Ÿ‡ญ๐Ÿ‡ณ Honduras English French
๐Ÿ‡ญ๐Ÿ‡บ Hungary English German
๐Ÿ‡ฎ๐Ÿ‡ณ India English Hindi
๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesia English Japanese
๐Ÿ‡ฎ๐Ÿ‡ท Iran English German
๐Ÿ‡ฎ๐Ÿ‡ถ Iraq English French
๐Ÿ‡ฎ๐Ÿ‡ฑ Israel English Hebrew
๐Ÿ‡ฎ๐Ÿ‡น Italy English Italian
๐Ÿ‡ฏ๐Ÿ‡ต Japan English Korean
๐Ÿ‡ฏ๐Ÿ‡ด Jordan English French
๐Ÿ‡ฐ๐Ÿ‡ฟ Kazakhstan English French
๐Ÿ‡ฐ๐Ÿ‡ฎ Kiribati English Spanish
๐Ÿ‡ฐ๐Ÿ‡ผ Kuwait English French
๐Ÿ‡ฐ๐Ÿ‡ฌ Kyrgyzstan English Russian
๐Ÿ‡ฑ๐Ÿ‡ฆ Laos English Chinese
๐Ÿ‡ฑ๐Ÿ‡ป Latvia English German
๐Ÿ‡ฑ๐Ÿ‡ง Lebanon English French
๐Ÿ‡ฑ๐Ÿ‡พ Libya English French
๐Ÿ‡ฑ๐Ÿ‡ฎ Liechtenstein English Spanish
๐Ÿ‡ฑ๐Ÿ‡น Lithuania English Spanish
๐Ÿ‡ฑ๐Ÿ‡บ Luxembourg English French
๐Ÿ‡ฒ๐Ÿ‡ฌ Madagascar English French
๐Ÿ‡ฒ๐Ÿ‡ผ Malawi English French
๐Ÿ‡ฒ๐Ÿ‡พ Malaysia English Japanese
๐Ÿ‡ฒ๐Ÿ‡ป Maldives English Spanish
๐Ÿ‡ฒ๐Ÿ‡ฑ Mali English French
๐Ÿ‡ฒ๐Ÿ‡น Malta English Spanish
๐Ÿ‡ฒ๐Ÿ‡ท Mauritania English French
๐Ÿ‡ฒ๐Ÿ‡ฝ Mexico English French
๐Ÿ‡ฒ๐Ÿ‡ฉ Moldova English German
๐Ÿ‡ฒ๐Ÿ‡จ Monaco English French
๐Ÿ‡ฒ๐Ÿ‡ณ Mongolia English Korean
๐Ÿ‡ฒ๐Ÿ‡ช Montenegro English Spanish
๐Ÿ‡ฒ๐Ÿ‡ฆ Morocco English French
๐Ÿ‡ฒ๐Ÿ‡ฟ Mozambique English French
๐Ÿ‡ฒ๐Ÿ‡ฒ Myanmar English Japanese
๐Ÿ‡ณ๐Ÿ‡ต Nepal English Japanese
๐Ÿ‡ณ๐Ÿ‡ฑ Netherlands English Spanish
๐Ÿ‡ณ๐Ÿ‡ฎ Nicaragua English French
๐Ÿ‡ณ๐Ÿ‡ช Niger English French
๐Ÿ‡ด๐Ÿ‡ฒ Oman English French
๐Ÿ‡ต๐Ÿ‡ฐ Pakistan English Arabic
๐Ÿ‡ต๐Ÿ‡ฆ Panama English French
๐Ÿ‡ต๐Ÿ‡พ Paraguay English Portuguese
๐Ÿ‡ต๐Ÿ‡ช Peru English Portuguese
๐Ÿ‡ต๐Ÿ‡ฑ Poland English Spanish
๐Ÿ‡ต๐Ÿ‡น Portugal English French
๐Ÿ‡ถ๐Ÿ‡ฆ Qatar English Spanish
๐Ÿ‡ท๐Ÿ‡ด Romania English Spanish
๐Ÿ‡ท๐Ÿ‡บ Russia English German
๐Ÿ‡ท๐Ÿ‡ผ Rwanda English French
๐Ÿ‡ธ๐Ÿ‡ฒ San Marino English Italian
๐Ÿ‡ธ๐Ÿ‡น Sรฃo Tomรฉ and Prรญncipe English French
๐Ÿ‡ธ๐Ÿ‡ฆ Saudi Arabia English French
๐Ÿ‡ธ๐Ÿ‡ณ Senegal English Spanish
๐Ÿ‡ธ๐Ÿ‡จ Seychelles English Spanish
๐Ÿ‡ธ๐Ÿ‡ฌ Singapore English Japanese
๐Ÿ‡ธ๐Ÿ‡ฐ Slovakia English German
๐Ÿ‡ธ๐Ÿ‡ด Somalia English Arabic
๐Ÿ‡ฐ๐Ÿ‡ท South Korea English Japanese
๐Ÿ‡ธ๐Ÿ‡ธ South Sudan English French
๐Ÿ‡ช๐Ÿ‡ธ Spain English Spanish
๐Ÿ‡ฑ๐Ÿ‡ฐ Sri Lanka English Japanese
๐Ÿ‡ธ๐Ÿ‡ฉ Sudan English French
๐Ÿ‡จ๐Ÿ‡ญ Switzerland English French
๐Ÿ‡ธ๐Ÿ‡พ Syria English German
๐Ÿ‡น๐Ÿ‡ฏ Tajikistan English Russian
๐Ÿ‡น๐Ÿ‡ญ Thailand English Japanese
๐Ÿ‡น๐Ÿ‡ฑ Timor-Leste English Portuguese
๐Ÿ‡น๐Ÿ‡ฌ Togo English German
๐Ÿ‡น๐Ÿ‡ณ Tunisia English French
๐Ÿ‡น๐Ÿ‡ท Tรผrkiye English German
๐Ÿ‡น๐Ÿ‡ฒ Turkmenistan English Russian
๐Ÿ‡ฆ๐Ÿ‡ช United Arab Emirates English French
๐Ÿ‡บ๐Ÿ‡ฆ Ukraine English German
๐Ÿ‡บ๐Ÿ‡พ Uruguay English Portuguese
๐Ÿ‡บ๐Ÿ‡ฟ Uzbekistan English Russian
๐Ÿ‡ป๐Ÿ‡ช Venezuela English Portuguese
๐Ÿ‡ป๐Ÿ‡ณ Vietnam English Chinese
๐Ÿ‡พ๐Ÿ‡ช Yemen English French
๐Ÿ‡ง๐Ÿ‡ผ Botswana French Spanish
๐Ÿ‡จ๐Ÿ‡ฆ Canada French Spanish
๐Ÿ‡ธ๐Ÿ‡ฟ Eswatini French Spanish
๐Ÿ‡ฌ๐Ÿ‡ฒ Gambia French English
๐Ÿ‡ฌ๐Ÿ‡ญ Ghana French Spanish
๐Ÿ‡ฐ๐Ÿ‡ช Kenya French Spanish
๐Ÿ‡ฑ๐Ÿ‡ธ Lesotho French Spanish
๐Ÿ‡ฑ๐Ÿ‡ท Liberia French English
๐Ÿ‡ฒ๐Ÿ‡บ Mauritius French English
๐Ÿ‡ณ๐Ÿ‡ฌ Nigeria French Spanish
๐Ÿ‡ธ๐Ÿ‡ฑ Sierra Leone French Spanish
๐Ÿ‡น๐Ÿ‡ฟ Tanzania French Swahili
๐Ÿ‡บ๐Ÿ‡ฌ Uganda French English
๐Ÿ‡ป๐Ÿ‡บ Vanuatu French English
๐Ÿ‡ฟ๐Ÿ‡ฒ Zambia French English
๐Ÿ‡ฟ๐Ÿ‡ผ Zimbabwe French Spanish
๐Ÿ‡ง๐Ÿ‡ฆ Bosnia and Herzegovina German English
๐Ÿ‡ฒ๐Ÿ‡ฐ North Macedonia German English
๐Ÿ‡ณ๐Ÿ‡ฆ Namibia German Spanish
๐Ÿ‡ท๐Ÿ‡ธ Serbia German English
๐Ÿ‡ธ๐Ÿ‡ฎ Slovenia German Spanish
๐Ÿ‡ป๐Ÿ‡ฆ Vatican City Italian English
๐Ÿ‡ง๐Ÿ‡น Bhutan Japanese Korean
๐Ÿ‡ง๐Ÿ‡ณ Brunei Japanese Chinese
๐Ÿ‡ต๐Ÿ‡ผ Palau Japanese English
๐Ÿ‡ต๐Ÿ‡ญ Philippines Japanese English
๐Ÿ‡ฆ๐Ÿ‡ฌ Antigua and Barbuda Spanish French
๐Ÿ‡ฆ๐Ÿ‡บ Australia Spanish French
๐Ÿ‡ง๐Ÿ‡ธ Bahamas Spanish French
๐Ÿ‡ง๐Ÿ‡ง Barbados Spanish French
๐Ÿ‡ง๐Ÿ‡ฟ Belize Spanish English
๐Ÿ‡ฉ๐Ÿ‡ฐ Denmark Spanish German
๐Ÿ‡ฉ๐Ÿ‡ฒ Dominica Spanish French
๐Ÿ‡ซ๐Ÿ‡ฏ Fiji Spanish French
๐Ÿ‡ซ๐Ÿ‡ฎ Finland Spanish Finnish
๐Ÿ‡ฌ๐Ÿ‡ง United Kingdom Spanish French
๐Ÿ‡ฌ๐Ÿ‡ฉ Grenada Spanish French
๐Ÿ‡ฌ๐Ÿ‡พ Guyana Spanish English
๐Ÿ‡ฎ๐Ÿ‡ธ Iceland Spanish English
๐Ÿ‡ฎ๐Ÿ‡ช Ireland Spanish Irish
๐Ÿ‡ฏ๐Ÿ‡ฒ Jamaica Spanish French
๐Ÿ‡ฒ๐Ÿ‡ญ Marshall Islands Spanish Japanese
๐Ÿ‡ซ๐Ÿ‡ฒ Micronesia Spanish Japanese
๐Ÿ‡ณ๐Ÿ‡ท Nauru Spanish Chinese
๐Ÿ‡ณ๐Ÿ‡ฟ New Zealand Spanish French
๐Ÿ‡ณ๐Ÿ‡ด Norway Spanish Norwegian
๐Ÿ‡ต๐Ÿ‡ฌ Papua New Guinea Spanish English
๐Ÿ‡ผ๐Ÿ‡ธ Samoa Spanish French
๐Ÿ‡ธ๐Ÿ‡ง Solomon Islands Spanish English
๐Ÿ‡ฟ๐Ÿ‡ฆ South Africa Spanish French
๐Ÿ‡ฐ๐Ÿ‡ณ St. Kitts and Nevis Spanish French
๐Ÿ‡ฑ๐Ÿ‡จ St. Lucia Spanish French
๐Ÿ‡ป๐Ÿ‡จ St. Vincent and the Grenadines Spanish French
๐Ÿ‡ธ๐Ÿ‡ท Suriname Spanish English
๐Ÿ‡ธ๐Ÿ‡ช Sweden Spanish Swedish
๐Ÿ‡น๐Ÿ‡ด Tonga Spanish French
๐Ÿ‡น๐Ÿ‡น Trinidad and Tobago Spanish French
๐Ÿ‡น๐Ÿ‡ป Tuvalu Spanish French
๐Ÿ‡บ๐Ÿ‡ธ U.S. Spanish English

Aside from English, other popular languages include Spanish (#1 in 33 countries) and French (#1 in 16 countries). But there’s clearly an overwhelming preference.

According to Duolingo research, the top reasons people learn English are to support their education, connect with others, and boost their careers.

English’s journey to become the dominant global language goes all the way back to British colonialism that spread it around the world. Back then, other languages—also part of empires—competed for lingua franca status, like French and Spanish.

Rank Language Countries Where
Language is #1
on Duolingo
1 ๐Ÿ‡ฌ๐Ÿ‡ง English 134
2 ๐Ÿ‡ช๐Ÿ‡ธ Spanish 33
3 ๐Ÿ‡ซ๐Ÿ‡ท French 16
4 ๐Ÿ‡ฉ๐Ÿ‡ช German 5
5 ๐Ÿ‡ฏ๐Ÿ‡ต Japanese 4
6 ๐Ÿ‡ฎ๐Ÿ‡น Italian 1

However, with the Industrial Revolution, English gained economic influence, and after the U.S. emerged as a global superpower, it became the preferred language for trade.

American cultural exports—Hollywood movies, music, clothing brands—also increased the language’s prominence.

What to Do

Mike's Notes

Notes on the worldview of Paul Graham.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

19/04/2025

What to Do

By: Paul Graham
paulgraham.com: March 2025

Paul Graham is a programmer, writer, and investor. In 1995, he and Robert Morris started Viaweb, the first software as a service company. Viaweb was acquired by Yahoo in 1998, where it became Yahoo Store. In 2001 he started publishing essays on paulgraham.com, which now gets around 25 million page views per year. In 2005 he and Jessica Livingston, Robert Morris, and Trevor Blackwell started Y Combinator, the first of a new type of startup incubator. Since 2005 Y Combinator has funded over 3000 startups, including Airbnb, Dropbox, Stripe, and Reddit. In 2019 he published a new Lisp dialect written in itself called Bel.

Paul is the author of On Lisp (Prentice Hall, 1993), ANSI Common Lisp (Prentice Hall, 1995), and Hackers & Painters (O'Reilly, 2004). He has an AB from Cornell and a PhD in Computer Science from Harvard, and studied painting at RISD and the Accademia di Belle Arti in Florence..

What should one do? That may seem a strange question, but it's not meaningless or unanswerable. It's the sort of question kids ask before they learn not to ask big questions. I only came across it myself in the process of investigating something else. But once I did, I thought I should at least try to answer it.

So what should one do? One should help people, and take care of the world. Those two are obvious. But is there anything else? When I ask that, the answer that pops up is Make good new things.

I can't prove that one should do this, any more than I can prove that one should help people or take care of the world. We're talking about first principles here. But I can explain why this principle makes sense. The most impressive thing humans can do is to think. It may be the most impressive thing that can be done. And the best kind of thinking, or more precisely the best proof that one has thought well, is to make good new things.

I mean new things in a very general sense. Newton's physics was a good new thing. Indeed, the first version of this principle was to have good new ideas. But that didn't seem general enough: it didn't include making art or music, for example, except insofar as they embody new ideas. And while they may embody new ideas, that's not all they embody, unless you stretch the word "idea" so uselessly thin that it includes everything that goes through your nervous system.

Even for ideas that one has consciously, though, I prefer the phrasing "make good new things." There are other ways to describe the best kind of thinking. To make discoveries, for example, or to understand something more deeply than others have. But how well do you understand something if you can't make a model of it, or write about it? Indeed, trying to express what you understand is not just a way to prove that you understand it, but a way to understand it better.

Another reason I like this phrasing is that it biases us toward creation. It causes us to prefer the kind of ideas that are naturally seen as making things rather than, say, making critical observations about things other people have made. Those are ideas too, and sometimes valuable ones, but it's easy to trick oneself into believing they're more valuable than they are. Criticism seems sophisticated, and making new things often seems awkward, especially at first; and yet it's precisely those first steps that are most rare and valuable.

Is newness essential? I think so. Obviously it's essential in science. If you copied a paper of someone else's and published it as your own, it would seem not merely unimpressive but dishonest. And it's similar in the arts. A copy of a good painting can be a pleasing thing, but it's not impressive in the way the original was. Which in turn implies it's not impressive to make the same thing over and over, however well; you're just copying yourself.

Note though that we're talking about a different kind of should with this principle. Taking care of people and the world are shoulds in the sense that they're one's duty, but making good new things is a should in the sense that this is how to live to one's full potential. Historically most rules about how to live have been a mix of both kinds of should, though usually with more of the former than the latter. [1]

For most of history the question "What should one do?" got much the same answer everywhere, whether you asked Cicero or Confucius. You should be wise, brave, honest, temperate, and just, uphold tradition, and serve the public interest. There was a long stretch where in some parts of the world the answer became "Serve God," but in practice it was still considered good to be wise, brave, honest, temperate, and just, uphold tradition, and serve the public interest. And indeed this recipe would have seemed right to most Victorians. But there's nothing in it about taking care of the world or making new things, and that's a bit worrying, because it seems like this question should be a timeless one. The answer shouldn't change much.

I'm not too worried that the traditional answers don't mention taking care of the world. Obviously people only started to care about that once it became clear we could ruin it. But how can making good new things be important if the traditional answers don't mention it?

The traditional answers were answers to a slightly different question. They were answers to the question of how to be, rather than what to do. The audience didn't have a lot of choice about what to do. The audience up till recent centuries was the landowning class, which was also the political class. They weren't choosing between doing physics and writing novels. Their work was foreordained: manage their estates, participate in politics, fight when necessary. It was ok to do certain other kinds of work in one's spare time, but ideally one didn't have any. Cicero's De Officiis is one of the great classical answers to the question of how to live, and in it he explicitly says that he wouldn't even be writing it if he hadn't been excluded from public life by recent political upheavals. [2]

There were of course people doing what we would now call "original work," and they were often admired for it, but they weren't seen as models. Archimedes knew that he was the first to prove that a sphere has 2/3 the volume of the smallest enclosing cylinder and was very pleased about it. But you don't find ancient writers urging their readers to emulate him. They regarded him more as a prodigy than a model.

Now many more of us can follow Archimedes's example and devote most of our attention to one kind of work. He turned out to be a model after all, along with a collection of other people that his contemporaries would have found it strange to treat as a distinct group, because the vein of people making new things ran at right angles to the social hierarchy.

What kinds of new things count? I'd rather leave that question to the makers of them. It would be a risky business to try to define any kind of threshold, because new kinds of work are often despised at first. Raymond Chandler was writing literal pulp fiction, and he's now recognized as one of the best writers of the twentieth century. Indeed this pattern is so common that you can use it as a recipe: if you're excited about some kind of work that's not considered prestigious and you can explain what everyone else is overlooking about it, then this is not merely a kind of work that's ok to do, but one to seek out.

The other reason I wouldn't want to define any thresholds is that we don't need them. The kind of people who make good new things don't need rules to keep them honest.

So there's my guess at a set of principles to live by: take care of people and the world, and make good new things. Different people will do these to varying degrees. There will presumably be lots who focus entirely on taking care of people. There will be a few who focus mostly on making new things. But even if you're one of those, you should at least make sure that the new things you make don't net harm people or the world. And if you go a step further and try to make things that help them, you may find you're ahead on the trade. You'll be more constrained in what you can make, but you'll make it with more energy.

On the other hand, if you make something amazing, you'll often be helping people or the world even if you didn't mean to. Newton was driven by curiosity and ambition, not by any practical effect his work might have, and yet the practical effect of his work has been enormous. And this seems the rule rather than the exception. So if you think you can make something amazing, you should probably just go ahead and do it.

Notes

[1] We could treat all three as the same kind of should by saying that it's one's duty to live well — for example by saying, as some Christians have, that it's one's duty to make the most of one's God-given gifts. But this seems one of those casuistries people invented to evade the stern requirements of religion: you could spend time studying math instead of praying or performing acts of charity because otherwise you were rejecting a gift God had given you. A useful casuistry no doubt, but we don't need it.

We could also combine the first two principles, since people are part of the world. Why should our species get special treatment? I won't try to justify this choice, but I'm skeptical that anyone who claims to think differently actually lives according to their principles.

[2] Confucius was also excluded from public life after ending up on the losing end of a power struggle, and presumably he too would not be so famous now if it hadn't been for this long stretch of enforced leisure.

Thanks to Trevor Blackwell, Jessica Livingston, and Robert Morris for reading drafts of this.

Kent Beck on Empirical Software Design: When & Why

Mike's Notes

An ACM Tech Talk interview yesterday with Kent Beck, author of Tidy First.

Resources

References

  • Reference

Repository

  • Home > Handbook > 

Last Updated

18/04/2025

Kent Beck on Empirical Software Design: When & Why

By: Kent Beck and Margaret-Anne Storey
ACM Tech Talk: 18/04/2025

Kent Beck is an American software engineer and the creator of Extreme Programming, a software development methodology that eschews rigid formal specification for a collaborative and iterative design process. Beck was one of the 17 original signatories of the Agile Manifesto.

Beck pioneered Test-Driven Development, its successor TCR: Test && Commit || Revert, software design patterns, and 3X: Explore/Expand/Extract. He wrote the SUnit unit testing framework for Smalltalk, which spawned the xUnit series of frameworks, notably JUnit for Java, which Beck wrote with Erich Gamma. Beck popularized CRC cards with Ward Cunningham, the inventor of the wiki.

Margaret-Anne Storey is a Professor of Computer Science and a Canada Research Chair in Human and Social Aspects of Software Engineering at the University of Victoria. Together with her students and collaborators, she seeks to understand how software tools, communication media, data visualizations, and social theories can be leveraged to improve how software engineers and knowledge workers explore, understand, analyze, and share complex information and knowledge. She has published widely on these topics and collaborates extensively with high-tech companies and non-profit organizations to ensure real-world applicability of her research contributions and tools.

Since the publication of Parnas' "On the Criteria to Be Used in Decomposing Systems into Modules" we have had good advice on how to design software. However, most software is more difficult to change than it should be and that friction compounds over time. The Empirical Design Project seeks to resolve the seemingly-irresolvable tradeoff between short-term feature progress and long-term optionality, focusing on:

  • How is software actually designed? What can we learn from data about how software is designed?
  • When should software design decisions be made? What is the optimal moment given unclear and changing information & priorities?
  • How can we enhance the survival of software projects while expanding optionality?

Spoiler alert: make design decisions later and in small, safe steps.

Other talks

Tidy First? A Daily Exercise in Empirical Design • Kent Beck • GOTO 2024

Wikipedia Structured Contents

Mike's Notes

Good news from Kaggle and Wikimedia. An opportunity to get structured data.

"...

As part of Wikimedia's mission to make all knowledge freely accessible and useful, Wikimedia is publishing a beta version of its structured content on Kaggle in French and English. This release gives data scientists, researchers, and machine learning enthusiasts a new, streamlined way to explore and analyze this global information resource.

..." - Kaggle.com

Resources

References

  • Reference

Repository

  • Home > 

Last Updated

18/04/2025

Wikipedia Structured Contents

By: Wikimedia Enterprise Team
Wikimedia Enterprises: 16/04/2025

Wikimedia Enterprise has released a new beta dataset on Kaggle, featuring structured Wikipedia content in English and French. Designed with machine learning workflows in mind, this dataset simplifies access to clean, pre-parsed article data that’s immediately usable for modeling, benchmarking, alignment, fine-tuning, and exploratory analysis.

This release is powered by our Snapshot API’s Structured Contents beta, which outputs Wikimedia project data in a developer-friendly, machine-readable format. Instead of scraping or parsing raw article text, Kaggle users can work directly with well-structured JSON representations of Wikipedia content—making this ideal for training models, building features, and testing NLP pipelines.The dataset upload, as of 15 April 2025, includes high-utility elements such as abstracts, short descriptions, infobox-style key-value data, image links, and clearly segmented article sections (excluding references and other non-prose elements). Because all content is derived from Wikipedia, it is freely licensed under Creative Commons Attribution-Share-Alike 4.0 and the GNU Free Documentation License (GFDL), with some additional cases where public domain or alternative licenses may apply.

“As the place the machine learning community comes for tools and tests, Kaggle is extremely excited to be the host for the Wikimedia Foundation’s data. Kaggle is already a top place people go to find datasets, and there are few open datasets that have more impact than those hosted by the Wikimedia Foundation. Kaggle is excited to play a role in keeping this data accessible, available and useful." - Brenda Flynn, Partnerships Lead, Kaggle

As a beta release, this dataset is an invitation to explore, test, and improve. We welcome feedback, questions, and suggestions from the Kaggle community directly in the dataset’s discussion tab.

Get the Dataset

Access the dataset directly on Kaggle

About Kaggle

Kaggle is home to one of the world’s largest communities of machine learning practitioners, researchers, and data enthusiasts. With millions of users and an expansive ecosystem of datasets, notebooks, and competitions—including challenges like the Arc Prize—Kaggle provides an ideal environment for experimenting with open structured data like Wikimedia’s Structured Content. Whether you’re testing a new architecture, evaluating data quality, or building a pipeline from scratch, this Wikipedia dataset is ready to plug into your process.

More info at Google Blog

Neobrutalism: Definition and Best Practices

Mike's Notes

This article by Hayat Sheikh from the NN Group's newsletter is relevant to the Ajabbi Design System. This Design System has a simple, clunky design style, like the web was 20 years ago and will be the default style of the Ajabbi workplace apps, and the Ajabbi website. The UI priority is nice, simple, reliable, fast, secure, and functional.

Users can easily change the style sheets to use their own design system.

Resources

References


Repository

  • Home > Design System >
  • Home > Ajabbi Research > Library > Subscriptions > The NN/g Newsletter

Last Updated

16/04/2025

Neobrutalism: Definition and Best Practices

By: Hayat Sheikh
The NN/g Newsletter: 11/04/2025

Hayat Sheikh, a Senior Designer at Nielsen Norman Group, is celebrated for her award-winning designs and extensive experience from renowned agencies. She also teaches at Lebanese American University, focusing on branding and human-centric design, and manages her NFT collection 'The Self.'

Summary:

As a UI design style, neobrutalism focuses on raw, unrefined elements like bold colors, simple shapes, and intentionally "unfinished" aesthetics.

Emerging as a reaction against sleek, minimalistic designs, neobrutalism creates a striking (almost rebellious) visual style. But while neobrutalism draws attention, designers must carefully balance its distinctive look with usability to avoid ending up with an overwhelming or confusing interface.

Defining Neobrutalism

Neobrutalist website design blends bold colors and sharp contrast for striking interfaces.

Neobrutalism (or neubrutalism), an evolution of traditional brutalism, is a visual-design trend defined by high contrast, blocky layouts, bold colors, thick borders, and “unpolished” elements.

Brutalism vs. Neobrutalism

Brutalism and neobrutalism are both edgy visual-design styles that draw inspiration from the architectural movements they get their names from. In digital design, brutalism tends to appear raw, harsh, unfinished, or utilitarian. Brutalist websites might use plain HTML elements and limited color palettes.

For example, Drudge Report embodies brutalist aesthetics with its barebones HTML structure, monospaced headlines, and rigid table-based layout, evoking the look of the pre-CSS web.


Drudge Report’s website embraces a brutalist style with its stripped-down aesthetic.

In contrast, neobrutalism combines the brutalist design style with nostalgic 90s graphic-design elements. Unlike true brutalist web design, neobrutalist designs are likely to be more colorful and orderly.

A striking example is Look Beyond Limits by Halo Lab, which features oversized typography, bold dividers, thick strokes with a pop of bright colors.

Look Beyond Limits by Halo Lab embraces neobrutalism with its raw, structured layout, and oversized typography.

Characteristics of Neobrutalism

High Contrast and Bright Colors

Neobrutalist designs use bold, primary colors and high-contrast combinations to emphasize key functions and UI elements. This approach introduces striking, contrasting hues to capture attention and enhance visual impact. It also helps users focus on essential elements while creating an unconventional, memorable experience.


99percentoffsale.com embraces neobrutalism through bold, high-contrast colors for a striking visual style.

Thick Lines and Geometric Shapes

This style does not shy away from using thick borders, angular forms, and solid lines that create structure without relying on gradients or shadows.


byooooob.com: Thick borders, solid lines, bright colors, and striking, playful shapes are typical for the neobrutalist visual style.

Stark Drop Shadows

Unlike minimalism, neobrutalism encourages bold, striking shadows instead of soft, layered ones. It incorporates solid, single-color shadows (e.g., a black drop shadow offset by 4px) to add depth while maintaining the "raw" aesthetic.


Unlike minimalist designs, which usually emphasize a flat, simple look and feel, neobrutalism often features bold, solid shadows that create depth while preserving a raw aesthetic.

Bold Type

Neobrutalism promotes the use of bold, “unpolished” elements that often include quirky or slightly eccentric typefaces. Despite their expressive forms, these typeface choices are balanced by a generous use of whitespace, creating a visual rhythm that feels deliberate rather than overwhelming. Typography in neobrutalism serves both as a functional element and as a focal point of the overall design.


Tony’s Chocolonely eCommerce by Tinloof uses bold, quirky typography that reinforces the brand’s personality.

Skeuomorphic Elements

Neobrutalism might incorporate nostalgic elements from early digital interfaces, such as Windows 98-style buttons and monospace fonts. These features create a sense of familiarity while blending retro aesthetics with modern design. For example, a neobrutalist design might use UI elements from an old browser, with traditional iconic buttons and appearance mimicking early web experiences.

cyanbanister.com: Neobrutalism blends retro UI elements (such as Pixel art and old-style  browser windows) with modern design elements (such as contemporary typography and layout).

Examples of Neobrutalism in Practice

Many brands are embracing the bold, raw aesthetic of neobrutalism to create memorable experiences through striking contrasts, unconventional typography, and minimalistic design. This approach reflects a shift toward prioritizing purpose and functionality over excessive polish, allowing brands to stand out in a crowded digital landscape.

Brands like Figma and Gumroad incorporated bold, high-contrast colors and raw elements, with a focus on user experience and simplicity.

Figma's brand refresh, with its use of bold contrasts and unconventional typography, exemplifies neobrutalist design. Just like its tools, the refreshed identity emphasizes creative freedom, flexibility, and a dynamic user experience, allowing users to work in ways that feel authentic and engaging.


Figma’s bold, geometric design reflects creative freedom.

Similarly, Gumroad, an ecommerce platform for independent creators, uses neobrutalism's raw aesthetic to align with its ethos of empowering independent creators. By stripping away unnecessary polish and focusing on functionality over flourish, the platform emphasizes simplicity and accessibility, staying true to its purpose of providing creative freedom and a straightforward user experience.


Gumroad’s raw design empowers creators with simplicity.

Designing with Neobrutalism: Best Practices

While neobrutalism thrives on bold colors, heavy typography, and sharp contrasts, without balance, it can overwhelm users and hinder accessibility. These tips help create designs that are both visually striking and user-friendly.

Design with Usability at the Forefront

Prioritize usability with clear buttons, readable type, and ample whitespace to keep the experience intuitive and accessible, even within a bold, raw aesthetic.

The API World landing page Incorporates a neobrutalist aesthetic while maintaining usability through its clear search functionality and calls to action.

Contrast Ratios Matter

Bold colors must meet text-contrast standards. Avoid pairing vibrant hues like yellow and cyan that fail readability tests. Tools like Coolors' contrast checker ensure that combinations remain accessible while staying visually striking.


Although neobrutalism uses bright, contrasting colors, it still needs to meet readability and accessibility standards.

Limit Your Color Palette

Restrict your palette to 2–3 bold, high-contrast colors (e.g., black, neon green, electric blue) to avoid overwhelming users.


bieffeforniture.it uses 2 main high-contrast colors (electric blue and red) to help maintain clarity and avoid overwhelming users.

Prioritize Readability

Pair bold, unconventional headlines (e.g., a chunky sans-serif font) with clean, neutral body fonts like Roboto or Inter. Avoid overly decorative or condensed typefaces for paragraphs to maintain legibility across devices.

dodonut.com adopts a neobrutalist style while still maintaining clear buttons, readable text, and ample whitespace.

Use Whitespace Strategically

Offset dense geometric shapes and thick borders with generous padding (e.g., 24–32px margins) to create breathing room, prevent clutter, and guide users to key actions or content.


Content in a neobrutalist layout needs enough padding to create space and focus users’ attention on key elements.

Test Interactions

Ensure that interactive elements (buttons, links) remain recognizable. Use underlines on hover or subtle color shifts to indicate state changes. For example, a neon button could lighten on click to signal feedback without gradients or shadows.

Interactive elements in a neobrutalist layout need clear feedback. Use underlines or color shifts to signal interaction.

Avoid Oversimplification

Retain hierarchy through size variation (e.g., headlines twice as large as body text) and color intensity. Even in a minimalistic layout, ensure that clear calls to action and key usability elements stand out to create a seamless user interface.

sui.io/overflow#overview maintains visual hierarchy by using different font sizes for headers, page text, and button labels, thus ensuring that CTAs and interactive elements stand out.

Key Takeaways

Neobrutalism’s rebellious aesthetic can grab attention, but its success hinges on balancing boldness with usability. By grounding the style in accessibility principles and testing with users, designers can create interfaces that are both striking and functional.

Shadow Table Strategy for Seamless Service Extractions and Data Migrations

Mike's Notes

Here is an InfoQ article by Apoorv Mittal & Rafal Gancarz, referenced in Data Engineering Weekly. It covers a valuable way to migrate data while keeping critical production going. It is something to use in the future.

Resources

References


Repository

  • Home > Ajabbi Research > Library > Subscriptions >Data Engineering Weekly

Last Updated

14/04/2025

Shadow Table Strategy for Seamless Service Extractions and Data Migrations

By: Apoorv Mittal & Rafal Gancarz
InfoQ: 09/04/2025

Apoorv Mittal is a passionate Software Engineer at Block (CashApp) based in Seattle, WA. With over a decade of experience in distributed systems and fintech, he has led transformative projects at Block, Dropbox, and AWS. His expertise spans modernizing legacy systems into scalable microservices, architecting resilient financial infrastructures, and pioneering cloud security innovations, notably through his patented AWS Traffic Mirroring solution. You can find Apoorv on LinkedIn.

Key Takeaways

  • The shadow table strategy creates a synchronized duplicate of the data that keeps the production system fully operational during changes, enabling zero-downtime migrations.
  • Database triggers or change data capture frameworks actively replicate every change from the original system to the shadow table, ensuring data integrity.
  • The shadow table strategy supports diverse scenarios - including database migrations, microservices extractions, and incremental schema refactoring - that update live systems safely and progressively.
  • Shadow tables deliver stronger consistency and simplify recovery compared to dual-writes or blue-green deployments.
  • Industry case studies from GitHub, Shopify, and Uber demonstrate that the shadow table approach drives robust large-scale data migrations by actively maintaining continuous data integrity and offering rollback-friendly safeguards.

Introduction

Modern software systems often need to evolve without disrupting users. When you split a monolith into microservices or modify a database schema, you must migrate data with minimal downtime and risk. Shadow tables have emerged as a powerful strategy to achieve this. In a nutshell, the shadow table approach creates a duplicate of the data (a shadow version) and keeps it in sync with the original, allowing a smooth switchover once the new setup is ready.

This article explores how shadow tables help in different migration scenarios — database migrations, service extractions, and schema changes — while referencing real case studies and comparing this approach to alternatives like dual-writes, blue-green deployments, and event replay mechanisms.

What is the Shadow Table Strategy?

The shadow table strategy maintains a parallel copy of data in a new location (the "shadow" table or database) that mirrors the original system’s current state. The core idea is to feed data changes to the shadow in real time, so that by the end of the migration, the shadow data store is a complete, up-to-date clone of the original. At that point, you can seamlessly switch to the shadow copy as the primary source. In practice, implementing a shadow table migration typically follows a pattern:

  1. Create a Shadow Table: Prepare a new table (or database) with the desired schema or location. Although initially empty, you structure it to accommodate the migrated data.
  2. Backfill Initial Data: Copy existing records from the original data store into the shadow table, processing them in chunks to avoid overloading the system.
  3. Sync Ongoing Changes: As the system runs, apply every new write or update from the original data to the shadow. Use database triggers, change data capture (CDC) events, or application-level logic to propagate each INSERT, UPDATE, or DELETE from the source to the shadow to remain in sync.
  4. Verification: Optionally, run checks, such as comparing row counts or sample records, to confirm that the shadow’s data matches the source, giving you confidence that no data was missed.
  5. Cutover: Point the application to the shadow table (or perform a table rename/swapping in the database) once you verify it is up to date. The switch occurs with negligible downtime because you have kept the shadow current.
  6. Cleanup: Retire the old data store after cutover or keep it in read-only mode as a backup until you no longer need it. By using this approach, you can complete migrations with zero downtime. The production system continues running during the backfill and sync phases because reads and writes still hit the original data store while you build the shadow. When you are ready, you can quickly switch to the new store, often through a simple metadata update like a table rename or configuration change.

Figure 1: Data migration using the shadow table strategy

This strategy is sometimes also called the ghost table method (notably by GitHub’s schema migration tool gh-ost) because the new table is like a "ghost" of the original (gh-ost: GitHub's online schema migration tool for MySQL - The GitHub Blog).

Use Cases Where Shadow Tables Shine

Shadow tables offer a robust and flexible mechanism for managing complex migrations, service extractions, and schema refactorings while keeping production systems running uninterrupted. There are three common scenarios where shadow tables can be especially beneficial: database migrations with zero downtime, service extractions in a microservices transition, and incremental schema changes with data model refactoring.

Database Migrations with Zero Downtime

Modern applications often rely on large, heavily used production databases that cannot afford extended downtime for schema modifications or engine migrations. Direct alterations — like adding a new column, changing data types, or indexing — can cause long locking periods and stall critical operations. Shadow tables provide an alternative approach that minimizes the risk of disruption.

Begin by creating a new table that mirrors the structure of the production table while incorporating the desired schema changes. Although this shadow table starts empty or partially populated, you fill it using a controlled backfill procedure. A robust backfill procedure copies historical data from the production table into the shadow table in controlled batches, allowing the system to run concurrently.

After the backfill, set up a continuous synchronization mechanism by leveraging database triggers or CDC frameworks that propagate every new insertion, update, or deletion from the production table to the shadow table. This dual-write mechanism ensures that the shadow table remains an up-to-date replica of the production system.

Simultaneously, automated verification processes continuously compare key metrics between the two tables. Checksums, row counts, and deep object comparisons confirm data integrity and ensure that the shadow table accurately mirrors the production data. Only once these validations confirm that the shadow is consistent with the source can the final cutover be executed, often through a fast, atomic table rename or pointer switch. This approach enables the migration to be completed with minimal downtime, reducing risk and preserving user experience.

Service Extractions in a Microservices Transition

Transitioning from a monolithic architecture to a microservices-based system requires more than just rewriting code; you often must carefully migrate data associated with specific services. Extracting a service from a monolith risks inaccuracy if you do not transfer its dependent data accurately and consistently. Here, shadow tables play a crucial role in decoupling and migrating a subset of data without disrupting the existing system.

In a typical service extraction, the legacy system continues to handle all live operations while developers build a new microservice to handle a specific functionality. During extraction, engineers mirror the data relevant to the new service into a dedicated shadow database. Whether implemented through triggers or event-based replication, the dual-write mechanism ensures that the system simultaneously records every change made in the legacy system in the shadow database.

Once the new microservice processes data from the shadow database, engineers perform parallel validation to ensure that its outputs match expectations. A comparison framework automatically checks that the outputs of the new service match the expected results derived from the legacy system. This side-by-side validation allows engineers to identify discrepancies in real time and make adjustments as necessary.

Teams carefully manage the gradual transition of traffic from the legacy system to the new microservice. By initially routing only a small portion of user requests to the new service, teams can monitor performance, validate data consistency, and ensure that the new system behaves as expected.

Once the shadow database and the new microservice have proven to maintain the same level of data integrity and functionality as the legacy system, engineers execute a controlled, incremental cutover. Over time, they shift all operations to the new service and gradually reduce the legacy system’s role until they fully decommission it for that functionality. This phased approach mitigates risk and provides a built-in rollback mechanism if they detect any issues during the transition.

Incremental Schema Changes and Data Model Refactoring

Even for smaller-scale changes, such as refactoring a table or updating a data model, shadow tables offer a powerful way to mitigate risk. In many systems, evolving the data model is an ongoing challenge, whether splitting a single table into multiple logical parts, merging fields, or adding non-null constraints to previously optional columns.

Instead of applying changes directly to the live table, engineers create a shadow version to reflect the new design. The system simultaneously writes data to both the original and shadow tables, ensuring that it captures any update in real time across both structures. This dual-writing approach allows continuous validation of the new schema against the existing one, enabling engineers to compare outcomes and ensure that the refactored data model handles all business logic correctly. 

Automated comparison tools play an essential role during this phase. By continuously monitoring and comparing data between the old and new schemas, the tools can detect discrepancies early — whether they arise from differences in data type conversions, rounding issues, or unforeseen edge cases. Once engineers have thoroughly validated the shadow table and adjusted for anomalies, they can seamlessly switch the application to the new schema. They can then gradually phase out the original table, with the shadow table taking over as the primary data store.

This incremental approach to schema changes minimizes the need for extended maintenance windows and reduces the risk of data loss or service interruptions. It provides a controlled path to evolve the data model while maintaining full operational continuity.

Industry Examples and Best Practices

Successful migrations using shadow tables have been reported by many organizations, forming a set of best practices:

  • Online Schema Change Tools: Companies like GitHub and Facebook built tools (gh-ost and OSC) to perform online schema changes using shadow/ghost tables. These tools have become open-source solutions that others use. MySQL migrations now use the standard procedure of creating a shadow table, syncing changes, and then renaming (Zero downtime MySQL schema migrations for 400M row table). Similarly, used the open-source LHM gem in their Rails applications to safely add columns, as it "uses the shadow-table mechanism to ensure minimal downtime" Shopify (Safely Adding NOT NULL Columns to Your Database Tables - Shopify). The best practice here is to automate the shadow table process with rigorous checks (row counts, replication lag monitoring, etc.) and fallback paths if something goes wrong (for example, aborting the migration leaves the original table untouched, which is safer than a half-completed direct ALTER).
  • Strangler Pattern for Microservices: Combining the strangler fig pattern with shadow reads/writes has proven to be a successful approach for migrating from a monolith. Amazon, Netflix, and others have used the idea of routing a portion of traffic to a new system in shadow mode to build confidence. Over time, they shifted reads and finally writes to the new service, effectively strangling out the old component. Best practice here:  migrate in phases (e.g., shadow/dual-run, verify, then cutover) and use monitoring/metrics to ensure the accuracy of the new service’s data. The shadow phase can catch any discrepancies early, avoiding faulty migrations.
  • Data Pipeline and CDC Usage: When using event streams for migration, you must ensure ordering and idempotency. Teams often choose Kafka or similar durable logs to replay events to the shadow database. The order of events must match the source’s commit order to maintain consistency. Industry best practice recommends schema versioning and backward-compatible change events when using this method, so that the new system can process events even if the schema evolves during the migration. Decoupling the pipeline (so that the old and new systems communicate via the event log rather than direct dual writes) also reduces risk to the production load. However, teams should monitor the lag between source and shadow and have a way to reconcile differences if the pipeline falls behind.
  • Fallback and Rollback Plans: A migration is not truly safe without a rollback plan. In many cases, shadow table strategies lend themselves to easy rollback. If you find a problem during verification, simply discard the shadow table before switching over; this will not impact users. Even after a cutover, if the new system/table misbehaves, switch back to the old one (provided you kept it intact for a while). Uber’s migration post-mortem stresses having the ability to reverse traffic back to the old system if needed (Uber’s Billion Trips Migration Setup with Zero Downtime). As a best practice, keep the old system running in read-only mode for a short period after cutover, just in case you need to fall back. This safety net, combined with thorough monitoring, makes the migration resilient.

Comparing Shadow Tables to Alternative Migration Approaches

While shadow table (or shadow database) migrations are powerful, you should choose the right strategy for your situation.

Shadow Tables vs. Dual-Write Approach

Shadow table strategy often uses triggers or external pipelines to sync data, whereas a pure dual-write approach relies on the application to perform multiple writes. Dual-writing can achieve a similar goal of keeping two systems in sync, but the complexity of distributed transactions comes with it. 

Without careful design, dual writes can lead to race conditions or partial failures – for example, the app writes to the new database but crashes before writing to the old one, leaving data out of sync. To mitigate this, developers use patterns like the Outbox Pattern, where the application writes changes to the primary DB and also to a special outbox table in the same transaction; the application then asynchronously publishes these changes to the second system. 

In contrast, a trigger-based shadow table inherently ties the two writes into the source database’s transaction (the trigger runs inside the commit), and a CDC-based approach will capture the exact committed changes from the log. Such an approach often makes shadow table strategies more reliable for consistency than ad-hoc dual-write logic.

Figure 2: Data migration using dual-write approach

However,  when you control both systems, dual writes may be simpler to implement at the application level, and they avoid the need for database-level fiddling or extra tooling. In summary, dual writes give you more control in application code, but you must exercise extreme care to avoid inconsistency. In contrast, shadow table methods leverage the database or pipeline to guarantee consistency.

Shadow Tables vs. Blue-Green Deployments

Shadow table strategy complements blue-green setups: one can see the shadow table as part of the green environment being prepared. The key difference is that blue-green by itself doesn’t specify how to keep the data in sync – it assumes you have a way to copy and refresh data in the green environment. A full outage could do this (not ideal), or a shadow/copy process could. So, in many cases, shadow table migrations are an enabling technique to achieve a blue-green style cutover for databases.

Figure 3: Blue-green deployments working with shadow tables

The ability to test the entire new stack in parallel is the advantage of a blue-green deployment. For example, you might run a new version of your service against the shadow database (green) while the old version runs against the old database (blue). You can then switch over when ready, and even switch back if something fails, since the blue environment is still intact. The downside is cost and complexity: temporarily doubling your infrastructure. 

Maintaining two full environments (including databases) and keeping them in sync is not trivial. Shadow tables ease this by focusing on the data layer sync. If your migration is purely at the database layer (e.g., moving to a new database server or engine), a shadow table approach is a blue-green deployment of the database. If your migration also involves application changes, you might do a blue-green deployment of the app in tandem with the shadow table migration of the data.

Both strategies share the goal of a zero-downtime switch, and they pair well, but blue-green is a broader concept encompassing more than data. In contrast, the shadow table strategy is laser-focused on data consistency during the transition.

Shadow Tables vs. Event Replay (Rebuilding from Event Logs)

Event replay leverages an event log or sequence of change events to build up the state in a new system. It’s related to the CDC but slightly different in intent. In a replay scenario, you might start a brand new service by consuming a backlog of historical events (for example, reprocessing a Kafka topic of all transactions for the past year) to reconstruct its database state. Alternatively, if your system is event-sourced (storing an append-only log of changes), you can initialize a new read model or database by replaying all events from the start. This approach ensures that the new database’s state is equivalent to that of the old system, which is derived from the same sequence of inputs.

Figure 4: Data migration using event replay

Unlike shadow tables, event replay can be more time-consuming and is usually done offline or in a staging environment first because processing a considerable history of events can take a while. Shadow table migrations tend to operate on live data in real time, whereas you might use replay to bootstrap and then switch to a live sync method (like CDC) for the tail end. Another difference is that event replay might capture business-level events rather than low-level row changes. 

For example, instead of copying rows from a SQL table, you might replay a stream of "OrderPlaced" and "OrderShipped" events to rebuild the state. This approach can be useful if you’re also transforming the data model in the new system (since the new system can interpret events differently). However, if you miss any events or the event log isn’t a perfect record, you risk an incomplete migration.

In practice, engineers often use event replay in combination with shadow strategies: one might do an initial event replay to catch up a new system, then use incremental CDC or dual-writes to capture any new events that occur during the replay (so the shadow doesn’t fall behind). The combination yields the same outcome: a fully synced shadow ready to take over. The choice between using database-level shadow copy versus event-level replay often comes down to what data you have available. 

Replay might be straightforward if you have a clean event log (like an append-only journal). Otherwise, tapping into the database (via triggers or log capture) might be more manageable. Both approaches aim for eventual consistency, but shadow table syncing (especially trigger-based) will typically have the new store up-to-date within seconds of the original, whereas an event replay might apply changes in batches and catch up after some delay.

Conclusion

The shadow table strategy has proven effective in performing complex data migrations safely and incrementally. Teams keep a live replica of data changes; this enables them to migrate databases, extract services, or refactor schemas without halting the application. Companies apply this pattern to add columns without downtime, migrate enormous tables, or gradually siphon traffic to new microservices, all while preserving data integrity.

Of course, no single approach fits all situations. Shadow tables shine when you need up-to-the-second synchronization and confidence through parallel run comparisons. Alternatives like dual-writes or event replay might be more appropriate in systems built around event messaging or in simpler scenarios where a full shadow copy is overkill. Many real-world migrations end up using a blend of these techniques. For example, one might do an initial bulk load (replay), then switch to a live shadow sync, or use dual-writes in the app plus a trigger-based audit to double-check consistency.

It’s essential that software engineering teams plan migrations as first-class projects and leverage industry best practices: they should run systems in shadow mode to validate behavior, keep toggles or backstops for quick rollback, and monitor everything. When executed with discipline, the shadow table strategy provides a moderate complexity path to achieve significant changes with little downtime. It enables the evolutionary changes that modern software demands, all while keeping users blissfully unaware that anything changed under the hood.