James Simons: My Guiding Principles

Mike's Notes

I'm a fan of Jim Simons. A brilliant mathematician, he made a fortune and then donated billions to support mathematics and open scientific research through the Simons Foundation, which he and his wife Marilyn established. Marilyn was the President, and Jim was the Chair. 

All the research results are made freely available. The foundation publishes Quanta, which is free and better than Nature. I often reproduce Quanta articles on this website.

Jim died last year.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

17/05/2025

    James Simons: My Guiding Principles

    By: James H. Simons
    Simons Foundation: January 22, 2020

    The chair of the Simons Foundation describes his five principles for building a successful organization.

    The foundation began in 1994, and in its early years, it had no particular mission and was simply Marilyn’s and my vehicle for distributing charity. In 2003, after the foundation had grown considerably (in part due to principle 5!), Marilyn and I became interested in autism — its cause and possible methods of alleviating its symptoms. We didn’t know how to begin this effort, but in the summer of that year, a friend of ours agreed to convene a roundtable of outstanding neuroscientists as well as some people who were already working in the field. We learned a great deal from this: first, that autism is very largely genetic; and, second, that few great scientists were working in the field.

    Our mission seemed clear: Attract great scientists, and begin with genetics. This decision was consistent with principle 1, since no one to our knowledge was taking this combined approach. A second decision was to bring Mike Wigler of Cold Spring Harbor Lab, a friend of mine and one of the world’s greatest geneticists, into the game. We were off and running.

    Another decision Marilyn and I made that year was to focus the foundation almost exclusively on mathematics and science research. Again, this was consistent with principle 1, as very few U.S. foundations had such a focus, and we felt we could really make a difference.

    After a few years of my overseeing the autism project, Marilyn and I realized that we needed an outstanding scientist to head the project, and we were lucky enough to hire Gerry Fischbach, the first scientist on the foundation staff. Not only was Gerry terrific, but he was indeed fun to work with. Together with Mike Wigler and a small staff, he created the Simons Simplex Collection. This was a cohort of almost 3,000 families that included both parents, one and only one child with autism, and at least one unaffected sibling. A great deal of effort was expended in designing and implementing this collection, and it has led to an enormous amount of great science. I consider this collection a beautiful thing — a fine example of principle 3.

    The Simons Foundation Autism Research Initiative (SFARI) has now gone on for 16 years. When Gerry stepped down in 2013, Louis Reichardt, an outstanding scientist and leader, became its head. Our perseverance has borne fruit. Not only have we discovered numerous genetic causes of the condition, but our first drug trials have been initiated. Our patience to some extent exemplifies principle 4.

    The next area we focused on was math and physical sciences (MPS), and at Marilyn’s suggestion, we brought on David Eisenbud to head that effort. David’s first step was to invite to the foundation a group of mathematicians, a group of theoretical physicists and a group of computer scientists to tell us what the foundation could do to advance each of their respective fields. The math and physics groups proposed various grant programs, which were subsequently put in place, but the computer science group wanted only one thing: an institute for theoretical computer science. There was no such institute in the world.

    A competition was established among leading U.S. research universities, and after several winnowing rounds, Berkeley was selected. Not only was their proposed leader, Richard Karp, an outstanding and renowned computer scientist, but Berkeley also agreed to give us the entirety of a small building that they would renovate for our purposes. This institute is now in its sixth year and has been an outstanding success, attracting multitudes of visitors and being run in a beautiful manner. The creation of this institute exemplifies principles 1, 2 and 3.

    Guiding Principles

    IN LATE 2010, just after I stepped down from Renaissance and started full-time at the foundation, I gave a talk at MIT. Marilyn accompanied me to the talk, and on the way there, she suggested that at the end I discuss my values. I felt ‘values’ was not quite the right word, so I used ‘guiding principles’ instead. They are listed below, and I believe they have been useful in my life and careers. After setting them forth, I will give a number of examples of their effect in building the foundation.

    1. DO SOMETHING NEW; DON’T RUN WITH THE PACK. I am not such a fast runner. If I am one of N people all working on the same problem, there is very little chance I will win. If I can think of a new problem in a new area, that will give me a chance.
    2. SURROUND YOURSELF WITH THE SMARTEST PEOPLE YOU CAN FIND. When you see such a person, do all you can to get them on board. That extends your reach, and terrific people are usually fun to work with.
    3. BE GUIDED BY BEAUTY. This is obviously true in doing mathematics or writing poetry, but it is also true in fashioning an organization that is running extremely well and accomplishing its mission with excellence.
    4. DON’T GIVE UP EASILY. Some things take much longer than one initially expects. If the goal is worth achieving, just stick with it.
    5. HOPE FOR GOOD LUCK!

    After a few years, Yuri Tschinkel succeeded David as head of MPS. Yuri is an excellent mathematician and a pleasure to work with. Under his leadership, a number of excellent grant-making programs have been initiated, as well as several collaborations (discussed below).

    The last grant-making area to be established was life science, headed by Marian Carlson, a member of the National Academy of Sciences and a totally fun person to work with. She too has initiated some innovative programs and collaborations (discussed below).

    As the programs grew, particularly SFARI, there was increasing need for IT personnel. This problem was solved by bringing on board Alex Lash, a very senior IT type at Memorial Sloan Kettering. I learned the hospital was quite unhappy to lose him, but Marilyn and I were delighted to hire him, and he gradually built an excellent team of almost 30 people who serve both the science and administrative sides of the house.

    In 2012, we came up with a new approach to grant-making, what we called collaborations. These would be goal-driven efforts involving a fairly large number of investigators from various institutions around the world, lasting as long as 10 years, or perhaps even more. To determine whether this was a good idea, we convened a weekend meeting of distinguished scientists from a broad set of fields. By the end of the meeting, we had decided this was indeed a good idea, provided the goal was extremely important, there was a reasonable chance of achieving it, and the leadership was outstanding.

    The first of these collaborations was Origins of Life, which was quickly followed by the Global Brain, an effort to understand the dynamics of the brain as a whole. Today we have 18 active collaborations covering life science, math and physical sciences. By and large these are going very well. This program is a clear example of principle 1, as we know of no other foundation or government organization that funds such efforts.

    In the course of the weekend meeting on collaborations, it was suggested that we create an institute devoted to data science. I liked this idea very much, since the wealth I had amassed was based on data science in the financial world, but instead of creating an institute on a university campus, I thought to assemble such an activity in-house. Marilyn was in full agreement. We were fortunate to recruit Leslie Greengard to head this effort, which became known as the Simons Center for Data Analysis (SCDA). Leslie not only is an outstanding applied mathematician, being a member of both the National Academy of Sciences and the National Academy of Engineering, but also holds an M.D. from Yale. He focused SCDA on computational biology and built up an excellent team. I consider this another example of principle 1, as nothing quite like this seemed to exist.

    SCDA worked so well that Marilyn and I decided to generalize the idea and create the Flatiron Institute for computational science. For this, of course, more space was required, and happily it existed right across the street! We recruited David Spergel from Princeton to build Computational Astrophysics, and then Antoine Georges from the Collège de France to build Computational Quantum Physics (working closely with Andy Millis). Finally, we determined that the fourth unit be Computational Mathematics, and that Leslie step down from Computational Biology to head this new unit. After a fairly long search, we decided to promote Mike Shelley, a group leader in the biology unit, to be its director. Underpinning this variegated research effort is the Scientific Computing Core, a team brilliantly headed by Nick Carriero and Ian Fisk.

    The Flatiron Institute has run magnificently in every way, its staff producing outstanding research and convening meetings and workshops that inspire other scientists to do the same. It clearly exemplifies principles 1, 2 and 3: it has originality and great leadership and is truly a thing of beauty.

    According to its bylaws, the Simons Foundation is intended to focus almost entirely on research in mathematics and science and to exist in perpetuity. If future leadership abides by these guiding principles, Marilyn and I believe the foundation will forever be a force for good in our society.

    Enshittification, tarpits and other things your mother never told you about

    Mike's Notes

    Here are some valuable resources about stuff you may have to deal with one day.

    • Enshittification
    • Tarpits
    I am so disgusted with the degeneration of previously useful websites provided by the world's largest software companies that I took these measures to ensure that Pipi would never end up like that.
    • No Investors
    • No social media
    • No sales, only word-of-mouth
    • No ads
    • No moats
    • Open-source as much as possible
    • Ownership by a foundation
    • Only SaaS applications that are socially useful
    • Built for experienced developer teams
    • Open Handbook
    AI companies are ignoring robots.txt files when they go out and scrape websites. That sucks.

    Resources

    References

    • Reference

    Repository

    • Home > Ajabbi Research > Library >
    • Home > Handbook > 

    Last Updated

    17/05/2025

    Enshittification, tarpits and other things your mother never told you about

    By: Mike Peters
    On a Sandy Beach: 13/02/2025

    Mike is the inventor and architect of Pipi and the founder of Ajabbi.

      Enshittification

      "Enshittification, also known as crapification and platform decay, is the term used to describe the pattern in which online products and services decline in quality over time. Initially, vendors create high-quality offerings to attract users, then they degrade those offerings to better serve business customers, and finally degrade their services to users and business customers to maximize profits for shareholders." - Wikipedia.

      Examples

      • Google Search
      • Facebook
      • Amazon

      Tarpit

      "A tarpit is a service on a computer system (usually a server) that purposely delays incoming connections. The technique was developed as a defense against a computer worm, and the idea is that network abuses such as spamming or broad scanning are less effective, and therefore less attractive, if they take too long. The concept is analogous with a tar pit, in which animals can get bogged down and slowly sink under the surface, like in a swamp." - Wikipedia

      Examples

      • Nepenthes
      • Iocane

      The creative process behind Pipi

      Mike's Notes

      I discovered this superb, filmed interview with author J.K. Rowling about her writing process. It's worth watching how this talented individual consciously uses a creative process in her work.

      Besides being in awe of her writing ability, I was also thinking about my creative process.

      I'm a pattern thinker. That's visual thinking.

      My creative process is similar to J.K. Rowling's, except I use images, relationships, 3D and 4D models, and patterns instead of words. I, too, have many ideas pouring out of my head like a firehose, just not in words.

      I, too, keep a notebook beside my bed, so when I wake up at 2 a.m., I can write a word or draw a diagram.

      I keep my drawings in A4 3-hole ring binders. There must be almost 100 binders now, and cartons of drawings are waiting to be filed away correctly. I use colour highlighters a lot. Everything is colour-coded like this website.

      I also have to use assistive technology to write. Every word on this website has been rewritten 30 to 40 times by me and corrected by Grammarly. I'm a slow writer; writing a page takes a whole day.

      I find articles written by others I agree with and republish them on this engineering blog, expressing my thoughts in their words and adding a few notes to explain the connection.

      I also have multiple kinds of synesthesia, which Wikipedia defines as " a perceptual phenomenon in which stimulation of one sensory or cognitive pathway leads to involuntary experiences in a second sensory or cognitive pathway."

      It was the same when I was a sculptor, making docos, building sets, doing landscape projects, and everything else, where I didn't need to write words. Reading and talking are easy. I have given hundreds of talks, usually technical, at conferences and workshops, and been on TV, radio, press interviews, etc.

      But I could never describe this process in written words. Now I can. Watch J K Rowling describe it.

      I also copied this post on my filmmaker website so that people could understand how I work.

      Resources

      References

      • Reference

      Repository

      • Home > Ajabbi Research > Library >
      • Home > Handbook > 

      Last Updated

      17/05/2025

        J K Rowling on Writing

        By: J K Rowling
        Jkrowling.com: 12/02/2025

        J.K. Rowling talks in depth for the first time about her writing

        J.K. Rowling is often asked questions by fans and budding writers about her writing process: where she writes, how she writes, her inspiration and her research, how a book comes about, from the germ of an idea to the editing process and eventual publication.

        Here for the first time, she responds to those questions, talking openly and in depth about her writing including Harry Potter, her other children’s books The Ickabog and The Christmas Pig, as well as writing as Robert Galbraith, the Cormoran Strike crime fiction series.

        Filmed in her writing room in Edinburgh and in a London pub, these three On Writing films provide a personal insight into J.K. Rowling’s writing world.

        Part 1

        Part 2

        Part 3

        Pipi self-hosts and the chicken-or-egg problem

        Mike's Notes

        I wanted to write about why and how Pipi 9 is self-hosting and why this created a chicken-or-egg problem.

        Pipi 9 is still largely headless, and I need to use the UI to build a UI.

        Resources

        References

        • Reference

        Repository

        • Home > Ajabbi Research > Library >
        • Home > Handbook > 

        Last Updated

        17/05/2025

        Pipi self-hosts and the chicken or egg problem

        By: Mike Peters
        On a Sandy Beach: 11/02/2025

        Mike is the inventor and architect of Pipi and the founder of Ajabbi.

        According to Wikipedia;

        "An operating system is self-hosted when the toolchain to build the operating system runs on that same operating system. For example, Windows can be built on a computer running Windows.

        Before a system can become self-hosted, another system is needed to develop it until it reaches a stage where self-hosting is possible. When developing for a new computer or operating system, a system to run the development software is needed, but development software used to write and build the operating system is also necessary. This is called a bootstrapping problem or, more generically, a chicken or the egg dilemma.

        A solution to this problem is the cross compiler (or cross assembler when working with assembly language). A cross compiler allows source code on one platform to be compiled for a different machine or operating system, making it possible to create an operating system for a machine for which a self-hosting compiler does not yet exist. Once written, software can be deployed to the target system using means such as an EPROM, floppy diskette, flash memory (such as a USB thumb drive), or JTAG device. This is similar to the method used to write software for gaming consoles or for handheld devices like cellular phones or tablets, which do not host their own development tools.

        Once the system is mature enough to compile its own code, the cross-development dependency ends. At this point, an operating system is said to be self-hosted." - Wikipedia

        Part of the secret sauce of Pipi 9's success is its self-generating nature.

        It's like watching the workings of a living biological cell, where hundreds of processes maintain and interact with each other. 

        I had to manually create each engine and let them run against each other. It was a very slow and experimental process, like inventing a new cake recipe through trial and error. I stumbled across something that worked by accident.

        As a result, it is headless. It has also been designed to generate its own no-code front-end, an HTML user interface for humans.

        The current problem is that a no-code interface is needed to create the no-code interface. The lower layers of Pipi are supposed to help generate the UI and learn from user interactions, so a lot needs to be done, making it challenging to simulate.

        This means the front end will go through many generations as Pipi learns from feedback loops to improve the UI further. The backend will also self-evolve.

        So, it is slow, but progress is steady. And the more I get done, the easier and faster it will get.

        So, back to that darn chicken, or is it the egg?

        The half-life of code & the ship of Theseus

        Mike's Notes

        I found this via the excellent engineering newsletter from PostHog. It's an article by Erik Bernhardsson, originally published on his blog in 2016. He has many thoughtful articles.

        The article has a very enticing name.

        So, what is the half-life of Pipi?

        Resources

        References

        • Reference

        Repository

        • Home > Ajabbi Research > Library >
        • Home > Handbook > 

        Last Updated

        17/05/2025

          The half-life of code & the ship of Theseus

          By: Erik Bernhardsson
          erikbern.com: 5/12/2016

          As a project evolves, does the new code just add on top of the old code? Or does it replace the old code slowly over time? In order to understand this, I built a little thing to analyze Git projects, with help from the formidable GitPython project. The idea is to go back in history historical and run a git blame (making this somewhat fast was a bit nontrivial, as it turns out, but I'll spare you the details, which involve some opportunistic caching of files, pick historical points spread out in time, use git diff to invalidate changed files, etc).

          In moment of clarity, I named “Git of Theseus” as a terrible pun on ship of Theseus. I'm a dad now, so I can make terrible puns. It refers to a philosophical paradox, where the pieces of a ship are replaced for hundreds of years. If all pieces are replaced, is it still the same ship?

          The ship wherein Theseus and the youth of Athens returned from Crete had thirty oars, and was preserved by the Athenians down even to the time of Demetrius Phalereus, for they took away the old planks as they decayed, putting in new and stronger timber in their places, in so much that this ship became a standing example among the philosophers, for the logical question of things that grow; one side holding that the ship remained the same, and the other contending that it was not the same.

          It turns out that code doesn't exactly evolve the way I expected. There is a “ship of Theseus” effect, but there's also a compounding effect where codebases keep growing over time (maybe I should call it “Second Avenue Subway” effect, after the construction project in NYC that's been going on since 1919).

          Let's start by analyzing Git itself. Git became self-hosting early on, and it's one of the most popular and oldest Git projects:


          This plots the aggregate number of lines of code over time, broken down into cohorts by the year added. I would have expected more of a decay here, and I'm surprised to see that so much code written back in 2006 is still alive in the code base – interesting!

          We can compute the decay for individual commits too. If we align all commits at x=0, we can look at the aggregate decay for code in a certain repo. This analysis is somewhat harder to implement than it sounds like because of various stuff (mostly because newer commits have had less time, so the right end of the curve represents an aggregate of fewer commits).

          For Git, this plot looks like this:


          Even after 10 years, 40% of lines of code is still present! Let's look at a broader range of (somewhat randomly selected) open source projects:


          It looks like Git is somewhat of an outlier here. Fitting an exponential decay to Git and solving for the half-life gives approx ~6 years.


          Hmm… not convinced this is necessarily a perfect fit, but as the famous quote goes: All models are wrong, some models are useful. I like the explanatory power of an exponential decay – code has an expected life time and a constant risk of being replaced.

          I suspect a slightly better model would be to fit a sum of exponentials. This would work for a repo with some code that changes fast and some code that changes slowly. But before going down a rabbit hole of curve fitting, I reminded myself of von Neumann's quote: With four parameters I can fit an elephant, and with five I can make him wiggle his trunk. There's probably some way to make it work, but I'll revisit some other time.

          Let's look at a lot of projects in aggregate (also sampled somewhat arbitrarily):


          In aggregate, the half-life is roughly ~3.33 years. I like that, it's an easy number to remember. But the spread is big between different projects. The aggregate model doesn't necessarily have super strong predictive power – it's hard to point to a arbitrary open source project and expect half of it to be gone 3.33 years later.

          Moar repos

          Apache (aka HTTPD) is another repo that goes way back:

          Rails:




          Beautiful exponential fit!

          Node:



          Wanna run it for your own repo? Again, code is available here.

          The monster repo of them all

          Note that most of these repos took at most a few minutes to analyze, using my script. As a final test I decided to run it over the Linux kernel which is HUGE – 635,229 commits as of today. This is 16 times larger than the second biggest repo I looked at (rails) and took multiple days to analyze on my shitty computer. To make it faster I ended up computing the full git blame only for commits spread out at least 3 weeks and also limited it to .c files:


          The squiggly lines are probably from the sampling mechanism. But look at this beauty – a whopping 16M lines! The code contribution from each year's cohort is extremely smooth at this scale. Individual commits have absolutely no meaning at this scale – they cumulative sum of them is very predictible. It's like going from Newton's laws to thermodynamics.


          Linux also clearly exhibits more of a linear growth pattern. I'm speculating that this has to do with its high modularity. The drivers directory has by far the most number of files (22,091) followed by arch (17,967) which contains support for various architectures. This is exactly the kind of things you would expect to scale very well with complexity, since they have a well defined interface.

          Somewhat off topic, but I like the notion of how well a projects scales with complexity. A linear scalability is the ultimate goal, where each one marginal feature takes roughly the same amount of code. Bad projects scale superlinearly, and every marginal feature takes more and more code.

          It's interesting to go back and contrast Linux to something like Angular, which basically exhibits the opposite behavior:



          The half-life of a randomly selected line in Angular is about 0.32 years. Does this reflect on Angular? Is the architecture basically not as “linear” and consistent? You might say the comparison is unfair, because Angular is new. That's a fair point. But I wouldn't be surprised if it does reflect on some questionable design. Don't mean to be shitting on Angular here, but it's an interesting contrast.

          Half-life by repository

          A somewhat arbitrary sample of projects and their half-lifes:

          Project Half-life (years) First Commit
          angular 0.32 2014
          bluebird 0.56 2013
          kubernetes 0.59 2014
          keras 0.69 2015
          tensorflow 1.08 2015
          express 1.23 2009
          scikit-learn 1.29 2011
          luigi 1.3 2012
          backbone 1.48 2010
          ansible 1.52 2012
          react 1.66 2013
          node 1.76 2009
          underscore 1.97 2009
          requests 2.1 2011
          rails 2.43 2004
          django 3.38 2005
          theano 3.71 2008
          numpy 4.15 2006
          moment 4.54 2015
          scipy 4.62 2007
          tornado 4.8 2009
          redis 5.2 2010
          flask 5.22 2010
          httpd 5.38 1999
          git 6.04 2005
          chef 6.18 2008
          linux 6.6 2005

          It's interesting that moment has such high half-life, but the reason is that so much of the code is locale-specific. This creates a more linear scalability with a stable core of code and linear additions over time. express is an outlier in the other direction. It's 7 years old but code changes extremely quickly. I'm guessing this is partly because (a) lack of linear scalability in code (b) it's probably one of the first major Javascript open source projects to hit mainstream/popularity, surfing on the Node.js wave. Possibly the code base also sucks, but I have no idea 😊

          Has coding changed?

          I can think of three reasons why there's such a strong relationship between the year the project was initiated, and the half-life

          1. Code churns more early on in projects, and becomes more stable a while in
          2. Coding has changed from 2006 to 2016, and modern projects evolve faster
          3. There's some kind of selection bias where the only projects that survive are the scalable stable ones

          Interestingly, I don't find any clear evidence of #1 in the data. The half-life for code written earlier in old projects are as high as late code. I'm skeptical about #3 as well because I don't see why there would be a relation between survival and code structure (but maybe there is). My conclusion is that writing code has fundamentally changed in the last 10 years. Code really seems to change at a much faster rate in modern projects.

          By the way, see discussion on Hacker News and on Reddit!

          AWS Reveals Multi-Agent Orchestrator Framework for Managing AI Agents

          Mike's Notes

          Some interesting examples of orchestrating AI agents.

          The article copied below is from InfoQ and is by Daniel Dominguez.

          Pipi 9 has an agent orchestrator called the Conductor Engine (cnd), so I am interested in comparing all these examples to get more insights and make a few improvements.

          Resources

          References


          Repository

          • Home > Ajabbi Research > Library >
          • Home > Handbook >

          Last Updated

          17/05/2025

          AWS Reveals Multi-Agent Orchestrator Framework for Managing AI Agents

          By: Daniel Dominguez
          InfoQ: 02/12/2024

          AWS has introduced Multi-Agent Orchestrator, a framework designed to manage multiple AI agents and handle complex conversational scenarios. The system routes queries to the most suitable agent, maintains context across interactions, and integrates seamlessly with a variety of deployment environments, including AWS Lambda, local setups, and other cloud platforms.

          The framework supports dual-language implementation in Python and TypeScript and accommodates both streaming and non-streaming agent responses. It includes pre-built agents for rapid deployment and provides extensive features such as intelligent intent classification, robust context management, and the scalability to integrate new agents or customize existing ones. This makes it a versatile tool for enterprises managing diverse AI applications.

          The high-level architecture diagram illustrates the process starting with the user input, which is analyzed by a Classifier. This Classifier uses both the characteristics of the agents and their conversation history to determine the most appropriate agent for the task. Once the agent is selected, it processes the user input, and the orchestrator updates the agent’s conversation history before delivering the response back to the user.

          AWS has also published a demo application showcasing the orchestrator’s capabilities. The demo includes six specialized agents, such as those for travel, weather, math, and health. These agents demonstrate the system's ability to seamlessly switch between tasks and maintain coherence in multi-turn conversations. Additional projects, such as a multilingual chatbot for flight reservations and an AI-powered e-commerce support system, highlight the framework's flexibility in addressing specialized use cases.

          Beyond text-based interactions, the orchestrator supports voice-based systems by integrating with tools like Amazon Connect and Lex. This capability enhances its application for AI-driven customer service, call centers, and other domains requiring natural language and voice interaction.

          In the community, the launch has generated interest from key figures in AI and ML.

          Bob Xu, founder of Quest Flow, a company specializing in multi-agent orchestration, commented on X, stating:

          "Multi-agent is going mainstream now."

          And AI/ML tech sourcer Jung Kim shared on X:

          "AWS Labs' Multi-Agent Orchestrator open source project rethinks distributed computing to make it easier to build sophisticated, efficient, and cost-effective AI systems."

          AWS's Multi-Agent Orchestrator launches as part of the growing trend towards agent-based AI systems. Other frameworks in this space include Microsoft Research's Magentic-One, a generalist multi-agent system for solving open-ended tasks, and IBM's Bee Agent Framework, designed for scalable, agent-based workflows. OpenAI has also introduced Swarm, a system focused on building and deploying multi-agent configurations. These frameworks collectively reflect an industry-wide effort to enhance the orchestration and efficiency of multi-agent AI deployments.

          DeepSeek Debates: Chinese Leadership On Cost, True Training Cost, Closed Model Margin Impacts

          Mike's Notes

          This excerpt from a recent article from SemiAnalysis sheds more light on the recent AI hype. DeepSeek is a significant improvement, but not everything seems as good.

          The full article requires a login to SemiAnalysisResources.

          I have also added some other articles to the references.

          Resources

          References

          • Reference

          Repository

          • Home > Ajabbi Research > Library > Subscriptions > SemiAnalysis
          • Home > Handbook > 

          Last Updated

          17/05/2025

          DeepSeek Debates: Chinese Leadership On Cost, True Training Cost, Closed Model Margin Impacts

          By: Dylan Patel, AJ Kourabi, Doug O'Laughlin, Reyk Knuhtsen
          SemiAnalysis: 31/01/2025

          The DeepSeek Narrative Takes the World by Storm

          DeepSeek took the world by storm. For the last week, DeepSeek has been the only topic that anyone in the world wants to talk about. As it currently stands, DeepSeek daily traffic is now much higher than Claude, Perplexity, and even Gemini.

          But to close watchers of the space, this is not exactly “new” news. We have been talking about DeepSeek for months (each link is an example). The company is not new, but the obsessive hype is. SemiAnalysis has long maintained that DeepSeek is extremely talented and the broader public in the United States has not cared. When the world finally paid attention, it did so in an obsessive hype that doesn’t reflect reality.

          We want to highlight that the narrative has flipped from last month, when scaling laws were broken, we dispelled this myth, now algorithmic improvement is too fast and this too is somehow bad for Nvidia and GPUs.

          The narrative now is that DeepSeek is so efficient that we don't need more compute, and everything has now massive overcapacity because of the model changes. While Jevons paradox too is overhyped, Jevons is closer to reality, the models have already induced demand with tangible effects to H100 and H200 pricing.

          DeepSeek and High-Flyer

          High-Flyer is a Chinese Hedge fund and early adopters for using AI in their trading algorithms. They realized early the potential of AI in areas outside of finance as well as the critical insight of scaling. They have continuously increasing their supply of GPUs as a result. After experimentation with models with clusters of thousands of GPUs, High Flyer made an investment in 10,000 A100 GPUs in 2021 before any export restrictions. That paid off. As High-Flyer improved, they realized that it was time to spin off “DeepSeek” in May 2023 with the goal of pursuing further AI capabilities with more focus. High-Flyer self funded the company as outside investors had little interest in AI at the time, with the lack of a business model being the main concern. High-Flyer and DeepSeek today often share resources, both human and computational.

          DeepSeek now has grown into a serious, concerted effort and are by no means a “side project” as many in the media claim.  We are confident that their GPU investments account for more than $500M US dollars, even after considering export controls.

          Source: SemiAnalysis, Lennart Heim

          The GPU Situation

          We believe they have access to around 50,000 Hopper GPUs, which is not the same as 50,000 H100, as some have claimed. There are different variations of the H100 that Nvidia made in compliance to different regulations (H800, H20), with only the H20 being currently available to Chinese model providers today. Note that H800s have the same computational power as H100s, but lower network bandwidth. We believe DeepSeek has access to around 10,000 of these H800s and about 10,000 H100s. Furthermore they have orders for many more H20's, with Nvidia having produced over 1 million of the China specific GPU in the last 9 months. For more specific detailed analysis, please refer to our Accelerator Model.

          Source: SemiAnalysis

          Our analysis shows that the total server CapEx for DeepSeek is almost $1.3B, with a considerable cost of $715M associated with operating such clusters.

          DeepSeek has sourced talent exclusively from China, with no regard to previous credentials, placing a heavy focus on capability and curiosity. DeepSeek regularly runs recruitment events at top universities like PKU and Zhejiang, where many of the staff graduated from. Roles are not necessarily pre-defined and hires are given flexibility, with jobs ads even boasting of access to 10,000s GPUs with no usage limitations. They are extremely competitive, and allegedly offer salaries of over $1.3 million dollars USD for promising candidates, well a over big Chinese tech companies. They have ~150 employees, but are growing rapidly.

          As history shows, a small well-funded and focused startup can often push the boundaries of what’s possible. DeepSeek lacks the bureaucracy of places like Google, and since they are self funded can move quickly on ideas. However, like Google, DeepSeek (for the most part) runs their own datacenters, without relying on an external party or provider. This opens up further ground for experimentation, allowing them to make innovations across the stack.

          We believe they are the single best “open weights" lab today, beating out Meta’s Llama effort, Mistral, and others.

          DeepSeek’s Cost and Performance

          DeepSeek’s price and efficiencies caused the frenzy this week, with the main headline being the “$6M” dollar figure training cost of DeepSeek V3. This is wrong. This akin to pointing to a specific (and large) part of a bill of materials and attribute it as the entire cost. The pre-training cost is a very narrow portion of the total cost.

          Training Cost

          We believe the pre-training number is nowhere near close to the actual amount spent on the model. We are confident their hardware spend is well higher than $500M over the company history. To develop new architecture innovations, during the model development, there is a considerable spend on testing new ideas, new architecture ideas, and ablations. Multi-Head Latent Attention, a key innovation of DeepSeek, took several months to develop and cost a whole team of manhours and GPU hours.

          The $6M cost in the paper is attributed to just the GPU cost of the pre-training run, which is only a portion of total cost of the model. Excluded are important pieces of the puzzle like R&D and TCO of the hardware itself. For reference, Claude 3.5 Sonnet cost $10s of millions to train, and if that was the total cost Anthropic needed, then they would not raise billions from Google and tens of billions from Amazon. It's because they have to experiment, come up with new architectures, gather and clean data, pay employees, and much more.

          So how was DeepSeek able to have such a large cluster? The lag in export controls is the key, and will be discussed in the export section below.

          Closing the Gap - V3’s Performance

          V3 is no doubt an impressive model, but it is worth highlighting impressive relative to what. Many have compared V3 to GPT-4o and highlight how V3 beats the performance of 4o. That is true but GPT-4o was released in May of 2024. AI moves quickly and May of 2024 is another lifetime ago in algorithmic improvements. Further we are not surprised to see less compute to achieve comparable or stronger capabilities after a given amount of time. Inference cost collapsing is a hallmark of AI improvement.

          Source: SemiAnalysis

          An example is small models that can be run on laptops have comparable performance to GPT-3, which required a supercomputer to train and multiple GPUs to inference. Put differently, algorithmic improvements allow for a smaller amount of compute to train and inference models of the same capability, and this pattern plays out over and over again. This time the world took notice because it was from a lab in China. But smaller models getting better is not new.

          Source: SemiAnalysis, Artificialanalysis.ai, Anakin.ai, a16z

          So far what we've witnessed with this pattern is that AI labs spend more in absolute dollars to get even more intelligence for their buck. Estimates put algorithmic progress at 4x per year, meaning that for every passing year, 4x less compute is needed to achieve the same capability. Dario, CEO of Anthropic arguea that algorithmic advancements are even faster and can yield a 10x improvement. As far as inference pricing goes for GPT-3 quality, costs have fallen 1200x.

          When investigating the cost for GPT-4, we see a similar decrease in cost, although earlier in the curve. While the decreased difference in cost across time can be explained by no longer holding the capability constant like the graph above. In this case, we see algorithmic improvements and optimizations creating a 10x decrease in cost and increase in capability.

          Source: SemiAnalysis, OpenAI, Together.ai

          To be clear DeepSeek is unique in that they achieved this level of cost and capabilities first. They are unique in having released open weights, but prior Mistral and Llama models have done this in the past too. DeepSeek has achieved this level of cost but by the end of the year do not be shocked of costs fall another 5x.

          Is R1’s Performance Up to Par with o1?

          On the other hand, R1 is able to achieve results comparable to o1, and o1 was only announced in September. How has DeepSeek been able to catch so fast?

          The answer is that reasoning is a new paradigm with faster iteration speeds and lower hanging fruit with meaningful gains for smaller amounts of compute than the previous paradigm. As outlined in our scaling laws report, the previous paradigm depended on pre-training, and that is becoming both more expensive and difficult to achieve robust gains with.

          The new paradigm, focused on reasoning capabilities through synthetic data generation and RL in post-training on an existing model, allows for quicker gains with a lower price. The lower barrier to entry combined with the easy optimization meant that DeepSeek was able to replicate o1 methods quicker than usual. As players figure out how to scale more in this new paradigm, we expect the time gap between matching capabilities to increase.

          Note that the R1 paper makes no mention of the compute used. This is not an accident – a significant amount of compute is needed to generate synthetic data for post-training R1. This is not to mention RL. R1 is a very good model, we are not disputing this, and catching up to the reasoning edge this quickly is objectively impressive. The fact that DeepSeek is Chinese and caught up with less resources makes it doubly impressive.

          But some of the benchmarks R1 mention are also misleading. Comparing R1 to o1 is tricky, because R1 specifically doesn't mention benchmarks that they are not leading in. And while R1 is matches reasoning performance, it's not a clear winner in every metric and in many cases it is worse than o1.


          Source: (Yet) another tale of Rise and Fall: DeepSeek R1

          And we have not mentioned o3 yet. o3 has significantly higher capabilities than both R1 or o1. In fact, OpenAI recently shared o3’s results, and the benchmark scaling is vertical. "Deep learning has hit a wall", but of a different kind.

          Source: AI Action Summit

          Google’s Reasoning Model is as Good as R1

          While there is a frenzy of hype for R1, a $2.5T US company released a reasoning model a month before for cheaper: Google’s Gemini Flash 2.0 Thinking. This model is available for use, and is considerably cheaper than R1, even with a much larger context length for the model through API.

          On reported benchmarks, Flash 2.0 Thinking beats R1, though benchmarks do not tell the whole story. Google only released 3 benchmarks so it's an incomplete picture. Still, we think Google’s model is robust, standing up to R1 in many ways while receiving none of the hype. This could be because of Google’s lackluster go to market strategy and poor user experience, but also R1 is a Chinese surprise. 


          Source: SemiAnalysis

          To be clear, none of this detracts from DeepSeek’s remarkable achievements. DeepSeek’s structure as a fast moving, well-funded, smart and focused startup is why it's beating giants like Meta in releasing a reasoning model, and that's commendable.

          Technical Achievements

          DeepSeek has cracked the code and unlocked innovations that leading labs have not yet been able to achieve. We expect that any published DeepSeek improvement will be copied by Western labs almost immediately.  

          What are these improvements? Most of the architectural achievements specifically relate to V3, which is the base model for R1 as well. Let’s detail these innovations.

          Training (Pre and Post)

          DeepSeek V3 utilizes Multi-Token Prediction (MTP) at a scale not seen before, and these are added attention modules which predict the next few tokens as opposed to a singular token. This improves model performance during training and can be discarded during inference. This is an example of an algorithmic innovation that enabled improved performance with lower compute.   

          There are added considerations like doing FP8 accuracy in training, but leading US labs have been doing FP8 training for some time.

          DeepSeek v3 is also a mixture of experts model, which is one large model comprised of many other smaller models that specialize in different things. One struggle MoE models have faced has been how to determine which token goes to which sub-model, or “expert”. DeepSeek implemented a “gating network” that routed tokens to the right expert in a balanced way that did not detract from model performance. This means that routing is very efficient, and only a few parameters are changed during training per token relative to the overall size of the model. This adds to the training efficiency and to the low cost of inference.

          Despite concerns that Mixture-of-Experts (MoE) efficiency gains might reduce investment, Dario points out that the economic benefits of more capable AI models are so substantial that any cost savings are quickly reinvested into building even larger models. Rather than decreasing overall investment, MoE's improved efficiency will accelerate AI scaling efforts. The companies are laser focused on scaling models to more compute and making them more efficient algorithmically.             

          In terms of R1, it benefited immensely from having a robust base model (v3). This is partially because of the Reinforcement Learning (RL). There were two focuses in RL: formatting (to ensure it provides a coherent output) and helpfulness and harmlessness (to ensure the model is useful). Reasoning capabilities emerged during the fine-tuning of the model on a synthetic dataset. This, as mentioned in our scaling laws article, is what happened with o1. Note that in the R1 paper no compute is mentioned, and this is because mentioning how much compute was used would show that they have more GPUs than their narrative suggests. RL at this scale requires a considerable amount of compute, especially to generate synthetic data.

          Additionally a portion of the data DeepSeek used seems to be data from OpenAI’s models, and we believe that will have ramifications on policy on distilling from outputs. This is already illegal in the terms of service, but going forward a new trend might be a form of KYC (Know Your Customer) to stop distillation.  

          And speaking of distillation, perhaps the most interesting part of the R1 paper was being able to turn non-reasoning smaller models into reasoning ones via fine tuning them with outputs from a reasoning model. The dataset curation contained a total of 800k samples, and now anyone can use R1’s CoT outputs to make a dataset of their own and make reasoning models with the help of those outputs. We might see more smaller models showcase reasoning capabilities, bolstering performance of small models.

          Multi-head Latent Attention (MLA)

          MLA is a key innovation responsible for a significant reduction in the inference price for DeepSeek. The reason is MLA reduces the amount of KV Cache required per query by about 93.3% versus standard attention. KV Cache is a memory mechanism in transformer models that stores data representing the context of the conversation, reducing unnecessary computation.

          As discussed in our scaling laws article, KV Cache grows as the context of a conversation grows, and creates considerable memory constraints. Drastically decreasing the amount of KV Cache required per query decreases the amount of hardware needed per query, which decreases the cost. However we think DeepSeek is providing inference at cost to gain market share, and not actually making any money. Google Gemini Flash 2 Thinking remains cheaper, and Google is unlikely to be offering that at cost. MLA specifically caught the eyes of many leading US labs. MLA was released in DeepSeek V2, released in May 2024.

          DeepSeek has also enjoyed more efficiencies under inference with the H20, due to higher memory and bandwidth capacity compared to the H100. They have also announced partnerships with Huawei but very little has been done with them so far with Ascend compute.

          We believe the most interesting implications is specifically on margins, and what that means for the entire ecosystem. Below we have a view of the future pricing structure of the entire AI industry, and we detail why we think DeepSeek is subsidizing price, as well as why we see early signs that Jevons paradox is carrying the day. We comment on the implications on export controls, how the CCP might react with added DeepSeek domninance, and more.

          Super Colour Palette

          Mike's Notes

          Dave Gray from Explorations and Explanations mentioned this great tool in his recent Substack. It allows you to choose a colour palette that can be exported.

          Resources

          References

          • Reference

          Repository

          • Home > Ajabbi Research > Library >
          • Home > Handbook > 
          • Home > Ajabbi Design > Design System > Colour

          Last Updated

          17/05/2025

          Super Colour Palette

          By: Dave Gray
          Explorations and Explanations: 07/02/2025