Showing posts with label DevOps. Show all posts
Showing posts with label DevOps. Show all posts

Nestspace

Mike's Notes

Thursday, 09/04/2026 (NZ time) was the longest day. I couldn't make a single mistake, so I turned off the phone and the internet to concentrate from 8am to 9pm, then slept like a log. I was a walking zombie yesterday. I'm glad this will never need to be done again.

This is the successful culmination of months of mentally challenging preparatory work for the new Pipi Core data centre.

A typical day for me is

  • 50% learning
  • 40% thinking
  • 10% doing

The less I do, the more productive the solutions. The faster the progress.

Speed will come from Pipi-driven automation, not Mike working harder.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

23/07/2026

Nestspace

By: Mike Peters
On a Sandy Beach: 10/04/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

Inside the Pipi Core data centre, the careful migration and reorganisation of hundreds of thousands of Pipi-related files began and were completed yesterday. Each file or directory was placed in only one Nestspace, and then the historical Nestspaces were converted into zip​ archives

It is a good solution, but it took a month of trial and error to figure out. It had to be 100% correct to enable rapid, reliable data centre automation, self-managed by Pipi.

It was fascinating to watch the recent Lex Fridman interview with NVIDIA CEO Jensen Huang, who described NVIDIA's large-scale problem-solving process.

Problems to solve

A large number of problems had to be solved in parallel. Each problem affected the others.

  • Where to put self-organising Pipi swarms.
  • Pipi instances
    • Naming
    • Self-evolving
    • Versioning
  • How to use DevOps automation with all of the above.
  • The knock-on effects on
    • Namespaces
    • Backups
    • Replication
    • Accounts
    • Workspaces
    • i18n
    • Web URLs
    • Developer documentation
    • Training
    • etc.
  • Securing 100% security and privacy
  • Cross-platform and environment portability

This has led to changes in the underlying Pipi System Engine (sys) data model, which I will write about tomorrow.

Nestspace

The fundamental organising principle is to use a uniquely named directory, now named as a "Nestspace".

The unique name is a string combination made of

  • <pipi major version> (integer)
  • <pipi edition> (single lowercase letter)
  • <account type> (single lowercase letter)

Data Centre / DevOps

Here are some examples.

  • 6pg/
  • 9ae/

Backups

Customer backup examples.

  • pipi_9ae_ajabbi_data_pg_20260411
  • pipi_9ae_ajabbi_pipi_www_learn.ajabbi.com_20260411

Archiving

Here are some examples.

  • pipi_6pg_20180219.zip
  • pipi_9ae_20241211.zip

Next job

Now, I can start configuring the files and settings for each Nestspace. The necessary code has already been successfully tested. Eventually, this process will be completely automated.

  • Application.cfc
  • server.xml
  • Datasources
  • OS environment
  • Java
  • Application server
  • Cloud platform
  • REPL

Speed is King

Then, each Pipi will be back in business and can be left running 24x7 in its own Nestspace, thereby increasing DevOps speed by at least 10x. The priority now is to increase Pipi's Data Centre speed by 1000x using automation and keep going. So what previously took a year can be done in an hour. The over-optimistic complete deadline is June 2026.

But nothing happens as expected

It is hard to make completion predictions when this architecture is completely novel, and everything is uncharted territory and hellish difficult. Everything is a back-of-the-envelope guess. I'm slowly getting there ... 

"like a man riding a drunken donkey facing the tail, two steps forwards, one step backwards"

... enjoying the whole journey and it's a hell of a lot of fun. 

Future customers

The increase in speed will also directly benefit all future customers using the SaaS workspace applications. Deployments, configuration, and updates will also get the same 1000x speed increases at no extra cost.

Pipi Data Centre Operational

Mike's Notes

Very good news for Pipi.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

26/02/2026

Pipi Data Centre Operational

By: Mike Peters
On a Sandy Beach: 25/02/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

The long-planned migration of Pipi 9 to its own data centre has been completed. It took 2 weeks to execute. The existing setup was split into an office network connected to the internet and an isolated data centre that is not connected to the internet.

Starting from zero

The initial data centre consists of a single 45U rack and some other shelving, with mainly older equipment. It will do for a start and can grow as more racks are added, equipment upgraded, and more servers are added, etc.

External hard drives being used in the shift

Issues

  • Terabytes of data on backup hard drives to shift
  • Clean reinstalls of many operating systems
  • 14 machines to configure
  • Adobe CS4 does not like Windows 11
  • Making do with what is available now
  • Go slow, think twice and get there faster

Opportunities

  • Pipi on 24x7x365
  • All systems can be turned on using multiple servers
  • DevOps automation is now possible
  • The development cycle will speed up 10x
  • The road to Pipi 10 with BoxLang is now open

Whats next

  • Seat-of-the-pants experimenting to tune the setup
  • Stress test to build resiliency and reliability

Developer access to Pipi, is coming

Mike's Notes

The data is precise on this one. Unfortunately, the current developer interest in NZ and Australia is 3% and 0%, respectively. I also can't find anyone in NZ who has the slightest technical understanding of what I'm doing. But there are plenty overseas, especially in MLOps. We speak the same language, even if the architecture and algorithms are radically different. Also, top-grade mathematicians get it. Internally, Pipi 9 uses a lot of maths.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

08/12/2025

Developer access to Pipi is coming

By: Mike Peters
On a Sandy Beach: 02/12/2025

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

The problem

To launch Pipi 9 with limited resources in 2026, the effort needs to be highly focused on developers who build large enterprise systems and have experienced failure, cost overruns, and staggering complexity. This is for them.

Waste

The staggering annual global cost of IT failures on big projects is around $US 3 trillion. How many schools or knee replacements would that be?

  • 15% succeed
  • 25% makes no difference
  • 60% fail

Web traffic stats

The steadily growing web traffic statistics of public interest in Pipi 9 are becoming very clear.

Since early 2019, the total traffic stats by country are;

  • China 19%
  • Singapore 16%
  • United States 15%
  • Hong Kong 11%
  • Brazil 10%
  • TOTAL 71%

Developer Accounts

The initial paid Developer Accounts will be restricted to experienced DevOps teams with great internal culture in those five countries. That also affects the language, currency, hours of support availability, etc. Later, as interest and resources grow, that list of countries can be expanded.

Personal Accounts

The free Personal Accounts will initially use an English interface and will not be restricted by country of residence. They will get community support.

Enterprise Accounts

The initial paid Enterprise Accounts will be supported by their associated Developer Accounts, who can charge them whatever they want for that service. They will initially use an English interface and will not be restricted by country of residence. Developer Accounts will be able to translate UI and documentation into any language and writing system.

Pipi 9 is in hiding

This engineering blog and the many other Ajabbi documentation websites are deliberately hidden from search engines. People are visiting because they are curious, as I write notes to myself and build, learning as I go. It is not easy to find the technical documentation unless you are really interested, clever and very determined. That has helped me find some early, keen technical fans who provide testing and feedback. It has also protected me from being overwhelmed by enquiries.

Communication constraints

I am a very slow writer, using assistive technology, have hearing problems and prefer video chats with people who speak good, clear English. I don't speak any other language apart from tiny bits of Maori, French and Spanish.

SEO and GEO

The SEO/GEO settings will be fixed when

  • Pipi 9 matures and becomes ready for public use
  • Community support is in place
  • Bugs fixed
  • Enough self-help documentation to help people get started.

Developer Account waitlist

There will be a signup queue to control demand, so scaling is steady with a positive resource feedback loop to solve the chicken-and-egg problem. The small queue is growing now. I will pick the best candidates with the highest chance of success. They will gain a first-mover advantage in building large, custom enterprise systems faster and at a much lower cost. The first ones will get free unlimited support. 

Relying on word-of-mouth recommendations.

There will be no marketing or sales, just good, clear documentation, live demos, and regular bookable office hours (NZ daytime) for having a chat.

Inflexion point

In the future, as workspaces mature and Pipi 10 becomes even easier to work with, resource constraints will disappear, teams will grow, and an inflexion point will be reached. Anyone will then be able to sign up.

Updates to Versioning Engine

Mike's Notes

This week's changes to the Versioning Engine (ver).

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

16/10/2025

Updates to Versioning Engine

By: Mike Peters
On a Sandy Beach: 16/10/2025

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

The recent UI workspaces work necessitated changes to Pipi's Versioning Engine (ver). That engine's data model was changed.

Objects that can change

  • Industry models
  • Industry Schema
  • Workspace UI's
  • User training material
  • Developer Documentation
  • Etc

Versioned namespaces

Every object has a versioned namespace. When a logged change to an object occurs, the Pipi Build Number increments.

On any day, if there is one or more changes to the Build Number, the daily Patch Release automatically increases by 1. The changes are then published in the Changelog.

Minor Release

On the second Friday of every third month, the Minor Release increases by 1. A big picture summary of the preceding three months is published in the Changelog.

Objects in sync

Changes to the workspace UI, developer documentation, user training materials, and related materials need to be kept in sync.

DevOps

Making sure that the recently added DevOps Engine (dvp) runs this sync process might work. It could be done using workflows.

Summary

The engines involved are 

  • Versioning
  • DevOps
  • Learning Object
  • Workspace
  • CMS

A roadmap to Pipi 10

Mike's Notes

Pipi 10 will run on BoxLang in production.

I sent draft notes to Ortus Solutions yesterday to discuss future use of BoxLang. I then edited the notes and shared them with the growing Pipi community. I will add the answers as more is discovered about what needs to be done to make it possible for Pipi 10 to use BoxLang.

Training is available on BoxLang Academy.

Can you spot any errors in my assumptions? Do you have other ideas or queries?

Contact me.

Update 04/06/2024, 05/06/2025

Cristóbal Escobar Henríquez from Ortus Solutions provided excellent answers to the questions sent via email. These answers have been added to the questions at the end of this article, which are slightly different from the original ones (all my fault). Thanks :)

Update 09/06/2025

Luis Majano from Ortus Solutions provided additional answers to my follow-up questions, which I sent by email. They have been added to the post below. Thank you, Luis.

Update 12/06/2025

Luis Majano from Ortus Solutions provided additional answers to my follow-up questions, which I sent by email. They have been added to the post below. Thank you, Luis.

Update 13/06/2025

Luis confirmed that BoxLang cannot use CommandBoc modules. It appears that BoxLang can currently handle 90% of my needs. The only real gap is that Pipi cannot currently autonomously configure BoxLang itself via the command line, which is needed for scaling.

This is sufficient to make a start, with the hope that Ortus will address these issues in the future.

I expect Pipi 10 to be a 12-month build after Pipi 9 is in full production.

Update 14/06/2025

There will be an online meeting with people from Ortus next week to discuss and confirm details regarding the migration plan, paid support, and sponsorship of Ortus' open-source.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

14/06/2025

A roadmap to Pipi 10

By: Mike Peters
On a Sandy Beach: 21/05/2025

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

Draft roadmap

  • This is a draft, so expect it to change
  • This roadmap is probably for the next 2-3 years at least
  • The finalised plan will be published for the Pipi Community.
  • No money is available right now.
  • A plan is needed to avoid rework.
  • Don't expect Ortus to do work for free.
  • Realistic expectations of what is possible.
  • Good income will come later from paying customers for solving complex problems.
  • Focus on the core by outsourcing other work to experienced contractors.
  • Open-source all modules on GitHub.

Ajabbi

  • The common name of 
    • A company to handle payments
    • Non-profit foundation to support users (to be established)
    • Research Institute for ongoing development (to be established)
  • Description
    • Solved a huge, expensive problem with no existing competitors
    • Pre-revenue start-up with no $$$
    • Growing community
  • Purpose
    • Act like other foundations (Wikimedia, Apache and Linux) and use the income surplus to support
      • Minority languages i18n
      • Open-source efforts (especially CFML, including Ortus)
      • User groups
      • User conferences
      • Book publishing
      • Open research grants
      • Etc

    Pipi

    • The name of the software
    • Mike has been the architect since 1997
    • Versions 1-4 were in production eventually with 300K lines of code, an 800-table database, ESRI Web GIS, 500 classes, 3K methods, a 5 FTE team and a 3-rack data centre
    • Versions 6-8 were working, non-production, undergoing refactoring
    • Version 9 works, moving to production (current)
    • Version 10 (planned)

    Current Pipi 9

    • Works great
    • Runs in a dev environment
    • Unique architecture with 20+ layers and hundreds of autonomous agents
    • Uses nature-inspired algorithms
    • Going from headless to a Pipi-generated UI
    • Migrating to a 42U server rack completely isolated from the internet.
    • Using old Windows server boxes will do for now.
    • The CFML code runs on ColdFusion Server 2021 Developer Edition.
    • The CFML code could run on ColdFusion MX, Open Blue Dragon, or Lucee. 
    • Uses <CF> tags as in Ben Forta, et al. (don't tell Michaela, I don't want to die)
    • Uses 300+ embedded config databases (MS Access)
    • Uses PostgreSQL for production data
    • Writes CFML code and then executes it
    • Writes SQL and then executes it to create and populate tables
    • Static HTML generator (90K page websites)
    • Uses XML, XML Schema, etc, internally
    • Mostly self-managing
    • Self-documenting
    • Currently rendering a 20,000-page static website for developers
    • The code base will self-grow 10x larger over time

    Planned Pipi 9

    • Some money is coming soon
    • Generate revenue from niche batch render jobs
    • Code generators and the generated code must be rewritten to be BoxLang compatible, so no refactoring by Ortus is required.
    • Pipi 9 will create Pipi 10
    • Pipi 10 to be BoxLang native
    • BoxLang (+++++++) paid support
    • Pipi 10 to become geodatabase capable
    • BoxLang will enable the future use of Python, PHP, etc

      Planned Pipi 10 Systems

      • Render System
        • Functions as an automated AI-assisted development environment
        • Isolated from the internet
        • Uses VM
        • Run on BoxLang installed on Windows on new 64 hardware
        • JDK 21+
        • Ortus to supply a ready-to-run VM image.
        • 100% self-managing via CommandBox dynamic CFML scripts
        • Renders working enterprise system modules (database, code, UI, API, user documentation)
          • MS SQL edition
          • MySQL edition
          • Oracle edition
          • PostgreSQL edition
        • Exports modules to the Staging System via manual hard drive transfers
      • Staging System
        • Connected to the internet
        • Uses VM
        • Runs on BoxLang installed on Debian
        • JDK 21+
        • Ortus to manage the BoxLang platform VM on a long-term contract
        • Uses PostgreSQL
        • Adds i18n to modules
        • Creates updates
        • Integrates i18n modules into customised enterprise systems for paying customers
        • Exports
          • Github
            • i18n modules to GitHub for open-source
            • i18n user documentation
          • Cloud
            • Customised i18n enterprise systems
            • System updates
      • Cloud System
        • Connected to the internet
        • Uses VM
        • Runs on BoxLang installed on Debian
        • JDK 21+
        • Ortus to manage the BoxLang platform VM on a long-term contract
        • VM hosted
          • AWS
          • Azure
          • Digital Ocean
          • GCP
          • IBM
          • etc
          • private cloud.
        • The customer chooses their production RDMS database(s) for their data.
          • MS SQL
          • MySQL
          • Oracle
          • PostgreSQL

      Questions for Ortus Solutions

      Answered by Cristóbal Escobar Henríquez, and Luis Majano.
      • Is a list of supported (or not supported) CFML tags available?
      • Is <Cf tags> style coding supported?
        • Yes, BoxLang fully supports CFML tags when running in CFML compatibility mode. This ensures legacy CFML applications can continue to function with minimal changes.
      • Is there a performance hit from using CML compatibility mode?
        • There is of course a penalty to pay if you are transpiling on the fly.  It's negligible to the point that we are faster in CFML mode than all Adobe engines even in transpile mode.  However, you can transpile to BX as part of your build process and then this won't exist.
      • Does BoxLang use CommandBox modules?
        • No.  BoxLang Modules are something new.  Eventually maybe, but for now, no.
      • Can BoxLang accept (CFML) command-line modules generated by Pipi without needing to be restarted?
        • Yes.
        • The ModuleService has the necessary API for you to load modules a-la-carte.  All of our integration testing uses this approach.
        • You can also uninstall them
      • Will these CF tags become compatible with BoxLang in future?
        • cfftp
          • This is done in our bx-ftp module
        • cfregistry
          • This can be sponsored if needed.
        • cfschedule
          • This is in progress, but can be accelerated by sponsorship
      • Could BoxLang use MS Access via JDBC on the isolated Render System, and how much would it cost to create a module to do this?
        • We can enable MS Access integration by building a custom module using the UCanAccess JDBC driver.
        • Estimated time: 10–15 hours (2–3 business days)
        • Note: This estimate covers driver integration only and does not guarantee full MS Access compatibility. As an alternative, we recommend migrating the database to SQLite, Derby, or MySQL for improved stability and long-term support.
      • What kinds of modules can be developed by Ortus?
        • Ortus can develop any module you need, whether it's for BoxLang, application features, integrations, or infrastructure. We tailor modules to your requirements, from core language enhancements to platform-level tools.
      • What is the cost for Ortus to provide dedicated platform support?
        • We offer dedicated platform support through our BoxLang+ and BoxLang++ plans, which include guaranteed response times, SLAs, priority bug handling, and access to our core engineering team.
        • You can review the full details and pricing here: https://boxlang.io/plans
        • We’d be happy to tailor a plan based on your service level and usage expectations.
      • Can patched VM images be supplied?
        • Yes, we can provide patched VM images as part of a BoxLang++ subscription. If you're not on a ++ tier, this can still be delivered through a separate consulting plan.
      • What about the time zone differences between New Zealand, the US, and Spain?
        • Our team operates mainly on Central Standard Time, but we can easily align with your schedule. If needed, we can also assign a developer located in the EU (CET) for better time zone overlap.
      • Would a 100% dedicated Ortus developer be possible in the future?
        • Yes, we offer team augmentation services, including assigning a dedicated developer to your project. We can adapt the engagement based on scope, timeline, and availability.
      • Is a 24/7/365 platform supported by Ortus possible in the future?
        • Yes, that’s absolutely possible. We can support a 24/7/365 production environment using our teams in El Salvador and Spain. This would require a custom service agreement or license tier beyond the standard BoxLang offerings, but we’re happy to scope and price this accordingly based on your needs in the future.

      Questions for the Pipi Community

      • Would support also be available from independent CFML contractors?

      • Where is the boundary between community contributions and paid contractors?

      • What other new BoxLang Modules would be helpful, including other coding languages?

      The Shift Left Data Manifesto

      Mike's Notes

      The Shift Left Data Manifesto is written by Chad Sanderson, CEO & Co-Founder of Gable.ai

      It is reproduced below. Thoughtful ideas.

      Resources

      References


      Repository

      • Home > Ajabbi Research > Library >
      • Home >

      Last Updated

      27/03/2025

      The Shift Left Data Manifesto

      By: Chad Sanderson
      Gable.ai Blog: 25/03/2025

      Chad Sanderson is the CEO & Co-Founder of Gable.ai

      A core idea behind shifting Data Left is simple but often overlooked: data is code. Or more accurately—data is produced by code. It’s not just some downstream artifact that lives in tables and gets piped into dashboards and spreadsheets. Every record, event, or log starts somewhere—created, updated, or deleted by a line of code. And just like DevOps demonstrated, if you want to manage something well, you start at the point of creation.

      Data Management for Software Engineering Teams

      Hello everyone, my name is Chad Sanderson. I am the author of the blog Data Products and the CEO/Co-Founder of Gable.ai. Over the past few years, I’ve written quite a bit on data management, data quality, and data contracts. In 2024 I spent a little bit less time writing, and more time implementing. I’ve worked with dozens of enterprises during that time span and watched the evolution of data contracts from a nascent idea stemming from a few LinkedIn posts to driving real change for some of the largest companies in the world.

      The result of that experience has been I have become a bit of a data extremist. I believe there is a completely new domain of data management over the horizon, one that will altogether change how we think about the discipline, rewrite most/all of our common best practices, and bring the various stakeholders into a cohesive lifecycle of data management. This is a mix of opinions, combined with a description of the cutting-edge - truly game-changing companies that are pioneering how data management is done, oftentimes from unexpected places.

      Over the course of this manifesto, I will try to convince you of a few things I strongly believe:

      1. Most engineering teams are federated or becoming so
      2. The way we manage data is designed for centralized environments
      3. Data strategies will almost always fail, due to points 1 and 2
      4. Federated data management is possible, but requires a different approach
      5. That approach has been historically successful in other engineering disciplines

      Is this manifesto about data contracts? No. But they do feature prominently. Data contracts are a component of shifting left, among dozens of other components. These components work together to create an entirely new dynamic of data management, completely inverting the processes, tools, strategies, and adoption rates of data quality and data governance. There is so much content to cover, across such a wide variety of topics that I’m splitting the manifesto into two parts for my own sanity. The first section is going to cover where we are today, why I believe the state of data management is fundamentally flawed, why “culture shift” is almost always impossible without technology, and working solutions, and what we can learn from other industries that have solved the same type of problems. Let’s jump into it.

      Conway’s Law

      Conway’s Law is the observation that organizations tend to design systems that mirror their communication structure. A product designed by a three-person organization will likely have three components. If it is designed by a single team, it will likely all be built within a single large service.

      A media company with separate teams for video encoding, recommendation algorithms, and user interfaces might build a streaming platform where these components are loosely coupled, reflecting the team's structure. Hospitals and insurance companies have separate IT systems due to distinct legal and compliance teams. This results in fragmented medical records across different providers, forcing patients to manually transfer records or redo tests.

      There are three primary stakeholders in the data management value chain:

      1. Producers: The teams generating the data
      2. Platforms: The teams maintaining the data infrastructure
      3. Consumers: The teams leveraging the data to accomplish tasks

      Conway's Law would dictate that the data management, governance, and quality systems implemented in a company will reflect how these various groups work together.

      In most businesses, data producers have no idea who their consumers are or why they need the data in the first place. They are unaware of which data is important for AI/BI, nor do they understand what it should look like. Platform teams are rarely informed about how their infrastructure is being leveraged and have little knowledge of the business context surrounding data, while consumers have business context but don't know where the data is coming from or whether or not it's quality.

      Is it any wonder that data management programs are a complete, disjointed mess?

      The opposite side of the coin of Conway’s Law is the Law of Unintended Consequences Systems, or to summarize - "The purpose of a system is what it does" (POSIWID) – coined by Stafford Beer, a cybernetics researcher. The rule means that what a technology does is more illustrative of what its intended goal is, rather than any stated intent.

      For example, suppose data pipelines are consistently breaking and the data is always low quality. In that case, it means the point of your data ecosystem is not to produce high quality trusted data - it is actually to enable teams to move fast, ship without accountability, and tolerate breakages as an acceptable trade-off.

      Your data ecosystem is optimized for speed over reliability, manual firefighting over prevention, and short-term fixes over long-term quality. If high-quality data were truly the goal, the system would have built-in schema enforcement, automated validation, and clear ownership—but since those don’t exist (or are routinely bypassed), the real function of the system is to allow chaotic, ad-hoc data handling that prioritizes short-term delivery over long-term trust.

      If Conway’s Law helps explain how data got into such a sorry state, POSIWID explains why - the broader organization optimizes for manual effort over proactive, automated, comprehensive solutions.

      Federation Ate the World

      The early 2000s marked a fundamental shift in how software engineering teams were structured. As technology companies scaled, they recognized that high-quality software required rapid iteration and continuous delivery. Research on software development lifecycles—such as the work popularized in the "Accelerate" book by Forsgren, Humble, and Kim—demonstrated that teams capable of shipping frequently could identify and fix defects faster, improve reliability, and respond to business needs with greater agility. The more frequently features were pushed into production, the sooner feedback loops closed, leading to better user experiences and stronger business outcomes.

      To facilitate this velocity, companies embraced Agile methodologies, which dismantled the traditional, slow-moving hierarchical structures and replaced them with small, autonomous, cross-functional teams. Rather than requiring months or years of deliberation, these teams operated with localized decision-making authority, allowing them to experiment, iterate, and ship software much faster. This shift not only optimized for speed but also reduced the coordination overhead that had historically slowed large engineering organizations.

      Out of this decentralized model emerged federated engineering structures and the adoption of microservices architectures. Instead of monolithic applications where multiple teams shared responsibility for a single massive codebase, companies transitioned to a world where individual teams owned their services, databases, infrastructure, and deployment pipelines. Each team was empowered to make locally optimal decisions—choosing their own programming languages, data models, and release schedules, all in the name of speed.

      The trade-off, however, was that many centralized cost centers—teams and functions designed for a monolithic, tightly controlled architecture—struggled to adapt. Operations teams, for example, were historically responsible for managing deployments in a centralized, controlled manner. With a monolithic system, they could plan releases, monitor performance, and enforce best practices through well-established governance processes. But in a federated world, visibility disappeared. Now, hundreds of teams were shipping thousands of changes independently, overwhelming centralized ops teams who could no longer track, validate, or mitigate risk effectively.

      This same dynamic played out in the world of data. Historically, data teams had ownership over the organization’s entire data architecture—curating data models, defining schema governance, and managing a centralized data warehouse. But as engineering teams began making independent decisions about which events to log, what databases to use, and how to structure data, the once-cohesive data ecosystem fragmented overnight.

      Without centralized oversight, engineering teams optimized for their immediate needs rather than long-term data quality. Events were collected inconsistently, naming conventions varied wildly, and different teams structured their data models based on what was most convenient for their service, rather than what was best for the organization as a whole. This led to massive data silos, duplicated efforts, and an overall decline in data consistency.

      The response from data teams? The Data Lake.

      Instead of trying to force governance onto hundreds of independent teams, companies adopted a "dump now, analyze later" approach, instructing engineering teams to send all their raw data into a centralized cold storage repository. This led to the rise of data engineering, a discipline that emerged to clean, transform, and organize this messy, unstructured data into something usable. Data engineers became reactive firefighters, constantly wrangling broken schemas, cleaning up unexpected transformations, and trying to reconstruct meaning from fragmented event logs.

      This model was deemed acceptable by business leaders because it allowed engineering teams to move quickly, even though it meant the data team was perpetually stuck in a reactive mode.

      In the early days of the cloud, this reactive data engineering model was sufficient. Most organizations primarily used data for dashboarding and reporting, where occasional inconsistencies could be tolerated. But as the industry evolved, the stakes for data reliability grew exponentially.

      • Machine Learning & AI: With AI-driven decision-making, poor data quality no longer just caused bad reports—it directly impacted product functionality and user experience. A mislabeled dataset could lead to a faulty recommendation algorithm, an unreliable fraud detection system, or an inaccurate pricing model.
      • Data as a Revenue-Generating Product: Companies began monetizing data directly—either by selling insights, building customer-facing analytics, or enabling real-time personalization. Inaccurate data now had a direct impact on revenue.
      • Regulatory Compliance & Risk: As GDPR, CCPA, and other data protection regulations took effect, bad data practices became a legal liability. A single oversight—such as failing to properly delete user data upon request—could result in multimillion-dollar fines.

      With these shifts, the consequences of reactive data engineering became untenable. Data teams could no longer afford to be downstream janitors, constantly cleaning up after engineering decisions made without governance. Instead, something fundamental had to change.

      The federated model of software engineering isn’t going away—if anything, it has only expanded. However, as we’ve seen, the decentralization of engineering cannot come at the cost of operational visibility, data integrity, and compliance. Organizations now face a critical inflection point:

      • How do we reintroduce governance without reintroducing bottlenecks?
      • How do we enable engineering speed while ensuring data correctness and compliance?
      • How do we prevent reactive firefighting and create proactive, self-service data management?

      Just as DevOps introduced infrastructure as code to solve the challenges of federated operations, the next era of data engineering must be proactive, automated, and deeply integrated into the software development lifecycle. Federation ate the world. But now we must decide how to rebuild it—this time, with sustainability, accountability, and resilience at its core.

      Shifting Left

      As the cloud, microservices, and decoupled engineering teams grew the centralized model of cost center management became harder to maintain and justify. The concept of Shifting Left emerged as a mechanism of driving ownership across a decoupled engineering organization and ultimately became the go-to solution for developers. In our context, shifting left is designed to help data teams overcome the people, process, and cultural challenges created by gaps in communication around data management. Instead of data management solely being the responsibility of the downstream data organizations, the treatment of data is a shared responsibility across producers, data platform teams, and consumers.

      Simply put: Shifting Left means moving ownership, accountability, quality and governance from reactive downstream teams, to proactive upstream teams.

      While Shifting Left may sound too good to be true, this pattern has happened on three notable occasions in software engineering. The first is DevOps, second is DevSecOps, and third is Feature Management.

      DevOps first emerged as a concept between 2007-2008 meant to address the growing gaps between IT teams (Dev) and Operations (ops). Before DevOps, software development and IT operations worked in silos. Developers would write code and pass it to operations teams, who were responsible for deployment and maintenance. This led to:

      • Slow releases due to hand-offs and bottlenecks.
      • Frequent deployment failures caused by differences between development and production environments.
      • Blame culture, where development blamed operations for slow deployments, and operations blamed development for unstable code.

      With DevOps, IT teams have become more agile, automated, and collaborative. Development and operations now work closely together, using continuous integration and continuous deployment pipelines to automate testing and deployment, reducing human errors and accelerating release cycles. Infrastructure as Code (IaC) allows teams to manage infrastructure programmatically, ensuring consistency and scalability. Monitoring, logging, and observability tools provide real-time insights, enabling proactive issue resolution rather than reactive firefighting.

      These days it is incredibly rare to see an engineering organization operating at a meaningful scale without a DevOps function. Most developers in a company are responsible for writing their own unit and integration tests. Teams rally around version control tools like Github and GitLab, both for collaboration, code review, auditing, and more.

      Roughly 10 years later in 2015, we saw a similar pattern in DevSecOps. Security teams were reactive -dealing with fraud and hacking after the fact rather than taking proactive and preventative steps to ensure software was designed with security in mind. Like Ops teams, the Security organization was siloed and disconnected from value, and as a cost center, suffered from the same problems as operations.

      DevSecOps is more complex than simply integrating security into existing DevOps workflows because it requires security to be automated, continuous, and developer-friendly—something traditional security practices were not designed for. Unlike traditional security, which was often applied as a final step before deployment, DevSecOps embeds security checks throughout the entire software development lifecycle. This shift introduces several challenges:

      • Shift-Left Security Requires Developer Buy-In: Security teams normally operated as gatekeepers, reviewing code and infrastructure late in the process. DevSecOps requires developers to take ownership of security much earlier meaning they need security tools that are easy to use, fast, and integrated into their existing workflows. However, many security tools were designed for security experts, not developers.
      • Balancing Security and Speed: DevOps emphasizes fast, frequent releases, while security traditionally slows things down with rigorous reviews and manual testing. DevSecOps must balance both, requiring automation that can enforce security without blocking deployments. Achieving this requires integrating automated security scanning, policy enforcement, and runtime protection into CI/CD pipelines without causing excessive friction.
      • Automated Security Testing at Scale: Traditional security relied on periodic manual testing (e.g., penetration testing, compliance audits). DevSecOps requires continuous security testing, including:
        • Static Application Security Testing for code vulnerabilities.
        • Software Composition Analysis for third-party dependencies.
        • Dynamic Application Security Testing for runtime security risks.
        • Infrastructure as Code Scanning to prevent misconfiguration.
        • Secrets and Credential Scanning to detect exposed sensitive data.
        • Integrating these into DevOps pipelines without overwhelming teams with false positives is a major challenge.

      The overall takeaway? The more complex or multi-component a cost center’s workflows, the more sophistication is required to effectively shift left while managing the delicate balance of developer expectations, speed and accuracy, plus scale.

      And finally, the Shift Left has happened with Feature Management. Traditionally, feature rollouts, experiments, and instrumentation were handled late in the development cycle—often by downstream teams like product, analytics, or growth. This led to:

      • Engineers shipping features without proper instrumentation or user tracking.
      • Product teams struggling to get clean data on feature performance post-launch.
      • A/B testing requiring significant engineering support, slowing experimentation velocity.
      • Limited ability to control or roll back features without a full redeploy.

      By shifting Feature Management into the software development lifecycle, teams can build observability, experimentation, and rollout controls directly into the feature itself. Feature Flags allow engineers to ship code behind toggles, enabling controlled rollouts and fast reversions without redeployments. Instrumentation and product analytics are now added as part of the development process, not as a follow-up task. And experimentation frameworks are increasingly embedded into the codebase, letting product teams test and iterate without waiting on engineering.

      Just as DevOps brought deployment and infrastructure closer to development, Feature Management brings experimentation, rollout control, and measurement upstream—making feature delivery safer, faster, and more data-driven.

      All three of these disciplines follow the exact same pattern: A critical business function is siloed downstream (Ops, Security, Product/Growth). By pushing tools and methodologies to the left, it isn’t an incremental change in value, but an inversion of how the job is done. Experimentation becomes something that happens for every new feature deployment by default. Systems are secure by default. While every team is following their own maturity within these paths - many are more sophisticated in shifting left than others - it is not just theory. This has already happened, and we in the data space should stop speaking about what might happen, and more about what can.

      Shifting Data Left

      Unlike engineering and security, data was the last frontier in cloud migration. While cloud-native infrastructure transformed application development and security in the early 2010s, data teams lagged behind, facing unique and more complex challenges that made cloud adoption far more difficult.

      The primary reason for this delay lies in the inherent complexity and multi-faceted nature of data that makes the shift left incredibly complex, costly, and difficult to manage. Security, despite its wide operational scope, primarily deals with permissions, monitoring, and compliance enforcement within a codebase or infrastructure environment. Engineering, too, could migrate by lifting application workloads into cloud-hosted services. The product discipline had the easiest transition, given that front-end/full-stack engineers were already adding monitoring and instrumentation to their services with our without product managers asking for it. However, data does not exist in a single place, nor does it follow a single lifecycle. It moves across repositories, services, and storage technologies, often passing through multiple transformations before it can be used.

      A typical data workflow spans:

      1. Source data ingestion (from application databases, logs, APIs, and event streams).
      2. Storage across multiple environments (operational databases, data lakes, warehouses, object storage).
      3. ETL (Extract, Transform, Load) and ELT processes that modify and refine data.
      4. Aggregation into analytical databases or data warehouses.
      5. Further transformations inside the warehouse to clean, normalize, and structure data.
      6. Downstream consumption by dashboards, machine learning models, or data products.

      This multi-stage pipeline meant that migrating data to the cloud required far more than simply moving databases—it demanded rebuilding the entire data infrastructure stack from ingestion to transformation, storage, governance, and consumption. The cost, complexity, and dependencies across teams slowed down cloud adoption significantly.

      Now, 20 years into the cloud era, data teams are encountering the same organizational and technical bottlenecks that operations teams faced in the mid-2000s. Back then, software engineering moved to a decentralized, service-based model, which broke traditional operations workflows and required a complete rethink of deployment and monitoring strategies. Today, data teams face the same disaggregation problem—modern software development creates silos that fragment data, leading to inefficiencies and bottlenecks that restrict its flow and value within the organization.

      In a highly federated engineering environment, individual teams often manage their own databases without centralized coordination, emit event streams without consistent schema or semantic governance, and choose storage solutions optimized for local needs rather than global usability. These teams also tend to create ad hoc transformations that may duplicate or overwrite critical business logic.

      The result? Data integrity, consistency, and discoverability suffer. Much like operations before the rise of DevOps, data engineering has become a reactive cost center—constantly fixing inconsistent schemas and data drift, resolving duplicate or contradictory transformations, debugging downstream breakages caused by upstream changes, and responding to compliance incidents like untracked PII exposure.

      Just as DevOps emerged to address the chaos of decentralized operations, we now need a similar movement for data—one that rethinks how we approach governance, engineering, and automation. A Shift Left approach to data requires embedding quality, governance, and security at the source, not just patching issues downstream.

      This new data paradigm must deliver on several fronts:

      • Schema and contract enforcement at ingestion, to prevent breakages by validating structure at the point of creation.
      • Versioning and change management, applying DevOps principles to schema evolution and business logic to ensure traceability and control.
      • End-to-end lineage tracking, giving teams visibility into how data transforms across systems, helping them understand and reduce the blast radius of change.
      • Automated compliance enforcement, detecting and tagging sensitive data like PII or financial records at the source.
      • Observability and real-time monitoring, to catch anomalies and schema drift before they impact analytics or AI.

      Operationally, shifting data left also changes how teams work:

      • Engineering teams must own data quality, just as they own application reliability.
      • Governance must become proactive, driven by automation and policy enforcement, not just documentation.
      • Data contracts should be standardized to reduce fragmentation and ensure consistent expectations across teams.
      • Compliance checks must be embedded into CI/CD pipelines, mirroring how DevSecOps integrates security into the development lifecycle.

      The lessons from other shift-left approaches are clear: embedding quality, security, and governance at the source is the only scalable approach. The same applies to data. The future of data engineering is not about building bigger, more sophisticated reactive teams—it is about pushing responsibility upstream, empowering engineering teams, and enforcing quality at the point of data creation.

      Much like operations and security before it, data must shift left. Organizations that fail to adapt will face the same problems they always have, except this time the excuses of “its someone else’s problem won’t cut it when the success of a companies AI initiative is on the line.. Those that succeed will transform data into a true production-grade asset, enabling faster decision-making, higher reliability, and greater business value.

      Data as Code

      A core idea behind shifting Data Left is simple but often overlooked: data is code. Or more accurately—data is produced by code. It’s not just some downstream artifact that lives in tables and gets piped into dashboards and spreadsheets. Every record, event, or log starts somewhere—created, updated, or deleted by a line of code. And just like DevOps demonstrated, if you want to manage something well, you start at the point of creation.

      Imagine you're in charge of a machine that produces a high-value product every day. Your job is to ensure quality. If something starts breaking, you don’t sit around analyzing boxes of defective products—you inspect the machine. Code is the machine. It runs on rules, inputs, and constraints that result in some form of data—a CRM entry, an API payload, a Kafka event, a database write. Managing that data means managing the system that generates it.

      This is here it’s useful to separate data management into three different but interconnected layers:

      • DevOps is focused on the software development lifecycle—code, pipelines, and deployment.
      • Observability is about the data that’s already been produced—monitoring records, metrics, and aggregates.
      • Business glossaries operate at a higher level—covering domains, policies, compliance, and internal processes.

      Each of these layers has value, but they serve different personas and purposes. DevOps is the proactive software engineering layer; observability is the reactive data team layer; business glossaries are organizational scaffolding. If one of these layers is missing, it becomes incredibly difficult to connect the dots.

      For example: without data lineage, business processes can’t be tied back to any actual systems or datasets. Without code lineage, your data lineage is blind outside the warehouse—you have no idea where upstream data is coming from or what’s generating it.

      This is why data management patterns—catalogs, contracts, lineage, and monitors—shouldn’t be thought of as individual tools. They’re cross-cutting patterns that apply across personas:

      • Engineers need catalogs of code assets, source systems, and event producers.
      • Data teams need catalogs of tables, metrics, and dashboards.
      • Business teams need catalogs of domains, data products, and process workflows.

      Same goes for lineage, contracts, and monitors. It’s not enough to do these things in isolation—we need contract enforcement in CI/CD, not just in downstream pipelines. We need code-level lineage, not just column-level lineage. We need contract monitors that can detect schema and semantic breakage as early as possible.

      But it’s not just about capabilities—it’s about making sure those capabilities actually work for the right people. Effective data management requires alignment across three key groups: producers, consumers, and business teams. And each of these groups has very different needs.

      Before the rise of the modern data stack, we had legacy catalogs—manual systems designed for centralized data stewardship. These catalogs were maintained by data stewards who filled in process docs, definitions, and ownership by hand. That model worked when data governance was owned by a few people and everything moved slowly.

      Then modern data catalogs showed up—tools that connected directly to warehouses like Snowflake and Databricks and scanned tables, dashboards, and metrics. They gave data teams a lot more visibility into the artifacts they worked with day-to-day. But they also ran into friction with software engineering teams. The problem? These catalogs weren’t built for engineers. They didn’t expose anything about the code that produces data, the services that emit events, or the systems that control data generation. And when a tool doesn’t map to your responsibilities, it doesn’t get adopted.

      Engineers want to understand how the code they own creates data, where that data goes, and what depends on it. They care about breaking changes, contract violations, and runtime errors—but none of that is visible in warehouse-first tooling. A table in Snowflake doesn’t tell you what GitHub repo created it or what line of code owns the transformation logic. So engineers are left in the dark, and downstream teams are left managing the fallout.

      To bridge that gap, we need to bring in techniques that have worked elsewhere—DevOps, CI/CD, and even security. Software composition analysis gives us a model for understanding dependencies. Dataflow analysis gives us insight into how code transforms data. CI pipelines can block bad changes before they reach production. Contract tests can catch mismatches between producers and consumers early. The point is: we already know how to solve this in software—we just need to apply those lessons to data.

      Until we do, data will remain something engineers generate but don’t own—leaving the rest of the organization to clean up the mess downstream.

      The number one question I get is: “Chad, this all sounds great—and conceptually we’re on board—but how? Where do we start?

      The answer: find allies on the software engineering team—specifically, those who already think in terms of quality, contracts, and validation. One of the best places to start? QA engineers and automated testing teams. They’re on a shift-left journey of their own, working to push testing and validation closer to the source of truth: the code. They already use a familiar concept—contract testing—to enforce expectations between APIs. This process defines the structure and behavior of communication between systems before runtime. Sound familiar? That’s a data contract.

      This concept can—and should—be extended to cover all data ingress and egress points: where data enters a system, and where it leaves. Just like APIs, data systems need contracts at the edges. These contracts should validate everything from schema shape to semantic meaning. But more importantly, they shouldn’t just observe the data after it's been produced—they should observe the code that produces it. We need systems that catch breaking changes at the source, during development, not days later in a downstream dashboard.

      A system like that doesn’t just help QA—it scales to every engineer. It defines clear ownership for producers and gives data teams hooks into the creation process instead of chasing quality issues downstream. It embeds data quality inside code quality.

      And here’s the mindset shift: most data quality problems are code quality problems. There are really two types.

      1. A lapse in judgment—an engineer skips writing a test and a bug slips through.
      2. A broken dependency—an engineer unknowingly changes something a downstream team relied on.

      The second category is where things get interesting. Think of a backend engineer changing an API that silently breaks the frontend. That’s not just a code quality issue—it’s a data quality issue. A schema changed. Expectations weren’t communicated. If engineers adopted data contracts to protect themselves, they’d also protect everyone else: analysts, ML teams, finance reporting, compliance—you name it.

      And this approach isn’t limited to traditional pipelines. AI is about to turn this problem up to eleven.

      Autonomous agents making code changes might work when isolated to a single codebase. But LLMs struggle with system-wide context—how one service depends on another, how a change in a producer might cascade across APIs, databases, and pipelines. As AI starts making changes across systems, data becomes the medium of communication—not just between humans and machines, but between AIs themselves. And with it comes a combinatorial explosion of data dependencies and breakages. Without contracts and shift-left enforcement, we’ll be flying blind into that complexity.

      What Happens When We Get This Right

      1. Data teams move upstream

      When data quality enforcement happens downstream, data teams are left cleaning up issues they didn’t create and don’t control. But when contracts, lineage, and validation happen at the code level, data teams become active participants in the creation of data—not just its consumers. They can influence modeling decisions, track where data originates, and establish shared accountability with engineering. Instead of writing Slack messages about broken dashboards, they’re writing rules that prevent them from breaking in the first place.

      2. Compliance becomes code

      Policies don’t scale when they live in wikis or checklists. But they do scale when they’re encoded directly into CI/CD. Contracts as code let organizations define standards—PII tagging, schema validation, retention rules—and then automatically enforce them wherever data is produced. This moves compliance from something reactive and manual to something continuous and automated. No more chasing teams down before audits—violations get caught at the pull request.

      3. Engineers get early feedback

      Software engineers are used to fast feedback loops. They expect test failures, linting errors, and contract violations to be caught before they merge code—not after something explodes in production. Data should work the same way. If a change to a schema or data payload is going to break an ML model or a reporting pipeline, the engineer should know before they hit merge. That kind of signal creates trust—and makes it easier for engineering to take ownership of data quality.

      4. Quality becomes a shared responsibility

      Right now, data quality is everyone’s problem and no one’s responsibility. But when data contracts are treated like API contracts, and quality issues are caught in dev, the responsibility naturally shifts to the teams closest to the cause. Engineers own their breakages, data teams own their transformations, and business teams finally get transparency into what’s reliable. Everyone’s incentives align. Quality isn’t something you inspect later—it’s something you build from the start.

      5. The language changes

      When data issues are framed as code issues, the conversation changes. Engineers don’t have to learn new tools or vocabulary—they just see tests, contracts, and CI checks like they do in the rest of their codebase. And when the language is familiar, adoption skyrockets. Suddenly, “data quality” isn’t a data team problem—it’s a software engineering best practice. This is how data quality becomes something everyone owns—because it's finally framed in a language software teams understand.

      Conclusion

      You made it to the end. Thanks for the read, I know it was long. This is a subject I’m passionate about, and I believe that shifting left is the key to unlocking such a wide range of solutions and utility to data teams that it is hard to succinctly describe. The impact of this movement, in my opinion, will be equally as impactful on software development as the advent of DevOps. And in the age of AI where data matters more than ever, developing the systems and culture to manage data as code is more critical than ever.

      You may have noticed this article was light on details. That was intentional, for my own sanity. In the next article: The Engineering Guide to Shifting Data Left, we’ll get more hands on with real world examples, use cases, and implementations. Take care until then, and good luck.

      -Chad

      Chad Sanderson
      Gable.ai
      CEO & Co-Founder