Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Gathering the fundamental laws of physics

Mike's Notes

To test Pipi, I previously manually added a number of verified laws from science to constrain what Pipi can do, anchoring it in objective reality rather than probabilistic LLM hallucinations. All the tests were successful.

Now, as Pipi Core races to production, we need to provide Ajabbi Researcher Accounts with editing rights and an automated way to import and update the laws and constants used. 

All this will be made freely and publicly available in future on an Ajabbi website, along with references to research, publications, seminars and how and where they are being used.

So I'm looking at all options, and as usual, discovery will be by trial and error.

Starting with physics, there are;

  • Proven fundamental laws of physics
  • Physics constants with uncertanty/units
According to Google.

"If you need a much larger database (hundreds or thousands of physics rows), you can download these datasets and extract them to 

  • CSV:Kaggle's PhysicsFormulas Dataset: This is a curated dataset containing 400 real, well-known physics formulas structured specifically for analysis and databases. It is split cleanly into CSV rows containing categories, descriptions, and LaTeX representations. 
  • TheorIA (Theoretical Physics Dataset): A highly structured database repository mapped by theoretical physicists. It is natively stored in JSON format (one file per physics entry), which can be effortlessly flattened into a CSV using a quick Python script or an online JSON-to-CSV tool."

- Google Search AI Mode

The list of resource links below is from a quick Google Search AI Mode.

Many other laws and constants will also be added over time for use by Pipi and in the enterprise workspaces. Other fields include:

  • software engineering
  • biology
  • geology
  • electrical, mechanical and civil engineering
  • medicine
  • economics
  • etc.

Resources

References

  • The Cambridge Handbook of Physics Formulas, Graham Woan, University of Glasgow, Cambridge University Press. 2000.

Repository

  • Home > Ajabbi Research > Library > Subject > Physics
  • Home > Handbook > 

Last Updated

13/08/2026

Gathering the fundamental laws of physics

By: Mike Peters
On a Sandy Beach: 13/08/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

This is a start on gathering the laws and constants as reference data to load into databases. Important to use the correct terms. NIST should be a good source. It will be a good test of the system.

List of Constants from NIST

  • alpha particle mass
  • alpha particle mass energy equivalent
  • alpha particle mass energy equivalent in MeV
  • alpha particle mass in u
  • alpha particle molar mass
  • alpha particle relative atomic mass
  • alpha particle rms charge radius
  • alpha particle-electron mass ratio
  • alpha particle-proton mass ratio
  • Angstrom star
  • atomic mass constant
  • atomic mass constant energy equivalent
  • atomic mass constant energy equivalent in MeV
  • atomic mass unit-electron volt relationship
  • atomic mass unit-hartree relationship
  • atomic mass unit-hertz relationship
  • atomic mass unit-inverse meter relationship
  • atomic mass unit-joule relationship
  • atomic mass unit-kelvin relationship
  • atomic mass unit-kilogram relationship
  • atomic unit of 1st hyperpolarizability
  • atomic unit of 2nd hyperpolarizability
  • atomic unit of action
  • atomic unit of charge
  • atomic unit of charge density
  • atomic unit of current
  • atomic unit of electric dipole moment
  • atomic unit of electric field
  • atomic unit of electric field gradient
  • atomic unit of electric polarizability
  • atomic unit of electric potential
  • atomic unit of electric quadrupole moment
  • atomic unit of energy
  • atomic unit of force
  • atomic unit of length
  • atomic unit of magnetic dipole moment
  • atomic unit of magnetic flux density
  • atomic unit of magnetizability
  • atomic unit of mass
  • atomic unit of momentum
  • atomic unit of permittivity
  • atomic unit of time
  • atomic unit of velocity
  • Avogadro constant
  • Bohr magneton
  • Bohr magneton in eV/T
  • Bohr magneton in Hz/T
  • Bohr magneton in inverse meter per tesla
  • Bohr magneton in K/T
  • Bohr radius
  • Boltzmann constant
  • Boltzmann constant in eV/K
  • Boltzmann constant in Hz/K
  • Boltzmann constant in inverse meter per kelvin
  • characteristic impedance of vacuum
  • classical electron radius
  • Compton wavelength
  • conductance quantum
  • conventional value of ampere-90
  • conventional value of coulomb-90
  • conventional value of farad-90
  • conventional value of henry-90
  • conventional value of Josephson constant
  • conventional value of ohm-90
  • conventional value of volt-90
  • conventional value of von Klitzing constant
  • conventional value of watt-90
  • Copper x unit
  • deuteron g factor
  • deuteron magnetic moment
  • deuteron magnetic moment to Bohr magneton ratio
  • deuteron magnetic moment to nuclear magneton ratio
  • deuteron mass
  • deuteron mass energy equivalent
  • deuteron mass energy equivalent in MeV
  • deuteron mass in u
  • deuteron molar mass
  • deuteron relative atomic mass
  • deuteron rms charge radius
  • deuteron-electron magnetic moment ratio
  • deuteron-electron mass ratio
  • deuteron-neutron magnetic moment ratio
  • deuteron-proton magnetic moment ratio
  • deuteron-proton mass ratio
  • electron charge to mass quotient
  • electron g factor
  • electron gyromagnetic ratio
  • electron gyromagnetic ratio in MHz/T
  • electron magnetic moment
  • electron magnetic moment anomaly
  • electron magnetic moment to Bohr magneton ratio
  • electron magnetic moment to nuclear magneton ratio
  • electron mass
  • electron mass energy equivalent
  • electron mass energy equivalent in MeV
  • electron mass in u
  • electron molar mass
  • electron relative atomic mass
  • electron to alpha particle mass ratio
  • electron to shielded helion magnetic moment ratio
  • electron to shielded proton magnetic moment ratio
  • electron volt
  • electron volt-atomic mass unit relationship
  • electron volt-hartree relationship
  • electron volt-hertz relationship
  • electron volt-inverse meter relationship
  • electron volt-joule relationship
  • electron volt-kelvin relationship
  • electron volt-kilogram relationship
  • electron-deuteron magnetic moment ratio
  • electron-deuteron mass ratio
  • electron-helion mass ratio
  • electron-muon magnetic moment ratio
  • electron-muon mass ratio
  • electron-neutron magnetic moment ratio
  • electron-neutron mass ratio
  • electron-proton magnetic moment ratio
  • electron-proton mass ratio
  • electron-tau mass ratio
  • electron-triton mass ratio
  • elementary charge
  • elementary charge over h-bar
  • Faraday constant
  • Fermi coupling constant
  • fine-structure constant
  • first radiation constant
  • first radiation constant for spectral radiance
  • Hartree energy
  • Hartree energy in eV
  • hartree-atomic mass unit relationship
  • hartree-electron volt relationship
  • hartree-hertz relationship
  • hartree-inverse meter relationship
  • hartree-joule relationship
  • hartree-kelvin relationship
  • hartree-kilogram relationship
  • helion g factor
  • helion magnetic moment
  • helion magnetic moment to Bohr magneton ratio
  • helion magnetic moment to nuclear magneton ratio
  • helion mass
  • helion mass energy equivalent
  • helion mass energy equivalent in MeV
  • helion mass in u
  • helion molar mass
  • helion relative atomic mass
  • helion shielding shift
  • helion-electron mass ratio
  • helion-proton mass ratio
  • hertz-atomic mass unit relationship
  • hertz-electron volt relationship
  • hertz-hartree relationship
  • hertz-inverse meter relationship
  • hertz-joule relationship
  • hertz-kelvin relationship
  • hertz-kilogram relationship
  • hyperfine transition frequency of Cs-133
  • inverse fine-structure constant
  • inverse meter-atomic mass unit relationship
  • inverse meter-electron volt relationship
  • inverse meter-hartree relationship
  • inverse meter-hertz relationship
  • inverse meter-joule relationship
  • inverse meter-kelvin relationship
  • inverse meter-kilogram relationship
  • inverse of conductance quantum
  • Josephson constant
  • joule-atomic mass unit relationship
  • joule-electron volt relationship
  • joule-hartree relationship
  • joule-hertz relationship
  • joule-inverse meter relationship
  • joule-kelvin relationship
  • joule-kilogram relationship
  • kelvin-atomic mass unit relationship
  • kelvin-electron volt relationship
  • kelvin-hartree relationship
  • kelvin-hertz relationship
  • kelvin-inverse meter relationship
  • kelvin-joule relationship
  • kelvin-kilogram relationship
  • kilogram-atomic mass unit relationship
  • kilogram-electron volt relationship
  • kilogram-hartree relationship
  • kilogram-hertz relationship
  • kilogram-inverse meter relationship
  • kilogram-joule relationship
  • kilogram-kelvin relationship
  • lattice parameter of silicon
  • lattice spacing of ideal Si (220)
  • Loschmidt constant (273.15 K, 100 kPa)
  • Loschmidt constant (273.15 K, 101.325 kPa)
  • luminous efficacy
  • magnetic flux quantum
  • molar gas constant
  • molar mass constant
  • molar mass of carbon-12
  • molar Planck constant
  • molar volume of ideal gas (273.15 K, 100 kPa)
  • molar volume of ideal gas (273.15 K, 101.325 kPa)
  • molar volume of silicon
  • Molybdenum x unit
  • muon Compton wavelength
  • muon g factor
  • muon magnetic moment
  • muon magnetic moment anomaly
  • muon magnetic moment anomaly
  • muon magnetic moment to Bohr magneton ratio
  • muon magnetic moment to nuclear magneton ratio
  • muon mass
  • muon mass energy equivalent
  • muon mass energy equivalent in MeV
  • muon mass in u
  • muon molar mass
  • muon-electron mass ratio
  • muon-neutron mass ratio
  • muon-proton magnetic moment ratio
  • muon-proton mass ratio
  • muon-tau mass ratio
  • natural unit of action
  • natural unit of action in eV s
  • natural unit of energy
  • natural unit of energy in MeV
  • natural unit of length
  • natural unit of mass
  • natural unit of momentum
  • natural unit of momentum in MeV/c
  • natural unit of time
  • natural unit of velocity
  • neutron Compton wavelength
  • neutron g factor
  • neutron gyromagnetic ratio
  • neutron gyromagnetic ratio in MHz/T
  • neutron magnetic moment
  • neutron magnetic moment to Bohr magneton ratio
  • neutron magnetic moment to nuclear magneton ratio
  • neutron mass
  • neutron mass energy equivalent
  • neutron mass energy equivalent in MeV
  • neutron mass in u
  • neutron molar mass
  • neutron relative atomic mass
  • neutron to shielded proton magnetic moment ratio
  • neutron-electron magnetic moment ratio
  • neutron-electron mass ratio
  • neutron-muon mass ratio
  • neutron-proton magnetic moment ratio
  • neutron-proton mass difference
  • neutron-proton mass difference energy equivalent
  • neutron-proton mass difference energy equivalent in MeV
  • neutron-proton mass difference in u
  • neutron-proton mass ratio
  • neutron-tau mass ratio
  • Newtonian constant of gravitation
  • Newtonian constant of gravitation over h-bar c
  • nuclear magneton
  • nuclear magneton in eV/T
  • nuclear magneton in inverse meter per tesla
  • nuclear magneton in K/T
  • nuclear magneton in MHz/T
  • Planck constant
  • Planck constant in eV/Hz
  • Planck length
  • Planck mass
  • Planck mass energy equivalent in GeV
  • Planck temperature
  • Planck time
  • proton charge to mass quotient
  • proton Compton wavelength
  • proton g factor
  • proton gyromagnetic ratio
  • proton gyromagnetic ratio in MHz/T
  • proton magnetic moment
  • proton magnetic moment to Bohr magneton ratio
  • proton magnetic moment to nuclear magneton ratio
  • proton magnetic shielding correction
  • proton mass
  • proton mass energy equivalent
  • proton mass energy equivalent in MeV
  • proton mass in u
  • proton molar mass
  • proton relative atomic mass
  • proton rms charge radius
  • proton-electron mass ratio
  • proton-muon mass ratio
  • proton-neutron magnetic moment ratio
  • proton-neutron mass ratio
  • proton-tau mass ratio
  • quantum of circulation
  • quantum of circulation times 2
  • reduced Compton wavelength
  • reduced muon Compton wavelength
  • reduced neutron Compton wavelength
  • reduced Planck constant
  • reduced Planck constant in eV s
  • reduced Planck constant times c in MeV fm
  • reduced proton Compton wavelength
  • reduced tau Compton wavelength
  • Rydberg constant
  • Rydberg constant times c in Hz
  • Rydberg constant times hc in eV
  • Rydberg constant times hc in J
  • Sackur-Tetrode constant (1 K, 100 kPa)
  • Sackur-Tetrode constant (1 K, 101.325 kPa)
  • second radiation constant
  • shielded helion gyromagnetic ratio
  • shielded helion gyromagnetic ratio in MHz/T
  • shielded helion magnetic moment
  • shielded helion magnetic moment to Bohr magneton ratio
  • shielded helion magnetic moment to nuclear magneton ratio
  • shielded helion to proton magnetic moment ratio
  • shielded helion to shielded proton magnetic moment ratio
  • shielded proton gyromagnetic ratio
  • shielded proton gyromagnetic ratio in MHz/T
  • shielded proton magnetic moment
  • shielded proton magnetic moment to Bohr magneton ratio
  • shielded proton magnetic moment to nuclear magneton ratio
  • shielding difference of d and p in HD
  • shielding difference of t and p in HT
  • speed of light in vacuum
  • standard acceleration of gravity
  • standard atmosphere
  • standard-state pressure
  • Stefan-Boltzmann constant
  • tau Compton wavelength
  • tau energy equivalent
  • tau mass
  • tau mass energy equivalent
  • tau mass in u
  • tau molar mass
  • tau-electron mass ratio
  • tau-muon mass ratio
  • tau-neutron mass ratio
  • tau-proton mass ratio
  • Thomson cross section
  • triton g factor
  • triton magnetic moment
  • triton magnetic moment to Bohr magneton ratio
  • triton magnetic moment to nuclear magneton ratio
  • triton mass
  • triton mass energy equivalent
  • triton mass energy equivalent in MeV
  • triton mass in u
  • triton molar mass
  • triton relative atomic mass
  • triton to proton magnetic moment ratio
  • triton-electron mass ratio
  • triton-proton mass ratio
  • unified atomic mass unit
  • vacuum electric permittivity
  • vacuum magnetic permeability
  • von Klitzing constant
  • W to Z mass ratio
  • weak mixing angle
  • Wien frequency displacement law constant
  • Wien wavelength displacement law constant

Realworld AI Coding Agent Exercise

Mike's Notes

An open-source version of Virtuoso will be installed in the data centre to explore possibilities. Kingsley Uyi Idehen's articles are all fascinating. I first came across him via the Ontology Forum.

The original post on LinkedIn has more links.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

08/04/2026

Realworld AI Coding Agent Exercise

By: Kingsley Uyi Idehen
LinkedIn: 07/02/2026

Founder & CEO at OpenLink Software | Driving GenAI-Based AI Agents | Harmonising Disparate Data Spaces (Databases, Knowledge Bases/Graphs, and File System Documents).

This post explains how I used Claude Code (Pro level, powered by Opus 4.5) and Mistral Vibe (whenever Claude Code rate limits kicked in) to modernize the aesthetics of our uniquely powerful faceted search and browsing interface for knowledge graph exploration—essentially giving the UI around the core engine a facelift.

At OpenLink Software, we strongly believe that LLM-powered AI Agents are exactly the right tools for tackling this long-standing challenge—provided the RDF-based knowledge graph runs on a platform that isn’t constrained by dataset size. In other words, one that scales naturally in linear fashion—such as our Virtuoso multi-model platform for managing data spaces spanning databases, knowledge bases, filesystems, and APIs.

Situation Analysis

As with many aspects of RDF (Resource Description Framework), challenges in tooling creation often stem from general misunderstandings about the framework itself. RDF-based knowledge graph representation is one such area, and linear visualization is another.

In the case of linear visualization, the goal is to present the description of an entity of interest (the subject) along with its associated attributes—i.e., predicate–object pairings that express attribute names and values.

Compounding the difficulty is the fact that, although faceted search UI/UX patterns have long served as the conceptual foundation, implementing them at the scale typical of RDF-based knowledge graphs remains extremely challenging. This challenge is further amplified by the complexity of delivering such interfaces in HTML leveraging CSS and JavaScript.

Virtuoso's Faceted Search & Browsing System Workings

Fundamentally, this system allows you to perform text search, attribute-name lookup, or entity-identifier–based exploration across one or more knowledge graphs hosted in a Virtuoso instance. It provides a property-sheet–style interface that presents entities (subjects) alongside their associated attributes (relationship predicates) and values (objects).

Thanks to Linked Data principles, hyperlink-based denotation of entities, attributes, and values (optionally) creates a Web of data. This enables the same “click and explore” experience—also known as the follow-your-nose exploration pattern—that users enjoy when interacting with web pages through a browser.

Old System

Here are screenshots depicting the UI/UX for a simple search sequence along the following lines:

  • Text Search Input: Virtuoso.

  • Initial matches presented in a list, sorted by text score and entity rank (think: page rank for data).

  • Add filters by attributes such as "type" (rdf:type) and "name" (schema:name) where the name value is "Virtuoso".

  • Click on item from the result to obtain a description of Virtuoso via the entity description property sheet based page.

New System

Original resource: New Virtuoso Faceted Browser Showcase Screencast

Here's the same demonstration sequence, but experienced via the revamped UI/UX.

  • Text Search Input: Virtuoso.

  • Initial matches presented in a list, sorted by text score and entity rank (think: page rank for data).

  • Add filters by attributes such as "type" (rdf:type) and "name" (schema:name) where the name value is "Virtuoso".

  • Click on item from the result to obtain a description of Virtuoso via the entity description property sheet based page.

Additional Notable Faceted Search & Browsing Features

These include the following:

  1. Handling the fact that lots of attributes (thousands+) could be associated with an entity sources for a massive collection of source knowledge graphs during the filtering stage via a sticky scrollable paging control
  2. Handling the fact that a select entity description can also comprise lots of attributes (thousands+) via a stickly scrollable paging control
  3. Spreadsheet-like table (with resizable, moveable, and sort enabling columns) for handling query results from filtering or when presenting entity descriptions
  4. Ability to export the description of an entity in a variety of formats (JSON-LD, RDF-Turtle, RDF-XML, N-Triples, RDF/JSON, CSV, etc)
  5. Permalinking for sharing interaction state e.g., a filter page or entity description page
  6. Ability to reveal underling SPARQL query that drives filtering
  7. Metadata that provides information on source named graph(s) from which attributes and values have been sourced for an entity description by way of entity (subject) or value (object) role in the source graph triples
  8. Metadata that automatically identifies explicit (via owl:sameAs attribute values) coreferences (via values of attributes that are uniquely identifying i.e., inverse-functional e.g., email addresses or any other attribute with the inverse-functional designation in a loaded ontology)
  9. Settings for enabling or disabling reasoning and inference informed by built-in or custom inference rules

The Refactoring Process

I achieved this very difficult refactoring task, alongside my other daily duties, by prompting Claude Code and Mistral. Claude Code (using the Opus 4.5 model) handled most of the heavy refactoring and planning work through the initial working ports.

I brought Mistral on board once rate limits kicked in, which became an important part of the experiment—namely, determining what’s possible with the Pro edition of Claude Code.

That said, Mistral also impressed me to the point where I see it as the closest rival to Claude Code in the battle for coding agent market dominance. Its CLI aesthetics and assistance mode are top-notch.

Live Demonstration Instances

URIBurner

This is a live Virtuoso instance since 2008 functioning as a Linked Data utility showcase and bridge to the massive Linked Open Data Cloud Knowledge Graph collective.

  1. Text Search: Virtuoso
  2. Entity Types associated with text pattern: Virtuoso
  3. Attribute Filtering on Type, Name, and Value: Virtuoso Universal Server (Row Store & Cluster Server Edition)
  4. Selected Entity Description Page

Conclusion

Software development is evolving before our eyes. True power now comes from pairing capable AI Agents with human expertise—letting judgment guide automation, producing dependable outcomes, and delivering real-world value. The age of AI isn’t just about smarter tools; it’s about amplifying what humans do best.

Design for the People: The US Web Design System and the Public Sans Typeface

Mike's Notes

The article reproduced below is from a fascinating website and is about another useful Design System.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

16/03/2026

Design for the People: The US Web Design System and the Public Sans Typeface

By: Jon Keegan
On a Sandy Beach: 02/07/2024

Jon Keegan is an investigative data journalist who covers technology. His work has appeared in The Wall Street Journal, The Markup and MIT Technology Review. Jon’s work has won several journalism awards, including the Loeb Award, the Society of Professional Journalists’ Excellence in Journalism Award and the Society of News Design’s Best of Digital Gold Award.

The United States has an official web design system and a custom typeface that belongs to the people. This thoughtful public design system aims to make government websites not only look good, but to make them accessible and functional for all.

Before the internet, Americans may have interacted with the federal government by stepping into grand buildings adorned with impressive stone columns and gleaming marble floors. Today, the neoclassical architecture of those physical spaces has been replaced by the digital architecture of website design – HTML code, tables, forms, and buttons. 

While people visiting a government website to apply for student loans, research veterans’ benefits, or enroll in Medicare may not notice these digital elements, they play a crucial role. If a website is buggy or doesn’t work on your phone, taxpayers cannot access the services they have paid for. This can feel like walking up to a boarded-up government building with broken windows, creating a negative impression of the government itself.  

 In the US, there are about 26,000 federal websites. Early on, each site had its own designs, fonts, and login systems, creating frustration for the public, and wasting government resources.

 A survey of the many different styles of buttons from government websites as of 2015. Source: 18F / GSA 

The troubled launch Healthcare.gov in 2013 highlighted the need for a better way to build government digital services. In 2014, President Obama created two new teams to help improve government tech.

Within the General Services Administration (GSA), a new team called 18F (named for their Washington, DC office at 1800 F Street) was created to “collaborate with other agencies to fix technical problems, build products, and improve public service through technology.” The team was built to move at the speed of tech start-ups rather than lumbering bureaucratic agencies. 

The U.S. Digital Service (USDS) was tasked “to deliver better government services to the American people through technology and design.” In 2015, the two teams collaborated to build the US Web Design System (USWDS)—a style guide and collection of user interface components and design patterns to ensure a consistent user experience across government websites. “Inconsistency is felt, even if not always precisely articulated in usability research findings,” said Dan Williams, the USWDS program lead, in an email. 

Some of the sample design elements for the USWDS. Source: https://designsystem.digital.gov/

Today, the system defines 47 user interface components such as buttons, alerts, search boxes and forms each with their own design examples, sample code and guidelines such as “Be polite” and “Don’t overdo it.” The USWDS is now in its third iteration, and is used in 160 government websites. “As of September 2023, 94 agencies use USWDS code, and it powers about 1.1 billion pageviews on federal websites,” said Williams.

USWDS design principles include focusing on real users’ needs, earning trust and embracing accessibility. The system requires websites to be optimized for all users, including people with disabilities such as those using screen readers or those with color blindness. Williams said accessibility is important to the team’s efforts, noting that they “prioritize any accessibility-related bug or improvement we find (or is contributed by our community).”






Some federal websites that use the USWDS. Clockwise from top left: Va.gov, Medicaid.gov, Worker.gov, Supremecourt.gov

To ensure clear and consistent typography, the free and open-source typeface Public Sans was created for the US government. “It started as a design experiment,” said Williams, who designed the typeface, which was released in 2019. “We were interested in trying to establish an open source solution space for a typeface, just like we had for the other design elements in the design system,” said Williams. Based on the Libre Franklin typeface, Public Sans is described as “a strong, neutral, principles-driven, open-source typeface for text or display.” 


Both Public Sans and the USWDS embrace transparency and collaboration with government agencies and the public, inviting contributions to their development via the projects’ GitHub pages. 

To ensure that the hard-learned lessons of improving public technology aren’t forgotten, the projects embrace continuous improvement. One of Public Sans’ design principles offers key guidance in this area: “Strive to be better, not necessarily perfect.”

The Dark Data Tax: How Hoarding is Poisoning Your AI

Mike's Notes

This is a very big problem. Unstructured data.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Data Engineering Weekly
  • Home > Handbook > 

Last Updated

12/12/2025

The Dark Data Tax: How Hoarding is Poisoning Your AI

By: Ananth Packkildurai
Data Engineering Weekly: 19/11/2025

Ananth Packkildurai is a data engineering leader, writer, and author of Data Engineering Weekly, sharing insights on modern data platforms, large-scale pipelines, and AI-driven architectures.

Storage is cheap. Attention is finite. Hallucinations are expensive. It’s time to stop building Data Lakes and start managing Data Metabolism.

With the increased adoption of the Lakehouses, we removed the last constraint on data accumulation.

We didn’t realize we were removing the last constraint on data obesity. The numbers are staggering. Enterprises now store 2.5 times more data than they did in 2019, yet the velocity of decisions derived from that data hasn’t just slowed—it has flatlined. According to IDC, global data storage capacity is estimated to reach 175 zettabytes by 2025, with 80% of that data unstructured. Furthermore, IDC predicts that 90% of unstructured data will remain unanalyzed.

This is data obesity: the condition where an organization accumulates data faster than it can derive value from it. It’s not a storage problem. It’s a metabolic one.

When Storage Became Infinite, Attention Became Finite

The obesity crisis began with the Lakehouse. Built on the triumvirate of S3, ADLS, and GCS, and crowned with Delta Lake, Iceberg, and Hudi, the Lakehouse solved data engineering’s oldest constraint: where to put the data. Object storage made retention elastic and nearly free. The cost of a gigabyte has fallen by 80% over the past decade, while enterprise data volume has grown by 250%.

However, the Lakehouse didn’t just lower costs—it removed the psychological barrier to data collection. When a terabyte costs less than a pizza, no one asks hard questions before ingesting. When schema evolution is automatic, there’s no migration friction to discourage table sprawl. When time travel promises infinite rollback, deletion feels like the destruction of potential value.

The result is a modern manifestation of Jevons’ Paradox: as storage became more efficient, our appetite for data expanded even more rapidly. We’ve built systems that can collect anything, but can’t measure whether that collection matters.

The Hidden Cost of Free Storage

Here’s what Lakehouse architecture diagrams don’t show: storage accounts for only 8% of the total cost of ownership. The remaining 92% is the human and computational effort to transform bytes into decisions—costs that scale with complexity, not volume.

This is the first symptom of data obesity: operational debt. Dark datasets don’t just sit there; they demand maintenance, schema migrations, and compliance audits. They clog your information schema, slow down your metastore, and multiply the blast radius of every infrastructure change.

The LLM Hunger Games: When Predators Can’t Hunt

Datology AI’s seminal research demonstrated that model performance doesn’t scale with dataset size—it scales with signal density.

Their experiments demonstrated that a curated 100TB corpus consistently outperformed a raw 1PB dataset, resulting in a 40% reduction in training time and a 35% reduction in inference costs. Redundant, inconsistent, or low-quality data not only slows training but also actively degrades performance through gradient noise.

Despite this evidence, enterprise adoption of LLMs has amplified data obesity rather than containing it. The prevailing architectural pattern—vectorizing entire document corpora and connecting models through RAG pipelines—assumes maximal data exposure is optimal. This leads to embedding every PDF, support ticket, and communication archive without evaluating the utility of the information.

The result is a predator that consumes indiscriminately but digests inefficiently. When an LLM retrieves three conflicting definitions of “active user” from dark, ungoverned partitions, the model doesn’t resolve the conflict—it synthesizes them into authoritative-sounding errors. The 68% of dark data that plagues analytics now introduces hallucination vectors into AI systems.

The Dark Data Tax: Internal Data Poisoning

Consider the technical mechanics. An AI agent tasked with “analyzing customer churn” might query 200 tables across your Lakehouse, resulting in massive compute costs. But if 80% of those tables are dark—poorly documented, schema-drifted, or semantically redundant—the agent spends valuable inference cycles disambiguating rather than analyzing.

This is where the danger lies. Recent research on model robustness (such as findings from Anthropic) indicates that even a small fraction of incoherent data in the context window can disproportionately degrade output quality.

Dark data effectively acts as internal data poisoning:

  • Context Contamination: When an LLM retrieves three conflicting definitions of “active user”—one from a live table and two from dark, ungoverned tables—it lacks the context to discern the truth.
  • The Hallucination Vector: Instead of ignoring the noise, the model attempts to reconcile it, averaging the conflicting signals into authoritative-sounding nonsense.

This creates a severe form of cognitive debt. Dark data doesn’t just waste storage; it actively sabotages the insights derived from your good data. By introducing these “hallucination vectors,” you are effectively feeding your predator garbage, and the resulting insights are not just expensive—they are fundamentally compromised.

If data poisoning is the pathology of the individual model, we need a different framework to understand the collapse of the entire ecosystem. Ecology solved this problem a century ago.

The Predator-Prey Framework: Diagnosing Obesity

To transition from vague awareness to precise diagnosis, we require a model that captures how data systems become imbalanced. Ecology solved this problem a century ago.

In the 1920s, Alfred Lotka and Vito Volterra developed elegant equations that described predator-prey dynamics, explaining why ecosystems collapse when one species outnumbers the other. The model maps uncannily onto data systems:

Data Volume (x) = Prey. Reproduces through ingestion, replication, and instrumentation.

Data Value (y) = Predators. Analysts, models, and decision pipelines that rely on data to thrive.

The equations:

dx/dt = αx - βxy

dy/dt = δxy - γy

Where:

  • α = Prey growth rate (your ingestion velocity)
  • β = Predation rate (how easily analysts discover and query data)
  • δ = Conversion efficiency (how well analytics become decisions)
  • γ = Value decay (how quickly insights become stale)

The Emergence of Dark Data

Dark data isn’t just “unused data”—it’s the prey population that grows unchecked when predation fails to occur. In the equation dx/dt = αx - βxy, the term βxy represents data consumption and conversion into a value. When this term is too small relative to αx, the system experiences prey overpopulation.

Mathematically, dark data is the excess when:

αx >> βxy → dx/dt ≈ αx

In plain terms, your ingestion rate (α) is high, but your effective consumption rate (β) is low. The xy term depends on both available data (x) and active predators (y). When you have 10,000 tables but only 200 active analysts, even if each analyst works at maximum capacity, βxy cannot offset αx.

The Predator Starvation Crisis

The second equation, dy/dt = δxy - γy, predicts the consequence. When dark data grows, it doesn’t just increase x—it contaminates xy. The δ term (conversion efficiency) drops because analysts spend more time navigating noise. The γ term (value decay) increases because insights become obsolete more quickly in a chaotic environment.

The result: dy/dt becomes negative—your organization’s ability to create value from data actually declines, even as you add more data.

This framework transforms dark data from a storage problem into an ecological pathology. Your data isn’t just unused—it’s overpopulated prey, starving your predators.

The Data Sustainability Index: Measuring Metabolic Health

If we use predator-prey dynamics to diagnose data obesity, we need a metric to measure its severity. We cannot manage what we do not measure.

The Data Sustainability Index (DSI) metric quantifies the metabolic efficiency of your data ecosystem—measuring how much energy (compute) your data generates relative to the weight (cost & complexity) it adds.

The Formula

To calculate this, we track three specific variables:

The Predator Activity: Total Analytical Compute Hours (TACH)

This measures consumption, not storage. It captures every SQL warehouse second, Spark job, BI refresh, and model training cycle.

  • Why it matters: TACH is a proxy for demand. Even a failed query or an inefficient scan counts as “positive” here because it signals that a predator is attempting to feed.
  • The Rule: If TACH = 0, the dataset is biologically dead.

The Environmental Drag: Lakehouse Total Cost

We must account for the full burden of the ecosystem.

  • The Calculation: Storage + Ingestion Compute + Governance Overhead.
  • The Hidden Tax: Don’t just count cloud credits. A dataset that costs $500/month to store but requires 5 hours of engineering maintenance carries a massive “cognitive weight” that must be factored in.

The Punishment Multiplier: Active Dataset Ratio

This is the most critical component. It penalizes hoarding.

Formula: (Datasets with TACH > 0) / (Total Datasets)

If you have 10,000 tables but only query 1,200 of them, your Active Ratio is 12%. This mathematically crushes your DSI score, reflecting the reality that the 8,800 dark tables are creating noise that makes it harder to find the 12% of valuable data.

The Scorecard: How Healthy Are You?

Once you calculate your DSI, the results typically fall into three zones:

The Trajectory Warning: The absolute number matters less than the trend. If your DSI is declining month over month, it means your ingestion rate is outpacing your consumption rate. You aren’t scaling—you’re bloating.

Four Symptoms of Data Obesity

Dark data is the visible symptom of systemic metabolic failure. It manifests as four interlocking debts:

Operational Debt: The Cost of Unused Infrastructure

Unused datasets still demand pipeline maintenance, schema migrations, and dependency tracking. The obesity wasn’t just storage; it was engineering time diverted from value creation to waste management.

Cognitive Debt: The Tax on Every Decision

Analysts don’t just search for data—they sift through noise. When a data scientist needs “user_events” and finds 23 variations, the search cost becomes a tax on every decision. Studies in cognitive load theory show that decision quality degrades exponentially with choice overload. In obese ecosystems, analysts spend 40% of their time on data discovery and validation rather than analysis.

Compliance Risk: The Weight of Regulatory Exposure

GDPR and CCPA don’t care if you use personal data—only that you store it. A healthcare provider faced a $12M fine for retaining patient records in a “temporary” Lakehouse partition that had been dark for three years. Dark data is invisible to analytics but blindingly visible to auditors. Each unused dataset with PII is a latent liability.

Cultural Drift: The Normalization of Hoarding

Teams conflate “data-driven” with “data-hoarding.” Deleting a dataset feels riskier than keeping it because the cost of being wrong about deletion is visible (blame), while the cost of retention is invisible (slow decay). This creates a tragedy of the data commons, in which individual rationality produces collective dysfunction.

The Future: Autonomous Metabolic Management

As we instrument obesity metrics, we can automate weight management. Imagine a Data Obesity Controller:

Input Streams:

  • Real-time TACH telemetry
  • Decision attribution logs
  • Cost per table
  • Schema drift alerts

Control Actions:

  • Auto-archive dark datasets (reduce mass)
  • Suggest semantic model improvements (increase metabolism)
  • Alert on insight decay (prevent metabolic slowdown)
  • Recommend dataset deprecation (remove obesity)

This is the logical evolution of DataOps. We’ve spent a decade automating data production. The next decade belongs to automating data health.

Conclusion: The Healthiest Ecosystem is the Leanest

The obesity crisis teaches us a harsh truth: the organizations that thrive won’t be those with the largest lakes, but those with the highest metabolic rates.

The predator-prey framework gives us the language to diagnose this. The DSI gives us the metric to measure it. But the real work is cultural: we must stop celebrating data mass and start celebrating data muscle.

Wikipedia Structured Contents

Mike's Notes

Good news from Kaggle and Wikimedia. An opportunity to get structured data.

"...

As part of Wikimedia's mission to make all knowledge freely accessible and useful, Wikimedia is publishing a beta version of its structured content on Kaggle in French and English. This release gives data scientists, researchers, and machine learning enthusiasts a new, streamlined way to explore and analyze this global information resource.

..." - Kaggle.com

Resources

References

  • Reference

Repository

  • Home > 

Last Updated

18/04/2025

Wikipedia Structured Contents

By: Wikimedia Enterprise Team
Wikimedia Enterprises: 16/04/2025

Wikimedia Enterprise has released a new beta dataset on Kaggle, featuring structured Wikipedia content in English and French. Designed with machine learning workflows in mind, this dataset simplifies access to clean, pre-parsed article data that’s immediately usable for modeling, benchmarking, alignment, fine-tuning, and exploratory analysis.

This release is powered by our Snapshot API’s Structured Contents beta, which outputs Wikimedia project data in a developer-friendly, machine-readable format. Instead of scraping or parsing raw article text, Kaggle users can work directly with well-structured JSON representations of Wikipedia content—making this ideal for training models, building features, and testing NLP pipelines.The dataset upload, as of 15 April 2025, includes high-utility elements such as abstracts, short descriptions, infobox-style key-value data, image links, and clearly segmented article sections (excluding references and other non-prose elements). Because all content is derived from Wikipedia, it is freely licensed under Creative Commons Attribution-Share-Alike 4.0 and the GNU Free Documentation License (GFDL), with some additional cases where public domain or alternative licenses may apply.

“As the place the machine learning community comes for tools and tests, Kaggle is extremely excited to be the host for the Wikimedia Foundation’s data. Kaggle is already a top place people go to find datasets, and there are few open datasets that have more impact than those hosted by the Wikimedia Foundation. Kaggle is excited to play a role in keeping this data accessible, available and useful." - Brenda Flynn, Partnerships Lead, Kaggle

As a beta release, this dataset is an invitation to explore, test, and improve. We welcome feedback, questions, and suggestions from the Kaggle community directly in the dataset’s discussion tab.

Get the Dataset

Access the dataset directly on Kaggle

About Kaggle

Kaggle is home to one of the world’s largest communities of machine learning practitioners, researchers, and data enthusiasts. With millions of users and an expansive ecosystem of datasets, notebooks, and competitions—including challenges like the Arc Prize—Kaggle provides an ideal environment for experimenting with open structured data like Wikimedia’s Structured Content. Whether you’re testing a new architecture, evaluating data quality, or building a pipeline from scratch, this Wikipedia dataset is ready to plug into your process.

More info at Google Blog