How Evolution Gave Us Free Will

Mike's Notes

This book explains how the evolution of living things gave rise to free will.

"... a shift toward thinking about the brain as a complex dynamical system with emergent properties that defy reduction to simple elements." - Nicole Rust Transmitter

"Kevin Mitchell is associate professor of genetics and neuroscience at Trinity College Dublin in Ireland. He studies the genetics of brain wiring and its relevance to variation in human faculties, psychiatric disease and perceptual conditions such as synesthesia. His current research focuses on the biology of agency and the nature of genetic and neural information.

Mitchell completed his Ph.D. at the University of California, Berkeley, studying the genetic instructions that direct the development of the nervous system in the fruit fly, and his postdoctoral work at the University of California, San Francisco and Stanford University, exploring the same topic in mice. He is the author of “Innate: How the Wiring of Our Brains Shapes Who We Are” and “Free agents: How Evolution Gave Us Free Will.” He also writes the Wiring the Brain blog and is on X (formerly known as Twitter) @WiringtheBrain." - Transmitter

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > The Transmitter
  • Home > Handbook > 

Last Updated

17/05/2025

Player One: An edited excerpt from ‘Free Agents — How Evolution Gave Us Free Will’

By: Kevin Mitchell
The Transmitter: 13/11/2023

In his new book, neuroscientist Kevin Mitchell argues that, despite his field’s mechanistic models of cognition, we are all “Player One” in the game of life, the authors of our own actions.

Are we the authors of our own stories? Or is our apparent freedom of choice really an illusion? These questions were brought home to me recently as I was watching my son play a video game — one where you wander around an open world, meeting interesting denizens of one type or another (and killing quite a few of them). As I watched, his character entered a tavern and approached the bartender, who offered a generic greeting. The game then threw up some options for things you could say in reply to get information about the prospects for fortune and glory in the surrounding territory.

In this exchange, my son’s possibilities for action were limited by the game, but he was really making choices among them, and these choices then affected how the conversation went and what would subsequently unfold. His decisions were based on his overall goal in the game, the tension between his goals of taking some immediate action or to keep exploring, his need to have enough information to make a decision with confidence, the risk of biting off more than he could chew and losing his hard-won stuff: All these considerations fed into the decisions he made. He had his reasons and he acted on them, just like you or I do every day, all day long.

The bartender, in contrast, was not making choices. He was a classic “non-player character,” an NPC. His responses were completely determined by his programming: He had no degrees of freedom. His actions were merely the inevitable outcome of a flow of electrons through the circuits of the game console, constrained by the rules encoded in the software. Even the more sophisticated NPCs in the game, including the monster that eventually caramelized my son’s avatar, were similarly constrained. The monster’s actions — even in the fast-moving melee — were determined by the software programming and mediated by the electronic components in the console.

Thus the NPCs only appear to be making choices. They’re not autonomous entities like us: They’re just a manifestation of lots of lines of code, implemented in the physical structure of the computer chips. Their behavior is entirely determined by the inputs they get and their preprogrammed responses. We, in contrast, are causes of things in our own right. We have agency: We make our own choices and are in charge of our own actions.

At least it seems that way. It certainly feels like we have “free will,” like we make choices, like we are in control of our actions. That’s pretty much what we do all day — go around making decisions about what to do. Some are trivial, like what to have for breakfast; some are more meaningful, like what to say or do in social or professional situations; and some are momentous, like whether to accept a job offer or a marriage proposal. Some we deliberate on consciously, and others we perform on autopilot—but we still perform them. Of course, our options may be more or less constrained (or informed) by all kinds of factors at any given moment, but generally we feel like the authors of our own actions.

And we interpret other people’s behavior in terms of their reasons for selecting different actions — their intentions, beliefs and desires that make up the content of their mental states. We constantly analyze each other’s motives and habits and character, looking for explanations and predictors of their behavior and the decisions they make. Why people act the way they do is ultimately the theme of most entertainment, from Dostoyevsky to Big Brother. All this rests on the view that we are not just acted on — we are actors. Things don’t just happen to us, in the way they happen to rocks or spoons or electrons: We do things.

The problem is that if you think about this view for too long, it becomes difficult to escape a discomfiting thought. After all, like the NPCs, our decisions, however complex they may be, are mediated by the flow of electrical ions through the circuits of our brains and thus are constrained by our own “programming,” by how our circuits are configured. Unless you invoke an immaterial soul or some other ethereal substance or force that is really in charge — call it spirit or simply mind, if you prefer — you cannot escape the fact that our consciousness and our behavior emerge from the purely physical workings of the brain.

There is no shortage of evidence for this from our own experience. If you’ve ever been drunk, for example, or even just a little tipsy, you’ve experienced how altering the physical workings of your brain alters your choices and the way you behave. There is a whole industry of recreational drugs — from caffeine to crystal meth — that people take because of the way that physically tweaking the brain’s machinery in various ways makes them feel and act. The ultimate consequence in some cases is addiction — perhaps the starkest example of how our actions can sometimes be out of our control.

And, of course, if the machinery of your brain gets physically damaged — as occurs with head injuries, strokes, brain tumors, neurodegenerative disorders and a host of other kinds of insults — or its function is impaired in other ways, as in conditions such as schizophrenia, depression and mania, then your ability to choose your actions may also be impaired. In some situations, the integrity of your very self may be compromised.

We all like to think that we are Player One in this game of life, but perhaps we are just incredibly sophisticated NPCs. Our programming may be complex and subtle enough to make it seem as if we are really making decisions and choosing our own actions, but maybe we’re just fooling ourselves. Perhaps “we” are just the manifestations of genetic and neural codes, implemented in biological rather than computer hardware. Perhaps we are the victims of a cruel joke, tragic figures in the grip of the Fates. As Gnarls Barkley sang, “Who do you, who do you, who do you think you are? Ha ha ha, bless your soul, you really think you’re in control.”

It’s hard not to look at the growing body of work from neuroscience and see only the machine at work. Driving this circuit or that one either directly causes an action or influences the cognitive operations that the animal — mouse or human or anything else — uses to decide between actions. If we were dissecting a robot in this way, we would apply engineering approaches to understand the kinds of information being processed, the control mechanisms configured into the different circuits and the computations that lead to one output or another. There does not seem to be any need for something like a mind in that discussion. There is no real need for life, for that matter.

If the circuits just work on physical principles, then who cares what the patterns of activity mean? Why does it matter what the mental content associated with a particular pattern of neural activity is, if it is solely the physical configuration of the circuitry that is going to determine what happens next? We may have set out, as neuroscientists, to explain how the workings of the brain generate or realize psychological phenomena, but we are in danger of explaining those phenomena away.

”We make decisions, we choose, we act — we are causal forces in the universe. These are the fundamental truths of our existence and absolutely the most basic phenomenology of our lives."

If the neuroscientists have it bad, pity the poor physicists, whose existential angst must run much deeper. Where neuroscientists can at least hold onto the view that the circuits in the brain are doing things (whether “you” are or not), some physicists claim that even that functionality is an illusion. After all, the brain is made of molecules and atoms that must obey the laws of physics, just like the molecules and atoms in any other bit of matter.

These small bits of matter are pushed and pulled by all the forces acting on them — gravity, electromagnetism, the so-called strong and weak nuclear forces that hold atoms together — and where each atom goes is fully determined by the way those interactions play out. These processes are no doubt complicated, as they would be in any system with so many atoms simultaneously acting on each other, and in practice how the system will evolve is unpredictable — but it is still all driven by the physics. Even at the lower levels of subatomic particles, how the system evolves is captured by the equations of quantum mechanics in a way that many would argue theoretically leaves no room for any other causes to be at play.

So, then, what does it matter what you are thinking? You cannot push the atoms in your brain around with a thought. You cannot override the fundamental laws of physics or exert some ghostly control over the basic constituents of matter. According to this view, the very idea of mental causation — of the content of your thoughts and beliefs and desires mattering in some way — is a naive superstition, a conceptual hangover inherited from philosophers like the famous dualist Rene Descartes.

I am not willing to give up on free will so easily. In this book I argue that we really are agents. We make decisions, we choose, we act — we are causal forces in the universe. These are the fundamental truths of our existence and absolutely the most basic phenomenology of our lives. If science seems to be suggesting otherwise, the correct response is not to throw our hands up and say, “Well, I guess everything we thought about our own existence is a laughable delusion.” It is to accept instead that there is a deep mystery to be solved and to realize that we may need to question the philosophical bedrock of our scientific approach if we are to reconcile the clear existence of choice with the apparent determinism of the physical universe.

But if we want to solve this mystery, humans are the absolute worst place to start. It is a truism in biology to say that nothing makes sense except in the light of evolution — and this is surely true of agency. Instead of trying to understand it in its most complex form, I go back to its beginnings and ask how it emerged, what the earliest building blocks were, and what the basic concepts should be. How can we think about things like purpose and value and meaning without sinking into mysticism or vague metaphor? I argue that we can do so by locating these concepts in simpler creatures and then following how they were elaborated over the course of evolution, increasing in complexity and sophistication as certain branches of life developed ever-greater autonomy and self-directedness.

Indeed, before tackling the question of free will in humans, we have a much more fundamental problem to solve. How can any organism be said to do anything? Most things in the universe don’t make choices. Most things — like rocks or atoms or planets — don’t do anything at all, in fact. Things happen to them, or near them, or in them, but they are not capable of action. But you are. You are the type of thing that can take action, that can make decisions, that can be a causal force in the world: You are an agent. And humans are not unique in this capacity. All living things have some degree of agency. That is their defining characteristic, what sets them apart from the mostly lifeless, passive universe. Living beings are autonomous entities, imbued with purpose and able to act on their own terms, not yoked to every cause in their environment but causes in their own right.

To understand how this could be, we have to go right back to the beginning, to the very origins of life itself. From the chemistry of rocks and hydrothermal vents — the chemistry of the evolving planet itself — life emerged as systems of interacting molecules, interlocked in dynamic patterns that became self-sustaining. The ones that most robustly maintained their own dynamic organization persisted, replicated, evolved. They became enclosed in a membrane — a tiny subworld unto themselves — exchanging matter and energy with their environment while protecting an internal economy and reconfiguring their own metabolism to adapt to changing conditions. They became autonomous entities, causally sheltered from the thermodynamic storm outside and selected to persist.

A new trick was invented: action, the ability to move or affect things out in the environment. Information became a valuable commodity, and mechanisms evolved to gather it from the environment. With that came the crude beginnings of value and meaning. Movement toward or away from various things out in the world became good or bad for the persistence of the organism. These responses were selected for and became wired into the biochemical circuitry of simple creatures.

As multicellular creatures evolved, a class of cells — neurons — emerged that specialized in transmitting and processing information. Initial circuits acted as internal control systems, designed to coordinate the various muscles or other moving parts of the multicellular animal and defining a repertoire of useful actions. At the same time, neurons coupled various sensory signals to specific actions in this repertoire, hardwiring adaptive instincts for approach or avoidance.

With the elaboration of the nervous system, this kind of pragmatic meaning eventually led to semantic representations. Perception and action were decoupled by layers of intervening cells. Instead of being acted on singly and immediately like a reflex, multiple sensory signals could be simultaneously conveyed to central processing regions and operated on in a common space. Circuits were built that integrated, amplified, compared, filtered and otherwise processed those signals to extract information about what was out in the world and what that meant for the organism. More and more abstract concepts were extracted — not just things, but also types of things and types of relations between them. Creatures capable of understanding emerged.

Meaning became the driving force behind the choice of action by the organism. That choice is real: The fundamental indeterminacy in the universe means the future is not written. The low-level forces of physics by themselves do not determine the next state of a complex system. In most instances, even the details of the patterns of neural activity do not actually matter and are filtered out in transmission. What matters is what they mean — how they are interpreted by the criteria established in the physical configuration of the system. Animals were now doing things for reasons.

That causal power does not come for free: It is packed into the organism through evolution, through development and through learning. It is encoded in the genome by the actions of natural selection. And it is embodied in the physical structure of the nervous system in the strength of neuronal connections that express functional criteria in relation to a hierarchy of aims of the organism. There is nothing here that violates the laws of physics; it just demands a wider concept of causation over longer time frames and an understanding that the dynamic organization of a system, which encodes meaning, can constrain and direct the dynamics of its component parts.


And yes, your actions are at any given moment constrained by all those prior causes. Yet you could just as well say, more positively, that they are informed by prior experience. That is precisely the property that sets life apart from other types of matter: Living things literally incorporate their history into their own physical structure to inform future action. For those who would argue this impinges on the freedom of the self to decide at any moment, I counter that it is this very process that enables the self to exist at all. There is no self in a given moment: The self is defined by persistence over time.

And though you are configured in a certain way that reflects all this history, you are not hardwired. We humans have the remarkable capacity for introspection and metacognition. We can inspect our own programming, treating goals and beliefs and desires as cognitive objects that can be recognized and manipulated. We can think about our own thoughts, reason about our own reasons and communicate with each other through a shared language. We can access the machine code running in our brains by translating high-level abstract concepts into causally efficacious patterns of neural activity. This gives a physical basis for how decisions are made in real time, not just as the outcome of complex physical interactions but also for consciously accessible reasons, and it provides a firm footing for the otherwise troublesome concept of mental causation.

So, if you want to know what kind of thing you are, you are the kind of thing that can decide. Not just a collection of atoms pushed around by the laws of physics. Not a complex automaton whose movements are determined by the patterns of electrical activity zipping through its circuits. And not an NPC, unknowingly driven by its programming. You are a new type of thing in the universe — a self, a causal agent. In the game of your life, you are Player One.

Pipi and its databases

Mike's Notes

It is time to write about Pipi 9 and why it has so many internal databases. I get a lot of questions about this.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

17/05/2025

Pipi and its databases

By: Mike Peters
On a Sandy Beach: 05/01/2025

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

Introduction

Pipi 9 consists of many autonomous agents, each a software program containing at least one relational database.

Example

Engines are a type of autonomous agent. The Namespace Engine is an example. It provides a unique name for every part of Pipi, ensuring message delivery, version history, security, and more.

This namespace data is stored in a database part of the engine. There can be multiple copies of any engine, so the records in each database will be different based on where the copy of the engine is located.

In the Namespace Engine, there are 60+ tables.

Data Model

The databases inside Pipi 9 are designed by humans, built and populated by Pipi.

Evolution

Pipi 9 was designed to evolve and self-organise in response to its external environment.

These changes happen at the level of the autonomous agents and in their database records.

The database also stores instructional information on how any autonomous agent should behave. Combined with templates, this data can generate static code for later execution. As the data changes, new code is written.

Speed

Pipi 9 is slow because of its complex internal structure. It runs many batch processes and adapts slowly but is reliable, stable, and cheap. It is ideal for interacting with large, complex systems where minor course adjustments are often needed. It is more like a slime mould than an LLM—a super thermostat capable of slowly learning.

Emergent Properties

The result of all these interactions causes Pipi's emergent behaviour as a generative platform. I still don't fully understand how I got this to work—it was more accident than intent. Hopefully, as Pipi receives a front end, more light will shine on this question.

SaaS Applications

Q. What about the free applications available on GitHub in the future?

A. They will each use an industry-specific relational database model.

Inspiration

I got the idea when I read Markus Covert's 2014 article in Scientific American about successful computer simulations of Mycoplasma.

Summary

The databases play a key role in this process because the records store the history of the path taken, phase changes, current configuration, settings, constraints, etc., which can change or "evolve" over time. This acts as a feedback loop as the engines and other autonomous agents interact with their environment.

From reductionism to dynamical systems

Mike's Notes

"For The Transmitter’s first annual book, five contributing editors reflect on what subfields demand greater focus in the near future—from dynamical systems and computation to technologies for studying the human brain.an article by Nicole Rust on how the brain works." - Transmitter

One of those authors was Nicole Rust who wrote "The field is increasingly embracing the notion that the brain is a “complex dynamical system” where causes lead to effects that feed back as causes—this happens through feedback loops within the brain and interactions between the brain and the environment. From ecology, engineering and other fields, we know that when complex dynamical systems go awry, they can be exceedingly difficult to restore. Tackling that challenge will be the key to developing treatments for the billions of people with brain conditions of nearly every type, from Parkinson’s disease to psychosis." - Nicole Rust Transmitter

That got me curious, so I looked up what else she had written and discovered the article below.

"Nicole Rust is professor of psychology at the University of Pennsylvania in Philadelphia. Her research focuses on understanding the brain’s remarkable ability to remember the things we’ve seen and using that knowledge to develop new therapies to treat memory dysfunction. She is also writing a book on the types of understanding of the brain that will ultimately be required to treat neurological and psychiatric conditions. In it, she argues that effective progress in brain research will require ambitious and unprecedented multidisciplinary conversations of the type that will appear in The Transmitter.

Rust received her Ph.D. in neuroscience from New York University and completed her postdoctoral training at the Massachusetts Institute of Technology. She has been recognized by the Troland Research Award from the National Academy of Sciences, the McKnight Scholar Award, a CAREER Award from the National Science Foundation, a Sloan Research Fellowship, the Charles Ludwig Distinguished Teaching Award, and election to the Memory Disorders Research Society." - Transmitter

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > The Transmitter
  • Home > Handbook > 

Last Updated

17/05/2025

From reductionism to dynamical systems: How two books influenced my thinking across 30 years of neuroscience

By: Nicole Rust
Transmitter: 26/08/2024

I first read Francis Crick’s “The Astonishing Hypothesis,” published in 1994, as a disenchanted undergraduate student. I knew that I wanted to change my major from chemical engineering, but I was unsure about what I wanted to switch to. In Crick’s book, I found the answer in spades. In fact, 30 years later I can still recite key lines from “The Astonishing Hypothesis” from memory: “You, your joys and your sorrows, your memories and your ambitions, your sense of personal identity and free will, are in fact no more than the behavior of a vast assembly of nerve cells and their associated molecules.” Though I wasn’t completely convinced that the hypothesis was correct, the notion that I could make a career out of studying it hit me as a profound insight. I still have my original copy of Crick’s book, complete with a dozen or so Post-it notes marking the most important passages.

I eventually ended up studying vision and memory rather than consciousness per se, but that tattered copy has continued to be influential across my career. I read it most recently a few years ago, when I was contemplating writing a book myself. Though I still very much respect the brilliance of Crick’s book, what struck me on that recent reread was how outdated it had become. I view that not as an indictment but as evidence of the evolution of our field. “The Astonishing Hypothesis” reflects an era of 1990s neuroscience in which researchers shifted toward thinking about the wonders of the brain and mind in ways that were more mechanistic and scientifically falsifiable. That approach tended to try to reduce every phenomenon to a simple explanation, such as the expression of a gene or the activity of a brain area. For instance, Crick proposed that the decisions we make may not be freely decided but instead predetermined, and that our illusion of free will may arise from what happens in the anterior cingulate cortex.

What is the modern alternative? Enter Kevin Mitchell‘s 2023 book “Free Agents: How Evolution Gave Us Free Will,” which reflects a shift toward thinking about the brain as a complex dynamical system with emergent properties that defy reduction to simple elements. In “Free Agents,” Mitchell spells out how will that is truly free could exist in a dynamical brain in which agency has evolved across millions of years. (The Transmitter published an excerpt from “Free Agents” last year.)

The gist of Mitchell’s proposal is that higher-level mental states do not simply emerge from the activation of neurons and their interactions, but that they also influence how brain activity will evolve in the future. In Mitchell’s account, noise allows for multiple possible future brain states, and top-down causal influences determine which one happens. It’s this process that creates free will. A long-held objection to this type of idea is that mental states cannot both emerge from brain activity and also simultaneously cause it, because in that argument, causality is “circular.” Central to Mitchell’s argument is that top-down influences do not act instantaneously but instead shape the future of the system, like a “spiral” that unfolds in time. This provocative proposal has inspired me to think about how we might test ideas about “spiraling” causality—to explain not just free will, but also functions such as seeing, remembering and feeling.

Causality question: Mitchell ascribes free will to a process in which noise allows for multiple possible future brain states, and top-down causal influences determine which one happens. A long-held objection to this type of idea is that mental states cannot both emerge from brain activity and also simultaneously cause it, because in that argument, causality is “circular.” Central to Mitchell’s argument is that top-down influences do not act instantaneously but instead shape the future of the system, like a “spiral” that unfolds in time.

I’m very much looking forward to revisiting Mitchell’s book in 30 years to see how well it holds up. My hope is that we will have either verified that the framework is correct, or (as with Crick’s book) we will regard it as pioneering but dated, and we will have moved on to a better alternative. Either would reflect tremendous progress in our field. I can’t wait to find out.

Data Model and Ontology

Mike's Notes

An excellent discussion titled Design Pattern Ontology recently occurred on the Ontolog Forum. Igor Toujilov started the thread on November 26, 2024.

It covered a lot of ground, and a point made by John Sowa is copied below.

Eventually, it led to a discussion about the relationship between data models and ontologies. I have copied below the text of many of the points raised as a valuable reference for future work on Pipi.

The contributors are;

  • John Sowa
  • William "Bill" Burkett
  • Mike Peters
  • Michael DeBellis
  • David Eddy
  • Paul Tyson
  • Igor Toujilov
  • Alex Shkotin
  • Kingsley Idehen
  • Elisa Kendell
  • Mike Bennett

I will keep adding to these notes as more useful contributions are made.

Any errors or omissions are mine.

Resources

People mentioned

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

17/05/2025

Data Model and Ontology

By: Mike Peters
On a Sandy Beach: 01/12/2024

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

Design Pattern Ontology

Note taken for the Ontolog Forum

John Sowa

Since Matthew isn't with us, I'll summarize some of his points which we had discussed over many years in different ways.

He had been working at Shell Oil for years, and he had developed a detailed ontology for the oil industry.  He later generalized it to develop a more general top level, which I considered quite good.   We had discussed issues about generalizing it even farther,  but he was reluctant to go farther in levels of abstraction.  We agreed that was a reasonable point of view.

But we also agreed that his details for the oil industry might conflict with details for other industries, such as banking and farming.   Furthermore, he recognized that different oil companies had different ways of representing the same terms because they had developed different policies and procedures.

I'll also mention another widely used ontology, which evolved over a period of about 70 years, and it is unlikely to change for a long, long time:  the ontology for making reservations for airlines, which was later extended to cover anything related to airlines, such as hotels, car rentals, trains, taxis, etc., etc.

And that ontology began in the 1950s with IBM's project SAGE for the airplanes used in the Strategic Air Command  over North America.  In the 1960s, IBM adapted that ontology for American Airlines.  IBM later sold the software to other airlines.   And all the other additions were made to conform to the same basic ontology from the 1950s.  70 years later, that top level is so entrenched that it is never going away.

Another world-wide ontology that also developed in the 1960s includes the global weather patterns that were established by the world-wide weather simulation programs.  They use a different way of representing the world than the reservation systems.

Fundamental problem:  There will never be a single universal ontology for representing anything and everything in the world (or the universe).

That is a fact that any system of knowledge representation must deal with.

Data Model and Ontology

Abbreviated notes taken from the Ontolog Forum

01/01/2025 and ongoing

William Burkett

  • Hay, D. C. (1996). Data model patterns : conventions of thought. New York, Dorset House Pub.
  • Hay, D. C. (2006). Data model patterns : a metadata map. Amsterdam ; Boston, Elsevier Morgan Kaufmann.
  • Silverston, L. (2009). The data model resource book, Vols 1-3. New York, John Wiley.

Mike Peters

  • Hay, D. C. (2011). UML and Data Modeling: A Reconciliation, Technics Publications.

Mike Peters

Ontologies are represented in graph structures. Non-relational databases like semantic or graph databases are better suited for this job, and ontologists (I'm not one) have no problem working with them.

However, the workforce needs conventional everyday interfaces driven by relational databases. So, there is an import/export issue that David Hay could have written a book about, and I wish he had. His explanations are excellent. His book on UML and data modelling also bridged two different ways of looking at the world.

Michael DeBellis

RE: The distinction / difference between data models & ontologies is what...?

I think there is a difference but it isn't what many people in the ontology community seem to think. Most of the methodologies and guidelines I've seen for building ontologies present the ontology as a thing in itself. For example, they often minimize or even ignore questions such as loading data or integrating with existing systems. In reality these are some of the most important issues to face for ontologies used in the real world. Some of the most important distinctions between an E/R model and an ontology are:

1) E/R models are typically used for Online Transaction Processing (OLTP) ontology and knowledge graph models are typically used for Online Analytic Processing (OLAP). The design of an OLTP model needs to be optimized for response time. Thus, such models tend to be fairly sparse, leaving most of the domain knowledge in the code of the systems that use them. The design for an OLAP model can be much richer and include much more knowledge about the domain in the model itself rather than in the code because the users will tend to be working on client machines that have more processing power and because they are going to be executing complex queries will be a bit more patient than someone adding a post to Facebook or doing a funds transfer on their bank account with their phone.

2) E/R models need to worry about various normalization forms for peak efficiency depending on what needs to be maximized. Ontologies are implemented as graphs and don't need to consider the same kinds of issues. Essentially when you build an ontology you are able to work at the analysis level whereas for an E/R model you tend to work at the design level.

3) E/R models are relatively difficult to change at run time. Ontologies and knowledge graphs can easily be changed at run time, not just the instance values but the schemas themselves can be changed at run time. In this way they are more similar to NoSQL databases that have "schema on read" rather than E/R models that have schema on write.

Actually, I just remembered I'm putting together a table comparing relational databases, NoSQL (Hadoop) and OWL. Here's what I have so far. This is a work in progress:

Feature

Hadoop (HDFS + Pig/Hive)

SQL Databases

OWL Knowledge Graphs (e.g., AllegroGraph)

Schema

Schema on read

Schema on write

Schema on write but schema can be modified at run time

Data Model

File-based (raw formats)

Tables (rows/columns)

RDF triples or quads (subject-predicate-object)

Storage

Distributed (HDFS)

Centralized or sharded

Triple or quad stores with distributed options and sharding

in some triplestores (e.g., AllegroGraph)

Processing

Batch (MapReduce)

Transactional (ACID)

Semantic reasoning and SPARQL queries. Some triplestores also

support ACID transactions.

Querying

Pig Latin, HiveQL

SQL

SPARQL and Description Logic queries

Reasoning

None

None

SWRL, SHACL, OWL DL axioms

Flexibility

Flexible (unstructured, semi-structured, structured)

Rigid schema

Highly flexible with semantic annotations and Linked Data

Scalability

Horizontal (many nodes)

Vertical (more powerful nodes)

Horizontal or vertical, depending on implementation

Integration

Tools for ETL and analytics

Tight coupling with apps

Can integrate with ontologies, Linked Data, and other RDF graphs

Best For

Batch processing, raw data storage

Transactional workloads, OLTP

Semantic data, reasoning, and complex relationships

David Eddy

You might want to add this site to your reading list. https://www.db-engines.com/en/ranking list of DBMSs DBMS engine tally at 417 as of 2024-01-04 Their counting system gets a little wonky at 390+

Paul Tyson

2024 - P15Y = 2009 < 2012 (date of R2RML recommendation)

Igor Toujilov

Mike, I would not say "Ontologies are represented in graph structures" only. Ontologies can be represented in a wide range of formalisms, including graphs, which are just one possible representation. For example, there are tools to store the same ontology in different representation formats: RDF/XML, Turtle, OWL Functional Syntax, Manchester OWL Syntax, etc. Yes, RDF and Turtle are graph representations. But OWL Functional and Manchester syntaxes have nothing to do with graphs. And yet they represent the same ontology.

I also disagree that "the workforce needs conventional everyday interfaces driven by relational databases". It depends on your system architecture. Today many systems use No-SQL or graph databases successfully without any need for relational databases.

In real systems, the difference between data models and ontologies can be sharp or subtle. Some systems continue using relational databases while performing some tasks on ontologies. Other systems have ontologies that are tightly integrated in the production process, so sometimes it is hard to separate the ontologies from data. And of course, there is a wide range of systems in between of those extreme cases.

William Burkett

My take is that the difference is primarily one of intention. Ontology designers and “conceptual data model” designers are seeking to create representations of the “real world”. Logical/physical data designers are seeking to create specifications for actual data structures for applications to store/access/use data. This is distinction is, of course, very fuzzy and fluid because both sets of designers are usually pursuing both of these intentions simultaneously. If an “ontology” is specified in OWL, it is, of course, is a self-defining “data structure” that is processable by applications designed to use that structure – so, IMHO, there is no objective, practical difference between them. Unless we’re talking about box-and-line diagrams, most ontologies that we talk about here (I think) are just special kinds of data models.

Alex Shkotin

Your table reminds me "What Goes Around Comes Around… And Around… Michael Stonebraker and Andrew Pavlo"

We have discussed it about 2 and half hours at our last meeting SIGMOD-MoscowThey overview RDF and graph DB, but not OWL.

Kingsley Idehen

Yes, R2RML is the component of “the SemWeb stack” designed to describe how relations in a DBMS are represented as relations in RDF.

Here’s a post I wrote years ago, complete with live examples, demonstrating how to create an RDF-based Entity Relationship Graph from CSV files located on a local filesystem accessible to a Virtuoso instance:;

Virtuoso is a multi-model database management system in one of its guises. Specifically, it can operate on relations represented in the coarse-grained form typical of SQL-based DBMSs or in the finer-grained form facilitated by RDF via its Quad Store functionality.

In my view, Ontology and Data Model are different terms for the same concept—they’re essentially conceptual synonyms. That said, practitioners often view them differently depending on their operational lenses. For example, most ontologists express their worldviews on data modeling using formal graphic or linear notations, while data modelers (often from the RDBMS domain) favor graphical notations such as ERD (Entity-Relationship Diagrams) or UML (Unified Modeling Language) diagrams.

Links:

  • [2] Example Hyperlink Super-Key
    A hyperlink-based super-key resolving to a specific row in a CSV document (which could also represent any ODBC- or JDBC-accessible table in an RDBMS)

Mike Peters

So, what options are there to map RDF to relational data?

Michael DeBellis

RE: So, what options are there to map RDF to relational data?

I was going to list what little I knew but I'm working with ChatGPT right now, so I asked it. Here's what it said:

Implementing OWL (Web Ontology Language) in a relational database involves translating OWL's semantic capabilities into a relational schema and querying framework. 

Several tools and approaches allow you to achieve this:

Tools and Frameworks for OWL in Relational Databases

1. RDF and SPARQL Layers on Relational Databases
  • Apache Jena with RDBMS Integration:
    • Jena provides a framework to work with RDF and OWL.
    • It supports relational database backends via the TDB or RDB storage layers.
    • You can map RDF triples to relational tables and query them using SPARQL.
    • Virtuoso Universal Server:
      • Virtuoso supports RDF storage and SPARQL queries directly in a relational database.
        • It combines a traditional RDBMS with a semantic store.
        • Useful for large-scale ontologies.
    • D2RQ:
      • A tool for exposing relational databases as virtual RDF graphs.
      • You can use OWL over the RDF representation while keeping data in the relational schema.
    2. Ontology Mapping Tools
    • Ontop:
      • A powerful tool to map OWL ontologies to relational databases using OBDA (Ontology-Based Data Access).
      • Allows you to query relational data through SPARQL while maintaining OWL semantics.
      • Supports mappings through R2RML or native Ontop mapping syntax.
    • Karma:
      • A tool for integrating data from relational databases with OWL ontologies.
      • Supports mapping relational data to RDF and OWL constructs.
    3. Hybrid Database Systems
    • GraphDB (by Ontotext):
      • A triplestore that can integrate with relational databases.
      • It allows you to store and query OWL data seamlessly alongside relational data.
    • Oracle RDF Graph:
      • Part of Oracle Database, it supports RDF and OWL data alongside relational data.
      • Provides SPARQL querying and reasoning capabilities.
    4. Custom Implementation Approaches
    • Mapping OWL Axioms to Relational Rules:
      • Use SQL views or stored procedures to encode OWL semantics manually.
      • Define mappings for classes, properties, and relationships in relational terms.
    • DL-Lite Ontology Mapping:
      • If your OWL usage aligns with OWL 2 QL (a subset of OWL optimized for relational databases), you can use DL-Lite mappings for efficient querying.

    Considerations for Choosing a Tool

    1. Scale of the Ontology:
    • Large ontologies benefit from systems like Virtuoso or GraphDB, which are optimized for scalability.
    2. OWL Subset:
    • OWL 2 QL is specifically designed for integration with relational databases. Tools like Ontop are excellent for such use cases.
    3. Querying Needs:
    • If SPARQL is a requirement, tools like Ontop, Jena, or Virtuoso are natural fits.
    • For simple SQL querying, D2RQ or custom mapping may suffice.
    4. Reasoning Capabilities:
    • If you require reasoning over OWL axioms, focus on tools like GraphDB, Oracle RDF Graph, or Jena.

    Suggested Approach

    If you're starting out and want robust OWL support with minimal setup, Ontop is highly recommended. It supports direct integration with relational databases, enables SPARQL querying, and is well-documented.

    Would you like a walkthrough on setting up any of these tools, or a comparison of their features?

    Elisa Kendell

    Together with other colleagues, we’ve used a number of the tools listed for mapping FIBO to relational for large banks successfully, including Ontop and Jena tools, among others. FIBO does include axioms, particularly some cardinality restrictions, that are outside of OWL RL, but most of the semantics can be mapped without any issues. RDFOx supports OWL RL, which is more expressive than OWL QL, and ignores anything outside of the RL profile rather than throwing up on it. Their team has also provided some rules that extended OWL RL for us to test with, including approximations of the axioms in FIBO that are outside of RL, which worked well. Stardog supports all of the semantics encoded in FIBO, and most knowledge graph tools that support SPARQL 1.1 can import it, though they may not support all of the reasoning encoded in the ontologies.

    RDFOx and some other knowledge graph engines prefer either Turtle or JSON-LD to RDF/XML, which is the serialization we work in (primarily to see all of the warts in what we are publishing). But FIBO and the other ontology efforts I participate in publish in all three serializations – RDF/XML, Turtle, and JSON-LD, so that we can supply whatever is needed to a given tool/framework. Same is true of the Commons, MVF, LCC, and other ontologies we publish at OMG - in RDF/XML and Turtle, at a minimum. There is also a toolkit available from the EDM Council that we use to support transformations between serializations consistently, that we use for GitHub comparisons as well as for tool support, which we publish as open source at https://github.com/edmcouncil/rdf-toolkit. It’s a fairly complex Swiss army knife, with various options you can use to manage the transformation as needed.

    Kingsley Idehen

    Do you mean the reverse—creating SQL RDBMS relations from RDF-based relations? If so, note that the SPARQL query language includes a SELECT option for projecting query solutions as tables from RDF Graphs, which can then be fed into a SQL RDBMS.

    Mike Peters

    Yes, I do mean importing ontologies into relational databases. I'm not an ontologist, but I can see the great value in using ontologies, schema and taxonomies as read-only references in a working database.

    The question is how to reliably and effectively allow users to point at any ontology using a form (e.g., something on OBO Foundry or SnowMed) and import it into the relational database they are logged into.

    I was thinking OWL and RDF. Are both possible?

    Michael DeBellis

    As someone rightly pointed out in response to one of my answers, OWL is a logical, not a graph model, and not necessarily tied to RDF.  Since OWL is a subset of First  Order Logic (Description Logic) it can directly map to a relational database rather than going through RDF first. According to ChatGPT: 

    Ontop Supports mappings through R2RML or native Ontop mapping syntax.

    Mike Bennett

    Well this has been a very interesting sub-thread. I'll fork here from before the sub-thread on RDF to RDB etc. considerations.

    Is "Ontology" really synonymous with, or even necessarily a kind of, "Data Model"?

    I'd say emphatically not. There are kinds of ontology that are a kind of data model, of course, and much has been said about these in this sub-thread.

    But those are not the only things of which it can be said "This is an ontology".

    Any model has an "aboutness"; that is, "Of what is this a model"

    For some models, what it is about is data: each element of the model represents some element of data.

    For some models, the aboutness is that of things in the world.

    The model language or formalism, and the model aboutness, are orthogonal: it does not necessarily follow that a model in a given language must be about a given kind of thing. UML Class models are designed to represent Object Oriented class constructs (with both behavioral and structural elements), but some people use them to represent all sorts of other things (including sometimes, things in the world). Similarly, an OWL model may represent RDF data and usually does, since that is what it is intended for.

    Suppose someone wants to have a model of real things in the world. One would call this an "Ontology". However, as  soon as someone says they want an ontology, various people pop up and say "I can do you an ontology" when what they mean, as evidenced in this thread, is "I can do you an ontology of the sort that is a kind of data model".

    Maybe that's what the customer needed, maybe it's not. If the business needs something that formally defines the meanings of things, for example for management communication, reporting, common understandings (in place of word-dependent dictionaries or glossaries) and so on, or if they want something for AI to process, then the chances are they need an ontology of the sort that represents real things in the world. All too often they get given an Ontology-as-data-model because someone thinks that is the only sort of ontology there is.

    There are some questions, the answer to which is not a data model.

    Let's consider 2 things:

    1. Basic engineering best practice
    2. Practical examples of how these kinds of ontology are different.

    Good engineering follows a separation of concerns. Artifacts that represent the customer or business view, for example defining what the customer wants or what their world looks like, should always be expressed independently of any assumptions about the design techniques or technologies that will be used in crafting a solution. 

    For example a business process model represents the activities that the business carries out, independently of any software design to automate these.

    The reason is that (a) things are represented without presuming anything about the solution and (b) the solution can then be validated against that design-independend artifact. That's basic QA. 

    Similarly, a data model is a kind of design (typically done at 2 levels: Platform Independent and Platform Specific, both of which are still designs).

    The corresponding design-independent artifact is a kind of ontology: one in which the real-world meanings of the things of interest to the business are expressed. In other words, what does it mean to be this kind of thing?

    Traditionally that's been done with words. But words are slippery. Better to use formal logic.

    So there are ontologies which are a kind of data model, and there are ontologies which are a representation of things in the world. Both are needed, at these different levels in the development method, and with linkages between them.

    A practical example of the difference is the best way to understand this distinction between these kinds of ontology.

    And the difference is best illustrated with an example of where it went wrong.

    The difference is between what we call "Truth Makers" and data. Truth Makers are what it takes for something to be defined as being a member of a given class of Thing. These are the necessary and sufficient conditions for a thing in the world to be a member of that class. Most of these are either physical matters such as physics and chemistry, or legal and social constructs such as legal capacities, value and so on (mainly classifiable under Searle's Ontology of Social Constructs - an ontology which is definitely not a data model; it's a book). A very few things get their meaning from data itself, as a kind of thing.

    An ontology of things in the world (let's call this a Concept Ontology) defines things using those truth makers.

    An ontology as a kind of data model looks for data surrogates for those things in the world: what data can you expect to find when this or that legal capacity, physical quantity value etc. is in play?

    Example: suppose we consider what it means to be a bank. Very loosely, this is something with certain legal capacities, such as the capacity to take on funds, the capacity to disburse those funds and so on.

    In one project I was involved with, the class "Bank" was defined using a data element for "FDIC Insurance", a kind of insurance that all banks in the US must carry. Then it was noticed that the DTCC, a clearing house, also carries FDIC Insurance, and so a different data item was sought instead.

    There were two errors, one inside the other. The first, proximate error was that they chose the wrong data surrogate. The error inside the error (the ultimate error) was that they did not realize they were making a design decision for a data surrogate. Therefore, that design decision was not peer reviewed and was only discovered later (costing time etc. to fix which is another reason we have separation of concerns).

    The model was updated to use the more correct data surrogate of Banking License. This reliably exists whenever the legal capacities for something to be a bank exists, and doesn't when it doesn't.

    Of course that would not work in all use cases: if the requirement was to detect when some entity was acting as a bank when it should not be, you would look for data about the entity's behavior instead. Different use cases may give rise to different data surrogate design decisions.

    And that's why, while ontology-as-data-model is an extremely valuable kind of data model, the same kind of engineeering integrity should go into their design as into the design of anything else, including the provision of ontologies of the target domain subject matter (subject to scope), against which these can be designed, against which design decisions can be reviewed, and against which the end result can be tested.

    Those are ontologies too. These ideally use formal logic because words are too slippery to be relied upon. But just because a concept ontology is framed using formal logic, does not make it a data model (logic has been around a lot longer than computational data).

    Kingsley Idehen

    You can use SPARQL SELECT, as I described against a collection of relations (comprising terms from both the RDF and OWL ontologies/vocabularies) to insert data into relations (colloquially referred to as tables) managed by an RDBMS.

    John Sowa

    As Yogi Berrra said, these discussions are "Deja vu all over again."

    Re "No SQL":  The person who coined that term, rewrote it as Not Only SQL   The original SQL was designed for data that is best organized in a table.  The fact that other data might be better represented in other formats does not invalidate the use of tables for data that is naturally tabular.

    Re tree structure in ontologies:   A  tree structure for the NAMES  of an ontology does NOT imply that the named data happens to be a tree.   Some of the data might be organized in a tree, but other data might be better organized in a table, list, vector, matrix, tensor, graph, multidimensional shapes, or combinations of all of them.

    The following survey article was written about 40 years of developments from 1970 to 2010.  Some new methods have been invented since then, but 90% of the discussions are about new names for old ideas re-invented by people who didn't know the history.   I wrote the survey, but 95% of the links are to writings by other people;  https://jfsowa.com/ikl .

    And by the way, I agree with Bill Burkett (on the list down below).  He is one of the people I collaborated with on various committees in the past many years.  We viewed the Deja Vu over and over and over.  That's one reason why I don't get excited by new names.

    Alex Shkotin

    What is an idea to import ontology into RDB? Just to store it? Or to use it as a schema for RDB?

    And if the first do you need to keep it structurally or just as a blob? 

    Mike Peters

    The idea is to feed Pipi 9 structured data from versioned external references, such as ontologies, taxonomies, XML Schemas, CSV, etc.

    This data then becomes bits of relational database schema or is used to populate the tables.

    This needs to be an automated process that is highly reliable. It's like using an external API.

    So, using a silly made-up example of what I want to end up with.


    Ontology-Imports-Table
    ----------------------------------
    ID | Source | Version | Thing 1 | Relation | Thing 2
    1 | obofoundary-example.owl | 5 | elf | worksFor | Santa
    2 | obofoundary-example.owl | 5 | rudolf | isA | Reindeer
    3 | obofoundary-example.owl | 5 | mary | isA | Elf
    4 | obofoundary-example.owl | 6 | mary | isA | RetiredElf
    5 | obofoundary-periodicTable.rdf | 1 | Plutonium | isA | Chemical Element
    6 | movieLab.rdf | 5 | Camera | hasA | Camera Lens
    7| movieLab.xml | 10 | DSMC2 Gemini 5K S35 | isA | Camera

    Depending on user requirements, this could be used to generate;

    Camera-Table
    Camera-Lens-Table
    ChemicalElement-Table
    etc

    Or

    Populate a table with read-only records.

    Alex Shkotin

    Why not ontology about ontologies like discussed here