Why BoxLang?

Mike's Notes

Pipi 9 is programming language agnostic and could use most languages, including Python, Java, and C++.

However, it uses ColdFusion Markup Language (CFML) because that is the language I know really well. It does an excellent job and is fast for prototyping.

ColdFusion code has to be executed by an application server. The main options are:

  • ColdFusion Server (Adobe)
  • Lucee (Lucee Association)
  • Open Blue Dragon (New Atlanta)
  • BoxLang (Ortus Solutions)

The Adobe product is excellent but expensive: $10,000 per server. The rest are now open-source and free. Ortus offers paid support for its open-source product.

I'm using the free developer edition of Adobe ColdFusion to test everything, but going into live production presents a big problem, the cost, and I will need multiple servers.

BoxLang by Ortus Solutions looks like the best bet, even though it is still in beta. Ortus have a proven track record.

BoxLang can also be used in many ways, including Docker, without incurring extra costs.

I can start with the free version and manage it myself. When scaling occurs, I can pay Ortus for support and get them to manage the servers remotely.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subject > CFML
  • Home > Handbook > 

Last Updated

17/05/2025

Why BoxLang

By: 
Ortus Solutions: 13/01/2025

BOXLANG programming language. BOXLANG is the modern,  dynamically and loosely typed scripting language for the Java  Virtual Machine (JVM) that supports Object-Oriented (OO) and functional Programming (FP) Constructs.

It can be deployed on multiple platforms and all operating systems,  web servers, Java application servers, AWS lambda, iOS, Android, web assembly, and more.

BOXLANG combines many features from different programming languages, including Java, ColdFusion, Python, Ruby, Go, and PHP, to provide developers with a modern, fluent, and expressive syntax. BOXLANG has been designed to be a highly modular and dynamic language that takes advantage of all the modern features of the JVM.

Key features:

Dynamic Language

BoxLang is dynamically typed, which means there´s no need to declare types. It can perform type inference, auto-casting, and promotions between different types. The language adjusts to its deployed runtime and can add, remove, or modify methods and properties at runtime, making it highly flexible and adaptable.

Multi-Runtime Development

BoxLang can be used to write adaptive code for any operating system, JVM, servlet container web server, cloud lambda functions, iOS,  Android, or even the browser using our web assembly package. BoxLang is built on a versatile language core, making it deployable on almost any current or future platform.

Established Ecosystem

BoxLang has a well-established ecosystem. Every Java and  ColdFusion/CFML library is compatible with BoxLang. It comes with  CommandBox, which serves as its package manager, server manager, task manager, and REPL tool.

Java Interoperability

BoxLang offers 100% interoperability with Java. It can extend and  implement Java objects, use annotations, declare classes, imports, and more. Thanks to InvokeDynamic and the BoxLang Dynamic Object core, everything in BoxLang is fully interoperable with Java.

Low Verbosity Syntax

It is highly functional, fluent, and human-readable. It is highly expressive and has low ceremony.

Event Driven Language

BoxLang comes with an internal event bus that can be used to enhance the language's capabilities or those of any application. It´s able to listen to almost every aspect of the language, parser, and even the runtime, or collaborate using the user own events.

Enterprise Caching

The system has internal caching capabilities, which can be customized with providers, object stores, listeners, and statistics.

IDE + Tools

It comes with various tools to assist developers in their work. It includes a Visual Studio Code extension for the language, providing features such as syntax highlighting, code debugging, code insight, code documentation, formatting, LSP integration, and more. Subscribers to the premium version receive additional tools, including enhanced debuggers, ColdFusion/CFML transformers.

Scheduling & Task Framework

It provides a centralized and portable way to define and manage scheduled tasks for user servers and applications.

Functional Programming

It enables users to explore functional concepts such as immutability and higher-order functions to write robust code. BoxLang promotes code that is easier to maintain and less error-prone, thanks to its support for functional programming principles.

Modular Extensibility

It's possible to add new built-in functions, template components, or modify existing functions on classes. This language also supports a Runtime Debugger, AOP aspects, and event listening within the language itself.

Multi Parsers: CFML & BoxLang

BoxLang supports a dual parser and transpiler for executing ColdFusion/CFML code natively. It is possible to run all ColdFusion applications within BoxLang natively. Additionally, it provides tooling that automatically transpiles ColdFusion code to BoxLang for Subscription users.

Professional Open Source

Professional open-source project based on the Apache 2 license. Core features and Community Support.

Serverless Development

Allows developers to write code in functions that trigger based on events and leave server management to the cloud provider.

Meta-Programming

Allows developers to modify the behavior of BoxLang classes and objects at runtime. It involves manipulating the way classes and objects respond to method calls, property access.

General Information

BOXLANG OPEN-SOURCE

A professional open-source project based on the Apache 2 license. The BOXLANG open-source edition can be used to create and deploy commercial software without paying a single cent. We love and breathe open-source, but we also know that it comes with responsibilities.

BOXLANG +

A subscription professional license for the language and its runtimes (web, android, web assembly, etc.). It gives world-class support, customizable SLAs (Service Level Agreements), custom patches, security notifications, enhanced features, modules, and more. It has a transparent and no-fuzz licensing model no matter which runtime is used or where is deployed.

BOXLANG + +

Provides its customers with a range of services that guarantee superior Service Level Agreements (SLAs) and an exclusive Yearly Checkup. This entails a thorough analysis of product usage, installation, deployment, and other critical aspects by proficient language engineers and development team members via remote or on-site consultation, as well as a dedicated account manager and video call support.

Who is responsible for BOXLANG?

BOXLANG is the future of programming. Developed by Ortus Solutions, a company with over 18 years of experience in ColdFusion development, BOXLANG is the result of a team of 25+ experts who have been working tirelessly to create a language that will revolutionize the way we write code. But BOXLANG is more than just a language; it's an entire ecosystem, driven by a community of over 100 independent contributors who are passionate about making coding more efficient, more powerful, and more accessible to everyone. Join the BOXLANG community today and help us build a better future for programming!

Ortus Solutions was founded by Luis Majano in 2006 with the vision of empowering developers with great open-source tools and empowering clients with scalable and robust applications. We have a proven track record of successful web application development from small scale to mission critical applications, software architecture, website design, training and support services.

We know how difficult it is to do everything all by yourself therefore we are here to help you out. Our focus is to make your business more efficient. We are at our core, developers who know the value of time and money and we have seen firsthand how our products and services can help you save those two invaluable things.

Ortus Solutions offers comprehensive suite of products and services for empowering, building, running and managing ColdFusion applications and now BOXLANG applications too.

We're the team behind ColdBox, the de-facto enterprise ColdFusion MVC Framework, TestBox, a Testing and Behavior Driven Development (BDD) Framework, ContentBox Modular CMS, a highly modular and scalable Content Management System, CommandBox, the ColdFusion (CFML) CLI, package manager, etc., and many more.

At Ortus Solutions we have become ColdFusion leaders by changing the way we ALL develop.

We create innovative open-source tools and software for empowering, building, running, and managing ColdFusion applications. Our Team of expert developers has over 16 years of experience in software-hardware architecture and OO design which provides our customers with world-class support, mentoring and custom development services.

Notes on reading Terrence Deacon

Mike's Notes

My notes taken while reading Terrence Deacon.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

17/05/2025

Notes on reading Terrence Deacon

By: Mike Peters
On a Sandy Beach: 12/01/2025

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

I'm reading Terrence Deacon's Incomplete Nature: How Mind Emerged from Matter. It's over 600 pages long, and I find it fascinating. It is very detailed, has many new technical words, and is hard work. So far, I'm 35% done.

I'm a visual pattern learner, so I decided to step back and watch videos of his talks, which are easier to follow. He is a great speaker.

I also discovered a great talk by Jeremy Sherman, given at Google. It covered the same material but was for beginners.

That way, I will first get a big-picture overview to give context to all the details in the book. Then, I will read and reread the whole book, marvelling at all the details.
These notes will be added to and changed as I learn what Deacon says.

Summary of different properties of the levels of order

  Life Self-organising Thermodynamic
Term Teleodynamic Morphodynamic Homeodynamic
Example Bacterium Whirlpool Hot stone cooling
Environment Controls dependence on Constantly perturbed by Random
Internal constraint Preserves Generates Destroys
Entropy Export Optimum Maximum Minimum
Self Has self Self-eliminating None
Work Use gradients to create constraints by providing paths    
How Means & Ends Cause & Effect Cause & Effect
Emergence from  Self-organising  Thermodynamic  
       

Talk by Jeremy Sherman @ Google 2018

"Researcher Jeremy Sherman poses a huge question still unanswered despite all we've discovered in the physical and life/social sciences:  what accounts for the difference between matter and mattering, cause-and-effect chemistry and means-to-ends life? 

Sherman distils his book Neither Ghost nor Machine: The Emergence and Nature of Selves, proposing a new solution – organisms prevent their own degeneration. He shows how this prevention could emerge from anything-goes chemistry at the origins of life and suggests implications for computer science." - YouTube

Terrence Deacon -- Language and complexity: Evolution inside out. Talk given at University of British Columbia26 Aug 2010

"Webcast sponsored by the Irving K. Barber Learning Centre, and hosted by the Department of Language and Literacy Education and the Faculty of Education as part of the plenary session at the 37th International Systemic Functional Congress, Deacon explains the extravagant complexity of the human language and our competence to acquire it has long posed challenges for natural selection theory.

To answer his critics, Darwin turned to sexual selection to account for the extreme development of language. Many contemporary evolutionary theorists have invoked incredibly lucky mutation or some variant of the assimilation of acquired behaviors to innate predispositions in an effort to explain it.

Recent evodevo approaches have identified developmental processes that help to explain how complex functional synergies can evolve by Darwinian means. Interestingly, many of these developmental mechanisms bear a resemblance to aspects of Darwin's mechanism of natural selection, often differing only in one respect (e.g. form of duplication, kind of variation, competition/cooperation).

A common feature is an interplay between processes of stabilizing selection and processes of relaxed selection at different levels of organism function. These may play important roles in the many levels of evolutionary process contributing to language.

Surprisingly, the relaxation of selection at the organism level may have been a source of many complex synergistic features of the human language capacity, and may help explain why so much language." - YouTube

Terrence Deacon, Origins of life, Autogen Demonstration

Life before genetics: autogenesis, and the outer solar system - Terrence Deacon (SETI Talks)

Terrence Deacon - Self-organization is not enough: on beyond complex systems

T19 W117 Terrence Deacon (Keynote) | Thermodynamics 2.0 | 2020

Terrence Deacon. Interactivism in Perspective Preconference Reading Group and Talks

TSC 2016 Plenary 8 Evolution and Consciousness

The very important video: Professor Terrence W. Deacon: The Climate Challenge - in good thinking!

Terrence Deacon Playlist (58 videos)

Teleodynamics

"How does life emerge in a world that moves lawfully toward disorder? Some have answered this question by suggesting that life is a “self-organizing” process. Self-organizing phenomena, like a whirlpool in a bathtub, produce order by turning the law of increasing entropy against itself. They are systems in which ordered forms may emerge locally in order to increase the rate of entropy production globally. The problem is that these self-organizing processes are also self-undermining: they always exhaust the boundary conditions that make them possible. A whirlpool helps the water drain faster, making the vortex disappear. 

Life does precisely the opposite. Rather than increasing the rate of global entropy production, organisms do work in order to resist further entropy increase, allowing them to preserve their forms and maintain their structural integrity. This stark difference reflects a radically different form of organization that we refer to as “teleodynamics.” More specifically, teleodynamics is defined as a higher-order reciprocal relationship between self-organizing processes, in which each process creates the boundary conditions that makes the other possible. This mutual synergy ensures that each form-generating process is brought to a halt before it has the chance to dissipate. For this reason, order is spontaneously generated, preserved, and reproduced.

This simple transition has surprisingly far-reaching implications. Browse the publications on this website to see how teleodynamic theory can provide naturalistic solutions to questions that thinkers have pondered for millenia. These questions include:

  • How do normativity and value emerge from the physical world?
  • What is conscious selfhood and emotional experience?
  • How can physical processes come to “interpret” and “represent” the world? How does a molecule like DNA become “about” other molecules?
  • How should we define “information” and how does information produce physical effects?


Teleodynamics emerges when multiple self-organizing phenomena generate forms (constraints) that serve as the boundary conditions that make the other self-organizing processes possible, resulting in a spontaneous tendency for the self-generation and self-maintenance of the whole. " - Teleodynamics.org

Databases in 2024: A Year in Review

Mike's Notes

Below is a blog post by Andy Pavlo, a Professor at Carnegie-Mellon University. It provides a valuable heads-up on some of the fortunes of database products. Knowing what to avoid, what "open-source" products might not remain so, and where support may change is handy. The original blog article contains many links.

His faculty page includes useful course material about databases, much of which is publicly available. Most of the public talks are also available on YouTube.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

17/05/2025

Databases in 2024: A Year in Review

By: Andy Pavlo
Carnegie-Mellon University Blog: 01/01/2025

Like a shot to your dome piece, I'm back to hit you with my annual roundup of what happened in the rumble-tumble game of databases. Yes, I used to write this article on the OtterTune blog, but the company is dead (RIP). I'm doing this joint on my professor blog.

There is much to cover from the past year, from 10-figure acquisitions, vendors running wild in the streets with license changes, and the most famous database octogenarian splashing cash to recruit a college quarterback to impress his new dimepiece.

I promised my first wife that I would write more professionally this year. I have also been informed that some universities assign my annual blog articles as required reading in their database courses. Let's see how it goes.

Previous entries:

  • Databases in 2023: A Year in Review
  • Databases in 2022: A Year in Review
  • Databases in 2021: A Year in Review

It’s My Database and I’m Going to License it the Way I Want! 

We live in the golden era of databases. There are many excellent (relational) choices for all types of application domains. Many are open-source despite being built by for-profit companies backed by VC money.

But VCs want that money back and their trap full, so these companies turn out a hosted service for their DBMSs on the cloud. But the cloud makes open-source DBMSs a tricky business. If a system becomes too popular, then a cloud vendor (like Amazon) slaps it up as a service and makes more money than the company paying for the development of the software. This threat is why many database companies change to more restrictive source code licenses that protect against cloud vendors from reselling their products. MongoDB was among the first to do this in 2018 when they switched to their Server Side Public License (SSPL).

This past year was a turbulent one for license changes, and the most prominent two were Redis™ and Elasticsearch.

Redis:

Redis Ltd. (the company) is on an aggressive path towards its IPO. Originally starting as Redis Labs in 2011, they switched their name to Redis Ltd. in 2021 when they acquired the Redis trademark from its creator (Salvatore Sanfilippo), who Redis Labs had bankrolled. Over the last few years, Redis Ltd. has attempted to consolidate control over the Redis landscape. The company has also tried to cast off the perception that the system is primarily used as an in-memory cache by adding support for vectors and other data models.

In March 2024, Redis Ltd. announced they were switching from the system's original (very permissive) BSD-3 license to a dual license comprising the proprietary Redis Source Available License and MongoDB's SSPL. The company announced this change the same day they announced the acquisition of Speedb, an open-source fork of RocksDB.

The backlash to the Redis license move was quick. The same week the license changed, two forks were announced based on the original BSD-3 code line: Valkey and Redict. Valkey started at Amazon, but engineers at Google and Oracle quickly joined. In only one week, the Valkey project clapped back at Redis Ltd. when it became part of the Linux Foundation, and several major companies shifted their development efforts to it. Redis Ltd. did not help their perception of being up to no good when they got frisky with their beloved trademark and started taking over open-source Redis extensions.

In an obvious homage to when Bushwick Bill (RIP), Scarface, and Willie D got back together in 2015, the Redis' creator announced in December 2024 that he is in touch with the Redis Ltd. management and is looking to make a comeback to reunite the Redis community.

Elasticsearch:

Elastic N.V. is the for-profit company backing the development of the leading text-search Elasticsearch DBMS. In 2021, they announced their switch to a dual-license model of the Elastic License and MongoDB's SSPL. Again, this was in response to the rising prominence of Amazon's Elasticsearch offering, even though the service had been out since 2015. Amazon didn't take this switch kindly and announced their OpenSearch fork.

Three years later, Elastic N.V. announced in August 2024 that they reverted their license change and switched to the AGPL. Their blog article announcing this change references Kendrick Lamar's songs (e.g., Not Like Us). Amazon did not like being called the Drake of Databases and announced the following month that they were transferring ownership of the OpenSearch project to the Linux Foundation.

Andy’s Take:

This turmoil seems like a lot just over licenses, but remember, there is big money in databases. And this is just two systems! I didn't even discuss Greenplum quietly killing off their open-source repository after nine years and going proprietary. But people did not notice because nobody willingly runs Greenplum anymore. The only DBMS I know of that made the same open-source reversal is Altibase in 2023.

I'll be blunt: I don't care for Redis. It is slow, it has fake transactions, and its query syntax is a freakshow. Our experiments at CMU found Dragonfly to have much more impressive performance numbers (even with a single CPU core). In my database course, I use the Redis query language as an example of what not to do. Nevertheless, I am sympathetic to Redis Ltd.'s plight of being overrun by Amazon. However, the company is overestimating the barrier of entry to build a simplistic system like Redis; it is much lower than building a full-featured DBMS (e.g., Postgres), so there are several alternatives to the OG Redis. They are not in position of strength where such posturing will be tolerated by the community.

The Elasticsearch saga is the same story as Redis, except they are further along in the plot: the company announces license change, competitors create an open-source fork, and the company reverts to an open-source license to muted fanfare.

Notice that Redis and Elasticsearch are receiving more backlash compared to other systems that made similar moves. There was no major effort to fork off MongoDB, Neo4j, Kafka, or CockroachDB when they announced their license changes. CockroachDB even changed its license again in 2024 to make larger enterprises start paying up. It cannot be because the Redis and Elasticsearch install base is so much larger than these other systems, and therefore, there were more people upset by the change since the number of MongoDB and Kafka installations was equally as large when they switched their licenses. In the case of Redis, I can only think that people perceive Redis Ltd. as unfairly profiting off others' work since the company's founders were not the system's original creators. An analysis of Redis' source code repository also shows that a sizable percentage of contributions to the DBMS comes from outside the company (e.g., Tencent, Alibaba). This "stolen valor" was the reason for the ire HashiCorp received when they changed Terraform's license in 2023.

The overarching issue in these license shuffles is the long-term viability of open-source independent software vendors (ISVs) in the database market. The cloud vendors are behemoths with infinite money. If an open-source DBMS takes off, they will start hosting it and make more money than the ISV. Or they will add your DBMS's wire protocol as a front-end to an existing DBMS, like when AWS added InfluxDB v2 protocol support for their Timestream DBMS in March 2024. They can then be the girlfriend of the aforementioned Bushwick Bill and shoot you in the eye, like when AWS announced that their new Valkey-compatible services are 30% cheaper than their Redis-compatible services.

Update 2025-01-01: I previously stated that AWS added the InfluxDB v2 protocol on top of their existing Timestream DBMS. Instead of just copying the protocol, AWS is providing a managed service of the InfluxDB v2 DBMS in collaboration with Influx Data. [Credit]

Update 2025-01-01: I missed the news that ScyllaDB announced in December 2024 that they are killing off the open-source (AGPL) version of their DBMS and making the enterprise version "source available". [Credit]

The Databricks vs. Snowflake Gangwar Continues 

There continues to be no love lost between Databricks and Snowflake. This fight is a classic database war that has spilled out into the streets. The two companies’ previous feud over query performance has expanded to other areas of data management and become much more expensive.

Databricks took the first shot in March 2024 when they announced they spent $10 million to build their DBRX open-source LLM with 132 billion parameters. The Mosaic team led the development of the DBRX model, which Databricks acquired in 2023 for $1.3 billion. One month later, Snowflake rolled up to the same corner and lit it up using their Arctic open-source LLM with 480 billion parameters. Snowflake boasted they only spent $2 million to train their model while outperforming DBRX for “enterprise” tasks like SQL generation. You can tell that Snowflake cares most about taking a shot at Databricks because their announcement shows other LLMs doing better than them (e.g., Llama3), but they highlight how they are better than DBRX. One AI researcher was confused about why Snowflake focused so much on DBRX in their analysis and not the other models; this person does not know how much blood these two database rivals have spilled.

While the public LLM battle raged, another front in the war between Databricks and Snowflake opened up behind the scenes over catalogs. For most of the 2010s, Hive’s HCatalog has been the de facto catalog system on data lakes. Iceberg and Hudi emerged as replacements in the late 2010s from Netflix and Uber, respectively, and both became top-level Apache projects backed by VC-funded startups. These systems provide a meta-data service to track files and support transactional ingestion of new data on an object store (e.g., S3). Databricks has a proprietary catalog service called Unity, which works with its DeltaLake platform. Snowflake announced their initial integration with Iceberg-backed tables in 2022. They then expanded their support for Iceberg over the next few years. Then, they looked into acquiring Tabular, the main company behind Iceberg, to compete against Databricks with Unity and DeltaLake. The story goes that Snowflake was about to close the deal for $600 million. But then Databricks crashed the party and splashed $2 billion to acquire Tabular. Databricks announced the acquisition the same day as the Snowflake CEO’s conference keynote address, where he was announcing their new open-source Polaris catalog service in June 2024. Databricks continued to kick Snowflake in the teeth when they announced they were open-sourcing their Unity catalog the following week. Straight murdergram.

Andy’s Take:

What is interesting about this database battle is that it is not just about raw performance numbers. It isn’t like the old Oracle versus Informix shootout of the 1990s, where they were mostly boasting about faster query latencies. It is true that the battle also went beyond just benchmark shots when Informix sued Oracle (and later had to withdraw their suit) because Oracle poached some top Informix executives. Later, the world found out that the Informix CEO had cooked the company’s books to inflate revenue numbers to look better against Oracle and had to do a two-month bid in the federal clink.

Instead, the Snowflake versus Databricks battle has expanded to be about the ecosystem around the database. That is, it is about the infrastructure people use to get their data into a database and then the tools they use on that data. Vectorized execution engines for analytical queries are a commodity now. Databricks and every other OLAP vendor follow Snowflake’s architecture design from 2013, originally based on one of the Snowflake co-founders’ Ph.D. thesis. What matters now are quality-of-life facets (which are hard to monetize and compare with competitors), compatibility with other tools, and AI/LLM magic.

At least the competition between Snowflake and Databricks has an upside for consumers. Such ferocity means better products and technology for data (e.g., Snowflake’s Polaris is now an Apache project) and eventually (hopefully) lower prices. It’s not like the previous pissing match between Oracle and SalesForce CEOs, where it was two rich guys taking pop shots at each other during their expensive conferences.

Shoving Ducks Into Everything 

In the same way that Postgres is the default choice for anyone starting a new operational database, DuckDB has entered the zeitgeist as the default choice for someone wanting to run analytical queries on their data. Pandas previously held DuckDB's crowned position. Given DuckDB's insane portability, there are several efforts to stick it inside existing DBMSs that do not have great support for OLAP workloads. This year, we saw the release of four different extensions to stick DuckDB up inside Postgres.

The first announcement came in May 2024 when Crunchy Data revealed their proprietary bridge for rewiring Postgres to route OLAP queries to DuckDB. They later announced an expanded version of their extension to leverage DuckDB's geospatial capabilities to accelerate PostGIS queries.

In June 2024, ParadeDB announced their open-source extension (pg_analytics) that uses Postgres' foreign data wrapper API to call into DuckDB; they previously were using DataFusion in an earlier version (pg_lakehouse) but switched to the Duck.

Then, in August 2024, the next DuckDB-for-Postgres extension (pg_duck) came out. The source code for this extension is hosted under the DuckDB Labs GitHub organization. As such, this is the officially sanctioned DuckDB extension for Postgres. The original announcement touted this project as being a collaboration between MotherDuck, Hydra, Microsoft, and Neon. The latter two were (allegedly) kicked out of the mix over a dispute on development controls, similar to Arabian Prince leaving NWA. The repository now only lists it as a joint effort between MotherDuck and Hydra.

The latest DuckDB extension dropped in November 2024 with pg_mooncake. Mooncake differs from the other three because it supports writing data through Postgres into Iceberg tables with full transaction support.

Andy’s Take:

Most OLAP queries do not access that much data. Fivetran analyzed traces from Snowflake and Redshift and showed that the median amount of data scanned by queries is only 100 MB. Such a small amount of data means a single DuckDB instance is enough for most to handle most queries.

DuckDB's convenience and portability are the reasons for its proliferation in the Postgres community. Although ClickHouse has existed since 2016, it was not as easy to run as DuckDB until recently (see this blog article that discusses the steps to deploy ClickHouse in 2018). These DuckDB extensions are a single entry point to the broader data ecosystem. Users no longer need to install separate extensions to access data in Iceberg and separately for S3. DuckDB can handle all of that for you. It allows organizations to gain high-performance analytics without needing an expensive data warehouse.

Postgres' support for extensions and plugins is impressive. One of the original design goals of Postgres from the 1980s was to be extensible. The intention was to easily support new access methods and new data types and operations on those data types (i.e., object-relational). Since 2006, Postgres' "hook" API. Our research shows that Postgres has the most expansive and diverse extension ecosystem compared to every other DBMS. We also found that the DBMS's lack of guard rails means that extensions can interfere with each other and cause incorrect behavior.

Earlier projects that added columnar storage to Postgres (e.g., Citus, Timescale) only solved part of the problem. Columnar data formats improve retrieving data from storage. However, a DBMS cannot fully exploit those formats if it still uses a row-oriented query processing model (e.g., Postgres). Using DuckDB provides both columnar storage and vectorized query processing.

There is likely a turducken joke here involving an elephant, but I will not make it because I do not want to get fired or put on probation by the university (again).

Random Happenings 

Many one-off events happened with databases last year that you might have overlooked. Here is a quick summary of them:

Releases:

Amazon Aurora DSQL

There isn't much public information yet about how AWS implemented their new "Spanner-like" DBMS (see Mark Brooker's discussion about the DBMS architecture). The key ideas are a distributed log service (rumors were it was going to be based on now-defunct QLDB) and timestamp ordering via Time Sync. But this announcement shows you how much brand recognition the name "Aurora" carries in the database world because AWS used it for this new DBMS that seemingly shares no code with their flagship Aurora Postgres RDS offering.

CedarDB

Umbra is one of the most state-of-the-art DBMSs written by the world's greatest database systems researcher (Thomas Neumann). But Thomas is content with staying at his university to work on Umbra, remaining comfortably on top of the Clickbench leaderboard, and not worrying about pesky customers. That's why his top Ph.D. students forked his code and are commercializing it as CedarDB.

Google Bigtable

The only interesting part of this announcement is that the former vanguard of the NoSQL movement in the late 2000s now supports SQL in 2024.

Limbo

Turso has been working on the libSQL fork of SQLite for a while, but they went all out in 2024 by announcing a complete rewrite of SQLite in Rust. In their announcement, they correctly point out that the value of SQLite is not just from its code, but also from the insane test engineering that ensures it runs correctly everywhere. That is why the Limbo developers are working with a deterministic testing startup by ex-FoundationDB people. See FoundationDB's 2020 CMU-DB talk for more information on what this testing means.

Microsoft Garnet

This key-value store is the successor to the impressive FASTER system from Microsoft Research. It is compatible with Redis and supports inter-query parallelism, larger-than-memory databases, and real transactions. Redis should not be anybody's first choice these days.

MySQL v9

Six years after MySQL v8 went GA, the team turned v9 out on the streets. But people quickly found that it crashed if your database had more than 8000 tables. I am underwhelmed with the feature list in this new major version. Oracle is putting all its time and energy into its proprietary MySQL Heatwave service. MySQL is still widely used, but the excitement is not there anymore. Everyone has moved on to Postgres.

Prometheus v3

It has been seven years since the latest major version of Prometheus. There are so many compatible alternatives now that the OG Prometheus may not be the best option for some organizations.

Acquisitions:

Alteryx → Private Equity

I've never met anybody who uses Alteryx, and I don't have an opinion about them.

MariaDB → Private Equity

Hopefully the PE people buying the MariaDB Corporation can clean up the mess. See my analysis from last year about the MariaDB dumpster fire.

OrioleDB → Supabase

This purchase makes sense if you are one of the leading Postgres ISVs. Postgres has a great front-end but an outdated storage architecture. OrioleDB fixes that problem.

PeerDB → ClickHouse

Better ETL tooling to get data out of Postgres and into ClickHouse. This is a smart move by ClickHouse, Inc.

PopSQL → Timescale

They bought themselves a fancy SQL editor UI. It is a quality-of-life improvement.

Speedb → Redis Ltd.

See the discussion above. They are likely going to use Speedb to allow Redis to spill data to disk. Speedb's developers never explained what changes and improvements they made in their RocksDB fork (or I could not find it?). See Mark Callaghan's recent comparison of Speedb vs. RocksDB.

Rockset → OpenAI

This is big news for the company, but unfortunately they had to shutdown the DBaaS in September 2024. Rockset had a great engineering team with some of the best database engineers from Facebook. I just never liked how their DBMS stored three copies of your data in its indexes.

Tabular → Databricks

Again, see the discussion above. Iceberg is the standard (sorry Hudi); even Amazon S3 now supports it. It remains to be seen how the adoption of Polaris will evolve and whether they will be able to maintain compatibility in the long term.

Verta.ai → Cloudera

I guess Cloudera is still alive?

Warpstream → Confluent

Rewriting Kafka in golang but then making it spill to S3. I'm happy for the Warpstream team, but Confluent could have done this themselves.

Funding:

  • Databricks - $10 billion Series J
  • DBOS - $8.5 million Seed Round
  • LanceDB - $8 million Seed Round
  • SDF - $9 million Seed Round
  • SpiceDB - $12 million Series A
  • TigerBeetle - $24 million Series A

There are a few more raises from CedarDB, SpiralDB, and others but those amounts are not public yet.

Deaths:

  • Amazon QLDB
    If Amazon can't figure out how to make money on a blockchain database, then nobody can. And yes, I know QLDB is not a true P2P blockchain, but it's close enough.
  • OtterTune
    Dana, Bohan, and I worked on this research project and startup for almost a decade. And now it is dead. I am disappointed at how a particular company treated us at the end, so they are forever banned from recruiting CMU-DB students. They know who they are and what they did.

I want to also give special props to Andres Freund for discovering the xz backdoor in 2024 while working on Postgres at Microsoft. This attack was a two-year campaign to inject malicious code into an important compression library widely used in computing. Although the backdoor targeted SSH and not Postgres directly, it is another example of why database engineers are some of the best programmers in the world.

Andy’s Take:

Databricks has blown away all other fundraising in the world of databases for the second year in a row with a disgustingly brash $10 billion Series J round. This is after their $500 million Series I in 2023 and $1.6 billion Series H in 2021. What is different about this time is this funding was for buying stock from employees who were getting impatient about Databricks' inevitable IPO. CMU-DB has several alumni at Databricks, including a former #1 ranked Ph.D. student. I know some of them are anxiously awaiting the Databricks IPO before deciding what to do next.

The upcoming year is going to be the test of strength for many database startups. Nobody wants to be the next MariaDB Corporation, and thus several are waiting to ride Databricks' wake before going IPO themselves. Declining interest rates in the upcoming year may open up additional funding for several database companies that have raised large amounts more than two years ago (e.g., CockroachDB, Starburst, Imply, DataStax, SingleStore, Firebolt). The one standout from this crowd is dbtLabs, which I have heard is comfortably crushing it.

See also the Database of Databases list of new DBMSs released in 2024.

Can’t Stop, Won’t Stop 

Do you know who had their 80th birthday this year? The legendary Larry Ellison! Once again, we see that he is a man who refuses to settle down or be put into a box. First, Larry propelled himself up Forbes' Billionaires List to become the third richest person in the world. In March 2024, the Oracle stock rose so much that he made $15 billion in a single day. Flush with cash, Larry went shopping in July 2024 and signed a deal to purchase Paramount Studios at $6 billion for his only son (third wife). He then decided to relax by buying a Palm Beach resort for only $277 million. These moves happened in just one year, and databases paid for all of it. But these are mere trifles compared to Larry's most significant accomplishment in 2024.

Everyone I know was surprised when our Larry Ellison news alerts woke us up in the middle of the night in November 2024. The headlines were touting how Larry helped the University of Michigan football program recruit the premiere college quarterback. The university had previously announced that this player was transferring from Louisiana State to Michigan. Their press release included a curious acknowledgement to "Larry and his wife Jolin" for helping with the recruiting effort. Reporters soon confirmed that this "Larry" was the one and only Larry Ellison! Larry contributed $12 million to the booster campaign to bankroll the best quarterback's move to Michigan.

The bigger mystery in this story was the identity of this "Jolin" person. Investigators found older photos of Larry watching a tennis match with a woman wearing a Michigan hat. Then, two weeks later, a major news organization broke the story at 5:30 am (my alerts woke me up again) that the woman's identity was Jolin (Keren) Zhu and they confirmed that she was Larry's new wife.

Andy’s Take:

I am beaming with pride over what Larry accomplished in the past year. He famously did not graduate from any university and had no prior connection to the University of Michigan. And yet, because the love of his life went to Michigan about a decade ago, Larry made magic happen by writing a check for a measly $12 million (about 0.0055% of his net worth). I told Larry this especially means a lot to me because my former #1 ranked Ph.D. student is now a professor in Michigan's Computer Science department with their famous Database Group.

What is even more fantastic about this story is that Larry is again in love! Too many people struggle in today's world to find that special somebody. Dating apps are a mess, speed dating events are awkward, and it is now considered uncouth to hang around a playground to meet single parents when you do not have kids of your own. Then, just when you think you finally found the right person, it all falls apart when you learn they do not wash their socks regularly or like to put hot sauce on cold cereal. That is why everyone was telling me that Larry would never get married again after his 2010 divorce from romance novelist Melanie Craft (fourth wife). Those people were telling me the same thing after his 2020 divorce from Nikita Kahn (fifth wife). But I knew better, and Larry proved me right with his surreptitious marriage to Keren Zhu (sixth wife)!

Conclusion

I was planning on starting this article boasting how this is the first time in three years where I was celebrating NYE not sick. But then my biological daughter gave me COVID so I'm laid up with that. I got boosted back in September and they gave me Paxlovid, so I'll survive this.

I am disappointed that OtterTune is dead. But I learned a lot and got to work with many brilliant people. I am a big fan of Intel Capital and Race Capital for sticking with us to the end. I hope to announce our next start-up soon (hint: it’s about databases).

In the meantime, I am happy to be back full-time at Carnegie Mellon University. Jignesh Patel and I have some baller research projects that we hope to turn out this upcoming year. I am also looking forward to teaching a new course on query optimization this semester. I need to figure out to juice my stats because in September 2024, Wikipedia removed the article about me over not having enough citations.

We are staying true to DJ Mooshoo while he is locked up in Cook County. We hope to free him in 2025.

Lastly, I want to give a shout-out to ByteBase for their article Database Tools in 2024: A Year in Review. In previous years, they emailed me asking for permission to translate my end-of-year database articles into Chinese for their blog. This year, they could not wait for me to finish writing this one, so they jocked my flow and wrote their own off-brand article with the same title and premise.

NoSQL — A tale of four sisters

Mike's Notes

I came across this blog post in one of the engineering newsletters I subscribe to. It is written by Ani and discusses the different types of "no SQL" databases. I have copied the article below, but the original has more pictures.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

17/05/2025

NoSQL — A tale of four sisters

By: Anirban "Ani" Goswami
Medium: 14/012022

The acronym NoSQL was first used in 1998 by Carlo Strozzi while naming his lightweight, open-source “relational” database that did not use SQL. The name came up again in 2009 when Eric Evans and Johan Oskarsson used it to describe non-relational databases. Relational databases are often referred to as SQL systems. The term NoSQL can mean either “No SQL systems” or the more commonly accepted translation of “Not only SQL,” to emphasize the fact some systems might support SQL-like query languages.

The rise of NoSQL is an important event in computer science and in application development because SQL has been so dominant for so long. Many other forms of database technology have come and gone, but few have had the wide adoption of NoSQL.

Companies are finding that they can apply NoSQL technology to a growing list of use cases while saving money in comparison to operating a relational database. NoSQL databases invariably incorporate a flexible schema model and are designed to scale horizontally across many servers, which makes them appealing for large data volumes or application loads that exceed the capacity of a single server.

The popularity of NoSQL has been driven by the following reasons:

  • The pace of development with NoSQL databases can be much faster than with a SQL database.
  • The structure of many different forms of data is more easily handled and evolved with a NoSQL database.
  • The amount of data in many applications cannot be served affordably by a SQL database.
  • The scale of traffic and need for zero downtime cannot be handled by SQL.
  • New application paradigms can be more easily supported.

The bloodline

There are various ways to classify NoSQL databases, with different categories and subcategories, some of which overlap. What follows is a non-exhaustive classification by data model, with examples.

(See the original blog post or Wikipedia article for a table of classifications.)

The most successful sisters

Well there are a lot of family members out there but out of them four has established them as the most successful family members of all time.

  • Wide-Column Store
  • Document Store
  • Key-Value Data Store
  • Graph Store

Wide-Column Store

A wide-column store (or extensible record store) is a type of NoSQL database which uses tables, rows, and columns, but unlike a relational database, the names and format of the columns can vary from row to row in the same table. A wide-column store can be interpreted as a two-dimensional key–value store. Google’s Bigtable and Apache Cassandra are one of the most eligible examples of wide-column store.

wide-column-store.png

Wide-Column Store — Where to use?

Wide-column databases are ideal for use cases that require a large dataset that can be distributed across multiple database nodes, especially when the columns are not always the same for every row. Examples are as following

  • Log data
  • IoT (Internet of Things) sensor data
  • Time-series data, such as temperature monitoring or financial trading data
  • Attribute-based data, such as user preferences or equipment features
  • Real-time analytics

Document Store aka Document DB

A document-oriented database is a specific kind of database that works on the principle of dealing with “documents” rather than strictly defined tables of information.

The central concept of a document store is that of a “document”. While the details of this definition differ among document-oriented databases, they all assume that documents encapsulate and encode data (or information) in some standard formats or encodings. Encodings in use include XML, YAML, and JSON and binary forms like BSON. Documents are addressed in the database via a unique key that represents that document. Another defining characteristic of a document-oriented database is an API or query language to retrieve documents based on their contents.

There’s very few example who are better than our own very favourite Mongo DB.

Different implementations offer different ways of organizing and/or grouping documents:

  • Collections
  • Tags
  • Non-visible metadata
  • Directory hierarchies

Compared to relational databases, collections could be considered analogous to tables and documents analogous to records. But they are different: every record in a table has the same sequence of fields, while documents in a collection may have fields that are completely different.

Document DB use cases

Document DB is widely used for storing product information and details by finance and e-commerce companies. You can even store the product catalogue of your brand in it. It can also be used to store and model machine-generated data, continuous stream ingestion etc.

Key-Value Data Store

A key-value store, or key-value database is a simple database that uses an associative array (think of a map or dictionary) as the fundamental data model where each key is associated with one and only one value in a collection. This relationship is referred to as a key-value pair.

In each key-value pair the key is represented by an arbitrary string such as a filename, URI or hash. The value can be any kind of data like an image, user preference file or document. The value is stored as a blob requiring no upfront data modeling or schema definition.

The storage of the value as a blob removes the need to index the data to improve performance. However, you cannot filter or control what’s returned from a request based on the value because the value is opaque.

The two finest implementations of key value stores are Amazon Dynamo DB and Azure Cosmos DB.

Use Case of Key-Value Data Store

Several popular use cases for key-value databases are

  • Web applications may store user session details and preference in a key-value store. All the information is accessible via user key, and key-value stores lend themselves to fast reads and writes.
  • Real-time recommendations and advertising are often powered by key-value stores because the stores can quickly access and present new recommendations or ads as a web visitor moves throughout a site.
  • On the technical side, key-value stores are commonly used for in-memory data caching to speed up applications by minimizing reads and writes to slower disk-based systems. Hazelcast is an example of a technology that provides an in-memory key-value store for fast data retrieval.

Graph Store

Graph databases are purpose-built to store and navigate relationships. Relationships are first-class citizens in graph databases, and most of the value of graph databases is derived from these relationships. Graph databases use nodes to store data entities, and edges to store relationships between entities. An edge always has a start node, end node, type, and direction, and an edge can describe parent-child relationships, actions, ownership, and the like. There is no limit to the number and kind of relationships a node can have.

A graph in a graph database can be traversed along specific edge types or across the entire graph. In graph databases, traversing the joins or relationships is very fast because the relationships between nodes are not calculated at query times but are persisted in the database. Graph databases have advantages for use cases such as social networking, recommendation engines, and fraud detection, when you need to create relationships between data and quickly query these relationships.

Neo4J is one of the most popular graph databases in the world at this moment.

Graph Store Use Cases

As explained by the image above it is evident graph stores are highly used in social networking sites. It has solid usage in fraud detection, threat identification, vigilance, Recommendation Engines, Supply Chain Mapping etc.

Web Almanac 2024

Mike's Notes

Smashing Magazine's first issue for 2025 includes a helpful review of the Web Almanac. Here are some notes from it, including an excellent chapter summarising structured data.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Smashing Magazine
  • Home > Handbook > 

Last Updated

17/05/2025

Web Almanac 2024

By: Iris Lješnjanin
Smashing Magazine: Issue 489, 7 Jan 2025

Looking back, what was the state of the web in 2024? The Web Almanac combines the raw stats and trends of the HTTP Archive with the expertise of the web community. The results are backed up by accurate data taken from 86 people who have volunteered countless hours in the planning, research, writing, and production phases.

With 16.9M websites tested, the 2024 edition is comprised of 19 chapters spanning aspects of page content, UX, publishing, and distribution. An insightful look into the state of performance is included, too. A massive Thank You to all of the contributors for their time and dedication!

Part I Chapter 3 Structured data

The Web Almanac 2024: 11 Nov 2024

Written by Andrea Volpini
Reviewed by Jarno van Driel and Ryan Levering
Analyzed by Nurullah Demir
Edited by James Gallagher

Table of Contents

  • Introduction
  • The expanding landscape of structured data
  • Key developments in 2023-2024:
  • Beyond traditional implementation
  • Structured data in the age of AI and machine learning
  • What this chapter provides
  • Key concepts
  • Linked data and the semantic web
  • Open data and the 5 stars model
  • AI-Powered search, voice assistants, and digital assistants
  • Semantic search engines and AI-powered search
  • The role of structured data
  • The shift to AI-powered search and its implications
  • Rich results and knowledge panels
  • Knowledge graphs and Graph RAG
  • Difference between Labeled Property Graphs and RDF graphs
  • Data Commons
  • Digital Product Passports and GS1 Digital Link
  • AI, machine learning, and structured data
  • Semantic SEO and data quality
  • A year in review
  • Structured data usage trends (2022-2024)
  • Comparison of JSON-LD, Microdata, and RDFa usage
  • RDFa
  • Dublin Core
  • Open Graph
  • Twitter
  • Facebook
  • Microformats and Microformats2
  • Microdata
  • JSON-LD
  • JSON-LD relationships
  • sameAs
  • JSON-LD context
  • Emerging trends and future outlook
  • Looking ahead: the future of structured data
  • Conclusion