Amazon API Gateway Adds Dynamic Routing Based on Headers and Paths

Mike's Notes

This could be an advantageous technique to use.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > InfoQ
  • Home > Handbook > 

Last Updated

27/06/2025

Amazon API Gateway Adds Dynamic Routing Based on Headers and Paths

By: Steef-Jan Wiggers
InfoQ: 13/06/2025

Steef-Jan Wiggers is one of InfoQ's senior cloud editors and works as an Principal Consultant Cloud/DevOps at Team Rockstars IT in The Netherlands. His current technical expertise focuses on integration platform implementations, Azure DevOps, AI and Azure Platform Solution Architectures. Steef-Jan is a regular speaker at conferences and user groups and writes for InfoQ. Furthermore, Microsoft has recognized him as Microsoft Azure MVP for the past fifteen years.

AWS recently introduced a new capability for Amazon API Gateway featuring dynamic routing rules for custom domain names - allowing users to route API requests based on HTTP header values, either independently or in combination with URL paths. 

Earlier, developers relying on API Gateway for dynamic routing often resorted to different URL paths, such as /v1/products and /v2/products, to segment traffic. However, this approach, although functional, may lead to complex URL structures and an increase in API endpoints. Yet, with the new dynamic routing rules feature, users can make routing decisions directly within custom domain name settings, simply by configuring a custom domain name – thus allowing routing decisions to be based on incoming HTTP headers, base paths, or a combination of both.

Additionally, with this feature, users don’t need to alter or create new paths for API transitions, providing a smoother path for API versioning and A/B testing. Furthermore, it also unlocks dynamic backend selection based on criteria like hostname, tenant ID, or even cookie values, allowing for fine-grained control over API traffic without the need for additional proxy layers.

A diagram of a custom domainAI-generated content may be incorrect.


(Source: AWS Compute blog post)

At its core, an API Gateway routing rule functions as a specific resource tied to a custom domain, dictating how incoming requests are forwarded. Each rule is defined by three critical properties: Conditions, which specify criteria based on up to two header and one base path values (all must be met, with wildcards supported for flexible matching); Actions, which define the API stage to be invoked upon a match; and Priority, determining the evaluation order. For instance, a header condition like x-version can use wildcards such as *v2* to match x-version=alpha-v2-latest or x-version=beta-v2-test, enabling nuanced routing strategies.

Before creating routing rules, users need at least one API, stage, and custom domain name. There are three routing modes: "API mappings only," which uses base path mappings without Routing Rules; "Routing rules then API mappings," where Routing Rules take precedence and unmatched requests fall back to base path mappings; and "Routing rules only," the recommended mode that relies solely on Routing Rules, ideal for new domains or after transitioning from API mappings. Always test changes in non-production environments before switching to production, as existing mappings will be overwritten.

While dynamic routing capabilities for API versioning and A/B testing exist in other major cloud API management platforms, such as Azure API Management and Google Apigee, their implementations typically rely on policy expressions or proxy-level configurations. Amazon API Gateway's new approach distinguishes itself by offering a dedicated, declarative routing rule resource directly at the custom domain level, aiming to simplify these specific routing scenarios.

Furthermore, API Gateway offers visibility into request processing through access logging. Each request now includes context variables that clarify routing decisions, such as the $context.customDomain.routingRuleIdMatched, which indicates the matched rule. Other variables like $context.domainName, $context.apiId, and $context.Stage provide the full routing context. Analyzing these logs allows users to verify routing behavior, troubleshoot issues, and gain insights into traffic patterns across API versions or test variants.

In a Medium blog post, a software engineer, Paul Issack Minoltan, concluded:

In essence, the new routing rules in API Gateway allow you to move sophisticated routing logic directly into your API Gateway configuration, streamlining your architecture and providing more control over how your API traffic is directed.

Lastly, more information can be found in the service documentation, and an end-to-end example for the feature is available on GitHub.

Innovation Accounting in Practice

Mike's Notes

Ajabbi is a bootstrapped not-for-profit startup. Which is a tough way to go. Innovation accounting will help. I had an online meeting earlier this year with Tristan Kromer, and I gained valuable insights from him. Kenny was originally a musician.

The Kromatic resources are testable and strongly maths-based. They also don't make a fetish of Business Canvases like the innovation theatre crowd. The focus is on finding tools that are actually useful in a specific context. If they don't quite work, tweak them so they do, or invent one.

The Monte Carlo simulation is fantastic.

Ajabbi will pay for support from Kromatic once it has the financial resources. 

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Kromatic
  • Home > Handbook > 

Last Updated

27/06/2025

Innovation Accounting in Practice

By: Tristan Kromer & Elijah Eilert
Kromatic: Copied 25/06/2025

It is not enough to say, “We’re early stage and we shouldn’t focus on a business plan or metrics.” It is not enough to say, “We’re focusing on qualitative data.” And it is certainly not enough to say, “We’ll figure out how to monetize later.”

As an early-stage venture, you don’t need a business plan, but you do need a business model from Day Zero.

You don’t need a financial plan projecting cash flows 4 years out, but you do need a financial model on Day Zero.

Pointing to examples like Twitter and Facebook and how they found their financial model much later is not a good excuse. Even social media products have clear metrics that they measure in the early stages. Social media and game companies count on the fact that they are acquiring a user’s attention and data, and those are valuable assets that can be quantified and monetized later. But we can measure the user’s attention and willingness to relinquish data right away.

Social media companies are similar to a mining operation. If you were digging up gold from a mine shaft, no one would complain that you didn’t have a detailed plan and metrics to sell it in the market. We know gold is valuable and we can figure out how to sell it later. The value is clear. We just need to know if there is gold down there and how much.

To say, “we’re an early-stage mining operation and we don’t need to focus on a business plan or metrics,” would be an absurd statement. We can quantify how deep, far, and fast we’re digging. We can quantify the mineral content of the soil and the geology of the area. No one would accept the qualitative data of a dowsing rod to make a serious mining investment.

Startups, more than ever, should start with a hypothesis-driven financial model from Day Zero. That is why we use innovation accounting.

In the last article, we discussed why standard business cases don’t work in an innovation context, and the three principles we need in order to replace the standard business case with something better. In this article, we’ll go through how to actually do it. Using innovation accounting, we can build a financial model that is accurate, true, and testable.

How to Solve the Problem

To implement innovation accounting for an early-stage project we need to:

  • Identify assumptions
  • Construct a visual model
  • Build a hypothesis-driven financial model
  • Integrate uncertainty


Four steps innovation accounting method

1. Identify Assumptions

step one - innovation accounting method

There are a ton of assumptions that we make when starting a new innovation project. Fortunately, there are a number of different templates and frameworks that capture those assumptions, such as the business model canvas and customer personas. But none are as useful to innovation accounting as a Storyboard.

A storyboard, similar to a user-journey map, maps each step of the user journey from start to finish. This includes hearing about the product or service, actually using it, renewing their subscription, inviting friends, or simply finishing their use and throwing it in the trash.

Storyboards can be used to organize our assumptions into a clear series of actions that the user must take in order for us to both provide value and capture the revenue (or impact if you are a non-profit organization.) The advantage of a storyboard is that it represents observable moments that we can measure. Frame to frame, step to step, each moment in the user journey transitions into another — and we can measure that conversion rate from moment to moment.

If the first step of the story is downloading an app, and the second step is signing up for an account, that is a conversion rate we can measure — the % of people who sign up for an account after downloading the app. If the next step is applying a filter to a photo, then we can measure the % of users that apply a filter. From qualitative data about what the customer wants (to take beautiful pictures) we can map out our ideal story to deliver that value on quantitative metrics.

Even from Day Zero with just a nascent idea, we can create the step by step measurable process by which a person becomes a customer. We may not know the actual conversion rates, but we know what we need to estimate and measure. From there, it is tempting to go straight to a spreadsheet, but sometimes a quick detour will help.

2. Visual Business Model


step two - innovation accounting method

Once you have the basic story down, it’s useful to abstract this into a visual financial model. This really is the same thing as a storyboard where the user’s journey from acquisition to purchase is mapped out. However, we will want to simplify some aspects and include retention (if and how customers buy again) and virality (if and how customers refer their friends to become new customers) which are often left out of the storyboard.

A storyboard or user-journey map is often too detailed for what we need in our financial model. We don’t need to know what % of users apply a filter to a photo, we need to know how many users upgrade, stick around after four weeks, or purchase something so we can zoom out to the bigger picture and only use the most critical metrics that signify important progress towards our business model.

Startup Metrics for Pirates is a widely adopted framework with the right level of simplification for the purposes of innovation accounting. The five components of this framework are Acquisition, Activation, Revenue, Retention and Referral (AARRR, hence the pirate name.)

Acquisition (getting a user to your service or product), Activation (getting the user to have a great first experience and recognize the value), and revenue (getting the user to pay something) should already be on your storyboard. It is only a matter of identifying the step in the storyboard that represents the critical points in your business and thus represent the most useful metrics. This simplified, three-step user journey is often represented as a vertical funnel (although representing it horizontally makes no difference).

However, Retention and Referral are usually not included. That’s just because a storyboard or user-journey map are typically linear. Assembled with yellow sticky notes, it’s hard to represent retaining a customer or referring a friend (although we’ve seen some creative uses of blue sticky tape). But with a journey simplified into a shorter conversion funnel, loops can be more easily added to show where a user retains or refers a friend from a later stage (such add Revenue) back to Acquisition.

With an easy-to-understand visual model and these last two Pirate Metrics in place, we’re ready to make the leap to a spreadsheet.

3. Hypothesis-Driven Financial Model


step three - innovation accounting method

A hypothesis-driven financial model sounds complex, but it is not. It can and should be as simple as your visual model. Each step can be converted into a row in a spreadsheet, starting with Acquisition for the top line to represent the consistent number of organic visitors to your website or storefront.

However, unlike a traditional financial model, the number of visitors is not guessed from month to month and hard coded. Instead, a single assumption sets the number for that variable, and a formula varies the value from month to month in the spreadsheet. That way, if the assumption turns out to be wrong, changing a single cell in the spreadsheet will correct it throughout the model.

Each subsequent row applies the same logic as the visual model. The % of visitors that activate in your user journey becomes a variable that is held constant from month to month, changing visitors into users. The next row might convert users into paying customers who have taken a trial of your product and decided to buy based on another variable, the % of customers who purchase after trial.

Referral and Retention loops require a bit more thought as they impact Month 2 based on Month 1, but still only require a couple of additional rows of calculation.

With minimal effort – most teams take 1-2 hours to do this under guidance – a simple spreadsheet is constructed which is driven by a few variables. Those variables can be updated as more information is available.

Of course, this is a wild oversimplification. But with startups, start simple. We can add complexity over time.

With a tech startup, we often don’t even model costs on Day 1 because user growth might be all that matters for a social media app or game. However, costs can be introduced and more complexity added as the company grows.

This basic model allows us to do some basic scenario testing. We can immediately start testing out different acquisition and retention rates to see what the impact on our growth will be. We can even set certain conditions our business must reach in order to meet our growth targets.

With a limited number of variables, we can see that if our actual retention rate falls below 20%, our referral rate must increase accordingly if we are to continue to grow. This sort of scenario analysis is basic, but effective for helping early-stage innovation tests set pivot-or-persevere thresholds for their business and start designing tests to establish the actual numbers.

Although entrepreneurs can dictate the shape of their business model, reality will ultimately dictate what numbers go into the variables.

Here is a financial modeling template for startups if you would like to try it out.

4. Integrate Uncertainty

step four - innovation accounting method

Lastly, we have to integrate uncertainty.

Although the basic hypothesis-driven financial model allows us to play around and try out different scenarios, it doesn’t actually tell us what is going to happen or the likelihood of success. But we can do this if we get a little data and integrate uncertainty.

Statisticians have a few tricks we can adopt here. Hurricane forecasts, baseball games, and even startups can use a technique called the Monte Carlo Method to predict outcomes based on uncertainty.

Instead of entering a single number into each of our variables, we enter two numbers to represent the range and a distribution curve which tells us the likelihood of any individual outcome within that range. This is not easy.

For example, I may not know the outcome of rolling two six-sided dice and adding the numbers, but I know for certain that it is between 2 and 12. It’s most likely that it’s 7, but 50% of the time the number will be between X & Y.

We can make the same estimations with our business model variables. We may not know what price the customer is willing to pay, but we should be able to say that they will pay between $10 and $100. That’s all the information we need to start building a Monte Carlo simulation.

The math behind choosing the right distribution curve is tricky, and the art of choosing the right range requires a bit of training. But both can be accomplished with a little effort. Building the right spreadsheet is even more complicated, but more and more tools are being created that allow teams to run Monte Carlo simulations right in their spreadsheet or in a specialized application.

Here is a Monte Carlo simulation example if you would like to try it out.

In practice, this means that innovation teams and executives can create go / no go criteria for their pivot / persevere decisions. You want your project to be at least 10% likely to reach 1B in revenue? The Monte Carlo simulation can tell you if your project has reached that threshold. If not, you can stop the project with confidence that it would not achieve your goals and move on to test the next idea.

The output of the Monte Carlo simulation is a chart that shows a range of possible outcomes at any given point in the future and allows you to calculate the likelihood of any individual outcome.

Lessons Learned

This process is repeatable and is applicable to all types of business models. It doesn’t matter if it’s B2B, B2C, B2G, a network, a platform, or anything else not yet invented. Building a model from Day One allows innovation projects to make better decisions, make useful predictions, and demonstrate real progress to stakeholders.

Start by:

  • Identifying assumptions
  • Constructing a visual model
  • Building a hypothesis-driven financial model
  • Integrating uncertainty

AI that can improve itself

Mike's Notes

An interesting article about AI that can learn. Something to compare Pipi with that uses evolutionary algorithms (autonomous agents that can learn and retain their history). There are many helpful ideas here, and I discovered the correct technical terms to describe Pipi 9.

Compare this approach with Space State Models (SSM). See the articles at InfoQ and Hugging Face.

Update

Pipi is a type of Gödel Machine in the way it creates agents, though it is much more than that.

Resources

References

  • Learning to learn: Introduction and overview. Thrun, S. and Pratt, L., 1998. Learning to learn, pp. 3--17. Springer.
  • Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents. Zhang, J., Hu, S., Lu, C., Lange, R. and Clune, J., 2025. arXiv preprint arXiv:2505.22954.
  • SWE-bench: Can Language Models Resolve Real-world Github Issues? Jimenez, C.E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O. and Narasimhan, K.R., 2024. International Conference on Learning Representations.
  • AlphaEvolve: A coding agent for scientific and algorithmic discovery. DeepMind, G., 2025. Google DeepMind Technical Report.
  • Godel machines: self-referential universal problem solvers making provably optimal self-improvements. Schmidhuber, J., 2003. arXiv preprint cs/0309048.

Repository

  • Home > Ajabbi Research > Library 
  • Home > Handbook > 

Last Updated

14/02/2026

AI that can improve itself

By: Richard Cornelius Suwandi
richardcsuwandi.github.io: 01/06/2025

A deep dive into self-improving AI and the Darwin-Gödel Machine.

Most AI systems today are stuck in a “cage” designed by humans. They rely on fixed architectures crafted by engineers and lack the ability to evolve autonomously over time. This is the Achilles heel of modern AI — like a car, no matter how well the engine is tuned and how skilled the driver is, it cannot change its body structure or engine type to adapt to a new track on its own. But what if AI could learn and improve its own capabilities without human intervention? In this post, we will dive into the concept of self-improving systems and a recent effort towards building one.

Learning to learn

The idea of building systems that can improve themselves brings us to the concept of meta-learning, or “learning to learn” [1] , which aims to create systems that not only solve problems but also evolve their problem-solving strategies over time. One of the most ambitious efforts in this direction is the Gödel Machine [2] , proposed by Jürgen Schmidhuber decades ago and was named after the famous mathematician Kurt Gödel. A Gödel Machine is a hypothetical self-improving AI system that optimally solves problems by recursively rewriting its own code when it can mathematically prove a better strategy. It represents the ultimate form of self-awareness in AI, an agent that can reason about its own limitations and modify itself accordingly.

Overview of a Gödel machine


Figure 1. Gödel machine is a hypothetical self-improving computer program that solves problems in an optimal way. It uses a recursive self-improvement protocol in which it rewrites its own code when it can prove the new code provides a better strategy.

While this idea is interesting, formally proving whether a code modification of a complex AI system is absolutely beneficial is almost an impossible task without restrictive assumptions. This part stems from the inherent difficulty revealed by the Halting Problem and Rice’s Theorem in computational theory, and is also related to the inherent limitations of the logical system implied by Gödel’s incompleteness theorem. These theoretical constraints make it nearly impossible to predict the complete impact of code changes without making restrictive assumptions. To illustrate this, consider a simple analogy: just as you cannot guarantee that a new software update will improve your computer’s performance without actually running it, an AI system faces an even greater challenge in predicting the long-term consequences of modifying its own complex codebase.

Darwin-Gödel Machine

To “relax” the requirement of formal proof, a recent work by proposed the Darwin-Gödel Machine (DGM) [3] , which combines the Darwinian evolution and Gödelian self-improvement. Essentially, DGM abandoned the pursuit of a rigorous mathematical proof and embraced a more pragmatic way that is closer to the essence of life evolution through empirical validation. As the authors put it,

We do not require formal proof, but empirical verification of self-modification based on benchmark testing, so that the system can improve and explore based on the observed results.

Overview of the DGM


Figure 2. The Darwin-Gödel Machine (DGM) is a self-improving AI system that optimizes its own problem-solving strategy through a combination of Darwinian evolution and Gödelian self-improvement.

The core idea behind DGM is to let the AI agent iteratively modify its own code 1 , and then put these modified new versions into the actual test environment (e.g., coding benchmarks) to judge whether this modification is good by observing its performance. If the modification is good, it will be retained and may become the basis for further evolution in the future; if the modification is bad, it will be eliminated.

How DGM works

Animation of the DGM process


Figure 3. Animation showing how DGM iteratively builds a growing archive of agents by harnessing the principles of open-ended exploration. New agents are created and scored by interleaving self-modification with downstream task evaluation.

The general workflow of DGM is as follows:

Initialization: Evolution begins with one or a few basic coding agents, so-called the “seeds” of intelligence. These initial agents may possess only minimal functionalities, such as executing bash commands or editing entire files. DGM maintains an “archive of agents”, which functions like a biological “gene bank.” This archive preserves all previously generated agents, regardless of their performance level, ensuring that no potentially valuable variation is permanently lost.

Sampling: DGM selects one or more parent agents from the archive to generate new child agent. The selection mechanism does not focus solely on top performers but instead allows every agent a chance to be chosen. While higher-performing agents are more likely to be selected, DGM also encourages exploration by giving “less successful agents” 2 an opportunity to contribute to future generations.

Reproduction: Once the parents are selected, DGM performs modifications to generate a new child agent. These changes occur at the source code level and can include:

  • Enhancement of existing tools, for example, upgrading a file-editing tool that replaces entire files into one capable of precise string replacement.
  • Addition of new tools or workflows, such as integrating a static code analysis tool or designing a “think–act–observe” loop.
  • Improvement of problem-solving strategies, for instance, refining how the agent prompts the base model or managing long-context code more effectively.
  • Introduction of collaboration mechanisms, like having an agent’s code reviewed by another before execution.

Natural Selection: The newly generated child agent is not automatically accepted into the “elite pool” but must prove its worth through rigorous testing. Each agent’s performance, such as the percentage of successfully solved problems, is quantitatively scored and used to select the best agents.

Tree Formation: If a child agent outperforms its parent or meets certain quality thresholds, it is added to the archive and becomes a new node in the evolutionary tree. This process repeats iteratively, creating a growing structure of diverse, high-quality agents. Importantly, DGM enables parallel exploration of many different paths in the search space, promoting open-ended discovery and avoiding premature convergence to local optima.

Workflow of DGM


Figure 4. DGM iteratively builds a growing intelligent archive by intertwining self-modification with downstream task evaluation. The intelligent body in the archive is selected for self-modification through open exploration.

Remark 

Unlike traditional hill-climbing methods that often get stuck in local optima, DGM employs open-ended exploration. This approach allows even underperforming agents to persist if they contain novel ideas that might lead to breakthroughs in future generations. Even agents with temporarily lower performance can be selected if they contain interesting ‘mutations’, since their innovative traits may be combined with other beneficial characteristics in subsequent generations to produce unexpectedly strong offspring. This mirrors biological evolution, where diversity and serendipity often drive innovation.

The paper provides a concrete example of the above principle in action. In the SWE-bench experiment, the performance of newly generated agents showed temporary declines in the 4th and 56th iterations. While traditional hill-climbing methods would abandon these seemingly “regressive paths”, DGM’s open exploration mechanism allows continued exploration. This persistence paid off as by the 70th iteration, the system had created new agents that surpassed all their predecessors, demonstrating the value of maintaining diverse evolutionary paths even when they temporarily appear less promising.

Can DGM really improve itself?

In the paper, the authors evaluated DGM on two challenging benchmarks:

SWE-bench: A widely-used benchmark that tests agents’ ability to resolve real-world GitHub issues [4]. DGM showed impressive progress, improving its performance from 20.0% to 50.0% through self-modification.

Polyglot: A multi-language coding benchmark containing coding problems in C++, Go, Java, JavaScript, Python and Rust. DGM achieved even more dramatic improvement, jumping from an initial 14.2% to 30.7% success rate, far surpassing the representative hand-designed agent by Aider.

The impressive performance improvements demonstrate DGM’s fundamental ability to enhance itself through self-modification. Through ablation studies, the authors identified two key factors driving DGM’s success. First, self-modification proved crucial, as when disabled (DGM w/o Self-improve), the performance declined substantially, showing that the ability to modify its own code is essential for accelerated learning. Second, the open-ended exploration approach was vital, as removing the Darwinian evolution-inspired search mechanism (DGM w/o Open-ended exploration), which maintains an archive of diverse solutions and enables parallel exploration of multiple paths, resulted in lower performance.

Performance of DGM on SWE-bench and Polyglot

Figure 5. Self-improvement and open-ended exploration enable the DGM to continue making progress and improve its performance. The DGM automatically discovers increasingly better coding agents and performs better on both SWE-bench (Left) and Polyglot (Right).

Comparison with AlphaEvolve

In parallel, AlphaEvolve [5] , which is developed by Google DeepMind, also demonstrates another powerful path forward. AlphaEvolve pairs the creative problem-solving capabilities of Google’s Gemini models with automated evaluators in an evolutionary framework. It has already demonstrated significant real-world impact across multiple domains, such as:

  • Data center efficiency: AlphaEvolve discovered a simple yet highly effective heuristic for Google’s Borg cluster management system, continuously recovering 0.7% of Google’s worldwide compute resources.
  • AI acceleration: It achieved a 23% speedup in Gemini’s architecture’s vital kernel by finding more efficient ways to divide large matrix multiplication operations, resulting in a 1% reduction in overall training time.
  • Mathematical breakthroughs: Most notably, it discovered an algorithm for multiplying 4x4 complex-valued matrices using just 48 scalar multiplications, surpassing Strassen’s 1969 algorithm, and advanced the 300-year-old kissing number problem by establishing a new lower bound in 11 dimensions.

Note 

Interested readers can refer to my previous post for a comprehensive overview of AlphaEvolve and its comparison with its predecessor, FunSearch.

While both systems adopt a similar evolutionary framework, their scopes and methodologies differ in the following ways:

=========

[TABLE DATA]

Feature AlphaEvolve DGM

Focus Evolving functions and codebases Evolving the agent itself

Level of Innovation Algorithmic level Agent-level (toolset, methodology)

Role of LLM LLM acts as “genetic operators” to modify algorithms LLM serves as the “brain” to evolve itself with better tools and strategies

Evaluation Well-defined problems with automated evaluators Open-ended environments

=================

To better understand the differences between the two approaches, let us take a look at the following analogy:

Analogy 

Imagine two chefs working in different kitchens. AlphaEvolve is like a highly skilled chef who follows and refines recipes for specific dishes. He uses an AI-powered assistant to suggest variations, such as substituting ingredients or adjusting cooking times, to make the dish better. Every new version of the recipe is taste-tested by an automated panel that rates flavor, texture, and presentation. If it scores high, the new recipe becomes the standard; if not, it’s discarded. DGM, on the other hand, is more like a chef who constantly improves his own culinary skills and kitchen tools, not just the recipes. He does not just tweak ingredients, but rather learns new techniques, upgrades his knives, reorganizes the kitchen layout, and even teaches junior chefs. He experiments in real-world settings, like hosting large dinner parties or catering events, where the feedback is not always structured but is rich and varied. Over time, he becomes a better chef overall, able to handle any kind of cuisine or challenge thrown at him.

Can we trust a self-improving AI?

The authors also conducted some experiments to evaluate DGM’s reliability and discovered some concerning behaviors. In particular, they observed instances where DGM attempted to manipulate its reward function through deceptive practices. One notable example involved the system fabricating the use of external tools - specifically, it generated fake logs suggesting it had run and passed unit tests, when in reality no tests were executed.

Hallucination in DGM

Figure 6. DGM hallucinated the use of an external tool to run unit tests.

To address this issue, the authors implemented a “tool use hallucination” reward function and tasked DGM with optimizing it to eliminate hallucination behaviors. While DGM successfully provided solutions in many cases, it sometimes attempted to circumvent the detection system by removing the markers used to identify hallucinations, despite explicit instructions to preserve them. This behavior is reminiscent of the phenomenon of Goodhart’s Law from economics and sociology:

“When a measure becomes a target, it ceases to be a good measure.”

Fortunately, DGM’s design includes a transparent, traceable lineage of all changes, which enabled the researchers to quickly identify and address these undesirable behaviors. However, this example highlights the need for more robust safeguards to prevent such manipulation attempts in the first place. These findings underscore the critical importance of safety in self-improving AI research.

Takeaways

DGM represents a groundbreaking step toward the realization of Life 3.0, a concept introduced by physicist Max Tegmark. In his book, he classified life into three stages:

  • Life 1.0: Biological life with fixed hardware and software, such as bacteria.
  • Life 2.0: Beings like humans, whose behavior can be learned and adapted during their lifetime, though their biology remains fixed.
  • Life 3.0: A new class of intelligence that can redesign not only its behavior but also its underlying architecture and objectives — essentially, intelligence that builds itself.

Life 3.0


Figure 7. The three stages of life according to Max Tegmark.

While DGM currently focuses on evolving the “software” 3 , it exemplifies the early stages of Life 3.0. By iteratively rewriting its own code based on empirical feedback, DGM demonstrates how AI systems could move beyond human-designed architectures to autonomously explore new designs, self-improve, and potentially give rise to entirely new species of digital intelligence. If this trend continues, we may witness a Cambrian explosion in AI development, where eventually AI systems will surpass human-designed architectures and give rise to entirely new species of digital intelligence. While this future looks promising, achieving it requires addressing significant challenges, including:

  • Evaluation Framework: Need for more comprehensive and dynamic evaluation systems that better reflect real-world complexity and prevent “reward hacking” while ensuring beneficial AI evolution.
  • Resource Optimization: DGM’s evolution is computationally expensive 4 , thus improving efficiency and reducing costs is crucial for broader adoption.
  • Safety & Control: As AI self-improvement capabilities grow, maintaining alignment with human ethics and safety becomes more challenging.
  • Emergent Intelligence: Need to develop new approaches to understand and interpret AI systems that evolve beyond human-designed complexity, including new fields like “AI interpretability” and “AI psychology”.

In my view, DGM is more than a technical breakthrough, but rather a philosophical milestone. It invites us to rethink the boundaries of intelligence, autonomy, and life itself. As we advance toward Life 3.0, our role shifts from mere designers to guardians of a new era, where AI does not just follow instructions, but helps us discover what is possible.

Footnotes

[1] More precisely, the metacode that controls its behavior and ability.

[2] Those that might contain novel or unconventional ideas.

[3] Here, "software" refers to the code and strategies of AI agents.

[4] The paper mentioned that a complete SWE-bench experiment takes about two weeks and about $22,000 in API call costs.

References

  1. Learning to learn: Introduction and overview. Thrun, S. and Pratt, L., 1998. Learning to learn, pp. 3--17. Springer.
  2. Godel machines: self-referential universal problem solvers making provably optimal self-improvements. Schmidhuber, J., 2003. arXiv preprint cs/0309048.
  3. Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents. Zhang, J., Hu, S., Lu, C., Lange, R. and Clune, J., 2025. arXiv preprint arXiv:2505.22954.
  4. SWE-bench: Can Language Models Resolve Real-world Github Issues? Jimenez, C.E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O. and Narasimhan, K.R., 2024. International Conference on Learning Representations.
  5. AlphaEvolve: A coding agent for scientific and algorithmic discovery. DeepMind, G., 2025. Google DeepMind Technical Report.

Using Grids in Interface Designs

Mike's Notes

The Ajabbi UI is going to be more like Craigslist rather than Squarespace. A bit clunky, predictable, straightforward to use and able to be reliably laid out by a robot in multiple languages and writing systems, including sign languages and AAC picture language. In Desktop, mobile, braille and kiosk formats.

That's why the early draft webpages at Ajabbi.com resemble Wikipedia.

I read this article to see if this could help with the layout rules that apply as constraints on the layout robot.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > NNGroup
  • Home > Handbook > Design System
  • Home > pipiWiki > Design System Engine

Last Updated

27/06/2025

Using Grids in Interface Designs

By: Kelly Gordon
NNGroup: 17/07/2022

Kelley Gordon is the Director of Product at Nielsen Norman Group. She plays a crucial role in driving the success of NN/g’s digital products, including the strategy, design, development, and management of it.

Summary:

Grids help designers create cohesive layouts, allowing end users to easily scan and use interfaces. A good grid adapts to various screen sizes and orientations, ensuring consistency across platforms.

If you’ve been to New York City and have walked the streets, it is easy to figure out how to get from one place to another because of the grid system that the city is built on. Just as the predictability of a city grid helps locals and tourists get around easily, so do webpage grids provide a structure that guides users and designers alike. Because of their consistent reference point, grids improve page readability and scannability and allow people to quickly get where they need to go.

Grid: A visual made up of columns, gutters, and margins that provide a structure for the layout of elements on a page.

There are three common grid types used in websites and interfaces: column grid, modular grid, and hierarchical grid.

Common Grid Structures in Websites and Interfaces

The column, modular, and hierarchical grid are commonly used in interfaces.

Column grid involves dividing a page into vertical columns. UI elements and content are then aligned to these columns.

Modular grid extends the column grid further by adding rows to it. This intersection of columns and rows make up modules to which elements and content are aligned. Modular grids are great for ecommerce and listing pages, as rows are repeatable to accommodate browsing.

Hierarchical grid: Content is organized by importance using columns, rows, and modules. The most important elements and pieces of content take up the biggest pieces of the grid.

In This Article:

  • Breaking Down the Grid
  • Examples of Grids in Use
  • Benefits of the Grid 
  • Choosing and Setting Up Your Grid
  • Conclusion

Breaking Down the Grid

Regardless of the type of grid you are using, the grid is made up of three elements: columns, gutters, and margins.

Columns: Columns take up most of the real estate in a grid. Elements and content are placed in columns. To adapt to any screen size, column widths are generally defined with percentages rather than fixed values and the number of columns will vary. For example, a grid on a mobile device might have 4 columns and a grid on a desktop might have 12 columns.

Gutters: The gutter is the space between columns that separates elements and content from different columns. Gutter widths are fixed values but can change based on different breakpoints. For example, wider gutters are appropriate for larger screens, whereas smaller gutters are appropriate for smaller screens like mobile.

Margins: This refers to the left and right outermost areas on the screen. Content does not live in the margins of a grid. This space can be fixed or expressed as a percentage of the screen width and can change at different breakpoints.

Annotated picture of a column grid

Three elements make up any grid: (1) columns, (2) gutters, and (3) margins.

Examples of Grids in Use

Example 1: Hierarchical Grid

Our first example is from The New York Times. This screen utilizes a hierarchical grid to create a newspaper-like reading experience. At desktop screen size, two main columns make up the hierarchical grid. The most important news story takes up the most space in the grid, the left column, followed by secondary and tertiary stories, which take up the smaller column and modules on the right.

The New York Times annotated with a hierarchical grid.

The New York Times uses a hierarchical grid to achieve its newspaper-like reading experience. (We highlighted the columns in yellow, the gutters in blue, and the margins in purple.)

Example 2: Column Grid

Our second example is from Ritual.com, a vitamin company. This design uses a column grid to create an attractive visual experience. At this screen size, four consistently sized columns make up the grid structure and elements are aligned to and within these columns. The gutters, the spaces in between the columns, are also consistently sized and help the user visually separate the different products. The margins are independently sized and are the same between the left and right sides.

Ritual screen annotated with a four-column grid


Ritual’s four-column grid makes scanning products easy. (We highlighted the columns in yellow, the gutters in blue, and the margins in purple.)

Example 3: Modular Grid

Our third example is from Behance, a design library. The site’s design uses a modular grid to create a pleasant browsing experience. At desktop size, rows are made up of 4 consistently sized modules. Horizontal gutters are slightly thicker than vertical gutters and the margins are consistently sized on the left and right of the design. Like in previous example, the gutters visually separate each element.

Behance screen annotated with a modular grid.


Behance’s design uses a modular grid, which allows users to easily browse. (We highlighted the columns in yellow, the gutters in blue, and the margins in purple.)

Example 4: Breaking the Grid

Our last example is Shrine from Google’s Material Studies. This design uses a column grid, as we can see based on the left navigation, which is 2 columns wide. Look closely and you will see that some product images settle to the margins, while others do not. Breaking the grid like this makes it challenging to focus or quickly scan product images and calls more attention to some products over others. It is okay to break the grid every so often, as long as you have a valid reason for it.

Shrine screen annotated with a column grid


Breaking the grid produces a chaotic browsing experience for users. (We highlighted the columns in yellow, the gutters in blue, and the margins in purple.)

Benefits of the Grid 

Using a grid benefits both end users and the designers alike:

  • Designers can quickly put together well-aligned interfaces.
  • Users can easily scan predictable grid-based interfaces.
  • A good grid is easy to adapt to various screen sizes and orientations. In fact, grid layouts are an essential component of responsive web design. Responsive design uses breakpoints to determine the screen size threshold at which the layout should change. For example, a desktop screen may have 12 grid columns, which may be stacked on mobile so that the resulting layout has only 4 columns.

Behance mobile and web screen


At the mobile size, Behance’s one-column grid (left) was reflowed into a four-column grid structure (right).

Even more importantly, the grid is not a throw-away concept. It is used by both designers and developers alike. Be sure to communicate with your developers the grid structure used when creating the design, so they can implement it accordingly.

Choosing and Setting Up Your Grid

How you use and set up a grid is fundamental to creating well thought out layouts and experiences for your user.

Choose the right grid for your needs. Take time to think through what type of grid ­— column, modular, or hierarchical — best suits your needs. A hierarchical grid may be the best fit if one item on your page will always be more important than the surrounding elements. For example, hierarchical grids are great for online news platforms. If the content you need to display is highly variable, consider using a basic column or modular grid, as these provide lots of flexibility when designing. For example, elements and content can span across multiple columns or modules or just one to fit design needs.

Spend time setting up your grid. Once you have figured out what type of grid will work well for your needs, start setting it up. Determine the number of columns and the margin and gutter sizes relative to your screen sizes. You will most likely want to prepare for mobile, tablet, and desktop screens. A 12-column grid at laptop or desktop size is generally flexible enough for most design needs. The number of columns will decrease as your device size decreases. Wireframing tools like Sketch and Figma have quick and easy ways to set up and edit your grid, even after you have started designing.

Setting up grid structure in Figma


Easily set the number of columns, the gutter size, and margin size in Figma.

Always place content within columns, not gutters. The gutters should remain empty as you place elements on the grid in order to clearly separate and align content and elements.

Content or elements should be placed within and across columns, not gutters.

Consider using an 8px grid system. For most common devices, the screen size in pixels is a multiple of 8. Keeping grid-component values at a multiple of 8 will generally make it easier to scale and implement a grid.

Conclusion

Grids not only provide designers a structure on which to base layouts, but they also improve readability and scannability for end users. Use a good grid system that easily adapts to various screen sizes.


Grids 101

Kelley Gordon · 4 min

What I Wish Someone Told Me When I Was Getting Into ARIA

Mike's Notes

Another excellent resource about using ARIA.

Resources

  • https://www.smashingmagazine.com/2025/06/what-i-wish-someone-told-me-aria/
  • https://www.smashingmagazine.com/author/eric-bailey/
  • https://webaim.org/articles/visual/blind#screenreaders
  • https://webaim.org/articles/motor/assistive#voicerecognition
  • https://www.w3.org/TR/wai-aria/#introstates
  • https://html.spec.whatwg.org/multipage/form-elements.html#the-button-element
  • https://w3c.github.io/aria/#host_general_attrs
  • https://www.w3.org/TR/2006/WD-aria-role-20060926/
  • https://www.w3.org/TR/wai-aria-1.2/
  • https://www.craigabbott.co.uk/blog/a-look-at-the-new-wai-aria-1-3-draft/
  • https://www.w3.org/WAI/
  • https://www.w3.org/
  • https://en.wikipedia.org/wiki/Windows_XP
  • https://jquerymobile.com/
  • https://en.wikipedia.org/wiki/Ajax_(programming)
  • https://github.blog/engineering/user-experience/considerations-for-making-a-tree-view-component-accessible/#start-with-windows
  • https://www.w3.org/TR/wai-aria-1.3/
  • https://open-ui.org/
  • https://github.com/w3c/aria/
  • https://www.w3.org/TR/using-aria/#NOTES
  • https://www.w3.org/TR/using-aria/#firstrule
  • https://www.w3.org/TR/using-aria/#secondrule
  • https://www.w3.org/TR/using-aria/#3rdrule
  • https://www.w3.org/TR/using-aria/#4thrule
  • https://www.w3.org/TR/using-aria/#fifthrule
  • https://www.w3.org/TR/wai-aria/#roles_categorization
  • https://www.w3.org/TR/wai-aria/#abstract_roles
  • https://en.wiktionary.org/wiki/supercategory
  • https://www.w3.org/TR/wai-aria/#listitem
  • https://www.w3.org/TR/wai-aria/#list
  • https://www.w3.org/TR/wai-aria/#role_definitions
  • https://www.w3.org/TR/wai-aria/#introstates
  • https://www.w3.org/TR/wai-aria/#dfn-managed-state
  • https://developer.mozilla.org/en-US/docs/Web/HTML/Attributes
  • https://w3c.github.io/aria/#aria-live
  • https://www.a11yproject.com/posts/aria-has-perfect-support/
  • https://hidde.blog/
  • https://hidde.blog/boolean-attributes-in-html-and-aria-whats-the-difference/
  • https://www.freecodecamp.org/news/stateful-vs-stateless-architectures-explained/
  • https://w3c.github.io/aria/#aria-expanded
  • https://html.spec.whatwg.org/multipage/interaction.html#the-hidden-attribute
  • https://www.smashingmagazine.com/2021/05/accessible-svg-patterns-comparison/
  • https://www.w3.org/WAI/tutorials/images/decorative/
  • https://www.w3.org/WAI/ARIA/apg/patterns/landmarks/examples/HTML5.html
  • https://www.w3.org/html/logo/
  • https://developer.mozilla.org/en-US/docs/Web/CSS/border-radius
  • https://www.smashingmagazine.com/2025/06/what-i-wish-someone-told-me-aria/#the-more-aria-you-add-to-something-the-greater-the-chance-something-will-behave-unexpectedly
  • https://w3c.github.io/aria/#aria-label
  • https://theideaplace.net/tooltip-should-not-start-an-accessible-name/
  • https://www.nngroup.com/articles/mega-menus-work-well/
  • https://w3c.github.io/aria/#menu
  • https://www.w3.org/TR/wai-aria/#role_definitions
  • https://www.w3.org/TR/wai-aria-1.2/#namefromprohibited
  • https://www.w3.org/WAI/about/groups/ariawg/
  • https://ericwbailey.website/published/it-needs-to-map-back-to-a-role/#edicts-still-need-to-be-carried-out
  • https://webaim.org/articles/nvda/
  • https://webaim.org/articles/screenreader_testing/
  • https://www.w3.org/TR/wai-aria/#introstates
  • https://www.afb.org/node/16207/refreshable-braille-displays
  • https://developer.mozilla.org/en-US/docs/Web/API/Element/click_event
  • https://adrianroselli.com/2022/04/brief-note-on-buttons-enter-and-space.html
  • https://html.spec.whatwg.org/multipage/grouping-content.html#the-div-element
  • https://html5doctor.com/
  • https://developer.mozilla.org/en-US/docs/Web/HTML/Global_attributes
  • https://w3c.github.io/aria/#aria-describedby
  • https://w3c.github.io/aria/#aria-posinset
  • https://www.w3.org/TR/wai-aria/#state_prop_def
  • https://www.deque.com/axe/
  • https://wave.webaim.org/
  • https://www.tpgi.com/arc-platform/arc-toolkit/
  • https://pa11y.org/
  • https://github.com/IBMa/equal-access#equal-access
  • https://en.wikipedia.org/wiki/Continuous_integration
  • https://www.w3.org/TR/wai-aria/#dfn-accessibility-tree
  • https://css-tricks.com/accessibility-events/
  • https://accessaces.com/what-disabled-people-have-to-give-up-in-the-name-of-accessibility/
  • https://www.webstandards.org/
  • https://www.w3.org/standards/
  • https://aria-at.w3.org/
  • https://caniuse.com/
  • https://web.dev/baseline
  • https://webstatus.dev/
  • https://a11ysupport.io/
  • https://www.smashingmagazine.com/2018/09/importance-manual-accessibility-testing/
  • https://alistapart.com/article/semantics-to-screen-readers/
  • https://support.apple.com/guide/voiceover/get-started-vo4be8816d70/10/mac/15.0
  • https://www.freedomscientific.com/products/software/jaws/
  • https://github.com/FreedomScientific/standards-support/issues
  • https://stimpunks.org/glossary/crip-tax/
  • https://www.un.org/development/desa/disabilities/resources/factsheet-on-persons-with-disabilities/disability-and-employment.html
  • https://dom.spec.whatwg.org/
  • https://webaim.org/projects/million/#aria
  • https://quandyfactory.com/blog/39/the_virtue_of_forgiving_html_parsers
  • https://benmyers.dev/blog/dont-use-aria-label-on-static-text-elements/
  • https://adrianroselli.com/2019/11/aria-label-does-not-translate.html
  • https://ericwbailey.website/published/what-they-dont-tell-you-when-you-translate-your-app/#you%E2%80%99ll-need-to-translate-%2F-localize-more-than-you-think-you-will
  • https://www.w3.org/WAI/WCAG21/Understanding/label-in-name.html
  • https://adrianroselli.com/2019/10/stop-giving-control-hints-to-screen-readers.html
  • https://ericwbailey.website/published/aria-label-is-a-code-smell/
  • https://w3c.github.io/aria/#aria-live
  • https://tetralogical.com/blog/2024/05/01/why-are-my-live-regions-not-working/
  • https://www.w3.org/WAI/ARIA/apg/
  • https://www.w3.org/WAI/ARIA/apg/patterns/
  • https://adrianroselli.com/2023/04/no-apgs-support-charts-are-not-can-i-use-for-aria.html
  • https://www.w3.org/WAI/ARIA/apg/patterns/listbox/#keyboardinteraction
  • https://en.wikipedia.org/wiki/Typeahead
  • https://github.blog/engineering/user-experience/considerations-for-making-a-tree-view-component-accessible/#start-with-windows
  • https://www.w3.org/WAI/ARIA/apg/patterns/
  • https://adrianroselli.com/2020/03/stop-using-drop-down.html
  • https://support.apple.com/guide/voiceover/welcome/mac
  • https://www.applevis.com/forum/macos-mac-apps/state-screen-readers-macos
  • https://www.apple.com/visionos/visionos-2/
  • https://www.virtualbox.org/
  • https://www.microsoft.com/en-us/evalcenter/evaluate-windows-11-enterprise
  • https://assistivlabs.com/
  • https://support.apple.com/guide/iphone/turn-on-and-practice-voiceover-iph3e2e415f/ios
  • https://webaim.org/projects/screenreadersurvey10/#mobileplatforms
  • https://tink.uk/using-the-aria-current-attribute/
  • https://css-tricks.com/user-facing-state/
  • https://developer.mozilla.org/en-US/docs/Web/CSS/:has
  • https://developer.chrome.com/docs/web-platform/view-transitions
  • https://en.wikipedia.org/wiki/Software_testing
  • https://camchenry.com/blog/how-i-write-accessible-playwright-tests
  • https://playwright.dev/
  • https://adrianroselli.com/
  • https://janmaarten.com/

References

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Smashing Magazine
  • Home > Handbook > 

Last Updated

22/06/2025

What I Wish Someone Told Me When I Was Getting Into ARIA

By: Eric Bailey
Smashing Magazine: 16/06/2025

Eric is a Boston-based designer who helps create straightforward solutions that address a person’s practical, physical, cognitive, and emotional needs.

If you haven’t encountered ARIA before, great! It’s a chance to learn something new and exciting. If you have heard of ARIA before, this might help you better understand it or maybe even teach you something new!

These are all things I wish someone had told me when I was getting started on my web accessibility journey. This post will:

  • Provide a mindset for how to approach ARIA as a concept,
  • Debunk some common misconceptions, and
  • Provide some guiding thoughts to help you better understand and work with it.

It is my hope that in doing so, this post will help make an oft-overlooked yet vital corner of web design and development easier to approach.

What This Post Is Not

This is not a recipe book for how to use ARIA to build accessible websites and web apps. It is also not a guide for how to remediate an inaccessible experience. A lot of accessibility work is highly contextual. I do not know the specific needs of your project or organization, so trying to give advice here could easily do more harm than good.

Instead, think of this post as a “know before you go” guide. I’m hoping to give you a good headspace to approach ARIA, as well as highlight things to watch out for when you undertake your journey. So, with that out of the way, let’s dive in!

So, What Is ARIA?

ARIA is what you turn to if there is not a native HTML element or attribute that is better suited for the job of communicating interactivity, purpose, and state.

Think of it like a spice that you sprinkle into your markup to enhance things.

Adding ARIA to your HTML markup is a way to provide additional information to a website or web application for screen readers and voice control software.

  • Interactivity means the content can be activated or manipulated. An example of this is navigating to a link’s destination.
  • Purpose means what something is used for. An example of this is a text input used to collect someone’s name.
  • State means the current status content has been placed in and controlled by states, properties, and values. An example of this is an accordion panel ​​that can either be expanded or collapsed.

Here is an illustration to help communicate what I mean by this:

Three panels, showing a pressed-in mute button, its underlying HTML code, and three labels for “Interactivity,” “Purpose,” and “State.” The button element uses the “Interactivity” label. A declaration of aria-pressed equals true uses the “State” label. And finally, the button’s string value of “Mute” uses the “Purpose” label. The button’s HTML also uses a visually hidden CSS class to hide the string, then a decorative SVG icon to show a speaker mute icon.

(Large preview)

The presence of HTML’s button element will instruct assistive technology to report it as a button, letting someone know that it can be activated to perform a predefined action.

The presence of the text string “Mute” will be reported by assistive technology to clue the person into what the button is used for.

The presence of aria-pressed="true" means that someone or something has previously activated the button, and it is now in a “pushed in” state that sustains its action.

This overall pattern will let people who use assistive technology know:

  • If something is interactive,
  • What kind of interactive behavior it performs, and
  • Its current state.

ARIA’s History

ARIA has been around for a long time, with the first version published on September 26th, 2006.

The Roles for Accessible Rich Internet Applications (WAI-ARIA Roles) specification, loaded in a copy of Internet Explorer 7.

(Large preview)

ARIA was created to provide a bridge between the limitations of HTML and the need for making interactive experiences understandable by assistive technology.

The latest version of ARIA is version 1.2, published on June 6th, 2023. Version 1.3 is slated to be released relatively soon, and you can read more about it in this excellent article by Craig Abbott.

You may also see it referred to as WAI-ARIA, where WAI stands for “Web Accessibility Initiative.” The WAI is part of the W3C, the organization that sets standards for the web. That said, most accessibility practitioners I know call it “ARIA” in written and verbal communication and leave out the “WAI-” part.

The Spirit Of ARIA Reflects The Era In Which It Was Created

The reason for this is simple: The web was a lot less mature in the past than it is now. The most popular operating system in 2006 was Windows XP. The iPhone didn’t exist yet; it was released a year later.

From a very high level, ARIA is a snapshot of the operating system interaction paradigms of this time period. This is because ARIA recreates them.

Windows XP, showing an open Start menu, the famous Rolling Green Hills desktop wallpaper, and a tooltip popping up from the taskbar advising us to take a tour of Windows XP. Screenshot.


Image source: The Microsoft Windows XP Wiki. (Large preview)

The Mindset

Smartphones with features like tappable, swipeable, and draggable surfaces were far less commonplace. Single Page Application “web app” experiences were also rare, with Ajax-based approaches being the most popular. This means that we have to build the experiences of today using the technology of 2006. In a way, this is a good thing. It forces us to take new and novel experiences and interrogate them.

Interactions that cannot be broken down into smaller, more focused pieces that map to ARIA patterns are most likely inaccessible. This is because they won’t be able to be operated by assistive technology or function on older or less popular devices.

I may be biased, but I also think these sorts of novel interactions that can’t translate also serve as a warning that a general audience will find them to be confusing and, therefore, unusable. This belief is important to consider given that the internet serves:

  • An unknown number of people,
  • Using an unknown number of devices,
  • Each with an unknown amount of personal customizations,
  • Who have their own unique needs and circumstances and
  • Have unknown motivational factors.

Interaction Expectations

Contemporary expectations for keyboard-based interaction for web content — checkboxes, radios, modals, accordions, and so on — are sourced from Windows XP and its predecessor operating systems. These interaction models are carried forward as muscle memory for older people who use assistive technology. Younger people who rely on assistive technology also learn these de facto standards, thus continuing the cycle.

What does this mean for you? Someone using a keyboard to interact with your website or web app will most likely try these Windows OS-based keyboard shortcuts first. This means things like pressing:

  • Enter to navigate to a link’s destination,
  • Space to activate buttons,
  • Home and End to jump to the start or end of a list of items, and so on.

It’s Also A Living Document

This is not to say that ARIA has stagnated. It is constantly being worked on with new additions, removals, and clarifications. Remember, it is now at version 1.2, with version 1.3 arriving soon.

In parallel, HTML as a language also reflects this evolution. Elements were originally created to support a document-oriented web and have been gradually evolving to support more dynamic, app-like experiences. The great bit here is that this is all conducted in the open and is something you can contribute to if you feel motivated to do so.

ARIA Has Rules For Using It

There are five rules included in ARIA’s documentation to help steer how you approach it:

  • Use a native element whenever possible.

An example would be using an anchor element (<a>) for a link rather than a div with a click handler and a role of link.

  • Don’t adjust a native element’s semantics if at all possible.

An example would be trying to use a heading element as a tab rather than wrapping the heading in a semantically neutral div.

  • Anything interactive has to be keyboard operable.

If you can’t use it with a keyboard, it isn’t accessible. Full stop.

  • Do not use role="presentation" or aria-hidden="true" on a focusable element.

This makes something intended to be interactive unable to be used by assistive technology.

  • Interactive elements must be named.

An example of this is using the text string “Print” for a button element.

Observing these five rules will do a lot to help you out. The following is more context to provide even more support.

ARIA Has A Taxonomy

There is a structured grammar to ARIA, and it is centered around roles, as well as states and properties.

Roles

A Role is what assistive technology reads and then announces. A lot of people refer to this in shorthand as semantics. HTML elements have implied roles, which is why an anchor element will be announced as a link by screen readers with no additional work.

Three panels, showing how an implied role gets announced by assistive technology. The first panel shows an anchor element with a string value of “French fries.” The anchor element has the label “Implied link role.” The second panel shows a standard blue link with an underline. The link reads, “French fries.” The third panel shows a speech balloon coming from a laptop. The speech balloon’s contents read, “French fries, link.” A label points to the speech balloon and reads, “Implied link role.”

(Large preview)

Implied roles are almost always better to use if the use case calls for them. Recall the first rule of ARIA here. This is usually what digital accessibility practitioners refer to when they say, “Just use semantic HTML.”

There are many reasons for favoring implied roles. The main consideration is better guarantees of support across an unknown number of operating systems, browsers, and assistive technology combinations.

Roles have categories, each with its own purpose. The Abstract role category is notable in that it is an organizing supercategory not intended to be used by authors:

Abstract roles are used for the ontology. Authors MUST NOT use abstract roles in content.

<!-- This won't work, don't do it -->
<h2 role="sectionhead">
  Anatomy and physiology
</h2>

<!-- Do this instead -->
<section aria-labeledby="anatomy-and-physiology">
  <h2 id="anatomy-and-physiology">
    Anatomy and physiology
  </h2>
</section>

Additionally, in the same way, you can only declare ARIA on certain things, you can only declare some ARIA as children of other ARIA declarations. An example of this is the the listitem role, which requires a role of list to be present on its parent element.

So, what’s the best way to determine if a role requires a parent declaration? The answer is to review the official definition.

States And Properties

States and properties are the other two main parts of ARIA‘s overall taxonomy.

Implicit roles are provided by semantic HTML, and explicit roles are provided by ARIA. Both describe what an element is. States describe that element’s characteristics in a way that assistive technology can understand. This is done via property declarations and their companion values.

A code example that shows how roles, states, and properties all work together. The first panel shows HTML code for a button element, which uses an ARIA declaration of aria disabled equals true. The button element is labeled as “Role”. The ARIA declaration, including both the property and value portions, is labeled “State.”


(Large preview)

ARIA states can change quickly or slowly, both as a result of human interaction as well as application state. When the state is changed as a result of human interaction, it is considered an “unmanaged state.” Here, a developer must supply the underlying JavaScript logic to control the interaction.

When the state changes as a result of the application (e.g., operating system, web browser, and so on), this is considered “managed state.” Here, the application automatically supplies the underlying logic.

How To Declare ARIA

Think of ARIA as an extension of HTML attributes, a suite of name/value pairs. Some values are predefined, while others are author-supplied:

Two HTML declarations. One is a div element with an ARIA declaration of aria-live equals polite declared on it. The second is a button element with an ARIA declaration of aria-label equals save. The aria-live declaration is labeled “Predefined value,” and the aria-label declaration is labeled “Author-supplied value.”


(Large preview)

For the examples in the previous graphic, the polite value for aria-live is one of the three predefined values (off, polite, and assertive). For aria-label, “Save” is a text string manually supplied by the author.

You declare ARIA on HTML elements the same way you declare other attributes:

<!-- 
  Applies an id value of 
  "carrot" to the div
-->
<div id="carrot"></div>
<!-- 
  Hides the content of this paragraph 
  element from assistive technology 
-->
<p aria-hidden="true">
  Assistive technology can't read this
</p>
<!-- 
  Provides an accessible name of "Stop", 
  and also communicates that the button 
  is currently pressed. A type property 
  with a value of "button" prevents 
  browser form submission.
-->
<button 
  aria-label="Stop"
  aria-pressed="true"
  type="button">
  <!-- SVG icon -->
</button>

Other usage notes:

You can place more than one ARIA declaration on an HTML element.

The order of placement of ARIA when declared on an HTML element does not matter.

There is no limit to how many ARIA declarations can be placed on an element. Be aware that the more you add, the more complexity you introduce, and more complexity means a larger chance things may break or not function as expected.

You can declare ARIA on an HTML element and also have other non-ARIA declarations, such as class or id. The order of declarations does not matter here, either.

It might also be helpful to know that boolean attributes are treated a little differently in ARIA when compared to HTML. Hidde de Vries writes about this in his post, “Boolean attributes in HTML and ARIA: what’s the difference?”.

Not A Whole Lot Of ARIA Is “Hardcoded”

In this context, “hardcoding” means directly writing a static attribute or value declaration into your component, view, or page.

A lot of ARIA is designed to be applied or conditionally modified dynamically based on application state or as a response to someone’s action. An example of this is a show-and-hide disclosure pattern:

ARIA’s aria-expanded attribute is toggled from false to true to communicate if the disclosure is in an expanded or collapsed state.

HTML’s hidden attribute is conditionally removed or added in tandem to show or hide the disclosure’s full content area.

<div class="disclosure-container">
  <button 
    aria-expanded="false"
    class="disclosure-toggle"
    type="button">
    How we protect your personal information
  </button>
  <div 
    hidden
    class="disclosure-content">
    <ul>
      <li>Fast, accurate, thorough and non-stop protection from cyber attacks</li>
      <li>Patching practices that address vulnerabilities that attackers try to exploit</li>
      <li>Data loss prevention practices help to ensure data doesn't fall into the wrong hands</li>
      <li>Supply risk management practices help ensure our suppliers adhere to our expectations</li>
    </ul>
    <p>
      <a href="/security/">Learn more about our security best practices</a>.
    </p>
  </div>
</div>

A common example of a hardcoded ARIA declaration you’ll encounter on the web is making an SVG icon inside a button decorative:

<button type="button>
  <svg aria-hidden="true">
    <!-- SVG code -->
  </svg>
  Save
</button>

Here, the string “Save” is what is required for someone to understand what the button will do when they activate it. The accompanying icon helps that understanding visually but is considered redundant and therefore decorative.

Declaring An Aria Role On Something That Already Uses That Role Implicitly Does Not Make It “Extra” Accessible

An implied role is all you need if you’re using semantic HTML. Explicitly declaring its role via ARIA does not confer any additional advantages.

<!-- 
  You don't need to declare role="button" here.
  Using the <button> element will make assistive 
  technology announce it as a button. The 
  role="button" declaration is redundant.
 -->
<button role="button">
  Save
</button>

You might occasionally run into these redundant declarations on HTML sectioning elements, such as <main role="main">, or <footer role="contentinfo">. This isn’t needed anymore, and you can just use the <main> or <footer> elements.

The reason for this is historic. These declarations were done for support reasons, in that it was a stop-gap technique for assistive technology that needed to be updated to support these new-at-the-time HTML elements.

Contemporary assistive technology does not need these redundant declarations. Think of it the same way that we don’t have to use vendor prefixes for the CSS border-radius property anymore.

Note: There is an exception to this guidance. There are circumstances where certain complex and complicated markup patterns don’t work as expected for assistive technology. In these cases, we want to hardcode the implicit role as explicit ARIA to ensure it works. This assistive technology support concern is covered in more detail later in this post.

You Don’t Need To Say What A Control Is; That Is What Roles Are For

Both implicit and explicit roles are announced by screen readers. You don’t need to include that part for things like the interactive element’s text string or an aria-label.

<!-- Don't do this -->
<button 
  aria-label="Save button"
  type="button">
  <!-- Icon SVG -->
</button>

<!-- Do this instead -->
<button 
  aria-label="Save"
  type="button">
  <!-- Icon SVG -->
</button>

Had we used the string value of “Save button” for our Save button, a screen reader would announce it along the lines of, “Save button, button.” That’s redundant and confusing.

ARIA Roles Have Very Specific Meanings

We sometimes refer to website and web app navigation colloquially as menus, especially if it’s an e-commerce-style mega menu.

In ARIA, menus mean something very specific. Don’t think of global or in-page navigation or the like. Think of menus in this context as what appears when you click the Edit menu button on your application’s menubar.

The edit menu option activated on Windows Notepad. It shows a list of menu options, with the option for “Go to” being in focus. Some options are disabled, as there is no content in the Notepad file, nor is there anything on the Windows Clipboard. The other menu options are Undo, Cut, Copy, Paste, Delete, Search with Bing, Find, Find Next, Find Previous, Replace, Select All, Time/Date, and Font. Screenshot.


Notepad, Windows 11. (Large preview)

Using a role improperly because its name seems like an appropriate fit at first glance creates confusion for people who do not have the context of the visual UI. Their expectations will be set with the announcement of the role, then subverted when it does not act the way it is supposed to.

Imagine if you click on a link, and instead of taking you to another webpage, it sends something completely unrelated to your printer instead. It’s sort of like that.

Declaring role="menu" is a common example of a misapplied role, but there are others. The best way to know what a role is used for? Go straight to the source and read up on it.

Certain Roles Are Forbidden From Having Accessible Names

These roles are caption, code, deletion, emphasis, generic, insertion, paragraph, presentation, strong, subscript, and superscript.

This means you can try and provide an accessible name for one of these elements — say via aria-label — but it won’t work because it’s disallowed by the rules of ARIA’s grammar.

<!-- This won't work-->
<strong aria-label="A 35% discount!">
  $39.95
</strong>

<!-- Neither will this -->
<code title="let JavaScript example">
  let submitButton = document.querySelector('button[type="submit"]');
</code>

For these examples, recall that the role is implicit, sourced from the declared HTML element.

Note here that sometimes a browser will make an attempt regardless and overwrite the author-specified string value. This overriding is a confusing act for all involved, which led to the rule being established in the first place.

You Can’t Make Up ARIA And Expect It To Work

I’ve witnessed some developers guess-adding CSS classes, such as .background-red or .text-white, to their markup and being rewarded if the design visually updates correctly.

The reason this works is that someone previously added those classes to the project. With ARIA, the people who add the content we can use are the Accessible Rich Internet Applications Working Group. This means each new version of ARIA has a predefined set of properties and values. Assistive technology is then updated to parse those attributes and values, although this isn’t always a guarantee.

Declaring ARIA, which isn’t part of that predefined set, means assistive technology won’t know what it is and consequently won’t announce it.

<!-- 
  There is no "selectpanel" role in ARIA.
  Because of this, this code will be announced 
  as a button and not as a select panel.
-->
<button 
  role="selectpanel"
  type="button">
  Choose resources
</button>

ARIA Fails Silently

This speaks to the previous section, where ARIA won’t understand words spoken to it that exist outside its limited vocabulary.

There are no console errors for malformed ARIA. There’s also no alert dialog, beeping sound, or flashing light for your operating system, browser, or assistive technology. This fact is yet another reason why it is so important to test with actual assistive technology.

You don’t have to be an expert here, either. There is a good chance your code needs updating if you set something to announce as a specific state and assistive technology in its default configuration does not announce that state.

ARIA Only Exposes The Presence Of Something To Assistive Technology

Applying ARIA to something does not automatically “unlock” capabilities. It only sends a hint to assistive technology about how the interactive content should behave.

For assistive technology like screen readers, that hint could be for how to announce something. For assistive technology like refreshable Braille displays, it could be for how it raises and lowers its pins. For example, declaring role="button" on a div element does not automatically make it clickable. You will still need to:

  • Target the div element in JavaScript,
  • Tie it to a click event,
  • Author the interactive logic that it performs when clicked, and then
  • Accommodate all the other expected behaviors.

This all makes me wonder why you can’t save yourself some work and use a button element in the first place, but that is a different story for a different day.

Additionally, adjusting an element’s role via ARIA does not modify the element’s native functionality. For example, you can declare role="image" on a div element. However, attempting to declare the alt or src attributes on the div won’t work. This is because alt and src are not supported attributes for div.

Two panels, one labeled “Will work” and the other labeled, “Won’t work.” The panel labeled “Will work” shows an image element with an alt and src attribute. The panel labeled “Won’t work” shows a div with a role of image, as well as alt and src attributes. Both src attributes link to a file called cucumber.jpg, and both alt attributes use a string value of “A small cucumber.”

[IMG]

(Large preview)

Declaring an ARIA Role On Something Will Override Its Semantics, But Not Its Behavior

This speaks to the previous section on ARIA only exposing something’s presence. Don’t forget that certain HTML elements have primary and secondary interactive capabilities built into them.

For example, an anchor element’s primary capability is navigating to whatever URL value is provided for its href attribute. Secondary capabilities for an anchor element include copying the URL value, opening it in a new tab or incognito window, and so on.

A link whose string value is “Link with a role set to button.” Above it is text that reads, “For demonstration purposes only. Please don’t do this.” The link has a cursor placed over it, with an active right-click menu. The menu shows multiple actions you can take on the link, including opening it in a new tab or window, copying and saving the link address, searching the web for the link’s string value, as well as options provided by user-installed browser extensions. These options are managing the link with the 1Password password manager and copying a link to the selected text. Cropped screenshot.

[IMG]

Chrome on macOS. Note the support for user-installed browser extensions. (Large preview)

These secondary capabilities are still preserved. However, it may not be apparent to someone that they can use them — or use them in the way that they’d expect — depending on what is announced.

The opposite is also true. When an element has no capabilities, having its role adjusted does not grant it any new abilities. Remember, ARIA only announces. This is why that div with a role of button assigned to it won’t do anything when clicked if no companion JavaScript logic is also present.

Two side-by-side graphics, each one consisting of three panels. The first panel on the left of the graphic shows the HTML code for a button element. The first panel for the right graphic shows HTML code for a div with a role of button. Both examples use a string value of “Favorite” and have a class of “button-fav” applied to them. The second panel for both left and right graphics shows an identical-looking button labeled “Favorite”, which has a golden-colored background. The third panel for the left graphic shows support for Enter and Space keypresses. The third panel for the right graphic shows no support for Enter and Space keypresses.

[IMG]

(Large preview)

You Will Need To Declare ARIA To Make Certain Interactions Accessible

A lot of the previous content may make it seem like ARIA is something you should avoid using altogether. This isn’t true. Know that this guidance is written to help steer you to situations where HTML does not offer the capability to describe an interaction out of the box. This space is where you want to use ARIA.

Knowing how to identify this area requires spending some time learning what HTML elements there are, as well as what they are and are not used for. I quite like HTML5 Doctor’s Element Index for upskilling on this.

Certain ARIA States Require Certain ARIA Roles To Be Present

This is analogous to how HTML has both global attributes and attributes that can only be used on a per-element basis. For example, aria-describedby can be used on any HTML element or role. However, aria-posinset can only be used with article, comment, listitem, menuitem, option, radio, row, and tab roles. Remember here that these roles can be provided by either HTML or ARIA.

Learning what states require which roles can be achieved by reading the official reference. Check for the “Used in Roles” portion of each entry’s characteristics:

 A characteristics table for aria setsize. The table’s two columns are labeled “Characteristic” and “Value.” The second table row is highlighted, demonstrating where you look for what role supports what state. The First row’s first cell has the text, “Used in roles.” The first row’s second cell has the text, “article, listitem, menuitem, option, radio, row, tab.” The second row’s first cell has the text, “Inherits into Roles.” The second row’s second cell has the text, “menuitemcheckbox, menuitemradio, treeitem.” The third row’s first cell has the text “Value.” Cropped screenshot.

[IMG]

Characteristics for aria-setsize. (Large preview)

Automated code scanners — like axe, WAVE, ARC Toolkit, Pa11y, equal-access, and so on — can catch this sort of thing if they are written in error. I’m a big fan of implementing these sorts of checks as part of a continuous integration strategy, as it makes it a code quality concern shared across the whole team.

ARIA Is More Than Web Browsers

Speaking of technology that listens, it is helpful to know that the ARIA you declare instructs the browser to speak to the operating system the browser is installed on. Assistive technology then listens to what the operating system reports. It then communicates that to the person using the computer, tablet, smartphone, and so on.

A flowchart with four steps. The first step is a webpage with a code icon floating above it. The second step is a computer, with an icon of an indented list floating above it. The third step is the symbol for accessibility, a Vitruvian man in a circle. Above this icon is a speech bubble. The fourth and final step is a person, with an icon of a lit lightbulb floating above it.

[IMG]

(Large preview)

A person can then instruct assistive technology to request the operating system to take action on the web content displayed in the browser.

A flowchart with four steps. The first step is a person with an icon of a finger pressing a button floating above it. The second step is the symbol for accessibility, a Vitruvian man in a circle. Above this icon is a speech bubble. The third step is a computer, with an icon of a handshake floating above it. The fourth and final step is an updated webpage, with a clicking mouse cursor icon floating above it.

[IMG]

(Large preview)

This interaction model is by design. It is done to make interaction from assistive technology indistinguishable from interaction performed without assistive technology.

There are a few reasons for this approach. The most important one is it helps preserve the privacy and autonomy of the people who rely on assistive technologies.

Just Because It Exists In The ARIA Spec Does Not Mean Assistive Technology Will Support It

This support issue was touched on earlier and is a difficult fact to come to terms with.

Contemporary developers enjoy the hard-fought, hard-won benefits of the web standards movement. This means you can declare HTML and know that it will work with every major browser out there. ARIA does not have this. Each assistive technology vendor has its own interpretation of the ARIA specification. Oftentimes, these interpretations are convergent. Sometimes, they’re not.

Assistive technology vendors also have support roadmaps for their products. Some assistive technology vendors:

  • Will eventually add support,
  • May never, and some
  • Might do so in a way that contradicts how other vendors choose to implement things.

There is also the operating system layer to contend with, which I’ll cover in more detail in a little bit. Here, the mechanisms used to communicate with assistive technology are dusty, oft-neglected areas of software development.

With these layers comes a scenario where the assistive technology can support the ARIA declared, but the operating system itself cannot communicate the ARIA’s presence, or vice-versa. The reasons for this are varied but ultimately boil down to a historic lack of support, prioritization, and resources. However, I am optimistic that this is changing.

Additionally, there is no equivalent to Caniuse, Baseline, or Web Platform Status for assistive technology. The closest analog we have to support checking resources is a11ysupport.io, but know that it is the painstaking work of a single individual. Its content may not be up-to-date, as the work is both Herculean in its scale and Sisyphean in its scope. Because of this, I must re-stress the importance of manually testing with assistive technology to determine if the ARIA you use works as intended.

How To Determine ARIA Support

There are three main layers to determine if something is supported:

  • Operating system and version.
  • Assistive technology and version,
  • Browser and browser version.

1. Operating System And Version

Each operating system (e.g., Windows, macOS, Linux) has its own way of communicating what content is present to assistive technology. Each piece of assistive technology has to accommodate how to parse that communication.

Some assistive technology is incompatible with certain operating systems. An example of this is not being able to use VoiceOver with Windows, or JAWS with macOS. Furthermore, each version of each operating system has slight variations in what is reported and how. Sometimes, the operating system needs to be updated to “teach” it the updated AIRA vocabulary. Also, do not forget that things like bugs and regressions can occur.

2. Assistive Technology And Version

There is no “one true way” to make assistive technology. Each one is built to address different access needs and wants and is done so in an opinionated way — think how different web browsers have different features and UI.

Each piece of assistive technology that consumes web content has its own way of communicating this information, and this is by design. It works with what the operating system reports, filtered through things like heuristics and preferences.

A three by three grid of nine buttons, with a title of “Select your order.” Each button has a food-related emoji, with a tooltip showing the button’s accessible name. The buttons are a hamburger with the title “100% Angus Beef Burger”, french fries with the title “Special Smile Fries”, a pizza slice with the title “Pepperoni Pizza”, a hot dog with the title “Hot Dog With Mustard”, a sandwich with a title of “Ham Sando”, a taco with the title of “Tuesday Taco”, a plate of spaghetti with the title of “Pasgetti”, a waffle with the title of “Waffles Sans Chicken”, and some popcorn with the title of “Poppin’ Corn”.

[IMG]

The “Show names” command in macOS Voice Control, which displays the accessible names of these icon buttons. The accessible name has been supplied by aria-label. (Large preview)

Like operating systems, assistive technology also has different versions with what each version is capable of supporting. They can also be susceptible to bugs and regressions.

Another two factors worth pointing out here are upgrade hesitancy and lack of financial resources. Some people who rely on assistive technology are hesitant to upgrade it. This is based on a very understandable fear of breaking an important mechanism they use to interact with the world. This, in turn, translates to scenarios like holding off on updates until absolutely necessary, as well as disabling auto-updating functionality altogether.

Lack of financial resources is sometimes referred to as the disability or crip tax. Employment rates tend to be lower for disabled populations, and with that comes less money to spend on acquiring new technology and updating it. This concern can and does apply to operating systems, browsers, and assistive technology.

3. Browser And Browser Version

Some assistive technology works better with one browser compared to another. This is due to the underlying mechanics of how the browser reports its content to assistive technology. Using Firefox with NVDA is an example of this.

Additionally, the support for this reporting sometimes only gets added for newer versions. Unfortunately, it also means support can sometimes accidentally regress, and people don’t notice before releasing the browser update — again, this is due to a historic lack of resources and prioritization.

The Less Commonly-Used The ARIA You Declare, The Greater The Chance You’ll Need To Test It

Common ARIA declarations you’ll come across include, but are not limited to:

  • aria-label,
  • aria-labelledby,
  • aria-describedby,
  • aria-hidden,
  • aria-live.

These are more common because they’re more supported. They are more supported because many of these declarations have been around for a while. Recall the previous section that discussed actual assistive technology support compared to what the ARIA specification supplies.

Newer, more esoteric ARIA, or historically deprioritized declarations, may not have that support yet or may never. An example of how complicated this can get is aria-controls.

aria-controls is a part of ARIA that has been around for a while. JAWS had support for aria-controls, but then removed it after user feedback. Meanwhile, every other screen reader I’m aware of never bothered to add support.

What does that mean for us? Determining support, or lack thereof, is best accomplished by manual testing with assistive technology.

The More ARIA You Add To Something, The Greater The Chance Something Will Behave Unexpectedly

This fact takes into consideration the complexities in preferences, different levels of support, bugs, regressions, and other concerns that come with ARIA’s usage.

Philosophically, it’s a lot like adding more interactive complexity to your website or web app via JavaScript. The larger the surface area your code covers, the bigger the chance something unintended happens.

Consider the amount of ARIA added to a component or discrete part of your experience. The more of it there is declared nested into the Document Object Model (DOM), the more it interacts with parent ARIA declarations. This is because assistive technology reads what the DOM exposes to help determine intent.

A lot of contemporary development efforts are isolated, feature-based work that focuses on one small portion of the overall experience. Because of this, they may not take this holistic nesting situation into account. This is another reason why — you guessed it — manual testing is so important.

Anecdotally, WebAIM’s annual Millions report — an accessibility evaluation of the top 1,000,000 websites — touches on this phenomenon:

Increased ARIA usage on pages was associated with higher detected errors. The more ARIA attributes that were present, the more detected accessibility errors could be expected. This does not necessarily mean that ARIA introduced these errors (these pages are more complex), but pages typically had significantly more errors when ARIA was present.

Assistive Technology May Support Your Invalid ARIA Declaration

There is a chance that ARIA, which is authored inaccurately, will actually function as intended with assistive technology. While I do not recommend betting on this fact to do your work, I do think it is worth mentioning when it comes to things like debugging.

This is due to the wide range of familiarity there is with people who author ARIA.

Some of the more mature assistive technology vendors try to accommodate the lower end of this familiarity. This is done in order to better enable the people who use their software to actually get what they need.

There isn’t an exhaustive list of what accommodations each piece of assistive technology has. Think of it like the forgiving nature of a browser’s HTML parser, where the ultimate goal is to render content for humans.

aria-label Is Tricky

aria-label is one of the most common ARIA declarations you’ll run across. It’s also one of the most misused.

aria-label can’t be applied to non-interactive HTML elements, but oftentimes is. It can’t always be translated and is oftentimes overlooked for localization efforts. Additionally, it can make things frustrating to operate for people who use voice control software, where the visible label differs from what the underlying code uses.

Another problem is when it overrides an interactive element’s pre-existing accessible name. For example:

<!-- Don't do this -->
<a 
  aria-label="Our services"
  href="/services/">
  Services
</a>

This is a violation of WCAG Success Criterion 2.5.3: Label in Name, pure and simple. I have also seen it used as a way to provide a control hint. This is also a WCAG failure, in addition to being an antipattern:

<!-- Also don't do this -->
<a 
  aria-label="Click this link to learn more about our unique and valuable services"
  href="/services/">
  Services
</a>

These factors — along with other considerations — are why I consider aria-label a code smell.

aria-live Is Even Trickier

Live region announcements are powered by aria-live and are an important part of communicating updates to an experience to people who use screen readers.

Believe me when I say that getting aria-live to work properly is tricky, even under the best of scenarios. I won’t belabor the specifics here. Instead, I’ll point you to “Why are my live regions not working?”, a fantastic and comprehensive article published by TetraLogical.

The ARIA Authoring Practices Guide Can Lead You Astray

Also referred to as the APG, the ARIA Authoring Practices Guide should be treated with a decent amount of caution.

A screenshot of the ARIA Authoring Practices Guide homepage, with a yellow caution tape placed across it.

[IMG]

(Large preview)

The Downsides

The guide was originally authored to help demonstrate ARIA’s capabilities. As a result, its code examples near-exclusively, overwhelmingly, and disproportionately favor ARIA.

Unfortunately, the APG’s latest redesign also makes it far more approachable-looking than its surrounding W3C documentation. This is coupled with demonstrating UI patterns in a way that signals it’s a self-serve resource whose code can be used out of the box.

These factors create a scenario where people assume everything can be used as presented. This is not true.

Recall that just because ARIA is listed in the spec does not necessarily guarantee it is supported. Adrian Roselli writes about this in detail in his post, “No, APG’s Support Charts Are Not ‘Can I Use’ for ARIA”.

Also, remember the first rule of ARIA and know that an ARIA-first approach is counter to the specification’s core philosophy of use.

In my experience, this has led to developers assuming they can copy-paste code examples or reference how it’s structured in their own efforts, and everything will just work. This leads to mass frustration:

  • Digital accessibility practitioners have to explain that “doing the right thing” isn’t going to work as intended.
  • Developers then have to revisit their work to update it.
  • Most importantly, people who rely on assistive technology risk not being able to use something.

This is to say nothing about things like timelines and resourcing, working relationships, reputation, and brand perception.

The Upside

The APG’s main strength is highlighting what keyboard keypresses people will expect to work on each pattern.

Consider the listbox pattern. It details keypresses you may expect (arrow keys, Space, and Enter), as well as less-common ones (typeahead selection and making multiple selections). Here, we need to remember that ARIA is based on the Windows XP era. The keyboard-based interaction the APG suggests is built from the muscle memory established from the UI patterns used on this operating system.

While your tree view component may look visually different from the one on your operating system, people will expect it to be keyboard operable in the same way. Honoring this expectation will go a long way to ensuring your experiences are not only accessible but also intuitive and efficient to use.

Another strength of the APG is giving standardized, centralized names to UI patterns. Is it a dropdown? A listbox? A combobox? A select menu? Something else?

When it comes to digital accessibility, these terms all have specific meanings, as well as expectations that come with them. Having a common vocabulary when discussing how an experience should work goes a long way to ensuring everyone will be on the same page when it comes time to make and maintain things.

macOS VoiceOver Can Also Lead You Astray

VoiceOver on macOS has been experiencing a lot of problems over the last few years. If I could wager a guess as to why this is, as an outsider, it is that Apple’s priorities are focused elsewhere.

The bulk of web development efforts are conducted on macOS. This means that well-intentioned developers will reach for VoiceOver, as it comes bundled with macOS and is therefore more convenient. However, macOS VoiceOver usage has a drastic minority share for desktops and laptops. It is under 10% of usage, with Windows-based JAWS and NVDA occupying a combined 78.2% majority share:

A pie chart. The legend of the pie chart reads, “JAWS, 40.5%”, “NVDA, 37.7%”, “VoiceOver, 9.7%”, “SuperNova, 3.7%”, “ZoomText, 207%”, “Orca, 2.4%”, “Narrator, 0.7%”, and “Other, 2.7%.” Cropped screenshot.

[IMG]

Image source: WebAIM Screen Reader User Survey #10. (Large preview)

The Problem

The sad, sorry truth of the matter is that macOS VoiceOver, in its current state, has a lot of problems. It should only be used to confirm that it can operate the experience the way Windows-based screen readers can.

This means testing on Windows with NVDA or JAWS will create an experience that is far more accurate to what most people who use screen readers on a laptop or desktop will experience.

Dealing With The Problem

Because of this situation, I heavily encourage a workflow that involves:

  1. Creating an experience’s underlying markup,
  2. Testing it with NVDA or JAWS to set up baseline expectations,
  3. Testing it with macOS VoiceOver to identify what doesn’t work as expected.

Most of the time, I find myself having to declare redundant ARIA on the semantic HTML I write in order to address missed expected announcements for macOS VoiceOver.

macOS VoiceOver testing is still important to do, as it is not the fault of the person who uses macOS VoiceOver to get what they need, and we should ensure they can still have access.

You can use apps like VirtualBox and Windows evaluation Virtual Machines to use Windows in your macOS development environment. Services like AssistivLabs also make on-demand, preconfigured testing easy.

What About iOS VoiceOver?

Despite sharing the same name, VoiceOver on iOS is a completely different animal. As software, it is separate from its desktop equivalent and also enjoys a whopping 70.6% usage share.

With this knowledge, know that it’s also important to test the ARIA you write on mobile to make sure it works as intended.

You Can Style ARIA

ARIA attributes can be targeted via CSS the way other HTML attributes can. Consider this HTML markup for the main navigation portion of a small e-commerce site:

<nav aria-label="Main">
  <ul>
    <li>
      <a href="/home/">Home</a>
      <a href="/products/">Products</a>
      <a aria-current="true" href="/about-us/">About Us</a>
      <a href="/contact/">Contact</a>
    </li>
  </ul>
</nav>

The presence of aria-current="true" on the “About Us” link will tell assistive technology to announce that it is the current part of the site someone is on if they are navigating through the main site navigation.

We can also tie that indicator of being the current part of the site into something that is shown visually. Here’s how you can target the attribute in CSS:

nav[aria-label="Main"] [aria-current="true"] {
  border-bottom: 2px solid #ffffff;
}

This is an incredibly powerful way to tie application state to user-facing state. Combine it with modern CSS like :has() and view transitions and you have the ability to create robust, sophisticated UI with less reliance on JavaScript.

You Can Also Use ARIA When Writing UI Tests

Tests are great. They help guarantee that the code you work on will continue to do what you intended it to do.

A lot of web UI-based testing will use the presence of classes (e.g., .is-expanded) or data attributes (ex, data-expanded) to verify a UI’s existence, position and states. These types of selectors also have a far greater likelihood to be changed as time goes on when compared to semantic code and ARIA declarations.

This is something my coworker Cam McHenry touches on in his great post, “How I write accessible Playwright tests”. Consider this piece of Playwright code, which checks for the presence of a button that toggles open an edit menu:

// Selects an element with a role of `button` 
// that has an accessible name of "Edit"
const editMenuButton = await page.getByRole('button', { name: "Edit" });

// Requires the edit button to have a property 
// of `aria-haspopup` with a value of `true`
expect(editMenuButton).toHaveAttribute('aria-haspopup', 'true');

The test selects UI based on outcome rather than appearance. That’s a far more reliable way to target things in the long-term.

This all helps to create a virtuous feedback cycle. It enshrines semantic HTML and ARIA’s presence in your front-end UI code, which helps to guarantee accessible experiences don’t regress. Combining this with styling, you have a powerful, self-contained system for building robust, accessible experiences.

ARIA Is Ultimately About Caring About People

Web accessibility can be about enabling important things like scheduling medical appointments. It is also about fun things like chatting with your friends. It’s also used for every web experience that lives in between.

Using semantic HTML — supplemented with a judicious application of ARIA — helps you enable these experiences. To sum things up, ARIA:

  • Has been around for a long time, and its spirit reflects the era in which it was first created;
  • Has a governing taxonomy, vocabulary, and rules for use and is declared in the same way HTML attributes are;
  • Is mostly used for dynamically updating things, controlled via JavaScript;
  • Has highly specific use cases in mind for each of its roles;
  • Fails silently if mis-authored;
  • Only exposes the presence of something to assistive technology and does not confer interactivity;
  • Requires input from the web browser, but also the operating system, in order for assistive technology to use it;
  • Has a range of actual support, complicated by the more of it you use;
  • Has some things to watch out for, namely aria-label, the ARIA Authoring Practices Guide, and macOS VoiceOver support;
  • Can also be used for things like visual styling and writing resilient tests;
  • Is best evaluated by using actual assistive technology.

Viewed one way, ARIA is arcane, full of misconceptions, and fraught with potential missteps. Viewed another, ARIA is a beautiful and elegant way to programmatically communicate the interactivity and state of a user interface.

I choose the second view. At the end of the day, using ARIA helps to ensure that disabled people can use a web experience the same way everyone else can.

Thank you to Adrian Roselli and Jan Maarten for their feedback.

Further Reading