Showing posts with label documentation. Show all posts
Showing posts with label documentation. Show all posts

No posts for a wee while

Mike's Notes

I was on holiday for the last few weeks and am back now. There will be no blog posts, newsletters or meetings until Pipi Core is back up and running.

Update 27/05/2026

Lots of surprises. Making rapid progress. The peace and quiet are bliss.

Update 31/05/2026

The problem and solution are how things are named. Pipi auto-generates thousands of code names using multiple pattern languages, and all the naming conventions require many minor fixes for several unexpected reasons after migrating from a developer laptop to a production server environment. Everything else is absolutely fine.

Other naming problems are also being solved now, including:

  • The rapid development of Boxlang by Ortus has brought forward another challenge. Pipi 10 will be migrated to run on top of Boxlang in 2027 to support multiple languages, including C++, CFML, COBOL, Go, Java, JavaScript, PHP, Python, Rust, etc.
  • Future integration with cloud-based LLMs.
  • Future integrations with Office365, Google Workspace, Zoho, LibreOffice, etc.

The common solution is to create standardised naming systems that are simple, stable, robust, schema-based, versioned, self-documenting, and extensible to meet unanticipated future needs.

This is done by replacing code-based naming rules with database-driven ones that can be easily edited in the future via an admin UI.

90% of these names are internal, hidden in the closed core, and how they work and what they are will not be discussed here. The rest will be publicly and fully documented as part of the open-source workspaces for developers to work with.

Update 02/06/2026

I'm changing the disclosure boundary between the Pipi closed-core and open-source workspaces. Previously, "disclose everything unless there is a security reason not to". This is now changed to "disclose on the basis of need to know".

Closed-core accounts for 90% and open-source workspaces for 10% of lines of code, databases, etc.

This will reduce the documentation burden, given Pipi's vast scale. So, the open-source workspaces will be fully shared and documented on GitHub, etc, without restriction. This includes;

  • Standards schema
  • Ontologies
  • Parameters
  • Laws of physics
  • HTML + CSS
  • Algorithms
  • Module DDD models
  • Workflow diagrams
  • Documentation
  • API schema
  • UI code
  • etc

This also means some existing technical documentation about the closed-core will become hidden and only available internally.

Update 07/06/2026

Pipi Core is the IDE used to edit Pipi Core (AKA: which came first, the chicken or the egg?). Temporary UIs have been created and are being used across multiple engines to edit the names in use. This is much faster than directly editing data, which had to be done initially. The next step will be turning auto-generation back on. Once that's done, temporary UIs will be used to build permanent UIs. More automation will then be enabled via the UIs, and so on, as Pipi Core builds itself with a human in the loop.

Update 08/06/2026

The list of code cases available to use now for auto-generated naming, I/O translation, etc with examples, includes;

  • camelCase: userProfilePicture
  • kebab-case: user-profile-picture
  • PascalCase: UserProfilePicture
  • snake_case: user_profile_picture
  • SCREAMING_SNAKE_CASE: USER_PROFILE_PICTURE
  • Train-Case: User-Profile-Picture
  • flatcase: userprofilepicture
  • UPPER-CASE-KEBAB-CASE: USER-PROFILE-PICTURE
  • Sentence case: User profile picture
  • Title Case: User Profile Picture
  • middot·case: user·profile·picture
  • dot.case: user.profile.picture
  • UPPER CASE: USER PROFILE PICTURE
  • lowercase: user profile picture

Update 12/06/20026

Checking that these changes to variable names and internal messaging do not clash with the Gรถdel Machine.

Update 17/06/2026

The DevOps Engine (dvp) has unexpectedly proven to be critical to solving this puzzle. Mostly fixed last night. Watching the rather excellent live Google talk, Beyond the GPU: Maximising goodput with self-healing AI infrastructure, this morning has given me valuable insights into how to fix the remaining issues by reviewing Google HPC YAML files. ๐Ÿ˜Ž๐Ÿ˜Ž Sometimes insights come from the strangest places.

Update 01/07/2026

The main work now is rapidly configuring Pipi for production and full autonomous automation. Using Google Search AI Mode (Gemini) and then Grammarly Pro makes the work easier and 100x faster.

  • I have decided to have Pipi re-render the many Ajabbi draft public websites with the new and missing developer information. (20K pages)
  • The website's .robot.txt file will then be unlocked to enable search engines.
  • The HTML will be updated to make it easier for AI to read.
  • This blog will be imported into Pipi, cleaned up, re-exported from Pipi, and published to Blogger via the API.
  • The new posts created in Pipi will return to A Sandy Beach to discuss something already built rather than being built.

Update 02/07/2026

The DevOps and IaC engines are getting rapid data model overhauls. The IaC engine is a great test for the variable names. I'm building a capability into Pipi to autonomously and automatically run OpenTofu and Ansible, initially targeting the Pipi Data Centre, then GCP and AWS for deployments. It's going very well and making rapid progress.

Update 05/07/2026

Pipi will initially run the open-source enterprise applications on Google Cloud Run and Google Cloud Storage (GCS). The code is complete and will be very low-cost to run, giving Ajabbi, a bootstrapping-purpose startup, a very long runway.

Update 18/07/2026

The job has now shifted to configuring, networking and deploying many physical servers. Installing software, including Pipi, labelling cables and rack gear, throwing out junk, tidying, etc., leaving nothing to chance. Shipping delays are holding up part deliveries.

Update 28/07/2026

Most of the equipment has arrived, and the small data centre setup is coming together. More deliveries later this week. It's already running a lot better and is much more productive.

Update 31/07/2026

Work on Pipi has reached a tipping point or system phase change as Pipi takes over tasks using autonomous automation. Pipi now has deadlines, not me. Soon it will set the deadlines. It's now a downhill run; daily posts from me resume tomorrow, and much more will come.

In hindsight. This whole project has been systematic trial and error, spending 10 years learning how to crack a hard problem.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

31/07/2026

No posts for a wee while

By: Mike Peters
On a Sandy Beach: 15/05/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

I was on a no-coding holiday for the last few weeks to clear my mind, and it has been great. I am back on the job today.

Suspended

Until the closed-source Pipi Core is back up and running 100% on autopilot, 10x faster, the following are suspended.

  • New posts "On a Sandy Beach
  • All newsletters, including the weekly Friday Report and the monthly Ajabbi Research Newsletter.
  • The fortnightly online Open R&D meeting.

Rapid refocus

  • A new developer area with five coding screens, designed to be more productive for hypervisual learners.
  • A better library has been set up for my A4 drawings in ring binders, the many reference books I use, and more bookshelves are on the way.
  • The server rack has been moved to a better location.
  • The light levels have been adjusted.
  • A big office tidy is almost done. An office-work-only desk has yet to be set up with a cat bed included.
  • A separate area with no screens for the happy cat, coffee, music, reading and drawing.

Less is more

Minimise screen time to be more productive at work. The new setup is also much less tiring.

Get the job done

The good thing is that, with a holiday and lots of drawing, I now have mental clarity about what needs fixing and how to fix it. Mainly, quite delicate changes here and there, organised into a list of steps. Now, I need to concentrate on one thing only: go as fast as possible, without meetings, post-deadlines, phone calls, or other distractions.

How

1. Use an AI workforce

Be the architect, and AI fills in the dots to make it happen.

Use Google Search AI mode (Gemini) to generate 99% of the code in one-page chunks (including references) to copy and paste, then manually change the variable names and SQL. Careful, test everything, resulting in 100x faster progress. Know how everything works and rapidly raise personal skill level.

2. Then build a cathedral

Make a wooden scale model of a cathedral for the builders. Google Search AI mode (Gemini) makes each brick, and Pipi Core assembles the bricks into floors, arches, walls, and vaults...

Speed is king

With the 100x coding productivity gains from Google Search AI mode (Gemini), plus the 10x10x10x speedup of Pipi Core currently underway over the next few months, what previously took a year will be done in hours and better.

Phase transitions

Once these initial migration issues from laptop to server are resolved, further transitions can be anticipated as the number of engines rapidly increases beyond 20. Increasing the number of engines slowly changes the whole system's behaviour from deterministic to probabilistic and adaptive.

Here is a partial list of transitions expected as the number of engines increases from 0 to 200. The actual numbers are a bit of a guess.

  • 20 engines enable Pipi 9 Core in a simple, deterministic structure.
  • 40 engines enable a workspace with a UI for administering Pipi Core.
  • 60 engines enable self-generation of user documentation.
  • 80 engines enable REPL and IAC (infrastructure-as-code).
  • 100 engines enable Workspaces for different user accounts.
  • Different Pipi 9 editions are made with the same engines, which recombine differently in response to the external environment.
  • And so on until...
  • 200 engines self-organise into a multi-layered complex fluid structure with probabilistic behaviour and emergent properties, as engines also act as agents.
  • 200+ engines enable Pipi 10 to interact with externally cloud-hosted LLMs, combining the very different strengths of both.

Creating 18 Engine descriptions

Mike's Notes

18 working engines are currently being imported into Pipi Core, configured, and tested.

Alex has sent me a DeepSeek chat that generated descriptions of those engines and how they worked together by analysing the existing web page "20 Engines". DeepSeek also produced a Mermaid Diagram from Markdown code based on those descriptions. DeepSeek was partially correct in some descriptions.

The Mermaid diagram was wrong, but what a great tool for Pipi to self-document with CL and accurate Markdown descriptions. It will be widely used in future as a plugin.

Feedback and suggestions are very welcome as always.

Update 23/05/2026

Java and CGI Engines added for interoperability by Pipi.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

23/05/2026

Creating 18 20 Engine descriptions

By: Mike Peters
On a Sandy Beach: 06/05/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

Most Pipi engines have historically been poorly described. This is an attempt to write better descriptions for the 18 20 engines currently being imported.  I had a go, then used Gemini to come up with better wording, then Mrs Grammarly did her bit. ๐Ÿ˜Ž

Much later, once the workspace UI are working, a one-page summary about each engine can be written for the pipiWiki. The Wiki Engine (wik) in the resources above has an example of such a summary.

Note: These are listed in the order they are being imported.

Descriptions

Each engine has;

  • Unique Name (Unique 3-letter code)
  • S: Short description suitable for tooltips under 100 characters.
  • D: Description under 255 characters.

System Engine (sys)

  • S: System identity and lifecycle controller.
  • D: Manages the vital signs and lifecycle of Pipi’s dynamic engine ecosystem, fostering complex emergent behaviours through seamless interaction.

Nest Engine (nst)

  • S: Host and environment interface bridge.
  • D: Serves as the foundational gateway between the host OS, JVM, the application server, CGI, and the internal Pipi environment.

JVM Engine (jvm)

  • S: JVM interoperability.
  • D: Provides Pipi with a way to work directly with an external JVM.

CGI Engine (cgi)

  • S: CGI interoperability.
  • D: Provides Pipi with a way to work directly with external CGI.

Namespace Engine (nsp)

  • S: Global identification and addressing.
  • D: Enforces a conflict-free global naming convention, ensuring every system element is uniquely addressable across the entire platform.

Render Engine (rnd)

  • S: Static file and asset rendering.
  • D: Processes and renders static resources, including HTML, CSS, source code, and databases.

Template Engine (tem)

  • S: Reusable pattern templates.
  • D: Manages reusable structural pattern templates used by the CMS to generate database-driven pages and components.

Variables Engine (var)

  • S: Centralised variables library.
  • D: Provides a centralised repository for managing variables used across templates, system configurations, and executable logic.

Log Engine (log)

  • S: Universal logging and telemetry controller.
  • D: Aggregates and configures logging parameters across all active engines to provide system-wide transparency and diagnostics.

Data Engine (dta)

  • S: Database lifecycle and CRUD operations.
  • D: Generates SQL to command the creation, evolution, and deletion of databases and their underlying data objects with full administrative control.

Configuration Engine (cnf)

  • S: Engine blueprint and manufacturing settings.
  • D: Supplies the precise DNA and configuration parameters required for the automated fabrication of individual Pipi engines.

Versioning Engine (ver)

  • S: Semantic versioning and update tracker.
  • D: Maintains system integrity by aggregating incremental updates from all engines into a unified semantic versioning timeline.

Code Engine (cde)

  • S: Internal code generation.
  • D: Facilitates automated code generation, including class libraries and logic synthesis directly within the Pipi platform.

Conductor Engine (cnd)

  • S: Internal process regulator.
  • D: Operates as the high-level orchestrator for major internal system processes and synchronisation.

Directory Engine (dir)

  • S: CMS file system and path manager.
  • D: Manages the logical and physical file system directories generated and utilised by the CMS.

Node Engine (nde)

  • S: Hierarchical template tree architect.
  • D: Maintains the addressable tree structure of templates to define content hierarchy within the CMS.

CMS Engine (cms)

  • S: Digital content management.
  • D: Powers the end-to-end lifecycle of digital content, from initial creation and editing to final publishing.

Core Engine (cor)

  • S: Primary system driver and logic hub.
  • D: The central engine that powers fundamental system behaviours and executes the primary logic that keeps Pipi running.

Factory Engine (fac)

  • S: Automated engine fabrication.
  • D: Assembles and deploys engines based on stored configuration files and real-time updates.

Page Engine (pge)

  • S: Semantic relationship and metadata mapper.
  • D: Maps external relationships for pages, managing keywords, references, and "See Also" semantic connections.

Using slide presentations to describe Pipi

Mike's Notes

Thoughts on how to give useful slide presentation talks about Pipi, record them, and make them available on YouTube as a way to explain how Pipi works.

Resources

References

  • Content Management Bible 2nd Ed., by Bob Boiko. Wiley. 2005.

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

09/02/2026

Using slide presentations to describe Pipi

By: Mike Peters
On a Sandy Beach: 09/02/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

I gave a slide talk last night at the regular Open Research Group online meeting about future blog posts being created by a human using a Workspace, transferred to the CMS Engine (cms), processed, and then automatically published to Google Blogger. Creating the slides made me realise the opportunity available to use this format to visually explain the many parts of Pipi simply.

I will give a slide presentation on the Workspace Engine (wsp) at the next meeting. I will also give one on the Workspaces for Screen to the local Film Industry group this month.

Backstory

When I was a young adult, I went around with a group of good people, many of whom have since become lifelong friends, who encouraged me to give some talks. The only problem was that their approach was to write the talk in advance and then read it aloud to the audience. They were all very good at it, and I was hopeless.

  • The first problem was that I found it impossible to write.
  • The second problem was reading out loud what was written. I tripped over the words.

Years later, I had this idea that just talking about something in front of me might work a lot better, a picture, a map, a physical gadget, for example. I have no problems talking about something I understand.

When I became National President of NZERN, I had to give many talks, and that was the method: show slides and just talk about the pictures or diagrams without using notes, unless there was a name or date to remember, often using a whiteboard to draw answers for people who asked questions in a meeting.

I ended up giving hundreds of talks at conferences and workshops across NZ. The longest was 2 1/2 hours, given to the South Island DOC IMU workshop about the NZERN GIS project using ESRI software, and it was highly technical. No notes, just 50 slides.

Using computers like a typewriter has been a tremendous help because of cut-and-paste, which is much easier than shuffling bits of paper. Besides, I use Arial 16pt, which is much easier to read than my handwriting.

Mrs Grammarly

Later, I learned to use assistive technology to help me write. Grammarly Pro rewrites every single sentence that has my name on it, including this post. Grammarly is set on formal British business English. I hope my personal secretary, Mrs Grammarly, is doing a good job.

Big Challenge

Pipi is largely undocumented because it was designed and built visually. There must be thousands of hand coloured drawings on A4 paper, some neatly filed in 50+ 3-hole A4 ring binders, and the rest in many cartons waiting to be filed. Pipi needs to be documented so others can use it. There is a steadily growing interest in Pipi worldwide.

Solutions

  1. Getting Pipi to self-document is well underway, using structured templates that render from hundreds of databases. A rough estimate is that 20,000 web pages of developer technical documentation will be required due to the scale and scope of this enterprise platform.
  2. Setting up a community forum where users can ask questions and provide answers will take the load off me.
  3. I also need to explain verbally the more complicated bits that I find too difficult to write about. Give a slide presentation and record it to share on YouTube.
  4. Use screen capture to record live demos of Pipi in use.
  5. Provide regular Office Hours that can be booked for video chats via Google Meet or Zoom. I'm doing that a lot, and it seems to work.
  6. Record video interviews with the people who wrote most of the articles that I have copied and republished on this engineering blog, On a Sandy Beach. They could be two-way and a chance to discuss some deep issues.
  7. Teaching someone something complex by making it simple is the best way to learn it. So, giving many talks will also help me understand more clearly.

Slide Presentations

Here is a possible list of some overview talks about just one engine as an example. Then there could be more detailed talks on the same subjects. There are hundreds of Agent Engines. Each talk could have about 10-20 slides.

CMS Engine

  • 101 Introduction
  • 102 Content Management System
  • 103 Publication
  • 104 Website
  • 105 Blog
  • 106 Wiki
  • 107 Docs
  • 108 Help
  • 109 Workspace

Next Steps

Once I get into the swing of it, it should get easy. I need to learn to speak more slowly, develop a visual style for the slides, establish a simple slide-naming convention, and address related details. Each slide set will need a webpage for downloading the PDF/PowerPoint/Google Slides, watching the YouTube Video, a printable PDF handout, and links to related information.

The recorded slides, talks, and demos could all be organised using the existing Diataxis framework and Learning Objects, which Pipi uses elsewhere.

Understanding the Hidden Markov Model

Mike's Notes

A Markov Chain is a collection of states with transition probabilities between them.

Markov is used extensively in Pipi 9. Here is some excellent background explanations from Vivek Vinushanth Christopher on Built In.

Vivek is a fantastic teacher with an excellent YouTube channel on Data Science.

While the core code of Pipi will be closed source, the weights, algorithms, possible states, ontologies used, and related materials will be made freely available. The front-end workspaces will be shared on GitHub, GitLab, and other platforms.

The planned user documentation will include references to popular books, articles, and easy-to-understand videos from people like Vivek.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

24/01/2026

Understanding the Hidden Markov Model

By: Vivek Vinushanth Christopher, Updated By: Brennan Whitfield
Built In: 14/08/2024

Vivek Vinushanth Christopher is a senior software engineer for WSO2 where he has worked since 2019. He has also served as a teaching assistant at University of Moratuwa. Christopher holds a bachelor’s degree in computer science and engineering. 

Hidden Markov models are probabilistic frameworks where the observed data are modeled as a series of outputs generated by one of several (hidden) internal states. Both Markov and hidden Markov models are engineered to handle data that can be represented as a sequence of observations over time.

Hidden Markov Model Definition

A hidden Markov model is a probabilistic framework used to predict the results of an event based on a series of observations with one or several hidden internal states.

While a Markov model or Markov chain concerns stochastic (random) process states that are visible to the observer, a hidden Markov model pertains to stochastic processes where states can be hidden or not directly visible to the observer.

What Is a Hidden Markov Model (HMM)?

A hidden Markov model (HMM) is utilized when we can’t observe the states of a stochastic process themselves, but only the result of some probability function (observation) of the states. HMM is a statistical Markov model in which the system being modeled is assumed to be a Markov process with unobserved (hidden) states.

Hidden Markov models can be used to identify underlying patterns or structures in sequential data. This makes it applicable for research and tasks in machine learning (including natural language processing and speech recognition), bioinformatics and gene analysis as well as time-series forecasting. 

For example, in speech recognition tasks, a hidden Markov model algorithm may be implemented to measure the probability of a certain word or lack of words occurring in a given audio recording. In this case, the occurrence of specific words or silence in the recording can represent states, and volume of speech throughout the recording can represent observations. By knowing observations (volume), this information can be used by the algorithm to determine the likelihood of hidden states — or words and lack of words in this example — and predict the most probable word being spoken.

Mathematically, here is how Markov models and hidden Markov models differ:

  • Markov model: Series of (hidden) states z={z_1,z_2………….} drawn from state alphabet S ={s_1,s_2,…….๐‘ _|๐‘†|} where z_i belongs to S.
  • Hidden Markov model: Series of observed output x = {x_1,x_2,………} drawn from an output alphabet V= {๐‘ฃ1, ๐‘ฃ2, . . , ๐‘ฃ_|๐‘ฃ|} where x_i belongs to V.

Hidden Markov Model Assumptions

A hidden Markov model is built on several assumptions, including:

1. Output Independence Assumption

Output observation is conditionally independent of all other hidden states and all other observations when given the current hidden state.

Eq. 5: Output independence assumption. | Image: Vivek Vinushanth Christopher

2. Emission Probability Matrix

Probability of hidden state generating output v_i given that state at the corresponding time was s_j.

Markov Model Assumptions

Markov models are developed based on two assumptions:

1. Limited Horizon Assumption

Probability of being in a state at a time t depend only on the state at the time (t-1).

Equation 1: Limited horizon assumption. | Image: Vivek Vinushanth Christopher

That means state at time t represents enough summary of the past to reasonably predict the future. This assumption is an order-1 Markov process. An order-k Markov process assumes conditional independence of state z_t from the states that are k + 1-time steps before it.

2. Stationary Process Assumption

Conditional (probability) distribution over the next state, given the current state, doesn’t change over time.

Equation 2: Stationary process assumption. | Image: Vivek Vinushanth Christopher

That means states keep on changing over time but the underlying process is stationary.

Notation Convention

  • There is an initial state and an initial observation z_0 = s_0
  • s_0: Initial probability distribution over states at time 0.
  • Initial state probability: (ฯ€)
  • At t=1, probability of seeing first real state z_1 is p(z_1/z_0).
  • Since z0 = s0:

Notation convention. | Image: Vivek Vinushanth Christopher

State Transition Matrix

State transition matrix. | Image: Vivek Vinushanth Christopher

๐€๐ข,๐ฃ: probability of transitioning from state i to state j at any time t.

The following chart is a state transition matrix of four states, including the initial state:

Fig. 2: State transition matrix chart. | Image: Vivek Vinushanth Christopher

2 Questions Answered in a Markov Model

  1. What is the probability of particular sequences of state z?
  2. How do we estimate the parameter of state transition matrix A to maximize the likelihood of the observed sequence?

Probability of Particular Sequences in a Markov Model

Eq.4: Finding probability of particular sequence. | Image: Vivek Vinushanth Christopher

Consider the state transition matrix above. Let’s determine the probability of sequence:

{z1 = s_hot , z2 = s_cold , z3 = s_rain , z4 = s_rain , z5 = s_cold}

P(z) = P(s_hot|s_0 ) P(s_cold|s_hot) P(s_rain|s_cold) P(s_rain|s_rain) P(s_cold|s_rain)

= 0.33 x 0.1 x 0.2 x 0.7 x 0.2 = 0.000924

Hidden Markov Model Example

Consider the example given below, which elaborates how a person feels in different climates.

Fig.3: Markov model as finite state machine. | Image: Vivek Vinushanth Christopher 

  • Set of states: (S) = {Happy, Grumpy}
  • Set of hidden states: (Q) = {Sunny , Rainy}
  • State series over time: = z∈ S_T
  • Observed states for four day: = {z1=Happy, z2= Grumpy, z3=Grumpy, z4=Happy}

The feeling that you understand from a person emoting is called the observations, since you observe them. The weather that influences the feeling of a person is called the hidden state, since you can’t observe it.

Emission Probabilities

In the above example, feelings (“Happy” or “Grumpy”) can be only observed. A person can observe that a person has an 80 percent chance to be “happy” given that the climate at the particular point of observation is sunny. Similarly there’s a 60 percent chance of a person being “grumpy” given that the climate is rainy. The 80 percent and 60 percent are emission probabilities since they deal with observations.

Transition Probabilities

When we consider the climates (hidden states) that influence the observations, there are correlations between consecutive days being sunny or alternate days being rainy. There is an 80 percent chance for the Sunny climate to be in successive days, whereas there’s a 60 percent chance for it to be rainy on consecutive days. The probabilities that explain the transition to/from hidden states are transition probabilities.

How Does a Hidden Markov Model Work?

A hidden Markov model answers three primary questions:

  1. What is the probability of an observed sequence?
  2. What is the most likely series of states to generate an observed sequence?
  3. How can we learn the values for the HMMs parameters A and B given some data?

Probability of Observed Sequence

We have to add up the likelihood of the data x given every possible series of hidden states. This will lead to a complexity of O(|S|)^T. Hence, two alternate procedures were introduced to find the probability of an observed sequence.

Forward Procedure

Calculate the total probability of all the observations (from t_1) up to time t:

๐›ผ_๐‘– (๐‘ก) = ๐‘ƒ(๐‘ฅ_1 , ๐‘ฅ_2 , … , ๐‘ฅ_๐‘ก, ๐‘ง_๐‘ก = ๐‘ _๐‘–; ๐ด, ๐ต)

Backward Procedure

Similarly calculate total probability of all the observations from final time (T) to t:

๐›ฝ_i (t) = P(x_T , x_T-1 , …, x_t+1 , z_t= s_i ; A, B)

A tutorial on how the hidden markov model works. | Video: ritvikmath

Hidden Markov Model Using Forward Procedure

Below is an example of a hidden Markov model using forward procedure.

  • S = {hot,cold}
  • v = {v1=1 ice cream ,v2=2 ice cream, v3=3 ice cream}, where V is the Number of ice creams consumed in a day.
  • Example Sequence: = {x1=v2,x2=v3,x3=v1,x4=v2}

Fig.4: Given data as matrices. | Image: Vivek Vinushanth Christopher

Generated finite state machines for HMM. | Image: Vivek Vinushanth Christopher

We first need to calculate the prior probabilities, that is, the probability of being hot or cold previous to any actual observation. This can be obtained from S_0 or ฯ€. From Fig.4, S_0 is provided as 0.6 and 0.4, which are the prior probabilities. Then based on Markov and HMM assumptions, we follow the steps in the figures below to calculate the probability of a given sequence.

1. First Observed Output x1=v2

Fig. 6: Step 1 of the HMM. | Image: Vivek Vinushanth Christopher

2. Observed Output x2=v3

Fig. 7: Step 2 of HMM illustrated. | Image: Vivek Vinushanth Christopher

3. Observed Output x3 and x4

Similarly for x3=v1 and x4=v2, we have to simply multiply the paths that lead to v1 and v2.

Fig. 8: Step 3 and 4 of HMM. | Image: Vivek Vinushanth Christopher

4. Maximum Likelihood Assignment

For a given observed sequence of outputs ๐‘ฅ ๐œ– ๐‘‰_๐‘‡, we intend to find the most likely series of states ๐‘ง ๐œ– ๐‘†_๐‘‡. We can understand this with an example found below.

Fig.9: Data for example two. | Image: Vivek Vinushanth Christopher

Fig.10: Markov model as a finite state machine from Fig.9. data. | Image: Vivek Vinushanth Christopher

The Viterbi algorithm is a dynamic programming algorithm similar to the forward procedure which is often used to find maximum likelihood. Instead of tracking the total probability of generating the observations, it tracks the maximum probability and the corresponding state sequence.

Consider the sequence of emotions: H,H,G,G,G,H for six consecutive days. Using the Viterbi algorithm we will find out the more likelihood of the series.

Fig.11: The Viterbi algorithm requires to choose the best path. | Image: Vivek Vinushanth Christopher

There will be several paths that will lead to sunny on Saturday, and many paths that lead to rainy on Saturday. Here, we intend to identify the best path to sunny or rainy Saturday and multiply with the transition emission probability of “Happy,” since Saturday makes the person feel “Happy.”

Let’s consider a sunny Saturday. The previous day, Friday, can be sunny or rainy. Then we need to know the best path up to Friday, and then multiply with emission probabilities that lead to a grumpy feeling. Iteratively, we need to figure out the best path at each day ending up in more likelihood of the series of days.

Fig.12: Step 1. | Image: Vivek Vinushanth Christopher

Fig. 13: Step 2. | Image: Vivek Vinushanth Christopher

Fig.14. Iterate the algorithm to choose the best path. | Image: Vivek Vinushanth Christopher

 The algorithm leaves you with maximum likelihood values and we now can produce the sequence with a maximum likelihood for a given output sequence.

5. Learn the Values for the HMMs Parameters A and B

Learning in HMMs involves estimating the state transition probabilities A and the output emission probabilities B that make an observed sequence most likely. Expectation-Maximization algorithms are used for this purpose. A commonly-used algorithm known as the Baum-Welch algorithm falls under this category and uses the forward algorithm.

Frequently Asked Questions

What is a hidden Markov model?

A hidden Markov model is a statistical model in which the system being modeled is assumed to be a Markov process with unobserved (hidden) states. It’s used when you can’t observe the states themselves but only the result of a probability function of the states

What’s the difference between a hidden Markov model vs. a Markov model?

    • A hidden Markov model is a probabilistic model used when the results are the product of one of several hidden internal states. 
    • A Markov model is a probabilistic model used to predict a sequence of events when the internal states can be observed.   

What are the basic problems of hidden Markov model?

The three basic problems of a hidden Markov model (HMM) include:

    1. Evaluation or Likelihood Problem: Given an HMM ฮป = (A, B) and an observation sequence of O = O_1 O_2, … O_T, determine the likelihood P(O'ฮป).
    2. Decoding Problem: Given an HMM ฮป = (A, B) and an observation sequence of O = O_1 O_2, … O_T, determine the optimal hidden state sequence of Q (set of finite states).
    3. Learning Problem: Given an observation sequence of O = O_1 O_2, … O_T, learn and adjust the HMM parameters A and B to maximize the probability P(O'ฮป).

Ontology Documentation Tools and Workflows

Mike's Notes

Marcelo Xavier started a fascinating discussion on the Ontolog Forum about creating online documentation for an Ontology. My friend Alex Shkotin used Gemini to identify a possible solution. Nico Matentzoglu also made valuable points, including suggestions around using DiataxisOBOOK (Open Biological and Biomedical Ontologies Organised Knowledge) is a fantastic resource and a great example of how to organise successful training material.

I have copied the whole thread here, along with the resources, for future reference. I can use all of this in future development of the Pipi Ontology Engine (ont).

Update

Michael DeBellis provided an alternative approach and an example for documenting using Widoco.

Dear community,

I would like to kindly request suggestions on the best way to create online documentation for an ontology currently being developed in WebProtรฉgรฉ.

We are looking for tools or approaches that enable us to share the progress and structural details of the ontology in a clear and accessible manner for all stakeholders involved.

Thank you very much in advance for your help and recommendations.

Sincerely,

Marcelo Xavier


"Dear Marcelo,

I am personally for self documented code, then we need one or another rendering engine. Did you ask OBO Foundry and [protege-user] list?

I asked Gemini, have a look https://gemini.google.com/share/6e091b60d275

Best,

Alex"

https://www.linkedin.com/in/ashkotin/


Dear Marcelo,
Unfortunately I cannot answer your question for WebProtege specifically; Note that independently of your issue (and despite this: https://github.com/protegeproject/webprotege/issues/284), I would urge anyone developing an ontology regardless of where it is curated/edited to use a standard version control system like GitHub or GitLab for community engagement and "open science best practice" (standard workflows, etc).
A lot of OBO ontologies use the Ontology Development Kit (ODK) which comes with some built-in functions to generate an mkdocs (material-themed) documentation scaffolding which can then be extended by the ontology team. This is usually deployed on github.io (which you seem to have some personal experience with as well). Some of our docs pages are very detailed, others vanilla, see for example:
https://obophenotype.github.io/uberon/
https://obophenotype.github.io/human-phenotype-ontology/
https://oborel.github.io/obo-relations/
We mostly curate our documentation manually, but claude code or similar can, with some guidance and careful review, generate quite reasonable pages as well. 
On a more personal note:
- I like the clarity of the diataxis framework (https://diataxis.fr/) for organising docs, which we more or less try to follow in OBO Academy (https://oboacademy.github.io/obook/). 
- I really like it if modelling patterns in the ontology are documented explicitly using something like DOSDP: https://github.com/monarch-initiative/mondo/blob/master/src/patterns/dosdp-patterns/autoimmune.yaml. It is trivially possible to generate documentation pages from these to have something like this: https://mondo.readthedocs.io/en/latest/editors-guide/patterns/ (I find this is, if I may be so bold, the most important part of ontology documentation - even though hardly anyone does it).
Good luck!

Nico


I'm not sure I understand the question. There isn't anything special about an ontology developed in WebProtege. Well except that (at least IMO) developing an ontology and ONLY using WebProtege is a really bad idea. WebProtege doesn't support reasoners so you can't define axioms on classes, SWRL rules, etc. If you aren't going to use the cool stuff from OWL why use OWL at all? Go with something like just RDF or Neo4J. It's like asking "are there any tools for documenting code written using PyCharm?" The IDE shouldn't dictate how you document your code and the tool you use to develop your ontology shouldn't dictate how you document it. 

But IMO, the clear answer... at least a very good answer is Widoco which is the same tool that many people use already to document ontologies: https://github.com/dgarijo/Widoco You can export your ontology from WebProtege and use Widoco. One thing to look out for, at least this confused me for a long time, is if your ontology is just on your local file system (although it sounds like that isn't an issue for you) and you use Widoco the results will be virtually empty documentation. This isn't because of a limitation of Widoco, actually I guess in some ways it is, it is because Widoco is designed for ontologies that are hosted on the Internet and if you don't currently host on the Internet but on your local file system Widoco will generate your documentation but when you try to look at it your Operating System will block it and it will look mostly  empty. One way around that is to host Widoco documentation on GitHub. Which IMO is a much better solution in the long term than leaving it in WebProtege. WebProtege is a great collaboration tool but doesn't have the support for real software development like branching, merging, issues, etc. that GitHub does. Here's an example where I did that for an ontology I developed last year on the social science research called Climate Obstruction: https://mdebellis.github.io/Climate_Obstruction/ 

Also, in my revision of the Pizza Tutorial: https://www.michaeldebellis.com/post/new-protege-pizza-tutorial I added a chapter (chapter 11) that talks about Web Protege and how you can go back and forth between WebProtege and Desktop Protege which is what I recommend. That way you get the collaboration features of WebProtege but you can write axioms and get other features only available on the Desktop version of Protege. The nice thing is that WebProtege was designed to do what we used to call in the CASE world "round trip engineering" I.e., you can export the ontology, make changes in Desktop Protege (and/or other tools) then re-import the ontology into WebProtege and the changes are a part of the WebProtege history as if you made them in WebProtege. Also, WebProtege does some odd things to IRIs that I talk about in the tutorial.

Cheers,

Michael

https://www.michaeldebellis.com/blog 

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Ontolog Forum
  • Home > Handbook > 

Last Updated

29/01/2026

Ontology Documentation Tools and Workflows

By: Alex Shkotin
Gemini: 14/01/2026

Alex is an independent researcher in Moscow.
https://www.linkedin.com/in/ashkotin/

Gemini

Some responses below were generated according to the creator's saved info and preferences. Learn moreOpens in a new window 14 January 2026 at 21:51

Published on 14 January 2026 at 21:53

Prompt

request suggestions on the best way to create online documentation for an ontology currently being developed in WebProtรฉgรฉ.

We are looking for tools or approaches that enable us to share the progress and structural details of the ontology in a clear and accessible manner for all stakeholders involved.

Output

Creating online documentation for an ontology in WebProtรฉgรฉ is a common challenge, as the platform is excellent for editing but doesn't have a built-in "one-click" public documentation site.

To share structural details with stakeholders, the best approach is to export your ontology from WebProtรฉgรฉ and then use a specialized documentation generator. Here are the most effective tools and methods for 2026.

1. Top Documentation Tools

Once you export your ontology (usually in Turtle or RDF/XML format), use one of these tools to generate a human-readable website.

Tool Best For Key Features
WIDOCO Standard-compliant Docs A "Wizard" that generates a complete HTML site with metadata, cross-references, and a visual diagram (via WebVOWL).
LODE Quick Previews A web service that turns an OWL file into a single, clean HTML page. Great for non-technical stakeholders to read class definitions.
WebVOWL Visual Stakeholders Provides an interactive, graphical view of the ontology. Excellent for showing relationships and hierarchy visually.
OnToology Automated Workflows If you use GitHub, this tool automatically runs WIDOCO and WebVOWL every time you push a change to your repository.

2. Recommended Workflow: The "Professional" Route

If you want a live site that stays updated as your team works in WebProtรฉgรฉ, follow these steps:

  1. Export from WebProtรฉgรฉ: Go to Project > Export and select the Turtle (.ttl) or RDF/XML format.
  2. Upload to GitHub: Create a repository for your ontology.
  3. Use WIDOCO: Run the WIDOCO wizard. It will ask you for a title, description, and logo. It then generates a folder full of HTML files.
  4. Host on GitHub Pages: Enable GitHub Pages on your repository to turn that folder into a public URL (e.g., https://your-org.github.io/ontology-doc).

3. Collaborative Alternatives (Internal)

If you aren't ready to publish a full website yet, you can use WebProtรฉgรฉ’s internal features to keep stakeholders in the loop:

  • Sharing & Permissions: You can invite stakeholders to the WebProtรฉgรฉ project with "View Only" or "Comment Only" permissions. This allows them to see the structure without accidentally changing anything.
  • Threaded Comments: Use the "Comments" tab in WebProtรฉgรฉ to discuss specific classes or properties. Stakeholders can leave feedback directly on the entity they are concerned about.
  • Project Feed: This provides a timeline of changes, which is helpful for stakeholders to see the "velocity" of the project and what has been added recently.

4. Enriching Your Documentation

For any documentation tool to work well, stakeholders need clear descriptions. Ensure your team fills out these Annotation Properties in WebProtรฉgรฉ for every class and property:

  • rdfs:label: A human-friendly name (e.g., "FinancialTransaction" vs "financial_transaction").
  • rdfs:comment or skos:definition: A clear, non-circular definition of the concept.
  • dc:creator: To show who is responsible for that part of the model.

Would you like me to help you draft a set of "Documentation Best Practices" for your team to follow within WebProtรฉgรฉ to ensure the final output is high quality?