Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

No posts for a wee while

Mike's Notes

I was on holiday for the last few weeks and am back now. There will be no blog posts, newsletters or meetings until Pipi Core is back up and running.

Update 27/05/2026

Lots of surprises. Making rapid progress. The peace and quiet are bliss.

Update 31/05/2026

The problem and solution are how things are named. Pipi auto-generates thousands of code names using multiple pattern languages, and all the naming conventions require many minor fixes for several unexpected reasons after migrating from a developer laptop to a production server environment. Everything else is absolutely fine.

Other naming problems are also being solved now, including:

  • The rapid development of Boxlang by Ortus has brought forward another challenge. Pipi 10 will be migrated to run on top of Boxlang in 2027 to support multiple languages, including C++, CFML, COBOL, Go, Java, JavaScript, PHP, Python, Rust, etc.
  • Future integration with cloud-based LLMs.
  • Future integrations with Office365, Google Workspace, Zoho, LibreOffice, etc.

The common solution is to create standardised naming systems that are simple, stable, robust, schema-based, versioned, self-documenting, and extensible to meet unanticipated future needs.

This is done by replacing code-based naming rules with database-driven ones that can be easily edited in the future via an admin UI.

90% of these names are internal, hidden in the closed core, and how they work and what they are will not be discussed here. The rest will be publicly and fully documented as part of the open-source workspaces for developers to work with.

Update 02/06/2026

I'm changing the disclosure boundary between the Pipi closed-core and open-source workspaces. Previously, "disclose everything unless there is a security reason not to". This is now changed to "disclose on the basis of need to know".

Closed-core accounts for 90% and open-source workspaces for 10% of lines of code, databases, etc.

This will reduce the documentation burden, given Pipi's vast scale. So, the open-source workspaces will be fully shared and documented on GitHub, etc, without restriction. This includes;

  • Standards schema
  • Ontologies
  • Parameters
  • Laws of physics
  • HTML + CSS
  • Algorithms
  • Module DDD models
  • Workflow diagrams
  • Documentation
  • API schema
  • UI code
  • etc

This also means some existing technical documentation about the closed-core will become hidden and only available internally.

Update 07/06/2026

Pipi Core is the IDE used to edit Pipi Core (AKA: which came first, the chicken or the egg?). Temporary UIs have been created and are being used across multiple engines to edit the names in use. This is much faster than directly editing data, which had to be done initially. The next step will be turning auto-generation back on. Once that's done, temporary UIs will be used to build permanent UIs. More automation will then be enabled via the UIs, and so on, as Pipi Core builds itself with a human in the loop.

Update 08/06/2026

The list of code cases available to use now for auto-generated naming, I/O translation, etc with examples, includes;

  • camelCase: userProfilePicture
  • kebab-case: user-profile-picture
  • PascalCase: UserProfilePicture
  • snake_case: user_profile_picture
  • SCREAMING_SNAKE_CASE: USER_PROFILE_PICTURE
  • Train-Case: User-Profile-Picture
  • flatcase: userprofilepicture
  • UPPER-CASE-KEBAB-CASE: USER-PROFILE-PICTURE
  • Sentence case: User profile picture
  • Title Case: User Profile Picture
  • middot·case: user·profile·picture
  • dot.case: user.profile.picture
  • UPPER CASE: USER PROFILE PICTURE
  • lowercase: user profile picture

Update 12/06/20026

Checking that these changes to variable names and internal messaging do not clash with the GΓΆdel Machine.

Update 17/06/2026

The DevOps Engine (dvp) has unexpectedly proven to be critical to solving this puzzle. Mostly fixed last night. Watching the rather excellent live Google talk, Beyond the GPU: Maximising goodput with self-healing AI infrastructure, this morning has given me valuable insights into how to fix the remaining issues by reviewing Google HPC YAML files. 😎😎 Sometimes insights come from the strangest places.

Update 01/07/2026

The main work now is rapidly configuring Pipi for production and full autonomous automation. Using Google Search AI Mode (Gemini) and then Grammarly Pro makes the work easier and 100x faster.

  • I have decided to have Pipi re-render the many Ajabbi draft public websites with the new and missing developer information. (20K pages)
  • The website's .robot.txt file will then be unlocked to enable search engines.
  • The HTML will be updated to make it easier for AI to read.
  • This blog will be imported into Pipi, cleaned up, re-exported from Pipi, and published to Blogger via the API.
  • The new posts created in Pipi will return to A Sandy Beach to discuss something already built rather than being built.

Update 02/07/2026

The DevOps and IaC engines are getting rapid data model overhauls. The IaC engine is a great test for the variable names. I'm building a capability into Pipi to autonomously and automatically run OpenTofu and Ansible, initially targeting the Pipi Data Centre, then GCP and AWS for deployments. It's going very well and making rapid progress.

Update 05/07/2026

Pipi will initially run the open-source enterprise applications on Google Cloud Run and Google Cloud Storage (GCS). The code is complete and will be very low-cost to run, giving Ajabbi, a bootstrapping-purpose startup, a very long runway.

Update 18/07/2026

The job has now shifted to configuring, networking and deploying many physical servers. Installing software, including Pipi, labelling cables and rack gear, throwing out junk, tidying, etc., leaving nothing to chance. Shipping delays are holding up part deliveries.

Update 23/07/2026

On the basis of open collaboration and credits for experimentation, I was going to offer Google exclusive use of Pipi for a period (as a thank you) before Pipi open-source is donated to the Cloud Native Computing Foundation for all to use.

Make money to provide a service.

I'm getting exasperated with XWF. They are the external sales contractors to Google, and since 2021, they regularly contact me.

  • Selling GCP products (No need; I'm already convinced).
  • Acting as gatekeepers to any contact with Google Engineers to discuss novel integration options, which is the actual issue. How to combine Gemini (an LLM) and Pipi (non-LLM) to make something much better.
  • They are all very nice, but a complete waste of my time. No more XWF meetings, folks.

So, I have decided to target integration with OpenRouter (and its alternatives) instead of Gemini and open up the Pipi developer platform (it is big and coming 😎) to enable developers from Alibaba, Alice AI, Anthropic, AWS, Azure, ByteDance, DeepSeek, Google, IBM, Meta, Mistral, Moonshot AI, Naver, OpenAI, Oracle, Palantir, Sarvam AI, xAI, etc, and anyone else, to enable integrations that are optimal, 100% secure and vetted, with everything publicly verifiable.

Pipi closed-core will never be for sale; this year it's getting a non-profit foundation behind it, a bit like Patagonia. I'm open to all genuine offers of assistance, collaboration and experimentation with no strings attached. Contact me.

Don't send sales engineers

Send a senior, highly experienced engineer/architect/chief scientist who loves a big fat problem and has time for an open chat without a pitch or an agenda, and just see where it goes.

If you want to meet in person, expect to work collaboratively at a whiteboard or blackboard like a real mathematician. Plus coffee, of course. 😎 To see how this works, watch the seminars at the London Institute of Mathematical Sciences, or the recorded physics seminars at Perimeter.

Pipi is rooted in biology and the laws of physics, so you need a very solid background in advanced sciences (microbiology, biochemistry, mathematics, philosophy, particle physics, thermodynamics, complex adaptive systems, etc).

Please, no venture capitalists or private equity. You're wasting your time. Go find something else to plunder. Pipi is a gift to humanity.

Update 28/07/2026

Most of the equipment has arrived, and the small data centre setup is coming together. More deliveries later this week. It's already running a lot better and is much more productive.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

28/07/2026

No posts for a wee while

By: Mike Peters
On a Sandy Beach: 15/05/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

I was on a no-coding holiday for the last few weeks to clear my mind, and it has been great. I am back on the job today.

Suspended

Until the closed-source Pipi Core is back up and running 100% on autopilot, 10x faster, the following are suspended.

  • New posts "On a Sandy Beach
  • All newsletters, including the weekly Friday Report and the monthly Ajabbi Research Newsletter.
  • The fortnightly online Open R&D meeting.

Rapid refocus

  • A new developer area with five coding screens, designed to be more productive for hypervisual learners.
  • A better library has been set up for my A4 drawings in ring binders, the many reference books I use, and more bookshelves are on the way.
  • The server rack has been moved to a better location.
  • The light levels have been adjusted.
  • A big office tidy is almost done. An office-work-only desk has yet to be set up with a cat bed included.
  • A separate area with no screens for the happy cat, coffee, music, reading and drawing.

Less is more

Minimise screen time to be more productive at work. The new setup is also much less tiring.

Get the job done

The good thing is that, with a holiday and lots of drawing, I now have mental clarity about what needs fixing and how to fix it. Mainly, quite delicate changes here and there, organised into a list of steps. Now, I need to concentrate on one thing only: go as fast as possible, without meetings, post-deadlines, phone calls, or other distractions.

How

1. Use an AI workforce

Be the architect, and AI fills in the dots to make it happen.

Use Google Search AI mode (Gemini) to generate 99% of the code in one-page chunks (including references) to copy and paste, then manually change the variable names and SQL. Careful, test everything, resulting in 100x faster progress. Know how everything works and rapidly raise personal skill level.

2. Then build a cathedral

Make a wooden scale model of a cathedral for the builders. Google Search AI mode (Gemini) makes each brick, and Pipi Core assembles the bricks into floors, arches, walls, and vaults...

Speed is king

With the 100x coding productivity gains from Google Search AI mode (Gemini), plus the 10x10x10x speedup of Pipi Core currently underway over the next few months, what previously took a year will be done in hours and better.

Phase transitions

Once these initial migration issues from laptop to server are resolved, further transitions can be anticipated as the number of engines rapidly increases beyond 20. Increasing the number of engines slowly changes the whole system's behaviour from deterministic to probabilistic and adaptive.

Here is a partial list of transitions expected as the number of engines increases from 0 to 200. The actual numbers are a bit of a guess.

  • 20 engines enable Pipi 9 Core in a simple, deterministic structure.
  • 40 engines enable a workspace with a UI for administering Pipi Core.
  • 60 engines enable self-generation of user documentation.
  • 80 engines enable REPL and IAC (infrastructure-as-code).
  • 100 engines enable Workspaces for different user accounts.
  • Different Pipi 9 editions are made with the same engines, which recombine differently in response to the external environment.
  • And so on until...
  • 200 engines self-organise into a multi-layered complex fluid structure with probabilistic behaviour and emergent properties, as engines also act as agents.
  • 200+ engines enable Pipi 10 to interact with externally cloud-hosted LLMs, combining the very different strengths of both.

Context Engineering for Coding Agents - Fausto's Amsterdam workshop

Mike's Notes

The MLOps Community is fantastic, and it has a regular newsletter from Demetrios.

100% better than anything coming out of NZ  or Australia. I attend MLOps events remotely whenever possible. Now I know where to find the talent capable of working on Pipi in future.

This is taken from a recent newsletter. Interesting how deterministic and probabilistic contexts are handled in generative AI.

All Pipi Engines are both deterministic and probabilistic as required.

The April 21 2026, MLOps Community Netherlands workshop video and slides are available.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > MLOps Community
  • Home > Handbook > 

Last Updated

10/05/2026

Context Engineering for Coding Agents - Fausto's Amsterdam workshop

By: Demetrios Brinkmann
MLOps Community: 07/05/2026

Demetrios founded the largest community dealing with productizing AI and ML models. In April 2020, he fell into leading the MLOps community (more than 75k ML practitioners come together to learn and share experiences), which aims to bring clarity around the operational side of Machine Learning and AI. Since diving into the ML/AI world, he has become fascinated by Voice AI agents and is exploring the technical challenges that come with creating them.

Fausto's workshop focused on the part of agent systems you can control: what gets injected into the context window, when it gets loaded, and what should live somewhere else. You are not training the model on a daily basis, so the engineering work shifts to managing context, memory, tools, and retrieval.

His rule of thumb was to keep context usage under 25%, regardless of whether you are working with a 200k or 1M token model. Past that, things get slower, more expensive, and more error-prone, with cleaner options than letting the context window bloat.

The lens for the rest of the talk was a brain analogy. Attention is finite, memory shapes attention, and knowledge does not get embedded in a vacuum. An agent will only notice the right things if you have given it the right priors.

From there, Fausto split context into three categories.

  • Deterministic context includes CLAUDE.md, project rules, hooks on lifecycle events, auto-memory, and scheduled loops. His point on rules was especially practical: coding conventions belong in path-scoped or file-extension-scoped rules, not dumped into CLAUDE.md.
  • Human context covers chat turns, slash commands, and references.
  • Probabilistic context covers sub-agents, retrieval, MCPs, skills, and observers. Sub-agents are useful because they do not inherit CLAUDE.md, memory, or the default system prompt, which makes them better suited for specific non-coding tasks. Skills are markdown plus optional scripts, and can also wrap calls to other models when Claude cannot handle the job. Fausto's example was routing native video analysis through Gemini.

Two practical tips came out of this section. Turn on deferred tool loading, a single flag that reduces what gets injected at session start. And lean toward project scope over user scope for skills and MCPs, so the agent has the full descriptions it needs to choose well at runtime.

The second half made the case for treating long-term memory as a folder of markdown files rather than defaulting to a vector store, inspired by Karpathy's wiki memory idea.

The structure is an index plus raw source files, processed summaries, and a policy that decides what to ingest and retrieve. Concepts that recur get weighted up. Concepts that go unused decay over time. An observer agent watches the session and either pulls relevant knowledge into the active context or pushes new findings into the wiki.

To show the difference in practice, Fausto ran two Claude Code sessions on the same cellular automata task. Same model, same skills, same sub-agents, same CLAUDE.md. The only difference was that one had a populated wiki and the other did not. Under a five-minute timer, the wiki run pulled the concepts it needed and produced a working visualization. The default run fell back on parametric memory and live search, then ran out of time.


Context Engineering for Coding Agents: Hands-on Lab + Agent Build-Off

21 April 2026
Prosus NV
Gustav Mahlerplein 5, 1082 MS Amsterdam, Netherlands

A hands-on workshop on configuring Claude Code for real work.
​Claude Code is powerful out of the box. But there's a gap between using it and mastering it. Closing that gap is mostly a matter of context: what the agent knows, where its memory lives, which rules it follows, and how its output gets verified before it ships.

​This workshop is three hours of hands-on practice to optimize your Claude environment. After a short primer you join a pre-assigned team and build out a Claude Code setup in a prepared sandbox — working on memory, rules, hooks, retrieval, and prompt quality. In the final stretch we drop an unknown technical drawing on the screen and your system gets one run at it. Scores go on a live leaderboard, winners are announced in the room.

​What you'll practice

  • ​Shaping what the agent sees and remembers
  • ​Writing rules and hooks that actually fire
  • ​Designing memory so it finds what matters
  • ​Judging output before it leaves the loop

Format

​One continuous team build, one unknown challenge, one live leaderboard, closing with pizza and drinks. — The system you build in 60 minutes is the score you receive.

​Agenda

  • ​5:00 PM — Walking dinner & Drinks
  • ​please be on time :)
  • ​5:30 – 5:45 PM — Welcome by Prosus
  • ​Opening and framing for the evening.
  • ​5:45 – 6:00 PM — Opening
  • ​6:00 – 7:00 PM — Theory + Mini-Demos
  • ​7:00 – 8:00 PM — Build Lab
​Six phases: mission lock, sandbox orientation, second brain, skills + guardrails, self-test, freeze. Teams build their Claude Code system against the run contract. Facilitators float; the deck runs a silent timer.
  • ​8:00 – 8:15 PM — Drawing Reveal + One-Shot Run
  • ​An unknown technical drawing is broadcast into every team sandbox. Each system gets a single run. Whatever it produces is what gets scored.
  • ​8:15 – 8:20 PM — Processing Break
  • ​Last runs finish, evaluator collects output.
  • ​8:20 – 8:30 PM — Live Leaderboard + Winners
  • ​Scores land on screen. Jury walks through the top runs and the most interesting architectural choices.
  • ​8:30 – 9:00 PM — Pizza + Networking
  • ​Pizza, drinks, Q&A.

​Who this is for

​Technical practitioners who have already used Claude Code (or a comparable coding agent) and want to push past the default setup. Comfortable on the command line, and with Github, fluent enough in Python to read and edit a small repo.

​What we provide

  • ​A pre-configured cloud sandbox per team, with Claude Code installed and API access included. Minimum requirement is to have a Claude Pro account — and bring your laptop.
  • ​Registering is not a confirmed seat — we curate the room and send confirmations separately
​Delivered by Fausto Albers · GenAI R&D, AUAS · WonderWhy.ai

YouTube 2:27:28

The Neural Harness: The new CPU

Mike's Notes

Some deep insights here from Will Schneck. Asking more questions than he answers. Especially deterministic vs probabilistic. Where does emergence emerge? 😎😎

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > The Focus AI
  • Home > Handbook > 

Last Updated

03/05/2026

The Neural Harness: The new CPU

By: Will Schenk
The Focus AI: 01/05/2026

I am a father, entrepreneur, technologist and aspiring woodsman.

My wife Ksenia and I live in the woods of Northwest Connecticut with our four boys and one baby girl. I have a lumber mill and all the kids love using the tractor.

I’m currently building The Focus AI, Umwelten, and Cornwall Market.

"Coding agents will build their own tools and their own agents. Agents will be used by non-engineers to manage other agents to manage parts of the org chart."


I'm on my second Claude Max plan. That's in addition to Cursor, Codex, Gemini, and a healthy Amp habit. Not to mention a Jetson AGX Thor I'm about to plug in at the office — more on that one later.

Overnight jobs parsing financial deal structures, ops stuff, research, monitoring logs, responding to events, all the little background things. The first plan tapped out, I added another, that one tapped out too, and now I'm provisioning a third the way you'd add a build runner. Mundane.

A new entry in an old list

Look at the paragraph I just wrote. Overnight jobs parsing financial deal structures, ops stuff, research, monitoring logs, responding to events. Half of those words are themselves names of native units of computing. Logs — log aggregators. Events — event streams. Research — search indexes. Ops — schedulers, orchestrators, deployment systems. Jobs — queues. The lede is already a list of older units I'm wiring into.

Computing has been accreting native units forever, and the way you build the next layer is by composing the units underneath it.

You combine adders and accumulators to make a CPU. You combine CPUs and memory and a bus to make a machine. You combine logic gates and clocks to make registers. You combine Boolean functions and a process model to make an operating system. You combine lexers and parsers and code generators to make a compiler. You combine source files and a compiler to make a program. You combine programs and a network stack to make a service. You combine services and a database to make an application. You combine applications and a queue to make a pipeline. You combine pipelines and a stream processor to make a real-time system. You combine streams and a log aggregator to make observability. You combine logs and a metric and an anomaly model to make a monitor. You combine all of it and a scheduler and you have a system that runs without you watching it.

Flat-color treemap of the computing stack: small blocks for adders, clocks, registers, cpu, memory, bus growing diagonally up and to the right through machine, os, compiler, program, service, application, pipeline, stream, observability, monitor — culminating in a large block for scheduler.

Each layer is just the layer below, composed. That's what a native unit is — the thing you stop writing yourself, the thing you wire to. You don't write a compiler. You don't write a Postgres. You don't write a Kafka or a Kubernetes or a Lucene or a git. You pick the unit, you combine it with other units, you build on top.

Now look at that list again. Everything on it is sitting on top of Boolean logic. Silicon, gates, arithmetic, state machines, if/then. Numbers, types, queries, schedules, indexes — all of it is deterministic logic resolving down to ones and zeros. You can climb that stack pretty high, but you don't get out of it.

19th-century geological cross-section: layered strata of fossilized circuit traces — gates, clocks, registers, CPU, OS, compiler, program, service, application — opening into a newly excavated neural floor below, soft coral and cream tones, neuron tendrils threading up into the rock. The new floor under the old stack.

Neural nets aren't more of that. They're a different kind of logic. Pattern, association, similarity, fuzzy matching, generation. The thing silicon-and-Boolean was bad at, that we kept failing to solve with cleverer rules, the neural net does natively. We added a new floor — GPUs, TPUs, the Cerebras inference fabric, the Jetson on my desk — and a new kind of computation running on it that doesn't reduce to if A and B then C.

By themselves these things predict tokens. They don't loop, they don't read files, they don't remember. To get computation out of one you wrap it. A loop, some tools, file access, a shell, a way to manage context. That wrapper is the harness. The harness is the unit that turns "predicts the next token" into "does the work" — and lets the new kind of logic compose with the old kind.

The neural harness is to neural nets what the compiler was to source code. New entry on the list, joining the family rather than replacing it. The work I'm running on these two-going-on-three Max plans is mostly the harness wiring into the older units — tailing logs, querying state, watching streams, kicking off jobs, hitting indexes. New unit, old units, composed.

That's why the second Max plan isn't weird. The bill scales with how much work you're doing in the new unit. I'm doing a lot of work in the new unit.

How it shows up in a day

It really has stopped being a tool I reach for; its just the tool.

When I'm coding, I'm in a harness. When I'm reading a PDF I needed to read anyway, the harness is the thing reading it. Operations folder — SOWs, invoices, content ideas, project status — that's a harness. Parsing 20 financial deal docs and writing me a summary while I sleep — harness. Family infographics, fasting tracker, oura ring trends — harness, harness, harness. Different work, same unit.

Small-multiples grid: the same harness icon repeated across sixteen everyday domains — code editor, PDF reading, invoice, SOW draft, content ideas, project status, financial deal, overnight job, log monitor, email triage, family infographic, fasting tracker, Oura trends, calendar, research note, ops dashboard. One unit, many domains.

Coding was just the first place this paid off, because the feedback loop is tightest. Compile or don't, test or don't, the world tells you you're wrong inside a second. So that's where the harness got tuned first. That's why the unit is called a "coding agent" right now. But "coding" is vestigial. The thing isn't a coding agent. It's a harness around a model, and what runs in it is whatever you have tools for.

Rick Blalock said it in AI Engineering Miami — coding agent as universal software primitive. A 60-year-old in Texas replaced a $10k/month HubSpot bill by pointing one of these at the problem for three months. A 24-year-old window cleaner in Florida runs marketing, sales, and estimating off the same primitive. Both of them bought MacMinis. Tim Cook didn't have that on his bingo card.

The model question is below the harness question

Here's something I noticed about my own behavior: I'm mainly on Claude. Have been for months. I dip in and out of GPT and Grok and Gemini, but just sort of end up back here. Not because I reasoned out a model strategy — because Claude Code defaults to it and now I'm on Opus all day every day. Amp has its opinion and I try to set Cursor to super max mode, but really the model picked itself by way of the harness picking it for me.

So the perennial "Opus vs GPT-5 vs Gemini 3" argument is pitched one floor below where the action is. It's not model-vs-model. It's harness-with-default-model vs other-harness-with-default-model. The harness drives the model choice, often without telling you.

And underneath that, there's a whole zoo. Frontier reasoning models. Cheap fast models. Code-specific fine-tunes. Local models that run on the GPU you already own. Cerebras-fast inference at 1,200 tokens/sec, a different regime entirely. And the inside-the-harness thing: Tejas Bhakta at Miami called it "everything is models" — a compaction model running every two seconds, a code-search model at 80k tokens/sec, a frontier model doing only the heavy reasoning, all stitched together. Software 3.5, he called it. The harness picks all of that for you, or doesn't, depending on which harness.

Da Vinci anatomical-plate: a single mechanical harness apparatus on top labeled HARNESSIS — UNITAS SUPERIOR with four tool-attachments (LEGERE, SCRIBERE, IMPERARE, ITERARE), and below it a labeled menagerie of seven model 'species' — Frontier, Velox, Codicis, Localis, Compactionis, Quaerens, Cerebras — drawn as small mechanical creatures on aged parchment.

Which means the harness is a model strategy. Picking a harness on purpose means picking which models do which jobs inside it.

So which harness?

A separate post coming soon — each one deserves its own treatment and the conversation moves week to week. The shape of it:

You can build your own in a weekend. About 50 lines gets you the loop. Highly recommend, even if you never use it. Claude Code is the one everyone uses, and — by Anthropic's own model on Anthropic's own benchmark — the worst Claude harness on offer. (Niels Rogge posted Terminal-Bench 2: same Opus 4.6, Claude Code last, ForgeCode and Capy at 70-75%. Twenty-five points of accuracy from picking a different harness.) Picode is Mario Zechner's minimal, self-modifying one — four tools, the agent writes its own extensions, hot-reloads in the session. The most fun one to play with right now. Amp is the one I'm most fascinated with — though to be clear, I'm editing this post in Cursor. The multimodel thing actually works now. In January I wrote that Amp "should be better, but, you know, isn't." Four months later: it is.

Tufte-style horizontal bar chart of Terminal-Bench 2 scores on the same Opus 4.6: ForgeCode 74%, Capy 71%, Picode 62%, Amp 55%, Claude Code 49% — outlined in vermilion. Annotation on the right: 25-point gap from picking a different harness.

The point of this post is the unit, not the catalog.

What I'm still circling

Da Vinci notebook spread with marginalia: a Jetson AGX Thor on a small workbench labeled MACHINA LOCALIS, a brass token-cost gauge labeled STIPENDIUM TOKEN — quanto?, a half-configured harness with question marks labeled HARNESSIS CONFIGURATA — UNITAS NAVIS?, and a small chart of a rising line labeled LINEA NOVA IN STATU FINANCIALI. Sepia ink on parchment, inkwell and quill in the corner.

What's the unit of shipping? Ben Davis's claim in Miami was that it's becoming a directory of skill files plus a coding-agent runtime. That feels right. But the runtime is also moving — picode's bet is that it should be malleable inside the session, so you can't pin it. Maybe the unit is even smaller. Maybe the unit is the harness, configured.

What about the Jetson on my desk. The other thing the bill is about to teach us is that some of this work shouldn't be paying a subscription at all. Local models on local hardware — gpt-oss, Qwen, MiniMax, whatever's frontier-enough for the job — running on the GPU you already own, or the Jetson, or the laptop. Cheap as electricity. No data leaving the building. The harness doesn't care which model it's calling. The bill cares a lot. I think a real chunk of what's running on the second Max plan ends up local by the end of the year.

When the bill becomes a real line item — and it will — what does that conversation sound like? "Cloud spend" took ten years to become its own column on the financial statement. "Token spend" might take less. We're paying for a unit of computation, not for software. Different shape entirely.

I'll get the third Max plan tomorrow. There's another job.

Anthropic Mythos -- We've Opened Pandora's Box

Mike's Notes

This is why Pipi Core is in its own data centre, physically isolated from the internet, to ensure 100% security and protect people's privacy.

I endorse Steve Blank's conclusion. The risks are enormous and growing.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Steve Blank
  • Home > Handbook > 

Last Updated

01/05/2026

Anthropic Mythos -- We've Opened Pandora's Box

By: Steve Blank
The Cipher Brief: 23/04/2026

Adjunct Professor at Stanford and Co-founder of the Gordian Knot Center for National Security Innovation

Steve Blank is an adjunct professor at Stanford and co-founder of the Gordian Knot Center for National Security Innovation. His book, The Four Steps to the Epiphany is credited with launching the Lean Startup movement. He created the curriculum for the National Science Foundation Innovation Corps. At Stanford, he co-created the Department of Defense Hacking for Defense and Department of State Hacking for Diplomacy curriculums. He is co-author of The Startup Owner's Manual.

EXPERT OPINION

For a decade the cybersecurity community was predicting a cyber apocalypse tied to a single event - the day a Cryptographically Relevant Quantum Computer could run Shor’s algorithm and break the public-key cryptography systems most of the internet runs on. We braced for a one-time shock we would absorb and adapt to. The National Institute for Standards and Technology (NIST) has already published standards for the first set of post-quantum cryptography codes.

It’s possible that the first cybersecurity apocalypse may have come early. Anthropic Mythos now tilts the odds in the cybersecurity arms race in favor of attackers - and the math of why it tilts, and how long it stays tilted, is different from anything our institutions were built to handle.

In 2013, Edward Snowden changed what people knew

In 2013, Edward Snowden changed what people understood about nation-state cyber capabilities. In the decade that followed disclosures and leaks of nation state cyber tools reduced uncertainty and accelerated the diffusion of cyber tradecraft.

The defensive playbook that followed - compartmentalization, need-to-know, leak-surface reduction, clearance reform, “worked” because the Snowden leaks and those that followed were one-time disclosures, absorbed over a decade, with the system returning to something like equilibrium.

We got good at responding to the shocks of disclosures. It became doctrine. It was the right doctrine for the wrong future.

Pandora's Box

In 2026, Anthropic Mythos (and similar AI systems) is changing what people can do. Mythos found Zero-day vulnerabilities and thousands of “bugs” that were not publicly known to exist (a must read article here.) Many of these were not just run-of-the-mill stack-smashing exploits but sophisticated attacks that required exploiting subtle race conditions, KASLR (Kernel Address Space Layout Randomization) bypasses, memory corruption vulnerabilities and logic flaws in cryptographic libraries in cryptography libraries, and bugs in TLS, AES-GCM, and SSH.

The reality is a number of these were not “bugs.” There were nation-state exploits built over decades.

What this means is that Anthropic Mythos, and the tools that will certainly follow, has exposed hacking tools previously only available to nation-states and transformed into tools that Script Kiddies will have within a few months (and certainly within a year.) No expertise will be required to apply that tradecraft, compressing both the learning curve and the execution barrier.

All Government’s Will Scramble

When Mythos-class systems are used to analyze the code in critical infrastructure and systems, the hidden sophisticated zero-day exploits that are already in use, (including ones nation-states have been sitting on for years) will be found and patched. That means intelligence agency sources of how to collect information will go dark as companies and governments patch these vulnerabilities.

Every serious intelligence service will scramble, likely with their own AI, to find new access before the visibility gap costs them something they cannot replace. A new generation of AI-driven exploits will rise to replace the ones that have been burned.This will build an arms race with a new generation of AI-driven cyber exploits looking to replace the ones that have been discovered. Whichever side sustains faster AI adoption - not just “procures” it, but ships it into operational systems, holds a widening advantage measured in powers of two every four months.

The binding constraint is not budget. Not authority. Not access to models. It is institutional capacity for change - the rate at which a defender organization can actually change what it deploys.

The Long Tail Will Not Be Patched

Anthropic has given companies early access to secure the world’s most critical software. That will help Fortune 100 companies. But the Fortune 100 is not just a small part of the software attack surface.

The attack surface includes the unpatched county water utility, the regional hospital, the third-tier defense supplier, the school district, the state Department of Motor Vehicles, the municipal 911 system, and the small-town electric co-op. Tens of thousands of systems running software nobody has time to patch, maintained by teams that have never heard of KASLR.

Every one of those systems is now exposed to nation-state-grade tradecraft, wielded by attackers with no expertise required. Mythos-class hardening at the top of the pyramid does not trickle down. The long tail will stay unpatched for years.

Attackers Advantage - For Now

Under continuous exponential growth of AI designed cyberattacks, a cyber defender using traditional tools can't just respond just once and stabilize their systems. They’ll need to keep investing at a rate that matches the offense's growth rate itself. A one-time defensive shock like compartmentalization might work against a sudden attack, but it will fail against sustained exponential pressure because there's no stable equilibrium to return to. The defender's investment rate has to track the offense's growth rate.

Ultimately and hopefully, the next generation of AI driven cyber-defense tools will create a new equilibrium.

What We Need to Do

Mythos and its follow-ons will change how we think about cyber-defense. We can’t just build a set of features to catch every exploit x or y. We need to build cyber systems that can maintain or exceed the capability rate of the attackers.

Here are the three tools governments and cyber defense companies need to build now:

  1. Measure the Gap Between Attackers and Defenders. We need to know the gap between what the attackers can do and what we can defend against. We need to develop instrumented red/blue exercises (a simulation of a cyberattack, where two teams – the red team and the blue team – are pitted against each other) to estimate the number of new vulnerabilities vs cyber defense mitigation. (This can be built in six months, with a small team.)
  2. Measure the Defender Response Time. For each corporate or government mission system, measure how long it takes to implement a change from identification to production deployment. Treat each organizational obstacle as equivalent to technical debt that needs to be remediated.
  3. Specify Speed, Not Features. Any new Cyber Defense tools and architecture - including the next-generation cloud-native systems sitting in review right now - should have explicit ‘rate’ requirements. Claims of “our product delivers X capability is now the wrong specification. “Closes detection gap at rate greater than or equal to the offense growth rate” is the right one.

Buckle up. It's going to be a wild ride - for companies, for defense and for government agencies.

Mythos is a sea change. It requires a different response than what the current cyber security ecosystem was built for, and one the current system is not built to produce. We are not behind yet. The gap between Mythos and what we can build to defend is small enough today that a serious response can still match it. A year from now, the same response will be eight times too slow. Two years, sixty-four.

By the way, the only thing left in Pandora’s Box was hope.

Vibing, Harness and OODA loop

Mike's Notes

Wise words from Oskar Dudycz. Subscribe to Architecture Weekly, it's awesome. 😎

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Architecture Weekly
  • Home > Handbook > 

Last Updated

30/04/2026

Vibing, Harness and OODA loop

By: Oskar Dudycz
Architecture Weekly: 27/04/2026

Software Architect & Consultant / Helping fellow humans build Event-Driven systems. /Blogger at https://event-driven.io/ . OSS contributor https://github.com/oskardudycz /Proud owner of Amiga 500.

On why Vibing and Harness are not new and why feedback loops are important.

Hey, have a look at what I made during the weekend. I had some time, grabbed a beer, turned on the computer and tried to code this feature. If I could do so much during the weekend, how much could you and your team do with it in 2 weeks?

It’s almost a 1:1 quote of what I heard from the startup founder I worked with over 10 years ago. I’m sure that you’ve heard similar phrases from people you worked with. We all know the annoying type of person who doesn’t code anymore but thinks, “I still got it!”. Then they threw a piece of stuff at you to “just fine-tune it a bit and do final touches”. Then they’re the first ones to ask “Why so long?“.

Nowadays, the Internet is full of such people. They shout about what they did with Claude or how much progress LLM tools have made. Some even predict the end of coding. I already wrote that this is wrong perspective. I won’t repeat that, but I want to say that…

Vibing isn’t new and isn’t always an issue.

I’m saying that LLM tools are an appraisal for ignorance. The more ignorant we are of the topic we’re working with, the better we see the outcomes. And that, by itself, is not always bad, as there’s power in ignorance if we focus on getting it done with the simplest tools we have.

Still, this can be terrible if we fall in love too much with what we’ve vibed.

To understand why that “weekend beer” energy is both a superpower and a liability, we need to look at the OODA Loop.


Disclaimer, it’s not a competition for Ralph Wiggum Loop. It’s much older and generic.

Military strategist John Boyd developed the OODA loop (Observe, Orient, Decide, Act) for fighter pilots. In a dogfight, the pilot who cycles through these four stages the fastest and most accurately survives.

In software, the “dogfight” is the gap between your intent and the production-ready feature.

OODA loop is built from four steps:

  1. Observe - This is the intake of raw, unfiltered information. In our world, this means looking at the state of the system.
  2. Orient - This is the most critical and difficult stage. It’s where you filter your observations through your experience, culture, and technical knowledge.
  3. Decide - Based on your orientation, you formulate a hypothesis.
  4. Act - You execute.

Getting back to my favourite founder and LLM-based tools.

The reason founder could build a PoC in a weekend while the team needed more than two weeks is that he bypassed the Observe and Orient phases. He went straight from a vague idea to Act.

If we skip or brush past the observation step, it feels like lightning speed. If the fancy UI grid is there and it does something we wanted, we move on. We’ve outsourced Orientation to our own ego. It’s too easy to assume that because we wrote it, it works.

Observation is the intake of raw data. In a professional environment, our eyes aren’t enough. We need a Harness. If we don’t have automations, tests, integration tests, and pristine traces, we aren’t observing the system; we’re just looking at it. If the inputs are messy, our observation is clouded.

But real engineering, the kind that takes those “two weeks”, is about closing the loop properly. That’s also where we need different perspectives and knowledge sharing.

Orientation is where you process those observations. This is the part where LLMs make us feel smarter than we are. If we don’t understand how a database handles concurrent connections, our “orientation” of a generated script will be shallow. We’ll see code that “looks” right, decide it’s fine, and act by deploying it.

The “I still got it” crowd loves the Decide and Act phases because that’s where the visible progress happens. LLM tools have made these phases nearly instantaneous. We can decide to build a feature and have the code for it in ten seconds.

The problem is that the faster we Act, the faster we need to Observe. If our “Act” phase takes seconds but our “Observe” phase requires a manual weekend of clicking around and drinking beer, our OODA loop is broken. We’re just generating a pile of stuff that we haven’t actually verified.

That’s why the team usually needs more than an imaginary “two weeks”. They are not “fine-tuning” the single-brilliant-dude masterpiece. They are building the infrastructure required to make the OODA loop sustainable.

And to make that possible, they need to run the full loop: Observe, Orient, Decide, Act. And do it multiple times. That takes time, but it’s required to assess the direction, automate what needs to be automated, and ensure they can iterate further and run this loop sustainably. That’s critical for delivering the outcome at the expected pace.

Of course, there’s a danger here, overfocusing on the Orient and Decide can lead to overengineering, building stuff we don’t need. That’s where ignorance can be blissful, especially when we connect it with humility. Being humble about what we don’t know and trying things the easiest way, then learning and making enhancements. Still, humility fails under deadline pressure. The harness doesn’t.

Let me give you…

The example

I’m adding proper Observability and Open Telemetry to Emmett right now. I spent some time working on it and instrumented the first component: Command Handling.

Of course, I had tests to prove it works, but I don’t trust them enough, and I wanted to try it on a real sample, since you never know until you run it. Even the best test suite won’t tell you all.

So I decided to plug it into the sample. See if it works, how ergonomic the API is and how it fits conventions in this area.

To do it, I decided to use Grafana stack and set it up with Docker Compose. So, stable, boring stack. Not going to lie, I vibed the config. Not that there are no docs, but I intentionally wanted to see the typical config people use.

If someone says LLM-based tools are great at proof of concepts, they don’t run the stuff they vibed. If I made the observation based on the initial config, then an oriented decision would be that it won’t work. Of course, then I did the typical back-and-forth, with the LLM tool doing some Linux command Voodoo to make it work. Once. Then, if you try to repeat it, you won’t know how to do it without doing Voodoo again.

Again, that’s not much different from the other stuff we do. I’m sure that you had multiple cases, when someone didn’t use Continuous Deployment tools, but clicked through Azure, AWS, GCP portal, deployed the stack, and then there was no trace on how to set it up again (e.g. to have a different environment for testing or demos for customers).

So, we need a harness, we need a leash to keep our process on track.

How to do the harness? My advice is to start simple. We may ask LLMs to give us shell scripts, and we may ask them to run them multiple times. We also need experience and knowledge of what we want to achieve and the tools we use. It’s fine not to remember all the YAML config to set up the Grafana stack, but it’s not fine not to understand why you even use it, how it relates, and how to set it up.

Still, our first loop can close on the first working solution, even a manually vibed one. But that’s not even a PoC. We need to automate them.

I asked LLM to take notes on what issues it had, and it solved them. Then, based on that, I asked to research how to code it in TypeScript. And to use tools I know, used in past, validating if there are no new more modern ones. For instance, I was a big fan of Gulp.js and Bullseye in the past, but they’re mostly dead. I wanted to have something in the same spirit, using native, maintained tooling.

I ended up with the following tools:

  • execa for running shell scripts,
  • native fetch for calling http endpoints,
  • native Node.js test tools for checking if the stack works as expected.

Then I asked it to create the script to automate the shell Voodoo they did to make Grafana stack and Docker Compose work.

Essentially, it should:

  1. Run Docker Compose script starting up services (Grafana, Prometheus, Loki, Tempo, PostgreSQL, etc.).
  2. Wait for them to check when they’re ready (it usually takes some time).
  3. Start the application and make a request.
  4. Check if the predefined dashboard with Emmett metrics appears, and shows expected traces and metrics.

Initial diagnostic tools looked like that

async function fetchWithDiag(label: string, url: string, init?: RequestInit) {
  const res = await fetch(url, init);
  if (!res.ok) {
    const body = await res.text().catch(() => '(could not read body)');
    console.error(`\n  ✗ ${label} → HTTP ${res.status}\n  body: ${body}\n`);
  }
  return res;
}
async function diagnoseCollector() {
  const text = await fetch(URLS.otelCollectorMetrics)
    .then((r) => r.text())
    .catch(() => 'unreachable');
  const emmett = text
    .split('\n')
    .filter((l) => l.startsWith('emmett_') && !l.startsWith('#'))
    .slice(0, 5);
  console.log(
    emmett.length
      ? `\n  collector /metrics (emmett lines):\n  ${emmett.join('\n  ')}`
      : '\n  collector /metrics: no emmett_* lines found',
  );
}
async function diagnosePrometheus() {
  const json = await fetch(
    `${URLS.prometheus}/api/v1/label/__name__/values`,
  )
    .then((r) => r.json() as Promise<{ data: string[] }>)
    .catch(() => ({ data: [] as string[] }));
  const emmett = json.data.filter((n) => n.startsWith('emmett_'));
  console.log(
    emmett.length
      ? `\n  Prometheus emmett_* metrics: ${emmett.join(', ')}`
      : '\n  Prometheus: no emmett_* metrics found yet',
  );
}
async function diagnoseLoki() {
  const labels = await fetch(`${URLS.loki}/loki/api/v1/labels`)
    .then((r) => r.json() as Promise<{ data?: string[] }>)
    .catch(() => ({ data: [] as string[] }));
  console.log(`\n  Loki labels: ${(labels.data ?? []).join(', ') || '(none)'}`);
}
async function diagnoseDockerLogs(service: string, lines = 10) {
  const { stdout } = await execa('docker', [
    ...COMPOSE,
    'logs',
    '--tail',
    String(lines),
    service,
  ]).catch(() => ({ stdout: '(could not get logs)' }));
  console.log(`\n  docker logs ${service} (last ${lines}):\n  ${stdout.split('\n').join('\n  ')}`);
}

Are they pretty? No. Can they be improved? Yes. Do they have to be improved at this specific moment? No.

The setup uses test infrastructure

const CLEANUP = process.env['CLEANUP'] === '1' || process.env['CLEANUP'] === 'true';
const CLEANUP_AFTER = process.env['CLEANUP_AFTER'] === '1' || process.env['CLEANUP_AFTER'] === 'true';
const NO_START = process.env['NO_START'] === '1' || process.env['NO_START'] === 'true';
// ─── configuration ───────────────────────────────────────────────────────────
const COMPOSE = ['compose', '-f', 'docker-compose.yml', '--profile', 'observability'];
const URLS = {
  app: 'http://localhost:3000',
  prometheus: 'http://localhost:9090',
  tempo: 'http://localhost:3200',
  loki: 'http://localhost:3100',
  grafana: 'http://localhost:3001',
  otelCollectorMetrics: 'http://localhost:8889/metrics',
};
// Fresh client per run — avoids stale cart state from previous runs.
const SERVICE_NAME = 'expressjs-with-postgresql';
const CLIENT_ID = randomUUID();
const CART_ENDPOINT = `${URLS.app}/clients/${CLIENT_ID}/shopping-carts/current/product-items`;
const CONFIRM_ENDPOINT = `${URLS.app}/clients/${CLIENT_ID}/shopping-carts/current/confirm`;
// Matches the .http file — unitPrice is resolved server-side.
const ADD_PRODUCT_BODY = JSON.stringify({ productId: randomUUID(), quantity: 10 });

before(async () => {
  console.log(`\n▶ client ID for this run: ${CLIENT_ID}\n`);
  if (NO_START) {
    console.log('▶ --no-start: skipping docker compose and app startup');
    return;
  }
  if (CLEANUP) {
    console.log('▶ --cleanup: killing port 3000 and tearing down stack (down -v)…');
    await execa('bash', ['-c', 'fuser -k 3000/tcp 2>/dev/null || true']).catch(() => {});
    await new Promise((r) => setTimeout(r, 500));
    await execa('docker', [...COMPOSE, 'down', '-v', '--remove-orphans'], {
      stdio: 'inherit',
    });
  }
  const stackReady = await fetch(`${URLS.prometheus}/-/ready`)
    .then((r) => r.ok)
    .catch(() => false);
  if (stackReady) {
    console.log('▶ observability stack already up — skipping docker compose up');
  } else {
    console.log('▶ starting observability stack…');
    await execa('docker', [...COMPOSE, 'up', '-d'], { stdio: 'inherit' });
  }
  console.log('▶ waiting for backends…');
  await waitFor(() => checkUrl('Prometheus', `${URLS.prometheus}/-/ready`), {
    timeout: 90_000, label: 'Prometheus',
  });
  await waitFor(() => checkUrl('Grafana', `${URLS.grafana}/api/health`), {
    timeout: 90_000, label: 'Grafana',
  });
  await waitFor(() => checkUrl('Tempo', `${URLS.tempo}/ready`), {
    timeout: 90_000, label: 'Tempo',
  });
  await waitFor(() => checkUrl('Loki', `${URLS.loki}/ready`), {
    timeout: 90_000, label: 'Loki',
  });
  // /health returns { status: 'ok', service: 'expressjs-with-postgresql' } —
  // checking service name lets us distinguish our app from other processes on :3000.
  const checkOurApp = () =>
    checkUrl('app /health', `${URLS.app}/health`, async (res) => {
      const json = (await res.json().catch(() => ({}))) as { service?: string };
      if (json.service !== SERVICE_NAME) {
        console.log(
          `    app /health: service="${json.service ?? '(missing)'}", expected="${SERVICE_NAME}"`,
        );
        return false;
      }
      return true;
    });
  const appIsOurs = stackReady && (await checkOurApp());
  if (appIsOurs) {
    console.log('▶ app already running and healthy — skipping npm start');
  } else {
    const portTaken = await fetch(URLS.app).then(() => true).catch(() => false);
    if (portTaken) {
      // Port is occupied but not by our app — stale process or unrelated service.
      console.error(
        '\n  ✗ Port 3000 is occupied by a process that is not this app.\n' +
          '  It may be a stale version of this app (connected to a wiped database)\n' +
          '  or a completely different service.\n' +
          '  Fix: run  npm run verify:observability:cleanup  to kill it and restart,\n' +
          '  or manually free port 3000.\n',
      );
      process.exit(1);
    }
    console.log('▶ starting app…');
    app = execa('npm', ['start'], { stdio: 'inherit' });
    await waitFor(checkOurApp, { timeout: 60_000, label: 'app /health' });
  }
  console.log('▶ setup complete\n');
});

As you see, nothing fancy, the cleanup is even simpler

after(async () => {
  if (app) {
    console.log('\n▶ stopping app…');
    app.kill('SIGTERM');
    await app.catch(() => {});
    console.log('▶ app stopped');
  }
  if (CLEANUP_AFTER) {
    console.log('▶ tearing down stack (down -v)…');
    await execa('docker', [...COMPOSE, 'down', '-v', '--remove-orphans'], {
      stdio: 'inherit',
    });
    console.log('▶ stack torn down');
  } else {
    console.log('▶ stack is still running');
    console.log('▶ to clean up: npm run verify:observability:cleanup');
  }
});

Having that we can run tests:

test('successful command returns x-trace-id header', async () => {
  const res = await fetchWithDiag('POST add product', CART_ENDPOINT, {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: ADD_PRODUCT_BODY,
  });
  assert.equal(res.status, 204, `Expected 204 — body logged above`);
  const header = res.headers.get('x-trace-id');
  if (!header) {
    console.error(
      '  ✗ x-trace-id missing — verify the wrapper app in src/index.ts ' +
        'adds it via @opentelemetry/api before mounting the emmett app',
    );
  }
  assert.ok(header, 'x-trace-id header missing');
  assert.match(header, /^[0-9a-f]{32}$/, `"${header}" is not a 32-hex trace ID`);
  traceId = header;
  console.log(`  trace ID: ${traceId}`);
});
test('OTel collector exposes Emmett metrics on port 8889', async () => {
  // Send a few more requests so metrics are definitely recorded.
  for (let i = 0; i < 5; i++) {
    await fetch(CART_ENDPOINT, {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: ADD_PRODUCT_BODY,
    });
  }
  try {
    await waitFor(
      async () => {
        let text: string;
        try {
          const res = await fetch(URLS.otelCollectorMetrics);
          text = await res.text();
        } catch {
          console.log('    collector :8889: connection refused');
          return false;
        }
        const emmettLines = text.split('\n').filter((l) => l.startsWith('emmett_') && !l.startsWith('#'));
        if (emmettLines.length === 0) {
          const allFamilies = [...new Set(text.split('\n').filter((l) => !l.startsWith('#') && l).map((l) => l.split('{')[0]))].slice(0, 5);
          console.log(`    collector :8889: no emmett_* metrics yet. Present: ${allFamilies.join(', ') || '(none)'}`);
          return false;
        }
        return true;
      },
      { timeout: 90_000, interval: 5_000, label: 'emmett metrics on collector :8889' },
    );
  } catch (err) {
    await diagnoseCollector();
    await diagnoseDockerLogs('otel-collector');
    throw err;
  }
});

I put it into a single file that can be run as a regular Node.js script.

It already showed me (and Claude) that what they initially did wasn’t working if you try to run it multiple times. It also showed that doing a full cleanup and rebuild, and making it reproducible, needs more work.

Is it done? Not yet; it takes too much time and resources to run it continuously throughout the pipeline. The code is a bit messy, so it needs to be organised. It’s segmented into blocks, includes basic automation and tests, and has already gone through some failures to get it done.

Could I do it better? Sure, and I will improve it, but that’s not the point. I wanted to show you my findings during weekend vibing (without beer tho), the real, not polished iteration, before I run the next one.

The main idea behind OODA loops is not to be perfect, but to iterate quickly, gather feedback as soon as possible, learn from it, develop another theory, and verify it through action.

It’s not about vibing, but it’s also not about analysis paralysis.

I hope you’re now better equipped to think about when vibing — with beer or without, with LLMs or without — actually helps, and when it doesn’t.

Vibe coding is just high-frequency steering. It only works if you have a Harness: a mechanical way to observe and orient, so you don’t steer the whole project into a wall.

Act takes seconds now. Observe takes as long as it always did. Without a harness, you’re not going faster; you’re just making more stuff you haven’t checked.

Harness is not magic, a new discipline, or the next buzzword; I hope I showed you that a bit in this article on what it may look like.

So iterate fast, but wisely remembering to do the full loop. It’s great that LLMs can help us make Acting faster, but we should not skip other steps. We should aim for a fast feedback loop to iterate in the right direction and achieve continuous improvement, to deliver proper value.

Just like Vibing isn’t new, we shouldn’t abandon old engineering practices. We should also not replace collaboration with solitary self-high fives.

Check also:

  • Emmett Pull Request with mentioned changes
  • Interactive Rubber Ducking with GenAI
  • The End of Coding? Wrong Question
  • A few tricks on how to set up related Docker images with docker-compose
  • Docker Compose Profiles, one the most useful and underrated features

Cheers!

Oskar