Oliver Sacks on the Three Essential Elements of Creativity

Mike's Notes

Very cool. Maria Popova's original post includes many more reference links. Everyone has creativity, some more than others, and it can be cultivated.

I love Oliver Sacks' colour-highlighted notebooks.

The River of Consciousness compiles the following essays:

  • Darwin and the Meaning of Flowers
  • Speed
  • Sentience: The Mental Lives of Plants and Worms
  • The Other Road: Freud as a Neurologist
  • The Fallibility of Memory
  • Mishearings
  • The Creative Self
  • A General Feeling of Disorder
  • The River of Consciousness
  • Scotoma: Forgetting and Neglect in Science

Resources

References

  • The River of Consciousness, Oliver Sacks, October 2017. Pan Macmillan.

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

05/09/2026

Oliver Sacks on the Three Essential Elements of Creativity

By: Maria Popova
The Marginalian: 09/11/2017

Maria Popova is a Bulgarian-born, American-based essayist, book author, poet, and writer of literary and arts commentary and cultural criticism that has found wide appeal both for her writing and for the visual stylistics that accompany it.

...

“And don’t ever imitate anybody,” Hemingway cautioned in his advice to aspiring writers. But in this particular sentiment, the otherwise insightful Nobel laureate seems to have been blind to his own admonition against the dangers of ego, for only the ego can blind an artist to the recognition that all creative work begins with imitation before fermenting into originality under the dual forces of time and consecrating effort.

Imitation, besides being the seedbed of empathy and our experience of time, is also, paradoxically enough, the seedbed of creativity — not only a poetic truth but a cognitive fact, as the late, great neurologist and poet of science Oliver Sacks (July 9, 1933–August 30, 2015) argues in a spectacular essay titled “The Creative Self,” published in the posthumous treasure The River of Consciousness (public library).


Oliver Sacks captures a thought in his journal at Amsterdam’s busy train station (Photograph by Lowell Handler from On the Move)

In his impressive handwritten notes on creativity and the brain, which became the basis of the essay, Sacks had enthused about — in two colors, underlined — the “buzzing, blooming chaos” of the mind engaged in creative work. But, contrary to the archetypal myth of the lone genius struck with a sudden Eureka! moment, this chaos doesn’t occur in a vacuum. Rather, it coalesces from a particulate cloud of influences and inspirations without which creativity — that is, birthing of something meaningful that hadn’t exist before — cannot come about.

With the illustrative example of Susan Sontag — herself a writer of abiding wisdom on the art of storytelling — Sacks traces the inevitable trajectory of creative development from imitation to originality:

Susan Sontag, at a conference in 2002, spoke about how reading opened up the entire world to her when she was quite young, enlarging her imagination and memory far beyond the bounds of her actual, immediate personal experience. She recalled,

When I was five or six, I read Eve Curie’s biography of her mother. I read comic books, dictionaries, and encyclopedias indiscriminately, and with great pleasure…. It felt like the more I took in, the stronger I was, the bigger the world got…. I think I was, from the very beginning, an incredibly gifted student, an incredibly gifted learner, a champion child autodidact…. Is that creative? No, it wasn’t creative…[but] it didn’t preclude becoming creative later on…. I was engorging rather than making. I was a mental traveler, a mental glutton…. My childhood, apart from my wretched actual life, was just a career in ecstasy.

[…]

I started writing when I was about seven. I started a newspaper when I was eight, which I filled with stories and poems and plays and articles, and which I used to sell to the neighbors for five cents. I’m sure it was quite banal and conventional, and simply made up of things, influenced by things, I was reading…. Of course there were models, there was a pantheon of these people…. If I was reading the stories of Poe, then I would write a Poe-like story…. When I was ten, a long-forgotten play by Karel Čapek, R.U.R., about robots, fell into my hands, so I wrote a play about robots. But it was absolutely derivative. Whatever I saw I loved, and whatever I loved I wanted to imitate — that’s not necessarily the royal road to real innovation or creativity; neither, as I saw it, does it preclude it…. I started to be a real writer at thirteen.

Sontag’s experience, Sacks argues, reflects the common pattern in the natural cycle of creative evolution — we learn our own minds by finding out what we love; these models integrate into a sensibility; out of that sensibility arises the initial impulse for imitation, which, aided by the gradual acquisition of technical mastery, eventually ripens into original creation. He writes:

If imitation plays a central role in the performing arts, where incessant practice, repetition, and rehearsal are essential, it is equally important in painting or composing or writing, for example. All young artists seek models in their apprentice years, models whose style, technical mastery, and innovations can teach them. Young painters may haunt the galleries of the Met or the Louvre; young composers may go to concerts or study scores. All art, in this sense, starts out as “derivative,” highly influenced by, if not a direct imitation or paraphrase of, the admired and emulated models.

When Alexander Pope was thirteen years old, he asked William Walsh, an older poet whom he admired, for advice. Walsh’s advice was that Pope should be “correct.” Pope took this to mean that he should first gain a mastery of poetic forms and techniques. To this end, in his “Imitations of English Poets,” Pope began by imitating Walsh, then Cowley, the Earl of Rochester, and more major figures like Chaucer and Spenser, as well as writing “Paraphrases,” as he called them, of Latin poets. By seventeen, he had mastered the heroic couplet and began to write his “Pastorals” and other poems, where he developed and honed his own style but contented himself with the most insipid or clichéd themes. It was only once he had established full mastery of his style and form that he started to charge it with the exquisite and sometimes terrifying products of his own imagination. For most artists, perhaps, these stages or processes overlap a good deal, but imitation and mastery of form or skills must come before major creativity.


A page from Dr. Sacks’s wild and wondrous handwritten notes on creativity and the brain.

Curiously, Sacks points out, many creators don’t make the leap from mastery to such “major creativity” — something Schopenhauer considered in his incisive distinction between talent and genius. Often, creators — be they artists or scientists — content themselves with reaching a level of mastery, then remaining at that plateau for the rest of their careers, comfortably creating more of what they already know well how to create. Sacks examines what set those who soar apart from those who plateau:

Why is it that of every hundred gifted young musicians who study at Juilliard or every hundred brilliant young scientists who go to work in major labs under illustrious mentors, only a handful will write memorable musical compositions or make scientific discoveries of major importance? Are the majority, despite their gifts, lacking in some further creative spark? Are they missing characteristics other than creativity that may be essential for creative achievement — such as boldness, confidence, independence of mind?

It takes a special energy, over and above one’s creative potential, a special audacity or subversiveness, to strike out in a new direction once one is settled. It is a gamble as all creative projects must be, for the new direction may not turn out to be productive at all.

Much of the gamble, Sacks argues, is a kind of patient gestation at the unconscious level — something Einstein touched upon in explaining how his mind worked. Echoing T.S. Eliot’s insistence on the necessity of “a long incubation” in creative work, Sacks adds:

Creativity involves not only years of conscious preparation and training but unconscious preparation as well. This incubation period is essential to allow the subconscious assimilation and incorporation of one’s influences and sources, to reorganize and synthesize them into something of one’s own…. The essential element in these realms of retaining and appropriating versus assimilating and incorporating is one of depth, of meaning, of active and personal involvement.


Illustration by Maurice Sendak from Open House for Butterflies by Ruth Krauss

He illustrates the detrimental absence of such a gestational period with an example from his own experience:

Early in 1982, I received an unexpected packet from London containing a letter from Harold Pinter and the manuscript of a new play, A Kind of Alaska, which, he said, had been inspired by a case history of mine in Awakenings. In his letter, Pinter said that he had read my book when it originally came out in 1973 and had immediately wondered about the problems presented by a dramatic adaptation of this. But, seeing no ready solution to these problems, he had then forgotten about it. One morning eight years later, Pinter wrote, he had awoken with the first image and first words (“Something is happening”) clear and pressing in his mind. The play had then “written itself” in the days and weeks that followed.

I could not help contrasting this with a play (inspired by the same case history) which I had been sent four years earlier, where the author, in an accompanying letter, said that he had read Awakenings two months before and been so “influenced,” so possessed, by it that he felt impelled to write a play straightaway. Whereas I loved Pinter’s play — not least because it effected so profound a transformation, a “Pinterization” of my own themes — I felt the 1978 play to be grossly derivative, for it lifted, sometimes, whole sentences from my own book without transforming them in the least. It seemed to me less an original play than a plagiarism or a parody (yet there was no doubting the author’s “obsession” or good faith).

In a testament to his uncommon empathic might and his endearing generosity of interpretation in regarding others, Sacks reflects on the deeper phenomena at play:

I was not sure what to make of this. Was the author too lazy, or too lacking in talent or originality, to make the needed transformation of my work? Or was the problem essentially one of incubation, that he had not allowed himself enough time for the experience of reading Awakenings to sink in? Nor had he allowed himself, as Pinter did, time to forget it, to let it fall into his unconscious, where it might link with other experiences and thoughts.

The unfortunate playwright seems to have embodied the lamentation which poet Mary Oliver so beautifully articulated in her meditation on the creative life: “The most regretful people on earth are those who felt the call to creative work, who felt their own creative power restive and uprising, and gave to it neither power nor time.”

Sacks points to three essential elements in a creative breakthrough, be it a great play or a deep mathematical insights: time, “forgetting,” and incubation. More than a century after Mark Twain declared that “substantially all ideas are second-hand, consciously and unconsciously drawn from a million outside sources,” Sacks — who had previously written at length about our unconscious borrowings — adds:

All of us, to some extent, borrow from others, from the culture around us. Ideas are in the air, and we may appropriate, often without realizing, the phrases and language of the times. We borrow language itself; we did not invent it. We found it, we grew up into it, though we may use it, interpret it, in very individual ways. What is at issue is not the fact of “borrowing” or “imitating,” of being “derivative,” being “influenced,” but what one does with what is borrowed or imitated or derived; how deeply one assimilates it, takes it into oneself, compounds it with one’s own experiences and thoughts and feelings, places it in relation to oneself, and expresses it in a new way, one’s own.

...

Complement this fathom of The River of Consciousness, thoroughly resplendent in its totality, with physicist and poet Alan Lightman on the psychology of creative breakthrough in art and science, then revisit Bill Hayes’s loving remembrance of Oliver Sacks and Sacks himself on what the poet Thom Gunn taught him about creativity.

Notes on USB drive formatting

Mike's Notes

Taking Pipi security seriously. A collection of notes copied from here and there about formatting USB drives. The sources are listed in the resources.

In the meantime, measures include air-gapping and turning off all wifi.

Later, we will employ data diodes and many other measures.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

04/09/2026

Notes on USB drive formatting

By: Mike Peters & Google
On a Sandy Beach: 01/01/2026

Mike is the inventor and architect of Pipi and the founder of Ajabbi.

Time required to format a USB drive

The time it takes to fully format a USB drive depends entirely on the drive's total storage size and the speed of the USB port, not the volume of files currently stored on it. A 64GB USB drive will take the exact same amount of time to fully format whether it is completely empty or completely full. Because a full format writes zeros to every gigabyte of storage, speed is limited by the USB drive's write performance.

  • 16 GB Drive: Takes roughly 2 to 5 minutes.
  • 32 GB Drive: Takes roughly 5 to 10 minutes.
  • 64 GB Drive: Takes roughly 10 to 20 minutes.
  • 128 GB Drive: Takes roughly 20 to 45 minutes.
  • 256 GB+ Drive: Can take over an hour.

Two factors determine where the drive falls on that time estimate:

  • USB Generation: A USB 3.0 or 3.2 drive plugged into a matching blue or Type-C port will format significantly faster than an older USB 2.0 drive, which is capped at very slow data transfer speeds.
  • Hardware Quality: Cheap, promotional USB drives use low-grade flash memory with incredibly slow write speeds, meaning they can take two to three times longer to format than high-quality name-brand drives.

What is a BadUSB Attack and How to Prevent It?

BadUSB, as the name suggests, is a crafty cybersecurity attack that acts as a puppeteer, controlling your USB devices at will. A BadUSB attack occurs when a USB device has a built-in firmware vulnerability that lets it disguise itself as a human interface device. Once connected to its target computer, a BadUSB could then discreetly execute harmful commands or inject malicious payloads.

A common type of BadUSB attack involves a MalDuino device. It uses a programmable USB device that mimics a keyboard when plugged into a system. This device can be pre-configured to automatically inject numerous malicious keystrokes into an unsuspecting user’s computer, enabling attackers to execute commands and compromise the system within seconds.

Within organisations, a preventative measure such as USB blocking software is a necessity because BadUSB attacks, if undetected or unstopped, could result in the unauthorised execution of commands that instigate security bypass incidents, privilege escalation, DDoS attacks, or malware infections of the host computers, which could then spread to target entire networks.

A BadUSB attack is incredibly dangerous because the malware does not live in the storage partition where your files are kept; it lives inside the USB controller chip's firmware.

When you perform a standard Windows format (even a full format), the computer only interacts with the flash memory storage blocks. It completely ignores the controller chip that dictates how the USB device talks to the computer.

How a BadUSB Attack Works.

  • The Disguise: The compromised firmware tricks your computer into thinking the USB drive is not a storage device at all, but rather a USB keyboard or network card.
  • The Execution: The moment you plug it in, the chip sends rapid, invisible keystrokes to your computer.
  • The Result: It can open your command prompt, download malware from the internet, and compromise your system in under three seconds—all before you even have a chance to open File Explorer or click "Format."

What are the types of BadUSB?

The different types of BadUSB found commercially available are:

  • MalDuino
  • WiFi-enabled BadUsb
  • BadUsb Cables

MalDuino

MalDuino is an open-source Arduino-based BadUSB that injects malicious payloads into a target computer. Gaining widespread attention recently, MalDuino packs several more features than regular BadUSB devices thanks to its onboard computer.

MalDuino devices commonly found today support Micro SD cards and include a set of DIP switches that let users toggle between stored programs on the card.

WiFi-enabled BadUsb

This type of BadUSB is similar to MalDuino in that an Arduino board serves as the base for the device but is specially designed with WiFi capabilities. Once plugged into a target system, these devices allow attackers to introduce malicious payloads into a victim's computer using the WiFi protocol.

WiFi-enabled BadUSB can take different forms depending on the exact purpose and function they serve. The common iterations of this device used today are as follows.

  • WiFi-enabled keystroke injectors
  • WiFi keyloggers
  • WiFi deauthers

WiFi-enabled keystroke injectors

These are the most common types of WiFi BadUSB found today. When plugged into a target computer, these devices remain dormant until an attacker makes further contact through a smartphone or a neighbouring system. Connecting to this BadUSB is as simple as connecting to a WiFi access point.

Once connected, the attacker can inject keystrokes using a suitable scripting language. These devices often come with their own applications that allow hackers to execute scripts remotely.

WiFi keyloggers

WiFi keyloggers are modern hacker hardware that particularly targets desktop computers. They work as a bridge between the keyboard's USB terminal and the computer itself. This BadUSB intercepts input signals from the keyboard and relays them to the hacker's computer.

WiFi keyloggers capture everything the user types, including sensitive information and passwords, without raising suspicion. On desktop computers, these devices can be completely hidden from plain sight and do not affect keyboard performance while in use.

WiFi deauthers

WiFi Deauthers are malicious devices that leverage flaws present in the WiFi protocol to force all users of a WiFi network to disconnect automatically. Subsequently, WiFi Deauthers prevent users from reconnecting as long as the device is active.

WiFi deauthers differ from other forms of BadUSB in that they do not directly affect the systems they connect to; instead, they use them as a power source to disrupt network connectivity and cause downtime.

BadUSB Cables

Gaining widespread attention recently, BadUSB cables look and function like any other USB cable, but they are secretly malicious devices that inject scripts and malware into a computer without the user's knowledge.

Also known as USB Ninja and USB Harpoon, these generic-looking cables hide a BadUSB within their internal circuitry and are more deceptive than many other variants. A BadUSB cable can support functions such as charging and data transfer while malicious activity happens in the background.

How to Deal with a Suspected BadUSB

If you believe a USB drive has compromised firmware, formatting it from Windows will not make it safe. You have three real options:

  1. Physical Destruction (Safest): Throwing the drive in the trash or physically destroying it is the only 100% reliable fix for everyday users. Flash drives are cheap; your data security is not.
  2. Firmware Flashing: You would need to find the exact manufacturer's tools for that specific controller chip and overwrite the firmware. This is highly technical, risky, and often impossible for generic drives.
  3. Hardware Write-Blockers: IT security professionals use specialised hardware to inspect suspicious drives without allowing data to flow back to the PC.

BadUSB Removal

No software-based open-source security tools can safely "remove" or clean BadUSB malware from a standard flash drive. Because the malicious code is hardcoded directly into the hardware's internal controller chip, a computer's operating system cannot reach, overwrite, or clean it through traditional software tools.

However, the open-source community has developed powerful tools to detect and block these attacks, as well as complex technical frameworks used by hardware reverse engineers to overwrite the controller chip entirely.

Open-Source Tools to Detect & Block Attacks

Rather than fixing the drive, these open-source tools sit on your computer and intercept a BadUSB device the second it tries to emulate a keyboard or inject malicious keystrokes:

  • Anti-BadUSB (Python-based): A popular open-source Python script hosted on GitHub. It continuously monitors keyboard inputs across your operating system. If a newly inserted USB device begins typing commands at superhuman speeds (keystroke injection), it instantly flags and blocks the input before the script can execute.
  • USB Auth Guard: A lightweight, open-source Linux security tool. It locks down your system's USB ports by default. When a new device is plugged in, it forces a security authentication prompt (via polkit) before the operating system can interact with the device. This completely stops human interface device (HID) exploits.
  • Linux Kernel udev Monitoring: Built-in open-source Linux subsystems can be configured to catch BadUSB activity. By opening a terminal and running udevadm monitor, you can see the exact moment a device is plugged in. If a USB storage stick suddenly triggers a kernel event trying to install a "HID Keyboard," you will catch the disguise instantly.

Making a start on Hilbert Spaces and matrix calculations using CPU

Mike's Notes

Working notes on finding a way for Pipi to use genetic algorithms with Hilbert Spaces and matrix calculations. A desktop exercise with Google at the moment. Computational experiments on CPU to come. Some generated Python and CFML sample code below to start playing with. 

It looks promising; it might help address limitations of other methods for using fitness landscapes. The output curves do look better.

This might void the need to use a Cerebras wafer for these fitness landscape calculations.

The next step is to watch free maths lectures on YouTube, then do a lot of drawings and process diagrams. Then play with code and see what happens. 😎

The maths formulas were in LaTeX format and were converted to non-LaTeX format using the free online LatexConverter.

Note

Nothing has been verified or tested. 😎

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

03/09/2026

Making a start on Hilbert Spaces and matrix calculations using CPU

By: Mike Peters & Google Search AI Mode
On a Sandy Beach: 30/08/2026

Mike Peters: Mike is the inventor and architect of Pipi and the founder of Ajabbi.

Google Search AI Mode: Free and surprisingly good.

...

Hilbert Spaces

Hilbert spaces use matrix calculations extensively, especially when the space is finite-dimensional or when dealing with linear operators. Here is how they connect:

  1. Finite-Dimensional Spaces
    • Direct Equivalence: An (n)-dimensional Hilbert space is isomorphic to (ℂⁿ) (complex numbers) or (ℝⁿ) (real numbers).
    • Vectors: Elements in the space are represented as column vectors.
    • Operators: Linear operators acting on the space are represented as (n × n) matrices.
    • Inner Product: The inner product is calculated using matrix multiplication (specifically, the conjugate transpose of a vector multiplied by another vector).
  2. Infinite-Dimensional Spaces
    • Infinite Matrices: In infinite-dimensional Hilbert spaces (like L²), the space of square-integrable functions, operators can be represented as infinite matrices.
    • Basis Dependence: By choosing an orthonormal basis, any linear operator can be written as a matrix with infinitely many rows and columns.
  3. Quantum Mechanics (Matrix Mechanics)
    • State Vectors: Quantum states are vectors in a Hilbert space.
    • Observables: Physical properties (like position, momentum, or spin) are Hermitian operators.
    • Heisenberg's Formulation: Werner Heisenberg's formulation of quantum mechanics is literally called "matrix mechanics" because it uses matrix calculations to solve Hilbert space problems.

In machine learning, Hilbert spaces—specifically Reproducible Kernel Hilbert Spaces (RKHS)—rely heavily on matrix calculations to make complex, non-linear data linearly separable.

Instead of explicitly mapping data points to an infinite-dimensional space, ML algorithms use the Kernel Trick to compute everything via an N × N matrix, where N is the number of data samples.

The Gram (Kernel) Matrix

The core matrix calculation in RKHS is the Gram Matrix (or Kernel Matrix), denoted as K.

  • Formula: (Kᵢⱼ = k(xᵢ, xⱼ) = ⟨ Φ(xᵢ), Φ(xⱼ) ⟩_{H})
  • Meaning: Each entry i,j is the inner product of data points (xᵢ) and (xⱼ) in the Hilbert space (H).
  • Properties: It is always a symmetric and positive semi-definite (PSD) matrix.

Core Matrix Operations in RKHS

Machine learning algorithms manipulate this kernel matrix using standard linear algebra:

  • Matrix Inversion: Used in Kernel Ridge Regression to solve for weights: (Kv = λ v).
  • Eigenvalue Decomposition: Used in Kernel PCA to find non-linear principal components by solving (Kv = λ v).
  • Quadratic Programming: Used in Support Vector Machines (SVMs) to maximise margins using matrix-vector products involving (K).

Computational Bottlenecks

Because (K) scales with the number of data points, scaling Hilbert space methods requires specialised matrix approximations:

  • Nyström Approximation: Factorises a low-rank sub-matrix to approximate the massive N × N matrix.
  • Random Fourier Features: Avoids the exact matrix calculation entirely by mapping data into a randomised, lower-dimensional space.

Graph image

Genetic Algorithms

Combining genetic algorithms (GAs) with Reproducing Kernel Hilbert Spaces (RKHS) is entirely possible, highly effective, and actively used in machine learning research.

In this hybrid framework, the fitness landscape is modelled as a continuous function belonging to an RKHS. This approach addresses a major limitation of traditional GAs: evaluating an individual's true fitness can be computationally expensive (e.g., running a physics simulation or a wet-lab experiment). The RKHS serves as an intelligent surrogate model.

Graph image

Core Mathematical Engines At Work

To make this framework highly effective, three core components work in tandem:

  • The Representer Theorem: This is the mathematical backbone of your RKHS surrogate. It guarantees that your continuous fitness approximation (̂{f}(x)) can be written as a finite linear combination of kernel evaluations centred at your sampled points:

(ᶠ(x)=∑ ᵢ₌₁ᴺα ᵢK(x,xᵢ))

  • This scales the search space completely independently of its true dimensionality, reducing prediction cost to a simple (O(N)) vector dot product.
  • Informed Exploration vs. Exploitation: Instead of letting the GA evaluate individuals blindly on the surrogate, you can use the RKHS variance (uncertainty) to construct an Acquisition Function (like Expected Improvement or Upper Confidence Bound). The GA then maximises this acquisition function rather than the raw estimated fitness, forcing the algorithm to intelligently search unmapped areas of the landscape.
  • Gram Matrix Regularisation: As new true data points are evaluated, they enter the Gram Matrix. To prevent numerical instability or overfitting as (N) grows, a small regularisation ridge ((λ I)) is maintained, smoothly adjusting the landscape's rigidity without rebuilding the framework from scratch.

How the Combination Works

[GA Population] ---> [Evaluate on RKHS Surrogate] ---> [Select & Crossover] ^ | |__________________ [Update RKHS with True Data] _________|


  1. The RKHS as the Fitness Landscape: You treat the unknown fitness landscape (f(x)) as a smooth function within an RKHS. By evaluating a small set of initial points, you use Kernel Ridge Regression or Kriging (Gaussian Process Regression) to build a continuous, analytical approximation of the landscape.
  2. Genetic Search on the Surrogate: The GA searches this approximated RKHS landscape. Because evaluating the RKHS kernel matrix is mathematically cheap (O(N)) for a new prediction, the GA can evolve through thousands of generations in seconds.
  3. Adaptive Sampling (Bayesian Optimisation): The best candidates found by the GA are evaluated using the true, expensive fitness function. These new data points are added back into the Gram matrix, updating the RKHS landscape to make it more accurate where it matters most.

Key Advantages of This Hybrid Approach

  • Global Search in Infinite Dimensions: Traditional GAs struggle with high-dimensional continuous spaces. By mapping the search space into an RKHS via a kernel (like the Gaussian RBF kernel), the GA implicitly searches a highly expressive, infinite-dimensional feature space.
  • Guaranteed Smoothness: Functions in an RKHS are mathematically smooth and bounded [6]. This eliminates extreme, discontinuous noise in the fitness landscape, allowing the GA to find gradient-like paths toward global optima without getting trapped in microscopic local minima.
  • Exact Kernel Alignment: You can use a GA to optimise the hyperparameters of the RKHS kernel itself (e.g., finding the optimal length-scale (σ ) of an RBF kernel) to perfectly align the geometry of the Hilbert space with your actual problem landscape.

Real-World Use Cases

  • Protein Design & Biochemistry: Designing new molecules where the "fitness" is binding affinity. The RKHS models the chemical landscape based on known proteins, and the GA mutates amino acid sequences to find optimal binders.
  • Hyperparameter Optimisation: Optimising complex neural network architectures where the evaluation landscape is highly non-linear and expensive to compute.
  • Robotics Control: Evolving locomotion parameters for robotics where physical testing is slow or dangerous.

Implementation Challenges

  • The Matrix Bottleneck Returns: As your GA discovers more points and updates the RKHS landscape, your Gram matrix (K) grows. Inverting this matrix to update the landscape scales at O(N³). You will need to use Nyström approximations or sparse online kernel methods if your total evaluation points exceed 10,000.
  • Exploration vs. Exploitation: GAs are aggressive exploiters. If your RKHS surrogate model is inaccurate in an unexplored region, the GA might get stuck optimising a "fake" peak. You must include an uncertainty metric (like the variance in a Gaussian Process) to force the GA to explore unknown areas.

Kernel Ridge Regression or Kriging (Gaussian Process Regression)

Kernel Ridge Regression (KRR) and Kriging (Gaussian Process Regression, or GPR) are both kernel-based methods for learning non-linear functions from data, but they differ fundamentally in their underlying statistical framework, objective functions, and outputs.

Key Differences

  • Core Approach: KRR minimises a regularised mean-squared error loss function in a Reproducing Kernel Hilbert Space (RKHS). GPR uses a probabilistic (Bayesian) approach, defining a Gaussian process prior over functions and updating it with a likelihood function based on observed data.
  • Uncertainty Estimation: KRR outputs point predictions only. GPR (and Kriging) naturally quantifies uncertainty, providing full posterior distributions, variance estimates, and confidence intervals.
  • Hyperparameter Optimisation: KRR typically optimises kernel parameters using grid search with cross-validation on a loss function. GPR optimises hyperparameters via gradient ascent on the marginal likelihood.
  • Terminology & Origin: Kriging originated in geostatistics (mining) to find the Best Linear Unbiased Predictor (BLUP), whereas GPR stems from machine learning and stochastic processes. Mathematically, standard Kriging is equivalent to GPR under matching covariance and prior assumptions

When to Use Which

  • Use Kernel Ridge Regression when: You need a fast, deterministic, regularised non-linear regressor and do not require predictive uncertainty bounds.
  • Use Kriging / Gaussian Process Regression when: You need confidence intervals for predictions, want to optimise hyperparameters via marginal likelihood, or are modelling spatial/geostatistical data.

Python Code (not tested)

Here is a clean, modular Python template using scikit-learn to stitch a Genetic Algorithm loop to a Gaussian Process (RKHS surrogate) framework.

import numpy as np
from sklearn.gaussian_process import GaussianProcessRegressor
from sklearn.gaussian_process.kernels import Matern

# --- 1. CONFIGURATION & SIMULATION BACKEND ---
BOUNDS = np.array([[-5.0, 5.0], [-5.0, 5.0]])  # 2D search space boundaries
N_DIM = BOUNDS.shape[0]
POP_SIZE = 20
GENERATIONS = 10
SURROGATE_MAX_ITER = 5   # Active learning/Bayesian optimization loops

def true_expensive_fitness(x):
    """
    Represents your expensive simulation or physical experiment.
    Expects a 1D array of shape (N_DIM,). Returns a scalar.
    """
    # Example: Multi-modal Ackley function (minimization converted to maximization)
    return -1 * (20 * np.exp(-0.2 * np.sqrt(0.5 * np.sum(x**2))) + 
                 np.exp(0.5 * np.sum(np.cos(2 * np.pi * x))) - 20 - np.e)

# --- 2. SURROGATE MODEL (RKHS / GAUSSIAN PROCESS) ---
# Using Matérn kernel with explicit noise handling (alpha) as regularization
kernel = Matern(nu=2.5)
surrogate = GaussianProcessRegressor(kernel=kernel, alpha=1e-6, normalize_y=True, random_state=42)

# --- 3. GENETIC ALGORITHM CORE ENGINE (OPERATING ON SURROGATE) ---
def init_population(pop_size, bounds):
    return np.random.uniform(bounds[:, 0], bounds[:, 1], size=(pop_size, len(bounds)))

def crossover(parent1, parent2):
    # Blend crossover (BLX-alpha)
    alpha = 0.5
    gamma = (1 + 2 * alpha) * np.random.random(size=N_DIM) - alpha
    return parent1 + gamma * (parent2 - parent1)

def mutate(individual, bounds, rate=0.2, scale=0.5):
    for i in range(len(individual)):
        if np.random.random() < rate:
            individual[i] += np.random.normal(0, scale)
            individual[i] = np.clip(individual[i], bounds[i, 0], bounds[i, 1])
    return individual

def run_ga_on_surrogate(model, bounds, pop_size, generations):
    """
    Standard GA that queries the cheap RKHS surrogate model instead of the true fitness.
    """
    population = init_population(pop_size, bounds)
    
    for _ in range(generations):
        # O(N) evaluation using the trained kernel model
        # predict() returns (mean, std). We maximize the predicted mean fitness.
        fitnesses, _ = model.predict(population, return_std=True)
        
        # Rank-based selection
        idx = np.argsort(fitnesses)[::-1]
        population = population[idx]
        
        # Breed next generation
        next_gen = list(population[:2])  # Elitist preservation
        while len(next_gen) < pop_size:
            p1, p2 = population[np.random.randint(0, 5)], population[np.random.randint(0, 5)]
            child = crossover(p1, p2)
            child = mutate(child, bounds)
            next_gen.append(child)
            
        population = np.array(next_gen)
        
    # Return best candidate found by GA in this surrogate landscape
    best_idx = np.argmax(model.predict(population, return_std=False))
    return population[best_idx]

# --- 4. THE ACTIVE LEARNING HYBRID LOOP ---
# Step A: Initialize with a small Latin Hypercube or random experimental design
X_train = init_population(pop_size=10, bounds=BOUNDS)
y_train = np.array([true_expensive_fitness(x) for x in X_train])

print("Starting Hybrid GA-RKHS Optimization Engine...\n")

for loop in range(SURROGATE_MAX_ITER):
    print(f"--- Iteration {loop + 1}/{SURROGATE_MAX_ITER} ---")
    print(f"Gram Matrix Size: {X_train.shape[0]} points")
    
    # Step B: Fit/Update RKHS surrogate model
    surrogate.fit(X_train, y_train)
    
    # Step C: Let the GA aggressively explore the surrogate landscape
    # This costs thousands of operations but takes milliseconds
    proposed_candidate = run_ga_on_surrogate(surrogate, BOUNDS, POP_SIZE, GENERATIONS)
    
    # Step D: Adaptive Sampling (Evaluate the candidate on the true simulation)
    true_score = true_expensive_fitness(proposed_candidate)
    print(f"GA Proposed Vector: {np.round(proposed_candidate, 4)}")
    print(f"True Fitness Score: {true_score:.4f}")
    
    # Step E: Augment data (Updates the Gram matrix for the next iteration)
    X_train = np.vstack([X_train, proposed_candidate])
    y_train = np.append(y_train, true_score)
    print(f"Current Best Score Found: {np.max(y_train):.4f}\n")

print("Optimization Complete.")

Enhancements 

  • Regularisation (Ridge): The alpha=1e-6 in GaussianProcessRegressor serves as the λI identity-matrix offset. This prevents numerical inversion crashes if the GA proposes points structurally close to existing data.
  • Normalisation Engine: normalize_y=True transforms unstable fitness functions to a standard normal distribution mapping, stabilising kernel hyperparameter tuning during .fit().
  • Scalability Hand-off: run_ga_on_surrogate can run for massive populations or generations; it never interacts with true_expensive_fitness, protecting your heavy computation budget.

CFML Code (not tested)

<cfscript>
/**
 * CFML Hybrid GA-RKHS Optimization Engine
 */

// --- 1. CONFIGURATION & CONFIG MATRIX ---
variables.bounds = [ {min: -5.0, max: 5.0}, {min: -5.0, max: 5.0} ]; // 2D Search Space
variables.popSize = 20;
variables.generations = 10;
variables.surrogateMaxIter = 5;

/**
 * Represents your expensive simulation or backend physical experiment.
 * Minimisation problem converted to Maximisation.
 */
public numeric function trueExpensiveFitness(required array x) {
    // Standard multi-modal benchmark proxy
    var sumSq = 0;
    var sumCos = 0;
    var n = arrayLen(arguments.x);
    
    for (var i = 1; i <= n; i++) {
        sumSq += arguments.x[i] ^ 2;
        sumCos += cos(2 * pi() * arguments.x[i]);
    }
    
    var term1 = -20.0 * exp(-0.2 * sqrt(0.5 * sumSq));
    var term2 = -exp(0.5 * sumCos);
    return -1 * (term1 + term2 - 20 - exp(1));
}

// --- 2. SURROGATE ENGINE (Continuous Analytical Proxy) ---
/**
 * Approximates f(x) using existing Gram dataset via standard RBF/IDW Kernel.
 */
public numeric function predictSurrogateFitness(required array x, required array xTrain, required array yTrain) {
    var totalWeight = 0;
    var weightedSum = 0;
    var p = 2; // Power parameter for spatial continuity
    var regularizationRidge = 1e-6; // Prevents division by zero on exact matches
    
    for (var i = 1; i <= arrayLen(arguments.xTrain); i++) {
        var distSq = 0;
        for (var d = 1; d <= arrayLen(arguments.x); d++) {
            distSq += (arguments.x[d] - arguments.xTrain[i][d]) ^ 2;
        }
        var dist = sqrt(distSq) + regularizationRidge;
        var weight = 1.0 / (dist ^ p);
        
        totalWeight += weight;
        weightedSum += weight * arguments.yTrain[i];
    }
    
    return weightedSum / totalWeight;
}

// --- 3. GENETIC ALGORITHM CORE ENGINE (OPERATING ON CHEAP SURROGATE) ---
public array function initPopulation(required numeric popSize, required array bounds) {
    var pop = [];
    for (var i = 1; i <= arguments.popSize; i++) {
        var ind = [];
        for (var d = 1; d <= arrayLen(arguments.bounds); d++) {
            arrayAppend(ind, rand() * (arguments.bounds[d].max - arguments.bounds[d].min) + arguments.bounds[d].min);
        }
        arrayAppend(pop, ind);
    }
    return pop;
}

public array function crossover(required array p1, required array p2) {
    var child = [];
    var alpha = 0.5; // Blend Crossover parameter
    for (var d = 1; d <= arrayLen(arguments.p1); d++) {
        var gamma = (1 + 2 * alpha) * rand() - alpha;
        arrayAppend(child, arguments.p1[d] + gamma * (arguments.p2[d] - arguments.p1[d]));
    }
    return child;
}

public array function mutate(required array individual, required array bounds, numeric rate=0.2, numeric scale=0.5) {
    var mutated = duplicate(arguments.individual);
    for (var d = 1; d <= arrayLen(mutated); d++) {
        if (rand() < arguments.rate) {
            // Box-Muller transform for normal distribution mutation
            var u1 = rand(); var u2 = rand();
            if(u1 == 0) u1 = 0.0001;
            var normalRandom = sqrt(-2.0 * log(u1)) * cos(2.0 * pi() * u2);
            
            mutated[d] += normalRandom * arguments.scale;
            // Clip boundaries
            if (mutated[d] < arguments.bounds[d].min) mutated[d] = arguments.bounds[d].min;
            if (mutated[d] > arguments.bounds[d].max) mutated[d] = arguments.bounds[d].max;
        }
    }
    return mutated;
}

public array function runGaOnSurrogate(required array xTrain, required array yTrain, required array bounds, required numeric popSize, required numeric generations) {
    var population = initPopulation(arguments.popSize, arguments.bounds);
    
    for (var gen = 1; gen <= arguments.generations; gen++) {
        // Evaluate complete generation instantly on the analytical surrogate proxy
        var scoredPop = [];
        for (var i = 1; i <= arrayLen(population); i++) {
            arrayAppend(scoredPop, {
                vector: population[i],
                score: predictSurrogateFitness(population[i], arguments.xTrain, arguments.yTrain)
            });
        }
        
        // Rank-based Sort Descending
        arraySort(scoredPop, function(a, b) {
            return b.score > a.score ? 1 : (b.score < a.score ? -1 : 0);
        });
        
        // Rebuild next generation
        var nextGen = [ scoredPop[1].vector, scoredPop[2].vector ]; // Elitist strategy
        while (arrayLen(nextGen) < arguments.popSize) {
            var parent1 = scoredPop[randRange(1, 5)].vector;
            var parent2 = scoredPop[randRange(1, 5)].vector;
            var child = crossover(parent1, parent2);
            child = mutate(child, arguments.bounds);
            arrayAppend(nextGen, child);
        }
        
        // Update population pointer
        for (var k = 1; k <= arrayLen(nextGen); k++) {
            population[k] = nextGen[k];
        }
    }
    
    // Return best vector from final generation iteration
    return population[1];
}

// --- 4. EXECUTING THE HYBRID SEED LOOP ---
writeOutput("Starting CFML Hybrid GA-RKHS Optimization Engine...<br><br>");

// Step A: Initialize Gram Dataset
variables.X_train = initPopulation(10, variables.bounds);
variables.y_train = [];
for (var i = 1; i <= arrayLen(variables.X_train); i++) {
    arrayAppend(variables.y_train, trueExpensiveFitness(variables.X_train[i]));
}

// Step B: Active Learning Global Intercept Loop
for (var loop = 1; loop <= variables.surrogateMaxIter; loop++) {
    writeOutput("--- Iteration #loop#/#variables.surrogateMaxIter# ---<br>");
    writeOutput("Gram Matrix Dataset Size: " & arrayLen(variables.X_train) & " points<br>");
    
    // Step C: Let the GA scan the math surrogate space aggressively
    var proposedCandidate = runGaOnSurrogate(variables.X_train, variables.y_train, variables.bounds, variables.popSize, variables.generations);
    
    // Step D: Run costly physical simulation check on candidate
    var trueScore = trueExpensiveFitness(proposedCandidate);
    
    writeOutput("GA Proposed Vector: [" & arrayToList(proposedCandidate, ", ") & "]<br>");
    writeOutput("True Real-World Fitness Score: " & numberFormat(trueScore, "0.0000") & "<br>");
    
    // Step E: Augment data matrices
    arrayAppend(variables.X_train, proposedCandidate);
    arrayAppend(variables.y_train, trueScore);
    
    writeOutput("Current System Best Score: " & numberFormat(arrayMax(variables.y_train), "0.0000") & "<br><br>");
}

writeOutput("Optimization Execution Cycle Terminated.");
</cfscript>

Key Architectural Shifts in the CFML Variant

  • Script-Only Architecture: Written purely within <cfscript> blocks to remove legacy tag overhead and maximise processing throughput for vector arithmetic loops.
  • Native Box-Muller Generator: Because ColdFusion’s native rand() produces a flat uniform distribution, a standard Box-Muller transformation was written directly into the mutate() method to correctly inject normal (Gaussian) distribution mutations.
  • Surrogate Mapping Array Matrix: Replaces Python arrays with native array structs. The predictSurrogateFitness method models continuous field behaviours natively via distances weighted dynamically across the current known points.

Lec 50 Kernel Ridge Regression (KRR)

NPTEL - Indian Institute of Science, Bengaluru

YouTube: 09/2025 35:13