ManageMyHealth breach: Patients at risk of identity theft, extortion - experts

Mike's Notes

There is never an excuse for these security breaches. Sloppy work created the vulnerability.

"Private health records, linked to the Manage My Health ransomware attack, appear to have already surfaced on the dark web, revealing patients’ most delicate medical details online.

Screenshots seen by The Post appear to show about 30 patient files, seemingly from multiple individuals, including intimate details of a 2018 head injury, a July 2025 vaginal swab, and a December 2025 heart attack." - Stuff

 NZ uses an opt‑in model.

  • Health information is governed by the Health Information Privacy Code 2020, which requires explicit patient consent for new forms of access.
  • Identity verification is required before enabling online access, which naturally fits an opt‑in workflow

NZ portal vendors include

  • ConnectMed
  • Health 365
  • ManageMyHealth
  • MyIndici
  • Vensa
  • MediMap

Updates

  • Second health provider, Canopy Health, hit in major cyber attack - RNZ
  • Patient data changed as major NZ health app MediMap hacked - RNZ

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

25/02/2026

ManageMyHealth breach: Patients at risk of identity theft, extortion - experts

By: Ruth Hill
RNZ: 5/01/2026

Reporter

What I cover: I mainly cover health stories. It's a critically important subject that touches all of our lives, and the goal of my coverage is to shine a light on inequities where they exist and examine some of the complexities in a way that makes sense. At their heart, health stories are always about people, and I love giving a voice to patients and the incredible people working in frontline health services.

My background: I joined RNZ in 2009 as a senior reporter, based in the Wellington newsroom.

Contact: If you have an idea for a story or feedback for me, feel free to get in touch - ruth.hill@rnz.co.nz.

This ransom post screenshot is from a popular hacking forum. Photo: Supplied

  • Hackers say ManageMyHealth ransomware attack about 'business'
  • Company has until Tuesday morning to pay up or 400,000 patient documents released
  • Cyber security experts fear some patients at risk of blackmail or identity theft
  • Patient health portal criticised for sluggish response

Thousands of patients caught up in the ManageMyHealth ransomware attack could be at risk of identity theft or extortion, cyber security experts are warning.

The hackers, calling themselves "Kazu", posted on Sunday morning that unless the company paid a ransom within 48 hours, they would leak more than 400,000 files in their possession.

In a post on Telegram, the group purporting to be behind the breach said it had brought forward the deadline from 15 January in part because ManageMyHealth had responded faster than expected, but mainly to "put pressure on the company".

"Their ignorance of our emails and messages, along with their failure to acknowledge users or explain exactly what happened, is the main issue. Many MMH users have been asking the company for an explanation, but they've either ignored them or responded with vague statements."

This deadline escalation statement was shared in a Telegram channel run by Kazu regarding the ManageMyHealth data breach.This deadline escalation statement was shared in a Telegram channel run by Kazu. Photo: Supplied

Kazu said it had opted for a low-ball ransom demand of $60,000 "to protect the data and quickly close the deal".

"But it seems the company doesn't care about their users' data."

The hackers indicated they were prepared to leak the "valuable" data just to make a point.

"We know exactly how valuable health data is and how sensitive it can be.

"Even if the company doesn't pay the ransom, we can still find buyers for this data.

"To prove our claims and increase the chances of successful deals in the future, we decided to leak the data for free if they don't pay the ransom."

Kazu said they were "not a hacktivist group with political motives".

"We're doing this as a business. Our main goal is money and building a good reputation in the community."

The hackers claimed to have successfully extracted ransom money from many healthcare companies in Asia and Africa over the last two months.

"Once the company pays, we send them a copy of the data, delete it from our servers and never post anything related to the company again."

Patients at risk

Samples for potential "buyers" included clinical notes, lab results, vaccination records, medical photographs and personal identification details, including names, birth dates, addresses, emails and phone numbers.

IT consultant and Hornby community board member Cody Cooper was signed up to ManageMyHealth through his GP.

"My clinic has got 20,000 patients so there's a real push for online. It's seen as convenient, but patients don't have a lot of choice."

He went online to verify the veracity of the claims and was horrified by what he found.

"There's people's passports, there's people's ADHD documents from a psychiatrist, there's pictures of people unclothed. It's very personal data. And my concern as a patient would be, will someone blackmail people? Or try to extort them personally as well, if they don't pay up?"

From what had been made visible so far, it did not appear the data had been encrypted, Cooper said.

"You can infer this fairly safely because resetting passwords doesn't cause users to 'lose' their stored documents. If the data had been encrypted properly with keys tied to credentials, access would break when credentials change."

He also questioned why ManageMyHealth took so long to respond.

"The hack was published around 10pm on 29 December, the MMH website notice appeared on the afternoon of 31 December, but the site wasn't taken offline until that evening."

Furthermore, the company was taking too long to inform affected clinics and patients, he said.

"It should have been able to determine the extent of the breach relatively quickly. The fact that, days later there is no clear confirmation about what was accessed or copied is worrying."

However, there was no guarantee that giving in to the hackers' demands would solve the problem for MMH, he said.

"They may still release the data anyway, they may still contact people, we have no way of knowing if they will honour it.

"Furthermore, if that person is from a country with sanctions, there are laws and treaties that forbid that payment from being made legally as well."

Patients were just collateral damage, he said.

"I will personally probably look to close my account. I can't really have confidence in the system after this. Hopefully my clinic will find a solution that's better."

'Big wakeup call' - Health Minister

The Health Minister said the cyber breach of the country's largest patient information portal was a "big wakeup call".

Simeon Brown told Morning Report he was incredibly concerned.

"It's a deeply serious situation," he said.

"I've been briefed a number of times by health officials who are working very closely with ManageMyHealth in regard to the notification process."

He said ManageMyHealth was also working with the Privacy Commissioner and the National Cyber Security Centre, who were providing them with advice around the notification process.

Brown said his expectation was that they do it as quickly as possible, but they also had to do it accurately as well, and in compliance with the Privacy Act.

"There's a number of processes they have to go through. My expectation is that they do that as quickly as possible so that patients who have had data breached are aware of that and of what data has been breached," he said.

Brown said the advice he's received was that the cyber hackers had only released a very small portion of data as part of their attempt in order to receive a ransom payment.

There was a forensic process underway at the moment to go through and identify who's been impacted and then the process of notification, which is what Manage My Health was doing, he said.

Brown said the group were using hacked information in order to receive a financial reward, but they did not know where they were operating from.

"The reality is that here is a big wakeup call in terms of the protection of private health data and their need for that to be held in the most secure form possible so that patients can have confidence in how it is being used," he said.

Hackers building their 'brand'

Data journalist Keith Ng said the hackers appeared to be using ManageMyHealth to leverage a bigger payout from one of their other targets: Saudi Icon Ransom.

"They're implying they've got their hands full and don't want to be distracted by small fry here, that's their explanation for wanting this over quickly - and if they don't get their ransom they will release data for free."

For Kazu, it was an exercise in brand management.

"They want to establish themselves as a 'trustworthy' ransomware group. By that they mean 'If you pay us, we'll delete the data and you'll never hear from us again. If you don't pay us, bad things will happen to you'.

"So they want to build up their business and use the New Zealand dataset to make an example out of, so people will take them more seriously in the future."

Unfortunately, the ManageMyHealth breach was unlikely to be the result of a sophisticated hacking operation, Ng said.

"This is probably a couple of days work for a couple of people. It's not like an elite hacking crew, it's about volume and they want to make sure they've got targets on the hook all the time.

"They poke around and try to find common vulnerabilities, flaws, they're really looking for low hanging fruit - and if they don't find it, they move on quickly to the next target."

Over and above the technical question of which part of ManageMyHealth's system was not secure, the more important question was what processes it had in place, whether it was having regular independent security audits and taking action to fix the problems identified, he said.

"A business that sets itself up as a health information management system has a lot of incentive to do things right because when they fail, really catastrophic things like this happen, and it is an existential risk for them.

"So we should expect better from these businesses and the fact they let this one slip past them, they should be held accountable."

In its public statements, ManageMyHealth appeared to be trying to minimise the scale of the problem, Ng said.

"They're saying only 7 percent of users were affected, but 7 percent of 1.8 million is quite a big number. The other thing they've said is 'only one component' of the site is affected, not the core database. But it's the kind of things in there - medical photos, test results - which make it so sensitive and damaging for people who are affected.

"It's probably the worst data breach that I recall seeing in New Zealand so far."

Aura Information Security's Patrick Sharp said medical records were hugely valuable to criminals.

The Medibank ransomware attack in Australia in 2022 resulted in many thousands - "maybe even hundreds of thousands" of real financial crimes, he said.

"It's quite likely that the 126,000 or so people affected - depending on the kind of information involved - may suffer at the hands of criminal gangs, lots of scams, blackmail, those kind of things."

ManageMyHealth has been approached for comment.

The GraphBLAS

Mike's Notes

 Alex found this. It might be useful. For future reference.

The original webpage has the links, and the graphic looks better.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

05/01/2026

The GraphBLAS

By: 
GraphBlas: 5/01/2026

.

The GraphBLAS Forum is an open effort to define standard building blocks for graph algorithms in the language of linear algebra.

An example graph and adjacency matrix

We believe that the state of the art in constructing a large collection of graph algorithms in terms of linear algebraic operations is mature enough to support the emergence of a standard set of primitive building blocks. We believe that it is critical to move quickly and define such a standard, thereby freeing up researchers to innovate and diversify at the level of higher level algorithms and graph analytics applications. This effort was inspired by the Basic Linear Algebra Subprograms (BLAS) of dense Linear Algebra, and hence our working name for this standard is “the GraphBLAS”.

A key insight behind this work is that when a graph is represented by a sparse incidence or adjacency matrix, sparse matrix-vector multiplication is a step of breadth first search. By generalizing the pair of scalar operations involved in the linear algebra computations to define a semiring, we can extend the range of these primitives to support a wide range of parallel graph algorithms.

More information

  • The GraphBLAS Wikipedia Page
  • The C reference implementation is SuiteSparse:GraphBLAS, which implements the version 2.1.0 (final) C API.
  • Our 2013 manifesto for this project can be found here.
  • The mathematical definition of the GraphBLAS can be found here.
  • Background information about graphs in the language of linear algebra can be found in the book: Graph Algorithms in the Language of Linear Algebra, edited by J. Kepner and J. Gilbert, SIAM, 2011.
  • The Mathematics of Big Data by J. Kepner and H. Jananthan is the first book to present the common mathematical foundations of big data analysis across a range of applications and technologies.
  • A straw man proposal for the GraphBLAS can be found here
  • Gabor Szarnyas maintains a list of GraphBLAS pointers with lots of tutorial material.

Application Program Interface (API)

Current versions

  • GraphBLAS C API, version 2.0.0 (November 15, 2021)
  • GraphBLAS C API, version 2.1.0 (December 22, 2023)

Legacy versions

Version 1.0 (provisional) of the C language API was released on May 29, 2017 at the GABB workshop here. Version 1.1.0 (provisional) released on November 14, 2017. Version 1.2.0 was released on May 18, 2018. Version 1.3.0 was released on September 25, 2019.

Projects developing implementations of the GraphBLAS

  • SuiteSparse GraphBLAS (Texas A&M)
  • IBM GraphBLAS
  • GraphBLAS Template Library, GBTL (CMU-SEI/Indiana/PNNL)
  • GraphBLAST (UC Davis and LBNL)
  • MPI/C++ Combinatorial BLAS (CombBLAS)
  • Java Graphulo
  • Matlab/Octave D4M
  • GraphPad (Intel)

Programming Language Interfaces to The GraphBLAS API

  • MATLAB (comes with SuiteSparse). MATLAB R2021a and later uses SuiteSparse:GraphBLAS v3.1 for C=A*B when A and B are sparse. Release Notes, under Performance.
  • forGraphBLASGo - Go binding for SuiteSparse:GraphBLAS
  • pygraphblas Python library
  • python-graphblas Python library
  • pggraphblas Postgres extension
  • Julia library

Graph analysis systems that integrate GraphBLAS

  • FalkorDB - a queryable Property Graph database - formerly RedisGraph

Workshops and conferences featuring the GraphBLAS (reverse chronological)

  • Graphs, Architectures, Programming, and Learning (GrAPL) @IPDPS
  • High Performance Extreme Computing (HPEC)
  • GraphChallenge.org
  • SIAM CSE’21 GraphBLAS Minisymposium Session 1
  • SIAM CSE’21 GraphBLAS Minisymposium Session 2
  • SIAM CSE’21 GraphBLAS Tutorial Session 1
  • SIAM CSE’21 GraphBLAS Tutorial Session 2
  • HPEC 2020
  • HPEC 2019
  • HPEC 2018
  • GABB 2018 @IPDPS
  • HPEC 2017
  • GABB 2017 @IPDPS
  • HPEC 2016
  • GABB 2016 @IPDPS
  • HPEC 2015
  • GABB 2015 @IPDPS
  • HPEC 2014
  • GABB 2014 @IPDPS
  • HPEC 2013

Videos and other interesting discussions on GraphBLAS

  • GraphBLAS Forum Update at SC’25 (November 19, 2025)
  • GraphBLAS Forum Update at SC’22 (November 15, 2022)
  • Graph Analytics, by Tim Mattson and Henry Gabb, Intel
  • Short video description of GraphBLAS
  • HPEC’20 presentation on GraphBLAS in Python and MATLAB
  • Presentation at UT Austin
  • A good discussion thread on dropping explicit zeros (fall 2019)
  • A YouTube video playlist on GraphBLAS topics

GraphBLAS mailing list

If you wish to join our effort (or just watch it), please send an email message to our mailing list coordinator.

Steering Committee (alphabetical)

  • David Bader (NJIT)
  • Aydin Buluc (Berkeley Lab)
  • John Gilbert (UC Santa Barbara)
  • Jeremy Kepner (MIT Lincoln Laboratory Supercomputing Center)
  • Tim Mattson (Intel)
  • Henning Meyerhenke (KIT)

The GraphBLAS is supported by the following organizations


The GraphBLAS logo is licensed under CC BY 4.0 (designer: Jakab Rokob)

Vector Database vs Graph Database: Key Technical Differences

Mike's Notes

Pipi 9 uses graph databases and, so far, has no need for vector databases.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library >
  • Home > Handbook > 

Last Updated

04/01/2026

Vector Database vs Graph Database: Key Technical Differences

By: Guy Korland
FalkorDB: 27/10/2024

Guy Korland serves as CEO at FalkorDB, where he drives graph database architecture for generative AI and retrieval-augmented generation workflows. He holds a PhD in Computer Science from Tel Aviv University and brings over 20 years of experience in database engineering. He previously led Redis’ incubation arm as SVP & CTO, oversaw platform architecture as GM & CTO at Stor.ai (Self-Point), co-founded and served as CTO of Shopetti, and directed R&D as VP at GigaSpaces.

Unstructured data is all the data that isn’t organized in a predefined format but is stored in its native form. Due to this lack of organization, it becomes more challenging to sort, extract, and analyze. More than 80% of all enterprise data is unstructured, and this number is growing.

This type of data comes from various sources such as emails, social media, customer reviews, support queries, or product descriptions, which businesses seek to extract meaningful insights from. The rapid growth of unstructured data presents both a challenge and an opportunity for businesses.

To extract insights from unstructured data, the modern approach involves leveraging large language models (LLMs) along with one of two powerful database systems for efficient data retrieval: vector databases or graph databases. These systems, combined with LLMs, enable organizations to structure, search, and analyze unstructured data. 

Understanding the difference between the two is crucial for developers looking to build modern AI applications or architectures like Retrieval-Augmented Generation (RAG). 

In this article, we dive deep into the concepts of vector databases and graph databases, exploring the key differences between them. We also examine their technical advantages, limitations, and use cases to help you make an informed decision when selecting your technology stack.

What is a Vector Database?

Vector databases excel at handling numerical representations of unstructured data — called embeddings — which are generated by machine learning models known as embedding models, unlike traditional databases that focus on structured data like rows and columns. These embeddings capture the semantic meaning (or, features) of the underlying data. Vector databases store, index, and retrieve data that has been transformed into these high-dimensional vectors or embeddings. 

You can convert any type of unstructured or higher-dimensional data into a vector embedding – text, image, audio, or even protein sequences – and this makes vector databases extremely flexible. When this data is converted into vector embeddings, the data points that are similar to each other are embedded closer in the embedding space. This allows for similarity (or, dissimilarity) searches, where you can find similar data using their corresponding vector representations. 

In that sense, vector databases are search engines designed to efficiently search through the higher dimensional vector space. 

For example, in a word embedding space, words with similar meanings or those that are often used in similar contexts would be closer together. The words “cat” and “kitten” would likely be near each other, while “automobile” would be farther away. In contrast, “automobile” might be close to words like “car” and “vehicle”.

The vector representation of these words might look like this:

"cat": [0.43, -0.22, 0.75, 0.12, ...]
"kitten": [0.41, -0.21, 0.76, 0.13, ...]
"automobile": [0.01, 0.62, -0.33, 0.94, ...]
"car": [0.02, 0.60, -0.30, 0.91, ...]

In this context, the vector representations of the words “cat” and “kitten” are closer to each other in the vector space due to their semantic similarity, while “automobile” and “car” would be farther from them but positioned closer to each other.

illustration of a vector representations of words

How does this help build retrieval systems in LLM-powered applications?

An example is a Vector RAG system, where a user’s query is first converted into a vector and then compared against the vector embeddings in the database of existing data. The vectors closest to the query vector are retrieved through a similarity search algorithm, along with the data they represent. This result data is then presented to the LLM to generate a response for the user.

Vector databases are valuable because they help uncover patterns and relationships between high-dimensional data points. 

However, they have a significant limitation: interpretability. The high-dimensional nature of vector spaces makes them difficult to visualize and understand. As a result, when a vector search yields incorrect or suboptimal results, it becomes challenging to diagnose and troubleshoot the underlying issues.

What is a Graph Database?

Graph databases work fundamentally differently from vector databases. 

Rather than using numerical embeddings to represent data, graph databases rely on knowledge graphs to capture the relationships between entities. 

In a knowledge graph, nodes represent entities, and edges represent the relationships between them. This structure allows for complex queries about relationships and connections, which is invaluable when the links between entities are as important as the entities themselves.

In the context of our earlier example involving “cat,” “kitten,” “automobile,” and “car,” each of these concepts would be stored as nodes in a knowledge graph. The relationship between “cat” and “kitten” (e.g., “is a type of”) would be represented as an edge connecting those two nodes. Similarly, “automobile” and “car” might have an edge representing a “synonym” relationship. This would capture the “subject”-“object”-“predicate” triples that form the backbone of knowledge graphs.

Nodes: "cat", "kitten", "automobile", "car"
Edges:
(kitten) -[: IS_A]-> (cat)
(automobile) -[: SYNONYM]-> (car)

Graph databases are ideal when your data contains a high degree of interconnectivity and where understanding these relationships is key to answering business questions. Also, unlike vector databases, knowledge graphs stored in a graph database can be easily visualized. This allows you to explore intricate relationships within your data. 

Modern graph databases support a query language known as Cypher, which allows you to query the knowledge graph and retrieve results. Let’s look at how Cypher works using the example of a slightly more complex knowledge graph.

knowledge graph flowchart of Barcelona FC and La Liga

To create the graph shown in the above image, you will need to construct the nodes and relationships that represent the different entities and their connections. You can use a graph database like FalkorDB to test the queries below. 

Here’s how we create the nodes:

// Creating Player nodes
CREATE (:PLAYER {name: 'Pedri'}), (:PLAYER {name: 'Lamine Yamal'});

// Creating Manager node
CREATE (:MANAGER {name: 'Hansi Flick'});

// Creating Team node
CREATE (:TEAM {name: 'Barcelona'});

// Creating League node
CREATE (:LEAGUE {name: 'La Liga'});

// Creating Country node
CREATE (:COUNTRY {name: 'Spain'});

// Creating Stadium node
CREATE (:STADIUM {name: 'Camp Nou'});

You can now create the relationships using Cypher in the following way: 

// Players play for a team
MATCH (p:PLAYER {name: 'Lamine Yamal'}), (t:TEAM {name: 'Barcelona'})
CREATE (p)-[:PLAYS_FOR]->(t);
MATCH (p:PLAYER {name: 'Pedri'}), (t:TEAM {name: 'Barcelona'})
CREATE (p)-[:PLAYS_FOR]->(t);

// Manager manages a team
MATCH (m:MANAGER {name: 'Hansi Flick'}), (t:TEAM {name: 'Barcelona'})
CREATE (m)-[:MANAGES]->(t);

// Team plays in a league
MATCH (t:TEAM {name: 'Barcelona'}), (l:LEAGUE {name: 'La Liga'})
CREATE (t)-[:PLAYS_IN]->(l);

// Team is based in a country
MATCH (t:TEAM {name: 'Barcelona'}), (c:COUNTRY {name: 'Spain'})
CREATE (t)-[:BASED_IN]->(c);

// Players have nationality
MATCH (p:PLAYER {name: 'Lamine Yamal'}), (c:COUNTRY {name: 'Spain'})
CREATE (p)-[:NATIONALITY]->(c);
MATCH (p:PLAYER {name: 'Pedri'}), (c:COUNTRY {name: 'Spain'})
CREATE (p)-[:NATIONALITY]->(c);

// Team's home stadium
MATCH (t:TEAM {name: 'Barcelona'}), (s:STADIUM {name: 'Camp Nou'})
CREATE (t)-[:HOME_STADIUM]->(s);

As you can see, Cypher queries are easily readable and self-explanatory. You can query the graph using the following example, where we search for players who play for Barcelona, along with their nationalities.

MATCH (p:PLAYER)-[:PLAYS_FOR]->(t:TEAM {name: 'Barcelona'})-[:BASED_IN]->(c:COUNTRY)
RETURN p.name AS Player, c.name AS Nationality;

Here’s the example output you will get: 

Player Nationality
Lamine Yamal Spain
Pedri Spain

Graph databases are purpose-built to efficiently store, query, and navigate complex knowledge graphs. Designed for handling large-scale knowledge graphs, they offer advanced search and querying capabilities. 

These databases are especially effective for applications requiring deep relationship analysis, such as GraphRAG systems, where knowledge graphs can be integrated with LLMs.

Key Differences between Vector Database and Graph Database

As we saw above, vector databases are optimized for similarity searches across high-dimensional data using vector embeddings generated by machine learning models. In contrast, graph databases are designed to model relationships between entities, making them ideal for tasks that require analyzing and understanding the connections between data points.

Here is a detailed breakdown of the key differences:

Feature Vector Database Graph Database
Data Model Represents data as vectors in a high-dimensional space. Represents data points as nodes (entities) connected by edges (relationships).
Query Capabilities Efficiently handles similarity search based on vector representations Effective for navigating & managing relationships. Involves graph traversal, subgraph matching, and shortest-path algorithms.
Performance Considerations Well-suited for large-scale, real-time similarity searches. Optimized for graph-based operations, such as network analysis and graph traversals.
Scalability Can scale horizontally to handle massive datasets and high-throughput queries. Scales with the number of data points. Can scale both horizontally and vertically to accommodate large graphs and complex queries. As it doesn’t have any schema, data can be easily added and modified. Scales with complexity and relationships of added data.
Indexing Vector databases rely heavily on ANN search for grouping the closest data points. Graph databases may use a combination of inverted indexes and graph-specific methods like adjacency matrix or GraphBLAS.


Key Similarities between Vector Database and Graph Database

Despite their differences in data representation and use cases, vector databases and graph databases share several core similarities, especially in how they support modern AI-driven applications and handle complex datasets. 

Both systems are designed to go beyond traditional relational databases, allowing developers to extract deeper insights from more complex and often unstructured data.

Here is a breakdown of their similarities.

Feature Vector Database Graph Database
Advanced Querying Capabilities Enables similarity search via Approximate Nearest Neighbor (ANN) algorithms. Allows relationship-based queries using traversal algorithms.
Handling Complex and Large Datasets Designed for large, high-dimensional datasets like embeddings. Optimized for complex, highly interconnected datasets with numerous relationships.
Optimized for Modern AI Applications Frequently used in AI/ML applications such as recommendation systems, semantic search, etc. Ideal for applications requiring knowledge representation.
Support for Low-Latency Queries Provides low-latency similarity search using efficient ANN algorithms. Optimized for real-time graph traversals and querying of relationships between entities.
Powering Recommendation and Search Systems Powers similarity-based recommendations and semantic search. Powers relationship-based recommendations and complex search queries.
Integration with AI Models Seamlessly integrates with AI models (e.g., LLMs) to transform data into vector embeddings, and also convert user queries into vectors for similarity search.

Seamlessly integrates with LLMs to transform data into knowledge graphs during ingestion, and convert natural language queries to Cypher during retrieval.


Vector Database vs Graph Database: Use Cases

When choosing between vector databases and graph databases, the decision largely depends on the nature of your data and the types of queries you need to perform. Below are key use cases for both, along with specific examples illustrating their advantages across various fields.

Fraud Detection

Graph Databases:

  • Graph databases are highly effective in fraud detection due to their ability to model complex relationships between entities such as users, transactions, accounts, and devices.
  • In financial systems, fraud often occurs within networks of interactions, where suspicious behavior is revealed through unusual patterns.
  • A graph database can analyze these relationships to identify potential fraud by traversing the network and detecting anomalies, such as unusual fund transfers or connections between seemingly unrelated accounts.
  • For instance, a query might explore the paths between accounts to uncover suspiciously interconnected transactions indicative of a money laundering scheme.

Vector Databases:

  • While vector databases are less commonly used for direct fraud detection, they can contribute by detecting anomalous behavior based on historical data patterns.
  • By embedding user behavior (e.g., browsing history, transaction patterns) as vectors, vector databases can identify instances where behavior deviates significantly from typical patterns through dissimilarity search. These deviations might suggest fraud and prompt further investigation.

Scientific Research

Graph Databases:

  • In scientific research, graph databases are invaluable for modeling complex systems where relationships between entities are critical.
  • For example, in biological research, entities like proteins, genes, and diseases are represented as nodes, while interactions between them (e.g., protein-protein interactions) are represented as edges.
  • Researchers can use graph traversal algorithms to uncover hidden connections between diseases and genetic markers, leading to new insights in genomics and drug discovery.
  • Knowledge graphs are also used in academic networks to trace citations and collaborations between researchers, identifying influential papers or emerging trends in a field.

Vector Databases:

  • Vector databases can also be applied in scientific research, particularly in fields like bioinformatics, where high-dimensional data such as DNA sequences or protein structures are common.
  • By converting these biological structures into vector embeddings, researchers can perform similarity searches to identify patterns in large datasets.
  • For instance, vector databases can be used to compare protein structures, searching for similar sequences across vast biological datasets to identify evolutionary relationships or potential drug targets.

eCommerce

Graph Databases

  • In ecommerce, graph databases are highly effective for recommendation systems and customer journey analysis.
  • By modeling the relationships between customers, products, and transactions, ecommerce platforms can generate personalized recommendations by traversing the graph to find connections between users with similar purchasing histories or interests.
  • Additionally, graph databases can track inventory, supplier relationships, and logistics, optimizing the entire supply chain by analyzing relationships across the network.

Vector Databases

  • Vector databases enhance ecommerce applications by enabling personalized recommendations based on user behavior and product similarities.
  • By converting user interactions (e.g., clicks, purchases) and product descriptions into vector embeddings, ecommerce platforms can use vector databases to identify products similar to those users have interacted with.
  • This technique is widely used in product recommendation engines, where users are presented with items similar to their previous searches or purchases, boosting engagement and conversion rates.

Media and Entertainment

Graph Databases

  • The media and entertainment industry benefits from graph databases by modeling content recommendation networks and social relationships.
  • For example, streaming platforms like Netflix and Spotify use graph databases to map user preferences, social connections, and content relationships (e.g., actors, genres, directors).
  • These platforms can then traverse the graph to recommend new movies or songs based on the preferences of similar users or related content. Additionally, graph databases can manage complex relationships between media assets (e.g., episodes, seasons, franchises) and their metadata.

Vector Databases

  • In media and entertainment, vector databases enable content-based search and recommendation systems by using vector embeddings for media content.
  • For instance, a vector database can store embeddings of movies, TV shows, or songs, capturing their semantic features.
  • Users can search for media by uploading images, audio, or even descriptions, and the vector database will return content that is semantically similar.
  • In applications like music discovery, vector databases help recommend songs with similar audio features, while in video search, they enable finding visually similar content based on user preferences or searches.

structure of a knowledge graph

How to Choose between Vector Database and Graph Database

Choosing between a vector database and a graph database depends on several key factors, including the nature of your data, your application’s requirements, and how you intend to query and use the data. 

Below are the most important considerations to guide your decision-making process:

Understand Your Data

The first step in choosing between a vector or graph database is understanding the type of data you are working with.

  • Vector Database: If your data is high-dimensional, such as images, multilingual text, audio, or video, then a vector database is a better fit. For instance, if you are working with embeddings from image recognition models, a vector database allows you to store these vectors and efficiently perform similarity searches between them.
  • Graph Database: If your data is knowledge-oriented and the relationships between entities are of primary importance, then a graph database is the right choice. For example, if you are modeling social networks, supply chains, or recommendation systems, where the relationships between entities (nodes) drive your queries and insights, graph databases are optimized for these scenarios.

Performance and Scalability Needs

Both vector and graph databases are designed to scale, but they have to be managed differently as the dataset grows.

Vector Database: Vector databases excel in low-latency searches even with millions of vectors. Techniques like Approximate Nearest Neighbor (ANN) algorithms ensure that similarity searches can be performed in near real-time, making them ideal for large-scale AI/ML applications. If your application requires fast retrieval of items based on vector similarity, and you expect the dataset to grow continuously, vector databases are optimized for this.

Graph Database: Graph databases, while scalable, face more challenges with performance as the graph becomes more interconnected and deeper. If your application requires complex, multi-hop queries across deeply connected data, you will need to ensure your graph database can handle the load. However, for applications that involve exploring relationships (e.g., shortest paths, friend-of-a-friend queries), graph databases offer performance advantages over relational models. Be mindful that as the graph grows, advanced partitioning and optimization strategies may be needed to maintain performance. In such scenarios, you should consider a graph database known for its low latency and scalability.

Evaluate the Specific Advantages of Each Technology

Weigh the advantages and trade-offs of each database based on the technical requirements of your application.

Graph Database: If relationship analysis and graph traversal are core to your application, then graph databases are unmatched in their ability to model and query complex, interrelated data. The flexibility to modify schema on the fly and the power to model rich, interconnected data make graph databases the best choice for knowledge-centric applications.

Vector Database: Offers clear advantages for AI-powered applications that rely on embeddings. However, they lack interpretability and are not ideal for applications that require understanding relationships between data points.

An Integrated Solution with FalkorDB

FalkorDB is a low-latency graph database graph with select vector capabilities. It offers high-speed performance for both graph traversals and vector similarity searches. 

Some key features of FalkorDB include:

  1. Integrated Data Management: FalkorDB’s unified structure allows for concurrent storage and querying of graph relationships and vector embeddings. This integration eliminates the need for multiple specialized databases, simplifying data architectures.
  2. Advanced Query Processing: The system employs algorithms to optimize queries that involve both graph connections and vector similarities.
  3. Robust Scalability: FalkorDB maintains rapid response times even as data volumes expand, making it suitable for evolving data needs and streaming data.
  4. Streamlined Operations: By combining graph and vector functionalities, FalkorDB reduces the complexity associated with managing and synchronizing separate database systems.

This approach offers a compelling solution for organizations seeking to leverage both semantic relationships and vector-based similarity in their data operations, all within a single, powerful platform.

Knowledge Graph Ecosystem

Additionally, FalkorDB comes with an ecosystem of tools that simplify the process of building applications that derive insights from unstructured data. Here are some: 

GraphRAG-SDK

  • This SDK is designed to simplify the creation of Graph Retrieval-Augmented Generation (GraphRAG) systems. It integrates with FalkorDB and LLMs like OpenAI’s GPT and Google’s Gemini. It enables developers to build knowledge graphs from unstructured data and query them using LLM-generated Cypher queries.
  • The SDK is particularly useful for building AI systems that require reasoning over complex data relationships, such as in finance, legal, or healthcare domains.

FalkorDB-Browser

  • This tool is a visualization interface for exploring and managing graph data stored in FalkorDB. It allows users to interactively navigate through nodes and edges, facilitating data exploration in large knowledge graphs.
  • The browser is ideal for users who need to visually understand the structure of their data or monitor real-time changes in a dynamic graph system​

FalkorDB CodeGraph

  • This tool transforms a codebase into a knowledge graph that visualizes relationships between different code entities like classes, functions, and variables.
  • By analyzing the structure of the code, developers can gain insights into dependencies, detect bottlenecks, and optimize software projects.

Knowledge Graph Ecosystem

Based on the detailed walkthrough above, you now have a comprehensive understanding of vector databases and graph databases. This knowledge equips you to choose the most suitable database type for your project, depending on your specific data structures and query requirements. 

To get started, here are the links to the documentation, cloud platform, and community channels of FalkorDB.

How Science Actually Uses Machine Learning in 2025

Mike's Notes

A great summary of ML resources other than Generative AI. It was published in the Turing Post newsletter. The Turing Post is worth subscribing to.

Resources

References

  • Reference

Repository

  • Home > Ajabbi Research > Library > Subscriptions > Turing Post
  • Home > Handbook > 

Last Updated

213/01/2026

How Science Actually Uses Machine Learning in 2025

By: Marktechpost
Turing Post: 18/12/2025

In this guest post, Marktechpost shares findings from the ML Global Impact Report 2025, an interactive research landscape now available at airesearchtrends.com. Based on an analysis of 5,000+ papers from Nature journals, the report shows that while the world obsesses over generative AI, classical methods still dominate real scientific research.

In fact, 77% of machine learning applications in science rely on traditional techniques like Random Forest, XGBoost, and CatBoost, not transformers or diffusion models. The gap between AI headlines and what actually runs in labs is much larger than most people realize.

State of current media

If you open any tech newsletter or media house, you’ll be convinced that science runs on GPT-4, diffusion models, and whatever architecture dropped last week.

What we found looks very different.

When we analyzed over 5,000 scientific articles published in Nature family journals between January and September 2025, a different picture emerged entirely.

Classical ML methods, like Random Forest, support vector machines (SVMs), and Scikit-learn workflows, account for 47% of all ML use cases in scientific research.

Moreover, adding established ensemble methods like XGBoost, LightGBM, and CatBoost, and traditional approaches, represents a dominant 77% of how scientists actually deploy ML models today.

But for neural networks and deep learning, it’s roughly 23%.

This has nothing to do with slow adoption but rather how research teams put forward their research.

In practice, researchers prioritize methods they can clearly justify during peer review. Novelty matters, but only after reliability, interpretability, and dataset sanity checks are satisfied. A Random Forest that reliably predicts protein binding affinity beats a transformer that might work better but introduces unpredictable failure modes.

When a paper depends on reproducibility and collaborators need to understand every step, novelty quickly becomes a liability. And in that context, reliability matters more than novelty.

For practitioners, the risk is obvious. Optimizing your stack around what trends on AI Twitter actively pulls you away from the methods that dominate real scientific workflows.

But who creates the tools versus who uses them?

The most consequential imbalance in ML research is who builds the tools versus who publishes with them. Nearly 90% of open-source ML tools referenced in 2025 scientific literature originate from the United States, where the foundational frameworks are powering imaging, genomics, and environmental science worldwide.

Yet the heaviest users aren’t American. China accounts for 43% of all ML-enabled papers in our dataset, compared to just 18% from the United States.

India holds third place and is climbing fast.

This imbalance shapes who sets research agendas and who executes them. Essentially, American institutions and companies build the infrastructure, while Chinese researchers publish the papers.

That said, non-US contributions to the tooling ecosystem are significant and growing. For instance:

  • Scikit-learn emerged from France.
  • U-Net, the workhorse of medical image segmentation, came from Germany.
  • CatBoost was developed in Russia by Yandex.
  • Canada produced foundational GAN and RNN architectures.

The US still dominates this layer of the stack, but it is not the only contributor.

What does this mean for the AI research community? The countries that define the frameworks shape how problems get formulated. When you inherit someone else’s tools, you often inherit their assumptions about what problems matter and how they should be solved.

Applied sciences and health research lead adoption

ML hasn’t penetrated all scientific fields equally. The heaviest adoption appears in applied sciences and health research, which are disciplines where prediction, classification, and pattern recognition map cleanly onto existing workflows.

Consider the specific applications we observed across thousands of papers:

  • In medical sciences, ML enables early-detection imaging, precision diagnostics, genomic sequence mapping, and mutation tracking. Cancer screening pipelines now routinely incorporate ensemble classifiers. Biomarker discovery increasingly depends on feature extraction from high-dimensional datasets.
  • In physical and materials sciences, researchers apply ML to advanced robotics, materials engineering, and complex physical simulations. Nanotechnology labs use neural networks to predict material properties before synthesis, saving months of experimental iteration.
  • In environmental sciences, large-scale Earth-observation analytics rely on ML for everything from monsoon prediction to disaster preparedness. Climate models incorporate machine learning components for pattern recognition in satellite imagery. Remote sensing has been transformed.
  • In agriculture, crop-yield forecasting and food systems resilience modeling now depend on ML pipelines trained on decades of historical data.

In most papers, ML appears as one step in an experimental pipeline, not the object of study.

In most domains, ML is not the focus of the paper itself. Instead, researchers publish papers about cancer or climate or crop yields that happen to use machine learning as methodology, showcasing that ML has transitioned from the research subject to the research infrastructure used in downstream problem-solving.

Collaborative research is predominant!

Contributions from the research community and our analysis make it clear that the lone genius scientist is a myth.

Most ML-enabled papers in our dataset include 2-15 institutional affiliations. Major studies rarely originate from a single organization.

But collaboration patterns differ dramatically by region.

To put this into figures:

  • American papers average 4.1 organizations per study, reflecting a highly distributed ecosystem spanning universities, hospitals, national laboratories, and private research institutions. Harvard Medical School leads in ML-enabled research volume, but the broader US landscape is remarkably decentralized.
  • Chinese papers average 2.6 organizations per study, indicating a more concentrated, high-throughput model where fewer institutions produce more output. China achieves its 43% global share through density rather than breadth, with 72.8 articles per university compared to 39.6 in the US.
  • Lastly, new collaboration corridors are emerging. India, Saudi Arabia, and the United States increasingly partner on applied sciences, materials research, engineering, and computer vision.

For AI practitioners building research relationships, these patterns matter. A collaboration strategy that works in the American context (distributed, multi-institutional) may need adjustment for Chinese partnerships (concentrated, institution-centric).

India’s quiet rise to third place

The report’s third-largest contributor doesn’t get enough attention: India has achieved a steep upward trajectory in ML-driven scientific output, now ranking behind only China and the United States.

What makes India’s approach distinctive is its focus on practical, scalable innovation rather than hype-driven experimentation. Indian ML research concentrates on socially relevant applications, like health diagnostics, climate resilience, fraud detection, agricultural optimization, and materials science, with local relevance.

Participation is broadening beyond elite institutions. Tier 1 and Tier 2 universities now contribute meaningfully to India’s ML research ecosystem, and a rapidly growing startup landscape is translating academic research into applied research and output.

For organizations looking to build global research partnerships, India represents an increasingly attractive option: strong technical talent, growing institutional capacity, and a pragmatic orientation toward problems that matter.

What this means for your ML strategy?

The gap discussed above mostly creates confusion, especially for practitioners who mistake visibility for impact.

If you are a researcher, don’t let trend-chasing distract from methodology that works. If Random Forest solves your problem, publish with confidence. The Nature portfolio clearly agrees.

If you are a practitioner, the tools that matter for production systems, like ensemble methods, classical feature engineering, and interpretable models, remain the tools that matter for scientific research. Investment in boring ML capabilities pays compounding returns.

If you are an ML organization, Geographic asymmetries in tool development versus research output suggest strategic opportunities. Countries consuming US-built frameworks may eventually develop alternatives.

If you are an investor or policymaker, the classical-vs-generative distribution should inform research priorities. Generative AI captures headlines, but traditional ML delivers the bulk of scientific impact.

ML is the infra layer of modern science

The takeaway is straightforward. Machine learning has stopped being the subject of research and has become the machinery behind it.

Scientists don’t write papers about Random Forest any more than they write papers about healthcare, for instance. They write papers about what Random Forest helped them discover.

As Asif Razzaq, Editor and Co-Founder of Marktechpost, put it: 

“This report shows that machine learning isn’t just reshaping AI, it’s reshaping science itself. The real story here is not hype, but impact, which tells that ML is now a fundamental instrument of modern scientific work.”

The 77% figure should humble anyone convinced that transformers and diffusion models define the field.

The geographic asymmetries should interest anyone thinking about how power flows through technical ecosystems. And the collaboration patterns should inform anyone building research networks in an increasingly multipolar scientific landscape.

The full interactive data is available at airesearchtrends.com, where you can explore country-level breakdowns, disciplinary distributions, and tool-by-tool analysis across 125+ countries and 5,000+ papers.

The next time someone tells you AI research is all about the latest foundation model, you’ll have 5,000 data points suggesting otherwise.

Thanks for reading!

This post was written by the Marktechpost team, specifically for Turing Post. We thank Marktechpost for sharing the ML Global Impact Report 2025 findings and supporting Turing Post’s mission to cut through the froth.