Blog

Graph Intelligence: The Hidden Power of Clinical Trial Comparisons and Knowledge Discovery

See how AI-powered trial similarity combines knowledge graphs, NLP and protocol data to find meaningful comparisons and support more informed trial design.
July 24, 2026

Share:

A key challenge in clinical trial design is benchmarking and identifying similar trials to enhance trial feasibility and likelihood of success. This requires navigating fragmented, heterogeneous data sources: protocols, registries, and internal assets. Manual approaches are slow, inconsistent, and error prone.

What is a clinical trial knowledge graph?

Before we get into the weeds of how a knowledge graph can address this challenge, let’s start with a definition of the term. A knowledge base is a dataset representing real-world facts and semantic relations. Edges represent the relationships between data points, wherein each entity is a node in the graph with certain attributes or characteristics. A knowledge graph is, thus, informed by domain-specific ontologies, which map key attributes of the entities, such as diseases, to standardized codes, such as ICD-10 CM.

For our purposes, the knowledge graph shows the relationships among multiple entities in a clinical trial, such as investigators, hospitals, sites, diseases, procedures, and more. Luca Parisi, Director of Data Science, says the clinical trial itself could be considered an entity and can, thus, be modeled as a digital twin. Next there is a site, such as a hospital that runs the trial, and an investigator who runs or participates in a trial. A site may run multiple trials, and an investigator may be the primary investigator of this specific trial but may also be the secondary or associated investigator of another trial.

Parisi said the knowledge graph can help answer questions such as: Is this investigator working with or likely to work with other investigators, or more likely to succeed if working with other investigators? Have they worked well in the past or could they work well in the future? What were or could be their outcomes?

“It’s like a map of all these entities tied together by domain-relevant relationships and described or characterized by relevant attributes. It’s important to capture all these relationships and key attributes defining the fingerprint of an entity given a particular use case or set of use cases and user personas considered.” — Luca Parisi, Director of Data Science

“The set of weights and the factors that we’ve identified,” Parisi continues, “they are robust enough to give our users a very good starting point for any clinical trial designs, which then can be fine-tuned and tailored to the specific trial design and/or user persona.”

Chief Data Scientist Hassan Malik notes that because the system is compatible with any AI model format, if a new AI model were to be released, this process would simply require very minor fine-tuning to incorporate it to inform the knowledge graph or derive predictive or prescriptive insights from it, not a complete overhaul.

How does clinical trial similarity search work?

Comparing one clinical trial to other trials is a two-step process, Parisi explains. The knowledge graph focuses on a selection of similar trials related to the specific trial. That selection is based on the key attributes or characteristics defining the fingerprint of a trial entity, i.e., uniquely defining and characterizing it and the relationships between similar trials. Therefore, a knowledge graph reveals a population of similar trials, which may include further trials that a user was not expecting and could, thus, complement their trial landscape and inform their prospective trial design.

“We don’t stop there,” Parisi says, noting that a second check involves leveraging natural language processing (NLP) and large language models (LLMs) to determine which trials are more similar or less similar with respect to the specific trial. “We dive deeper into the texts that define those protocols, for example, inclusion criteria, primary endpoints, treatment plans, and more. We can pinpoint exactly what is similar and what is different, to what extent, because of our domain expert-curated categories within our database.”

“It really goes back to the rigor that is already built into the process,” said Malik. Using the analogy of Google, he continued: “Anyone can open Google and perform a search or ask a question, but then it becomes very difficult to know that you have covered everything.”

Trial similarity vs. traditional trial search: what’s the difference?

When asked how this new Trial Similarity component of Trialtrove+ compares to the existing search capabilities of Citeline’s Trialtrove solution, Parisi replied: “It is a more flexible search. In Trialtrove, customers must have a very clear idea of what they’re looking for. They need to do a lot of work beforehand to filter through trials. But what we offer here is a complementary lens, a bit more at the beginning of that process for a broader search on similar trials, not necessarily just those that match specific filters.”

Accuracy and validation: 92% expert alignment

The trial similarity process can be likened to a lab setting with a highly magnified microscope providing minute detail, coupled with the oversight of a lab technician. “Through human evaluation and validation, we achieved 92% alignment with content and clinical domain experts across Citeline and our parent company Norstella,” Parisi says. Parisi explains that the approach was tested against state-of-the-art AI tools and was consistently better by 20%.

“The secret sauce of this is that 8% deviation from content and clinical expert arbitration. It’s GXP compliant, interpretable, auditable, near-human levels of performance.” — Skye Hodson, VP of Clinical Solutions

“I think that the secret sauce of this is that 8% deviation from content and clinical expert arbitration,” says Skye Hodson, VP of Clinical Solutions. “It’s GXP compliant, interpretable, auditable, near-human levels of performance.”

Surfacing what you didn’t know to look for

“I think the beauty of the knowledge graph is that it captures clinically meaningful relationships, but it can also surface something that the customer may have not thought about before,” Parisi says. “It may surface new trial design indicators that the client may have not considered earlier, but they may be seen as potential epidemiological and/or operational factors substantiating a certain trial design versus another.”

Hodson agrees. “The beauty of the AI-driven knowledge graph approach,” he says, is the ability to surface similarities that aren’t immediately obvious, to unearth hidden relationships, to analytically find those similarities that might surprise you. That’s where you may find competitors or benchmarks you might not have accounted for, trials that have operated in this space that you might not have considered. That is the value.

“It’s like, oh gosh, that’s a trial I would never have considered, and it would have torpedoed my program. Or oh wow, that’s an interesting analogous benchmark for a trial running in an ambulatory care setting I hadn’t considered.

“The benefit is you can find trials that you maybe would never have considered to be similar through this approach. They might be in a similar care setting. They might be having a similar trial design consideration. They might have a similar outcome, a similar schedule of events, some similarity that is not easy to find or immediately obvious. That’s what I find interesting about this technique.”

Why proprietary data is the foundation

Returning to the power of the knowledge graph, Hodson says, “Not everyone can create a knowledge graph of this caliber. It’s an expression of the value of our data, the ontologies we’ve built, the quality of our data, the AI readiness of our data.”

“Anyone could build a knowledge graph, but nobody can build it as good as we can from a clinical standpoint, because we have the right data foundations. And those right data foundations are Trialtrove, Sitetrove, and Pharmaprojects.” — Luca Parisi, Director of Data Science

Reiterating Hodson’s point, Parisi says: “Anyone could build a knowledge graph, but nobody can build it as good as we can from a clinical standpoint, because we have the right data foundations. And those right data foundations are Trialtrove, Sitetrove, and Pharmaprojects.”

Contributors
Luca Parisi, Director of Data Science, Citeline
Skye Hodson, VP of Clinical Solutions, Citeline
Hassan Malik, Chief Data Scientist, Citeline

Recent Insights and News

"Built for Decisions, Not Demos": A Q&A with Kris Kaneta on Norstella Atlas
Norstella Launches Atlas, a Biopharma Agentic Platform that Can 'Train You'
The AI Productivity Paradox: What's Holding Pharma Back from Reaching More Patients?
Norstella Launches Atlas, an Agentic AI Platform Built on Biopharma's Most Connected Data
S2 E1 Busting Real-World Data and AI Myths
How Norstella is embedding artificial intelligence to improve pipeline to patient decision-making
Norstella’s AI bet: Clinical trials are often won or lost before the first patient enrolls

Get started

See what Norstella AI can do for your team

Connect today to speak with an expert!

Frequently asked questions

A clinical trial knowledge graph is a structured data model that maps the relationships between key entities in a trial, including investigators, sites, diseases, procedures, and outcomes. Informed by domain-specific ontologies and standardized codes such as ICD-10 CM, it enables AI to reason across complex clinical relationships rather than matching simple keyword filters.
Trial similarity search uses a two-step process. First, a knowledge graph identifies a population of trials sharing key attributes with the target trial. Second, NLP and LLMs analyse protocol texts, including inclusion criteria, primary endpoints, and treatment plans, to rank those trials by degree of similarity and surface the specific dimensions in which they differ.
Trialtrove+ is Citeline’s AI-enhanced tier of its Trialtrove clinical trial intelligence platform. It extends traditional filter-based search with capabilities such as Trial Similarity, which uses AI knowledge graphs and NLP to surface clinically relevant trials a user may not have considered through conventional search methods.
Norstella’s Trial Similarity approach achieved 92% alignment with content and clinical domain experts across Citeline and Norstella. The methodology was independently benchmarked against state-of-the-art AI tools and outperformed them consistently by 20%. The system is also GXP compliant, interpretable, and auditable.