A key challenge in clinical trial design is benchmarking and identifying similar trials to enhance trial feasibility and likelihood of success. This requires navigating fragmented, heterogeneous data sources: protocols, registries, and internal assets. Manual approaches are slow, inconsistent, and error prone.
What is a clinical trial knowledge graph?
Before we get into the weeds of how a knowledge graph can address this challenge, let’s start with a definition of the term. A knowledge base is a dataset representing real-world facts and semantic relations. Edges represent the relationships between data points, wherein each entity is a node in the graph with certain attributes or characteristics. A knowledge graph is, thus, informed by domain-specific ontologies, which map key attributes of the entities, such as diseases, to standardized codes, such as ICD-10 CM.
For our purposes, the knowledge graph shows the relationships among multiple entities in a clinical trial, such as investigators, hospitals, sites, diseases, procedures, and more. Luca Parisi, Director of Data Science, says the clinical trial itself could be considered an entity and can, thus, be modeled as a digital twin. Next there is a site, such as a hospital that runs the trial, and an investigator who runs or participates in a trial. A site may run multiple trials, and an investigator may be the primary investigator of this specific trial but may also be the secondary or associated investigator of another trial.
Parisi said the knowledge graph can help answer questions such as: Is this investigator working with or likely to work with other investigators, or more likely to succeed if working with other investigators? Have they worked well in the past or could they work well in the future? What were or could be their outcomes?
“The set of weights and the factors that we’ve identified,” Parisi continues, “they are robust enough to give our users a very good starting point for any clinical trial designs, which then can be fine-tuned and tailored to the specific trial design and/or user persona.”
Chief Data Scientist Hassan Malik notes that because the system is compatible with any AI model format, if a new AI model were to be released, this process would simply require very minor fine-tuning to incorporate it to inform the knowledge graph or derive predictive or prescriptive insights from it, not a complete overhaul.
How does clinical trial similarity search work?
Comparing one clinical trial to other trials is a two-step process, Parisi explains. The knowledge graph focuses on a selection of similar trials related to the specific trial. That selection is based on the key attributes or characteristics defining the fingerprint of a trial entity, i.e., uniquely defining and characterizing it and the relationships between similar trials. Therefore, a knowledge graph reveals a population of similar trials, which may include further trials that a user was not expecting and could, thus, complement their trial landscape and inform their prospective trial design.
“We don’t stop there,” Parisi says, noting that a second check involves leveraging natural language processing (NLP) and large language models (LLMs) to determine which trials are more similar or less similar with respect to the specific trial. “We dive deeper into the texts that define those protocols, for example, inclusion criteria, primary endpoints, treatment plans, and more. We can pinpoint exactly what is similar and what is different, to what extent, because of our domain expert-curated categories within our database.”
“It really goes back to the rigor that is already built into the process,” said Malik. Using the analogy of Google, he continued: “Anyone can open Google and perform a search or ask a question, but then it becomes very difficult to know that you have covered everything.”
Trial similarity vs. traditional trial search: what’s the difference?
When asked how this new Trial Similarity component of Trialtrove+ compares to the existing search capabilities of Citeline’s Trialtrove solution, Parisi replied: “It is a more flexible search. In Trialtrove, customers must have a very clear idea of what they’re looking for. They need to do a lot of work beforehand to filter through trials. But what we offer here is a complementary lens, a bit more at the beginning of that process for a broader search on similar trials, not necessarily just those that match specific filters.”
Accuracy and validation: 92% expert alignment
The trial similarity process can be likened to a lab setting with a highly magnified microscope providing minute detail, coupled with the oversight of a lab technician. “Through human evaluation and validation, we achieved 92% alignment with content and clinical domain experts across Citeline and our parent company Norstella,” Parisi says. Parisi explains that the approach was tested against state-of-the-art AI tools and was consistently better by 20%.
“I think that the secret sauce of this is that 8% deviation from content and clinical expert arbitration,” says Skye Hodson, VP of Clinical Solutions. “It’s GXP compliant, interpretable, auditable, near-human levels of performance.”
Surfacing what you didn’t know to look for
“I think the beauty of the knowledge graph is that it captures clinically meaningful relationships, but it can also surface something that the customer may have not thought about before,” Parisi says. “It may surface new trial design indicators that the client may have not considered earlier, but they may be seen as potential epidemiological and/or operational factors substantiating a certain trial design versus another.”
Hodson agrees. “The beauty of the AI-driven knowledge graph approach,” he says, is the ability to surface similarities that aren’t immediately obvious, to unearth hidden relationships, to analytically find those similarities that might surprise you. That’s where you may find competitors or benchmarks you might not have accounted for, trials that have operated in this space that you might not have considered. That is the value.
“It’s like, oh gosh, that’s a trial I would never have considered, and it would have torpedoed my program. Or oh wow, that’s an interesting analogous benchmark for a trial running in an ambulatory care setting I hadn’t considered.
“The benefit is you can find trials that you maybe would never have considered to be similar through this approach. They might be in a similar care setting. They might be having a similar trial design consideration. They might have a similar outcome, a similar schedule of events, some similarity that is not easy to find or immediately obvious. That’s what I find interesting about this technique.”
Why proprietary data is the foundation
Returning to the power of the knowledge graph, Hodson says, “Not everyone can create a knowledge graph of this caliber. It’s an expression of the value of our data, the ontologies we’ve built, the quality of our data, the AI readiness of our data.”
Reiterating Hodson’s point, Parisi says: “Anyone could build a knowledge graph, but nobody can build it as good as we can from a clinical standpoint, because we have the right data foundations. And those right data foundations are Trialtrove, Sitetrove, and Pharmaprojects.”
Luca Parisi, Director of Data Science, Citeline
Skye Hodson, VP of Clinical Solutions, Citeline
Hassan Malik, Chief Data Scientist, Citeline