Guest Column | July 22, 2026

Building The RNA Knowledge Graph: How Better Annotation Is Accelerating Therapeutic Discovery

By Giulia Antonazzo, Ph.D., Research Associate, University of Cambridge

drug delivery, nanotechnology technology-GettyImages-1337413453

The future of RNA therapeutics will depend not only on our ability to design new RNA molecules but also on our ability to understand the complex biological systems that determine how those molecules function.

Over the past decade, RNA biology has expanded far beyond the traditional view of RNA as a messenger between DNA and protein. Researchers have uncovered thousands of regulatory RNA molecules, including microRNAs, long non-coding RNAs (lncRNAs), small interfering RNAs (siRNAs), circular RNAs (circRNAs), and other emerging RNA classes, that influence gene expression through diverse mechanisms.

This expanding landscape has created extraordinary opportunities for therapeutic innovation. RNA-based medicines can now be designed to silence genes, alter splicing, replace missing proteins, modulate immune responses, and potentially regulate disease-associated pathways in ways that were previously inaccessible. However, this progress has also introduced a significant challenge: how do we organize, interpret, and connect the enormous amount of biological information being generated?

As RNA research becomes increasingly data-driven, high-quality biological annotation is becoming an essential foundation for discovery. Without accurate and standardized ways to describe RNA functions and regulatory relationships, even the most advanced computational approaches will struggle to uncover meaningful biological insights.

Our recent work within the Gene Ontology Consortium focused on addressing this challenge by improving how non-coding RNA-mediated regulation of gene expression is represented. By creating a more comprehensive framework for describing these biological processes, we aim to help researchers better interpret complex data sets and accelerate discoveries across molecular biology and medicine.

The Growing Complexity Of RNA Biology

For many years, biological databases and computational frameworks were primarily built around protein-coding genes. While this reflected the focus of molecular biology at the time, it did not fully capture the complexity of gene regulation that has emerged through advances in genomics and transcriptomics.

We now understand that gene expression is controlled by a vast network of regulatory mechanisms involving both coding and non-coding elements. Non-coding RNAs play critical roles throughout these networks. They can regulate transcription, influence RNA stability, control translation, modify chromatin states, and coordinate cellular responses to environmental and developmental signals.

From a therapeutic perspective, these regulatory functions are increasingly important. Many emerging RNA medicines are designed not simply to replace a missing protein but to influence the regulatory systems controlling gene expression. Antisense oligonucleotides, RNA interference technologies, and other RNA-based approaches rely on a detailed understanding of these molecular relationships.

Yet as the number of known RNA regulators continues to increase, accurately representing their biological roles becomes more challenging.

Why Biological Annotation Matters

Biological annotation provides a structured way to describe the functions of genes, RNAs, proteins, and cellular processes. At its core, annotation allows researchers to move beyond individual experiments and connect discoveries across different studies, organisms, and biological contexts. This is particularly important in RNA biology, where the same RNA molecule may participate in multiple regulatory processes depending on the cell type, developmental stage, or disease state.

Without consistent terminology and structured representation, important biological connections can remain difficult to identify. For example, two research groups may study related regulatory mechanisms but describe them using different terminology. Even when the underlying biology is connected, computational systems may not recognize those relationships unless the information is represented in a standardized way. Improved annotation helps overcome this challenge by providing a common language for describing biological knowledge.

Expanding The Gene Ontology For RNA Regulation

The Gene Ontology (GO) provides one of the most widely used frameworks for representing biological knowledge. It enables researchers to describe gene product functions, biological processes, and cellular components in a standardized manner.

As our understanding of non-coding RNA biology advanced, it became clear that GO needed to better represent the mechanisms through which regulatory RNAs influence gene expression. Our work focused on expanding and refining these representations, allowing researchers to more accurately capture processes involving regulatory RNAs, including mechanisms mediated by microRNAs, small interfering RNAs, and other non-coding RNA classes.

This effort is not simply about adding new terminology. It is about creating a framework that reflects the biological complexity uncovered by modern research. By improving how these relationships are represented, we can support better integration of experimental results and enable more powerful computational analyses.

Building A Knowledge Graph For RNA Discovery

One way to think about biological annotation is as the foundation of a growing RNA knowledge graph. A knowledge graph connects biological entities — including genes, RNAs, proteins, pathways, diseases, and cellular processes — through defined relationships. Instead of viewing individual discoveries as isolated pieces of information, knowledge graphs allow researchers to identify connections across large and complex data sets.

For RNA therapeutics, this capability is increasingly valuable. The development of an RNA medicine often requires answering complex questions:

  • Which RNA regulators influence a disease pathway?
  • Which targets are most likely to produce a therapeutic benefit?
  • How does an RNA molecule behave in different cellular contexts?
  • What biological pathways may contribute to efficacy or toxicity?

Answering these questions requires integrating information from multiple sources, including genomics, transcriptomics, functional screens, and clinical studies.

Better biological annotation provides the structure needed to make those connections possible.

Why AI Depends On Better Biological Knowledge

Artificial intelligence is rapidly transforming biomedical research. Machine learning approaches are being applied to RNA design, target discovery, biomarker identification, and therapeutic optimization. However, AI systems are only as effective as the biological knowledge available to them.

Large data sets alone do not guarantee meaningful predictions. Computational models need context — the ability to understand relationships between biological entities and distinguish functional connections from unrelated correlations.

Structured biological resources such as Gene Ontology provide this foundation. By representing biological relationships in a consistent and interpretable way, annotation frameworks can help AI systems better understand the meaning behind complex data sets. As AI becomes increasingly integrated into RNA therapeutics, improving biological knowledge representation will become an important component of responsible and effective discovery.

Supporting The Next Generation Of RNA Medicines

The diversity of RNA therapeutic approaches continues to expand. Researchers are exploring new strategies involving messenger RNA, circular RNA, antisense oligonucleotides, RNA interference, RNA editing, and regulatory RNA pathways. Each modality introduces new biological questions.

A deeper understanding of RNA regulation will be essential for identifying appropriate targets, predicting therapeutic outcomes, and designing molecules with improved precision.

Better annotation also supports collaboration across disciplines. RNA therapeutics increasingly brings together molecular biologists, computational scientists, clinicians, data scientists, and drug developers. A shared framework for describing biological knowledge enables these communities to communicate more effectively and accelerate translation from discovery to therapy.

The Future Of RNA Discovery Is Knowledge-Driven

The next generation of RNA therapeutics will be shaped by innovation across multiple areas: chemistry, delivery technologies, manufacturing, computational modeling, and biology. Behind each of these advances lies a common requirement: a detailed understanding of how RNA functions within complex biological systems. Building better biological knowledge frameworks may not receive the same attention as a new delivery platform or therapeutic candidate, but it represents a critical foundation for progress.

Every carefully curated relationship, every improved annotation, and every standardized description of RNA function strengthens our ability to transform biological data into actionable insights. As RNA therapeutics continues to mature, the field will require not only new technologies but also better ways to organize and interpret the knowledge those technologies generate.

The future of RNA medicine will depend on our ability to decode RNA biology at unprecedented scale — and building the knowledge infrastructure to support that discovery is an essential step toward making the next generation of RNA therapeutics possible.

About The Author

Giulia Antonazzo, Ph.D., is a computational biologist and researcher with the Gene Ontology Consortium at the University of Cambridge, where she contributes to the development and refinement of standardized frameworks for representing biological knowledge. Her work focuses on improving the annotation and interpretation of gene and RNA regulation, enabling researchers to better integrate complex biological data sets and accelerate discovery. Antonazzo’s research interests center on bioinformatics, functional genomics, knowledge representation, and the computational analysis of molecular mechanisms. She has contributed to efforts that expand how non-coding RNA-mediated regulation is represented within biological databases, helping ensure that emerging discoveries in RNA biology can be accurately captured and leveraged by the scientific community. Antonazzo supports the development of resources that underpin modern computational biology, including data integration, machine learning applications, and systems-level approaches to understanding gene regulation.