Guest Column | July 31, 2026

Beyond Better Models: Why The Future Of AI-Designed mRNA Depends On More Relevant Biological Data

By Devan Shah, founder & CEO, RNAV8 Bio

GettyImages-2284767177.jpg

The remarkable success of synthetic mRNA therapeutics over the last five years has transformed how the industry thinks about programmable medicines. COVID-19 vaccines scaled the platform at unprecedented speed. Since then, the field has rapidly expanded into gene editing, protein replacement, and in vivo cell engineering, with clinical milestones ranging from hereditary angioedema to transthyretin amyloidosis to cancer vaccines to the first in vivo CAR-T programs.

Yet despite this clinical momentum, one of the most discussed topics at this year’s TIDES meeting was artificial intelligence. Nearly every organization developing mRNA therapeutics is exploring AI-assisted sequence design. The promise is obvious: use machine learning to identify sequence architectures that maximize protein expression while improving stability, reducing dose, and ultimately expanding the therapeutic window.

But there is a fundamental problem: Today’s AI models are being trained on datasets that often bear little resemblance to real therapeutic mRNAs.

The Data Problem Hiding Behind The AI Revolution

Over the past several years, numerous machine learning approaches have emerged to optimize untranslated regions (UTRs), codon usage, and translation efficiency. However, a review of major published models reveals several recurring limitations.

Many are trained using short reporter constructs rather than therapeutic payloads. Others rely on unmodified RNA instead of N1-methylpseudouridine (m1Ψ)-modified molecules used clinically. Most evaluate lipid-free transfection rather than LNP delivery, and the overwhelming majority use immortalized cell lines instead of therapeutically relevant primary cells or tissues.

These differences matter because every one of these variables changes the biology. Translation kinetics, RNA structure, intracellular trafficking, immune sensing, and degradation pathways all behave differently in therapeutic settings than in simplified experimental systems.

A Self-Driving Car Trained Only On Sunny Highways

During my presentation, I used an analogy from autonomous driving.

Imagine training a self-driving vehicle exclusively on sunny California highways or even worse, an F1racetrackk. The model may perform beautifully during development, but it will fail when deployed into chaotic neighborhood streets, snowstorms, construction zones, heavy traffic, or poorly marked roads.

The failure is not necessarily the neural network. It is the distribution mismatch between training data and deployment.

The same challenge exists for therapeutic mRNA. Current AI systems are often trained using simplified experimental conditions, while therapeutic deployment involves long m1Ψ-modified transcripts delivered by LNPs into complex tissues such as liver, dendritic cells, T cells, or hematopoietic stem cells. Improving model architecture alone cannot overcome a dataset that fails to represent therapeutic reality.

Why Protein AI Advanced Faster Than RNA AI

Another important distinction is the enormous asymmetry in available biological data. Protein biology benefits from decades of structural biology, hundreds of thousands of experimentally solved structures, hundreds of millions of protein sequences, and mature structure-function relationships. These resources enabled transformative advances such as AlphaFold.

Therapeutic mRNA lacks an equivalent foundation. Experimental structural data remain comparatively sparse, therapeutic-scale expression datasets are limited, and the relationship between RNA sequence, chemistry, higher-order structure, and biological function remains incompletely understood.

This is less an AI problem than a biology problem.

New Biology Is Exposing Blind Spots In Existing Models

One particularly compelling example comes from recent work identifying TENT4-recruiting viral regulatory elements. These sequence elements increase mRNA stability by extending poly(A) tails with mixed nucleotides that resist deadenylation. Importantly, this represents an entirely new class of regulatory features that most current AI models simply cannot recognize because those features were absent from their training datasets.

As our understanding of RNA biology expands, AI models will need to incorporate new biological features rather than simply learning better representations of old ones.

Building The Right Datasets

The opportunity for the field is not simply to generate larger datasets; it is to generate biologically relevant datasets. That means studying therapeutic-length mRNAs, clinically relevant chemical modifications, authentic delivery systems, diverse primary cell types, and multiple functional readouts including translation, stability, immunogenicity, manufacturability, and in vivo performance.

Rather than treating experimentation and AI as separate disciplines, they should form a continuous feedback loop in which empirical data informs computational models, which in turn guide the next generation of experiments.

Looking Ahead

The next decade of mRNA therapeutics will likely be defined less by whether AI is used, and more by what biological knowledge those AI systems learn from.

The industry has already demonstrated that mRNA can become a versatile therapeutic modality. The next challenge is making its design increasingly deterministic rather than empirical. Achieving that goal will require richer datasets, deeper understanding of RNA biology, and tighter integration between computational modeling and experimental validation.

The future of mRNA/LNP will not be built solely through superior algorithms. It will be built by creating more therapeutically relevant biological datasets.

About The Author

Devan Shah is founder and CEO of RNAV8 Bio (“Renovate”), an mRNA engineering and design platform focused on radically improving mRNA therapeutic performance across rare and common diseases. His background spans finance/VC, BD, nucleic acid manufacturing, cell/gene therapy, and computational biology. Shah was recently the founding head of BD and founding business head of the Nucleic Acids and Cell Therapy Franchises at National Resilience (“Resilience”). At Resilience, he built the Nucleic Acids business from $0 to over $100M in contracted revenue with publicly disclosed clients such as Intergalactic Therapeutics and Moderna. Before Resilience, Devan led business development at Stanford Medical School’s Center for Definitive and Curative Medicine (CDCM), where he negotiated industry partnerships with biotechs and manufacturers in the cell, gene, and antibody fields to accelerate the translation of these novel Stanford-developed therapies into the clinic. Devan began his career on Wall Street in healthcare investment banking and life science VC at Citigroup and New Leaf Venture Partners.