CovaSyn
All Articles
Use Case7 min readJuly 15, 2026

Ionizable Lipid Screening Straight From SMILES

Ionizable lipid screening without burning stock: rank custom lipids from SMILES, pick the most informative next run, and see the real size trade-off.

OK

Oliver Kraft

CovaSyn

Ionizable Lipid Screening Straight From SMILES

Custom ionizable lipids are expensive and slow to make, and most screening campaigns burn a meaningful fraction of a precious batch on formulations that were never going to win. A 12-lipid microfluidic screen with RiboGreen, DLS and a cell readout is a week of hands-on work plus mRNA you cannot get back. The question worth answering before you touch a pipette is narrow: given the lipids I actually have, which one deserves the next run?

What "screening from SMILES" actually means

CovaLNP takes a formulation as structures plus ratios, not as names. You give it the ionizable lipid SMILES, the helper lipid, cholesterol, PEG-lipid, the molar ratios and the N/P ratio. It returns a predicted property vector: encapsulation efficiency, particle size, PDI and an expression score.

That matters for one specific reason. A name lookup only works for lipids that exist in someone's table. Your in-house lipid, the one with the branched tail you made last month, is not in any table. If the model is SMILES-native, a novel lipid enters the workflow exactly like MC3 does. There is no "unknown compound" path.

One caveat up front, and it is the honest limit of the whole approach: these are predictions trained on published LNP datasets. They are triage to rank your candidate list, not a substitute for RiboGreen, DLS or a transfection assay.

The three calls that replace a first-pass screen

Step 1: score the benchmark

DLin-MC3-DMA is still the reference most groups anchor to. Run it first so every later number has something to beat.

covalnp_predict on an MC3 mRNA-LNP returned:

PropertyPredicted value
Encapsulation efficiency86.3 %
Particle size94.9 nm
PDI0.150
Expression score (HeLa)3.56

That run used a 50 / 10 / 38.5 / 1.5 molar ratio at N/P 6.

Here is a useful thing that fell out of running the same lipid twice. A second covalnp_predict call on MC3 with a different composition returned 77.3 % encapsulation and 94.4 nm. Same ionizable lipid, different ratios, roughly nine points of encapsulation apart. The model is responding to composition, not just to the lipid. So fix your ratios and your N/P before you compare lipids, or you will attribute a formulation effect to chemistry.

Bar chart of predicted encapsulation efficiency in ionizable lipid screening: MC3 at 86.3 and 77.3 percent on two compositions, ALC-0315 at 84.4 percent, Pareto front best 88.0 percent.
The same MC3 lipid shifted roughly nine points of encapsulation on a change of ratios alone. Fix ratios and N/P before you compare lipids, or you will read a formulation effect as chemistry. Source: covalnp_predict, covalnp_compare and covalnp_optimize runs described in this article, all from ionizable lipid SMILES plus ratios.

The expression figure needs its own label. transfection_* outputs are a unitless predicted expression score where higher is better. It is not a luciferase RLU value and it does not carry a cell-line-specific unit. Treat it as a ranking signal.

Step 2: let active learning pick the next lipid

The instinct is to run a grid. The better move is to run the experiment that tells you the most.

covalnp_suggest ranks candidates by expected information gain rather than by predicted performance alone. Its top pick was ALC-0315, at a predicted expression of 3.81 with an expected information gain of 10.97. CKK-E12 trailed at a predicted 3.70.

Read that ordering carefully. It is not simply "the highest predicted number wins". Information gain rewards candidates that are both promising and in a region where the model is uncertain, because those runs move the model most. A candidate you are already confident about buys you a confirmation, not knowledge.

ALC-0315 went in as a structure, not a name:

CC(CCCCCCCC)OC(=O)CCCCN(CCCCCCCCCCCCCC)CCCCC(=O)OC(C)CCCCCCCC

Your proprietary lipid goes in the same field, the same way.

Step 3: confirm head to head, including the part you will not like

covalnp_compare puts both lipids at matched ratios so the only variable is the ionizable lipid:

MetricMC3ALC-0315Delta
Expression score (HeLa)3.563.81+7.0 %
Encapsulation efficiency77.3 %84.4 %+7.1 points
Particle size94 nm113 nm+19 nm

The size number is the one worth arguing about internally. A 19 nm increase is not automatically a loss, but if your spec sheet has an upper size limit, or you have a biodistribution reason to stay under 100 nm, then the encapsulation and expression gains come with a cost you now have to weigh. A screening tool that only reports the two metrics that improved is not helping you.

Why this ordering saves material

The sequence is deliberate: baseline, then most-informative candidate, then confirmation. It is the same logic as a sequential DoE. You are not trying to map the whole lipid space in silico. You are trying to decide which two or three lipids get bench time this month.

A practical shape for that:

  • Score every lipid you physically have, plus the literature references you want to beat, in one batch of covalnp_predict calls.
  • Drop anything that fails a hard spec outright, for example predicted size well outside your target window.
  • Run covalnp_suggest over what remains to get the ranked next experiment.
  • Take the top two or three to the bench, not the top ten.

For multi-objective work, covalnp_optimize walks the ionizable / helper / PEG / N/P space with NSGA-II and returns a constraint-valid Pareto front. In the run behind this article, encapsulation across that front spanned 82.1 % to 88.0 % against an MC3 starting point of 77.3 %. One honest note on that tool: in the same run its predicted_transfection field came back as 0.0, an internal quirk, so we used it for encapsulation only and took expression from predict and compare.

Honest limits

What this does not tell you:

  • No in vivo claim. The expression score is an in vitro-flavoured ranking signal. It says nothing about liver tropism, endosomal escape efficiency in a given tissue, or dose response in an animal.
  • No apparent pKa or ionisation profile. The models predict formulation outcomes, not the lipid's titration behaviour, which is often the property you actually want to design against.
  • Name resolution is not reliable. SM-102 supplied as a name did not resolve to distinct features in our testing, so we do not quote SM-102 numbers anywhere. Supply structures.
  • No feature attribution in the standard tier. covalnp_explain is enterprise-gated, so we make no SHAP-style claims about which substructure drove a prediction.
  • Applicability domain. Predictions on lipid chemistries far from the published training distribution, novel head groups in particular, should be treated as weak priors and confirmed early.
  • Stability is out of scope. Encapsulation at t=0 is not four-week stability at 5 degrees C, and nothing here speaks to lipid hydrolysis or mRNA integrity over time.

Frequently asked questions

What is ionizable lipid screening?

Ionizable lipid screening is the process of ranking candidate ionizable lipids for an mRNA or siRNA lipid nanoparticle before committing material to formulation work. Traditionally it means formulating each candidate by microfluidic mixing and measuring encapsulation, size, PDI and transfection. In silico screening ranks the same candidates from their chemical structures first, so only the top few reach the bench.

Can I screen a proprietary lipid that is not in any database?

Yes, if the tool is structure-native. CovaLNP takes the ionizable lipid as SMILES together with helper lipid, cholesterol, PEG-lipid, molar ratios and N/P ratio, so an unpublished in-house lipid is handled exactly like a reference lipid. Name-based lookups fail here: in our testing SM-102 supplied as a name did not resolve to distinct features, which is why structures are the supported input.

How does ALC-0315 compare with MC3 in prediction?

In a matched-ratio covalnp_compare run, ALC-0315 predicted higher than DLin-MC3-DMA on both encapsulation efficiency, 84.4 % against 77.3 %, and expression score, 3.81 against 3.56, a 7.0 % relative gain. Particle size increased from 94 nm to 113 nm. These are model predictions for triage, not measured assay values, and the size increase may matter for your spec.

What is expected information gain and why not just pick the highest prediction?

Expected information gain scores how much a proposed experiment would reduce model uncertainty, not just how good the outcome looks. In our covalnp_suggest run ALC-0315 came top with a predicted expression of 3.81 and an information gain of 10.97, ahead of CKK-E12 at 3.70. Running only your most confident candidate confirms what you already believe and teaches the model very little.

How accurate are LNP property predictions?

Accurate enough to rank candidates, not accurate enough to release a batch. The models are trained on published LNP formulation data, so predictions for chemistries close to that distribution are more trustworthy than novel head groups. Predicted encapsulation for the same lipid shifted from 86.3 % to 77.3 % when composition changed, so hold ratios and N/P constant when you compare lipids.

Does formulation composition change the prediction, or only the lipid?

Composition changes it substantially. Two covalnp_predict runs on the same MC3 lipid at different molar ratios returned 86.3 % and 77.3 % encapsulation. That is the intended behaviour, since ratios and N/P genuinely drive encapsulation, but it means a lipid-versus-lipid comparison is only meaningful at matched composition. Use the compare tool rather than two separate ad hoc predictions.

Related reading

Bring your own lipid SMILES and score it on the free tier before your next formulation run. - LNP Formulation Optimization for mRNA

Tools for this topic

Use these in your AI agent right away.

  • CovalnpLNP formulation for mRNA: encapsulation, size, N/P.
Ionizable Lipid Screening Straight From SMILES | CovaSyn