Antibody Developability Prediction From Sequence
Antibody developability prediction from sequence alone: real CovaBio output on trastuzumab re-finds its documented CDR and Fc liabilities before wet-lab.
Oliver Kraft
CovaSyn

You have twenty lead candidates and budget to characterise three. The liabilities that will hurt you later - a deamidation motif in a CDR, an isomerisation site next to the paratope, a viscosity wall at subcutaneous concentration - are all encoded in sequences you already have. The question is whether an in-silico triage is good enough to decide which three molecules get the DSC run and the accelerated stability study.
Here is one way to test that: run a marketed antibody whose liabilities are already documented in the literature, and see whether the tool finds them without being told.
What sequence-only developability screening actually computes
CovaBio's covabio_developability takes an amino acid sequence and returns a composite 0-100 score plus per-dimension detail: aggregation-prone windows, deamidation motifs (NG, NS, NT, NH), isomerisation motifs (DG, DS, DT), oxidation-prone Met and Trp with an exposure call, N-glycosylation sequons, cysteine pairing, pyroglutamate and C-terminal Lys heterogeneity, an HIC retention index and a heuristic aggregation temperature.
None of this needs a structure, an assay or a homology model. It is pattern chemistry plus windowed hydrophobicity. That is exactly why it is cheap enough to run on every candidate in a panel - and exactly why it is triage rather than evidence.
Worked example: trastuzumab, sequence in, liabilities out
We ran the public trastuzumab (anti-HER2 IgG1) heavy and light chain sequences through three CovaBio tools. Every number below is a verbatim tool output from a live run on 2026-07-25.
Step 1: which chain owns the risk
covabio_developability on each chain separately:
| Chain | Score | Class | Flags returned |
|---|---|---|---|
| Heavy | 29 / 100 | poor | 5 aggregation hotspots, 3 critical NG deamidation sites, 4 exposed-Met oxidation sites, 1 DP cleavage site, unpaired Cys |
| Light | 72 / 100 | good | 3 aggregation hotspots, 1 exposed-Met oxidation site, unpaired Cys |
Two calls and the engineering attention goes to the heavy chain. That is the entire point of a triage layer.
Step 2: are the liabilities in the CDRs or in forgiving framework
A liability in FR3 is a nuisance. The same motif in CDR-H2 is a potency risk. covabio_antibody (Kabat scheme) returned for the heavy chain:
- CDR-H1: DTYIH, risk low
- CDR-H2: RIYPTNGYTRYADSVK, risk moderate, 1 deamidation motif
- CDR-H3: WGGDGFYAMDY, risk moderate, 2 oxidation-prone residues
Cross-referencing the liability positions from step 1:
- NG at position 55, risk "critical", context YPTNGYTR. That NG sits inside CDR-H2.
- DG at position 102, risk "high", context WGGDGFYA. That DG sits inside CDR-H3.

Asp isomerisation in trastuzumab's CDR-H3 is a documented, published critical quality attribute - it was characterised in the peer-reviewed literature and is controlled in manufacture. The tool re-found it from 450 characters of sequence, with no knowledge of the molecule.
Step 3: the Fc liabilities everyone already controls
The same run flagged exposed Met oxidation at linear positions 255, 361 and 431 (risk "high"), plus Met83 in the VH. Positions 255 and 431 correspond to the classic EU-numbered Met252 and Met428 Fc oxidation sites that drive FcRn binding loss. Note the numbering: the tool reports linear sequence index, not EU or Kabat numbering. If you compare against a literature position you have to convert.
It also returned the single N-glycosylation sequon at position 300 (NST, context EQYNSTYRV) - the conserved CH2 glycan - and a "high" C-terminal Lys clipping flag, which is the textbook source of mAb charge heterogeneity.
Step 4: formulation risk before formulation exists
covabio_viscosity at 150 mg/mL, pH 6.0 (a subcutaneous target) returned:
- estimated viscosity 18.97 cP, risk level "moderate"
- drivers: charge-patch asymmetry 0.33 (2 positive, 4 negative patches), largest hydrophobic patch length 6
- estimated pI 8.02, delta pH-pI 2.02, pI-proximity risk "low"
- recommendation: excipient screening (arginine, proline) or lower concentration
- model basis: heuristic, Sharma 2014
And a heuristic aggregation temperature from the developability call: Tagg 69.5 C for the heavy chain, 69.1 C for the light chain, both labelled confidence: heuristic by the tool itself.
How to read the output without fooling yourself
Three rules we apply internally.
Rank, do not certify.
A 29 versus a 72 is a useful ordering. "29/100" is not a specification. Use the score to sort a panel, then let the assay decide.

Weight by location.
A critical NG in CDR-H2 and a critical NG at position 318 (DWLNGKEY, in the Fc) are the same motif with very different consequences. The CDR call from covabio_antibody is what turns a motif list into a decision.
Read the tool's own hedges.
The viscosity model says heuristic_sharma_2014. The Tagg says Empirical estimate based on sequence composition, not structure. Those strings are in the output for a reason.
Honest limits: what this does not tell you
- No structure means no real solvent exposure. "Exposed" and "buried" are sequence-context heuristics. A motif buried in the folded protein can be scored as exposed and vice versa.
- Per-chain analysis misreads inter-chain disulfides. Both chains came back with an "unpaired cysteine" flag and an odd Cys count. That is an artefact of scoring each chain in isolation: the trastuzumab hinge and the H-L cysteines pair across chains. The disulfide pairing in the output is a distance heuristic, not a prediction of the real bonding pattern. Do not act on that flag for a normal IgG.
- Tagg and viscosity are estimates, not measurements. 69.5 C and 18.97 cP are sequence-composition heuristics. They are useful to rank ten candidates against each other. They are not a substitute for DSC, DLS or a cone-and-plate rheometer, and they should never appear in a filing as data.
- Motif presence is not modification rate. An NG motif is a susceptibility, not a degradation rate. Actual deamidation depends on local flexibility, pH, buffer and temperature. Forced degradation plus peptide mapping is what quantifies it.
- No immunogenicity certainty.
covabio_immunogenicityreturns MHC-II binding-window statistics (for trastuzumab VH: epitope density 0.39, 112 nine-mer windows across 8 HLA-DRB1 alleles, 24 hotspots). That is an in-silico prior, not a clinical ADA rate. - One antibody is not a validation set. Trastuzumab is a credibility check, not a benchmark. It shows the tool re-discovers known liabilities; it does not establish a prospective hit rate on novel scaffolds.
Where this fits in a real workflow
Run it at the point where you have sequences and no material. A typical pass is: developability score per chain, CDR mapping to separate paratope liabilities from framework noise, viscosity at your intended dose concentration, immunogenicity screen, then a shortlist. Everything after that is wet lab - expression titre, SEC, DSC, forced degradation, peptide mapping.
The value is not that the model is right. It is that it is cheap, reproducible, and it puts the argument on paper before you spend the money.
Frequently asked questions
Can antibody developability be predicted from sequence alone?
Partly. Sequence-only tools reliably detect chemical liability motifs - NG and NS deamidation, DG isomerisation, exposed Met and Trp oxidation, N-glycosylation sequons, C-terminal Lys clipping - and map them to CDR or framework regions. They estimate aggregation propensity, viscosity and aggregation temperature only heuristically. Use sequence screening to rank and shortlist candidates; use DSC, SEC, HIC and forced degradation to confirm.
What did CovaBio find on trastuzumab?
Running the public sequences through covabio_developability and covabio_antibody gave a heavy chain score of 29/100 ("poor") versus 72/100 ("good") for the light chain. It flagged a critical NG deamidation motif at position 55 inside CDR-H2 (RIYPTNGYTR), a high-risk DG isomerisation motif at position 102 inside CDR-H3 (WGGDGFYA), and exposed Met oxidation at linear positions 255, 361 and 431 in the Fc.
How accurate is sequence-based viscosity prediction?
It is a heuristic, not a measurement. covabio_viscosity uses a sequence-based model (Sharma 2014) driven by charge-patch asymmetry, surface hydrophobicity and pI proximity. For trastuzumab at 150 mg/mL and pH 6.0 it returned 18.97 cP, risk level "moderate", with charge-patch asymmetry 0.33 and a length-6 hydrophobic patch as drivers. Treat the number as a relative ranking signal between candidates, not an absolute cP value.
Why do both antibody chains show an unpaired cysteine?
Because each chain is scored in isolation. An IgG1 heavy chain and its light chain both contain cysteines that pair across chains - the hinge disulfides and the H-L bond. A per-chain analysis sees an odd cysteine count and flags aggregation risk. For a conventional IgG this flag is an artefact and should be ignored; it matters only for engineered constructs with genuinely free thiols.
What is a good developability score?
There is no universal cutoff, and treating one as absolute is the main way these tools get misused. The composite score is ordinal. Compare candidates from the same panel scored the same way, look at which dimensions drive a low score, and weight liabilities in CDRs far more heavily than the same motifs in framework or constant regions.
Does an NG motif in a CDR disqualify a candidate?
No. It flags a control strategy problem, not a dead molecule. Several marketed antibodies carry CDR deamidation and isomerisation sites and manage them through formulation pH, excipients and release testing. The decision you should take from the flag is to run forced degradation and peptide mapping early, and to check whether the modified form loses target binding.
Related reading
- Forced degradation study design under ICH Q1A(R2)
- Mass balance in forced degradation studies
- ICH Q1E shelf-life calculation from accelerated data
- Design space, NOR and PAR explained (ICH Q8)
Run the same three calls on your own sequence on the CovaSyn free tier and see where your candidate sits.
Tools for this topic
Use these in your AI agent right away.
- CovabioAntibodies, peptides, mRNA, siRNA, ADCs.
