Immunogenicity Prediction for Antibodies: MHC-II Screening
Immunogenicity prediction for antibodies: how MHC-II epitope screening works, real CovaBio output on trastuzumab VH/VL, and what in-silico ADA risk cannot.
Oliver Kraft
CovaSyn

Anti-drug antibodies do not usually kill a program at discovery. They kill it in the clinic, after the cell line, the tox package and two years of work. By then the only lever left is a new molecule.
Sequence-level T-cell epitope screening is the cheap lever. It runs in seconds, it runs before you commit to a lead, and it is wrong often enough that you have to know exactly what it is telling you. This article covers both halves.
Why MHC class II, and not "immunogenicity" in general
Clinically meaningful, persistent ADA responses against a protein therapeutic are usually T-cell dependent. The chain is: the drug is taken up by an antigen-presenting cell, proteolysed, loaded onto MHC class II, presented to CD4+ T helper cells, and if a T cell engages, B cells get help and class-switch to high-affinity IgG.
That means the class II binding step is the bottleneck you can actually predict from sequence. It is also the step that deimmunization targets: remove or weaken the peptide-MHC interaction and the help never arrives.
What sequence-level prediction does not cover: aggregates and particulates (a major real-world ADA driver), host-cell impurities, formulation and route, dose regimen, and the patient population. Those are handled elsewhere in your development plan.
How the screen works
covabio_immunogenicity slides a 9-mer window along the sequence and scores every window against 8 common HLA-DRB1 alleles using positional scoring matrices. Nine residues is the length of the MHC class II binding core; the P1/P4/P6/P9 anchor pockets do most of the binding work, while P2/P3/P5/P8 face the T-cell receptor.
Outputs are: a per-window score per allele, merged hotspot regions, per-allele binder counts, an overall epitope density, and a set of point mutations aimed at the TCR-facing positions.
Worked example: trastuzumab variable domains
Trastuzumab is a humanized IgG1 with a well documented clinical record, which makes it a useful calibration case: we know roughly what answer a screen ought to give.
Running the published VH and VL sequences through covabio_immunogenicity (live call, 2026-07-25):
| Metric | VH (heavy variable) | VL (light variable) |
|---|---|---|
| 9-mer windows scanned | 112 | 99 |
| Epitope density | 0.3929 | 0.2828 |
| Risk classification | moderate | moderate |
| Merged hotspot regions | 24 | 25 |
| Top allele by binder count | DRB1*0101, 46 binders | DRB1*0101, 32 binders |
| Lowest allele | DRB1*1101 and DRB1*1501, 29 binders | DRB1*1101, 26 binders |
| Deimmunization points suggested | 8 | 8 |
Epitope density here is the flagged-window fraction: 0.3929 x 112 = 44 flagged windows on the heavy chain, 0.2828 x 99 = 28 on the light chain (arithmetic derived from the tool output, not a separately reported field). Both land in the tool's "moderate" band. That is the right answer for a humanized antibody: not clean, not alarming.
For context, the trastuzumab US prescribing information reports treatment-emergent anti-therapeutic antibodies in roughly 0.1% of tested patients. So "moderate" epitope density at the sequence level clearly does not translate into a measured ADA rate. It is a relative triage signal between candidates, not an incidence forecast. More on that below.
Where the hotspots sit
The per-allele picture on VH is flat: DRB1*0101 leads with 46 binders, DRB1*1101 and DRB1*1501 trail at 29, mean scores span only 1.554 to 1.866. No allele is a dramatic outlier, which argues against a narrow HLA-restricted liability.
The strongest merged regions on VH, all binding all 8 alleles:
- Positions 1-12
VQLVESGGGLVQ, score 9.0 - Positions 67-75
FTISADTSK, score 9.0 - Positions 69-77
ISADTSKNT, score 8.0 - Positions 85-94
LRAEDTAVYY, score 8.0

Every one of those is framework, not CDR. That is the practically important observation, and it cuts both ways. Framework hotspots that sit in germline-derived stretches are shared with the patient's own antibody repertoire, so central and peripheral tolerance should suppress the response - a pure binding matrix has no tolerance filter and will flag them anyway. Treat germline framework hits as low priority until you check them against the closest human germline.
The CDR-proximal hits are the ones worth engineering attention: positions 50-59 IYPTNGYTRY (score 7.0, all 8 alleles) overlaps CDR-H2, and 98-106 WGGDGFYAM (score 4.0, 6 alleles) overlaps CDR-H3. These are non-germline, murine-derived in origin, and therefore not covered by tolerance.
The deimmunization suggestions, and their catch
The tool returned 8 point mutations for VH, all of the same class: conservative substitutions at TCR-facing P2 positions with a predicted binding reduction of +1.0 score unit each.
| Position | Original | Suggested | Sits in |
|---|---|---|---|
| 22 | A | S | Framework 1 |
| 23 | A | S | Framework 1 |
| 57 | T | S | CDR-H2 |
| 60 | A | S | CDR-H2 |
| 68 | T | S | Framework 3 |
| 73 | T | S | Framework 3 |
| 87 | A | S | Framework 3 |
| 105 | A | S | CDR-H3 region |
Read that column on the right carefully. Positions 57, 60 and 105 are inside or adjacent to CDRs. Trastuzumab's CDR-H2 RIYPTNGYTRYADSVK and CDR-H3 WGGDGFYAMDY were confirmed by covabio_antibody under Kabat numbering. Mutating a CDR to reduce epitope score is a direct trade against HER2 affinity, and the tool does not model affinity. Framework substitutions (22, 23, 68, 73, 87) are the low-risk starting set; CDR substitutions go on the list only if you are prepared to re-measure binding.
The other catch: a "+1.0 predicted binding reduction" is a shift on the tool's internal matrix scale. It is not a measured shift in peptide-MHC IC50, and it is not a measured drop in T-cell proliferation.
How to actually use this in a workflow
1. Screen every candidate in the panel at the same time, with the same allele set. The comparison across candidates is the reliable part. 2. Split hits into germline-framework and non-germline (CDR, junction, engineered linker, fusion junction). Prioritise the second group. 3. Check whether hotspots cluster on one allele. A broad, flat allele profile (as in the trastuzumab VH above) is a different engineering problem than a single-allele spike. 4. Design conservative substitutions outside the CDRs first, re-run the screen, and confirm the density actually drops. 5. Send the shortlist to a wet-lab assay. In-silico screening decides what you test, not whether it is safe.
Sequence liabilities do not travel alone. The same trastuzumab heavy chain scored 29/100 on covabio_developability with 5 aggregation-prone regions, 3 critical NG deamidation sites and 4 exposed methionines, and covabio_viscosity predicted 18.97 cP at 150 mg/mL, pH 6. Aggregation and immunogenicity are linked in practice, so it is worth looking at both reports side by side before you commit to a variant.
Honest limits: what this screen does not tell you
- It is a positional scoring matrix, not a trained neural predictor. It approximates class II binding preference. It does not reproduce NetMHCIIpan-class accuracy and we do not claim it does.
- No tolerance model. Self and germline-derived peptides are scored exactly like foreign ones. This inflates apparent risk for humanized and fully human frameworks.
- No antigen processing model. Real epitopes must survive endosomal proteolysis and compete for loading. Binding affinity alone overpredicts.
- DRB1 only, 8 alleles. No DP, no DQ, no DRB3/4/5. Population coverage is partial and skewed toward common European and East Asian alleles.
- No ADA incidence prediction. There is no calibration from epitope density to a clinical percentage, and the trastuzumab example shows why one should not be inferred.
- No B-cell or conformational epitopes, and nothing about aggregates, impurities, dose, route or population - which is where much of the observed clinical ADA signal comes from.
- Scores are internal and relative. Use them to rank candidates and to compare a variant with its parent. Do not report them as absolute risk.
The confirmatory work this feeds into is unchanged: HLA-peptide binding assays, MAPPs to see what is actually presented, and PBMC or T-cell proliferation assays on the shortlist. CovaBio is triage that tells you which three of thirty candidates deserve that spend.
Frequently asked questions
What is immunogenicity prediction for antibodies?
It is a sequence-level screen for T-cell epitopes: 9-amino-acid windows are slid along the antibody sequence and scored for binding to MHC class II alleles, usually HLA-DRB1. Windows that bind strongly across many alleles are flagged as hotspots. The output ranks candidates by relative anti-drug-antibody risk and points to residues for deimmunization. It does not predict a clinical ADA incidence.
Why MHC class II and 9-mers?
Persistent, high-affinity ADA responses are typically CD4+ T-cell dependent, so MHC class II presentation is the rate-limiting step you can predict from sequence. The class II binding groove is open at both ends and accommodates a 9-residue core, with anchor pockets at P1, P4, P6 and P9. Scoring every 9-mer window therefore covers all possible binding registers in the sequence.
What does an epitope density of 0.39 mean?
It is the fraction of scanned 9-mer windows flagged as predicted binders. On trastuzumab VH, covabio_immunogenicity returned 0.3929 across 112 windows, classified "moderate". It is a relative metric for comparing candidates or comparing a variant with its parent. There is no published calibration from epitope density to a percentage of patients who will develop ADA.
Can I just mutate every predicted epitope?
No. Many hotspots sit in germline framework, where tolerance likely suppresses a response and mutation only adds risk. Others sit in CDRs, where substitution trades directly against target affinity - on trastuzumab VH, 3 of the 8 suggested mutations fall in or next to CDRs. Start with conservative framework substitutions outside the paratope, re-screen, and re-measure binding on anything CDR-adjacent.
How accurate is in-silico immunogenicity prediction?
Matrix-based MHC-II screening is a triage tool, not a validated assay. It has no tolerance model, no antigen-processing model, and covers only 8 HLA-DRB1 alleles, so it overpredicts for humanized frameworks and misses non-T-cell drivers such as aggregates and impurities. Use it to rank and to guide engineering, then confirm the shortlist with MAPPs or T-cell proliferation assays.
When in the workflow should this run?
Before lead selection, alongside developability. The screen costs seconds per sequence, so run the whole panel, not the survivor. Its value is in deselecting candidates while switching is still cheap - once you have a cell line and a tox package, an immunogenicity finding means a new molecule.
Related reading
- Antibody developability triage from sequence
- Predicting antibody viscosity at high concentration
- AlphaFold3 vs OpenDDE, ESMFold, Boltz and Chai
Screening a candidate panel takes a sequence and a few seconds - the free tier is enough to run your own VH and VL and see where the hotspots land. - Retention Time Prediction in HPLC Method Development
Tools for this topic
Use these in your AI agent right away.
- CovabioAntibodies, peptides, mRNA, siRNA, ADCs.
