pLDDT Explained: When to Trust a Predicted Structure
pLDDT interpretation for practitioners: what the bands mean, why a global mean hides the region you care about, and real CovaFold numbers showing what MSA.
Oliver Kraft
CovaSyn

You get a predicted structure back with a single number attached: pLDDT 83. Someone wants to know whether the binding site can be docked against. The global mean cannot answer that, and in one of our own runs the residue everyone cared about scored in the low 30s while the whole-chain average looked survivable. pLDDT is useful, but only if you read it per residue and know what drove it.
What pLDDT actually measures
pLDDT is the model's predicted local distance difference test score: a per-residue, 0 to 100 estimate of how well the local environment of that residue would agree with an experimental structure. Two things follow immediately.
It is local, not global. A pLDDT of 95 on residue 12 says the neighbourhood of residue 12 is likely right. It says nothing about whether the domain it sits in is correctly placed relative to another domain. For inter-domain or inter-chain placement you need PAE (predicted aligned error), pTM or ipTM instead.
It is a self-assessment, not a validation. The model is predicting how good its own answer is. That estimate is well calibrated on targets resembling its training distribution and degrades outside it - which is exactly what the numbers below show.
The whole-chain figure people quote is just the mean over residues. It is a summary of a distribution, and like any mean it is easy to fool.
The bands
These are the conventional AlphaFold-derived bands, and CovaFold's covafold_quality uses the same thresholds internally (high-confidence fraction counts residues at or above 70; low-confidence residues are those below 50).
| pLDDT | Band | What you can do with it |
|---|---|---|
| > 90 | Very high | Backbone and most side-chain rotamers reliable. Usable for docking, mutation modelling, pocket analysis. |
| 70 - 90 | Confident | Backbone generally reliable, side chains less so. Fine for fold and topology, treat rotamer-level detail with care. |
| 50 - 70 | Low | Treat as a hypothesis. Backbone may be roughly right, may not. Do not dock into it and report the pose as if it were meaningful. |
| < 50 | Very low | Not a structural prediction. Often a strong predictor of intrinsic disorder, not an error per se. |
The last row is the one most often misread, so it gets its own section.
Low pLDDT is not always a failure
For intrinsically disordered regions, a sub-50 score is the model being right. There is no single native conformation to predict, and the published AlphaFold work established that low pLDDT correlates with disorder well enough to be used as a disorder predictor in its own right. A flexible loop, a terminal tail, a linker between two folded domains - these routinely come back below 50 while the folded cores sit above 90.
So before you call a model bad, ask which residues are low and whether those residues have any business being ordered. covafold_quality reports low_confidence_residues, largest_low_conf_region and total_low_conf_residues alongside the global mean precisely so you can answer that without opening the PDB by hand. A model with a 40-residue disordered tail at pLDDT 35 and a core at 95 is a good model with a global mean that lies to you.
The failure mode looks different: low confidence scattered *through* a region that should be a folded, conserved domain. That is the model telling you it does not know the fold.
Worked example: the same construct, three engines, four numbers
This is CovaSyn's own measured data on KRAS G12C (189 aa), all via covafold_fold, and it is the clearest illustration we have of what actually drives pLDDT.
| Setup | Mean pLDDT | High-confidence fraction (>= 70) | pLDDT at residue 12 (the G12C site) |
|---|---|---|---|
| OpenDDE engine, no MSA (single sequence) | 49.4 (a repeat control in a later session: 51.22) | 1 - 12 % | 32 - 37 |
| ESMFold (single-sequence by design) | 83.1 | 89 % | - |
| OpenDDE + self-hosted MSA, depth 3,913 | 94.74 | 98.9 % | 95.47 |
| OpenDDE + MSA, depth 8,112 / 16,917 | 94.59 / 94.74 | 98.4 % | 95.61 |
Same sequence. Same weights in rows 1, 3 and 4. The only variable is whether the model got a multiple sequence alignment.
Read what this does to interpretation. In the no-MSA row, the mean of 49.4 is honest - almost nothing in that model is trustworthy, and the mutation site itself is at 32 to 37. But note that this is *not* an intrinsically low-confidence target: with coevolutionary signal the same residue reaches 95.5. Had you seen only the global mean around 50 you might have shrugged and called KRAS a hard target. It is not. The pipeline was starved.

The second lesson is that MSA depth saturates. Going from 3,913 to 16,917 sequences moved the mean by 0.15 pLDDT points, well inside run-to-run noise. Depth matters enormously up to a point and then buys nothing.
For contrast, on the same engine and the same no-MSA configuration, a ubiquitin control folded to 93.3 with 97 % high-confidence residues. Small, compact, heavily represented - the engine was never broken. The signature was target-dependent. That is why "our folding is good, we measured 93" is a claim you should always ask for the target list behind.
Our current documented monomer runs, post-MSA-enable:
| Target | Length | Mean pLDDT | MSA depth |
|---|---|---|---|
| Ubiquitin | 76 aa | 97.19 | 9,640 |
| KRAS G12C | 189 aa | 94.61 | 7,513 |
| CDK2 | 298 aa | 94.79 | 9,523 |
These are monomer folds measured on CovaSyn infrastructure, three targets, all well-represented proteins. They are not a benchmark and should not be read as an expected accuracy for a novel or orphan sequence.
A practical decision procedure
1. Look at the per-residue profile before the mean. If your tool only gives you a mean, get a different tool.
2. Score the region you will actually use. Docking a pocket? Read the pLDDT of the pocket residues. Modelling a point mutation? Read that residue. A global 83 with a 35 at your site is a no.
3. Ask whether the low regions are supposed to be ordered. Termini and linkers below 50 are expected. A low-confidence beta sheet in a conserved domain is not.
4. Check the MSA depth reported with the run. If it is absent or tiny, you are looking at a single-sequence prediction and the confidence estimate is out of its calibrated regime.
5. For complexes, stop using pLDDT. Use ipTM for interface confidence and PAE for relative placement. For reference, our KRAS G12C complex run via covafold_cofold reported ipTM 0.99, and the ubiquitin complex 0.63 - same pipeline, completely different interface confidence. The monomer pLDDT would not have told you that.
What pLDDT does not tell you
- Whether the conformation is the biologically relevant one. Apo vs holo, active vs inactive kinase conformation, which of several accessible states - all invisible to pLDDT. A 95 model can be confidently the wrong state.
- Whether domains are correctly arranged. That is PAE's job.
- Anything about the interface in a complex. That is ipTM's job.
- Whether a ligand will bind. Structure confidence is not affinity. Downstream docking scores have their own, separate uncertainty.
- Whether the prediction is right. It is a calibrated estimate, and calibration is a property of the training distribution. Designed, synthetic or orphan sequences with no homologs sit outside it.
- Anything you can put in a filing. Predicted structures are triage and design input. Experimental determination remains the evidence.
Frequently asked questions
What is a good pLDDT score?
Above 90 is very high confidence: backbone and most side chains are reliable enough for docking and mutation modelling. 70 to 90 is confident, with the backbone generally right and side-chain detail less certain. 50 to 70 should be treated as a hypothesis only. Below 50 is not a usable structural prediction, though it is often a correct signal of intrinsic disorder rather than a failed run.
Is a pLDDT of 70 good enough for docking?
Generally no, not on its own. What matters is the pLDDT of the pocket residues, not the chain mean. A model averaging 70 can have a binding site at 85 (workable) or at 40 (not). Extract the per-residue scores for the residues lining the site and judge those. Docking into a low-confidence pocket produces poses that look plausible and mean nothing.
Why is my pLDDT low even though the protein is well studied?
The most common cause is a missing or shallow multiple sequence alignment. On CovaSyn's own KRAS G12C runs via covafold_fold, the same construct scored a mean pLDDT of 49.4 without an MSA and 94.74 with one at depth 3,913 - the mutation site itself moved from roughly 32 to 95.5. Check the MSA depth reported with your run before concluding the target is hard.
What is the difference between pLDDT, pTM and ipTM?
pLDDT is per-residue and local: it estimates how correct the immediate environment of each residue is. pTM estimates global fold accuracy for a whole chain. ipTM estimates the accuracy of the interface between chains in a complex. For a monomer read pLDDT; for domain arrangement read PAE; for a protein-protein or protein-ligand interface read ipTM. They are not interchangeable.
Does low pLDDT mean the region is disordered?
Often, yes. Low pLDDT correlates strongly with intrinsic disorder, and the AlphaFold authors documented this well enough that the score is used as a disorder predictor. Terminal tails and inter-domain linkers routinely score below 50 in otherwise excellent models. The concerning pattern is low confidence spread through a region that should be a compact, conserved, folded domain.
Does a deeper MSA always give a better structure?
No. Depth matters greatly up to a point and then saturates. On CovaSyn KRAS G12C runs, an MSA depth of 3,913 gave mean pLDDT 94.74 while 16,917 gave 94.74 as well - a difference inside run-to-run noise. Getting from zero homologs to a few thousand is transformative. Getting from a few thousand to sixteen thousand generally is not.
Related reading
Fold a sequence with covafold_fold on the free tier and read the per-residue profile from covafold_quality rather than the headline mean.
- AlphaFold3 Alternatives in 2026: An Honest Comparison
- Binding Site Druggability Prediction, Honestly Read
- Co-Folding vs Docking: Which Question Are You Asking?
Tools for this topic
Use these in your AI agent right away.
- CovafoldProtein folding and structure prediction.
