CovaSyn
All Articles
Explainer8 min readJuly 15, 2026

Docking Score Interpretation: What -10 kcal/mol Means

Docking score interpretation, honestly: what a -10 kcal/mol Vina score does and does not tell you, using a real CovaDock run on KRAS G12C.

OK

Oliver Kraft

CovaSyn

Docking Score Interpretation: What -10 kcal/mol Means

Someone hands you a docking result: -10.35 kcal/mol, labelled "very strong binding". The number has energy units, so it looks like a binding free energy. It is not one. If you treat it as one, you will over-promise a compound to your project team and be wrong by orders of magnitude in Kd.

This is what that number actually is, what it can carry, and where it breaks.

The worked example: sotorasib into KRAS G12C

We ran sotorasib (AMG 510) against KRAS G12C through covadock_dock on GPU. The receptor was the deposited crystal structure, PDB 6OIM. Ligand structure came from covabasic_pubchem (CID 137278711, C30H30F2N6O3, MW 560.6) and was checked with covabasic_validate (InChIKey NXQKSXLFSAEQCZ-SFHVURJKSA-N).

Verbatim outputs:

FieldValueTool
Best pose score-10.35 kcal/molcovadock_dock
Label"very strong"covadock_dock summary
Protein-ligand interactions, top pose61covadock_results
Key contactsGly24, Ala25, Gly74, Gln75 (switch-II)covadock_results
ReceptorPDB 6OIM (crystal, not a predicted model)run config

That is a good result. Sotorasib is a real approved KRAS G12C drug, and the pose lands in the switch-II pocket the drug class is designed for. The contacts are chemically the right ones. So this is close to the best case for docking: known drug, high-quality co-crystal receptor, well-defined pocket.

And even in that best case, the number is not a binding free energy.

Where the "kcal/mol" comes from

The default engine behind covadock_dock is AutoDock Vina (default exhaustiveness=16, 10 poses per engine, PoseBusters validity check on the top poses). Vina's scoring function is an empirical sum of weighted terms - steric, hydrophobic, hydrogen bonding, plus a rotatable-bond penalty - with weights fitted to reproduce measured affinities across a training set.

Two consequences follow directly:

1. The output has energy units because the fit target had energy units. There is no partition function, no explicit solvent, no entropy calculation. It is a regression output wearing a thermodynamic costume. 2. Its accuracy is bounded by the fit, not by physics. Off the training distribution, it degrades without warning.

The label CovaDock prints is a fixed threshold scale, published in the tool docstring so you can see the cut points: weak (> -6), moderate (-6 to -8), strong (-8 to -10), very strong (< -10). Sotorasib at -10.35 clears the last bin by 0.35 kcal/mol. That margin is far inside the method's error bar.

Threshold chart for docking score interpretation: sotorasib scores -10.35 kcal/mol against CovaDock's very strong cut point of -10, a margin of 0.35 kcal/mol inside a 2.4 kcal/mol error bar.
Sotorasib clears the very strong bin by 0.35 kcal/mol, while the method's residual scatter is about 2.4 kcal/mol. The bin boundary is finer than the score can resolve. Source: covadock_dock run of sotorasib (PubChem CID 137278711, C30H30F2N6O3, MW 560.6) into KRAS G12C on crystal receptor PDB 6OIM; label cut points from the tool docstring.

How big is the error bar? The benchmark numbers

The standard reference is CASF-2016 (Su et al.), a core set of 285 protein-ligand complexes across 57 targets with 5 ligands each. Published AutoDock Vina results on that set:

  • Docking power (does the top-ranked pose reproduce the crystal pose within 2.0 A RMSD): 90.2% top-1 success. Increases to 94.4% when structural waters are included.
  • Scoring power (correlation of score with measured affinity): Pearson R = 0.604, standard deviation 1.73 log-affinity units.
  • Ranking power (ordering ligands correctly within one target): Spearman rho = 0.528.
Bar chart of AutoDock Vina on CASF-2016 for docking score interpretation: 90.2% docking power for pose geometry against Pearson R 0.604 scoring power and Spearman rho 0.528 ranking power.
Vina reproduces the crystal pose about nine times in ten, but its score tracks measured affinity at only R 0.604. Trust the geometry more than the number. Source: CASF-2016 core set, 285 protein-ligand complexes across 57 targets with 5 ligands each; published AutoDock Vina values as reported in Lin_F9 (PMC8478859).

Read those three lines together, because they say three different things.

Vina finds the right pose about nine times in ten on well-behaved crystal structures. That is the part that works, and it is why the 61 switch-II contacts in the sotorasib run are worth looking at.

Vina's affinity number carries R = 0.604 with an SD of 1.73 log units. At 298 K, one log unit of Kd is RT ln(10) = 1.36 kcal/mol, so 1.73 log units is roughly 2.4 kcal/mol of residual scatter. Sotorasib's -10.35 therefore sits inside a distribution that comfortably spans -8 to -13 in equivalent terms. A 0.35 kcal/mol difference between two compounds is noise. A 3 kcal/mol difference is a weak signal.

Three-step arithmetic for docking score interpretation: 1.73 log units of CASF-2016 scatter times 1.36 kcal/mol per log unit gives 2.4 kcal/mol, so a -10.35 score spans -8 to -13.
The benchmark scatter in kcal/mol is why a 0.5 kcal/mol gap between two compounds is noise. Use docking to cut a list, then measure the survivors. Source: CASF-2016 AutoDock Vina scoring power standard deviation of 1.73 log-affinity units (Lin_F9, PMC8478859), converted at RT ln(10) = 1.36 kcal/mol per log unit of Kd at 298 K.

Spearman 0.528 is the number most people should care about most, because within-target ranking is what a hit list actually is. Roughly: the ordering is better than random and clearly worse than reliable.

Enrichment versus affinity: the distinction that matters

These are two different jobs, and docking is only good at one of them.

Enrichment

is: given 100,000 compounds, does the top 1% contain more actives than a random 1% would? Here a noisy score is still useful. You are asking for a population shift, not a per-compound truth. Docking earns its keep in virtual screening precisely because a mediocre correlation applied across a large library still concentrates actives.

Affinity prediction

is: is this molecule 30 nM or 3 uM? Docking scores cannot answer that with any usable confidence interval. The residual scatter above is larger than the potency window most medicinal chemistry decisions turn on.

Practical translation:

  • Use scores to cut a large list. Good.
  • Use scores to order the top 20 for synthesis. Weak; combine with interaction patterns, pose sanity and chemical judgement.
  • Use scores to predict Kd or to claim compound A beats compound B by 0.5 kcal/mol. Do not.

The covalent problem, illustrated by this exact compound

Sotorasib works by forming a covalent bond between its acrylamide warhead and Cys12. That bond is the entire mechanism.

The CovaDock run scored it non-covalently. The scoring function has no term for bond formation. So -10.35 kcal/mol describes only the reversible recognition step that positions the warhead - it understates the real thermodynamics of the drug by a large and unquantified amount.

This cuts both ways and is worth internalising:

  • For a covalent series, the docking score is a proxy for warhead placement, not potency. Check whether the reactive atom sits near the target cysteine in the pose. That is the readout with information in it.
  • Comparing a covalent compound against a non-covalent one on docking score alone is meaningless. You are comparing two different physical quantities.

What the interaction count does and does not add

61 interactions sounds like strong evidence. It is useful, but not as a magnitude.

Interaction counts depend on the geometric cutoffs the analyser uses, and they scale with ligand size. A 560 Da molecule will make more contacts than a 300 Da fragment, whether or not it binds better. What the list is genuinely good for is identity: the contacts landed at Gly24, Ala25, Gly74 and Gln75, which is the switch-II region. That tells you the pose is in the right place. A high score with contacts scattered outside the known pocket is a red flag no number will surface for you.

Same logic applies to ligand efficiency. Sotorasib has 41 heavy atoms (from the C30H30F2N6O3 formula returned by covabasic_pubchem). LE = 10.35 / 41 = 0.25 kcal/mol per heavy atom. Respectable for a large late-stage molecule, unimpressive for a fragment. The raw score alone hides that entirely.

Honest limits

What a CovaDock score does not tell you:

  • Not a Kd, Ki or IC50. No conversion factor exists that makes it one.
  • No covalent chemistry. Warhead reactivity, k_inact and residence time are outside the model.
  • Receptor-dependent. The 90.2% docking power figure comes from curated crystal structures. Dock into a predicted model or a poorly resolved pocket and both pose and score degrade. Our run used crystal 6OIM deliberately.
  • Rigid-receptor by default. Induced fit, cryptic pockets and large loop rearrangements are not sampled.
  • Waters and ions dropped. CASF shows a 90.2% to 94.4% swing on that choice alone.
  • No selectivity, no ADMET, no developability. A -10 score says nothing about whether the compound is safe or gets anywhere near the target in vivo. That is a separate screen (covatox_structural_alerts, covatox_ich_m7).

CovaDock is triage. It narrows what you make and test. It does not replace a binding assay, and no docking output belongs in a regulatory filing as evidence of affinity.

Frequently asked questions

What does a -10 kcal/mol docking score mean?

It means the scoring function ranked that pose in its top bin - CovaDock labels anything below -10 kcal/mol "very strong". It does not mean the compound binds with a free energy of -10 kcal/mol. AutoDock Vina scores correlate with measured affinity at Pearson R = 0.604 with 1.73 log units of scatter on the CASF-2016 core set of 285 complexes, so treat it as a rank, not a measurement.

Can I convert a docking score to Kd or Ki?

No. The apparent conversion (1.36 kcal/mol per log unit of Kd at 298 K) is arithmetically valid but scientifically misleading, because the input carries roughly 2.4 kcal/mol of residual error. That propagates to nearly two orders of magnitude of uncertainty in Kd. Use docking to prioritise compounds for measurement, then measure.

Is a difference of 0.5 kcal/mol between two compounds real?

Almost certainly not. Vina's within-target ranking on CASF-2016 reaches Spearman rho = 0.528, and its scoring standard deviation is 1.73 log units. Differences smaller than roughly 2 kcal/mol should be treated as ties. Break those ties with pose quality, interaction location, ligand efficiency and synthetic accessibility rather than with the score.

Why is docking still useful if the scores are that noisy?

Because enrichment and affinity prediction are different tasks. Across a large library, even a moderately correlated score concentrates actives in the top-ranked fraction, which is all a screening funnel needs. Vina also reproduces the crystal pose within 2 A about 90% of the time on well-resolved structures, so the geometry it gives you is more trustworthy than the number.

Does CovaDock model covalent inhibitors?

No. The sotorasib run scored the compound non-covalently, so the -10.35 kcal/mol result excludes the covalent bond to Cys12 that drives its mechanism and understates the true thermodynamics. For covalent series, read the score as a measure of warhead positioning and check that the reactive atom sits near the target residue in the returned pose.

What should I look at besides the score?

The interaction list and its location. Our KRAS G12C run returned 61 contacts at Gly24, Ala25, Gly74 and Gln75, which is the switch-II pocket the drug class targets. A high score with contacts outside the known pocket usually means an artefact. Also check PoseBusters validity, which CovaDock runs on the top poses by default.

Related reading

Run your own target and ligand through covadock_dock on the free tier and read the score the way this article suggests: pose first, contacts second, number last.

Tools for this topic

Use these in your AI agent right away.

  • CovadockMolecular docking, binding energies.
Docking Score Interpretation: What -10 kcal/mol Means | CovaSyn