Skip to content
CovaSyn
Try the MCP free

MolecularIQ, arXiv:2601.15279

Benchmark and methodology

We measure CovaSyn on MolecularIQ, a benchmark we did not design (Bartmann et al., Klambauer Lab, Institute for Machine Learning, JKU Linz). Frontier models score 14 to 41% on chemical structure analysis there; with CovaSyn MCP attached they reach 76 to 92%. This page sets out the numbers, the methodology and the limits of the measurement.

Create an account, get your API key, use the tools in Claude, ChatGPT or Cursor.

Result per modelMolecularIQ, exact match
Model aloneWith CovaSyn MCP0%25%50%75%100%Haiku 4.5+64.2 pp · 4.0×Haiku 4.5, model alone: 21.18%Haiku 4.5, with CovaSyn MCP: 85.38%21.2%85.4%Opus 4.7+50.8 pp · 2.2×Opus 4.7, model alone: 40.75%Opus 4.7, with CovaSyn MCP: 91.51%40.8%91.5%GPT-5.5+67.6 pp · 4.0×GPT-5.5, model alone: 22.29%GPT-5.5, with CovaSyn MCP: 89.92%22.3%89.9%Gemini 3.5 Flash+62.0 pp · 5.5×Gemini 3.5 Flash, model alone: 13.68%Gemini 3.5 Flash, with CovaSyn MCP: 75.66%13.7%75.7%
Accuracy on the test split, symbolically verified: Haiku 4.5 on all 3,540 tasks, the other models on 910 each. Same questions without and with CovaSyn. pp = percentage points.

The result

The same models, once without and once with CovaSyn.

4 models, 3,540 verified tasks, 12,540 responses. Only an exact match with the stored solution counts.

Models alone
14 to 41%
With CovaSyn MCP
76 to 92%
MolecularIQ, arXiv:2601.15279, on arXiv

Source: Bartmann C., Schimunek J., Ielanskyi M., Seidl P., Klambauer G., Luukkonen S. (2026). MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs. arXiv:2601.15279. Snapshot: 2026-05-17.

Results per model

The effect is not tied to one model.

Baseline means the same model without tools. The MCP column is the same model with the CovaSyn MCP server attached. Haiku 4.5 was evaluated on the full test split; the other three models on a proportionally stratified subset.

Accuracy per model without and with CovaSyn MCP
ModelBaselineWith CovaSyn MCPDifference (percentage points)Factor
Claude Haiku 4.521.18%85.38%+64.204.03×
Claude Opus 4.740.75%91.51%+50.762.25×
OpenAI GPT-5.522.29%89.92%+67.634.03×
Gemini 3.5 Flash13.68%75.66%+61.985.53×

Source: Bartmann C., Schimunek J., Ielanskyi M., Seidl P., Klambauer G., Luukkonen S. (2026). MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs. arXiv:2601.15279. Snapshot: 2026-05-17.

Four models, three providers, the same direction. The effect of the integration is not tied to a specific model but reproducible across frontier models. If you work with Claude today and switch to GPT-5.5 or Gemini tomorrow, the effect carries over.

Methodology

What was measured, and how.

Benchmarks a vendor designs for itself deserve scepticism: whoever tests their own tools has an incentive to look good. That is why we measure on a benchmark we did not design and whose task selection we had no influence over.

Independent benchmark
MolecularIQ was created by Bartmann et al. (Klambauer Lab, Institute for Machine Learning, JKU Linz) and published as a preprint on arXiv (arXiv:2601.15279). CovaSyn neither designed nor contributed to the benchmark; we measured our tools against it.
Symbolic verification
Every answer is verified symbolically against the stored solution: exact match or wrong. No language model acting as judge, no partial credit, no room for interpretation. That strictness makes the result robust against cherry-picking, and it matters more than any single number.
Tasks
We evaluated the test split of MolecularIQ with 3,540 tasks, restricted to question types that can be verified symbolically.
Models and scope
4 models: Claude Haiku 4.5, Claude Opus 4.7, OpenAI GPT-5.5 and Gemini 3.5 Flash, each without tools and with CovaSyn MCP. 12,540 model answers evaluated in total (one answer without and one with CovaSyn MCP per task and model); Haiku 4.5 on the full test split, the other three on a stratified sample of 910 questions per model to keep cost reasonable.
Tools
Chemistry primitives from the CovaBasicChem suite: deterministic cheminformatics that returns exact values, which the model then uses in its answer.
Snapshot and reproducibility
Data snapshot of 2026-05-17. Dataset and evaluation code are public and a CovaSyn account is free, so every number on this page can be recomputed.

Paper, data, code

The PDF is hosted by arXiv, not on our servers. Dataset and code belong to the benchmark's authors.

Citation

If you use the numbers on this page in a publication, a poster or an internal report, please cite the authors' MolecularIQ paper:

BibTeX
@misc{bartmann2026moleculariq,
  title         = {MolecularIQ: Characterizing Chemical Reasoning Capabilities
                   Through Symbolic Verification on Molecular Graphs},
  author        = {Bartmann, C. and Schimunek, J. and Ielanskyi, M. and
                   Seidl, P. and Klambauer, G. and Luukkonen, S.},
  year          = {2026},
  eprint        = {2601.15279},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2601.15279}
}
Plain text
Bartmann C., Schimunek J., Ielanskyi M., Seidl P., Klambauer G., Luukkonen S. (2026). MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs. arXiv:2601.15279.

In depth

What this means in cost terms.

Accuracy is one half, cost per question is the other. Teams that run chemistry queries at scale do not think in single tokens but in thousands of tool calls per week. With CovaSyn you can often run the cheaper model without giving up accuracy.

Accuracy, cost per question and latency per configuration
ConfigurationAccuracyUS dollars per questionLatency
Haiku 4.5 without tools21.18%0.000692.1 s
Haiku 4.5 with CovaSyn MCP85.38%0.007815.8 s
Opus 4.7 without tools40.75%0.025295.1 s
Opus 4.7 with CovaSyn MCP91.51%0.125367.4 s
GPT-5.5 without tools22.29%0.027507.9 s
GPT-5.5 with CovaSyn MCP89.92%0.030059.4 s
Gemini 3.5 Flash without tools13.68%0.009405.5 s
Gemini 3.5 Flash with CovaSyn MCP75.66%0.0217010.8 s

The point in one sentence

Haiku 4.5 with CovaSyn reaches more than twice the accuracy of Opus 4.7 without tools at roughly a third of the cost, and is about 16 times cheaper than Opus 4.7 with CovaSyn while giving up 6 percentage points of accuracy. Gemini 3.5 Flash with CovaSyn shows the largest relative lift (5.53 times, from 13.7 to 75.7%) at roughly 2.3 times its baseline cost and twice its baseline latency; the obvious option if you already run Gemini.

At ten thousand questions a month, Opus 4.7 with CovaSyn costs roughly 1,250 US dollars, Haiku 4.5 with CovaSyn roughly 78 US dollars and Gemini 3.5 Flash with CovaSyn roughly 217 US dollars.

Cost and accuracyUS dollars per question
Model aloneWith CovaSyn MCP02550751000.0010.010.1US dollars per question (log scale)Accuracy in %2.1× accuracyat 31% of the costHaiku 4.5, alone: 21.18%, 0.00069 USDHaiku 4.5, with CovaSyn: 85.38%, 0.00781 USDHaiku 4.5Haiku 4.5Opus 4.7, alone: 40.75%, 0.02529 USDOpus 4.7, with CovaSyn: 91.51%, 0.12536 USDOpus 4.7Opus 4.7GPT-5.5, alone: 22.29%, 0.02750 USDGPT-5.5, with CovaSyn: 89.92%, 0.03005 USDGPT-5.5GPT-5.5Gemini 3.5 Flash, alone: 13.68%, 0.00940 USDGemini 3.5 Flash, with CovaSyn: 75.66%, 0.02170 USDGemini 3.5 FlashGemini 3.5 Flash
Top left is good: high accuracy at low cost. Each line joins the same model without and with CovaSyn. Dashed: Haiku 4.5 with CovaSyn versus Opus 4.7 alone.

Categories

Where the integration has the strongest effect.

Mean lift per question category, averaged over Haiku 4.5, Opus 4.7 and GPT-5.5. The current snapshot has no per-category values for Gemini 3.5 Flash.

Mean accuracy per question category without and with CovaSyn MCP
CategoryBaselineWith CovaSyn MCPDifference (percentage points)
Scaffold and fragments18.0%86.5%+68.4
Rings and topology29.4%93.2%+63.8
Bonds and chains17.6%80.9%+63.3
Multi-feature questions27.3%88.4%+61.1
Atom and formula counts38.7%98.3%+59.7
Stereochemistry28.7%86.0%+57.4
Electronics and H-bonds31.2%81.5%+50.3

These are the building blocks of daily medicinal chemistry. A scaffold has to be identified exactly, a ring analysis must not drift into guessing, and a multi-constraint query such as “all branch points within two bonds of a halogen” has to resolve exactly.

Lift per categoryMean of three models
Model aloneWith CovaSyn MCP0%25%50%75%100%Scaffold and fragments+68.4 ppScaffold and fragments, model alone: 18.00%Scaffold and fragments, with CovaSyn MCP: 86.50%18.0%86.5%Rings and topology+63.8 ppRings and topology, model alone: 29.40%Rings and topology, with CovaSyn MCP: 93.20%29.4%93.2%Bonds and chains+63.3 ppBonds and chains, model alone: 17.60%Bonds and chains, with CovaSyn MCP: 80.90%17.6%80.9%Multi-feature questions+61.1 ppMulti-feature questions, model alone: 27.30%Multi-feature questions, with CovaSyn MCP: 88.40%27.3%88.4%Atom and formula counts+59.7 ppAtom and formula counts, model alone: 38.70%Atom and formula counts, with CovaSyn MCP: 98.30%38.7%98.3%Stereochemistry+57.4 ppStereochemistry, model alone: 28.70%Stereochemistry, with CovaSyn MCP: 86.00%28.7%86.0%Electronics and H-bonds+50.3 ppElectronics and H-bonds, model alone: 31.20%Electronics and H-bonds, with CovaSyn MCP: 81.50%31.2%81.5%
Averaged over Haiku 4.5, Opus 4.7 and GPT-5.5, sorted by gain.

Limits

Where the remaining gap sits.

Accuracy is not 100%, and we do not hide that. We assigned every wrong answer with CovaSyn MCP to one of three causes. This is how the wrong answers break down per model, and where it pays to look closer in your own validation.

Wrong answersShare per cause
Each bar = all wrong answers of one model with CovaSyn (100%). The failure analysis is a separate classification with its own counting basis; for accuracy, the table above applies.
Composition of wrong answers with CovaSyn MCP per model, share of all wrong answers
CauseHaiku 4.5Opus 4.7GPT-5.5Gemini 3.5 Flash
Tool result overridden81%86%66%25%
Format error1%1%25%73%
Tool value off18%13%9%3%
Tool result overridden
The model ignores a correct tool result and produces its own answer. This behaviour depends on the model, not the tool. Clearer prompt structure reduces it noticeably but does not remove it entirely.
Tool value off
The tool value does not fit the question. These are real workflow errors, mostly edge cases with complex stereochemistry or unusual structures. We work on them continuously.
Format error
The model output cannot be parsed. This class is technically easy to catch because an output check surfaces it immediately.

At most 18% of a model's wrong answers come from a tool value that does not fit. The rest arises in the model: it overrides a correct result or returns output that cannot be parsed. Both can be caught in the workflow.

Implications

What this means for research and development.

The model is no longer automatically the bottleneck.
Teams paying about 0.03 US dollars per chemistry query today can move to about 0.008 US dollars at equal or better accuracy. That noticeably changes the cost of medicinal chemistry workflows.
The lift is model-independent.
Building an MCP layer does not lock you to Anthropic, OpenAI or Google. When the next frontier model arrives, the layer goes directly in front of it and the effect carries over.
The gap is disclosed.
For validation in regulated environments an open error distribution is worth more than a polished promise. Your QA team can derive risk assessment and acceptance criteria directly from the failure classes.

FAQ

Frequently asked questions

What is the MolecularIQ benchmark?

A benchmark published in 2026 by the Klambauer Lab, Institute for Machine Learning, JKU Linz (arXiv:2601.15279) that measures how well language models solve structured chemistry analysis. 3,540 tasks in the evaluated test split, symbolic verification without a language model acting as judge.

Why does accuracy rise so much with CovaSyn MCP?

Language models often guess on structural chemistry tasks such as atom counts or scaffold extraction because they lack a deterministic mechanism. CovaSyn returns exact values from cheminformatics libraries via tool calls, which the model then uses in its answer.

How can Haiku 4.5 with CovaSyn be cheaper than Opus 4.7 without?

Haiku 4.5 is the much smaller model: in this measurement a question without tools cost about 0.0007 US dollars, with Opus 4.7 about 0.025. The tool calls lift Haiku to about 0.008 US dollars per question, roughly a third of Opus 4.7 without tools. Because the MCP layer supplies structural correctness, the small model is left with phrasing and reasoning. Result: about 85 instead of 41% accuracy.

Can the result be reproduced?

Yes. The dataset is public on Hugging Face, the evaluation code is on GitHub, the models are available through their official APIs, and you can create a CovaSyn MCP account for free.

Why does Gemini 3.5 Flash post the lowest score but the largest lift?

Gemini 3.5 Flash starts at 13.68%, the lowest of the four models, because it is built for speed and cost rather than structural chemistry. That is exactly why its relative lift, a factor of 5.53, is the largest: the MCP layer supplies the structural correctness the model lacks on its own. At 0.02170 US dollars per question, Gemini with CovaSyn beats Opus 4.7 without tools by more than 30 percentage points, at clearly lower cost than Opus 4.7 with CovaSyn.

What about the other CovaSyn tools?

The numbers on this page come from one part of our tools. We are measuring the others against matching datasets. Once results are in, they will appear on this page, with full methodology and the full distribution.

Seal of the German research allowance certification body, research and development 2026

Research and development

The development of the CovaSyn MCP platform is certified as research and development under the German Research Allowance Act by the certification body BSFZ.

The certificate assesses the research character of the development work. It is not a statement about product properties and not a state product approval. Bescheinigungsstelle Forschungszulage.

Security and privacy

Next step

Measure it yourself.

Test the tools behind these numbers on your own molecule, free.

Create an account, get your API key, use the tools in Claude, ChatGPT or Cursor.

Benchmark and methodology | CovaSyn