MolecularIQ, arXiv:2601.15279
Benchmark and methodology
We measure CovaSyn on MolecularIQ, a benchmark we did not design (Bartmann et al., Klambauer Lab, Institute for Machine Learning, JKU Linz). Frontier models score 14 to 41% on chemical structure analysis there; with CovaSyn MCP attached they reach 76 to 92%. This page sets out the numbers, the methodology and the limits of the measurement.
Create an account, get your API key, use the tools in Claude, ChatGPT or Cursor.
The result
The same models, once without and once with CovaSyn.
4 models, 3,540 verified tasks, 12,540 responses. Only an exact match with the stored solution counts.
- Models alone
- 14 to 41%
- With CovaSyn MCP
- 76 to 92%
Source: Bartmann C., Schimunek J., Ielanskyi M., Seidl P., Klambauer G., Luukkonen S. (2026). MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs. arXiv:2601.15279. Snapshot: 2026-05-17.
Results per model
The effect is not tied to one model.
Baseline means the same model without tools. The MCP column is the same model with the CovaSyn MCP server attached. Haiku 4.5 was evaluated on the full test split; the other three models on a proportionally stratified subset.
| Model | Baseline | With CovaSyn MCP | Difference (percentage points) | Factor |
|---|---|---|---|---|
| Claude Haiku 4.5 | 21.18% | 85.38% | +64.20 | 4.03× |
| Claude Opus 4.7 | 40.75% | 91.51% | +50.76 | 2.25× |
| OpenAI GPT-5.5 | 22.29% | 89.92% | +67.63 | 4.03× |
| Gemini 3.5 Flash | 13.68% | 75.66% | +61.98 | 5.53× |
Source: Bartmann C., Schimunek J., Ielanskyi M., Seidl P., Klambauer G., Luukkonen S. (2026). MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs. arXiv:2601.15279. Snapshot: 2026-05-17.
Four models, three providers, the same direction. The effect of the integration is not tied to a specific model but reproducible across frontier models. If you work with Claude today and switch to GPT-5.5 or Gemini tomorrow, the effect carries over.
Methodology
What was measured, and how.
Benchmarks a vendor designs for itself deserve scepticism: whoever tests their own tools has an incentive to look good. That is why we measure on a benchmark we did not design and whose task selection we had no influence over.
- Independent benchmark
- MolecularIQ was created by Bartmann et al. (Klambauer Lab, Institute for Machine Learning, JKU Linz) and published as a preprint on arXiv (arXiv:2601.15279). CovaSyn neither designed nor contributed to the benchmark; we measured our tools against it.
- Symbolic verification
- Every answer is verified symbolically against the stored solution: exact match or wrong. No language model acting as judge, no partial credit, no room for interpretation. That strictness makes the result robust against cherry-picking, and it matters more than any single number.
- Tasks
- We evaluated the test split of MolecularIQ with 3,540 tasks, restricted to question types that can be verified symbolically.
- Models and scope
- 4 models: Claude Haiku 4.5, Claude Opus 4.7, OpenAI GPT-5.5 and Gemini 3.5 Flash, each without tools and with CovaSyn MCP. 12,540 model answers evaluated in total (one answer without and one with CovaSyn MCP per task and model); Haiku 4.5 on the full test split, the other three on a stratified sample of 910 questions per model to keep cost reasonable.
- Tools
- Chemistry primitives from the CovaBasicChem suite: deterministic cheminformatics that returns exact values, which the model then uses in its answer.
- Snapshot and reproducibility
- Data snapshot of 2026-05-17. Dataset and evaluation code are public and a CovaSyn account is free, so every number on this page can be recomputed.
Paper, data, code
- Paper on arXiv (abstract)https://arxiv.org/abs/2601.15279
- Paper PDF on arXivhttps://arxiv.org/pdf/2601.15279
- Dataset on Hugging Facehttps://huggingface.co/datasets/ml-jku/moleculariq-v0.0
- Evaluation code on GitHubhttps://github.com/ml-jku/moleculariq
The PDF is hosted by arXiv, not on our servers. Dataset and code belong to the benchmark's authors.
Citation
If you use the numbers on this page in a publication, a poster or an internal report, please cite the authors' MolecularIQ paper:
@misc{bartmann2026moleculariq,
title = {MolecularIQ: Characterizing Chemical Reasoning Capabilities
Through Symbolic Verification on Molecular Graphs},
author = {Bartmann, C. and Schimunek, J. and Ielanskyi, M. and
Seidl, P. and Klambauer, G. and Luukkonen, S.},
year = {2026},
eprint = {2601.15279},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2601.15279}
}Bartmann C., Schimunek J., Ielanskyi M., Seidl P., Klambauer G., Luukkonen S. (2026). MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs. arXiv:2601.15279.
In depth
What this means in cost terms.
Accuracy is one half, cost per question is the other. Teams that run chemistry queries at scale do not think in single tokens but in thousands of tool calls per week. With CovaSyn you can often run the cheaper model without giving up accuracy.
| Configuration | Accuracy | US dollars per question | Latency |
|---|---|---|---|
| Haiku 4.5 without tools | 21.18% | 0.00069 | 2.1 s |
| Haiku 4.5 with CovaSyn MCP | 85.38% | 0.00781 | 5.8 s |
| Opus 4.7 without tools | 40.75% | 0.02529 | 5.1 s |
| Opus 4.7 with CovaSyn MCP | 91.51% | 0.12536 | 7.4 s |
| GPT-5.5 without tools | 22.29% | 0.02750 | 7.9 s |
| GPT-5.5 with CovaSyn MCP | 89.92% | 0.03005 | 9.4 s |
| Gemini 3.5 Flash without tools | 13.68% | 0.00940 | 5.5 s |
| Gemini 3.5 Flash with CovaSyn MCP | 75.66% | 0.02170 | 10.8 s |
The point in one sentence
Haiku 4.5 with CovaSyn reaches more than twice the accuracy of Opus 4.7 without tools at roughly a third of the cost, and is about 16 times cheaper than Opus 4.7 with CovaSyn while giving up 6 percentage points of accuracy. Gemini 3.5 Flash with CovaSyn shows the largest relative lift (5.53 times, from 13.7 to 75.7%) at roughly 2.3 times its baseline cost and twice its baseline latency; the obvious option if you already run Gemini.
At ten thousand questions a month, Opus 4.7 with CovaSyn costs roughly 1,250 US dollars, Haiku 4.5 with CovaSyn roughly 78 US dollars and Gemini 3.5 Flash with CovaSyn roughly 217 US dollars.
Categories
Where the integration has the strongest effect.
Mean lift per question category, averaged over Haiku 4.5, Opus 4.7 and GPT-5.5. The current snapshot has no per-category values for Gemini 3.5 Flash.
| Category | Baseline | With CovaSyn MCP | Difference (percentage points) |
|---|---|---|---|
| Scaffold and fragments | 18.0% | 86.5% | +68.4 |
| Rings and topology | 29.4% | 93.2% | +63.8 |
| Bonds and chains | 17.6% | 80.9% | +63.3 |
| Multi-feature questions | 27.3% | 88.4% | +61.1 |
| Atom and formula counts | 38.7% | 98.3% | +59.7 |
| Stereochemistry | 28.7% | 86.0% | +57.4 |
| Electronics and H-bonds | 31.2% | 81.5% | +50.3 |
These are the building blocks of daily medicinal chemistry. A scaffold has to be identified exactly, a ring analysis must not drift into guessing, and a multi-constraint query such as “all branch points within two bonds of a halogen” has to resolve exactly.
Limits
Where the remaining gap sits.
Accuracy is not 100%, and we do not hide that. We assigned every wrong answer with CovaSyn MCP to one of three causes. This is how the wrong answers break down per model, and where it pays to look closer in your own validation.
| Cause | Haiku 4.5 | Opus 4.7 | GPT-5.5 | Gemini 3.5 Flash |
|---|---|---|---|---|
| Tool result overridden | 81% | 86% | 66% | 25% |
| Format error | 1% | 1% | 25% | 73% |
| Tool value off | 18% | 13% | 9% | 3% |
- Tool result overridden
- The model ignores a correct tool result and produces its own answer. This behaviour depends on the model, not the tool. Clearer prompt structure reduces it noticeably but does not remove it entirely.
- Tool value off
- The tool value does not fit the question. These are real workflow errors, mostly edge cases with complex stereochemistry or unusual structures. We work on them continuously.
- Format error
- The model output cannot be parsed. This class is technically easy to catch because an output check surfaces it immediately.
At most 18% of a model's wrong answers come from a tool value that does not fit. The rest arises in the model: it overrides a correct result or returns output that cannot be parsed. Both can be caught in the workflow.
Implications
What this means for research and development.
- The model is no longer automatically the bottleneck.
- Teams paying about 0.03 US dollars per chemistry query today can move to about 0.008 US dollars at equal or better accuracy. That noticeably changes the cost of medicinal chemistry workflows.
- The lift is model-independent.
- Building an MCP layer does not lock you to Anthropic, OpenAI or Google. When the next frontier model arrives, the layer goes directly in front of it and the effect carries over.
- The gap is disclosed.
- For validation in regulated environments an open error distribution is worth more than a polished promise. Your QA team can derive risk assessment and acceptance criteria directly from the failure classes.
FAQ
Frequently asked questions
What is the MolecularIQ benchmark?
A benchmark published in 2026 by the Klambauer Lab, Institute for Machine Learning, JKU Linz (arXiv:2601.15279) that measures how well language models solve structured chemistry analysis. 3,540 tasks in the evaluated test split, symbolic verification without a language model acting as judge.
Why does accuracy rise so much with CovaSyn MCP?
Language models often guess on structural chemistry tasks such as atom counts or scaffold extraction because they lack a deterministic mechanism. CovaSyn returns exact values from cheminformatics libraries via tool calls, which the model then uses in its answer.
How can Haiku 4.5 with CovaSyn be cheaper than Opus 4.7 without?
Haiku 4.5 is the much smaller model: in this measurement a question without tools cost about 0.0007 US dollars, with Opus 4.7 about 0.025. The tool calls lift Haiku to about 0.008 US dollars per question, roughly a third of Opus 4.7 without tools. Because the MCP layer supplies structural correctness, the small model is left with phrasing and reasoning. Result: about 85 instead of 41% accuracy.
Can the result be reproduced?
Yes. The dataset is public on Hugging Face, the evaluation code is on GitHub, the models are available through their official APIs, and you can create a CovaSyn MCP account for free.
Why does Gemini 3.5 Flash post the lowest score but the largest lift?
Gemini 3.5 Flash starts at 13.68%, the lowest of the four models, because it is built for speed and cost rather than structural chemistry. That is exactly why its relative lift, a factor of 5.53, is the largest: the MCP layer supplies the structural correctness the model lacks on its own. At 0.02170 US dollars per question, Gemini with CovaSyn beats Opus 4.7 without tools by more than 30 percentage points, at clearly lower cost than Opus 4.7 with CovaSyn.
What about the other CovaSyn tools?
The numbers on this page come from one part of our tools. We are measuring the others against matching datasets. Once results are in, they will appear on this page, with full methodology and the full distribution.
Research and development
The development of the CovaSyn MCP platform is certified as research and development under the German Research Allowance Act by the certification body BSFZ.
The certificate assesses the research character of the development work. It is not a statement about product properties and not a state product approval. Bescheinigungsstelle Forschungszulage.
Next step
Measure it yourself.
Test the tools behind these numbers on your own molecule, free.
Create an account, get your API key, use the tools in Claude, ChatGPT or Cursor.
