Skip to content
CovaSyn

AI for chemistry

AI for chemistry and drug discovery, the deterministic tool layer.

Frontier LLMs reach 14 to 41% on pure chemistry tasks. With a deterministic MCP tool layer attached: 76 to 92%, measured on MolecularIQ, a benchmark we did not design (arXiv:2601.15279). This page explains where AI for chemistry actually works today, where it doesn't, and how the Model Context Protocol (MCP) closes the gap.

What “AI for chemistry” actually means in 2026

The term gets used loosely. Three very different things sit inside it, and you shouldn't conflate them:

  • AI agents with chemistry tools. An LLM like Claude, GPT, or Gemini calls deterministic chemistry functions through a standard interface. This is the stack that went into production in 2025 and 2026. Details below, and on the MCP platform page.
  • LLMs that “do” chemistry themselves. Asking an untooled LLM for logP, pKa, or an ICH M7 triage. Works for intuitive answers, fails reproducibly on hard numerical or categorical questions. On MolecularIQ, the benchmark from the Klambauer Lab at JKU Linz (arXiv:2601.15279), four frontier models land at 14 to 41% in our evaluation.
  • Specialised models for chemistry (foundation models, fine-tuned LLMs). AlphaFold and RoseTTAFold for protein structures, BioGPT for biomedical text, ChemLLM as a language model for chemistry. Strong in narrowly defined areas but not an end-to-end research stack. Anyone deploying “AI for chemistry” combines these specialised models as tools, not as the only answer source.

The architecture that works in 2026 is LLM agent plus deterministic tool layer. The LLM understands the question and plans the steps. The tool layer computes. The two are kept separate. To connect them, Anthropic published the Model Context Protocol (MCP) as an open standard in 2024.

Three problems AI alone doesn't solve in chemistry

  • Hallucination in the value space

    The LLM makes up plausible but wrong numbers: logP 2.3 instead of 4.1, pKa “around 5” instead of 5.23. In a submission, such a value does not hold.

  • Missing audit trail

    Non-reproducible answers cannot be substantiated under EU GMP Annex 11 and 21 CFR Part 11. Pilots get caught out at the QA audit, if not sooner.

  • Inconsistency between runs

    Even at temperature 0, an LLM can give different answers on repetition. A validation protocol can't be closed out that way.

More on the three failure modes and why “better data management” is the wrong answer: Data quality isn't the real bottleneck.

How MCP for chemistry closes the gap

The Model Context Protocol is the standardised interface between an LLM and a tool function. An MCP server for chemistry exposes functions like covatox_assess_ichm7_batch or covastab_design_study as deterministic tools with a clearly defined input/output contract. The LLM calls them, gets a reproducible answer back, and either communicates it or chains it into another step.

Three consequences for pharma and chemistry teams

Deterministic tools, not LLM estimates
When the agent needs ADMET, a tox endpoint, or a stability prediction, it calls a deterministic function. The thousandth call returns the same result as the first.
Audit trail built in
Every tool call is logged with timestamp, tool, version and result status, and versions are pinned. Reproducibility is a technical property of the layer.
Standardised interface
The Model Context Protocol is an open standard supported by many clients, including Claude, ChatGPT, Cursor and VS Code. No vendor lock-in on the client side.

Background and technical detail on the MCP platform page and the market overview “The 5 leading chemistry MCP servers for pharma R&D compared”.

Foundation models and tool layers, two roads for AI in chemistry

Pharma AI is splitting into two separate architectural camps. Anyone making a platform decision should understand the difference, because it determines which workflows are reproducible and which are not.

Camp A

Foundation models

One large model computes everything itself: from text prompt to 3D structure to quantum properties to retrosynthesis, in one end-to-end stack. Examples: ChemLLM, AlphaFold successor models, multimodal pharma models.

Strong at:
Exploration of new compound classes, materials science, generative tasks where creativity matters more than reproducibility.
Weak at:
Audit trail, reproducibility, regulatory submissions. Same prompt, two runs, two answers.

Camp B (where CovaSyn sits)

LLM plus deterministic tool layer

A frontier LLM (Claude, GPT, Gemini) plans and communicates. A separate deterministic layer (CovaSyn MCP tools, RDKit, OpenMS) computes. Connected via the Model Context Protocol.

Strong at:
Reproducibility, support for validation, ICH M7 and Q1 workflows, regulated environments, cost at scale.
Weak at:
Pure exploration without clear tool contracts. If the answer doesn't come from an existing function, the layer returns nothing.

When which architecture

The two camps don't exclude each other, they cover different phases of the R&D lifecycle:

Early discovery, novel modalities
Foundation models. If you're working in white space, you need generative freedom.
Lead optimisation, tox triage, stability, submissions
Tool-layer architecture. When QA and auditors are in the loop, you need reproducibility.
Real pipelines
Both in parallel. The foundation model proposes a candidate, the tool layer recomputes it.

CovaSyn is explicitly the tool-layer half. We don't compete with foundation-model providers, we sit underneath, recomputing the proposals of such models in a reproducible form. More on practical use on the MCP platform page.

Tool areas at CovaSyn

Deterministic tools for NMR, MS, IR, UV, stability, toxicology, solubility and more, each callable from Claude Desktop, Cursor, VS Code or your own agent stack.

  • Cheminformatics (covabasic, covachem)

    SMILES, InChI and MOL handling, druglikeness, fingerprints, scaffold analysis, ADMET, pKa, tautomers: the standard layer for med-chem AI workflows.

  • Toxicology (covatox)

    ICH M7 batch assessment, Tox21 endpoints, structural alerts, CYP450, ecotoxicology. Built for regulatory triage.

  • Mass spectrometry (covams)

    Formula prediction, fragment annotation, impurity profiling, metabolite ID, retention-time prediction. Analysis with deterministic backends.

  • NMR (covnmr)

    1D and 2D NMR analysis, predict and assign, sudoku solver, verification, identification.

  • Stability (covastab)

    Arrhenius models, shelf-life estimation under ICH Q1E, OOS/OOT detection, batch variability.

  • Bio (covabio)

    Antibody profiling, peptide, ADC, mRNA, oligo, siRNA, immunogenicity, developability, for biotech workflows.

  • Folding and structure (covafold, covadock)

    Protein and RNA folding, binding sites, mutation analysis, docking. Its own tool layer, not just an AlphaFold wrapper.

  • DoE and optimisation (covadoe, covaopt)

    Design of experiments with agent guidance, response surface modelling, process optimisation, robust conditions, for CDMO workflows.

See the full tool list

Proof

Measured on a benchmark we did not design

On MolecularIQ from the Klambauer Lab, Institute for Machine Learning, JKU Linz, 3,540 verified chemistry tasks (arXiv:2601.15279): 4 frontier LLMs with and without CovaSyn MCP attached.

Without CovaSyn
14 to 41%
With CovaSyn MCP
76 to 92%
ModelWithout CovaSynWith CovaSyn MCP
Claude Haiku 4.521.2%85.4%
Claude Opus 4.740.8%91.5%
OpenAI GPT-5.522.3%89.9%
Gemini 3.5 Flash13.7%75.7%

For every model evaluated, the hit rate with CovaSyn is above the best value without CovaSyn. The gain comes from the deterministic tool layer, not from the model.

Source: Bartmann C., Schimunek J., Ielanskyi M., Seidl P., Klambauer G., Luukkonen S. (2026). MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs. arXiv:2601.15279. Snapshot: 2026-05-17.

Full methodology, gaps and reproduction steps

Real workflows with AI chemistry

  • AI for drug discovery

    The agent generates compound candidates, filters them with covachem_adme and covatox_assess_ichm7_batch, and hands the top hits to the med-chem team.

  • AI for stability studies

    The agent designs a stress scheme following ICH Q1A with covastab_design_study, estimates shelf life via covastab_estimate_shelf and checks it against real data.

  • AI for ICH M7 mutagenicity triage

    The agent screens an impurity library with covatox_assess_ichm7_batch, assigns classes 1 to 5 and returns an auditable justification per compound. Expert review remains in place.

  • AI for spectral analysis

    A mass spectrum comes in, the agent calls covams_identify and covams_fragment_annotate and returns a structured identification with a confidence score. NMR likewise with covnmr_identify.

Who this is built for

  • Pharma R&D teams

    Med-chem, ADMET, regulatory triage. If you already know ICH M7, Q1 and Annex 11, you know our vocabulary.

  • Biologics teams

    Antibodies, mRNA, ADCs, oligos. Bio tools with developability, immunogenicity, viscosity.

  • CDMOs and contract research

    Time-critical workflows from quote to release, where standard calculations should not wait in a queue.

  • Pharma AI engineering teams

    If you run your own Claude deployment or an agent stack and want a chemistry MCP server to plug in without building it yourself.

FAQ on AI for chemistry and drug discovery

What's the best AI stack for chemistry in 2026?

A frontier LLM (Claude, GPT, Gemini) with a deterministic MCP chemistry tool layer underneath. The LLM understands and plans, the MCP tools compute. On MolecularIQ (arXiv:2601.15279), a benchmark we did not design, LLMs without tools land at 14 to 41%, with CovaSyn MCP at 76 to 92%.

What's the difference between an LLM for chemistry and an AI agent for chemistry?

An LLM for chemistry generates text (including numbers) based on its training data: it doesn't compute, it remembers. An AI agent for chemistry uses an LLM as a language and planning layer and calls deterministic tools alongside it for the actual computation. The latter is reproducible and auditable, the former is not.

How does AI for drug discovery differ from classic cheminformatics?

Classic cheminformatics is a deterministic function with a clearly defined input and output (RDKit, Open Babel, OpenMS). AI for drug discovery layers an AI agent on top that orchestrates workflows, interprets data and communicates with humans. They need each other: the deterministic layer provides reproducible computed results, the agent makes the experience natural.

Can I use Claude or GPT directly for ICH M7 triage?

Without a tool layer, not for regulatory submissions: the answers are neither reproducible nor traceably substantiated. With a deterministic QSAR backend attached via MCP (for example CovaSyn covatox_assess_ichm7_batch), the answers are reproducible and auditable; expert review stays.

What data location does AI for chemistry need in EU pharma?

The GDPR does not prescribe a storage location, but it requires safeguards for transfers to third countries. Many pharma companies therefore prefer data held inside the EU, and many IT security teams ask for operation on their own infrastructure. CovaSyn offers both: hosting at Hetzner in Nuremberg and operation as a container in-house.

What does AI for chemistry cost in practice?

Scope ranges from evaluating individual workflows up to a company rollout operated in-house. We work this out individually in a 30-minute initial call, tailored to your workflow.

Next step

Make AI verifiable in your chemistry.

Create an account, get your API key, use the tools in Claude, ChatGPT or Cursor.

AI for chemistry: the deterministic tool layer | CovaSyn