
4chems vs. "Just Another IUCLID Platform": How We Think About AI and Cloud Risk
The debate around cloud and AI tools in REACH compliance usually revolves around one word: risk.
Read More →Have questions? We're here to help.
A Eurotox 2026 study found LLMs agree more consistently than expert toxicologists do. Here is what that means for legacy IUCLID dossier data.

Have questions? We're here to help.
Toxicology runs on some of the strictest lab standards in science: OECD guidelines, GLP protocols, validated assays. A study presented at Eurotox 2026 found that the inconsistent part isn’t the lab work. It’s the reading.
At the HESI CITE Lecture, Jess Ewald’s team at EMBL-EBI faced a familiar problem: over 100,000 free-text toxicology findings, collected across decades of animal studies from different labs, in different formats, under different conventions. All of it needed to be mapped onto one standardized ontology (the Mammalian Phenotype Ontology) before it could be reused at scale.
Manual mapping at that volume isn’t realistic, so the team built an LLM-assisted pipeline and tested it against seven expert toxicologists doing the same mapping by hand.
Measured in mean pairwise agreement (Jaccard similarity):
| Comparison | Agreement |
|---|---|
| LLM ↔ LLM | 0.85 |
| Human ↔ LLM | 0.60 |
| Human ↔ Human | 0.55 |
Every difference was significant at p < 0.001. Every individual model landed closer to the average human answer than any single human expert did to their own colleagues.
The unreliable readers weren’t the models. They were the experts.
We spend a lot of energy asking whether AI is reliable enough for regulatory work. This result asks the harder question the other way: how reliable is the expert judgment that existing dossiers are already built on?
That question isn’t abstract for anyone managing a legacy IUCLID instance. Every dossier accumulated over years holds the same kind of material: endpoint summaries and classification rationale, written by different reviewers, in different years, in slightly different words. None of it is technically wrong. Very little of it agrees perfectly with the record next to it.
Nobody flags this as an error, because it isn’t one — it’s drift. And drift stays invisible until someone has to read two files side by side and decide which version is actually correct. During a migration, an audit, or a submission review, that decision has real consequences.
This is precisely the gap that structured, well-managed infrastructure is built to close. A single authoritative dossier per substance, current picklists, and one record that flows outward instead of getting re-keyed at every step reduces the number of places drift can hide in the first place. It doesn’t replace expert judgment — it gives that judgment a consistent record to work from.
If you’re evaluating how much inconsistency is sitting in your own legacy dossiers, that’s exactly the kind of review we do before a migration to 4chems.com. Talk to us about what a structured, current IUCLID environment would surface in your data.
Source: Jess Ewald, HESI CITE Lecture, “Connecting chemical effects across biological scales with AI,” Eurotox 2026, Vienna.

The debate around cloud and AI tools in REACH compliance usually revolves around one word: risk.
Read More →
K-REACH. KKDİK. UK REACH. Four jurisdictions that looked at EU REACH and built their own version — each with its own …
Read More →
One of the questions regulatory affairs teams keep running into — usually right after something went wrong: do we need a …
Read More →