AI and Analysis-ready Data Data Curation and Management AI Solutions and Augmentation Bioinformatics & Data Science

Onboarding Your AI: Why Agents Need Ontologies Like Employees Need Training

I saw this posted recently on Reddit. It's an AI-generated poster spotted in a London restaurant. At first glance it looks fine, but upon closer inspection...not so much. Anyone for a Spaghett?

The "Confidently Wrong" Problem

LLMs are trained on vast amounts of web data, which makes them great at sounding authoritative. But life sciences content is sparse in that training mix. So they've learned to speak the language without necessarily understanding the science.

For a highly regulated industry like pharma where patient safety, regulatory compliance, and scientific reproducibility are on the line, "sounds right" is potentially dangerous.

Why Agentic AI + Ontologies Changes Everything

The solution isn't ever larger, more complex models - it's grounding them in reality.

Agentic AI uses specialized agents working together in orchestrated workflows, often leveraging existing tools. But the special sauce is structured knowledge in the form of ontologies.

Ontologies codify what is known. In life sciences, we are already lucky enough to have rich ontologies covering everything from chemical structures to disease classifications. For AI agents, these can serve as:

- Training resources grounded in verified knowledge
- Reference materials to verify accuracy
- Frameworks for understanding business structure and data assets

Think of it Like Onboarding

Just like a new employee needs to understand your organization—departments, locations, therapeutic portfolio, data repositories—AI agents need that same grounding to work effectively.

Without it, they're generating nonsense pasta shapes.
With it, they're making contextually appropriate decisions anchored in your actual business reality.

This intersection of agentic AI and ontologies is at the heart of what we do at Rancho Biosciences—we help pharma organizations adopt agentic AI systems while building and maintaining the ontologies and structured knowledge that ensure accuracy and trustworthiness.

---

You can hear more at my presentation at the Pistoia USA Conference next week in Boston, hope to see you there! Pistoia Alliance Rancho BioSciences https://lnkd.in/e7kVNFKw

Frequently Asked Questions

What causes AI hallucinations in life sciences content?
Hallucinations happen when a model has learned the linguistic patterns of a domain without the underlying factual structure. Life sciences content is sparse relative to general web text in most training data, so models can reproduce the vocabulary of biology, chemistry, and clinical research while generating claims that no source supports.
Why do large language models sound authoritative even when they are wrong?
LLMs are optimized to produce fluent, confident text, not to signal uncertainty. Fluency and accuracy are separate properties, so a model can return a well-formed sentence with a fabricated gene target, dosage, or trial result. This is why output that "sounds right" is not evidence that it is right.
What is an ontology in life sciences?
An ontology is a structured, machine-readable model of a domain that defines entities and the relationships between them. Life sciences ontologies cover chemical structures, anatomy, diseases, phenotypes, and experimental factors, giving software a shared vocabulary and a formal map of how concepts connect.
How do ontologies reduce AI hallucinations?
Ontologies ground models in verified knowledge rather than statistical association. They serve three roles for AI agents: a training resource anchored in curated facts, a reference layer for validating generated claims against known relationships, and a framework for representing an organization's own structure and data assets.
What are examples of widely used life sciences ontologies?
Commonly used resources include ChEBI for chemical entities, MONDO and the Disease Ontology for disease classification, Uberon for anatomy, HPO for phenotypes, EFO for experimental factors, NCIt for oncology terminology, and MeSH for biomedical indexing. Most organizations also need internal ontologies mapped to these public standards.
Can larger models solve the hallucination problem on their own?
Scale alone does not solve it. Increasing model size improves fluency and general reasoning, but a model with no connection to authoritative sources still has no mechanism for verifying a biomedical claim. Accuracy comes from grounding, retrieval, and validation against structured knowledge.
Next Step

Talk To Our Team

Get In Touch
Jane Lomax
About The Author
Jane Lomax
Ontologies | Terminologies | Data Standards | Data Modelling | Life Science | Biocuration | Professional Services | PhD
LinkedIn
Tags: #AI and Analysis-ready Data #Data Curation and Management #AI Solutions and Augmentation #Bioinformatics & Data Science
Share: LinkedIn X Facebook

Comments