Health
AI System Trained on Pathologists’ Search Patterns Achieves 100%
A new AI system named “Pathology-o3”, trained on pathologists’ visual search behavior—not just diagnostic outcomes—achieved 100% sensitivity in detecting colorectal cancer in lymph node tissue slides, though with a 15.5% false-positive rate.

A newly developed artificial intelligence system—trained to emulate how pathologists visually scan tissue slides rather than merely learning from final diagnostic labels—demonstrated enhanced ability to identify suspicious regions within biopsy samples, according to a recently published study.
How Pathologists Search—And Why It Matters
Most existing AI tools in digital pathology divide tissue slides into fixed-size tiles for analysis. In contrast, human pathologists adopt a dynamic, multi-scale approach: they first survey the entire slide at low magnification, then adjust zoom levels and navigate flexibly across regions before focusing intensively on areas appearing abnormal. This method is critical because a single tissue slide may contain billions of pixels, while cancerous indicators can occupy only a minuscule fraction of that space.
Chih-Huang, an assistant professor of pathology and laboratory medicine at the University of Pennsylvania and a co-author of the study, compared this process to a helicopter searching for missing persons: clinicians do not begin by scrutinizing a narrow patch of ground but instead conduct an initial broad overview before narrowing in on specific locations.
Building “Pathology-o3” From Real-World Behavior
Huang and his colleagues designed a training framework that captures pathologists’ actual scanning behavior—not just their end diagnoses. They collected eye-tracking and navigation data from eight board-certified pathologists as they examined lymph node tissue slides for signs of colorectal cancer. Researchers recorded movement patterns across slides and magnification adjustments used during active search for suspicious areas.
They filtered out incidental movements and isolated behaviors indicating deliberate attention—such as prolonged dwell time or repeated revisiting of a region. To validate alignment with visual focus, the team cross-referenced these motion patterns with independent eye-tracking data. Additionally, the AI model was required to generate brief textual explanations for why each flagged region was deemed significant; pathologists were then invited to accept, modify, or reject those rationales.
Using this behavioral dataset, the researchers built an AI system they named “Pathology-o3”. The system operates in stages: it first scans the full slide at low resolution, identifies candidate regions warranting closer inspection, and then directs high-resolution imaging of those specific zones to a secondary AI model for detailed analysis.
Performance Against OpenAI’s “o3” Model
The team tested “Pathology-o3” on tissue slides from patients diagnosed with colorectal cancer, specifically focusing on lymph node specimens. Its performance was benchmarked against OpenAI’s “o3” model.
In the primary test set, “Pathology-o3” correctly identified all slides containing cancer—achieving 100% sensitivity for positive cases. However, 15.5% of slides classified as cancer-positive were, in fact, negative. By comparison, the “o3” model achieved 87.5% sensitivity but misclassified 53.3% of its positive predictions as false positives.
When evaluated on an independent, unseen dataset, “Pathology-o3” maintained 97.6% sensitivity for detecting cancerous slides, though 37.1% of its positive calls were false positives.
Interpreting the False-Positive Rate
Researchers noted that the elevated false-positive rate likely reflects the system’s design priority: minimizing missed cancers—even at the cost of over-flagging regions for further review. The architecture intentionally favors conservative selection of candidate areas to reduce the risk of overlooking subtle malignancies.
Expert Commentary: Utility, Not Replacement
Mohammad Asadi, a data scientist at Stanford University who was not involved in the study, observed that the results suggest “Pathology-o3” may generalize across diverse slide sources—but cautioned that the findings do not yet demonstrate improved diagnostic speed or accuracy among practicing pathologists using the tool.
The study did not aim to prove that “Pathology-o3” could outperform or replace pathologists. Rather, it sought to test whether modeling the clinician’s search strategy—not just diagnostic outcomes—could improve AI performance.
Huang emphasized that the core value lies in leveraging underutilized behavioral data already generated routinely in hospital laboratories. Future experiments will assess whether integrating the system alongside pathologists increases detection rates and accelerates workflow.
Asadi suggested the most pragmatic current application is as a pre-screening aid—highlighting regions requiring deeper clinical scrutiny. He stressed the need for multicenter trials to evaluate real-world accuracy, processing speed, false-alert volume, and added workload burden on pathologists.
The system remains limited relative to full diagnostic practice, which often involves reviewing multiple slides, applying varied staining techniques, and incorporating patient medical history. Huang underscored that autonomous diagnosis is not the objective: “I would not claim it should diagnose independently.”
Latest news

Gold Holds Steady at $4,300.96/oz Ahead of Fed Rate Decision

Meghan Markle’s First Government-Backed Security Review Since 2020 Begins

Hidden Mouth Signs May Signal Serious Diseases


