COLLECTED LIGHT · ORIGINAL STORY
Pathology Never Assumed Experts Couldn't Be Wrong

How AI Made Me Understand Pathology All Over Again
Today I was chatting with several colleagues from eBay and TikTok about platform governance and trust.
In marketplace governance, a 3% false positive can wrongly punish thousands of legitimate sellers. A single false negative, on the other hand, may expose consumers to significant harm and erode trust in the platform.
Naturally, the conversation turned to AI.
Can AI solve this?
Can we keep pushing accuracy even higher?
At one point, someone was surprised to learn that I graduated from medical school and had previously worked in pathology, focusing on lymphoma, immunohistochemistry, and molecular diagnostics.
That immediately brought me back to my days in Professor Gao Zifen's pathology department.
When I was young, I never understood why pathologists spent so much time discussing a single case.
Why did the immunohistochemistry panel keep expanding?
Why did experts constantly seek second opinions—and sometimes openly disagree with one another?
Why did nobody ever seem completely certain?
Years later, I think I finally understand.
Pathology never assumed experts couldn't be wrong.
Quite the opposite.
It starts with the assumption that people will disagree, drift in judgment, become overconfident, and sometimes even lead an entire community gradually off course without realizing it.
That's why multidisciplinary consultations, double reading, immunohistochemistry, molecular diagnostics, digital pathology, external quality assessment, and today's increasingly important proficiency testing are all expressions of the same philosophy at different stages of technological progress.
Their purpose has never been to prove that we are right.
Their purpose is to keep a complex system continuously calibrated.
The greatest success of pathology proficiency testing is the mistakes that never happen.
In The Black Swan, Nassim Nicholas Taleb describes the idea of Via Negativa.
We naturally reward achievements that happened.
We rarely reward disasters that never occurred because someone prevented them.
Viewed this way, every issue uncovered through proficiency testing represents a real-world error that was eliminated before it could ever reach a patient.
In a sense, proficiency testing records a kind of counterfactual history—events that never happened, yet mattered immensely.
The most dangerous enemy is rarely ignorance.
It's overconfidence.
Donald Rumsfeld famously divided uncertainty into known knowns, known unknowns, and unknown unknowns.
Many of our greatest risks don't come from problems we already recognize.
They come from the problems we haven't yet realized exist.
In complex systems, the greatest threat is often not making mistakes.
It's becoming increasingly confident while quietly drifting away from the truth.
Knowledge eventually outlives individuals.
Systems become its memory.
Pathologists retire.
Young physicians take their place.
Technology evolves—from digital pathology to spatial transcriptomics to AI.
Yet experience continues to be encoded into the system through feedback and calibration, allowing knowledge to survive far longer than any individual expert.
In that sense, external quality assessment and proficiency testing are forms of collective memory.
Great systems don't depend on extraordinary individuals.
They depend on the continuous accumulation, preservation, and transfer of experience.
The greatest strength of an excellent system isn't that it trusts experts.
It's that it refuses to trust anyone completely—including the experts themselves.
Karl Popper wrote in Conjectures and Refutations:
Science progresses not by proving ourselves right, but by exposing ourselves to being wrong.
Science advances not because we continuously prove ourselves correct, but because we continuously create opportunities to discover where we might be wrong.
Looking back now, I think the greatest invention of proficiency testing was never higher accuracy.
It was turning humility into a system.
Institutionalized humility.
Perhaps we've been asking the wrong questions all along.
For years I kept asking:
How can we reduce false positives?
How can we reduce false negatives?
How can we make AI even more accurate?
Today, I'm no longer convinced those are the first-principles questions.
Complex systems will always encounter new risks and new unknowns.
The real challenge has never been eliminating error.
It's discovering mistakes earlier, correcting them faster, and preventing an entire system from slowly drifting in the wrong direction.
Rather than pursuing perfect prediction, perhaps we should pursue continuous calibration.
Rather than relying on increasingly intelligent models, perhaps we should build increasingly honest feedback mechanisms.
Today, I find myself placing less faith in any single expert, any single team, or even any single AI model.
Instead, I trust systems that continuously invite challenge, expose their own weaknesses, and learn from every correction.
Perhaps the true purpose of a complex system is not to prove that it is right.
It is to create endless opportunities to discover where it may be wrong.
Because the deepest form of safety in any complex system never comes from perfection.
It comes from its ability to continually correct itself.
The greatest danger isn't making mistakes.
It's when an entire community slowly drifts away from reality—and no one notices.
And the most dangerous mistakes are often the ones that make us feel increasingly certain that we're right.