Case 01
Pranik.ai · AI evals and clinical safety
Hat 01 · AI evals, models and clinical safety
Ranking AI errors by clinical risk: the evals loop behind a medical AI.
I led the design of the evals loop that takes every doctor correction from diff to root cause to the right team, and built the model evaluation framework leadership uses for AI decisions. Drug-name errors fell ~70%, and prescription latency dropped from 2+ minutes to ~35 seconds.
- ~70%fewer drug-name errors, to ~10% of mentions
- 2+ min → ~35 saverage prescription latency
- 10root-cause categories in the rubric
- Same daygo or no-go on a new model
- 01Doctors correct the AI’s case sheets every day. I led the design of a loop that turns each correction into a diagnosis: what broke, in which layer, and which team fixes it.
- 02An error-classification rubric with edit-diff rules and a 10-category root-cause matrix routes each fix to prompt, context, memory, validation, RAG, model, or product and UX.
- 03Fixes are ranked by clinical risk, not frequency. A left-right flip outranks a common, harmless typo.
- 04A combination of four fixes cut drug-name errors ~70%, to about 10% of mentions. I owned the correction UX.
- 05Set the latency target and chose the pipeline that took prescriptions from 2+ minutes to ~35 seconds. Both results were verified in blind doctor tests.
- 06Built the model evaluation framework: tiered tests for every LLM, ASR and RAG pipeline, and a same-day go or no-go on any new model.
Why generic evals weren’t enough
Our AI listens to a doctor and patient talking in a mix of Telugu, Hindi, English and other languages, often in a noisy OPD, and turns it into a prescription and case sheet. When it gets something wrong, the mistake can be harmless (a symptom spelled oddly) or dangerous (the wrong drug, the wrong dose, the wrong side of the body).
Benchmarks and generic accuracy scores flatten those two into one number. They also miss our richest signal: every time a doctor edits the AI’s output, they are telling us exactly what the right answer was.
So the question I set out to answer was simple to say and hard to build: how do we turn thousands of doctor edits into the right fix, in the right layer, owned by the right team, in the right order?
The loop
The clinical evals loop
running on the live product
Every doctor edit travels the loop. Step 04 is where the fix finds its owner.
- 01 · Edit Doctor corrects the AI case sheet Every single edit is a signal. One edit is the counting unit, not one case sheet.
- 02 · Diff Expected vs captured Edit-diff rules compare the doctor’s version with the AI’s.
- 03 · Classify One edit, one or more error types A 10-category root-cause matrix. Automated rules pre-classify at scale.
- 04 · Root cause Which layer broke? Prompt, context, memory, validation, RAG or grounding, model, or product and UX.
- 05 · Triage Clinical risk first Safety failures jump the queue, however rare they are.
- 06 · Fix + verify Ship, then check blind Fast UX patch now, deeper fix later. Doctors confirm the change without being told.
I led the design of the loop and its rubric, pulling together inputs from our data scientists, leadership and testers. It runs on the live clinical product.
- The counting unit is a single edit, not a case sheet. Every instance counts, which gives us a cumulative error rate per scenario.
- One correction can span several error types. A dose recorded wrongly in the wrong language is two problems, not one.
- Each error gets a root cause and a layer: prompt, context, memory, validation, RAG or grounding, the model itself, or product and UX. That decides which team owns it.
- Automated rules classify at scale. A long passage deleted outright suggests the AI made something up. Corrected text that is much longer than the original suggests it left something out. A diff on numbers and units flags dose errors.
Clinical risk beats frequency
If you sort the backlog by count, the patient-safety failures sink to the bottom. So the rubric is a form of severity-weighted quality: it triages every issue by clinical risk first, then by UX impact.
Severity-weighted quality
Illustrative placement
The most common edits are rarely the most dangerous ones.
Clinical risk ↑
Rare, dangerous
- Left-right flip
- Missed red flag
Fix first, however rare
Common, dangerous
- Sound-alike drug names
Fix now, with a fast UX patch
Rare, harmless
- Odd formatting
Watch it
Common, harmless
- Vernacular idioms
- Spelling of symptoms
Batch it into the data roadmap
How often it happens →
The difference is easier to feel than to read. Here is a short version of the triage, with synthetic examples:
Try it · triage four corrections
You’re on the evals loop today. Four doctor corrections came in.
0/4 classified right
What went wrong?
This week
Next sprints
Now sort the backlog. Which order would you ship in?
Where my decisions sit
Our data science team decides the technical approaches: which method, which inference strategy. My job in the loop is different, and I own it:
- Pull signals for each issue from the evals platform, product metrics and the doctors themselves.
- Classify it as a data fix, a UI fix or an AI fix, or more often a mix.
- Decide which function owns it and in what sequence. A typical call: ship a UX patch within a week while data science builds the deeper fix over a month.
- Own the UX fix myself.
- Decide when data science’s tested approach clears the ship bar.
Case: drug names that sound alike
Similar-sounding drugs were being captured as each other. The metric: drug mentions with an error, divided by all drug mentions the system captured.
The fix was four changes shipped together over about two months by a data science lead and two interns. No single one of them would have done it alone.
Layered defense · sound-alike drug names
Four changes, shipped together
No single layer fixed it. Together they did.
- 01 · CapturePrompt restructure, shorter audio chunksBetter context stitching, so the name is heard right more often.
- 02 · GroundKnowledge graph of approved medicinesMedicines approved in India, linked to the diseases they treat.
- 03 · RankSound-and-meaning search, rerankedEvery mention is checked against similar names, then against the chief complaint.
- 04 · RecoverCorrection UX for the doctorTap a wrong name to see the next-ranked options, or type freely.
Drug-name errors fell ~70%, to about 10% of mentions, and it has held since.
Case: prescription latency
Doctors will not wait two minutes for a prescription with the next patient already in the room. I set the end-to-end latency target, then worked with engineering through R&D proofs of concept at each pipeline stage, benchmarking variants of prompting, chunking and output structure.
The hard part was cascading effects: a faster stage upstream could quietly lower quality downstream. We weighed quality against latency at every stage before I chose the pipeline now in use.
End of consultation to prescription
Average prescription latencycurrent pipeline; some runs go above it
The model evaluation framework
Models change every few weeks. Before this framework, every new release meant ad hoc retesting with no SOP. I created the framework that leadership’s AI decisions now rest on.
Eval gates for every LLM, ASR and RAG pipeline
Cheap checks first, human checks last, a decision the same day.
- AI team
- QA
It also drives quality-versus-cost calls on vendors. When our primary LLM provider deprecated a model our voice pipeline depended on, the replacement performed worse for our use. The framework let us keep the stronger model in production through its deprecation window while we rebuilt the pipeline around a different model in development.
Data operations: a golden dataset
The data science roadmap is set by our Director of AI, Prof. Gaurav Raina (IIT Madras), with the founders. I operationalize it: daily sprints, progress tracking, prioritization, and coordination with the professors, including Prof. Praveen Tammana (IIT Hyderabad), who mentors our data science team and our IIT Madras and IIT Hyderabad interns.
I set up our clinical data operations: the SOPs and framework, the team structure, onboarding and training. I also designed the annotation and evals portal in Claude Design, which engineering then built. The design goal was to reduce doctors’ cognitive load and the time each case takes.
Today 4 doctors and 2 medical annotators are turning 6,000+ consultations into a gold-standard Telugu medical dataset. The same portal handles doctor feedback and side-by-side ranking of outputs from different pipelines.
Shared at a public-safe level. Vendor names, internal data and patient examples stay out. Happy to go deeper in a conversation.