← All work

Case 01

Pranik.ai · AI evals and clinical safety

Hat 01 · AI evals, models and clinical safety

Ranking AI errors by clinical risk: the evals loop behind a medical AI.

I led the design of the evals loop that takes every doctor correction from diff to root cause to the right team, and built the model evaluation framework leadership uses for AI decisions. Drug-name errors fell ~70%, and prescription latency dropped from 2+ minutes to ~35 seconds.

Role
Led the design of the evals loop and its rubric. Owns triage, sequencing and the UX fixes.
Timeline
Sep 2025 to present
Team
Data science lead and interns, AI and backend engineers, QA, 4 doctors and 2 medical annotators
Skills
LLM evalsClinical safetyASRRAGModel trade-offs
  • ~70%fewer drug-name errors, to ~10% of mentions
  • 2+ min → ~35 saverage prescription latency
  • 10root-cause categories in the rubric
  • Same daygo or no-go on a new model

Why generic evals weren’t enough

Our AI listens to a doctor and patient talking in a mix of Telugu, Hindi, English and other languages, often in a noisy OPD, and turns it into a prescription and case sheet. When it gets something wrong, the mistake can be harmless (a symptom spelled oddly) or dangerous (the wrong drug, the wrong dose, the wrong side of the body).

Benchmarks and generic accuracy scores flatten those two into one number. They also miss our richest signal: every time a doctor edits the AI’s output, they are telling us exactly what the right answer was.

So the question I set out to answer was simple to say and hard to build: how do we turn thousands of doctor edits into the right fix, in the right layer, owned by the right team, in the right order?

The loop

The clinical evals loop

running on the live product

Every doctor edit travels the loop. Step 04 is where the fix finds its owner.

  1. 01 · Edit Doctor corrects the AI case sheet Every single edit is a signal. One edit is the counting unit, not one case sheet.
  2. 02 · Diff Expected vs captured Edit-diff rules compare the doctor’s version with the AI’s.
  3. 03 · Classify One edit, one or more error types A 10-category root-cause matrix. Automated rules pre-classify at scale.
  4. 04 · Root cause Which layer broke? Prompt, context, memory, validation, RAG or grounding, model, or product and UX.
  5. 05 · Triage Clinical risk first Safety failures jump the queue, however rare they are.
  6. 06 · Fix + verify Ship, then check blind Fast UX patch now, deeper fix later. Doctors confirm the change without being told.
Routed to a layer and a team → PromptContextMemoryValidationRAGModelProduct / UX
FIG. Each doctor correction travels the loop. The real rubric and matrix are confidential; this shows the shape.

I led the design of the loop and its rubric, pulling together inputs from our data scientists, leadership and testers. It runs on the live clinical product.

  • The counting unit is a single edit, not a case sheet. Every instance counts, which gives us a cumulative error rate per scenario.
  • One correction can span several error types. A dose recorded wrongly in the wrong language is two problems, not one.
  • Each error gets a root cause and a layer: prompt, context, memory, validation, RAG or grounding, the model itself, or product and UX. That decides which team owns it.
  • Automated rules classify at scale. A long passage deleted outright suggests the AI made something up. Corrected text that is much longer than the original suggests it left something out. A diff on numbers and units flags dose errors.

Clinical risk beats frequency

If you sort the backlog by count, the patient-safety failures sink to the bottom. So the rubric is a form of severity-weighted quality: it triages every issue by clinical risk first, then by UX impact.

Severity-weighted quality

Illustrative placement

The most common edits are rarely the most dangerous ones.

Clinical risk ↑

Rare, dangerous

  • Left-right flip
  • Missed red flag

Fix first, however rare

Common, dangerous

  • Sound-alike drug names

Fix now, with a fast UX patch

Rare, harmless

  • Odd formatting

Watch it

Common, harmless

  • Vernacular idioms
  • Spelling of symptoms

Batch it into the data roadmap

How often it happens →

“stomach ulcer” → “peptic ulcer” Off the chart: not an error at all. It is the doctor’s style, so it becomes a rule, not a bug.
FIG. Sorted by count, the top row sinks to the bottom of the backlog. The rubric sorts by clinical risk first, then by UX impact.

The difference is easier to feel than to read. Here is a short version of the triage, with synthetic examples:

Try it · triage four corrections

You’re on the evals loop today. Four doctor corrections came in.

0/4 classified right

AI wrote

Doctor changed it to

What went wrong?

FIG. Synthetic, generic examples and illustrative counts. The categories echo the real matrix, which stays confidential.

Where my decisions sit

Our data science team decides the technical approaches: which method, which inference strategy. My job in the loop is different, and I own it:

  1. Pull signals for each issue from the evals platform, product metrics and the doctors themselves.
  2. Classify it as a data fix, a UI fix or an AI fix, or more often a mix.
  3. Decide which function owns it and in what sequence. A typical call: ship a UX patch within a week while data science builds the deeper fix over a month.
  4. Own the UX fix myself.
  5. Decide when data science’s tested approach clears the ship bar.

Case: drug names that sound alike

Similar-sounding drugs were being captured as each other. The metric: drug mentions with an error, divided by all drug mentions the system captured.

The fix was four changes shipped together over about two months by a data science lead and two interns. No single one of them would have done it alone.

Layered defense · sound-alike drug names

Four changes, shipped together

No single layer fixed it. Together they did.

  1. 01 · CapturePrompt restructure, shorter audio chunksBetter context stitching, so the name is heard right more often.
  2. 02 · GroundKnowledge graph of approved medicinesMedicines approved in India, linked to the diseases they treat.
  3. 03 · RankSound-and-meaning search, rerankedEvery mention is checked against similar names, then against the chief complaint.
  4. 04 · RecoverCorrection UX for the doctorTap a wrong name to see the next-ranked options, or type freely.
FIG. Drug-name errors fell ~70%, to about 10% of mentions, measured in the weeks after the combined fix with no model swap in that window. Lines are illustrative.

Drug-name errors fell ~70%, to about 10% of mentions, and it has held since.

Case: prescription latency

Doctors will not wait two minutes for a prescription with the next patient already in the room. I set the end-to-end latency target, then worked with engineering through R&D proofs of concept at each pipeline stage, benchmarking variants of prompting, chunking and output structure.

The hard part was cascading effects: a faster stage upstream could quietly lower quality downstream. We weighed quality against latency at every stage before I chose the pipeline now in use.

End of consultation to prescription

Average prescription latencycurrent pipeline; some runs go above it

Before2+ min
After~35 s
FIG.Both this and the drug-name result were verified in blind tests by doctors, not just on our dashboards.

The model evaluation framework

Models change every few weeks. Before this framework, every new release meant ad hoc retesting with no SOP. I created the framework that leadership’s AI decisions now rest on.

Eval gates for every LLM, ASR and RAG pipeline

Cheap checks first, human checks last, a decision the same day.

InputA new model, or a change to any pipeline
  1. Scenarios A workbook of every scenario, by priority, per use case Written testing SOPs and stored results, updated continuously.
  2. Tier 1 Automated checks Run by the AI team. A cheap first gate.
    AI team
  3. Tier 2 Scripted human tests Languages, accents and noise conditions that can’t be fully automated.
    QA
OutputSame-day go or no-go

It also drives quality-versus-cost calls on vendors. When our primary LLM provider deprecated a model our voice pipeline depended on, the replacement performed worse for our use. The framework let us keep the stronger model in production through its deprecation window while we rebuilt the pipeline around a different model in development.

Data operations: a golden dataset

The data science roadmap is set by our Director of AI, Prof. Gaurav Raina (IIT Madras), with the founders. I operationalize it: daily sprints, progress tracking, prioritization, and coordination with the professors, including Prof. Praveen Tammana (IIT Hyderabad), who mentors our data science team and our IIT Madras and IIT Hyderabad interns.

I set up our clinical data operations: the SOPs and framework, the team structure, onboarding and training. I also designed the annotation and evals portal in Claude Design, which engineering then built. The design goal was to reduce doctors’ cognitive load and the time each case takes.

Today 4 doctors and 2 medical annotators are turning 6,000+ consultations into a gold-standard Telugu medical dataset. The same portal handles doctor feedback and side-by-side ranking of outputs from different pipelines.

Shared at a public-safe level. Vendor names, internal data and patient examples stay out. Happy to go deeper in a conversation.

Pending launch