2026.08.17

Offline Evaluations to Improve Our Systems

Catherine Chen, Samir Khan,
Michael Oberst, Alex Chouldechova
Introduction

To protect patient and clinician privacy, in all our blog posts we use representative synthetic examples.

All new features and product improvements at Abridge undergo “offline” pre-deployment evaluation. These evaluations are conducted on systems that are not yet deployed with our partners (not yet “online”), enabling experimentation without risking adverse impacts.

Offline evaluations enable us to compare new system versions to existing ones across a range of quality dimensions. While improvements to existing products often directly target a subset of quality dimensions—such as reducing misattribution rates—it is important to evaluate more broadly to be able to detect and address unexpected regressions along non-focal quality dimensions.

A core signal used in offline evaluations is expert feedback, often coming from clinicians or professional medical coders. We regularly ask experts to assess the quality of proposed system versions along different dimensions (for more info about these blinded head-to-head comparisons, please see “Pioneering the Science of AI Evaluation”). However, this approach is constrained by the cost and speed of human assessments; this impedes rapid system iteration cycles. One way we address this challenge is to conduct research on sample-efficient methods for combining expert annotations with automated annotations from LLM judges (for more details, please see “No Free Lunch: The Risks and Rewards of Incorporating LLM Judges”). The second way is to leverage expertise not for annotation but for developing automated offline evaluations.

This blog post describes one of our automated evaluation strategies, in which we create internal rubric-based benchmarks (for more details about our other offline evaluation strategies, please see “AI Evaluation at Abridge”). Our rubric-based benchmarks consist of encounter-specific, clinician-written rules that a high quality note must satisfy. These benchmarks codify clinicians’ judgment of what makes a high quality note for a given encounter into a set of rules that LLMs can automatically verify for any generated note.  This strategy provides high-quality feedback in a fraction of the time, facilitating rapid product iteration against a high-quality evaluation signal.

Constructing Our Rubric-Based Benchmarks

Our rubric-based benchmarking approach is similar to those used elsewhere to evaluate our clinical decision support systems, and to approaches in the literature for evaluating clinical chatbot conversations.1, 2, 3, 4 To ensure our benchmarks are reflective of the deployment setting, we construct them using de-identified data from Abridge encounters. These examples thus reflect all of the ambiguity and complexity of real-world clinical settings where Abridge is used, and provide the strongest validation of our systems.

The characteristics of a high-quality note are often specialty-dependent. For example, if in an encounter the patient mentions that they have been using Dove soap, this must be documented for a Dermatology encounter, but it should not be documented for an Ophthalmology encounter. Therefore we construct separate benchmarks for different specialties.

To build each benchmark, we enlist clinicians who have expertise in the respective specialty. For each encounter, the clinicians create a set of rules that describe the criteria a high-quality note documenting the encounter must meet. These rules form encounter-specific rubrics that can be re-used to grade notes for the same encounter but produced by different systems. On average, encounters in our benchmarks contain ten rules, but longer and more complex encounters can include up to fifty rules. By using LLMs to automate the grading, this approach allows us to evaluate new note generation system versions without requiring further expert annotation.

For example, we may have an encounter in which the patient presents with abdominal pain:

Doctor
Transcript
Patient
Good morning, Ms. Alvarez. What brings you in today?
I've been having this pain in my upper stomach, right here in the middle, just below my ribs. It's been bothering me for a while now.
Can you describe what the pain feels like?
It's a burning feeling. Almost like something hot sitting right there.
When did this start?
About three months ago. At first it was only once in a while, so I didn't think much of it.
Has anything changed recently?
Yes, actually — over the last two weeks it's been happening a lot more often. That's really why I made the appointment.
Is there anything that makes it worse?
It's definitely worse after I eat. Especially bigger meals. And it gets worse when I lie down — at night when I go to bed, that's when it really flares up.
Anything that makes it better?
Tums help a little, but not for long.
Any nausea, vomiting, blood in your stool, or black tarry stools?
No, none of that.
Any unintentional weight loss or trouble swallowing?
No.
Do you drink alcohol or caffeine?
Coffee every morning, maybe two cups. Wine with dinner a few nights a week.
Do you take ibuprofen, aspirin, or other anti-inflammatories regularly?
Just occasionally for headaches.
Okay, let me examine you. Lie back for me... I'm going to press on your belly. Tell me if anything hurts.
Ow — yes, right there, when you press in the middle up top. It's tender, but not terrible.
Does it hurt more when I let go quickly?
No, just when you press.
Okay. Your abdomen is soft — I don't feel any rigidity or guarding, and there's no rebound tenderness. Just some mild tenderness in that epigastric area where you've been feeling the pain.
So what do you think it is?
A few things could explain this — acid reflux, or GERD, is high on the list given the burning quality and the fact that it's worse after meals and lying down. Gastritis, which is inflammation of the stomach lining, is also possible. And we need to consider a peptic ulcer, especially since it's been going on for three months and getting more frequent.
Is there a test for that?
Yes. I'd like to check a stool test for H. pylori — that's a bacteria that can cause ulcers and gastritis. In the meantime, I'm going to start you on omeprazole, one tablet daily, which reduces stomach acid. But you need to wait to start the omeprazole until we collect the stool, as it can affect the result.
Okay.
I also want you to make some dietary changes: avoid eating late at night — try to finish dinner at least three hours before lying down — and cut out the caffeine and alcohol for now, since both can make this worse.
That'll be hard, but okay.
And I'm going to refer you to gastroenterology. If your symptoms don't improve on the omeprazole and the dietary changes, they can evaluate whether you need an upper endoscopy — a camera study to look directly at your stomach lining.
Got it. Thank you, doctor.
We'll have you follow up in four weeks, sooner if anything gets worse — especially any vomiting blood, black stools, or severe pain. Take care.

For this transcript, clinicians may write a rubric with the following rules:

Section
Rule
Criteria
HPI
01
The note must include that the patient describes epigastric location of pain.
02
The note must include the character of the epigastric pain as burning.
03
The note must include the onset of the pain as 3 months ago.
04
The note must include that the pain is worse after meals.
05
The note must include that the pain is worse when lying down.
06
The note must include that the pain has been increasing in frequency over the last two weeks.
Exam
07
The note must include the physical exam finding of mild epigastric tenderness to palpation.
08
The note must include the absence of rebound tenderness.
09
The note must include the absence of guarding.
A&P
10
The note must include the plan to obtain H. pylori stool antigen testing.
11
The note must include the referral to gastroenterology.
12
The note must include that the gastroenterology referral is for evaluation for upper endoscopy if symptoms do not improve.
13
The note must include GERD in the differential diagnosis.
14
The note must include gastritis in the differential diagnosis.
15
The note must include peptic ulcer disease in the differential diagnosis.
16
The note must include the initial treatment plan of a trial of omeprazole daily.
17
The note must include the plan for dietary modification.
18
The note must include avoiding late meals as part of the dietary modification.
19
The note must include avoiding caffeine as part of the dietary modification.
20
The note must include avoiding alcohol as part of the dietary modification.

Different candidate systems will produce different notes:

History of Present Illness
Maria Alvarez is a 54 year old female who presents with epigastric abdominal pain. She describes a burning pain located in the epigastrium that began approximately 3 months ago. The pain is worse after meals, particularly larger meals, and worse when lying down at night. Over the last two weeks, episodes have been increasing in frequency, which prompted today's visit. Antacids provide brief partial relief. She denies nausea, vomiting, hematochezia, melena, dysphagia, and unintentional weight loss. She drinks two cups of coffee daily and wine with dinner several nights per week. She uses NSAIDs occasionally for headaches.
Medications
- Ibuprofen PRN headaches
- Calcium carbonate (Tums) PRN
Social History
- Coffee: 2 cups daily
- Alcohol: wine with dinner several nights per week
- Denies tobacco use
Physical Exam
VITALS: BP- 122/76 HEENT: Normocephalic, atraumatic. NECK: Neck supple, no thyromegaly, no cervical lymphadenopathy. CHEST: Lungs clear to auscultation bilaterally, no wheezes, rales, or rhonchi. CARDIOVASCULAR: Heart regular rate and rhythm, no murmurs. ABDOMEN: Abdomen soft, non-distended, mild epigastric tenderness to palpation, no rebound tenderness, no guarding, bowel sounds normal. EXTREMITIES: Extremities warm and dry, peripheral pulses 2+ bilaterally. SKIN: Skin without rashes or pallor.
Assessment & Plan
Epigastric Abdominal Pain
Burning epigastric pain of 3 months duration, worse postprandially and when lying down, increasing in frequency over the past two weeks. Differential diagnosis includes GERD, gastritis, and peptic ulcer disease.
- Obtain H. pylori stool antigen testing.
- Begin trial of omeprazole daily.
- Dietary modification: avoid late meals, avoid caffeine, avoid alcohol.
- Referral to gastroenterology.
- Follow up in 4 weeks; return precautions reviewed for hematemesis, melena, or severe pain.
History of Present Illness
Maria Alvarez is a 54 year old female who presents with epigastric abdominal pain. She describes a burning pain located in the epigastrium. The pain is worse after meals, particularly larger meals. Antacids provide brief partial relief. She denies nausea, vomiting, hematochezia, melena, dysphagia, and unintentional weight loss. She drinks two cups of coffee daily and wine with dinner several nights per week. She uses NSAIDs occasionally for headaches.
Medications
- Ibuprofen PRN headaches
- Calcium carbonate (Tums) PRN
Social History
- Coffee: 2 cups daily
- Alcohol: wine with dinner several nights per week
- Denies tobacco use
Physical Exam
VITALS: BP- 122/76 HEENT: Normocephalic, atraumatic. NECK: Neck supple, no thyromegaly, no cervical lymphadenopathy. CHEST: Lungs clear to auscultation bilaterally, no wheezes, rales, or rhonchi. CARDIOVASCULAR: Heart regular rate and rhythm, no murmurs. ABDOMEN: Abdomen soft, non-distended, mild epigastric tenderness to palpation, no rebound tenderness, no guarding, bowel sounds normal. EXTREMITIES: Extremities warm and dry, peripheral pulses 2+ bilaterally. SKIN: Skin without rashes or pallor.
Assessment & Plan
Epigastric Abdominal Pain
Burning epigastric pain, worse postprandially. Differential diagnosis includes GERD, gastritis, and peptic ulcer disease.
- Obtain H. pylori stool antigen testing.
- Begin trial of omeprazole daily.
- Dietary modification: avoid late meals, avoid caffeine, avoid alcohol.
- Referral to gastroenterology for evaluation for upper endoscopy if symptoms do not improve.
- Follow up in 4 weeks; return precautions reviewed for hematemesis, melena, or severe pain.

The rubric can be re-used across these candidate systems to detect errors in the notes.

#
Rule
Note A
Note B
01
The note must include that the patient describes epigastric location of pain.
PASS
PASS
02
The note must include the character of the epigastric pain as burning.
PASS
PASS
03
The note must include the onset of the pain as 3 months ago.
PASS
FAIL
04
The note must include that the pain is worse after meals.
PASS
PASS
05
The note must include that the pain is worse when lying down.
PASS
FAIL
06
The note must include that the pain has been increasing in frequency over the last two weeks.
PASS
FAIL
07
The note must include the physical exam finding of mild epigastric tenderness to palpation.
PASS
PASS
08
The note must include the absence of rebound tenderness.
PASS
PASS
09
The note must include the absence of guarding.
PASS
PASS
10
The note must include the plan to obtain H. pylori stool antigen testing.
PASS
PASS
11
The note must include the referral to gastroenterology.
PASS
PASS
12
The note must include that the gastroenterology referral is for evaluation for upper endoscopy if symptoms do not improve.
FAIL
FAIL
13
The note must include GERD in the differential diagnosis.
PASS
PASS
14
The note must include gastritis in the differential diagnosis.
PASS
PASS
15
The note must include peptic ulcer disease in the differential diagnosis.
PASS
PASS
16
The note must include the initial treatment plan of a trial of omeprazole daily.
PASS
PASS
17
The note must include the plan for dietary modification.
PASS
PASS
18
The note must include avoiding late meals as part of the dietary modification.
PASS
PASS
19
The note must include avoiding caffeine as part of the dietary modification.
PASS
PASS
20
The note must include avoiding alcohol as part of the dietary modification.
PASS
PASS

We’ve created rubric-based benchmarks for 20+ specialties and we continue to add new specialties. Collectively, these benchmarks include thousands of encounters. These benchmarks are designed to be challenging for our systems–this ensures that our benchmarks are far from saturated, and allows us to detect improvements to our systems. To ensure that the benchmarks remain challenging, we continually add new encounters and rubrics, focusing on encounters for which earlier system versions have produced unsatisfactory notes.

Using Rubric-Based Benchmarks

The specificity of the rubrics in our benchmarks provides interpretable, actionable evaluations of our notes. We can examine the specific rules our notes fail, identify consistent patterns, and apply targeted mitigations.

To enable rapid iteration, we use an LLM-based rule evaluator to assess generated notes against each rubric rather than relying on clinician review. The evaluator takes in the generated note and rubric for each encounter, and produces a pass/fail score for each rule. To ensure this evaluator accurately reflects expert judgments, we validated it on clinician verdicts for 300 rules across 40 notes. The rule-checking model had high agreement with clinicians (0.78 TPR, 0.07 FPR). While the rule-checking model is far from perfect, it provides a directional quality signal that is useful for system development.

We use the benchmarks to iteratively improve our systems.

Naively iterating against the rubrics in a benchmark risks over-fitting: we might build systems that work very well on this specific set of encounters, but perform poorly when used in production on new encounters. To prevent over-fitting, we split each dataset into a train set for iteratively improving system quality, and a test set for evaluating quality on new encounters. These benchmarks have helped us measurably improve our systems. In one instance, we used one of our benchmarks to iterate on an agent that pulls information from prior notes in order to improve the note for the current encounter. The rubric-based evaluation revealed specific errors in which the agent pulled unnecessary prior content, placed information in the wrong section, and omitted important prior content. Using this offline quality signal, we rapidly iterated on the agent, and in particular iterated much faster than a clinician review between each iteration would have allowed. After these iterations led to a final candidate system, we enlisted clinicians to directly validate the improvements. The clinicians strongly preferred the outputs from the updated agent: in a blinded comparison between the original and updated outputs, clinicians preferred the updated outputs in 91% of comparisons.

Conclusion

Automated offline evaluations enable us to rapidly iterate on our systems to improve quality.  Our internal rubric-based benchmarks are just one piece of our offline evaluation strategy, which also includes clinician review, LLM judges, and adversarial testing. Each of these pieces provides a different view of system quality. Together, they help us improve our systems and prevent errors before they ever reach our partners.