suki.ai logo

Whitepaper

One score, nine dimensions, zero signal: The problem with holistic Likert ratings

The way most organizations evaluate clinical AI scribes fails to measure meaningful quality.

PDQI-9, the most widely used clinical note quality tool, was validated in 2012 on 9 inpatient notes from 3 charts. It's now the de facto gatekeeper for ambient AI scribes writing millions of notes a year.

We ran it across 84 paired notes in 4 specialties, with two production-candidate LLM models. The results:

  • No one could agree on what counts as a "hallucination"
  • Even "accuracy", the most basic quality check, was unreliable
  • The same raters flip-flopped depending on which AI wrote the note
  • The worst disagreement was originally hidden by a data-formatting bug

This isn't noisy data. It's a structural mismatch. PDQI-9 scores holistically, looking at factors like organization, conciseness, and cohesiveness, which is exactly the wrong lens for LLM errors. For example, an incorrect dose can sit inside a beautifully organized, internally consistent note and vanish from the score entirely. LLM error phenomenology is structurally different from the deficiencies PDQI-9 was built to detect.

We looked at the alternatives too (PDSQI-9, SCRIBE, FActScore, VeriFact, CREOLA, and more). Real progress in places, but none combine sentence-level error detection, inter-rater reliability as a first-class metric, and a statistical procedure for model release decisions.

Download the full whitepaper from Head of Clinical AI Informatics Karl Swanson, and Senior Machine Learning Engineer Nikita Gupta, to discover why holistic rubrics cannot detect LLM errors, and what this means in practice. 

Fill out our form to access full whitepaper now.

Privacy - Terms

suki.ai logo
AICPA | SOC logoHippa Compliant
Visit Suki’s Trust Portal to learn more.
Privacy Policy Terms of Service
© Copyright 2026. Suki AI, Inc. All rights reserved.