← Research

Deployment Risk Probe Framework (DRPF)

Complementing CER with actionable failure modes

Ming-Kun Chou, Yu-Tai Lo, Wen-Hsiang Lu

ROCLING 2026, Archival Track. Oral presentation, November 2026.

The paper is written in Chinese with an English abstract. This page summarizes it; the proceedings version will be linked here once it is in the ACL Anthology.

Why CER is not enough

In a voice interface for advance care planning (ACP), an older adult may answer a question such as whether they would want intubation with one short phrase. If 「不要了」 (“no more”) is transcribed as 「服藥了」 (“took the medication”), the meaning is reversed, yet character error rate (CER) counts it as two wrong characters.

Two systems with similar CER can also fail in different ways. One tends to output something wrong, which a person has to check. The other tends to output nothing, in which case the right response is to ask again. CER does not show which situation you are in.

Three error types, three follow-ups

DRPF sorts errors by the kind of input, and each type maps to one deployment response. It does not replace CER; it is reported alongside CER and a keyword error rate for medical terms (KW-ERR).

InputSystem outputError typeFollow-up
Background noise only, nobody speakingA line of text(i) Text on non-speech inputFilter the input or tune voice activity detection
「嗯」(no output)(ii) Empty output on speechAsk the speaker to repeat; never treat it as consent
「嗯啊」「二十」(iii) Incorrect content on speechManual verification
「不要了」「服藥了」(iii), counted by PCPVerify first

Polarity-conflict probe (PCP)

Within type (iii), PCP counts the clearest high-risk case: the speaker gives a negated short answer and the transcript contains a medical action word such as 服藥 (take medication), 插管 (intubation) or 急救 (resuscitation). The trigger words were fixed during development and then checked on held-out sentences that played no part in choosing them, so that the rule is not simply fitted to the test answers.

Setup

Two deployment pipelines were run on the same recordings and scored the same way, changing one setting at a time: a cloud pipeline (remote Whisper large-v3-turbo with built-in voice activity detection) under two language settings, called lang-med and lang-base, and an edge pipeline (Nemotron 3.5 ASR 0.6B).

The short-answer probes are 400 clips from four speakers: confirmations, negations, other short replies, and hesitations. Of these, 76 are negated answers used for PCP. A separate set of 200 clinic-style sentences provided by a physician gives the CER reference.

Results

Setting Clinic-sentence CER Short-answer mismatch (ii) Empty (iii) Incorrect PCP
Cloud, lang-med11.2%259/40016/400243/40020/76
Cloud, lang-base3.9%128/40016/400112/4000/76
Edge (Nemotron)7.6%132/40072/40060/4000/76
Mismatch is (ii) plus (iii). Counts are those reported in the paper.
Cloud, lang-med 259/400
Cloud, lang-base 128/400
Edge (Nemotron) 132/400
  • (iii) incorrect content
  • (ii) empty output
  • Short-answer errors out of 400 clips
  • Nearly equal error counts, different handling. Cloud with lang-base and the edge pipeline miss about the same number of short answers (128 and 132 of 400). For the cloud pipeline, close to nine in ten of those are incorrect content, which needs manual verification. For the edge pipeline, more than half are empty output, where asking the speaker to repeat is the fitting response.
  • A high-risk error that CER does not show. Under lang-med, 20 of 76 negated short answers were rewritten into text containing a medical action word. On held-out negated sentences the count was 24 of 120, and the error appeared for all four speakers.
  • One setting changes the risk profile. Switching the cloud language setting from lang-med to lang-base removed these rewrites (0 of 76), while the Taiwanese lexical hit rate fell from 86.1% to 11.2%.
  • The scoring carries over to other domains. With the trigger words replaced by finance and law terms, the same program measured PCP at 10 of 72 and 6 of 72 under lang-med.

Limitations

  • This is a proof of concept on scripted utterances from four speakers. The results show that the error types can be told apart; they are not estimates of how often each error occurs in a population or statements about clinical safety.
  • The probes assume clear turns, one speaker at a time, and non-streaming input. Overlapping speech, interruptions, and streaming recognition were not covered, and the speech of real older patients was not included.
  • Only a small set of recognition systems was compared. Whether the same pattern holds for newer large audio-language models and for models trained specifically for Traditional Chinese remains to be tested.
  • Type (iii) errors are currently handed to manual verification. Detecting or correcting them semi-automatically is outside this evaluation design.

Code and data

The repository contains the scoring code, text normalization, the digital-silence clips, manifests for the public noise set, and the sentence scripts for the short-answer and hesitation probes. Speaker recordings, hospital background noise, and the physician-provided sentences are not released because they contain personal data or are restricted by the institutions involved.

This work was supported by the National Science and Technology Council, Taiwan (NSTC 115-2314-B-006-045).

Citation

@inproceedings{chou2026drpf,
  title     = {Deployment Risk Probe Framework ({DRPF}): Complementing {CER} with Actionable Failure Modes},
  author    = {Chou, Ming-Kun and Lo, Yu-Tai and Lu, Wen-Hsiang},
  booktitle = {Proceedings of the 38th Conference on Computational Linguistics and Speech Processing (ROCLING 2026)},
  year      = {2026}
}