AI models should output confidence levels for all outputs.
Candidate Thesis
Output Confidence Levels
Summary
The case for including this
Attaching a confidence level to outputs operationalizes honesty about uncertainty: it helps users calibrate how much to trust a given answer, discourages blind over-reliance, and is especially valuable in high-stakes settings like medicine or law where 'how sure are you?' changes what the user should do next.
The case for changing or excluding this
Including this trades a reassuring appearance of precision against actual reliability: today's models are often poorly calibrated, so a confident-looking number can mislead users more than a plain answer would, lending false authority to wrong outputs. Requiring it for every output is also impractical and noisy — many responses have no meaningful single 'confidence' — and the figure can be gamed or misread as a guarantee of accuracy.
Discussion
Sign in to join the discussion.