External Resource
XSTest: Identifying Exaggerated Safety Behaviours in LLMs
Röttger et al., NAACL 2024
Shows models systematically over-refuse clearly safe prompts, withholding legitimate help — the 'over-blocking' harm a strict no-harm rule can cause.
Theses this informs
- Do No HarmChallenges