AI Constitution

External Resource

XSTest: Identifying Exaggerated Safety Behaviours in LLMs

Röttger et al., NAACL 2024

View at source ↗

Shows models systematically over-refuse clearly safe prompts, withholding legitimate help — the 'over-blocking' harm a strict no-harm rule can cause.

Theses this informs