Navigating the OverKill in Large Language Models

Shi, Chenyu; Wang, Xiao; Ge, Qiming; Gao, Songyang; Yang, Xianjun; Gui, Tao; Zhang, Qi; Huang, Xuanjing; Zhao, Xun; Lin, Dahua

Computer Science > Computation and Language

arXiv:2401.17633 (cs)

[Submitted on 31 Jan 2024]

Title:Navigating the OverKill in Large Language Models

Authors:Chenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang, Xuanjing Huang, Xun Zhao, Dahua Lin

View PDF

Abstract:Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer benign queries. In this paper, we investigate the factors for overkill by exploring how models handle and determine the safety of queries. Our findings reveal the presence of shortcuts within models, leading to an over-attention of harmful words like 'kill' and prompts emphasizing safety will exacerbate overkill. Based on these insights, we introduce Self-Contrastive Decoding (Self-CD), a training-free and model-agnostic strategy, to alleviate this phenomenon. We first extract such over-attention by amplifying the difference in the model's output distributions when responding to system prompts that either include or omit an emphasis on safety. Then we determine the final next-token predictions by downplaying the over-attention from the model via contrastive decoding. Empirical results indicate that our method has achieved an average reduction of the refusal rate by 20\% while having almost no impact on safety.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2401.17633 [cs.CL]
	(or arXiv:2401.17633v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2401.17633

Submission history

From: Xiao Wang [view email]
[v1] Wed, 31 Jan 2024 07:26:47 UTC (7,807 KB)

Computer Science > Computation and Language

Title:Navigating the OverKill in Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Navigating the OverKill in Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators