FROMANNUAL REVIEWS

CogSci 2025

•

August 01, 2025

•

San Francisco, United States

keywords:

computer-based experiment

interactive behavior

computer science

artificial intelligence

natural language processing

Large Language Models (LLMs) have been proven useful for various tasks but remain vulnerable to malicious exploitation. Attackers can bypass LLM safety restrictions ("jail") through carefully crafted "jailbreaking" prompts. To evaluate LLMs' security, researchers proposed various jailbreak techniques based on optimization, obfuscation, or persuasive strategies. However, these methods treat LLMs as passive persuasion targets, which overlooks LLMs' ability to reason actively. We propose Persu-Agent, a novel jailbreak framework based on Greenwald's Cognitive Response Theory. We focus more on LLM's internal cognitive processing of a prompt than the prompt itself. Persu-Agent uses the self-persuasion strategy to guide LLMs in generating justifications and rationalizing responses to harmful queries. The experimental results on advanced open-source and commercial LLMs revealed that Persu-Agent achieved an average jailbreak success rate of 84%, surpassing existing SOTA methods. Our work provides valuable insights into understanding LLMs' cognitive traits and contributes to developing safer LLMs.

Downloads

PaperTranscript English (automatic)

Next from CogSci 2025

Cognitive Insights into Document Comprehension: The Role of Reading Order and Visual Attention in Human and Large Language Models
poster

Cognitive Insights into Document Comprehension: The Role of Reading Order and Visual Attention in Human and Large Language Models

CogSci 2025

Qingxuan Wang and 1 other author

01 August 2025

Similar lecture

Distract Large Language Models for Automatic Jailbreak Attack
poster

Distract Large Language Models for Automatic Jailbreak Attack

EMNLP 2024

Guanhua Chen
+1
Zeguan Xiao and 3 other authors

13 November 2024

Stay up to date with the latest Underline news!

Select topic of interest (you can select more than one)

PRESENTATIONS

  • All Presentations
  • For Librarians
  • Resource Center
  • Free Trial
Underline Science, Inc.
1216 Broadway, 2nd Floor, New York, NY 10001, USA

© 2026 Underline - All rights reserved

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.