Lecture image placeholder

Premium content

Access to this content requires a subscription. You must be a premium user to view this content.

Monthly subscription - $9.99Pay per view - $4.99Access through your institutionLogin with Underline account
Need help?
Contact us
Lecture placeholder background

EMNLP 2025

November 09, 2025

Suzhou, China

Would you like to see your presentation here, made available to a global audience of researchers?
Add your own presentation or have us affordably record your next conference.

Sycophancy is a key behavioral risk in LLMs, yet is often treated as an isolated failure mode that occurs with a single causal mechanism. We instead propose modeling it as geometric and causal compositions of psychometric traits such as emotionality, openness, and agreeableness, similar to factor decomposition in psychometrics. Using Contrastive Activation Addition (CAA) (Panickssery et al., 2024), we map activation direction to these factors and study how different combinations may give rise to sycophancy (e.g., high extraversion combined with low conscientiousness). This perspective allows for interpretable and compositional vector based interventions like addition, subtraction and projection; that may be used to mitigate safety-critical behaviors in LLMs.

Next from EMNLP 2025

Long Context Benchmark for the Russian Language
workshop paper

Long Context Benchmark for the Russian Language

EMNLP 2025

+5
Murat Apishev and 7 other authors

09 November 2025

Stay up to date with the latest Underline news!

Select topic of interest (you can select more than one)

PRESENTATIONS

  • All Presentations
  • For Librarians
  • Resource Center
  • Free Trial
Underline Science, Inc.
1216 Broadway, 2nd Floor, New York, NY 10001, USA

© 2026 Underline - All rights reserved