FROMANNUAL REVIEWS

CogSci 2025

•

August 02, 2025

•

San Francisco, United States

keywords:

language understanding

pattern recognition

artificial intelligence

neural networks

natural language processing

How do vision-language (VL) transformer models ground verb phrases and do they integrate contextual and world knowledge in this process? We introduce the CV-Probes dataset, containing image-caption pairs involving verb phrases that require both social knowledge and visual context to interpret (e.g., `beg'), as well as pairs involving verb phrases that can be grounded based on information directly available in the image (e.g., sit"). We show that VL models struggle to ground VPs that are strongly context-dependent. Further analysis using explainable AI techniques shows that such models may not pay sufficient attention to the verb token in the captions. Our results suggest a need for improved methodologies in VL model training and evaluation. The code and dataset will be available https://github.com/ivana-13/CV-Probes.

Downloads

Paper

Next from CogSci 2025

Boosting Cognitive Modelling for Human Reasoning
poster

Boosting Cognitive Modelling for Human Reasoning

CogSci 2025

Marco Ragni
Meghna Bhadra and 1 other author

02 August 2025

Similar lecture

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
poster

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

AAAI 2026

+6
Yifan Fang and 8 other authors

25 January 2026

Stay up to date with the latest Underline news!

Select topic of interest (you can select more than one)

PRESENTATIONS

  • All Presentations
  • For Librarians
  • Resource Center
  • Free Trial
Underline Science, Inc.
1216 Broadway, 2nd Floor, New York, NY 10001, USA

© 2026 Underline - All rights reserved

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.