Lecture image placeholder

Premium content

Access to this content requires a subscription. You must be a premium user to view this content.

Monthly subscription - $9.99Pay per view - $4.99Access through your institutionLogin with Underline account
Need help?
Contact us
Lecture placeholder background

EMNLP 2025

November 07, 2025

Suzhou, China

Would you like to see your presentation here, made available to a global audience of researchers?
Add your own presentation or have us affordably record your next conference.

keywords:

masked diffusion language model

visual representation learning

vision language pretraining

Learning from paired vision-language data is a key direction in building large-scale foundation models. In this work, we revisit image captioning as a framework for visual representation learning and demonstrate its effectiveness in leveraging rich linguistic context. We propose masked diffusion captioner (MDC), a masked diffusion language model conditioned on the image, designed to overcome the limitations of BERT-style and parallel decoding masking strategies. MDC learns significantly stronger visual features and generates higher-quality captions. Extensive experiments across benchmarks show that MDC remains competitive in overall feature quality with traditional autoregressive models. Further studies are done to provide insights into the design and scalability of MDC. Our findings position captioning based pretraining as a promising paradigm for vision-language representation learning.

Downloads

Paper
access premium content

Next from EMNLP 2025

MPRF: Interpretable Stance Detection through Multi-Path Reasoning Framework
poster

MPRF: Interpretable Stance Detection through Multi-Path Reasoning Framework

EMNLP 2025

+2Jiafeng GuoXueqi Cheng
Jade Chang and 4 other authors

07 November 2025

Stay up to date with the latest Underline news!

Select topic of interest (you can select more than one)

PRESENTATIONS

  • All Presentations
  • For Librarians
  • Resource Center
  • Free Trial
Underline Science, Inc.
1216 Broadway, 2nd Floor, New York, NY 10001, USA

© 2026 Underline - All rights reserved