EMNLP 2025

November 05, 2025

Suzhou, China

Would you like to see your presentation here, made available to a global audience of researchers?
Add your own presentation or have us affordably record your next conference.

We introduce Spire, a speech-augmented language model (LM) capable of both translating and transcribing speech input from English into 10 other languages as well as translating text input in both language directions. Spire integrates the speech modality into an existing multilingual LM (MLM) via speech discretization and continued pre-training using only 42.5K hours of speech. In particular, we adopt the pretraining framework of MLMs and treat discretized speech input as an additional translation language. This approach not only equips the MLM with speech capabilities, but also preserves its strong text-only performance. We achieve this using significantly less data than existing speech LMs, demonstrating that discretized speech input integration as an additional language is feasible during LM adaptation. We will make our code and models available to the community.

Downloads

SlidesPaperTranscript English (automatic)

Next from EMNLP 2025

DeepNote: Note-Centric Deep Retrieval-Augmented Generation
poster

DeepNote: Note-Centric Deep Retrieval-Augmented Generation

EMNLP 2025

+9Zhiyuan Liu
Yuxuan Chen and 11 other authors

05 November 2025

Stay up to date with the latest Underline news!

Select topic of interest (you can select more than one)

PRESENTATIONS

  • All Presentations
  • For Librarians
  • Resource Center
  • Free Trial
Underline Science, Inc.
1216 Broadway, 2nd Floor, New York, NY 10001, USA

© 2025 Underline - All rights reserved