Lecture image placeholder

Premium content

Access to this content requires a subscription. You must be a premium user to view this content.

Monthly subscription - $9.99Pay per view - $4.99Access through your institutionLogin with Underline account
Need help?
Contact us
Lecture placeholder background
VIDEO DOI: https://doi.org/10.48448/sdt4-2a63

workshop paper

ACL 2024

August 15, 2024

Bangkok, Thailand

Do LLMs Speak Kazakh? A Pilot Evaluation of Seven Models

keywords:

kazakh

llm

large language models

evaluation

We conducted a systematic evaluation of seven large language models (LLMs) on tasks in Kazakh. Kazakh is a Turkic language spoken by approximately 13 million native speakers in Kazakhstan and abroad. We used six datasets corresponding to different tasks -- questions answering, causal reasoning, middle school math problems, machine translation, and spelling correction. Three of the datasets were prepared for this study. As expected, the quality of the LLMs on the Kazakh tasks is lower than on the parallel English tasks. GPT-4 shows the best results, followed by Gemini and Aya. In general, LLMs perform better on classification tasks and struggle with generative tasks. Our results provide valuable insights into the applicability of currently available LLMs for Kazakh. We will publish the data collected for this study, which will be a good start for an LLM benchmark focused on Kazakh.

Downloads

Transcript English (automatic)

Next from ACL 2024

An Intelligent Tutor to Support Teaching and Learning of Tatar
workshop paper

An Intelligent Tutor to Support Teaching and Learning of Tatar

ACL 2024

+2Anisia Katinskaia
Alsu Zakirova and 4 other authors

15 August 2024

Stay up to date with the latest Underline news!

Select topic of interest (you can select more than one)

PRESENTATIONS

  • All Lectures
  • For Librarians
  • Resource Center
  • Free Trial
Underline Science, Inc.
1216 Broadway, 2nd Floor, New York, NY 10001, USA

© 2023 Underline - All rights reserved