Evaluating Multilingual LLMs at Scale

90K

human evaluations

30

models

10

Indian languages

3

weeks

Evaluation of multilingual LLMs is challenging due to insufficient linguistic diversity, benchmark contamination and the lack of local, cultural nuances in translated benchmarks. Karya’s data experts can evaluate models based on an array of benchmarks, including testing for linguistic acceptability, hallucinations, reasoning, and creativity.

Related