Back to Catalog

skMTEB

Dataset

SkMTEB is the first comprehensive MTEB-style text embedding benchmark for Slovak — a low-resource West Slavic language with ~5 million speakers. The benchmark comprises 31 datasets across 7 task types, covering nearly 4× the depth of existing multilingual benchmark coverage for Slovak (compared to 8 tasks in MMTEB).

Contributors
SlovakNLP Community
License
CC-BY-4.0
Language
sk
Modality
Text
Domain
Other
Task
Benchmarking