What's New in the SLAIH Asset Catalog? Five Resources Worth Exploring
One of the main goals of SLAIH is to make it easier to discover high-quality resources for Slovak language AI. In this edition, we'd like to highlight a few recent additions that we believe deserve your attention. Whether you're building an LLM application, training an ASR system, or experimenting with multilingual models, these resources are well worth exploring.
MiniMax-M3: A Strong New Multimodal Foundation Model
One of the most exciting recent additions is MiniMax-M3, a state-of-the-art multimodal foundation model capable of processing text, images, and video. The model was released as an open-weight model for non-commercial use and has attracted significant attention thanks to its long-context capabilities and strong reasoning performance.
What makes MiniMax-M3 particularly interesting for our community is its impressive multilingual performance. Early evaluations indicate that it performs remarkably well on Slovak, making it a compelling option for developers looking beyond the usual closed-source alternatives.
If you're experimenting with multilingual AI assistants, document understanding, or multimodal applications, MiniMax-M3 is definitely a model to keep an eye on.
SkMTEB Datasets: More Than Just a Benchmark
Many of you have already heard about SkMTEB, the first comprehensive benchmark for Slovak text embeddings that was recently presented at ACL 2026. The benchmark evaluates embedding models across 31 datasets covering seven different task types, providing a standardized way to compare embedding models for Slovak.
What is sometimes overlooked, however, is that the datasets themselves are now available.
Instead of starting from scratch, developers now have access to a curated collection of Slovak evaluation datasets covering a wide range of real-world problems.
GEST: Measuring Gender Bias in Language Models
As AI systems become part of our everyday lives, evaluating their capabilities is no longer enough—we also need to understand their limitations.
GEST is a manually created benchmark designed to measure gender-stereotypical reasoning in language models and machine translation systems. Unlike many existing bias benchmarks that focus solely on English, GEST covers English and nine Slavic languages, including Slovak, making it a valuable resource for studying fairness in multilingual AI.
The dataset contains examples representing 16 common gender stereotypes—such as "Women are beautiful" or "Men are leaders"—that were carefully designed in collaboration with gender experts. It enables researchers to systematically evaluate whether models reinforce harmful stereotypes when generating or translating text.
SloPalSpeech: One of the Largest Slovak Speech Corpora
Speech technologies have seen enormous progress over the last few years, but high-quality training data remains one of the biggest bottlenecks for low-resource languages.
SloPalSpeech addresses this challenge by providing a large-scale Slovak speech corpus created from parliamentary recordings. The dataset contains more than 2,800 hours of aligned Slovak speech and transcripts, making it one of the largest publicly available Slovak speech datasets to date. If you're working on speech recognition for Slovak, this is one of the first resources you should look at.
Piper: High-Quality Open-Source Text-to-Speech
Another interesting addition to the catalog is Piper, an open-source neural text-to-speech system that supports Slovak alongside many other languages. Piper is lightweight, runs locally, and is suitable for applications where privacy, low latency, or offline deployment are important. For developers building voice assistants, accessibility tools, or spoken interfaces, Piper provides a practical open-source alternative that is easy to integrate into existing workflows.
Discover More
These five resources are only a small sample of what is available in the SLAIH catalog. New models, datasets, benchmarks, and tools are continuously being added as the Slovak AI ecosystem grows.
Our goal is for SLAIH to become the first place you visit when looking for Slovak language AI resources—whether you're a researcher searching for evaluation datasets, a developer looking for a production-ready model, or a company exploring AI solutions for Slovak.
We encourage you to browse the catalog, try out the resources, and let us know if there's a model, dataset, or tool that should be included. Together, we can continue building a stronger and more connected ecosystem for Slovak NLP and speech technologies.