Slovak STS Synthetic
This is a synthetic Slovak dataset intended for the semantic textual similarity (STS) task. Each pair consists of a sentence from the Slovak Summarization dataset and a generated sentence produced by GPT-5. GPT-5 was instructed to generate six sentences per original sentence (from SlovakSum) to each similarity score (0-5). Each original sentence had to contain at least 60 characters and no more than 200 characters, and they had to consist of a single sentence. A total of 6 594 sentence pairs were generated. After the generation of sentences, a random sample of 300 sentence pairs was selected from the dataset. Fifty pairs were drawn from each similarity score. The sample was annotated by three annotators. Each annotator received information about the task and explanations for each STS score, as well as five example sentences with explanations in the annotation guidelines. All three annotators independently annotated the same set of 300 sentence pairs. The final dataset is divided into two parts: a train split consisting of the non-annotated part and the test split consisting of the annotated part. Train split: { "sentence1": "Vrtuľník lietal nad národným parkom bez povolenia ochranárov, sťažovať sa chcú aj obce", "sentence2": "Vrtuľník lietal nad národným parkom bez súhlasu ochranárov a obce sa chcú tiež sťažovať.", "similarity_score": "4.6666666667" }
- Contributors
- License
- CC-BY-NC-4.0
- Language
- sk
- Modality
- Text
- Task
- Semantic Textual Similarity