Back to Catalog
MultiSocial
Dataset
we propose the first multilingual (22 languages) and multi-platform (5 social media platforms) dataset for benchmarking machine-generated text detection in the social-media domain, called MultiSocial. It contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs - aya-101, Mistral-7B-Instruct-v0.2, gpt-3.5-turbo-0125, v5-Eagle-7B-HF, vicuna-13b, opt-iml-max-30b, Llama-2-70b-chat-hf
Gated access
- Contributors
- Language
- sk and other
- Modality
- Text
- Domain
- News
- Task
- Machine-generated text detection