Back to Catalog

MultiSocial

Dataset

we propose the first multilingual (22 languages) and multi-platform (5 social media platforms) dataset for benchmarking machine-generated text detection in the social-media domain, called MultiSocial. It contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs - aya-101, Mistral-7B-Instruct-v0.2, gpt-3.5-turbo-0125, v5-Eagle-7B-HF, vicuna-13b, opt-iml-max-30b, Llama-2-70b-chat-hf

Gated access
Contributors
Language
sk and other
Modality
Text
Domain
News
Task
Machine-generated text detection