Resources
Datasets and tools I have built and released, mostly around Arabic NLP and sarcasm detection.
iSarcasmEval
-
The dataset from SemEval-2022 Task 6: iSarcasmEval, the first shared task on intended sarcasm detection in English and Arabic, co-located with NAACL 2022. Attracted 60 participating teams.
-
Unlike most sarcasm datasets, the data is author-annotated (first-party): authors label their own text as sarcastic or not, avoiding the noise of third-party annotation.
-
Covers English and Arabic, with three subtasks: sarcasm detection, sarcasm category classification, and pairwise sarcasm identification.
ArSarcasm-v2
-
Around 15.5K Arabic tweets labelled for sarcasm, sentiment, and dialect. An extension of the original ArSarcasm, and the standard dataset for the WANLP 2021 shared task on sarcasm and sentiment detection in Arabic. I recommend ArSarcasm-v2 over v1.
ArSarcasm (v1)
-
Around 10.5K Arabic tweets labelled for sarcasm, sentiment, and dialect, built by re-annotating existing Arabic sentiment analysis datasets.
Mazajak (archived)
-
A free online Arabic sentiment analysis tool and API, along with Arabic word embeddings for social media: word2vec vectors (CBOW and skip-gram) trained on 250M tweets.
-
The hosted service at the University of Edinburgh analysed more than 40M sentences, but it is no longer online.