Resources

Datasets and tools I have built and released, mostly around Arabic NLP and sarcasm detection.

iSarcasmEval

  • The dataset from SemEval-2022 Task 6: iSarcasmEval, the first shared task on intended sarcasm detection in English and Arabic, co-located with NAACL 2022. Attracted 60 participating teams.

  • Unlike most sarcasm datasets, the data is author-annotated (first-party): authors label their own text as sarcastic or not, avoiding the noise of third-party annotation.

  • Covers English and Arabic, with three subtasks: sarcasm detection, sarcasm category classification, and pairwise sarcasm identification.

  • Dataset on GitHub · Task overview paper, SemEval 2022


ArSarcasm-v2

ArSarcasm (v1)


Mazajak (archived)

  • A free online Arabic sentiment analysis tool and API, along with Arabic word embeddings for social media: word2vec vectors (CBOW and skip-gram) trained on 250M tweets.

  • The hosted service at the University of Edinburgh analysed more than 40M sentences, but it is no longer online.

  • Paper, WANLP at ACL 2019