Nigeria

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

Shamsuddeen Hassan Muhammad

,

David Ifeoluwa Adelani

,

Sebastian Ruder

,

Ibrahim Said Ahmad

,

et al.

January 21, 2022

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria (Hausa, Igbo, Nigerian-Pidgin, and Yorùbá ) consisting of around 30,000 annotated tweets per language (and 14,000 for Nigerian-Pidgin)

Download paper

Abstract

Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria (Hausa, Igbo, Nigerian-Pidgin, and Yorùbá ) consisting of around 30,000 annotated tweets per language (and 14,000 for Nigerian-Pidgin), including a significant fraction of code-mixed tweets. We propose text collection, filtering, processing and labeling methods that enable us to create datasets for these low-resource languages. We evaluate a range of pre-trained models and transfer strategies on the dataset. We find that language-specific models and language-adaptive fine-tuning generally perform best. We release the datasets, trained models, sentiment lexicons, and code to incentivize research on sentiment analysis in under-represented languages.

Paper: https://arxiv.org/abs/2201.08277

Github: https://github.com/hausanlp/NaijaSenti

‍

Authors

Shamsuddeen Hassan Muhammad

,

David Ifeoluwa Adelani

,

Sebastian Ruder

,

Ibrahim Said Ahmad

,

et al.

Countries

Nigeria

Related research

All research

All African countries

November 3, 2022

Rapport technique n°1: Conception de projets de données sensibles au genre

Alex Berryhill

,

Lorena Fuentes

,

All African countries

April 1, 2022

Rapport technique n°4: Relier les données sur le genre à l’action

Alex Berryhill

,

Lorena Fuentes

,

All African countries

November 4, 2021

Rapport technique n°3 : L’implication des parties prenantes pour une recherche en santé sensible au genre

Alex Berryhill

,

Lorena Fuentes

,

All African countries

November 2, 2021

Nota técnica 3: Participación de las partes interesadas en la investigación en salud con sensibilidad de género

Alex Berryhill

,

Lorena Fuentes

,

All African countries

April 6, 2022

Nota técnica 4: Conectar los datos de género con la acción

Alex Berryhill

,

Lorena Fuentes

,

All African countries

August 5, 2021

Rapport technique n°2: Guide pour une recherche en santé plus sensible au genre

Alex Berryhill

,

Lorena Fuentes

,

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

Abstract

Authors

Countries

Site links

Contact