WikiAnn is a dataset for cross-lingual name tagging and linking based on Wikipedia articles in 295 languages.
54 PAPERS • 7 BENCHMARKS
MasakhaNER is a collection of Named Entity Recognition (NER) datasets for 10 different African languages. The languages forming this dataset are: Amharic, Hausa, Igbo, Kinyarwanda, Luganda, Luo, Nigerian-Pidgin, Swahili, Wolof, and Yorùbá.
45 PAPERS • 2 BENCHMARKS