Semantic Role Labeling with Pretrained Language Models for Known and Unknown Predicates

RANLP 2019 · Daniil Larionov, Artem Shelmanov, Elena Chistova, Ivan Smirnov ·

We build the first full pipeline for semantic role labelling of Russian texts. The pipeline implements predicate identification, argument extraction, argument classification (labeling), and global scoring via integer linear programming. We train supervised neural network models for argument classification using Russian semantically annotated corpus {--} FrameBank. However, we note that this resource provides annotations only to a very limited set of predicates. We combat the problem of annotation scarcity by introducing two models that rely on different sets of features: one for {``}known{''} predicates that are present in the training set and one for {``}unknown{''} predicates that are not. We show that the model for {``}unknown{''} predicates can alleviate the lack of annotation by using pretrained embeddings. We perform experiments with various types of embeddings including the ones generated by deep pretrained language models: word2vec, FastText, ELMo, BERT, and show that embeddings generated by deep pretrained language models are superior to classical shallow embeddings for argument classification of both {``}known{''} and {``}unknown{''} predicates.

PDF Abstract