A First Dataset for Film Age Appropriateness Investigation

LREC 2020  ·  Emad Mohamed, Le An Ha ·

Film age appropriateness classification is an important problem with a significant societal impact that has so far been out of the interest of Natural Language Processing and Machine Learning researchers. To this end, we have collected a corpus of 17000 films along with their age ratings. We use the textual contents in an experiment to predict the correct age classification for the United States (G, PG, PG-13, R and NC-17) and the United Kingdom (U, PG, 12A, 15, 18 and R18). Our experiments indicate that gradient boosting machines beat FastText and various Deep Learning architectures. We reach an overall accuracy of 79.3{\%} for the US ratings compared to a projected super human accuracy of 84{\%}. For the UK ratings, we reach an overall accuracy of 65.3{\%} (UK) compared to a projected super human accuracy of 80.0{\%}.

PDF Abstract

Datasets


  Add Datasets introduced or used in this paper

Results from the Paper


  Submit results from this paper to get state-of-the-art GitHub badges and help the community compare results to other papers.

Methods