Reassessing Claims of Human Parity and Super-Human Performance in Machine Translation at WMT 2019

Antonio Toral*

    We reassess the claims of human parity and super-human performance made at the news shared task of WMT2019 for three translation directions: English > German, English > Russian and German >English. First we identify three potential issues in the human evaluation of that shared task: (i) the limited amount of intersentential context available, (ii) the limited translation proficiency of the evaluators and (iii) the use of a reference translation. We then conduct a modified evaluation taking these issues into account. Our results indicate that all the claims of human parity and super-human performance made at WMT2019 should be refuted, except the claim of human parity for English > German. Based on our findings, we put forward a set of recommendations and open questions for future assessments of human parity in machine translation.
