LSDC - A comprehensive dataset for Low Saxon Dialect Classification

Janine Siewert*, Yves Scherrer, Martijn Wieling, Jörg Tiedemann

*Bijbehorende auteur voor dit werk

    Onderzoeksoutput: Conference contributionAcademicpeer review

    37 Downloads (Pure)

    Samenvatting

    We present a new comprehensive dataset for the unstandardised West-Germanic language Low Saxon covering the last two centuries, the majority of modern dialects and various genres, which will be made openly available in connection with the final version of this paper. Since so far no such comprehensive dataset of contemporary Low Saxon exists, this provides a great contribution to NLP research on this language. We also test the use of this dataset for dialect classification by training a few baseline models comparing statistical and neural approaches. The performance of these models shows that in spite of an imbalance in the amount of data per dialect, enough features can be learned for a relatively high classification accuracy.
    Originele taal-2English
    TitelProceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialects
    RedacteurenMarcos Zampieri, Preslav Nakov, Nikola Ljubešić, Jörg Tiedemann, Yves Scherrer
    UitgeverijInternational Committee on Computational Linguistics (ICCL)
    Pagina's25-35
    Aantal pagina's11
    StatusPublished - 2020
    Evenement7th Workshop on NLP for Similar Languages, Varieties and Dialects - Online
    Duur: 13-dec-2020 → …

    Workshop

    Workshop7th Workshop on NLP for Similar Languages, Varieties and Dialects
    Verkorte titelVarDial 2020
    Periode13/12/2020 → …

    Citeer dit