The project Digitizing Armenian Linguistic Heritage: Armenian Multivariational Corpus and Data Processing (DALiH) aims at building for the first time an open-access and open-source unified digital linguistic platform for the whole spectrum of Armenian language variation. Each language variety will be represented by a comprehensive text database which will be provided with full morphological annotation. More particularly, DALiH will design 1) a Classical Armenian corpus; 2) a Modern Western Armenian corpus; 3) a pilot corpus of Middle Armenian 4) three pilot corpora of dialects, and 5) an updated Modern Eastern Armenian annotated corpus. Deep-learning and rule-based natural language processing resources will be designed in order to process the databases, to develop grammatical annotation and Automatic speech recognition models and to cross-check their value for further corpus enlargement, in a context of multiparameter language variation for an under-resourced language.
