In the Ninth Century, the rich Arabic tradition of adab finds its way to Spain, in al-Andalus, which then played a central role in knowledge exchange from the Orient and then relayed to the West, by monasteries from the North of the Iberian Peninsula in the 11th and 12th C. In al-Andalus, the adab literature meets the Jewish sapiential tradition of the midrashic literature. New collections are composed, including original works in the 10th and 11th centuries and from the 12th century on, exempla and philosophers’ sayings are translated into Hebrew, Latin, and Romance languages. Much of this complex heritage is found in the extensive Spanish paremiological literature, which is at its highest in the 16th and 17th centuries, and in current Spanish, Judeo-Spanish and Maghrebian collections of proverbs. Although the main lines of these exchanges are known, we lack specific information on the circulation of these short sapiential statements (our basic research unit), on the successive translating choices made by the translators, the cultural reinterpretations, or the weight of a borrowing over another. If sapiential textual filiations and translation sequences should be treated cautiously, this is particularly true for the sapiential statements contained in these texts. Due to the difficulty of understanding them, these volatile elements, whose categorization varies with time and considered cultures, have never been subject to overall textual studies, which would recount their sources, circulation and evolution through the different spoken or written languages by the three cultures within the Iberian Peninsula, during the Middle-Ages. The paremiological studies have principally produced compilations of proverbs (thesauri); editions; erudite studies dedicated to a single work, a single language or a single culture, except for D. Gutas’ remarkable groundbreaking work on the Philosophical Quartet (1975). The few existing databases take into account contemporary “paremiae” corpora, most often unilingual or with a traductology perspective. Therefore, the aim of the ALIENTO project is to calculate matches even when partial, close or distant connections in order to reassess inter-textual relations by comparing a great quantity of data and intersecting encoded texts written in different languages. This I why the project, which needs a close interdisciplinary collaboration between computational researchers (ATILF) and the linguists and specialists of literature (MSH Lorraine + INALCO and the international network of collaborators), will develop a computational software transferable to other similar texts using a large corpus of reference composed of 8 related texts which circulated in the Iberian Peninsula (in Latin, Arabic, Hebrew, Spanish and Catalan), representing 582 pages for a number of sapiential statements evaluated at 9,570 units. The developed software will extract and connect brief sapiential units through matching generated by the specific encoding system elaborated scientifically and written in an encoding manual XML-TEI. The choice and the type of annotations used result from a collaborative reflexion between the members of the project, specialists of linguistic paremiology, ancient texts, design engineers of textual databases, computational researchers during special scientific sessions. It will evolve in a collaborative manner during the matching processes. At the end we will have: - a body of texts belonging to a multilingual corpus, digitized, tagged in XML/TEI and publicly accessible, linked to a set of data on the text and its author. - a set of brief sapiential units with their XML/TEI annotations, accessible free of charge. - a trilingual questioning interface, making it possible to display the matched statements contained in these works, with information which can be used to study them regardless of the language. - an encoding methodology and a software for matching data transferable to other similar corpora.
