textual documents. However, a large number of data sources provide multimedia documents (sound, images, audio-visual documents), for which the description techniques remain rudimentary, restricted to very specific types of sources (e.g. identity photos), and not very homogeneous because built according to particular needs. Our project consists in designing and experimenting generic and flexible techniques for content-based indexing and searching, dedicated to distributed sources of multimedia documents. The project relies on three complementary axes. 1. The first one aims at studying low level descriptors that may be automatically generated from multimedia documents, in order to use them as a support for content-based searching. By "low level descriptor", we mean vectors of values that characterize the content of a document, independently of any contextual information. We aim at characterizing these descriptors, as well as the extraction algorithms for producing them, as generically as possible, in order to cover a large palette of audio, video or audio-visual documents. Our goal in this axis is to exploit our complementary competences concerning the processing of non textual data. 2. The second axis consists in defining index structures and search operators for large collections of descriptors. By "index" we mean any structure (research tree) or technique (hashing) that allows restricting the research space, in order to avoid exhaustive exploring of a data collection. Here also we aim at factorizing as largely as possible techniques applicable to all the types of multimedia documents we considered. Our goal in this second axis is to complement the production of descriptors with the specification of a complete and generic toolkit for multimedia data processing. A content provider should be able to extend this toolkit in order to develop a search engine specific to its own collections. 3. The third axis concerns the distribution aspects of content-based search. We consider the case of institutions that wish to reference their collections and to benefit from a common indexing and searching system, based on the sharing of their descriptors. We intend to study in this axis the extension of the search structures and algorithms to the case of distributed sources. We will also exploit distribution to manage system scalability. The project also includes the implementation of a platform enabling to test on real data and in real environments all the technical proposals resulting from the three axes above. We do not include the problem of content distribution, which would raise problems related to access rights and ownership, but only that of references to content, each provider being free to define its own access rights policy. The consortium is composed of 5 partners – three public laboratories (Wisdom, INRIA Lille, IRCAM) and two content providers with different profiles: European Web Archive (archiving of free audio and audio-visual content collected on the Web) and the photo agency of the Réunion des Musées Nationaux (RMN), which will provide its collection of images. The three laboratories come with complementary competences on the management of audio documents (INRIA, IRCAM), images (Wisdom) and distributed search systems (Wisdom).
