DILIGENT Data Challenge on EGEE Infrastructure

The DILIGENT team used the EGEE computing Grid to process 37 million images from the online Flickr database in just 16 weeks. This computation generated approximately 112 million text and image objects—nearly 5 TB of data—containing more than 150 million extracted features. This is equivalent to an average processing capacity of over 300,000 images per day.

This unique collection will be used by the SAPIR project to develop new large-scale content-based data retrieval and automatic data classification techniques that combine both text and image content, expanding the limits of conventional search engines, which can only search text associated to images and audio-visual content.

The computational load required to generate this massive data collection was outsourced to DILIGENT, and then delegated to the EGEE Pre-Production Service (PPS) Grid infrastructure via the gLite middleware. A total of 44,333 gLite jobs were successfully executed by the EGEE PPS infrastructure resource broker. Each job processed approximately 1000 images.

The data challenge lasted for 116 days, from 16 June to 9 October 2007, and was organized in three different phases. During the initial preparation phase experimental jobs were submitted to some EGEE PPS sites to test the feature extraction application and optimize the number of images to process per day.

The next two phases involved actual execution of the data challenge, exploiting ten EGEE PPS sites that contributed their computational resources: University of Athens, Scuola Normale Superiore, ISTI-CNR, LIP, ESA-ESRIN, CERN, CESGA, University of Macedonia, Ben Gurion University, and CYFRONET. Four of these sites are maintained by DILIGENT partners.

Media Contact

Sarah Purcell alfa

All latest news from the category: Information Technology

Here you can find a summary of innovations in the fields of information and data processing and up-to-date developments on IT equipment and hardware.

This area covers topics such as IT services, IT architectures, IT management and telecommunications.

Back to home

Comments (0)

Write a comment

Newest articles

Creating good friction: Pitt engineers aim to make floors less slippery

Swanson School collaborators Kurt Beschorner and Tevis Jacobs will use a NIOSH award to measure floor-surface topography and create a predictive model of friction. Friction is the resistance to motion…

Synthetic tissue can repair hearts, muscles, and vocal cords

Scientists from McGill University develop new biomaterial for wound repair. Combining knowledge of chemistry, physics, biology, and engineering, scientists from McGill University develop a biomaterial tough enough to repair the…

Constraining quantum measurement

The quantum world and our everyday world are very different places. In a publication that appeared as the “Editor’s Suggestion” in Physical Review A this week, UvA physicists Jasper van…

Partners & Sponsors