Forum for Science, Industry and Business

Sponsored by:     3M 
Search our Site:

 

Pinning Down the Fleeting Internet: Web Crawler Archives Historical Data for Easy Searching

19.11.2008
University of Washington researchers are grabbing hold of the fleeting Web and storing historical Web sites that users can easily search using an intuitive application called Zoetrope.

The Internet contains vast amounts of information, much of it unorganized. But what you see online at any given moment is just a snapshot of the Web as a whole – many pages change rapidly or disappear completely, and the old data gets lost forever.

“Your browser is really just a window into the Web as it exists today,” said Eytan Adar, University of Washington computer science and engineering doctoral student. “When you search for something online you’re only getting today’s results.”

Now, Adar and his colleagues at UW and Adobe Systems Inc. are grabbing hold of the fleeting Web and storing historical sites that users can easily search using an intuitive application called Zoetrope.

“There are so many ways of finding and manipulating and visualizing data on what we call ‘the today Web’ that it’s kind of amazing that there’s no way to do anything similar to the ephemeral Web,” said Dan Weld, a UW computer science and engineering professor who also worked on the application. One service, the Internet Archive, has been capturing old versions of Web sites for years, but the records for the stored sites are inconsistent, Weld said. More importantly, there’s no easy way to search the archive.

With Zoetrope, anyone will be able to use easy keyword searches to find archived Web information or look for patterns over time. The research was presented Oct. 22 by Mira Dontcheva, the system’s co-creator and a recently graduated UW computer science and engineering doctoral student now at Adobe Systems Inc., at the ACM Symposium on User Interface Software and Technology in Monterey, Calif.

There are a variety of ways people might want to search the historical Internet. For example, to find a history of traffic patterns in the Seattle area, you’d have to sort through lengthy PDF files from the state Department of Transportation, Adar said. With Zoetrope, you could easily view past versions of any traffic Web site, and getting more specific, search for drive-times on Interstate 90 at 6 p.m. on rainy Fridays. Zoetrope can also capture and help analyze information that might otherwise not be available anywhere.

Sports fanatics could use the program to check historical rankings of their favorite teams or players, information that currently may not be easy to find. The application can do more than just simple keyword searches, Adar said. It also can be used to analyze historical data or link information from different sites. For example, Adar wondered whether air pollution conditions could affect the performance of Olympic athletes, so he used Zoetrope to find daily records of pollution levels in Beijing and the number of world records broken in the 2008 Olympics on each day, and looked to see whether fewer records were broken on days with high pollution levels.

“Zoetrope is aimed at the casual researcher,” Weld said. “It’s really for anyone who has a question.”

Zoetrope could eventually be built in to any other Web browser, Adar said. If you just want to browse the past versions of a given site, you drag a slider backwards to see older and older versions. Alternatively, you can draw a box around just one part of the site, if you’re interested in, say, the lead story on CNN.com but don’t care about the rest of the page. These boxes can be filtered by keyword searches or date, so you could look only for lead stories featuring Hollywood actors or stories that ran on Fridays.

Users can view historical data by moving the slider, but more sophisticated analyses are available as well. If you’re looking at something numerical, such as gas prices over time, the program can draw graphs for you. Or you can pull out images from specific times, such as traffic pictures, and compare them all side by side. These kinds of visualizations can be further organized in a timeline or by clustering – Zoetrope can make an image comparing traffic patterns on sunny days versus cloudy days, for example.

Right now, Zoetrope saves a new version of approximately 1,000 different sites every hour, Adar said. It’s been running for four months, so records go no further than that, but Adar hopes to eventually incorporate information from the Internet Archive’s nearly 14 years of records into the program.

He wants to figure out how to scale the program up from 1,000 Web pages to all pages in existence, and has run studies to figure how often each page would need to be captured. For example, a traffic site or stock-watching page would need versions saved much more often than every hour, but there are many unchanging pages that could be archived less frequently. Eventually, Zoetrope could automatically figure out how often to capture a page based on how frequently it changes, Adar said.

“This is really a new way to think about storing information on the Web,” he said.

The researchers hope to release Zoetrope free, and say it may be available as early as next summer.

The National Science Foundation, the Achievement Rewards for College Scientists Foundation and the Washington Research Foundation provided funding for Zoetrope. James Fogarty, UW computer science and engineering assistant professor, also worked on the application.

For more information, contact Eytan Adar at eadar@cs.washington.edu or (650) 799-8823, or Weld at weld@cs.washington.edu or (206) 543-9196.

Rachel Tompa | Newswise Science News
Further information:
http://www.washington.edu
http://uwnews.org/article.asp?articleID=45255

More articles from Information Technology:

nachricht Controlling robots with brainwaves and hand gestures
20.06.2018 | Massachusetts Institute of Technology, CSAIL

nachricht Innovative autonomous system for identifying schools of fish
20.06.2018 | IMDEA Networks Institute

All articles from Information Technology >>>

The most recent press releases about innovation >>>

Die letzten 5 Focus-News des innovations-reports im Überblick:

Im Focus: Temperature-controlled fiber-optic light source with liquid core

In a recent publication in the renowned journal Optica, scientists of Leibniz-Institute of Photonic Technology (Leibniz IPHT) in Jena showed that they can accurately control the optical properties of liquid-core fiber lasers and therefore their spectral band width by temperature and pressure tuning.

Already last year, the researchers provided experimental proof of a new dynamic of hybrid solitons– temporally and spectrally stationary light waves resulting...

Im Focus: Overdosing on Calcium

Nano crystals impact stem cell fate during bone formation

Scientists from the University of Freiburg and the University of Basel identified a master regulator for bone regeneration. Prasad Shastri, Professor of...

Im Focus: AchemAsia 2019 will take place in Shanghai

Moving into its fourth decade, AchemAsia is setting out for new horizons: The International Expo and Innovation Forum for Sustainable Chemical Production will take place from 21-23 May 2019 in Shanghai, China. With an updated event profile, the eleventh edition focusses on topics that are especially relevant for the Chinese process industry, putting a strong emphasis on sustainability and innovation.

Founded in 1989 as a spin-off of ACHEMA to cater to the needs of China’s then developing industry, AchemAsia has since grown into a platform where the latest...

Im Focus: First real-time test of Li-Fi utilization for the industrial Internet of Things

The BMBF-funded OWICELLS project was successfully completed with a final presentation at the BMW plant in Munich. The presentation demonstrated a Li-Fi communication with a mobile robot, while the robot carried out usual production processes (welding, moving and testing parts) in a 5x5m² production cell. The robust, optical wireless transmission is based on spatial diversity; in other words, data is sent and received simultaneously by several LEDs and several photodiodes. The system can transmit data at more than 100 Mbit/s and five milliseconds latency.

Modern production technologies in the automobile industry must become more flexible in order to fulfil individual customer requirements.

Im Focus: Sharp images with flexible fibers

An international team of scientists has discovered a new way to transfer image information through multimodal fibers with almost no distortion - even if the fiber is bent. The results of the study, to which scientist from the Leibniz-Institute of Photonic Technology Jena (Leibniz IPHT) contributed, were published on 6thJune in the highly-cited journal Physical Review Letters.

Endoscopes allow doctors to see into a patient’s body like through a keyhole. Typically, the images are transmitted via a bundle of several hundreds of optical...

All Focus news of the innovation-report >>>

Anzeige

Anzeige

VideoLinks
Industry & Economy
Event News

Munich conference on asteroid detection, tracking and defense

13.06.2018 | Event News

2nd International Baltic Earth Conference in Denmark: “The Baltic Sea region in Transition”

08.06.2018 | Event News

ISEKI_Food 2018: Conference with Holistic View of Food Production

05.06.2018 | Event News

 
Latest News

Graphene assembled film shows higher thermal conductivity than graphite film

22.06.2018 | Materials Sciences

Fast rising bedrock below West Antarctica reveals an extremely fluid Earth mantle

22.06.2018 | Earth Sciences

Zebrafish's near 360 degree UV-vision knocks stripes off Google Street View

22.06.2018 | Life Sciences

VideoLinks
Science & Research
Overview of more VideoLinks >>>