Forum for Science, Industry and Business

Sponsored by:     3M 
Search our Site:

 

New search engine ranks tables by title, document content, text reference

13.08.2007
Penn State researchers have developed a search engine-TableSeer-which not only can identify and extract tables from PDF documents but also can index and rank the search results using factors including the table's title, text references to the table and date of publication.

The engine's innovative ranking algorithm, TableRank, also can identify tables found in frequently cited documents and weigh that factor as well in the search results, said Prasenjit Mitra, an assistant professor in the Penn State College of Information Sciences and Technology (IST) and one of the lead researchers in the development of the search engine.

"TableSeer makes it easier for scientists and scholars to find and access the important information presented in tables, and as far as we know, is the first search engine for tables," Mitra said.

Tables are an important data resource for researchers. In a search of 10,000 documents from s and conferences, the researchers found that more than 70 percent of papers in chemistry, biology and computer science included tables. Furthermore, most of those documents had multiple tables.

But while some software can identify and extract tables from text, existing software cannot search for tables across documents. That means scientists and scholars must manually browse documents in order to find tables-a time-consuming and cumbersome process.

TableSeer automates that process and captures data not only within the table but also in tables' titles and footnotes. In addition, it enables column-name-based search so that a user can search for a particular column in a table.

In tests with documents from the Royal Society of Chemistry, TableSeer correctly identified and retrieved 93.5 percent of tables created in text-based formats, Mitra said.

Searching for tables has some unique challenges, as there is no standard table representation, so tables can appear in PDF, PowerPoint, HTML and Microsoft Word documents. The researchers chose to focus on PDF documents because of their growing popularity in digital libraries and because PDF documents had been overlooked in other table-search efforts.

"Tables can be made using a number of editor tools, and the techniques we are using in TableSeer should work with any text-based tool," said C. Lee Giles, professor of information sciences and technology and co-director of the IST Cyber-Infrastructure Lab where the research originated. "While we designed and developed TableSeer to facilitate searching of tables occurring in articles in the chemistry domain, it can be used in any domain where data is presented in tabular form including other scientific, technical, social and business areas."

The development of TableSeer is part of an open-source cyber-infrastructure project focusing on chemical document search for environmental chemistry and funded by the National Science Foundation. The grant awarded to the Penn State Department of Chemistry aims to enable automatic data analysis.

"Searching and extracting information from data tables is an essential component of data analysis in environmental science, where many research groups publish large amounts of kinetic data describing chemical changes in the environment," said Karl Mueller, professor of chemistry and principal investigator for the NSF grant.

"As we approach multidisciplinary problems within the Penn State Center for Environmental Kinetics Analysis, our students spend many days hunting down and compiling large amount of data from tables. The TableSeer tools will definitely increase the efficiency of this process and allow more time to be spent on creative scientific analysis," he added.

TableSeer can be tested online (see http://chemxseer.ist.psu.edu). The source code will be made available near the completion of the project, the researchers said.

In the meantime, research is ongoing to improve the ranking algorithm by adding additional features. The researchers also are working on a search engine that can identify, extract and rank figures found in documents, as figures are another important device for disseminating data and findings in the natural sciences.

Margaret Hopkins | EurekAlert!
Further information:
http://chemxseer.ist.psu.edu
http://www.psu.edu

More articles from Information Technology:

nachricht New technology enables 5-D imaging in live animals, humans
16.01.2017 | University of Southern California

nachricht Fraunhofer FIT announces CloudTeams collaborative software development platform – join it for free
10.01.2017 | Fraunhofer-Institut für Angewandte Informationstechnik FIT

All articles from Information Technology >>>

The most recent press releases about innovation >>>

Die letzten 5 Focus-News des innovations-reports im Überblick:

Im Focus: Interfacial Superconductivity: Magnetic and superconducting order revealed simultaneously

Researchers from the University of Hamburg in Germany, in collaboration with colleagues from the University of Aarhus in Denmark, have synthesized a new superconducting material by growing a few layers of an antiferromagnetic transition-metal chalcogenide on a bismuth-based topological insulator, both being non-superconducting materials.

While superconductivity and magnetism are generally believed to be mutually exclusive, surprisingly, in this new material, superconducting correlations...

Im Focus: Studying fundamental particles in materials

Laser-driving of semimetals allows creating novel quasiparticle states within condensed matter systems and switching between different states on ultrafast time scales

Studying properties of fundamental particles in condensed matter systems is a promising approach to quantum field theory. Quasiparticles offer the opportunity...

Im Focus: Designing Architecture with Solar Building Envelopes

Among the general public, solar thermal energy is currently associated with dark blue, rectangular collectors on building roofs. Technologies are needed for aesthetically high quality architecture which offer the architect more room for manoeuvre when it comes to low- and plus-energy buildings. With the “ArKol” project, researchers at Fraunhofer ISE together with partners are currently developing two façade collectors for solar thermal energy generation, which permit a high degree of design flexibility: a strip collector for opaque façade sections and a solar thermal blind for transparent sections. The current state of the two developments will be presented at the BAU 2017 trade fair.

As part of the “ArKol – development of architecturally highly integrated façade collectors with heat pipes” project, Fraunhofer ISE together with its partners...

Im Focus: How to inflate a hardened concrete shell with a weight of 80 t

At TU Wien, an alternative for resource intensive formwork for the construction of concrete domes was developed. It is now used in a test dome for the Austrian Federal Railways Infrastructure (ÖBB Infrastruktur).

Concrete shells are efficient structures, but not very resource efficient. The formwork for the construction of concrete domes alone requires a high amount of...

Im Focus: Bacterial Pac Man molecule snaps at sugar

Many pathogens use certain sugar compounds from their host to help conceal themselves against the immune system. Scientists at the University of Bonn have now, in cooperation with researchers at the University of York in the United Kingdom, analyzed the dynamics of a bacterial molecule that is involved in this process. They demonstrate that the protein grabs onto the sugar molecule with a Pac Man-like chewing motion and holds it until it can be used. Their results could help design therapeutics that could make the protein poorer at grabbing and holding and hence compromise the pathogen in the host. The study has now been published in “Biophysical Journal”.

The cells of the mouth, nose and intestinal mucosa produce large quantities of a chemical called sialic acid. Many bacteria possess a special transport system...

All Focus news of the innovation-report >>>

Anzeige

Anzeige

Event News

12V, 48V, high-voltage – trends in E/E automotive architecture

10.01.2017 | Event News

2nd Conference on Non-Textual Information on 10 and 11 May 2017 in Hannover

09.01.2017 | Event News

Nothing will happen without batteries making it happen!

05.01.2017 | Event News

 
Latest News

Water - as the underlying driver of the Earth’s carbon cycle

17.01.2017 | Earth Sciences

Interfacial Superconductivity: Magnetic and superconducting order revealed simultaneously

17.01.2017 | Materials Sciences

Smart homes will “LISTEN” to your voice

17.01.2017 | Architecture and Construction

VideoLinks
B2B-VideoLinks
More VideoLinks >>>