Forum for Science, Industry and Business

Sponsored by:     3M 
Search our Site:

 

Unpacking pecking orders to get the gist of web gab

08.06.2006


A USC Information Sciences Institute system pulls answers from online conversations by identifying the alpha chatterers.


The ISI study characterized online posts according a schema of speech acts. While some speech-act characterization was done by hand in this study to test the results, the ISI group has already developed effective software to accomplish the task automatically. Credit: USC Information Sciences Institute



The system, to be presented at a conference on human language technology on June 6, was developed to analyze technical conversations in which an objectively correct answer exists. But the method for statistically characterizing response by the group to individuals is generalizable.

Online communities are now firmly established in domains ranging from high school gossip to professional open-source software design discussions, generating huge repositories of records of human knowledge processing, pre-converted to digital form.


"For study of online natural language interaction, it’s the mother lode," says Eduard Hovy of the University of Southern California Information Sciences Institute.

Such sites provide raw material for a new method that may, among other things, enable Internet chat room users to get a statistical measurement of their influence in their room.

This research is one of the first quantitative studies in the field of natural language processing that takes account of the fact that chat conversations are structured interactions among a large number of people.

In the long term, research in this area will lead to the development of systems that can automatically produce reports and summaries of meetings, researchers hope.

It’s easy to simply harvest factoids from text, said Hovy (left), who holds an appointment as research associate professor in the USC Viterbi School of Engineering department of computer science in addition to his post as deputy director of the ISI Intelligent Systems Division and director of the ISI Natural Language Group.

But the fact that human conversation has an inherent structure, including temporal ordering, references to previous statements, labeled sourcing and other clues opens the door to much deeper machine-generated understanding.

To make use of the structure, the team used a graph-based algorithm called HITS (Hypertext Induced Topic Selection) originally used by Cornell computer scientist Joel Kleinberg to rank and classify web pages by their connections to each other.

In the study, connections between conversation participants replace the web links for the HITS analysis.

The interactions used in the study were threaded discussions from three semesters of a USC undergraduate course in computer science, including 2214 messages in 640 threads, all discussing class material and posing questions about problems.

The goal was to extract from the conversation the best answer to the questions discussed. And, according to the paper, the system works -- not perfectly, but much better than one that selects answers at random. Random selection got the answer (as determined by human inspection) right 87 out of 314 times, where the best implementation of the HITS system was correct 221 times.

The ISI implementation of HITS integrates three separate elements--speech act analysis, lexical similarity, and poster trustworthiness--to create links for interpretation for individual conversation participants.

Speech act analysis classifies the statements in the record according to what they do in the context of the discussion, assigning each to one of thirteen kinds of acts, grouped in three categories: inform, request, social-interaction.

The "inform" speech act category includes corrections, descriptions, elaborations, suggestions, and answers to questions, both simple and complex. "Requests" include not just requests for information but also for action, namely commands. "Social" speech acts include acknowledgements, thanks, compliments, criticisms, objections, and supportive statements.

Lexical analysis looks for similarities in the vocabulary of responses to see which are related to each other. From this the system can determine the threads of the conversation, and decide when new subtopics are split off.

Finally, poster trustworthiness measures the degree to which participants accept statements made by each individual. This is determined by scoring responses to a given person’s posts as either negative or positive. Over time, people whose statements are more positively viewed become more central and more trusted in the online community.

To test the method, part of the data (the classification of the speech acts) was initially human coded. After it was trained, the machine system was then applied to the same data, and its performance was compared to that of the human coder. It achieved accuracy of between 65% and 70% -- a figure that is likely to improve.

How soon will it be possible to download a version that can score a given poster’s influence in his/her chat community? "This technology has considerable potential for commercialization," said Hovy.

Eric Mankin | EurekAlert!
Further information:
http://www.usc.edu

More articles from Information Technology:

nachricht A novel hybrid UAV that may change the way people operate drones
28.03.2017 | Science China Press

nachricht Timing a space laser with a NASA-style stopwatch
28.03.2017 | NASA/Goddard Space Flight Center

All articles from Information Technology >>>

The most recent press releases about innovation >>>

Die letzten 5 Focus-News des innovations-reports im Überblick:

Im Focus: A Challenging European Research Project to Develop New Tiny Microscopes

The Institute of Semiconductor Technology and the Institute of Physical and Theoretical Chemistry, both members of the Laboratory for Emerging Nanometrology (LENA), at Technische Universität Braunschweig are partners in a new European research project entitled ChipScope, which aims to develop a completely new and extremely small optical microscope capable of observing the interior of living cells in real time. A consortium of 7 partners from 5 countries will tackle this issue with very ambitious objectives during a four-year research program.

To demonstrate the usefulness of this new scientific tool, at the end of the project the developed chip-sized microscope will be used to observe in real-time...

Im Focus: Giant Magnetic Fields in the Universe

Astronomers from Bonn and Tautenburg in Thuringia (Germany) used the 100-m radio telescope at Effelsberg to observe several galaxy clusters. At the edges of these large accumulations of dark matter, stellar systems (galaxies), hot gas, and charged particles, they found magnetic fields that are exceptionally ordered over distances of many million light years. This makes them the most extended magnetic fields in the universe known so far.

The results will be published on March 22 in the journal „Astronomy & Astrophysics“.

Galaxy clusters are the largest gravitationally bound structures in the universe. With a typical extent of about 10 million light years, i.e. 100 times the...

Im Focus: Tracing down linear ubiquitination

Researchers at the Goethe University Frankfurt, together with partners from the University of Tübingen in Germany and Queen Mary University as well as Francis Crick Institute from London (UK) have developed a novel technology to decipher the secret ubiquitin code.

Ubiquitin is a small protein that can be linked to other cellular proteins, thereby controlling and modulating their functions. The attachment occurs in many...

Im Focus: Perovskite edges can be tuned for optoelectronic performance

Layered 2D material improves efficiency for solar cells and LEDs

In the eternal search for next generation high-efficiency solar cells and LEDs, scientists at Los Alamos National Laboratory and their partners are creating...

Im Focus: Polymer-coated silicon nanosheets as alternative to graphene: A perfect team for nanoelectronics

Silicon nanosheets are thin, two-dimensional layers with exceptional optoelectronic properties very similar to those of graphene. Albeit, the nanosheets are less stable. Now researchers at the Technical University of Munich (TUM) have, for the first time ever, produced a composite material combining silicon nanosheets and a polymer that is both UV-resistant and easy to process. This brings the scientists a significant step closer to industrial applications like flexible displays and photosensors.

Silicon nanosheets are thin, two-dimensional layers with exceptional optoelectronic properties very similar to those of graphene. Albeit, the nanosheets are...

All Focus news of the innovation-report >>>

Anzeige

Anzeige

Event News

International Land Use Symposium ILUS 2017: Call for Abstracts and Registration open

20.03.2017 | Event News

CONNECT 2017: International congress on connective tissue

14.03.2017 | Event News

ICTM Conference: Turbine Construction between Big Data and Additive Manufacturing

07.03.2017 | Event News

 
Latest News

'On-off switch' brings researchers a step closer to potential HIV vaccine

30.03.2017 | Health and Medicine

Penn studies find promise for innovations in liquid biopsies

30.03.2017 | Health and Medicine

An LED-based device for imaging radiation induced skin damage

30.03.2017 | Medical Engineering

VideoLinks
B2B-VideoLinks
More VideoLinks >>>