A little known secret in data mining is that simply feeding raw data into a data analysis algorithm is unlikely to produce meaningful results, say the authors of a new Cornell University study.
From recognizing speech to identifying unusual stars, new discoveries often begin with comparison of data streams to find connections and spot outliers.
But most data comparison algorithms today have one major weakness – somewhere, they rely on a human expert to specify what aspects of the data are relevant for comparison, and what aspects aren't. But experts aren't keeping pace with the growing amounts and complexities of big data.
Cornell computing researchers have come up with a new principle they call "data smashing" for estimating the similarities between streams of arbitrary data without human intervention, and without access to the data sources. Hod Lipson, associate professor of mechanical engineering and computing and information science, and Ishanu Chattopadhyay, a former postdoctoral associate with Lipson and now at the University of Chicago, have described their method in Royal Society Interface, Oct. 1.
Data smashing is based on a new way to compare data streams. The process involves two steps. First, the data streams are algorithmically "smashed" to "annihilate" the information in each other. Then, the process measures what information remained after the collision. The more information remained, the less likely the streams originated in the same source.
Data smashing principles may open the door to understanding increasingly complex observations, especially when experts do not know what to look for, according to the researchers.
The authors demonstrated the application of their principle to data from real-world problems, including the disambiguation of electroencephalograph patterns from epileptic seizure patients; detection of anomalous cardiac activity from heart recordings; and classification of astronomical objects from raw photometry.
In all cases and without access to original domain knowledge, the researchers demonstrated performance on par with the accuracy of specialized algorithms and heuristics devised by experts.
The work in the paper, "Data smashing: Uncovering lurking order in data," was supported by the Defense Advanced Research Projects Agency and the U.S. Army Research Office.
Syl Kacapyr | Eurek Alert!
Diagnoses: When Are Several Opinions Better Than One?
19.07.2016 | Max-Planck-Institut für Bildungsforschung
High in calories and low in nutrients when adolescents share pictures of food online
07.04.2016 | University of Gothenburg
Researchers from the Institute for Quantum Computing (IQC) at the University of Waterloo led the development of a new extensible wiring technique capable of controlling superconducting quantum bits, representing a significant step towards to the realization of a scalable quantum computer.
"The quantum socket is a wiring method that uses three-dimensional wires based on spring-loaded pins to address individual qubits," said Jeremy Béjanin, a PhD...
In a paper in Scientific Reports, a research team at Worcester Polytechnic Institute describes a novel light-activated phenomenon that could become the basis for applications as diverse as microscopic robotic grippers and more efficient solar cells.
A research team at Worcester Polytechnic Institute (WPI) has developed a revolutionary, light-activated semiconductor nanocomposite material that can be used...
By forcefully embedding two silicon atoms in a diamond matrix, Sandia researchers have demonstrated for the first time on a single chip all the components needed to create a quantum bridge to link quantum computers together.
"People have already built small quantum computers," says Sandia researcher Ryan Camacho. "Maybe the first useful one won't be a single giant quantum computer...
COMPAMED has become the leading international marketplace for suppliers of medical manufacturing. The trade fair, which takes place every November and is co-located to MEDICA in Dusseldorf, has been steadily growing over the past years and shows that medical technology remains a rapidly growing market.
In 2016, the joint pavilion by the IVAM Microtechnology Network, the Product Market “High-tech for Medical Devices”, will be located in Hall 8a again and will...
'Ferroelectric' materials can switch between different states of electrical polarization in response to an external electric field. This flexibility means they show promise for many applications, for example in electronic devices and computer memory. Current ferroelectric materials are highly valued for their thermal and chemical stability and rapid electro-mechanical responses, but creating a material that is scalable down to the tiny sizes needed for technologies like silicon-based semiconductors (Si-based CMOS) has proven challenging.
Now, Hiroshi Funakubo and co-workers at the Tokyo Institute of Technology, in collaboration with researchers across Japan, have conducted experiments to...
14.10.2016 | Event News
14.10.2016 | Event News
12.10.2016 | Event News
21.10.2016 | Health and Medicine
21.10.2016 | Information Technology
21.10.2016 | Materials Sciences