Turning vast amounts of genomic data into meaningful information about the cell is the great challenge of bioinformatics, with major implications for human biology and medicine.
Researchers at the University of California, San Diego School of Medicine and colleagues have proposed a new method that creates a computational model of the cell from large networks of gene and protein interactions, discovering how genes and proteins connect to form higher-level cellular machinery.
The findings are published in the December 16 advance online publication of Nature Biotechnology.
"Our method creates ontology, or a specification of all the major players in the cell and the relationships between them," said first author Janusz Dutkowski, PhD, postdoctoral researcher in the UC San Diego Department of Medicine. It uses knowledge about how genes and proteins interact with each other and automatically organizes this information to form a comprehensive catalog of gene functions, cellular components, and processes.
"What's new about our ontology is that it is created automatically from large datasets. In this way, we see not only what is already known, but also potentially new biological components and processes – the bases for new hypotheses," said Dutkowski.
Originally devised by philosophers attempting to explain the nature of existence, ontologies are now broadly used to encapsulate everything known about a subject in a hierarchy of terms and relationships. Intelligent information systems, such as iPhone's Siri, are built on ontologies to enable reasoning about the real world. Ontologies are also used by scientists to structure knowledge about subjects like taxonomy, anatomy and development, bioactive compounds, disease and clinical diagnosis.
A Gene Ontology (GO) exists as well, constructed over the last decade through a joint effort of hundreds of scientists. It is considered the gold standard for understanding cell structure and gene function, containing 34,765 terms and 64,635 hierarchical relations annotating genes from more than 80 species.
"GO is very influential in biology and bioinformatics, but it is also incomplete and hard to update based on new data," said senior author Trey Ideker, PhD, chief of the Division of Genetics in the School of Medicine and professor of bioengineering in UC San Diego's Jacobs School of Engineering.
"This is expert knowledge based upon the work of many people over many, many years," said Ideker, who is also principal investigator of the National Resource for Network Biology, based at UC San Diego. "A fundamental problem is consistency. People do things in different ways, and that impacts what findings are incorporated into GO and how they relate to other findings. The approach we have proposed is a more objective way to determine what's known and uncover what's new."
In their paper, Dutkowski, Ideker and colleagues capitalized upon the growing power and utility of new technologies like high-throughput assays and bioinformatics to create elaborately detailed datasets describing complex biological networks. To test the approach, the scientists pulled together multiple such datasets, applied their method, and then compared the resulting "network-extracted ontology" to the existing GO.
They found that their ontology captured the majority of known cellular components, plus many additional terms and relationships, which subsequently triggered updates of the existing GO.
Neither Ideker nor Dutkowski say the new approach is intended to replace the current GO. Rather, they envision it as complementary high-tech model that identifies both known and uncharacterized biological components derived directly from data, something the current GO does not do well. Moreover, they note a network-extracted ontology can be continuously updated and refined with every new dataset, moving scientists closer to the complete model of the cell.
Co-authors are Michael Kramer, UCSD Departments of Medicine and Bioengineering; Michal A. Surma, Max Planck Institute of Molecular Cell Biology and Genetics, Dresden, Germany and UCSF Department of Cellular and Molecular Pharmacology; Rama Balakrishnan and J. Michael Cherry, Department of Genetics, Stanford University; and Nevan J. Krogan, UCSF Department of Cellular and Molecular Pharmacology and J. David Gladstone Institutes, San Francisco.
Funding for this research came, in part, from National Institutes of Health grants (P41-GM103504, P50-GM085764, R01-GM084448 and P50-GM081879).
Scott LaFee | EurekAlert!
New technique unveils 'matrix' inside tissues and tumors
29.06.2017 | University of Copenhagen The Faculty of Health and Medical Sciences
Designed proteins to treat muscular dystrophy
29.06.2017 | Universität Basel
Computer scientists use wave packet theory to develop realistic, detailed water wave simulations in real time. Their results will be presented at this year’s SIGGRAPH conference.
Think about the last time you were at a lake, river, or the ocean. Remember the ripples of the water, the waves crashing against the rocks, the wake following...
An international team of scientists has proposed a new multi-disciplinary approach in which an array of new technologies will allow us to map biodiversity and the risks that wildlife is facing at the scale of whole landscapes. The findings are published in Nature Ecology and Evolution. This international research is led by the Kunming Institute of Zoology from China, University of East Anglia, University of Leicester and the Leibniz Institute for Zoo and Wildlife Research.
Using a combination of satellite and ground data, the team proposes that it is now possible to map biodiversity with an accuracy that has not been previously...
Heatwaves in the Arctic, longer periods of vegetation in Europe, severe floods in West Africa – starting in 2021, scientists want to explore the emissions of the greenhouse gas methane with the German-French satellite MERLIN. This is made possible by a new robust laser system of the Fraunhofer Institute for Laser Technology ILT in Aachen, which achieves unprecedented measurement accuracy.
Methane is primarily the result of the decomposition of organic matter. The gas has a 25 times greater warming potential than carbon dioxide, but is not as...
Hydrogen is regarded as the energy source of the future: It is produced with solar power and can be used to generate heat and electricity in fuel cells. Empa researchers have now succeeded in decoding the movement of hydrogen ions in crystals – a key step towards more efficient energy conversion in the hydrogen industry of tomorrow.
As charge carriers, electrons and ions play the leading role in electrochemical energy storage devices and converters such as batteries and fuel cells. Proton...
Scientists from the Excellence Cluster Universe at the Ludwig-Maximilians-Universität Munich have establised "Cosmowebportal", a unique data centre for cosmological simulations located at the Leibniz Supercomputing Centre (LRZ) of the Bavarian Academy of Sciences. The complete results of a series of large hydrodynamical cosmological simulations are available, with data volumes typically exceeding several hundred terabytes. Scientists worldwide can interactively explore these complex simulations via a web interface and directly access the results.
With current telescopes, scientists can observe our Universe’s galaxies and galaxy clusters and their distribution along an invisible cosmic web. From the...
19.06.2017 | Event News
13.06.2017 | Event News
13.06.2017 | Event News
29.06.2017 | Physics and Astronomy
29.06.2017 | Life Sciences
29.06.2017 | Health and Medicine