Detecting millions of consistently unidentified spectra in vast tracts of proteomics data is possible with a new algorithm developed at EMBL-EBI
A new algorithm clusters the millions of peptide mass spectra in the PRIDE Archive public database, making it easier to detect millions of consistently unidentified spectra across different datasets. Published in Nature Methods, the new tool is an important step towards fully exploiting data produced in discovery proteomics experiments.
On average, almost three quarters of spectra measured in discovery proteomics experiments remain unidentified, regardless of the quality of the experiment, as they cannot be interpreted by standard sequence-based search engines.
Alternative approaches to improve the rate of identification exist, but are fraught with disadvantages including ambiguous results. In today's study, researchers working on the PRIDE Archive public repository of proteomics data present a large-scale 'spectrum clustering' solution that takes advantage of the growing number of mass spectrometry (MS) datasets to systematically study millions of unidentified spectra.
"MS experiments produce huge amounts of data, but identifying meaningful sequences that could be assigned to specific biological functions can be troublesome," says Johannes Griss, formerly at EMBL-EBI in the UK and now at the Medical University of Vienna, Austria.
"Discovery proteomics is a mature technology, and it's crucial that we are able to exploit the data efficiently."
One of the challenges with these technologies is that a large proportion of the data generated can't be interpreted, as they correspond to peptides that have not yet been observed and are not available in databases. Such spectra could correspond to peptide variants derived from individual generic variation, or to peptides containing post-translational modifications, which are essential for the biological functions of proteins.
"What we have now is an algorithm that shows us patterns, or groups of spectra, that we've consistently missed, and helps us figure out which ones are good enough to pursue," adds Johannes. "It's a valuable tool that helps us unpick what's going on in proteomics, so we can better understand basic biological processes."
The team used the approach to recognise 9 million consistently unidentified spectra, which can make post-translational modifications and peptides containing sequence variants more discoverable. They identified three distinct sets of spectra: those that have been incorrectly identified, those that are not of high enough quality to identify properly, and those that are truly unidentified. They also combined their new approach with other methods to identify roughly 20% of the originally unidentified spectra in the public archive.
"Discovery proteomics is a mature technology, and it's crucial that we are able to exploit the data efficiently - but creating a sensible subset of spectra to start an in-depth analysis of unidentified spectra has been very challenging," says Juan Antonio Vizcaíno, who leads the Proteomics team at EMBL-EBI. "We developed a comparatively lightweight computational approach that makes it much easier to detect sequences that have been incorrectly identified, or consistently observed but not identified. These ready-to-use collections of commonly unidentified spectra are a resource for the community, so that we can all pool our efforts to find lasting solutions for proteomics research."
The new algorithm will be used to improve quality control in the PRIDE Archive. The complete spectrum clustering results are available through the PRIDE Cluster resource, which aims to simplify further investigation into unidentified spectra.
Source article: Griss J., et al. (2016). Recognizing millions of consistently unidentified spectra across hundreds of shotgun proteomics datasets. Nature Methods (in press). DOI: 10.1038/nmeth.3902
Mary Todd Bergman | EurekAlert!
How brains surrender to sleep
23.06.2017 | IMP - Forschungsinstitut für Molekulare Pathologie GmbH
A new technique isolates neuronal activity during memory consolidation
22.06.2017 | Spanish National Research Council (CSIC)
An international team of scientists has proposed a new multi-disciplinary approach in which an array of new technologies will allow us to map biodiversity and the risks that wildlife is facing at the scale of whole landscapes. The findings are published in Nature Ecology and Evolution. This international research is led by the Kunming Institute of Zoology from China, University of East Anglia, University of Leicester and the Leibniz Institute for Zoo and Wildlife Research.
Using a combination of satellite and ground data, the team proposes that it is now possible to map biodiversity with an accuracy that has not been previously...
Heatwaves in the Arctic, longer periods of vegetation in Europe, severe floods in West Africa – starting in 2021, scientists want to explore the emissions of the greenhouse gas methane with the German-French satellite MERLIN. This is made possible by a new robust laser system of the Fraunhofer Institute for Laser Technology ILT in Aachen, which achieves unprecedented measurement accuracy.
Methane is primarily the result of the decomposition of organic matter. The gas has a 25 times greater warming potential than carbon dioxide, but is not as...
Hydrogen is regarded as the energy source of the future: It is produced with solar power and can be used to generate heat and electricity in fuel cells. Empa researchers have now succeeded in decoding the movement of hydrogen ions in crystals – a key step towards more efficient energy conversion in the hydrogen industry of tomorrow.
As charge carriers, electrons and ions play the leading role in electrochemical energy storage devices and converters such as batteries and fuel cells. Proton...
Scientists from the Excellence Cluster Universe at the Ludwig-Maximilians-Universität Munich have establised "Cosmowebportal", a unique data centre for cosmological simulations located at the Leibniz Supercomputing Centre (LRZ) of the Bavarian Academy of Sciences. The complete results of a series of large hydrodynamical cosmological simulations are available, with data volumes typically exceeding several hundred terabytes. Scientists worldwide can interactively explore these complex simulations via a web interface and directly access the results.
With current telescopes, scientists can observe our Universe’s galaxies and galaxy clusters and their distribution along an invisible cosmic web. From the...
Temperature measurements possible even on the smallest scale / Molecular ruby for use in material sciences, biology, and medicine
Chemists at Johannes Gutenberg University Mainz (JGU) in cooperation with researchers of the German Federal Institute for Materials Research and Testing (BAM)...
19.06.2017 | Event News
13.06.2017 | Event News
13.06.2017 | Event News
23.06.2017 | Physics and Astronomy
23.06.2017 | Physics and Astronomy
23.06.2017 | Information Technology