The work was reported in two recent papers in Genome Research, published online on July 3 and Sept. 27.
“Our goal is to understand how regulatory information is encrypted and to learn which sequence variations contribute to medical risks,” says Andrew McCallion, Ph.D., associate professor of molecular and comparative pathobiology in the McKusick-Nathans Institute of Genetic Medicine at Hopkins.
“We give data to a computer and ‘teach it’ to distinguish between data that has no biological value versus data that has this or that biological value. It then establishes a set of rules, which allows it to look at new sets of data and apply what it learned. We’re basically sending our computers to school.”
These state-of-the-art “machine learning” techniques were developed by Michael Beer, Ph.D., assistant professor of biomedical engineering at the Johns Hopkins School of Medicine, and by Ivan Ovcharenko, Ph.D., at the National Center for Biotechnology Information. The researchers began both studies by creating “training sets” for their computers to “learn” from. These training sets were lists of DNA sequences taken from regions of the genome, called enhancers, that are known to increase the activity of particular genes in particular cells.
For the first of their studies, McCallion’s team created a training set of enhancer sequences specific to a particular region of the brain by compiling a list of 211 published sequences that had been shown, by various studies in mice and zebrafish, to be active in the development or function of that part of the brain.
For a second study, the team generated a training set through experiments of their own. They began with a purified population of mouse melanocytes, which are the skin cells that produce the pigment melanin that gives color to skin and absorbs harmful UV rays from the sun. The researchers used a technique called ChIP-seq (pronounced “chip seek”) to collect and sequence all of the pieces of DNA that were bound in those cells by special enhancer-binding proteins, generating a list of about 2,500 presumed melanocyte enhancer sequences.
Once the researchers had these two training sets for their computers, one specific to the brain and another to melanocytes, the computers were able to distinguish the features of the training sequences from the features of all other sequences in the genome, and create rules that defined one set from the other. Applying those rules to the whole genome, the computers were able to discover thousands of probable brain or melanocyte enhancer sequences that fit the features of the training sets.
In the brain study, the computers identified 40,000 probable brain enhancer sequences; for melanocytes, 7,500. Randomly testing a subset of each batch of sequences, the scientists found that more than 85 percent of the predicted enhancer sequences enhanced gene activity in the brain or in melanocytes, as expected, verifying the predictive power of their approach.
The researchers say that, in addition to identifying specific DNA sequences that control the genetic activity of a particular organ or cell type, these studies contribute to our understanding of enhancers in general and have validated an experimental approach that can be applied to many other biological questions as well.
Authors on the brain paper include Grzegorz Burzynski, Xylena Reed, Zachary Stine, Takeshi Matsui and Andrew McCallion from The Johns Hopkins University, and Leila Taher and Ivan Ovcharenko from the National Center for Biotechnology Information.
Authors on the melanocyte paper include David Gorkin, Dongwon Lee, Xylena Reed, Christopher Fletez-Brant, Seneca Bessling, Michael Beer and Andrew McCallion from The Johns Hopkins University, and Stacie Loftus and William Pavan from the National Human Genome Research Institute.
This work was supported by grants from the National Institute of Neurological Disorders and Stroke (NS062972), the National Human Genome Research Institute’s Intramural Research Program, the National Library of Medicine, the National Institute of General Medical Sciences (GM07814, GM071648), the National Science Foundation and the Searle Scholars Program.
Catherine Kolf | Source: Newswise Science News
Further information: www.jhmi.edu
More articles from Life Sciences:
ASU researchers discover chameleons use colorful language to communicate
12.12.2013 | Arizona State University
Sleep-deprived mice show connections among lack of shut-eye, diabetes, age
12.12.2013 | University of Pennsylvania School of Medicine
A unique solar panel design made with a new ceramic material points the way to potentially providing sustainable power cheaper, more efficiently, and requiring less manufacturing time.
It also reaches a four-decade-old goal of discovering a bulk photovoltaic material that can harness energy from visible and infrared light, not just ultraviolet light.
Scaling up this new design from its tablet-size prototype to a full-size solar panel would be a large step toward making solar power affordable compared with ...
Atlantische Flohkrebse pflanzen sich jetzt auch in arktischen Gewässern fort
Biologen des Alfred-Wegener-Institutes, Helmholtz-Zentrum für Polar- und Meeresforschung (AWI), haben zum ersten Mal nachgewiesen, dass sich in den arktischen Gewässern westlich Spitzbergens auch Flohkrebse aus dem wärmeren Atlantik fortpflanzen.
Diese überraschende Entdeckung deute auf einen möglichen Wandel der arktischen Zooplankton-Gemeinschaft hin, berichten die Wissenschaftler und Wissenschaftlerinnen in der Fachzeitschrift Marine Ecology ...
The molecular architecture of three key proteins and their complexes reveals how plants fine-tune their immune response to pathogens
Plants rarely get sick in their natural environment. When the threat of infection arises, a quick decision is made about the necessary countermeasures. The course is set by a protein which forms complexes with its partner proteins for this purpose.
Jane Parker from the Max Planck Institute for Plant Breeding ...
Researchers studying speciation of butterfly orchids on the Azores have been startled to discover that the answer to a long-debated question "Do the islands support one species or two species?" is actually "three species".
Hochstetter's Butterfly-orchid, newly recognized following application of a battery of scientific techniques and reveling in a complex taxonomic history worthy of Sherlock Holmes, is arguably Europe's rarest orchid species. Under threat in its mountain-top retreat, the orchid urgently requires conservation recognition.
A lavishly illustrated publication, titled "Systematic revision of Platanthera in ...
Researchers from Brown University and the University of Hawaii have found some mineralogical surprises in the Moon's largest impact crater.
Data from the Moon Mineralogy Mapper that flew aboard India's Chandrayaan-1 lunar orbiter shows a diverse mineralogy in the subsurface of the giant South Pole Aitken basin.
The differing mineral signatures could be reflective of the minerals dredged up at the time of the giant impact 4 billion years ago, ...
12.12.2013 | Life Sciences
12.12.2013 | Earth Sciences
12.12.2013 | Studies and Analyses
11.12.2013 | Event News
10.12.2013 | Event News
05.12.2013 | Event News