Computational approaches to predict bacteriophage-host relationships

Robert A Edwards, Katelyn McNair, Karoline Faust, Jeroen Raes, Bas E Dutilh

Research output: Contribution to journalArticleAcademicpeer-review

Abstract

Metagenomics has changed the face of virus discovery by enabling the accurate identification of viral genome sequences without requiring isolation of the viruses. As a result, metagenomic virus discovery leaves the first and most fundamental question about any novel virus unanswered: What host does the virus infect? The diversity of the global virosphere and the volumes of data obtained in metagenomic sequencing projects demand computational tools for virus-host prediction. We focus on bacteriophages (phages, viruses that infect bacteria), the most abundant and diverse group of viruses found in environmental metagenomes. By analyzing 820 phages with annotated hosts, we review and assess the predictive power of in silico phage-host signals. Sequence homology approaches are the most effective at identifying known phage-host pairs. Compositional and abundance-based methods contain significant signal for phage-host classification, providing opportunities for analyzing the unknowns in viral metagenomes. Together, these computational approaches further our knowledge of the interactions between phages and their hosts. Importantly, we find that all reviewed signals significantly link phages to their hosts, illustrating how current knowledge and insights about the interaction mechanisms and ecology of coevolving phages and bacteria can be exploited to predict phage-host relationships, with potential relevance for medical and industrial applications.

Original languageEnglish
Pages (from-to)258-272
JournalFEMS Microbiology Reviews
Volume40
Issue number2
DOIs
Publication statusPublished - 2016

Keywords

  • phages
  • viruses of microbes
  • metagenomics
  • co-occurrence
  • CRISPR
  • oligonucleotide usage

Fingerprint

Dive into the research topics of 'Computational approaches to predict bacteriophage-host relationships'. Together they form a unique fingerprint.

Cite this