This book is developed for Codanics learners and is launched through https://www.codanics.com.
Viromics is the sequencing-based study of viral communities. It sits at the meeting point of biology, NGS, Linux, statistics, ecological interpretation, and careful reporting. Unlike bacterial amplicon profiling, viromics has no universal marker like 16S rRNA. That single fact reshapes the whole workflow: identification, quality control, clustering, and taxonomy all have to be rebuilt around evidence rather than a barcode.
The learner must think like a detective. A viral contig is not accepted because one tool flags it. Confidence grows when several independent lines of evidence agree — viral prediction, hallmark genes, CheckV quality, coverage, taxonomy, functional annotation, and biological context. Where the evidence is thin, the honest answer is an uncertain one, and this book teaches you to say so.
How this edition is organized
This edition moves in one continuous arc: it builds concepts first, then a clean computing environment, then a complete analysis pipeline, and finally visualization, a public-data mini project, and research ideas. Every chapter opens with clear learning objectives and closes with quizzes, and reusable scripts plus small example datasets let you practice each step without waiting on large downloads. Four reference appendices — a Linux cheatsheet, a tool summary, a troubleshooting guide, and a glossary — are meant to be consulted out of order whenever you need them.
What you need before starting
You should be comfortable with:
- Linux terminal basics
- FASTQ files
- paired-end sequencing
- conda or mamba
- FASTA files
- basic metagenomics
- simple R or Python plotting
Hardware assumptions
All runtime and storage estimates in this book use this reference machine:
| CPU |
12 threads |
| RAM |
32 GB |
| Storage |
SSD |
| OS |
Ubuntu 22.04 or 24.04 |
| Internet |
stable broadband |
Tool databases can be much larger than the example dataset. Keep databases on a large disk. For full DRAM, iPHoP, geNomad, and CheckV installations, plan hundreds of GB if you want all databases locally.
Citation style
The book uses Quarto citations from references.bib. Core references include CheckV (Nayfach et al. 2021), VirSorter2 (Guo et al. 2021), geNomad (Camargo et al. 2024), vConTACT2 (Bin Jang et al. 2019), VIBRANT (Kieft, Zhou, and Anantharaman 2020), DRAM-v (Shaffer et al. 2020), eggNOG-mapper (Cantalapiedra et al. 2021), MEGAHIT (Li et al. 2015), fastp (Chen et al. 2018), MAFFT (Katoh and Standley 2013), IQ-TREE 2 (Minh et al. 2020), and ICTV taxonomy resources (International Committee on Taxonomy of Viruses 2026).
Nayfach, Stephen, Antonio Pedro Camargo, Frederik Schulz, Emiley Eloe-Fadrosh, Simon Roux, and Nikos C. Kyrpides. 2021.
“CheckV Assesses the Quality and Completeness of Metagenome-Assembled Viral Genomes.” Nature Biotechnology 39: 578–85.
https://doi.org/10.1038/s41587-020-00774-7.
Guo, Jiarong, Benjamin Bolduc, Ahmed A. Zayed, Arvind Varsani, Gabriela Dominguez-Huerta, Tom O. Delmont, Akbar A. Pratama, et al. 2021.
“VirSorter2: A Multi-Classifier, Expert-Guided Approach to Detect Diverse DNA and RNA Viruses.” Microbiome 9: 37.
https://doi.org/10.1186/s40168-020-00990-y.
Camargo, Antonio Pedro, Simon Roux, Frederik Schulz, Michal Babinski, Yan Xu, Bin Hu, Patrick S. G. Chain, Stephen Nayfach, and Nikos C. Kyrpides. 2024.
“Identification of Mobile Genetic Elements with geNomad.” Nature Biotechnology 42: 1303–12.
https://doi.org/10.1038/s41587-023-01953-y.
Bin Jang, Ho, Benjamin Bolduc, Olivier Zablocki, Jens H. Kuhn, Simon Roux, Evelien M. Adriaenssens, J. Rodney Brister, et al. 2019.
“Taxonomic Assignment of Uncultivated Prokaryotic Virus Genomes Is Enabled by Gene-Sharing Networks.” Nature Biotechnology 37: 632–39.
https://doi.org/10.1038/s41587-019-0100-8.
Kieft, Kristopher, Zhichao Zhou, and Karthik Anantharaman. 2020.
“VIBRANT: Automated Recovery, Annotation and Curation of Microbial Viruses, and Evaluation of Viral Community Function from Genomic Sequences.” Microbiome 8: 90.
https://doi.org/10.1186/s40168-020-00867-0.
Shaffer, Michael, Mikayla A. Borton, Brendan B. McGivern, Ahmed A. Zayed, Sabina L. La Rosa, Lindsey M. Solden, Pengfei Liu, et al. 2020.
“DRAM for Distilling Microbial Metabolism to Automate the Curation of Microbiome Function.” Nucleic Acids Research 48 (16): 8883–8900.
https://doi.org/10.1093/nar/gkaa621.
Cantalapiedra, Carlos P., Ana Hernandez-Plaza, Ivica Letunic, Peer Bork, and Jaime Huerta-Cepas. 2021.
“eggNOG-Mapper V2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale.” Molecular Biology and Evolution 38 (12): 5825–29.
https://doi.org/10.1093/molbev/msab293.
Li, Dinghua, Chi-Man Liu, Ruibang Luo, Kunihiko Sadakane, and Tak-Wah Lam. 2015.
“MEGAHIT: An Ultra-Fast Single-Node Solution for Large and Complex Metagenomics Assembly via Succinct de Bruijn Graph.” Bioinformatics 31 (10): 1674–76.
https://doi.org/10.1093/bioinformatics/btv033.
Chen, Shifu, Yanqing Zhou, Yaru Chen, and Jia Gu. 2018.
“Fastp: An Ultra-Fast All-in-One FASTQ Preprocessor.” Bioinformatics 34 (17): i884–90.
https://doi.org/10.1093/bioinformatics/bty560.
Katoh, Kazutaka, and Daron M. Standley. 2013.
“MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability.” Molecular Biology and Evolution 30 (4): 772–80.
https://doi.org/10.1093/molbev/mst010.
Minh, Bui Quang, Heiko A. Schmidt, Olga Chernomor, Dominik Schrempf, Michael D. Woodhams, Arndt von Haeseler, and Robert Lanfear. 2020.
“IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era.” Molecular Biology and Evolution 37 (5): 1530–34.
https://doi.org/10.1093/molbev/msaa015.
International Committee on Taxonomy of Viruses. 2026.
“ICTV Taxonomy.” https://ictv.global/taxonomy.