Setup
This page covers everything you need before running either pipeline: installing ViralUnity, generating sample sheets for the example data, and downloading the reference databases. Most of it is one-time work; once you finish, you can move straight to either pipeline walkthrough.
1b. Build per-rule environments
Each pipeline rule runs in its own conda environment. ViralUnity caches these envs across runs so they only need to be built once; pre-warm the cache now:
viralunity setup --pipelines all
This materializes every per-rule env into ~/.cache/viralunity/conda-envs/ (override with --conda-prefix PATH). The subsequent pipeline runs in this tutorial — and every future run on this machine — reuse the cache and skip env creation. Restricting to a single pipeline (e.g. --pipelines consensus-illumina) is faster if you only plan to run one flavor; add the others later.
If setup fails with CreateCondaEnvironmentException, see Troubleshooting.
1c. Download the example data
The example dataset used throughout this tutorial is hosted as a single tarball. Download it once at the repo root:
curl -L -o my_test_data.tar.gz https://itps-nimbus.nyc3.cdn.digitaloceanspaces.com/my_test_data.tar.gz
tar -xzf my_test_data.tar.gz
rm my_test_data.tar.gz
You should now have:
my_test_data/
├── illumina_data/
│ ├── itps-0001_R1.fastq.gz
│ ├── itps-0001_R2.fastq.gz
│ ├── itps-0002_R1.fastq.gz
│ └── itps-0002_R2.fastq.gz
└── nanopore_data/
├── barcode05.itps-0003.fastq.gz
└── barcode09.itps-0004.fastq.gz
my_test_data/ is gitignored, so the download stays out of version control. The consensus tutorial additionally needs the ARTIC nCoV-2019 V3 reference FASTA and primer BED — grab them from the ARTIC primer-schemes repo (raw URLs: nCoV-2019.reference.fasta, nCoV-2019.scheme.bed) and place them somewhere you can point --reference / --primer-scheme at (the rest of this tutorial uses databases/refs/).
2. Generate sample sheets
A sample sheet is a no-header CSV that tells ViralUnity which FASTQ files belong to which sample:
Illumina — 3 columns:
sample_id,R1_path,R2_pathNanopore — 2 columns:
sample_id,fastq_path
You can write the CSV by hand, but viralunity create-samplesheet builds it for you by scanning a run directory.
The bundled test data has all FASTQs directly inside the run directory (rather than one subdirectory per sample), so we pass --level 0 and a small naming convention:
# Illumina: FASTQs in my_test_data/illumina_data/, sample ID = portion before first '_'
viralunity create-samplesheet \
--input my_test_data/illumina_data/ \
--output samples_illumina.csv \
--level 0 \
--separator _ \
--pattern R1
This writes:
itps-0001,my_test_data/illumina_data/itps-0001_R1.fastq.gz,my_test_data/illumina_data/itps-0001_R2.fastq.gz
itps-0002,my_test_data/illumina_data/itps-0002_R1.fastq.gz,my_test_data/illumina_data/itps-0002_R2.fastq.gz
# Nanopore: FASTQs in my_test_data/nanopore_data/, sample ID = portion before first '.'
viralunity create-samplesheet \
--input my_test_data/nanopore_data/ \
--output samples_nanopore.csv \
--level 0 \
--separator . \
--pattern barcode
Which writes:
barcode05,my_test_data/nanopore_data/barcode05.itps-0003.fastq.gz
barcode09,my_test_data/nanopore_data/barcode09.itps-0004.fastq.gz
Tip
The reference sample sheets that ship with this checkout are at my_results/samples_illumina.csv and my_results/samples_nanopore.csv — use them to compare against what you just generated.
For your own data, the same command works with --level 1 (one subdirectory per sample, the default) if your run is organised that way; see the Commands reference for the full table of options.
3. Decide which databases you need
Different pipelines need different reference data:
Pipeline |
Required |
Optional (enable extra features) |
|---|---|---|
|
reference FASTA |
primer-scheme BED (amplicon data), adapter FASTA (custom adapters) |
|
Kraken2 index, Krona taxonomy |
NCBI taxdump (filters), host FASTA or Deacon index (dehosting) |
|
Kraken2, Krona, taxdump, DIAMOND DB, DIAMOND taxids |
(same as above) |
|
the above, viral genomes FASTA + |
BLAST index (built automatically) — for the |
The consensus pipeline does not need any database downloads — only a reference FASTA you already have. The rest of this section sets you up for the metagenomic walkthrough.
4. Download the metagenomic databases
The shortcut command grabs the four most-used databases in one go (Kraken2 + Krona + taxdump + DIAMOND). The conda environment must be active because the Krona step needs ktUpdateTaxonomy.sh:
conda activate viralunity
viralunity get-databases all --path databases/ --threads 4
Plan for the download. Approximate sizes on disk after extraction:
Database |
Path created |
Approx. size |
Used by |
|---|---|---|---|
Kraken2 viral |
|
~1.1 GB |
|
Krona taxonomy |
|
~150 MB |
|
NCBI taxdump |
|
~600 MB |
|
DIAMOND viral |
|
~400 MB |
|
If you only plan to use Kraken2 on reads, you can skip get-databases diamond and only run the first three. Each subcommand is also available standalone — see [viralunity get-databases --help](../commands.md#viralunity-get-databases).
Three more databases for the full meta walkthrough
The all subcommand intentionally leaves out three databases that are large and only needed for specific features:
# Viral genomes + per-genome taxid map + BLAST index — required for --run-reference-assembly
viralunity get-databases virus-genome --path databases/
# Deacon human host index — recommended for dehosting (faster than minimap2 vs a host FASTA)
viralunity get-databases deacon-index --path databases/ --index-name panhuman-1
# NCBI nr protein database — required for --run-nr-validation (false-positive reduction).
# WARNING: 100+ GB, hours to download. aws/gcp mirrors are usually faster than ncbi.
viralunity get-databases nr --path databases/ --source gcp --threads 8
virus-genome ships with viral.genomes.fasta, genome2taxid.tsv, and the BLAST .nhr/.nin/.nsq index files (used for the similarity reference-selection strategy). The Deacon index lands at databases/deacon_indexes/panhuman-1.idx. nr downloads NCBI’s preformatted BLAST+ nr into databases/diamond/nr.* and prepares it for DIAMOND; pass --nr-diamond-database databases/diamond/nr. If you already have nr (or work air-gapped), use --from-blastdb <prefix> or --from-fasta <nr.faa> instead of downloading — see [commands.md](../commands.md#nr--ncbi-nr-database-for-nr-validation).
Tip
If you want to dehost against a non-human genome, use viralunity get-databases host-genome --accession <NCBI accession> to download a host FASTA, then either pass it directly via --host-reference (minimap2 dehosting) or build a Deacon index with viralunity build-deacon-index --input <fasta> for a faster path.
5. Verify the setup
After downloading, you should have:
databases/
├── kraken2/ # Kraken2 index files
├── krona/taxonomy/ # taxonomy.tab and friends
├── taxdump/ # nodes.dmp, names.dmp, …
├── diamond/ # viral.dmnd, protein2taxid.tsv
├── virus_genomes/ # viral.genomes.fasta(.nhr/.nin/.nsq), genome2taxid.tsv
└── deacon_indexes/ # panhuman-1.idx
A quick sanity check:
ls databases/
ls databases/diamond/ databases/virus_genomes/
You are now ready to: