Skip to content

Latest commit

 

History

History
98 lines (66 loc) · 6.54 KB

File metadata and controls

98 lines (66 loc) · 6.54 KB

mashwrapper

Nextflow run with conda run with docker run with singularity

Org: CDC/NCIRD/DBB/RDB/PSLB
Contact Email: jhamlin@cdc.gov
Exemption: None
Status: Maintenance

Introduction

mashwrapper is a wrapper around the program Mash and the NCBI Datasets command line tools (CLI). It identifies the most likely species from paired gzipped FASTQ reads using a Mash database.

You can provide the database for comparison in two ways:

  1. --get_database: Used when downloading and building a new Mash database from genomes
  2. --use_database: Used when you're skipping the build step and instead providing a prebuilt Mash database

The tool outputs a text file containing the top five matches from the Mash database for the input reads. This output includes standard Mash results, and the best species match is determined by a cutoff based on the Mash distance score. For Legionella, this cutoff is conservatively set to a Mash distance of < 0.05. If you're using the tool for a different species, you should adjust this cutoff value based on what is most appropriate for your organism.

The pipeline is built using Nextflow, a workflow tool to run tasks across multiple compute infrastructures in a very portable manner. It uses Docker/Singularity containers, making installation trivial and results highly reproducible.

Pipeline summary

  1. Confirm input sample sheet (--get_database or --use_database)
  2. Confirm input organism sheet [optional] (--get_database)
  3. Download genomes from NCBI using NCBI datasets CLI [optional] (--get_database)
  4. Format downloaded genomes to be Genus_Species_GenebankIdentifier.fna using NCBI dataformat CLI [optional] (--get_database)
  5. Build individual Mash sketches for all genomes [optional] (--get_database)
  6. Build Mash database from all Mash sketches [optional] (--get_database)
  7. Test FASTQ reads against a Mash database either built or provided (--get_database or --use_database)
  8. Collate results from each isolate of interest tested against the Mash database (--get_database or --use_database)

Quick Start

  1. Install Nextflow (>=21.10.3)

  2. Install either Docker or Singularity to ensure full pipeline reproducibility with Nextflow. Conda may be used as a last resort; see docs)

  3. Clone or download the pipeline and test it on a minimal dataset:

This repository includes a test dataset with the following files:

  • inputDB.txt - A plain text file of species to download when using the -profile testGet option. File does not include a header.
  • inputReads.csv - A CSV file listing paired-end read files. It has the following header: sample,fastq_1,fastq_2
  • myMashDatabase.msh - A prebuilt Mash database from isolates listed in inputDB.txt file and used with the -profile testUse option.
  • subERR125190_(1,2).fastq.gz - Subsampled reads (45,000 reads) from Legionella fallonii
  • subERR351242_(1,2).fastq.gz - Subsampled reads (45,000 reads) from Legionella pneumophila
  • subSRR10019387_(1,2).fastq.gz - Subsampled reads (45,000 reads) from Legionella longbeachae

Step-by-step example commands

 ## Step 1: Clone the repository
 git clone https://github.com/CDCgov/mashwrapper.git

 ## Step 2: Test downloading and building the databse
 ## "YOURPROFILE" is your preferred execution environment (Docker, Singularity or Conda)
 nextflow run mashwrapper -profile testGet,YOURPROFILE
 
 ## Step 3: Test using a prebuilt database
 ## "YOURPROFILE" is your preferred execution environment (Docker, Singularity or Conda)
 nextflow run mashwrapper -profile testUse,YOURPROFILE 

You will likely need to adjust the nfcore_custom.config file to work on your compute environment. To use it, specify the path to its directory using the --custom_config_base flag. This should point to the "conf" directory (i.e., ~/mashwrapper/conf).

  1. Start running your analysis!
 ## Build a Mash database for organism(s) of interest
 nextflow run nf-core/mashwrapper -profile <Docker/Singularity/Conda> --input samplesheet.csv --get_database organismsheet.txt --custom_config_base ~/mashwrapper/conf

## Use a prebuilt Mash database
 nextflow run nf-core/mashwrapper -profile <Docker/Singularity/Conda> --input samplesheet.csv --use_database myMashDatabase.msh --custom_config_base ~/mashwrapper/conf

Documentation

The nf-core/mashwrapper pipeline comes with documentation about the pipeline usage and parameters and output.

Credits

mashwrapper is based heavily on previous work by Jason Caravas with the current version written by Jenna Hamlin.

We thank the following people for their extensive assistance in the development of this pipeline:

Contributions and Support

If you would like to contribute to this pipeline, please file an Issue

Repository Usage and Legal Notices

Please see the notices page for detailed information