Org: CDC/NCIRD/DBB/RDB/PSLB
Contact Email: jhamlin@cdc.gov
Exemption: None
Status: Maintenance
mashwrapper is a wrapper around the program Mash and the NCBI Datasets command line tools (CLI). It identifies the most likely species from paired gzipped FASTQ reads using a Mash database.
You can provide the database for comparison in two ways:
--get_database: Used when downloading and building a new Mash database from genomes--use_database: Used when you're skipping the build step and instead providing a prebuilt Mash database
The tool outputs a text file containing the top five matches from the Mash database for the input reads. This output includes standard Mash results, and the best species match is determined by a cutoff based on the Mash distance score. For Legionella, this cutoff is conservatively set to a Mash distance of < 0.05. If you're using the tool for a different species, you should adjust this cutoff value based on what is most appropriate for your organism.
The pipeline is built using Nextflow, a workflow tool to run tasks across multiple compute infrastructures in a very portable manner. It uses Docker/Singularity containers, making installation trivial and results highly reproducible.
- Confirm input sample sheet (
--get_databaseor--use_database) - Confirm input organism sheet [optional] (
--get_database) - Download genomes from NCBI using NCBI datasets CLI [optional] (
--get_database) - Format downloaded genomes to be Genus_Species_GenebankIdentifier.fna using NCBI dataformat CLI [optional] (
--get_database) - Build individual Mash sketches for all genomes [optional] (
--get_database) - Build Mash database from all Mash sketches [optional] (
--get_database) - Test FASTQ reads against a Mash database either built or provided (
--get_databaseor--use_database) - Collate results from each isolate of interest tested against the Mash database (
--get_databaseor--use_database)
-
Install
Nextflow(>=21.10.3) -
Install either
DockerorSingularityto ensure full pipeline reproducibility with Nextflow.Condamay be used as a last resort; see docs) -
Clone or download the pipeline and test it on a minimal dataset:
This repository includes a test dataset with the following files:
- inputDB.txt - A plain text file of species to download when using the
-profile testGetoption. File does not include a header.- inputReads.csv - A CSV file listing paired-end read files. It has the following header: sample,fastq_1,fastq_2
- myMashDatabase.msh - A prebuilt Mash database from isolates listed in inputDB.txt file and used with the
-profile testUseoption.- subERR125190_(1,2).fastq.gz - Subsampled reads (45,000 reads) from Legionella fallonii
- subERR351242_(1,2).fastq.gz - Subsampled reads (45,000 reads) from Legionella pneumophila
- subSRR10019387_(1,2).fastq.gz - Subsampled reads (45,000 reads) from Legionella longbeachae
Step-by-step example commands
## Step 1: Clone the repository
git clone https://github.com/CDCgov/mashwrapper.git
## Step 2: Test downloading and building the databse
## "YOURPROFILE" is your preferred execution environment (Docker, Singularity or Conda)
nextflow run mashwrapper -profile testGet,YOURPROFILE
## Step 3: Test using a prebuilt database
## "YOURPROFILE" is your preferred execution environment (Docker, Singularity or Conda)
nextflow run mashwrapper -profile testUse,YOURPROFILE You will likely need to adjust the nfcore_custom.config file to work on your compute environment. To use it, specify the path to its directory using the --custom_config_base flag. This should point to the "conf" directory (i.e., ~/mashwrapper/conf).
- Start running your analysis!
## Build a Mash database for organism(s) of interest
nextflow run nf-core/mashwrapper -profile <Docker/Singularity/Conda> --input samplesheet.csv --get_database organismsheet.txt --custom_config_base ~/mashwrapper/conf
## Use a prebuilt Mash database
nextflow run nf-core/mashwrapper -profile <Docker/Singularity/Conda> --input samplesheet.csv --use_database myMashDatabase.msh --custom_config_base ~/mashwrapper/confThe nf-core/mashwrapper pipeline comes with documentation about the pipeline usage and parameters and output.
mashwrapper is based heavily on previous work by Jason Caravas with the current version written by Jenna Hamlin.
We thank the following people for their extensive assistance in the development of this pipeline:
If you would like to contribute to this pipeline, please file an Issue
Please see the notices page for detailed information