Skip to content

TaxTriage Documentation

DOI Nextflow

Docker Singularity

See it first

Interactive report demo Example ODR (PDF) Pathogen sheet

The demo is a live report built from the example dataset, its the same artifact the pipeline writes at the end of a run, complete with live filters and sample QC flags you can edit in the browser. The ODR is the static PDF deliverable. The pathogen sheet is every organism TaxTriage can flag, searchable and filterable.

About TaxTriage

TaxTriage is a flexible, containerized bioinformatics pipeline designed to identify pathogens within complex samples (e.g., respiratory swabs, lesion swabs, whole blood) using untargeted DNA or RNA sequencing data. It supports short-read (Illumina) and long-read (ONT, PacBio) platforms, incorporating quality control, organism classification, read mapping, and a unified confidence metric for all identified organisms.

The final analysis output is an Organism Discovery Report (ODR) - 2 files (a PDF and an interactive HTML) with summaries of all intermediate data supporting pathogen identification. TaxTriage is designed for broad deployment and early-stage outbreak investigations and is not intended as a standalone diagnostic capability.


Quick Navigation

Section Description
Installation Install Nextflow, Docker, and Singularity
Quick Start Run your first test pipeline in minutes
Samplesheet Format and prepare your input samplesheet
Running the Pipeline Commands, profiles, and execution modes
CLI Parameters Full reference for all pipeline parameters
Pipeline Modules Step-by-step breakdown of the workflow
Output Understanding output files and directories
TASS Scoring How the confidence score is calculated
Microbial Categories Primary, Opportunistic, Potential, Commensal - what they mean and how to read them
In-Silico Simulation Simulated read validation
Cloud & Seqera Running on AWS with Nextflow Tower / Seqera
Geneious Plugin Geneious Prime integration
Troubleshooting FAQ and common errors
Contributing How to contribute to TaxTriage
Citations Tools and publications to cite

Pipeline Overview

TaxTriage ingests raw FASTQ data and processes it through the following major stages:

  1. Read QC - FastQC / NanoPlot / pycoQC
  2. Trimming - Trimgalore (Illumina) / Porechop (ONT); optional fastp / fastplong quality filtering with --enable_fastp (off by default)
  3. Host Removal - Minimap2 against host reference
  4. Metagenomics Classification - Kraken2 (+ Krona plots)
  5. Top Hits Assignment - Selects organisms for alignment
  6. Reference Preparation - Downloads assemblies from NCBI
  7. Alignment - BWA-MEM2 (Illumina) / Minimap2 (ONT/PacBio)
  8. Stats - Coverage, depth, MAPQ via samtools
  9. TASS Scoring - Confidence metric calculation
  10. Report Generation - MultiQC + Organism Discovery Report PDF

TaxTriage Schematic


Collaborative Efforts

UW Madison - Geneious Prime Plugin

Dave O'Connor's Laboratory at the University of Wisconsin-Madison developed a custom Geneious Prime plugin to run TaxTriage analyses directly from within Geneious. See the Geneious Plugin page for details.


Citation

If you use TaxTriage, please cite:

Brian B. Merritt & Jeremy D. Ratcliff et. al, TaxTriage: an open-source metagenomic sequencing data analysis pipeline enabling putative pathogen detection. doi: 10.1093/bioinformatics/btag119/8571885

See the Citations page for the full list of tools to cite.


Copyright 2022 - 2026 The Johns Hopkins University Applied Physics Laboratory LLC. See LICENSE for details.