Skip to content

Interactive Report (all.odr.html)

TaxTriage builds a single self-contained, interactive HTML report that lets you compare every sample in a run side by side. It is written to report/all.odr.html and opens in any modern browser with no server required. All data and JavaScript are inlined into the file, so it can be emailed or hosted on GitHub Pages as-is.

The interactive report is multi-sample by design. The static, per-sample deliverable is the Organism Discovery Report PDF. The two are complementary: the PDF is the fixed record, the HTML is the exploratory view.


How It Is Built

The report is generated by bin/make_report.py (module CREATE_COMPARISON_REPORT) from the assets/heatmap.html template. Its preferred inputs are the *.paths.json files emitted per sample by ALIGNMENT_PER_SAMPLE (one per sample, also published under alignment/); it can fall back to a single merged TSV/XLSX tabular report when JSON is unavailable.

*.paths.json  ──┐
                ├──►  make_report.py  ──►  report/all.odr.html
heatmap.html ───┘      (+ optional VF/AMR annotation XLSX)

Optional protein-annotation XLSX files (from --annotate_proteins / --annotate_meta) are merged in to populate the VF/AMR views.


Offline Reports

The report's own data and JavaScript are always inlined, but a handful of third-party libraries (D3, SheetJS/xlsx, jsPDF, Leaflet, Leaflet.markercluster, and Font Awesome) are loaded from a public CDN when the report is opened. By default those <script> / <link> tags are left pointing at the CDN, so the machine viewing the report needs internet access the first time it loads. For air-gapped or intermittently connected environments, two parameters fold those libraries directly into all.odr.html at build time, producing a report that opens with no network at all.

Mode What happens When to use
default (neither flag) CDN <script>/<link> tags are kept; the browser fetches the libraries on load. The viewing machine has internet. Smallest report file.
--offline_report The build step downloads each library and embeds it inline. Requires internet on the machine running the pipeline (not on the viewer). The pipeline host is online but report viewers may be offline.
--offline_report_files <dir> Local copies of the libraries in <dir> are embedded inline, so no network access is needed at all. Takes precedence over --offline_report. Fully air-gapped builds, or to pin exact library versions.

In both offline modes the fonts and marker images those stylesheets reference (Font Awesome webfonts, Leaflet markers) are embedded as base64 data URIs too, so icons and the geographic map markers render offline as well. Only the interactive map's boundary tiles still come from a public map service, but the ranked geographic list works without them.

Preparing an --offline_report_files directory

A helper script downloads exactly the files the template references (reading the CDN URLs straight out of assets/heatmap.html, so it stays in sync with version bumps) and saves them, along with the fonts/images each stylesheet needs, into assets/offline_report_libs/:

python scripts/fetch_offline_report_libs.py

Run it once from a machine with internet access, commit or copy the resulting folder alongside the pipeline, then build offline reports with:

nextflow run . <your args> --offline_report_files assets/offline_report_libs

Files are matched to the template's CDN URLs by basename, so the internal folder layout does not matter. All that matters is that each referenced file (e.g. d3.min.js, all.min.css, fa-solid-900.woff2) is present somewhere under the directory. If a referenced file is missing, the build fails with an error naming the file and its URL. A repo ships with a ready-made assets/offline_report_libs/ (see its README.md) that you can use directly.

To test the embedding without running the whole pipeline:

python bin/report_template.py -t assets/heatmap.html \
    -o report_offline.html --offline_report_files assets/offline_report_libs

See CLI Parameters → Output, Reporting, and Visualization for the flag reference.


Granularity: Strain, Species, and Genus

A View level dropdown in the controls lets you switch the entire report (tables, heatmap, TASS and coverage plots) between three taxonomic granularities, meaning the level of taxonomic detail each row represents:

View level Internal key What it shows
Strain key One row per detected strain/accession (the default, most granular view).
Species subkey Strains rolled up to their species-level taxid.
Genus toplevelkey Species rolled up to genus (or whatever rank --rank is set to).

This matters for closely related organisms (e.g. several near-identical Salmonella strains). At the strain level, reads shared between strains are ambiguous and each strain's coverage and score are diluted; rolling up to species or genus recovers a confident call. See TASS Scoring → Species/Genus roll-up scoring for how the underlying metrics are recomputed.

The View-level dropdown only changes the tables/heatmap/TASS/coverage plots. The Sunburst and Summary views are unaffected. The dropdown is enabled for both JSON and flat (TSV/XLSX) inputs; with flat inputs the species and genus rows are built on the fly when you switch.

Roll-up rescue markers

When a strain's own TASS score is below the cutoff but its species or genus roll-up passes, the heatmap flags it with a corner triangle (◤) so the call is not silently lost:

  • Orange ◤ - rescued by a species roll-up
  • Purple ◤ - rescued by a genus roll-up

A legend appears only when rescued cells are present, and each marker carries a hover tooltip ("Orange corner = species rescue / Purple corner = genus rescue"). In the tables these rows are tagged with ↑ species / ↑ genus badges. A strain only earns a badge when its parent roll-up actually clears the cutoff; if neither the strain nor its roll-up passes, the row is simply hidden.

The "Roll up threshold" checkbox (Strain view only) is orthogonal to the View-level dropdown: the dropdown chooses which granularity you look at, while the checkbox toggles strict vs. rolled-up filtering within the Strain view.

Below-cutoff organisms with VF/AMR hits

When VF/AMR annotation is present (--annotate), an organism can fall below the TASS cutoff at the strain, species and genus level, and yet still have one or more Virulence-Factor or AMR genes detected for its genus in that same sample. By default these would be hidden by the score filter, which can mask a real pathogenicity signal.

A "Show below-cutoff organisms with VF/AMR hits" checkbox in the Filters sidebar (on by default; visible only when annotations are loaded) brings those organisms back as a faded, italicised row marked with a ↓ below cutoff badge, in both the Summary → Detections table and the full Table tab. Uncheck it to disable the behaviour.

Important scoping:

  • Only organisms that are already real detections (present in the data, just below the cutoff) are brought back. Nothing is invented from annotation hits alone.
  • The match is by genus within the same sample (the granularity at which VF/AMR hits are recorded).
  • These faded rows appear only in the detection tables. KPI cards, the heatmap, the TASS/coverage plots, and the Histogram selector stay restricted to detections that pass the cutoff, so summary statistics are not inflated.

Sub-threshold & novelty-supported organisms

A "Show sub-threshold & novelty-supported organisms" checkbox (off by default; shown when novelty data is loaded) re-surfaces two further kinds of hidden row, even in samples that already have passing detections:

  • ↓ sub-threshold (aligned) (blue rail): below the cutoff, but with reference reads aligned.
  • ✦ novelty, no alignment (orange rail): no reference alignment, but the active novelty backend placed this organism's genus/species in the same sample (the Novelty column shows the matching signal).

Rows already shown by the VF/AMR toggle are de-duplicated, all sidebar filters apply, and, as with the other rescues, these rows appear only in the detection tables. See Detection Rescue for the full description of all three rescue toggles.

Right panel and cross-sample charts

  • Every sample is listed in the right panel, including those with no hits / no alignments / everything filtered out, so they can still be coloured, hidden, renamed, or inspected. Samples with few reads or no alignments carry an amber ⚠ icon whose hover warns that their plots/scores may be unreliable.
  • Cross-sample charts show detail on hover rather than crowding the plot: the Feature Compare matrix cells are colour-only (value + per-sample breakdown on hover), the Co-occurrence cells list shared/detected samples, and the Genera Comparison chart has a Show values toggle that draws numbers only where they fit (no overlap), with the per-genus total at the bar end.

Specimen grouping (merge samples)

Multiple samples (library-level: e.g. a separate DNA and RNA extraction from one swab) can be treated as a single specimen (biological-unit level). Grouping is reversible and affects the view only. It never changes the underlying per-sample data, so toggling it off instantly restores the raw per-sample rows everywhere.

Where to turn it on

  • A Merge: On/Off toggle plus Group… and Combine all buttons sit above the sample list in the right panel (shown once a run has more than one sample).
  • The same Merge: On/Off toggle also appears in the Heatmap tab's controls, next to the Specimens label. It reflects and drives the same shared state, so flipping it in either place updates the whole report.
  • Group… opens a modal where you drag sample chips into named specimen boxes (or drop onto + New specimen); Combine all merges every loaded sample into one specimen in a single click. An Edit specimens table (from the heatmap controls) offers the same grouping as a plain list of sample → specimen name fields, for quick text edits instead of dragging.

Where grouping comes from

  • The default grouping is read from the samplesheet's specimen column (or specimen_id / specimen_group in --meta), carried into SAMPLE_META[sample].specimen. A sample with no specimen assignment is its own single-member specimen.
  • Any live edit made in the report (drag-and-drop, rename, or the Edit-specimens table) is stored as an in-browser override and takes priority over the samplesheet default. Reset to samplesheet in either editor clears these overrides.

What changes when merge is ON

Every cross-sample view (Summary, Heatmap, TASS, Coverage, Table, Explore, the Run Metadata analyses, Map markers, VF/AMR genus/property charts and table, and the top banner's specimen/organism/reads counts) switches from one row per sample to one row per specimen, aggregating the member samples' rows for the same organism at the current view level:

Field(s) Aggregation Rationale
# Reads Aligned, # Unique Reads Aligned, # Reads, K2 Reads, Mean Depth Sum across members Read counts and depth from independent libraries are additive.
TASS Score, Species TASS, Genus TASS, Coverage, Covered Bases, Breadth %, Breadth Score, Breadth of Coverage, Gini Coefficient, Minhash Score, MapQ Score, Disparity Score, Diamond Identity, MicrobeRT Probability Max across members The strongest evidence in any one library wins rather than being diluted.
Mean MapQ, Mean BaseQ Read-weighted mean Avoids letting a small library's quality figure misstate a much larger merged one.
% Reads, RPM, RPKM Recomputed from the merged aligned-read count against the specimen's combined input/classified reads (visible, non-hidden members only) Keeps normalized abundance consistent with the merged denominator instead of inheriting one member's ratio.
High Consequence, Passes Threshold OR (any member true ⇒ true) A flag detected in any one library should not be lost by merging.
Mol Type "both" if members mix DNA and RNA, else the shared value Surfaces mixed-library specimens explicitly.
Every other column Taken from the member row with the highest TASS score A single representative row supplies fields with no defined aggregation rule.

A specimen with only one contributing sample passes through unaggregated. Merged rows carry a small "merged ×N" badge (hover to see the member sample names) anywhere a specimen appears, and merged Summary/Table rows are tagged so a collapsed call is distinguishable from a directly-detected one.

Prevalence and TASS-cutoff interaction

  • Prevalence % (Explore, co-occurrence, etc.) is computed against the number of distinct specimens with at least one positive detection when merge is on, instead of the number of samples. That denominator is deliberately kept separate from the live TASS slider so it reflects everything that ever tested positive, not just what currently clears the cutoff.
  • The per-sample-type TASS cutoff (see TASS Cutoff below) is unaffected by merging: the cutoff is still applied to each row's Sample Type before rows are aggregated into their specimen.

Conflicting metadata

Merging can combine samples whose run metadata disagrees (e.g. different latitude/longitude, collection_time, or run_id). Before applying a merge, the Group… modal scans every field on the affected samples (skipping pipeline bookkeeping fields like total_reads or best_cutoffs) and, for any field where members disagree:

  • Single-valued fields (e.g. location, collection time) are flagged as a warning and you pick which member's value the merged specimen keeps, via a dropdown per field.
  • Multi-valued fields (e.g. host_disease, which can carry a comma-separated list of symptoms) are automatically union-merged, meaning the values from each sample are pooled into one combined list. There is no conflict to resolve, just a note showing the combined value.

Your choices are stored per specimen and re-applied automatically the next time metadata is edited, until you regroup.

Coloring a specimen

A merged specimen can be recolored the same way an individual sample can, using a small color swatch, from either of two places, and both stay in sync (the color is stored once, keyed by specimen name):

  • The colored envelope box that wraps a merged specimen's rows directly in the right-panel sample list (a swatch sits next to the drag handle in the box's header).
  • Each specimen's box inside the Group… modal.

Picking a new color updates the box, its member rows, the heatmap/legend, and any map markers immediately.


Tabs

Tab Purpose
Summary ODR-style summary table across all samples, plus KPI cards (e.g. High Consequence count). The primary at-a-glance view.
Heatmap Organism × sample TASS heatmap with the roll-up rescue markers described above.
TASS TASS-score distribution with the per-sample-type cutoff line.
Sunburst Hierarchical taxonomic composition; slice labels render below the slices.
Coverage Per-organism coverage profiles.
Histogram Per-position genome coverage histogram, pre-filtered to a chosen organism + sample. Hidden/disabled for run-level species or genus views where per-position data is not meaningful.
Explore Scatter/bubble and correlogram exploration of metrics across organisms and samples.
Table Full sortable, paginated data table with configurable columns (all columns shown by default, including MicrobeRT Probability / MicrobeRT Model when --microbert was used).
Proteins (VF/AMR) Virulence-factor and AMR gene annotation table. Hidden unless protein annotations are present.
Novelty Reference-free / open-set detection on the de novo contigs via the chosen backend (kaiju / bracken / mmseqs2): per-sample novelty score & flag, candidate genus+ taxa, a method-coverage breakdown (alignment/TASS vs backend rescue vs dark matter), and a genus TASS-vs-backend comparison. Column labels track the active backend. Hidden unless --novelty produced data. See Novelty Detection.
Map Geographic view; sample collection map plotted from latitude / longitude. Hidden unless run metadata supplies coordinates.
Run Metadata Editable per-sample metadata table plus four analysis sub-tabs (Longitudinal, Geographic Comparison, Host & Disease, Cross-Entry Comparison) driven by recognized metadata columns; see Run Metadata tab and metadata-driven views. Hidden unless samples or metadata are present.

Summary tab details

  • Group by Sample is on by default; rows sort by sample, then by TASS descending within each group.
  • Each organism links out to the NCBI Taxonomy browser (icon on the organism name in the Summary tab; the Taxonomic ID is linkified in the Table tab).
  • Sample Type is its own column (no longer folded into the keyword chips).
  • The sparkline in each row draws the genome per-position coverage profile; clicking it jumps to the Histogram tab pre-filtered to that organism + sample, and hovering previews coverage across positions.
  • The High Consequence KPI card lists the flagged organisms and the samples they were seen in on hover.

MicrobeRT columns

When the run used --microbert, two columns are available in the Table tab (shown by default):

  • MicrobeRT Probability: the model's confidence (0 - 100 %), averaged over the reads aligned to that organism's reference accession(s) and weighted by read support.
  • MicrobeRT Model: the model directory name used for prediction.

The value is linked to each organism by reference accession (parsed from the clustered-read IDs), not by the model's predicted taxonomy, so it populates even when the model classifies into a different taxonomic namespace than the detected organism. The value reflects what the chosen model was trained on; see CLI Parameters → MicroBERT for the model-domain caveat. The same value appears as mmbert% in the Organism Discovery Report PDF.


In-Silico suite tab

The In-Silico tab appears whenever a run produced series datasets - --sim_subsample on simulated reads, --background_reads on a real matrix, or --spikein_sheet for a spike-in series (see In-Silico → Choosing an experiment). Each subsample dataset flows through the pipeline as its own sample, so the report already holds the observed result at every dilution; the tab aggregates those into an expected-vs-reality view. It has four parts:

Parameters used. A provenance panel listing the subsampling mode (consistent / randomized), the read-count series, replicates, seed, master read count, ISS/NanoSim settings, and abundance source. These come from an insilico_params.json emitted at report time (with any gaps inferred from the subsample sample names). The panel header also carries Export suite buttons - CSV / TSV (two files: one per-dataset table, one long-format per-organism series), XLSX (one workbook with Parameters, Datasets and Organism series sheets) and JSON (the raw suite payload the report holds). These export every group at once, including columns not shown on screen (dataset id, seed, master pool size, log2 fold change); the per-plot and per-table download buttons still export a single figure or table.

Per-dataset expected vs reality. For each (parent sample × platform) group, one row per subsample dataset (count × replicate) shows the target vs actual read count, the observed aligned reads, and TP/FP/FN plus precision / recall / F1. "Expected" organisms are those recovered at full depth; a dataset that drops a low-abundance organism at shallow depth records it as a false negative. Authoritative target/actual/master counts come from the *_subsample_manifest.tsv files; when those are absent the target count is parsed from the dataset name.

Run metrics across the dilution series. Three charts per group, averaged over replicates at each target depth:

  • Performance vs depth - precision, recall and F1 as depth increases.
  • Detection composition vs depth - stacked mean TP / FP / FN organism counts, showing whether a shallow depth loses organisms (FN) or invents them (FP).
  • Read recovery vs depth - reads present in the dataset vs reads that actually aligned to a reference, on a √ scale.

Per-organism dilution series (limit of detection). For every simulated organism, a compact chart plots expected vs observed reads across the read-count series, with a per-count detection dot and a LoD badge marking the lowest read count at which the organism is still recovered (detection rate ≥ 50 % across replicates). This is the at-a-glance answer to "how deep do I need to sequence to still catch this organism?" A dashed rule marks the LoD depth; a detection dot is green when every replicate detected the organism, amber when only some did, and grey when none did.

A Level control above this section switches between the level the series was simulated at and a Genus rollup, which merges members of the same genus into one series (reads and expected share add, TASS takes the strongest member, and the genus counts as detected wherever any member was). The rollup is reflected in the card subtitles, in the member list on hover, and in the suite export, which always exports the level currently on screen.

The detection cutoff, and how a spike-in run can improve on it. The cutoff the tab detects against is the calculated recommendation - best_cutoffs.subkey.best_threshold, derived from historical sample-type data - not --min_conf/--mintass, which is only a fallback when no recommendation exists. For a spike-in run the tab adds a Detection cutoff card:

  • A slider that recomputes every organism's detection and limit of detection live, from the per-replicate TASS values in the payload. The dataset table and the three run-metric charts stay at the build cutoff and say so.
  • A recommendation derived from these LoD curves, chosen by maximising F1 over the whole series: a spiked organism called where it was spiked is a TP, missed is a FN, and a FP is either a spiked organism called in the level-0 blank (the score cannot separate spike from matrix at that cutoff) or a call that is neither spiked nor part of the matrix. Ties resolve to the middle of the widest tied run rather than to an edge. The F1-vs-cutoff curve is drawn with the run, recommended and current cutoffs marked.
  • Level 0 is never a limit of detection - nothing was spiked into it. An organism called there gets an in blank badge, because the LoD shown beside it is optimistic at that cutoff.

What counts as truth in a spike-in run. A dilution series takes its truth from the deepest dataset, which is right when the whole community was simulated. A spike-in series does not: the truth set is exactly what the sheet says was spiked, matched by taxid (resolved from the accession at fetch time, since the sheet's name column is optional). Everything else detected is the background matrix - genuinely present, so it is never counted as a false positive. The matrix is listed separately under Background matrix, from the level-0 control.

The real samples. Because a spike-in run analyses the background and its spiked copies, the samples the run is actually about would otherwise be absent from this tab. Two tables close that gap: Real samples vs this background (per sample: organisms called, how many are also in the matrix, how many are unique, which spiked organisms appear) and Spiked organisms in the real samples, which places each sample's load for each spiked organism against that organism's LoD at the current cutoff.

Cross-reference from the Detections table. When an organism in the Detections table was also part of the in-silico pool, a small ⚗ chip appears beside its name. Hovering shows a depth-matched comparison card; clicking opens the full comparison (reads-vs-depth and TASS-vs-depth panels plus the series datapoints).

The comparison is depth-relative, which is the point: the real sample is placed on the dilution axis at its own sequencing depth, not compared against the deepest dataset. If the series runs 5k / 8k / 10k / 15k / 20k / 30k and the sample carries 20k reads, the sample is drawn at 20k and measured against the series value interpolated at 20k. Three reference values are shown at that depth:

  • Series at this depth - the observed reads the dilution series actually produced there (piecewise-linear between the bracketing datasets; scaled through the origin below the shallowest and extrapolated along the last segment above the deepest, flagged when extrapolated).
  • Expected from pool share - this organism's fraction of the simulated pool times the sample's depth. Compositional, so it is exact at any depth.
  • Limit of detection - and how many times the sample's depth exceeds (or falls short of) it.

Levels: strains, species and genera. The suite is built at a single level (group.level - Species when the data supports it, else Strain), so a detection can sit below the level its series exists at: an Orthopoxvirus strain detection has no series of its own while its species does. Matching therefore walks the row's lineage - the row's own taxon, then its species, then its genus - and always says which rung it landed on:

  • The chip carries a rung marker: ⚗ sp for a strain shown against its species' series, ⚗ gen for a genus series.
  • The card and modal open with a banner reading, for example, "Strain detection · Species series - You are looking at a strain detection. The dilution series exists at species level (Monkeypox virus), so this plot is the species's series."
  • Where the report also holds the matching Species/Genus row for that sample (the rollup levels are all in the data), that row's reads and TASS are used for the comparison so both sides are at the same level, and the clicked row's own numbers are kept alongside - "This strain's reads 2,600 (50% of the species)". If no such row exists, the strain's own reads are plotted and the card says a strain is a subset of its species, so sitting below the series is arithmetic rather than a finding (the verdict is softened to match, instead of crying under-recovery).
  • A genus comparison has no directly simulated counterpart, so the series is assembled from every simulated member of that genus: reads and expected share add, TASS takes the strongest member, and the genus counts as detected wherever any member was. The banner names the members.

The chip is colour-coded by the verdict: green when the sample is within 0.5 - 2× of the series, amber with a ↓ when it is under-recovered, purple with a ↑ when over-recovered, and red when the sample was sequenced below the organism's LoD - the case worth acting on, since the series says this depth cannot reliably see this organism. TASS is compared the same way, against the run's detection cutoff.

Paired-end runs are reconciled automatically: the series counts read pairs while the report's read columns count each mate, so the real sample is halved before comparison and both sides are labelled in the group's unit. If the organism was simulated in more than one group, the sample's own group is used and the modal offers a selector for the others.

Hover statistics. Every dataset row, chart bar and chart point raises the shared report tooltip with the full statistics for that point - target vs actual vs aligned reads and alignment recovery, TP/FP/FN with precision / recall / F1 (plus the F1 standard deviation across replicates), and per organism the expected vs observed reads, recovery, log2 fold change, mean TASS, detection rate across replicates and the LoD. Every plot reserves its own axis gutter, detection strip and label row, so marks and axis labels never overlap.

Expected per-organism composition is modeled from the full-depth (deepest) dataset's observed proportions scaled to each target read count. This is a practical proxy for the simulated ground truth; if you need the exact simulated abundances instead, they are available in simulation/<sample>/ (the accession-level abundance.tsv).


Run Metadata tab and metadata-driven views

The Run Metadata tab shows an editable per-sample metadata table and unlocks several cross-sample analyses. Every metadata column you supply, either via extra samplesheet columns or a separate --meta CSV/XLSX (see Samplesheet → Metadata Support), appears in this table and is carried into the report. A set of recognized column names additionally power dedicated views; columns are matched by exact name (snake_case), so spelling matters.

The table itself is editable in-browser: Add column, Rows for all samples, Fill location from lat/long (reverse-geocodes latitude / longitude into country / state, needs internet for boundary data), and Export (XLSX / CSV / TSV). Edits are saved into the exported report state.

Recognized metadata columns and what they drive

Column Drives Notes
latitude, longitude Map tab, Geographic Comparison, Fill location from lat/long Numeric. Presence of any lat/lon pair enables the Map tab (has_geo).
collection_time Longitudinal Analysis sub-tab (X axis) ISO (YYYY-MM-DD[ HH:MM:SS]) or M/D/YYYY. A sample needs a parseable value to appear on the timeline.
sample_origin_country Geographic Comparison choropleth (Country level) Standardized origin field.
sample_origin_state_province_territory Geographic Comparison choropleth (State / Province level, which is the default) Standardized origin field.
host_scientific_name Host & Disease (group by host)
host_disease Host & Disease (default grouping) and the Symptom × Organism matrix Multi-value: a comma-separated cell (e.g. runny stool, cramps, cold-like symptoms) is split into individual symptoms.
environmental_site Host & Disease (group by environmental site) For environmental rather than host-associated samples.
run_id Groups samples by run Also used in the ODR PDF.
depth, salinity, location Displayed in the metadata table depth / salinity are numeric (m / PSU).
sequencing_instrument, sequencing_platform, library_preparation_kit, sequencing_protocol_primer_set Displayed with friendly labels Recognized so they render with readable headers.

Any other column you add is still shown in the table and is available for display/filtering; it just does not feed a dedicated chart.

Grouping and picker controls

Wherever the report lets you pick from a list that can grow with the run - the Group by legend (shared bar, map, Group Heatmap, Group Network) and the Longitudinal Analysis organism picker - the control is a collapsed dropdown with a search box rather than a row of chips. One line high whatever the run size, and you can type to find an entry instead of scanning a wrapped block.

For the group picker every group starts ticked (visible), and it keeps the three states the chips had:

Action Effect
Untick the checkbox Hidden - the group drops out of every grouped view and off the map
Click the row Cycles normal → highlight → hidden; highlight emphasises that group, fades the rest, and draws its similarity network on the map
Shift-click the row Steps backwards through the cycle
Alt-click the row Solo - show only that group (alt-click again to restore)
All / None Show or hide every group at once
Reset (n) Appears beside the dropdown once anything is off-default; returns all groups to normal

All four group pickers share one state, so hiding a group in the Group Heatmap is reflected immediately in the bar, the map and the network.

Analysis sub-tabs

The Run Metadata tab exposes four sub-tabs. Each shows an explanatory warning (instead of an empty chart) when the metadata it needs is absent:

  • Longitudinal Analysis: plots a chosen metric (TASS, Coverage, Breadth %, # Reads Aligned, Mean Depth, Minhash Score) over collection_time for selected organisms, with linear / log / sqrt Y scaling. Requires a parseable collection_time column.
  • Geographic Comparison: a choropleth aggregating the visible detections by sample_origin_country or sample_origin_state_province_territory, shaded by Detections, Samples, # Reads Aligned, or Mean TASS. Hover a country for its state breakdown; click to drill in; Shift-click / right-click to pin a region. Boundary outlines load from a public map CDN (the ranked list still works offline).
  • Host & Disease: bar aggregation of detections grouped by host_disease, host_scientific_name, or environmental_site, plus a Symptom × Organism matrix counting, per symptom (from host_disease), the samples in which each organism was detected. Click a symptom row to drill in.
  • Cross-Entry Comparison: compares detection profiles across metadata entries; available once at least two entries are present.

All four respect the sidebar filters and recompute live, so they reflect the current TASS cutoff and view level.


TASS Cutoff (per sample type)

The TASS cutoff is per sample type, not a single global value. The cutoff slider is pre-populated from the best_cutoffs recorded in the input data (the most conservative recommended threshold across loaded samples), drawn from assets/sampletype_best_thresholds.json (matched on platform and body site). Moving the slider re-filters every tab live. See CLI Parameters → --thresholds_json.


Sample QC flags

Every other filter in the report narrows detections. Sample QC narrows samples: a small rule engine that answers "does this whole sample look usable?" - enough reads, enough distinct organisms above a TASS cutoff, the right metadata - and marks the ones that fail.

The Sample QC / Flags block in the right-hand sidebar reports how many samples are flagged and opens the rule builder; the dialog itself lists exactly which samples tripped which rule. Each rule is one {field} {operator} {value} clause drawn from four sources:

Source Examples
Sample metrics Total reads, aligned reads, organism-assigned reads, species/strain key counts, platform, sample type, control type
Detection counts Distinct organisms above a TASS cutoff, detections passing threshold, high-consequence organisms, distinct genera, highest TASS
Metadata column Any column in the Metadata & Mapping table, including ones added or uploaded in the report - equals, contains, regex, is empty, …
Detection column (aggregated) Any numeric detection column (Coverage, Mean Depth, RPM …) rolled up per sample by max / min / mean / sum / count

Rules combine with any (flag when one matches) or all (flag only when every rule matches), and each carries its own action:

  • flag it - the sample stays fully visible and is marked everywhere.
  • flag & hide it - the sample is additionally removed from every chart and table, exactly as the sidebar eye icon hides one. Nothing is deleted: clearing the rule, or the Hide flagged checkbox, brings it straight back.

A Missing values count as a match switch decides whether a sample with no value at all for a field trips the rule; it is off by default, so a criterion can never flag a sample purely because the field was never populated.

Where flags appear

Tab Marker
Heatmap The sample's column is outlined in dashed amber with a solid cap under the axis, and its header carries a ⚑ glyph. The cell colours are left untouched, since they are the data.
Table The sample's group header is tinted, carries a left rule and a flagged badge (grouped view).
Metadata & Mapping The sample's row is tinted and its name carries the badge.
Summary A Flagged Samples KPI card (hover for the per-sample reasons), plus the same tinted group headers.
Right sidebar A flag icon on each sample row - hover for the reasons, click to open the rule builder.

Hovering any marker lists exactly which rules the sample tripped and what its actual values were.

Defaults from the pipeline

Rules can ship with the run: the --report_flag_* parameters are baked into the report so it opens with them already applied. Reset to pipeline defaults in the dialog restores that set at any time. Live edits are saved with the exported session state, so a saved-and-reloaded report comes back with your rules, not the pipeline's. See CLI Parameters → Report Sample-QC Flags.

Organism QC flags

Sample QC judges whole samples; Organism QC judges individual detections (one organism in one sample) and either highlights them or hides them from every view. The Organism QC / Flags block sits under Sample QC in the right-hand sidebar: an On switch for the whole rule set, the Filter Organisms button that opens the rule builder, and a dropdown that switches the report between:

  • Highlight flagged - every row stays; flagged ones are marked (a rule whose action is flag & hide it still hides its own matches).
  • Hide flagged - every flagged row is removed from every chart and table.
  • Only flagged - nothing but the flagged rows (handy for reviewing what a rule catches).

A rule is one or more conditions that must all hold; separate rules are independent (a row is flagged when any rule matches). That is what lets one rule say "genus is Streptococcus and fewer than 50 reads". Conditions come from two sources:

Source Fields
Detection column Any column of the row: # Reads Aligned, TASS Score, Genus, Family, Microbial Category, Breadth %, K2 Reads … Text columns take equals, contains, is one of (comma list), regex, …
In-sample context Compared with the other hits of the same sample at the same level: shared ANI % with another hit (and # hits sharing high ANI), qualified by stronger hit (only a partner with more reads, then higher TASS, counts - so of two near-identical references only the weaker is flagged) or any hit; genus reads in this sample; % of its genus's reads; rank within its genus (1 = top); # hits in the same genus; % of the sample's organism-aligned reads; aligned ÷ classifier (K2) reads; # samples detected in; taxonomy lineage text (any rank).

The ANI fields need a run with --enable_matrix; the pipeline only records partners at or above --ani_threshold (default 95 %). A detection without ANI data never matches an ANI comparison (use is empty to find those). Values are taken from the full dataset, never from what the current filters display, so a flag does not flicker as you move the TASS slider.

Presets cover the common cases: Few reads, Low TASS, Shared ANI with a stronger hit, Genus + few reads (type the genera), Minor member of a genus and Classifier-only. The dialog previews every flagged row with the reason and actual values, and counts the rows each rule catches.

Flagged rows carry an amber QC badge and a left rule in the Summary detections table and the Table tab, and an amber corner in the bottom-left of their heatmap cell (the cell tooltip lists the reasons). The Detections export gains an Organism QC Flag column. Pipeline defaults come from the --report_org_flag_* parameters (CLI Parameters → Report Organism-QC Flags); Reset to pipeline defaults restores them, and live edits are saved with the session state.


Export

One Export… button sits under the Filters heading in the right panel (or press Ctrl/⌘+Shift+E). It opens a single popup where the first choice is what to take away:

  • Data tables - the underlying tables from as many tabs as you like, as one file
  • Report PDF - the current filtered report as a printable layout

Data tables

The dialog lists every table the report holds, grouped by the tab it belongs to, with a live row count next to each; tables the run carries no data for are greyed out. Tick what you want, then pick an output shape:

Shape What you get
Excel workbook - one sheet per table One .xlsx, one sheet per selected table, plus an Export Info sheet recording the run, the filters in force and the row count of each sheet.
Excel / CSV - single joined table Every selected table folded into ONE table, joined on Specimen ID × Organism. Columns from each source are prefixed with its name (Coverage Summary · Mean Depth).
CSV - all tables stacked Every selected table in one CSV, one after another, with a leading Dataset column and the union of all columns.
Pivot - organisms × a metadata field A crosstab of the detections against run metadata. See below.

Choosing columns

A detections export is 53 columns wide; most reviews need a handful. Every table in the list carries a columns chip on its right (all 53 cols) - click it for a multi-select of that table's columns, the same show-all / hide-all pattern as a table's own Table Display dialog, plus a filter box for finding a column by name and Select filtered to take everything the filter matches. The chip then reads 7 of 53 cols and highlights, so a narrowed table is obvious at a glance.

The choice is per table and applies to every output shape, and the Export Info sheet records exactly which columns each table kept. Dropping a column never breaks the joined shapes: the Specimen ID × Organism join still keys on those fields internally even when you have taken them out of the printed output.

Apply current filters (on by default) exports what the sidebar filters, the TASS cutoff, the view level and sample visibility leave in view; turn it off for the complete underlying data. Either way the export is the full dataset behind a table, not just the page of it you can see - table pagination does not truncate it, and a table exports correctly even if you never opened its tab.

A few tables - per-gene VF/AMR hits, novelty candidates, in-silico series, per-contig coverage - carry several rows per organism, so they cannot be folded into a single joined row. Choosing a joined shape with one of those selected shows a warning naming them; use the workbook or the stacked CSV to include them.

Counting detections against metadata (pivot)

"How many hits to Influenza A came from each collection site?" is a question no single tab answers: the counts live in the detections, the site lives in the run metadata. The Pivot output shape joins the two and aggregates:

  • Rows - Detected Organism (default), Organism + Taxid, Genus, Microbial Category, Domain, Specimen ID or Sample Type
  • Columns - any run-metadata field the run carries: location, host type, host disease, country, sequencing platform, run id, and anything you uploaded on the Metadata tab. Each option shows how many distinct values it has; continuous numerics (read counts, coordinates) are left out, since a crosstab keyed on those would have one column per sample. Leave it on (none) for a plain rollup with totals only.
  • Count - # detections, # specimens, # distinct organisms, total reads aligned, mean TASS or max TASS
  • Layout - wide puts one column per metadata value, the crosstab shape you read in a spreadsheet; long emits one row per organism × value pair (Organism, Field, Measure, Value), the tidy shape R, pandas and most plotting libraries want

A live preview shows the first rows and columns before you download. Samples with nothing recorded for the chosen field are collected under (not recorded) rather than dropped, so the column totals still add up. The pivot is built from the detections plus metadata, so it ignores the table selection above it.

For the raw material behind it, tick Detections + Metadata in the table list: one row per detection with every metadata column joined on, ready to drop into your own pivot table.

Same tables, without opening the report

The pipeline can write all of this during the run with --export_data, into <outdir>/report/export_data/ - pivot included, via --export_data_pivot. The tables and column headers are identical, so a spreadsheet from the pipeline lines up with one exported by hand here. See CLI Parameters → Combined Data Export.

Detections and the TASS cutoff

Passes Threshold in the underlying data is always unset - the report decides pass/fail live against the cutoff. Every detections export therefore carries two extra columns: TASS Cutoff (the cutoff that applied to that sample) and Passes Cutoff (the verdict, honouring the species/genus rollup rescue when it is on).


Report PDF

The popup's second type renders the current report - with whatever filters, view level and tab state are active - to a static printable layout for sharing or archiving. Choose where the sample-colour and TASS-cutoff legend goes (cover page, every page, or nowhere), then Prepare PDF: the report walks every tab, then opens your browser's print dialog. Set the destination to Save as PDF; landscape orientation with Headers and footers off gives the best result.

Exporting a single plot, table or the map

Hovering a plot shows a small download button in its corner (tables get one permanently, next to the column/font settings button). It opens an Export dialog: plots offer PNG / JPEG / SVG / PDF / HTML at any width and height, tables offer XLSX / CSV / TSV with a choice of delimiter.

The Precise (lat/long) map in Mapping & Geography carries the same button, in the top-right control stack next to the cluster-mode box. The map is not an SVG like the other plots - it is a stack of raster basemap tiles plus HTML marker icons - so it is composed into an equivalent SVG first (tiles embedded as base64 images, markers redrawn as vectors, the current basemap's attribution stamped in) and then goes through the same formats and size fields as every other plot. The export captures the map exactly as framed on screen: current zoom, pan, marker colours/shapes, cluster bubbles and filter dimming. Embedding the basemap needs the tile server to allow a cross-origin read; if it refuses (or the report is offline), the export still writes the markers, on a plain background. The basemap picker defaults to Esri Light Gray; OpenStreetMap is offered but is not the default, because OSM's tile usage policy blocks viewers it cannot attribute - a report opened straight from disk sends no referrer and gets an "Access blocked" tile back. The map detects that refusal and falls back to the next basemap in the list.


Tooltips and Annotations

  • Organism-name cells show a help cursor and a detail tooltip on hover (only on the name cell, not other columns).
  • Right-panel / view-level controls carry static tooltips explaining what each option affects.
  • In the Summary table, top TASS scores for a species or genus show a tooltip breaking down the contributing strains.
  • VF / "# gene-hit" cells in the annotation table show a breakdown of the top gene / AMR hit names with counts (active only when protein annotations are present).
  • Rollup summary rows are annotated to distinguish an inferred species/genus call from a directly detected strain.