# Illumina® DRAGEN™ Secondary Analysis

Illumina DRAGEN (Dynamic Read Analysis for GENomics) secondary analysis was developed to address important challenges associated with analyzing NGS (Next Generation Sequencing) data for a range of applications, including genome, exome, transcriptome, and methylome studies. DRAGEN secondary analysis processes NGS data and enables tertiary analysis to drive insights. The available tools make up a highly accurate, comprehensive, and efficient solution that enables labs of all sizes and disciplines to do more with their genomic data.

**Product highlights**

**Accurate results:**

* Pangenome reference genome and machine learning drive unprecedented accuracy
* 99.89% accuracy score with the Precision FDA Truth Challenge V2 benchmark data (*2,3*)

**Comprehensive platform:**

* Analyze NGS data from whole genomes, exomes, methylomes, and transcriptomes
* Available on platform of choice and scalable based on needs

**Efficient analysis:**

* Process a 34x genome in \~ 30 minutes, with all supported callers with DRAGEN server v4 (*1*)
* Reduce FASTQ file sizes up to 5x with DRAGEN ORA Compression

\
\
*References:*

1. Illumina data on file, 2022.
2. Illumina DRAGEN Secondary Analysis is the first single platform to achieve 99.89% accuracy based on [PrecisionFDA v2 Truth Challenge Benchmark Data](https://precision.fda.gov/challenges/10). Details here [DRAGEN sets new standard for data accuracy in PrecisionFDA benchmark data](https://www.illumina.com/science/genomics-research/articles/dragen-shines-again-precisionfda-truth-challenge-v2.html). Accessed March 22, 2023
3. PrecisionFDA Truth Challenge V2: Calling Variants from Short and Long Reads in Difficult-to-Map Regions. [precision.fda.gov/challenges/10](https://precision.fda.gov/challenges/10). Accessed November 3, 2020.


# DRAGEN Applications

## Applications

DRAGEN analysis offers a large selection of application pipelines.

| Pipeline                                             | Description                                                                                                                                                                                                                                                                                                   | Variant Types Detected             | Metrics Provided                                                                     |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------------------------------------------ |
| DRAGEN Demultiplexing                                | Rapid demultiplexing of NGS analysis                                                                                                                                                                                                                                                                          | N/A                                | N/A                                                                                  |
| DRAGEN ORA Compression                               | DRAGEN ORA compression is optimized for high compression ratios of FASTQ files, as well as rapid compression and decompression, all while preserving data integrity.                                                                                                                                          | N/A                                | Compression Ratio Run Time                                                           |
| DRAGEN Map + Align                                   | The DRAGEN Map + Align can be run as a standalone or as part of DRAGEN’s suite of pipelines                                                                                                                                                                                                                   | N/A                                | Mapping metrics Duration Metrics Coverage Metrics                                    |
| DRAGEN Germline                                      | The DRAGEN Germline Pipeline provides end-to-end NGS analysis, including advanced error model calibration for increased accuracy, and repeat expansion detection and genotyping through Illumina Expansion Hunter.                                                                                            | SNV/Indel CNV SV Repeat Expansions | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Somatic                                       | The DRAGEN Somatic Pipeline includes tumor-only and tumor–normal modes, designed for detecting somatic variants in tumor samples. Both modes make no ploidy assumptions, enabling detection of low-frequency alleles.                                                                                         | SNV/Indel CNV SV TMB MSI HLA       | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Enrichment                                    | The DRAGEN Enrichment Pipeline combines DRAGEN’s germline and somatic callers into a pipeline designed specifically for analyzing enrichment samples. Includes a full suite of enrichment metrics and reporting.                                                                                              | SNV/Indel CNV SV                   | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN RNA                                           | The DRAGEN RNA Pipeline performs transcriptome analysis starting with splice junction discovery and alignment, followed by rapid alignment and splice junction mapping and quantification. For differential expression, Illumina recommends the DRAGEN Differential Expression app on BaseSpace Sequence Hub. | Gene fusion SNV/Indel              | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Single Cell RNA                               | The DRAGEN Single Cell RNA pipeline performs demultiplexing, cell-barcode and UMI error correction, sequence alignment, and quantification of gene expression.                                                                                                                                                | N/A                                | Mapping Metrics Duration Metrics Coverage Metrics Callability Report Cell Metrics    |
| DRAGEN Joint Genotyping                              | The DRAGEN Joint Genotyping/Population Pipeline calls variants jointly across multiple genomes and scales to large cohorts of samples at expedited speeds with uncompromising accuracy.                                                                                                                       | SNV/Indel CNV SV Repeat Expansions | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Methylation                                   | The DRAGEN Methylation Pipeline performs alignment, methyl calling, and calculates alignment and methylation metrics.                                                                                                                                                                                         | N/A                                | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Reference Builder                             | Accepts FASTA files, and builds the proprietary reference used by the DRAGEN apps.                                                                                                                                                                                                                            | N/A                                | N/A                                                                                  |
| DRAGEN TruSight Oncology 500 ctDNA Analysis Software | Secondary analysis support for Illumina’s TruSight Oncology 500 ctDNA. Available on the local DRAGEN Server version 3 and later.                                                                                                                                                                              | SNV/Indel CNV DNA fusions MSI TMB  | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Imputation                                    | The DRAGEN Imputation pipeline is an end to end user friendly tool that enables scalable low pass whole genome sequencing analysis                                                                                                                                                                            | N/A                                | Impute ≤100 samples simultaneously 1.7x faster compared to original GLIMPSE code     |

## Analysis Uses

DRAGEN analysis can be used in numerous fields in the biological sciences.

| Analysis                   | Description                                                                                                                    |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Genetic Diseases           | Reduce time required for genomic analysis, with high accuracy and comprehensiveness                                            |
| Oncology                   | Analyze tumor-only and tumor/normal samples with accuracy, comprehensiveness, and efficiency                                   |
| Cell and Molecular Biology | Advance understanding of cellular mechanisms with rapid analysis pipelines for bulk and single cell samples                    |
| Population Genomics        | Accurately and efficiently analyze sequenced genomes at scale. Accelerate re-analysis as computational tools improve over time |
| Infectious Disease         | Detect and characterize infectious diseases with a comprehensive solution                                                      |
| Agrigenomics               | Efficiently analyze animals and plants of varying genomic complexities with custom reference                                   |


# DRAGEN User Guides

Visit the links below for DRAGEN technical user guides.

[DRAGEN v4.5](https://help.dragen.illumina.com/dragen-v4.5)

[DRAGEN v4.5 for Illumina TruPath Genome](https://help.dragen.illumina.com/dragen-v4.5-trupath)

[DRAGEN v4.4](https://help.dragen.illumina.com/dragen-v4.4)

[DRAGEN v4.3](https://help.dragen.illumina.com/dragen-v4.3)


# Deployment Options

DRAGEN analysis is available on multiple platforms.

| Platform                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| DRAGEN on-premises server test                       | <p>DRAGEN on-premises server offers highly accurate secondary analysis in a fraction of time compared with a traditional CPU-based system.<br>- Analyze and store data locally<br>- Supports varying levels of command line interface<br>- Replace up to 30 traditional compute instances<br>- Fully process a 34× whole human genome in \~30 minutes. <em>(1)</em><br>- One unit supports two NovaSeq 6000 Systems running at full capacity</p>       |
| DRAGEN analysis on Illumina Connected Analytics      | Couples the accuracy and speed of the DRAGEN with the ability to customize analysis pipeline to operationalize informatics on a secure platform.                                                                                                                                                                                                                                                                                                       |
| DRAGEN on BaseSpace Sequence Hub (BSSH)              | Push button analysis capability in an intuitive, easy-to-use interface with compliance, and storage features of BaseSpace Sequence Hub and Amazon Web Services (AWS).                                                                                                                                                                                                                                                                                  |
| DRAGEN onboard NovaSeq X Series                      | <p>- Flexibly runs multiple secondary analysis pipelines in parallel.<br>- Performs up to four simultaneous applications per flow cell in a single run.<br>- Brings up to 5x lossless data compression, and analysis with supported applications<br>- Provides savings on analysis, which over five years can exceed the price of the sequencer</p>                                                                                                    |
| DRAGEN onboard NextSeq 1000 and NextSeq 2000 Systems | <p>- Provides access to select DRAGEN analysis informatics pipelines<br>- Enables users to generate results in as little as two hours<br>- Uses intuitive pipeline algorithms to reduce reliance on external informatics experts</p>                                                                                                                                                                                                                   |
| DRAGEN onboard MiSeq i100 Series                     | <p>Intuitive, ultra-rapid analysis including DRAGEN BCL convert, DRAGEN Library QC, DRAGEN small WGS and DRAGEN Microbial Enrichment Plus.<br>- Rapid results with comprehensive secondary analysis generated in two hours or less <em>(2)</em><br>- Highly efficient workflow with a single user touchpoint to VCF and/or html report and no intermediate file transfers<br>- Exceptionally easy with an intuitive interface for non-expert users</p> |
| DRAGEN on AWS, Azure                                 | DRAGEN supports the FPGA enabled instance types of AWS, Azure. Rpm installers and the Kernel driver can be installed on images managed by the user, and DRAGEN can be run by purchasing a license.                                                                                                                                                                                                                                                     |
| DRAGEN on AWS and Azure Marketplace                  | Pre-configured Amazon Machine Images (AMI) and Azure Virtual Machines with DRAGEN installed can be accessed from the respective marketplace offerings in a Pay-As-You-Use model.                                                                                                                                                                                                                                                                       |
| DRAGEN on GCP                                        | DRAGEN is made available on the Google Cloud Platform. Pre-configured instances with DRAGEN installed can be accessed through the GCP application interface. Limited availability. Please reach out to your Illumina representative for access.                                                                                                                                                                                                        |

> (1) HG002 from PrecisionFDA truth challenge V2 run with DRAGEN analysis v4.0 on DRAGEN server v4, all callers

> (2) When run according to sample recommendations


# Illumina Assay to DRAGEN Cloud Pipeline Mapping

This table identifies the appropriate DRAGEN pipeline for analyzing data output from Illumina assay kits. Visit [Illumina BioInsight Platform Core](https://help.ica.illumina.com/) or [Illumina BaseSpace Sequence Hub](https://help.basespace.illumina.com/) for more information on accessing DRAGEN pipelines in the cloud.

| Assay                                                                                                                                     | DRAGEN Pipeline                       |
| ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| Illumina Spatial Transcriptome Prep, FF (L) Cloud                                                                                         | DRAGEN Spatial Transcriptome          |
| Illumina Spatial Transcriptome Prep, FF (S) Cloud                                                                                         | DRAGEN Spatial Transcriptome          |
| Illumina Spatial Transcriptome Prep, FF (L) Cloud Starter                                                                                 | DRAGEN Spatial Transcriptome          |
| Illumina Spatial Transcriptome Prep, FF (S) Cloud Starter                                                                                 | DRAGEN Spatial Transcriptome          |
| ILMN Protein Prep 9.5k Plasma 96 Reactions                                                                                                | DRAGEN Protein Quantification         |
| ILMN Protein Prep 9.5k Serum 96 Reactions                                                                                                 | DRAGEN Protein Quantification         |
| Illumina 5-Base DNA Prep (24 Samples) Early Access                                                                                        | DRAGEN Germline: 5-Base               |
| Illumina 5-Base DNA Prep with Enrichment (24 Samples) Early Access                                                                        | DRAGEN Enrichment: 5-Base             |
| Illumina 5-Base DNA Prep (24 Samples)                                                                                                     | DRAGEN Germline: 5-Base               |
| Illumina 5-Base DNA Prep with Enrichment (24 Samples)                                                                                     | DRAGEN Enrichment: 5-Base             |
| Illumina Single Cell 3' RNA Prep, T2 (8 samples, 2000 cells/sample)                                                                       | DRAGEN Single Cell RNA: T2            |
| Illumina Single Cell 3' RNA Prep, T10 (8 samples, 10,000 cells/sample)                                                                    | DRAGEN Single Cell RNA: T10           |
| Illumina Single Cell 3' RNA Prep, T20 (4 samples, 20,000 cells/sample)                                                                    | DRAGEN Single Cell RNA: T20           |
| Illumina Single Cell 3' RNA Prep, T100 (2 samples, 100,000 cells/sample)                                                                  | DRAGEN Single Cell RNA: T100          |
| Illumina Single Cell CRISPR Prep, T10 (8 samples, 10,000 cells/sample)                                                                    | DRAGEN Single Cell CRISPR: T10        |
| Illumina Single Cell CRISPR Prep, T100 (2 samples, 100,000 cells/sample)                                                                  | DRAGEN Single Cell CRISPR: T100       |
| Illumina Single Cell CRISPR Prep, M1 (1 sample, 1 million cells/sample)                                                                   | DRAGEN Single Cell CRISPR: M1         |
| Illumina Single Cell CRISPR Prep, M10 (10 samples, 1 million cells/sample)                                                                | DRAGEN Single Cell CRISPR: M10        |
| The 8 sample Constellation SKU + Flowcell                                                                                                 | DRAGEN Germline: Constellation        |
| The 2 sample Constellation SKU + Flowcell                                                                                                 | DRAGEN Germline: Constellation        |
| Ribo-Zero Plus Mcrbm Dpltn Kit, 96                                                                                                        | DRAGEN Microbiome Metatranscriptomics |
| ILMN VSP V2 Kit, Set A (96 spl)                                                                                                           | DRAGEN Microbial Enrichment Plus      |
| ILMN VSP V2 Kit, Set B (96 spl)                                                                                                           | DRAGEN Microbial Enrichment Plus      |
| ILMN VSP V2 Kit, Set C (96 spl)                                                                                                           | DRAGEN Microbial Enrichment Plus      |
| ILMN VSP V2 Kit, Set D (96 spl)                                                                                                           | DRAGEN Microbial Enrichment Plus      |
| IMAP with MiniSeq Core Cons Kit 48 smp                                                                                                    | DRAGEN Microbial Amplicon             |
| IMAP with MiSeq Core Cons Kit 48 smp                                                                                                      | DRAGEN Microbial Amplicon             |
| Illumina Microbial Amplicon Prep, 48 sample                                                                                               | DRAGEN Microbial Amplicon             |
| lLMN Microb Amp Prep Infl. A/B (48 Spl)                                                                                                   | DRAGEN Microbial Amplicon             |
| Resp Virus Enrich Kit A (96 idx,96 spl)                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| Resp Virus Enrich Kit B (96 idx,96 spl)                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| Resp Virus Enrich Kit C (96 idx,96 spl)                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| Resp Virus Enrich Kit D (96 idx,96 spl)                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| RPIP Enrich Kit (RUO)(96 idx,96 spl)-A                                                                                                    | DRAGEN Microbial Enrichment Plus      |
| RPIP Enrich Kit (RUO)(96 idx,96 spl)-B                                                                                                    | DRAGEN Microbial Enrichment Plus      |
| RPIP Enrich Kit (RUO)(96 idx,96 spl)-C                                                                                                    | DRAGEN Microbial Enrichment Plus      |
| RPIP Enrich Kit (RUO)(96 idx,96 spl)-D                                                                                                    | DRAGEN Microbial Enrichment Plus      |
| RPIP with MiSeq Core Cons Kit 96 smp                                                                                                      | DRAGEN Microbial Enrichment Plus      |
| RPIP with NS1K2K Core Cons Kit 96 smp                                                                                                     | DRAGEN Microbial Enrichment Plus      |
| Urinary Path. ID/AMR Enrichment Kt SetA                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| Urinary Path. ID/AMR Enrichment Kt SetC                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| VSP w ILMN RNA Prep w Enrich A, 96 Rxns                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| VSP w ILMN RNA Prep w Enrich B, 96 Rxns                                                                                                   | DRAGEN Microbial Enrichment Plus      |
| VSP with MiniSeq Core Cons Kit 96 smp                                                                                                     | DRAGEN Microbial Enrichment Plus      |
| VSP with MiSeq Core Cons Kit 96 smp                                                                                                       | DRAGEN Microbial Enrichment Plus      |
| ILMN VSP V2 Panel (96 spl)                                                                                                                | DRAGEN Microbial Enrichment Plus      |
| RESP. VIRUS OLIGO PANEL v2                                                                                                                | DRAGEN Microbial Enrichment Plus      |
| Viral Surveillance Panel, RUO 96 rxns                                                                                                     | DRAGEN Microbial Enrichment Plus      |
| IDT ILMN UMI DNA/RNA UDI A Lig 96Id96Spl                                                                                                  | DRAGEN ILMN cfDNA prep w/Enrichment   |
| IDT ILMN UMI DNA/RNA UDI B Lig 96Id96Spl                                                                                                  | DRAGEN ILMN cfDNA prep w/Enrichment   |
| TruSight Oncology Comprehensive (EU) Kit                                                                                                  | TruSight Oncology Comprehensive (LRM) |
| TruSight Oncology DNA Control                                                                                                             | TruSight Oncology Comprehensive (LRM) |
| TruSight Oncology RNA Control                                                                                                             | TruSight Oncology Comprehensive (LRM) |
| TruSight Onco 500 DNA Kt (48 Spl)                                                                                                         | DRAGEN TSO 500                        |
| TruSight Onco 500 DNA Kt, NSQ (48 Spl)                                                                                                    | DRAGEN TSO 500                        |
| TruSight Onco 500 DNA/RNA Kt,NSQ(24 Spl)                                                                                                  | DRAGEN TSO 500                        |
| TruSightOnco 500 DNA Auto Kt,NSQ(64 Spl)                                                                                                  | DRAGEN TSO 500                        |
| TS Onc 500 DNA HT (144 Spl)                                                                                                               | DRAGEN TSO 500                        |
| TS Onc 500 DNA/RNA HT (72 Spl)                                                                                                            | DRAGEN TSO 500                        |
| TSO 500 ctDNA Kit (48 smpl) + Velsera                                                                                                     | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 (24 spls)                                                                                                                | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 + S2 (24 spls)                                                                                                           | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 + S4 (24 spls)                                                                                                           | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 Auto (48 spls)                                                                                                           | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 Auto + S4 (48 spls)                                                                                                      | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 DNA HT (48 Spl)+Velsera                                                                                                           | DRAGEN TSO 500                        |
| TSO 500 DNA/RNA Auto Kt,NSQ(32 Spl)                                                                                                       | DRAGEN TSO 500                        |
| TSO 500 DNA/RNA HT (24 Spl)+Velsera                                                                                                       | DRAGEN TSO 500                        |
| TSO 500 DNA/RNA HT (72 Spl)+Velsera                                                                                                       | DRAGEN TSO 500                        |
| TSO500 DNA Kit (48 samples)+Velsera                                                                                                       | DRAGEN TSO 500                        |
| TSO500 DNA Kit NextSeq (48Spl)+Velsera                                                                                                    | DRAGEN TSO 500                        |
| TSO500 DNA/RNA Auto Kt NSQ(32) + Velsera                                                                                                  | DRAGEN TSO 500                        |
| TSO500 DNA/RNA HT Auto (32 spls)+Velsera                                                                                                  | DRAGEN TSO 500                        |
| TSO500 DNA/RNA HT Auto(72 spls)+Velsera                                                                                                   | DRAGEN TSO 500                        |
| TSO500 DNA/RNA Kit (24Spl)+Velsera                                                                                                        | DRAGEN TSO 500                        |
| TSO500 DNA/RNA Kit NextSeq (24)Velsera                                                                                                    | DRAGEN TSO 500                        |
| TSO500 DNA/RNA Kit NextSeq (24) plus Illumina Connected Insights Software                                                                 | DRAGEN TSO 500                        |
| TSO500 DNA/RNA HT Auto (32 spls) plus Illumina Connected Insights Software                                                                | DRAGEN TSO 500                        |
| TSO 500 DNA/RNA HT (24 Spl) plus Illumina Connected Insights Software                                                                     | DRAGEN TSO 500                        |
| TSO500 DNA/RNA Auto Kt NSQ(32) plus Illumina Connected Insights Software                                                                  | DRAGEN TSO 500                        |
| TruSight Oncology 500 ctDNA v2 (24 samples) plus Connected Insights Interpretation Report                                                 | DRAGEN TruSight Oncology 500 ctDNA    |
| TruSight Oncology 500 ctDNA v2 for Automation (48 samples) plus Connected Insights Interpretation Report                                  | DRAGEN TruSight Oncology 500 ctDNA    |
| TruSight Oncology 500 ctDNA v2 plus Connected Insights Interpretation Report, for use with NovaSeq 6000 S2 (24 samples)                   | DRAGEN TruSight Oncology 500 ctDNA    |
| TruSight Oncology 500 ctDNA v2 plus Connected Insights Interpretation Report, for use with NovaSeq 6000 S4 (24 samples)                   | DRAGEN TruSight Oncology 500 ctDNA    |
| TruSight Oncology 500 ctDNA v2 Automation Kit plus Connected Insights Interpretation Report, for use with NovaSeq 6000 S2 (48 samples)    | DRAGEN TruSight Oncology 500 ctDNA    |
| TruSight Oncology 500 ctDNA v2 Automation Kit plus Connected Insights Interpretation Report, for use with NovaSeq 6000 S4 (48 samples)    | DRAGEN TruSight Oncology 500 ctDNA    |
| oncoReveal Myeloid Panel plus Illumina Connected Insights software                                                                        | DRAGEN Amplicon Pipeline              |
| oncoReveal Multi-Cancer CNV + Fusion plus Illumina Connected Insights software                                                            | DRAGEN Amplicon Pipeline              |
| oncoReveal BRCA1 & BRCA2 + CNV plus Illumina Connected Insights software                                                                  | DRAGEN Amplicon Pipeline              |
| oncoReveal Essential MPN Panel plus Illumina Connected Insights software                                                                  | DRAGEN Amplicon Pipeline              |
| TruSight Oncology Comprehensive                                                                                                           | TruSight Oncology Comprehensive (LRM) |
| TSO 500 ctDNA v2 +S4 +Velsera (24 spls)                                                                                                   | DRAGEN TruSight Oncology 500 ctDNA    |
| oncoReveal BRCA1 & BRCA2 + CNV plus Correlation Engine software (24 samples)                                                              | DRAGEN Amplicon Pipeline              |
| oncoReveal Essential MPN Panel plus Correlation Engine software (24 samples)                                                              | DRAGEN Amplicon Pipeline              |
| oncoReveal Multi-Cancer CNV + Fusion plus Correlation Engine software (48 samples)                                                        | DRAGEN Amplicon Pipeline              |
| oncoReveal Myeloid Panel plus Correlation Engine software (24 samples)                                                                    | DRAGEN Amplicon Pipeline              |
| oncoReveal Essential LBx Panel plus Correlation Engine software (24 samples)                                                              | DRAGEN Amplicon Pipeline              |
| oncoReveal Core LBx Panel plus Correlation Engine software (24 samples)                                                                   | DRAGEN Amplicon Pipeline              |
| oncoReveal Fusion LBx Panel plus Correlation Engine software (24 samples)                                                                 | DRAGEN Amplicon Pipeline              |
| oncoReveal Multi-Cancer + Fusion Panel plus Correlation Engine software (24 samples)                                                      | DRAGEN Amplicon Pipeline              |
| oncoReveal Solid Tumor v2 Panel plus Correlation Engine software (24 samples)                                                             | DRAGEN Amplicon Pipeline              |
| oncoReveal Essential LBx Panel plus Illumina Connected Insights software (24 samples)                                                     | DRAGEN Amplicon Pipeline              |
| oncoReveal Core LBx Panel plus Illumina Connected Insights software (24 samples)                                                          | DRAGEN Amplicon Pipeline              |
| oncoReveal Fusion LBx Panel plus Illumina Connected Insights software (24 samples)                                                        | DRAGEN Amplicon Pipeline              |
| oncoReveal Multi-Cancer + Fusion Panel plus Illumina Connected Insights software (24 samples)                                             | DRAGEN Amplicon Pipeline              |
| oncoReveal Solid Tumor v2 Panel plus Illumina Connected Insights software (24 samples)                                                    | DRAGEN Amplicon Pipeline              |
| TruSight Oncology 500 DNA High-Throughput Kit (144 Samples), plus Velsera                                                                 | DRAGEN TSO 500                        |
| TruSight Oncology 500 DNA Automation Kit plus Velsera interpretation report (16 indexes, 64 samples)                                      | DRAGEN TSO 500                        |
| TruSight Oncology 500 DNA/RNA Automation Kit plus Velsera interpretation report (16 indexes, 32 Samples)                                  | DRAGEN TSO 500                        |
| TruSight Oncology 500 High-Throughput DNA for Automation (64 Samples), Plus Velsera                                                       | DRAGEN TSO 500                        |
| TruSight Oncology 500 High-Throughput DNA for Automation (144 Samples), Plus Velsera                                                      | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit (24 samples)                                                                                         | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit (48 samples)                                                                                             | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit (32 samples)                                                                              | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Automation Kit (64 samples)                                                                                  | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit (96 samples)                                                                              | DRAGEN TSO 500                        |
| TSO 500 v2 DNA/RNA Kit NSQ (24 spl)                                                                                                       | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit, For Use with NextSeq 550 (48 samples)                                                                   | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit, For Use with NextSeq 550 (32 samples)                                                    | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Automation Kit, For Use with NextSeq 550 (64 samples)                                                        | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit, For Use with NextSeq 1000/2000 P2 (24 samples)                                                      | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit, For Use with NextSeq 1000/2000 P2 (48 samples)                                                          | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit, For Use with NextSeq 1000/2000 P2 (32 samples)                                           | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Automation Kit, For Use with NextSeq 1000/2000 P2 (64 samples)                                               | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit plus Velsera interpretation report (24 samples)                                                      | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit plus Velsera interpretation report (48 samples)                                                          | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Velsera interpretation report (32 samples)                                           | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Automation Kit plus Velsera interpretation report (64 samples)                                               | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Velsera interpretation report (96 samples)                                           | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit plus Velsera interpretation report, For Use with NextSeq 550 (24 samples)                            | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit plus Velsera interpretation report, For Use with NextSeq 550 (48 samples)                                | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Velsera interpretation report, For Use with NextSeq 550 (32 samples)                 | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Automation Kit plus Velsera interpretation report, For Use with NextSeq 550 (64 samples)                     | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit plus Velsera interpretation report, For Use with NextSeq 1000/2000 P2 (24 samples)                   | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit plus Velsera interpretation report, For Use with NextSeq 1000/2000 P2 (48 samples)                       | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Velsera interpretation report, For Use with NextSeq 1000/2000 P2 (32 samples)        | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit plus Illumina Connected Insights Software (24 samples)                                               | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit plus Illumina Connected Insights Software (48 samples)                                                   | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Illumina Connected Insights Software (32 samples)                                    | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Automation Kit plus Illumina Connected Insights Software (64 samples)                                        | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Illumina Connected Insights Software (96 samples)                                    | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit plus Illumina Connected Insights Software, For Use with NextSeq 550 (24 samples)                     | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit plus Illumina Connected Insights Software, For Use with NextSeq 550 (48 samples)                         | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Illumina Connected Insights Software, For Use with NextSeq 550 (32 samples)          | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Automation Kit plus Illumina Connected Insights Software, For Use with NextSeq 550 (64 samples)              | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Kit plus Illumina Connected Insights Software, For Use with NextSeq 1000/2000 P2 (24 samples)            | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA Kit plus Illumina Connected Insights Software, For Use with NextSeq 1000/2000 P2 (48 samples)                | DRAGEN TSO 500                        |
| TruSight Oncology 500 v2 DNA/RNA Automation Kit plus Illumina Connected Insights Software, For Use with NextSeq 1000/2000 P2 (32 samples) | DRAGEN TSO 500                        |
| TSO 500 ctDNA v2 + P4 (24 spls)                                                                                                           | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 Auto + S2 (48 spls)                                                                                                      | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 + Velsera (24 spls)                                                                                                      | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 Auto + Velsera(48 spls)                                                                                                  | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 +P4 +Velsera (24 spls)                                                                                                   | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 +S2 +Velsera (24 spls)                                                                                                   | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 Auto+S2+Velsera(48spls)                                                                                                  | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 ctDNA v2 Auto+S4+Velsera(48spls)                                                                                                  | DRAGEN TruSight Oncology 500 ctDNA    |
| TruSight Oncology 500 ctDNA v2 plus Illumina Connected Insight Software (24 spls)                                                         | DRAGEN TruSight Oncology 500 ctDNA    |
| TSO 500 v2 DNA Auto Kit ICI P2 (64 spl)                                                                                                   | DRAGEN TSO 500                        |
| TSO500 DNA Auto Kt NSQ(64Spl)+Velsera                                                                                                     | DRAGEN TSO 500                        |
| oncoReveal Myeloid Panel                                                                                                                  | DRAGEN Amplicon Pipeline              |
| oncoReveal Multi-Cancer CNV + Fusion                                                                                                      | DRAGEN Amplicon Pipeline              |
| oncoReveal BRCA1 & BRCA2 + CNV                                                                                                            | DRAGEN Amplicon Pipeline              |
| oncoReveal Essential MPN Panel                                                                                                            | DRAGEN Amplicon Pipeline              |
| oncoReveal Essential LBx Panel                                                                                                            | DRAGEN Amplicon Pipeline              |
| oncoReveal Core LBx Panel                                                                                                                 | DRAGEN Amplicon Pipeline              |
| oncoReveal Fusion LBx Panel                                                                                                               | DRAGEN Amplicon Pipeline              |
| oncoReveal Multi-Cancer Fusion Panel                                                                                                      | DRAGEN Amplicon Pipeline              |
| oncoReveal Solid Tumor v2 Panel                                                                                                           | DRAGEN Amplicon Pipeline              |


# Illumina® DRAGEN™ Secondary Analysis

Illumina DRAGEN (Dynamic Read Analysis for GENomics) secondary analysis was developed to address important challenges associated with analyzing NGS (Next Generation Sequencing) data for a range of applications, including genome, exome, transcriptome, and methylome studies. DRAGEN secondary analysis processes NGS data and enables tertiary analysis to drive insights. The available tools make up a highly accurate, comprehensive, and efficient solution that enables labs of all sizes and disciplines to do more with their genomic data.

**Product highlights**

**Accurate results:**

* Pangenome reference genome and machine learning drive unprecedented accuracy
* 99.89% accuracy score with the Precision FDA Truth Challenge V2 benchmark data (*2,3*)

**Comprehensive platform:**

* Analyze NGS data from whole genomes, exomes, methylomes, and transcriptomes
* Available on platform of choice and scalable based on needs

**Efficient analysis:**

* Process a 34x genome in \~ 30 minutes, with all supported callers with DRAGEN server v4 (*1*)
* Reduce FASTQ file sizes up to 5x with DRAGEN ORA Compression

\ <br>

*References:*

1. Illumina data on file, 2022.
2. Illumina DRAGEN Secondary Analysis is the first single platform to achieve 99.89% accuracy based on [PrecisionFDA v2 Truth Challenge Benchmark Data](https://precision.fda.gov/challenges/10). Details here [DRAGEN sets new standard for data accuracy in PrecisionFDA benchmark data](https://www.illumina.com/science/genomics-research/articles/dragen-shines-again-precisionfda-truth-challenge-v2.html). Accessed March 22, 2023
3. PrecisionFDA Truth Challenge V2: Calling Variants from Short and Long Reads in Difficult-to-Map Regions. [precision.fda.gov/challenges/10](https://precision.fda.gov/challenges/10). Accessed November 3, 2020.


# DRAGEN Applications

## Applications

DRAGEN analysis offers a large selection of application pipelines.

| Pipeline                                             | Description                                                                                                                                                                                                                                                                                                   | Variant Types Detected             | Metrics Provided                                                                     |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------------------------------------------ |
| DRAGEN Demultiplexing                                | Rapid demultiplexing of NGS analysis                                                                                                                                                                                                                                                                          | N/A                                | N/A                                                                                  |
| DRAGEN ORA Compression                               | DRAGEN ORA compression is optimized for high compression ratios of FASTQ files, as well as rapid compression and decompression, all while preserving data integrity.                                                                                                                                          | N/A                                | Compression Ratio Run Time                                                           |
| DRAGEN Map + Align                                   | The DRAGEN Map + Align can be run as a standalone or as part of DRAGEN’s suite of pipelines                                                                                                                                                                                                                   | N/A                                | Mapping metrics Duration Metrics Coverage Metrics                                    |
| DRAGEN Germline                                      | The DRAGEN Germline Pipeline provides end-to-end NGS analysis, including advanced error model calibration for increased accuracy, and repeat expansion detection and genotyping through Illumina Expansion Hunter.                                                                                            | SNV/Indel CNV SV Repeat Expansions | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Somatic                                       | The DRAGEN Somatic Pipeline includes tumor-only and tumor–normal modes, designed for detecting somatic variants in tumor samples. Both modes make no ploidy assumptions, enabling detection of low-frequency alleles.                                                                                         | SNV/Indel CNV SV TMB MSI HLA       | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Enrichment                                    | The DRAGEN Enrichment Pipeline combines DRAGEN’s germline and somatic callers into a pipeline designed specifically for analyzing enrichment samples. Includes a full suite of enrichment metrics and reporting.                                                                                              | SNV/Indel CNV SV                   | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN RNA                                           | The DRAGEN RNA Pipeline performs transcriptome analysis starting with splice junction discovery and alignment, followed by rapid alignment and splice junction mapping and quantification. For differential expression, Illumina recommends the DRAGEN Differential Expression app on BaseSpace Sequence Hub. | Gene fusion SNV/Indel              | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Single Cell RNA                               | The DRAGEN Single Cell RNA pipeline performs demultiplexing, cell-barcode and UMI error correction, sequence alignment, and quantification of gene expression.                                                                                                                                                | N/A                                | Mapping Metrics Duration Metrics Coverage Metrics Callability Report Cell Metrics    |
| DRAGEN Joint Genotyping                              | The DRAGEN Joint Genotyping/Population Pipeline calls variants jointly across multiple genomes and scales to large cohorts of samples at expedited speeds with uncompromising accuracy.                                                                                                                       | SNV/Indel CNV SV Repeat Expansions | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Methylation                                   | The DRAGEN Methylation Pipeline performs alignment, methyl calling, and calculates alignment and methylation metrics.                                                                                                                                                                                         | N/A                                | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Reference Builder                             | Accepts FASTA files, and builds the proprietary reference used by the DRAGEN apps.                                                                                                                                                                                                                            | N/A                                | N/A                                                                                  |
| DRAGEN TruSight Oncology 500 ctDNA Analysis Software | Secondary analysis support for Illumina’s TruSight Oncology 500 ctDNA. Available on the local DRAGEN Server version 3 and later.                                                                                                                                                                              | SNV/Indel CNV DNA fusions MSI TMB  | Mapping metrics Duration Metrics Coverage Metrics Variant Metrics Callability Report |
| DRAGEN Imputation                                    | The DRAGEN Imputation pipeline is an end to end user friendly tool that enables scalable low pass whole genome sequencing analysis                                                                                                                                                                            | N/A                                | Impute ≤100 samples simultaneously 1.7x faster compared to original GLIMPSE code     |

## Analysis Uses

DRAGEN analysis can be used in numerous fields in the biological sciences.

| Analysis                   | Description                                                                                                                    |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Genetic Diseases           | Reduce time required for genomic analysis, with high accuracy and comprehensiveness                                            |
| Oncology                   | Analyze tumor-only and tumor/normal samples with accuracy, comprehensiveness, and efficiency                                   |
| Cell and Molecular Biology | Advance understanding of cellular mechanisms with rapid analysis pipelines for bulk and single cell samples                    |
| Population Genomics        | Accurately and efficiently analyze sequenced genomes at scale. Accelerate re-analysis as computational tools improve over time |
| Infectious Disease         | Detect and characterize infectious diseases with a comprehensive solution                                                      |
| Agrigenomics               | Efficiently analyze animals and plants of varying genomic complexities with custom reference                                   |


# Deployment Options

DRAGEN analysis is available on multiple platforms.

| Platform                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| DRAGEN on-premises server                            | <p>DRAGEN on-premises server offers highly accurate secondary analysis in a fraction of time compared with a traditional CPU-based system.<br>- Analyze and store data locally<br>- Supports varying levels of command line interface<br>- Replace up to 30 traditional compute instances<br>- Fully process a 34× whole human genome in \~30 minutes. <em>(1)</em><br>- One unit supports two NovaSeq 6000 Systems running at full capacity</p>       |
| DRAGEN analysis on Illumina Connected Analytics      | Couples the accuracy and speed of the DRAGEN with the ability to customize analysis pipeline to operationalize informatics on a secure platform.                                                                                                                                                                                                                                                                                                       |
| DRAGEN on BaseSpace Sequence Hub (BSSH)              | Push button analysis capability in an intuitive, easy-to-use interface with compliance, and storage features of BaseSpace Sequence Hub and Amazon Web Services (AWS).                                                                                                                                                                                                                                                                                  |
| DRAGEN onboard NovaSeq X Series                      | <p>- Flexibly runs multiple secondary analysis pipelines in parallel.<br>- Performs up to four simultaneous applications per flow cell in a single run.<br>- Brings up to 5x lossless data compression, and analysis with supported applications<br>- Provides savings on analysis, which over five years can exceed the price of the sequencer</p>                                                                                                    |
| DRAGEN onboard NextSeq 1000 and NextSeq 2000 Systems | <p>- Provides access to select DRAGEN analysis informatics pipelines<br>- Enables users to generate results in as little as two hours<br>- Uses intuitive pipeline algorithms to reduce reliance on external informatics experts</p>                                                                                                                                                                                                                   |
| DRAGEN onboard MiSeq i100 Series                     | <p>Intuitive, ultra-rapid analysis including DRAGEN BCL convert, DRAGEN Library QC, DRAGEN small WGS and DRAGEN Microbial Enrichment Plus.<br>- Rapid results with comprehensive secondary analysis generated in two hours or less <em>(2)</em><br>- Highly efficient workflow with a single user touchpoint to VCF and/or html report and no intermediate file transfers<br>- Exceptionally easy with an intuitive interface for non-expert users</p> |
| DRAGEN on AWS, Azure                                 | DRAGEN supports the FPGA enabled instance types of AWS, Azure. Rpm installers and the Kernel driver can be installed on images managed by the user, and DRAGEN can be run by purchasing a license.                                                                                                                                                                                                                                                     |
| DRAGEN on AWS and Azure Marketplace                  | Pre-configured Amazon Machine Images (AMI) and Azure Virtual Machines with DRAGEN installed can be accessed from the respective marketplace offerings in a Pay-As-You-Use model.                                                                                                                                                                                                                                                                       |
| DRAGEN on GCP                                        | DRAGEN is made available on the Google Cloud Platform. Pre-configured instances with DRAGEN installed can be accessed through the GCP application interface. Limited availability. Please reach out to your Illumina representative for access.                                                                                                                                                                                                        |

> (1) HG002 from PrecisionFDA truth challenge V2 run with DRAGEN analysis v4.0 on DRAGEN server v4, all callers

> (2) When run according to sample recommendations


# DRAGEN v4.5


# Getting Started

DRAGEN provides tests you can run to make sure that your DRAGEN system is properly installed and configured. Before running the tests, make sure that the DRAGEN server has adequate power and cooling, and is connected to a network that is fast enough to move your data to and from the machine with adequate performance.

Please refer to the [Server Site Prep & Installation Guide](https://support.illumina.com/downloads/illumina-dragen-server-site-prep-guide.html) when installing a new system.

## On-premises Installation

The software can be installed on an on-premises server by executing the .run installer for the desired version. Installers are made available for all releases at the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/downloads.html).

Installation procedure:

* Download the desired installer from the support website and unzip the package
* The archive integrity can be checked using: `./<dragen .run file> --check`
* Install the appropriate release based on your Linux OS with the command: `sudo sh <dragen .run file>`

The .run file includes a script that administers un-installation of an existing software, integrity checking of the package and files, installation of the new DRAGEN software version. The DRAGEN software is installed in part by use of the Linux RPM Package Manager (rpm). Several rpm packages comprise the installation of a single DRAGEN software version. The RPM packages also configure the system for dragen, like raised user `ulimits`, and the .run script starts services needed for functionality, such as the Licensing daemon `dragen_licd`, and the hugepages daemon, `dragend_hp`.

> NOTE: Root privileges are required for the installation.

### Single Version Installation

Up to DRAGEN Software v4.2, only one version of the DRAGEN software can be installed at a time. Executing the .run file will remove any existing installed version and (re)install the new version.

After installation, the application and associated files are available at `/opt/edico`.

The single version installer will add `/opt/edico` to the Linux $PATH, so that the user can just call `dragen` without specifying the full path.

### Multi-Version Installation

Starting with DRAGEN Software v4.3 and later, multiple compatible versions of the DRAGEN software can be installed at a time. Executing the .run file will add the new version to the system.

After installation, the application files are available at `/opt/dragen/{version}/bin` and FPGA files are located at `/opt/bitstream/{bitstream version}`.

The multi-version installer will NOT add `/opt/dragen/{version}/bin` to the Linux $PATH, since multiple versions can be present at a given time. User should manage the desired paths to the specific version they want to run. When this guide provides command line examples, it will assume that the Linux $PATH is set to correct dragen version, and we will just refer to `dragen <options>`

Notes on multi-version installation:

* Installers released for DRAGEN v4.2 and earlier are single version packages
* Single version packages and multi-version packages can not be mixed
  * Installation of a prior single version package will remove all the multi-version packages
  * Installation of a multi-version package will remove any installed single version package
* After installing a multi-version package, see a list of installed versions at any time by running `/usr/bin/dragen_versions`
* To remove any multi-version package, call `yum remove` on its Path
* Adding PATH="/opt/dragen/{version}/bin:$PATH" to the last line of .bashrc file avoids the need to set the path upon each server login

Example:

```
$ dragen_versions
The output format of this command may change. Use --json for machine readable output.

Dragen Version           Size (MB)  Install Date         Path
4.3.2                    1378.03    2024-03-10 18:26:17  /opt/dragen/4.3.2
4.4.3                    1381.41    2024-03-18 20:56:39  /opt/dragen/4.4.3
4.3.5                    1379.25    2024-03-11 15:20:24  /opt/dragen/4.3.5

Bitstream Version        Size (MB)  Install Date         Path
07.031.732 (0x18101306)  598.95     2024-03-10 18:26:03  /opt/bitstream/07.031.732
07.031.745 (0x18101306)  598.95     2024-03-18 20:56:18  /opt/bitstream/07.031.745
 
To remove a dragen version, call `yum remove` on its Path.
```

### Location of `dragen` and resource files

| DRAGEN Version  | on-premises server      | cloud instance |
| --------------- | ----------------------- | -------------- |
| 4.3 and later   | `/opt/dragen/{version}` | `/opt/edico/`  |
| 4.2 and earlier | `/opt/edico/`           | `/opt/edico/`  |

Throughout this guide we will refer to `<INSTALL_PATH>` which will be either of the locations above

## Licensing

DRAGEN requires license(s) for most functionality, please refer to the [Licensing Reference Section](/dragen-v4.5/reference/licensing) for guidance on how to install and/or review your current licenses.

## Running the System Check

After turning on the server, you can make sure that your DRAGEN server is functioning properly by running `<INSTALL_PATH>/self_test/self_test.sh`, which does the following:

* Automatically indexes chromosome M from the hg19 reference genome
* Loads the reference genome and index
* Maps and aligns a set of reads
* Saves the aligned reads in a BAM file
* Asserts that the alignments exactly match the expected results

Each server ships with the test input FASTQ data for this script, which is located in `<INSTALL_PATH>/self_test`. The system check takes approximately 25--30 minutes.

The following example shows how to run the script and shows the output from a successful test.

```
$ /opt/dragen/4.3.4/self_test/self_test.sh
#############################################################
Logging to /var/log/dragen/self_test.1714627157_160164.0.details.log
Using dragen executables in /opt/dragen/4.3.4/bin
Using board(s): 0 
#############################################################
Running tests for board 0 (u200)
Using scratch directory /tmp/self_test.4BO0pfPST9/0
-------------------------------------------------------------
Board 0 test 1, FPGA MEMORY TEST
Loading DIAG bitstream
Running fpga memory test, this will take ~13 minutes
Board 0 test 1, FPGA MEMORY TEST: PASS
-------------------------------------------------------------
Board 0 test 2, BAR REGISTER ACCESS
Board 0 test 2, BAR REGISTER ACCESS: PASS
-------------------------------------------------------------
Board 0 test 3, FPGA TEMP REG ACCESS
FPGA Temperature: 27C  (Max Temp: 36C, Min Temp: 22C)
Board 0 test 3, FPGA TEMP REG ACCESS: PASS
-------------------------------------------------------------
Board 0 test 4, BOARD SERIAL # REG ACCESS
Serial Number: 2130069BM05V
Board 0 test 4, BOARD SERIAL # REG ACCESS: PASS
-------------------------------------------------------------
Board 0 test 5, DRAGEN GENOME LICENSE
Board 0 test 5, DRAGEN GENOME LICENSE: PASS
-------------------------------------------------------------
Board 0 test 6, CPLD DATE TEST
cpld date is n/a
Board 0 test 6, CPLD DATE TEST: PASS
-------------------------------------------------------------
Board 0 test 7, ENCRYPTION KEY EXISTENCE TEST
Board 0 test 7, ENCRYPTION KEY EXISTENCE TEST: PASS
-------------------------------------------------------------
Board 0 test 8, PARTIAL RECONFIGURATION
DNA-MAPPER: ok
RNA-MAPPER: ok
HMM: ok
ZIP: ok
UNZIP: ok
DIAG: ok
Board 0 test 8, PARTIAL RECONFIGURATION: PASS
-------------------------------------------------------------
Board 0 test 9, HASH TABLE GENERATION
Board 0 test 9, HASH TABLE GENERATION: PASS
-------------------------------------------------------------
Board 0 test 10, MAP AND ALIGNER
running mapper aligner: ok
unmapped input records percentages: ok
md5sum check dbam sorted: pass
Board 0 test 10, MAP AND ALIGNER: PASS
-------------------------------------------------------------
Board 0 test 11, VARIANT CALLER E2E
running variant caller: ok
md5sum check dbam sorted: ok
md5sum check VCF: ok
Board 0 test 11, VARIANT CALLER E2E: PASS
#############################################################
SELF TEST COMPLETED
SELF TEST RESULT : PASS
#############################################################
Log file at /var/log/dragen/self_test.1714627157_160164.0.details.log

```

If the output BAM file does not match expected results, then the last line of the above text is as follows:

`SELF TEST RESULT : FAIL`

If you experience a FAIL result after running this test script immediately after turning on your DRAGEN server, contact Illumina Technical Support.

## Running Your Own Test

When you are satisfied that your DRAGEN system is performing as expected, you are ready to run some of your own data through the machine, as follows:

* Load the reference table for the reference genome
* Determine location of input and output files
* Process input data

### Loading the Reference Genome

Before a reference genome can be used with DRAGEN, it must be converted from FASTA format into a custom binary format for use with the DRAGEN hardware. For more information, see [Prepare a Reference Genome](/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support/prepare-a-reference-genome).

The reference hash table specified on the command line is automatically loaded onto the board the first time you process data with a pipeline. You can manually load the hash table for your reference genome by using the following command:

`dragen -r <reference_hash-table_directory>`

Make sure that the reference hash table directory is on the fast file IO drive.

The default location for the hash table for hg19 is as follows.

`/staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149`

The command to load reference genome hg19 from the default location is as follows.

`dragen -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149`

This command loads the binary reference genome into memory on the DRAGEN board, where it is used for processing any number of input data sets. You do not need to reload the reference genome unless you restart the system or need to switch to a different reference genome. It can take up to a minute to load a reference genome.

DRAGEN checks whether the specified reference genome is already resident on the board. If it is, then the upload of the reference genome is automatically skipped. You can force reloading of the same reference genome using the `force-load-reference (-l)` command line option.

The command to load the reference genome prints the software and hardware versions to standard output. For example:

```
DRAGEN Host Software Version 01.001.035.01.00.30.6682 and

Bio-IT Processor Version 0x1001036
```

After the reference genome has been loaded, the following message is printed to standard output:

```
DRAGEN finished normally
```

### Determine Input and Output File Locations

The DRAGEN Pipeline is very fast, which requires careful planning for the locations of the input and output files. If the input or output files are on a slow file system, then the overall performance of the system is limited by the throughput of that file system. It is recommended that inputs and outputs are streamed directly from/to a mounted external storage system.

The DRAGEN system is preconfigured with at least one fast file system consisting of a set of fast SSD disks grouped with RAID-0 for performance. This file system is mounted at `/staging`. This name was chosen to emphasize the fact that this area was built to be large and fast, but is not redundant. Failure of any of the file system's constituent disks leads to the loss of all data stored there.

During processing, DRAGEN generates and reads back temporary files. With DRAGEN, it is highly recommended to always direct temporary files to the fast SSD (or `/staging`) by using the `--intermediate-results-dir` option. If the `--intermediate-results-dir` option is not provided, temporary files are written to the `--output-directory`. DRAGEN recommends streaming inputs and outputs using an mounted external storage system.

### Process Your Input Data

To analyze FASTQ data, use the dragen command. For example, the following command can be used to analyze a single-ended FASTQ file:

```
dragen \
-r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
-1 /staging/test/data/SRA056922.fastq \
--output-directory /staging/test/output \
--output-file-prefix SRA056922_dragen \
--RGID DRAGEN_RGID \
--RGSM DRAGEN_RGSM
```

For detailed information on the command line options, see [DRAGEN Host Software](/dragen-v4.5/product-guides/dragen-v4.5/dragen-host-software).

For recommended command lines in typical use cases, see [DRAGEN Recipes](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes).


# DRAGEN Host Software

You use the DRAGEN host software program *dragen* to build and load reference genomes, and then to analyze sequencing data by decompressing the data, mapping, aligning, sorting, duplicate marking with optional removal, and variant calling.

Invoke the software using the *dragen* command. The command line options are described in the following sections.

Command line options can also be set in a configuration file. For more information on configuration files, see [Configuration Files](#configuration-files) . If an option is set in the configuration file and is also specified on the command-line, the command line option overrides the configuration file.

## Command-line Options

The following are examples of frequently used command lines:

* Build Reference/Hash Table

  ```
  dragen --build-hash-table true --ht-reference <REF_FASTA> \
  --output-directory <REF_DIRECTORY>  [options]
  ```
* Run Map/Align and Variant Caller (\*.fastq to \*.vcf)

  ```
  dragen -r <REF_DIRECTORY> --output-directory <OUT_DIRECTORY> \
  --output-file-prefix <FILE_PREFIX> [options] -1 <FASTQ1> \
  [-2 <FASTQ2>] --RGID <RG0> --RGSM <SM0> --enable-variant-caller true
  ```
* Run Map/Align (\*.fastq to \*.bam)

  ```
  dragen -r <REF_DIRECTORY> --output-directory <OUT_DIRECTORY> \
  --output-file-prefix <FILE_PREFIX> [options] \
  -1 <FASTQ1> [-2 <FASTQ2>]  \
  --RGID <RG0> --RGSM
  ```
* Run Variant Caller Only (\*.bam to \*.vcf)

  ```
  dragen -r <REF_DIRECTORY> --output-directory <OUT_DIRECTORY> \
  --output-file-prefix <FILE_PREFIX> [options] -b <BAM> \
  --enable-map-align false \
  --enable-variant-caller true
  ```
* Re-map and Run Variant Caller (\*.bam to \*.vcf)

  ```
  dragen -r <REF_DIRECTORY> --output-directory <OUT_DIRECTORY> \
  --output-file-prefix <FILE_PREFIX> [options] -b <BAM> \
  --enable-map-align true \
  --enable-variant-caller true
  ```
* Run BCL Converter (BCL to \*.fastq)

  ```
  dragen --bcl-conversion-only true --bcl-input-directory <BCL_DIRECTORY> \
  --output-directory <OUT_DIRECTORY>
  ```
* Run RNA Map/Align (\*.fastq to \*.bam)

  ```
  dragen -r <REF_DIRECTORY> --output-directory <OUT_DIRECTORY> \
  --output-file-prefix <FILE_PREFIX> [options] -1 <FASTQ1> \
  [-2 <FASTQ2>] --enable-rna true
  ```

For recommended command lines in typical use cases, see [DRAGEN Recipes](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes).

### Reference Genome Options

Before you can use the DRAGEN system for aligning reads, you must load a reference genome and its associated hash tables onto the PCIe card. For information on preprocessing a reference genome's FASTA files into the native DRAGEN binary reference and hash table formats, see [Prepare a Reference Genome](/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support/prepare-a-reference-genome). You must also specify the directory containing the preprocessed binary reference and hash tables with the `-r [or --ref-dir]` option. This argument is always required.

Use the following command to load the reference genome and hash tables to DRAGEN card memory separately from processing reads.

`dragen -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149`

Use the `-l (--force-load-reference)` option to force the reference genome to load even if it is already loaded.

`dragen -l -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149`

The time needed to load the reference genome depends on the size of the reference, but for typical recommended settings, it takes approximately 30--60 seconds.

### Operating Modes

DRAGEN has two primary modes of operation, as follows:

* Mapper/aligner
* Variant caller

DRAGEN is capable of performing each mode independently or as an end-to-end solution. DRAGEN also allows you to enable and disable decompression, sorting, duplicate marking, and compression along the DRAGEN pipeline.

* **Full pipeline mode** To execute full pipeline mode, set `--enable-variant-caller` to `true` and provide input as unmapped reads in \*.fastq, \*.bam, or \*.cram formats. DRAGEN performs decompression, mapping, aligning, sorting, and optional duplicate marking and feeds directly into the variant caller to produce a VCF file. In this mode, DRAGEN uses parallel stages throughout the pipeline to drastically reduce the overall run time.
* **Map/align mode** Map/align mode is enabled by default. Input is unmapped reads in \*.fastq, \*.bam, or \*.cram format. DRAGEN produces an aligned and sorted BAM or CRAM file. To mark duplicate reads at the same time, set `‑-enable‑duplicate‑marking` to `true`.
* **Variant caller mode** To execute variant caller mode, set the `--enable-variant-caller` option to true, and set `--enable-map-align` option to false. The input must be a mapped and aligned BAM/CRAM file. DRAGEN produces a VCF file. DRAGEN will force-enable re-sorting of the BAM, because a number of read statistics and estimates are required for the Variant Caller to operate effectively. Setting `--enable-sort` to `false` will be overridden. BAM files cannot be duplicate marked in the DRAGEN pipeline prior to variant calling if they have not already been marked. Use the end-to-end mode of operation to take advantage of the mark-duplicates feature.
* **RNA-Seq data** To enable processing of RNA-Seq--based data, set `--enable-rna` to `true`. DRAGEN uses the RNA spliced aligner during the mapper/aligner stage. DRAGEN dynamically switches between the required modes of operation..
* **Bisulfite MethylSeq data** To enable processing of Bisulfite MethylSeq data, set the `--enable-methylation-calling` option to true. DRAGEN automates the processing of data for Lister (directional) and Cokus (nondirectional) protocols to generate a single BAM with bismark-compatible tags. Alternatively, you can run DRAGEN in a mode that produces a separate BAM file for each combination of the C->T and G->A converted reads and references. To enable this mode of processing, you need to build a set of reference hash tables with `--ht-methylated` enabled, and run DRAGEN with the appropriate `‑‑methylation-protocol` setting.

### Output Options

The following command line options for output are mandatory:

* `--output-directory <out_dir>`—Specifies the output directory for generated files.
* `--output-file-prefix <out_prefix>`-Specifies the output file prefix. DRAGEN appends the appropriate file extension onto this prefix for each generated file.
* `-r [--ref-dir ]`—Specifies the reference hash table.

The following examples do not include these mandatory options.

For mapping and aligning, the output is sorted and compressed into BAM format by default before saving to disk. The user can control the output format from the map/align stage with the `--output-format <SAM|BAM|CRAM>` option. If the output file exists, the software issues a warning and exits. To force overwrite if the output file already exists, use the `-f [ --force ]` option.

For example, the following commands output to a compressed BAM file, and then forces overwrite:

`dragen ... -f`

`dragen ... -f --output-format bam`

To generate a BAI-format BAM index file (\*.bai), set `--enable-bam-indexing` to true.

The following example outputs to a SAM file, and then forces overwrite:

`dragen ... -f --output-format sam`

The following example outputs to a CRAM file, and then forces overwrite:

`dragen ... -f --output-format cram`

DRAGEN only outputs lossless CRAM files. All QNAMEs and BAM tags are preserved in the CRAM.

#### Alignment tags

DRAGEN can generate mismatch difference (MD) tags, as described in the BAM standard. The feature is turned off by default because there is a small performance cost to generate these strings. To generate MD tags, set `--generate-md-tags` to true.

DRAGEN can also annotate additional information about alignments in a ZS:Z tag. The following are valid tag values:

| Tag        | Tag meaning                                                                                               |
| ---------- | --------------------------------------------------------------------------------------------------------- |
| `ZS:Z:R`   | Multiple alignments with similar score were found.                                                        |
| `ZS:Z:NM`  | No alignment was found.                                                                                   |
| `ZS:Z:QL`  | An alignment was found but it was below the quality threshold.                                            |
| `ZS:Z:NRD` | Alignment is to an auto-added decoy contig (not present in input FASTA).                                  |
| `ZS:Z:PAI` | Alignment is to an insertion encoded in a population based alternate contig (not present in input FASTA). |

By default, DRAGEN writes a ZS:Z:PAI tag in the output BAM for alignments that map completely inside insertions encoded in population based alternate contigs. To write ZS:Z alignment status tags for all other types described above, set `--generate-zs-tags` to true (false by default). These tags are only generated in the primary alignment and when a read has suboptimal alignments qualifying for secondary output (even if none were output because `--Aligner.sec-aligns` was set to 0).

To generate SA:Z tags, set `--generate-sa-tags` to true (the default). These tags provide alignment information (position, cigar, orientation) of groups of supplementary alignments, which are useful in structural variant calling.

To generate pair score in a ps:i tag, set `--generate-ps-tags` to true (false by default for DNA, true for RNA). The pair score is used in DRAGEN for computing MAPQ and can be used to check how well alignment candidate pairs score against each other.

DRAGEN can also output mate alignment tags. To generate the mate cigar (in the MC:Z tag), set `--generate-mc-tags` to true (this is the default). To generate the mate mapping quality (in the MQ:i) tag, set `--generate-mq-tags` to true (this is the default). To generate mate sequence (in the R2:Z tag) and mate base qualities (in the Q2:Z tag), set `--generate-r2-tags` to true (default is false) and set `--generate-q2-tags` to true (default is false) respectively. Please note that when enabled, R2:Z and Q2:Z tags are emitted only for improperly paired read alignments with fragment length atleast 1000 bp. Also, our methylation pipelines currently do not support the output of mate alignment tags.

DRAGEN also outputs a graph alignment tag ga:Z `--generate-ga-tags` (true by default for DNA, false for RNA) when applicable. This tag is used to describe the best alt contig alignment which improved the score of a primary-contig alignment at its liftover position. It can also be used to describe read alignments to alt contigs for which there is no liftover and the primary alignment is unmapped. For example, cases when the read maps best to an alt contig describing a novel long-insertion that is not present in the reference. In addition, read alignments that have been marked as unmapped because they map to auto-detected decoy contigs not present in the original user-provided FASTA also have their alignments described in the ga tag.

The ga tag uses the same format as the SA tag used to describe supplementary alignments.

#### CRAM Output

When CRAM is selected as output, DRAGEN generates a CRAM file with the following features:

* CRAM format V3.0 is produced by default, V3.1 can be enabled by using the option --cram-version 3.1
* The CRAM is lossless. Lossy compression is never employed and not optional
* Quality score compression is lossless. Read names are preserved
* Only the GZIP compression algorithm is employed for maximum compatibility. bgzip, lzma not employed. rANS is used for quality scores
* All input BAM tags are preserved
* The reference used to compress the CRAM file, is the DRAGEN Hash Table provided during the map/align run. When decompressing the CRAM with a FASTA file and 3rd party tools, the FASTA that was used to generate the Hash Table must be used. Refer to the Reference Files table in [Illumina DRAGEN Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) to find the correct FASTA file.
* A CRAM index is produced in .crai format
* CRAM output is only possible when sort is enabled. CRAM alignments will always be positionally sorted

The following list of default settings are used for the CRAM output

| CRAM option          | Value                 | Description                                     |
| -------------------- | --------------------- | ----------------------------------------------- |
| SEQS\_PER\_SLICE     | 2000                  | Max sequences per slice                         |
| BASES\_PER\_SLICE    | SEQS\_PER\_SLICE\*500 | Max bases per slice                             |
| SLICE\_PER\_CNT      | 1                     | Max slices per container                        |
| embed\_ref           | 0                     | Do not embed reference sequence                 |
| noref                | 0                     | Do not use non-referenced based encoding        |
| multiseq             | -1                    | Do not use multiple references per slice        |
| unsorted             | 0                     | Do not use unsorted mode                        |
| use\_bz2             | 0                     | Do not compress using bzip2                     |
| use\_lzma            | 0                     | Do not compress using lmza                      |
| use\_rans            | 1                     | Use rANS for quality score compression          |
| binning              | NONE                  | Qual score binning not used                     |
| preserve\_aux\_order | 1                     | Preserve all aux tags and order (incl RG,NM,MD) |
| preserve\_aux\_size  | 0                     | Aux tag sizes not preserved ('i', 's', 'c')     |
| lossy\_read\_names   | 0                     | Preserve read names                             |
| lossy                | 0                     | Do not enable Illumina 8 quality-binning system |
| ignore\_md5          | 0                     | Enable all checking of checksums                |
| decode\_md           | 0                     | Do not (re)generate MD and NM tags              |
| cram\_version        | 3.0                   | Default is CRAM v3.0.                           |

### Input Options

DRAGEN can process reads in FASTQ format or BAM/CRAM format. DRAGEN supports the following compression options for FASTQ input files.

* Uncompressed
* gzip or bgzip compression
* ORA compression. To use ORA compression, you must provide an ORA reference and reference directory. See ORA Compression and Decompression.

If your input FASTQ files are gzipped, DRAGEN automatically decompresses the files using hardware-accelerated decompression, and then streams the reads into the mapper. If your files end in \*.ora, DRAGEN automatically decompresses the files using ORA decompression, and then streams the reads into the mapper. The same FASTQ command-line options apply for all compression formats.

#### FASTQ Input Files

FASTQ input files can be single-ended or paired-end, as shown in the following examples.

* Single-ended in one FASTQ file (-1 option)

  ```
  dragen -r <REF_DIR> -1 <fastq> \
  --output-directory <OUT_DIR> -output-file-prefix <OUTPUT_PREFIX> \
  --RGID <RGID> --RGSM <RGSM>
  ```
* Paired-end in two matched FASTQ files(-1 and -2 options)

  ```
  dragen -r <REF_DIR> -1 <fastq1> -2 <fastq2> \
  --output-directory <OUT_DIR> --output-file-prefix <OUT_PREFIX> \
  --RGID <RGID> --RGSM <RGSM>
  ```
* Paired-end in a single interleaved FASTQ file(`--interleaved (-i)` option)

  ```
  dragen -r <REF_DIR> -1 <INTERLEAVED_FASTQ> -i \
  --RGID <RGID> --RGSM <RGSM>
  ```

Both bcl2fastq and the DRAGEN BCL command use a common file naming convention, as follows:

`<SampleID>_S<#>_<Lane>_<Read>_<segment#>.fastq.gz`

Older versions of bcl2fastq and DRAGEN could segment FASTQ samples into multiple files to limit file size or to decrease the time to generate them.

For Example:

```
RDRS182520_S1_L001_R1_001.fastq.gz

RDRS182520_S1_L001_R1_002.fastq.gz

...

RDRS182520_S1_L001_R1_008.fastq.gz
```

These files do not need to be concatenated to be processed together by DRAGEN. To map/align any sample, provide the first file in the series (`-1 <FileName>_001.fastq`). DRAGEN reads all segment files in the sample consecutively for both of the FASTQ file sequences specified using the -1 and -2 options for paired-end input and for compressed fastq.gz files. To turn the behavior off, set `‑‑enable-auto-multifile` to false on the command line.

DRAGEN can also optionally read multiple files by the sample name given in the file name, which can be used to combine samples that have been distributed across multiple BCL lanes or flow cells. To enable this feature, set the `--combine-samples-by-name` option to true

If the FASTQ files specified on the command-line use the Casava 1.8 file naming convention shown above and additional files in the same directory share that sample name, those files and all their segments are processed automatically. Note that sample name, read number, and file extension must match. Index barcode and lane number can differ.

To avoid impacting system performance, input files must be located on a fast file system.

**Note** The `--combine-samples-by-name` option is not functional since dragen v4.3. It is recommended to use the fastq-list input option instead.

#### Multiple FASTQ Input Files

To process multiple FASTQ input files as one sample, it is recommended that you use the `--fastq-list <csv file name>` option to specify the name of a CSV file containing the list of FASTQ files, instead of using the `--combine-samples-by-name` option.

For example:

```
dragen -r <ref_dir> --fastq-list <CSV_FILE> \
-fastq-list-sample-id <Sample_ID> -output-directory <OUT_DIR> 
--output-file-prefix <OUT_PREFIX>
```

Using a CSV file avoids having to concatenate the FASTQ files, for cases where there are multiple FASTQ files for a sample such as top-up scenarios or where FASTQ files are split across lanes. It also allows you to name the FASTQ input files, input from multiple subdirectories, and add BAM tags specified explicitly for each read group. DRAGEN automatically generates a CSV file of the correct format during BCL conversion to FASTQ. The CSV file is named `fastq_list.csv` and contains an entry for each FASTQ file or paired-end file pair produced during the run.

**FASTQ CSV File Format**

The first line of the CSV file specifies the title of each column, and is followed by one or more data lines. All lines in the CSV file must contain the same number of comma-separated values and should not contain white space or other extraneous characters.

Column titles are case-sensitive. The following column titles are required:

* RGID--Read Group
* RGSM--Sample ID
* RGLB--Library
* Lane--Flow cell lane
* Read1File--Full path to a valid FASTQ input file
* Read2File--Full path to a valid FASTQ input file. Required for paired-end input. If not using paired-end input, leave empty.

Each FASTQ file referenced in the CSV list can be referenced only once. All values in the Read2File column must be either nonempty and reference valid files, or they must all be empty.

When generating a BAM file using fastq-list input, one read group is generated per unique RGID value. The BAM header contains RG tags for the following read groups:

* ID (from RGID)
* SM (from RGSM)
* LB (from RGLB)

You can specify additional tags for each read group by adding a column title. The column title must be only four upper-case characters and begin with RG. For example, to add a PU (platform unit) tag, add a column named RGPU and specify the value for each read group in this column. All column titles must be unique.

A fastq-list file can contain files for more than one sample. If a fastq-list file contains only one unique RGSM entry, then no additional options need to be specified, and DRAGEN processes all files listed in the fastq-list file. If there is more than one unique RGSM entry in a fastq-list file, `--fastq-list-sample-id <SampleID>` must be used in addition to `--fastq-list <filename>` to process only a specific sample from the CSV file. Only the entries in the fastq-list file with an RGSM value that match the specified SampleID are processed.

* Independent processing and output for multiple individual samples in one run is not supported.
* To process all listed files together as one sample, regardless of the RGSM value, the option `--fastq-list-all-samples=true` can be used instead of `--fastq-list-sample-id`.

Note

For a single run, only one BAM and VCF output file are produced because all input read groups are expected to belong to the same sample. To process multiple samples independently from one BCL conversion run, DRAGEN must be run multiple times using different values for the \`--fastq-list-sample-id\` option.

There is no option to specify groupings or subsets of RGSM values for more complex filtering, but the fastq-list file can be modified to achieve the same effect.

The following is an example FASTQ list CSV file with the required columns:

```
RGID,RGSM,RGLB,Lane,Read1File,Read2File
CACACTGA.1,RDSR181520,UnknownLibrary,1,/staging/RDSR181520_S1_L001_R1_001.fastq,
/staging/RDSR181520_S1_L001_R2_001.fastq
AGAACGGA.1,RDSR181521,UnknownLibrary,1,/staging/RDSR181521_S2_L001_R1_001.fastq,
/staging/RDSR181521_S2_L001_R2_001.fastq
TAAGTGCC.1,RDSR181522,UnknownLibrary,1,/staging/RDSR181522_S3_L001_R1_001.fastq,
/staging/RDSR181522_S3_L001_R2_001.fastq
AGACTGAG.1,RDSR181523,UnknownLibrary,1,/staging/RDSR181523_S4_L001_R1_001.fastq,
/staging/RDSR181523_S4_L001_R2_001.fastq
```

If you use the `--tumor-fastq-list` option for somatic input, use the `--tumor-fastq-list-sample-id SampleID>` option to specify the sample ID for the corresponding FASTQ list, as shown in the following example:

```
dragen -r <ref_dir> --tumor-fastq-list <csv_file> \
--tumor-fastq-list-sample-id <Sample_ID> \
--output-directory <out_dir> \
--output-file-prefix <out_prefix> --fastq-list <csv_file_2> \
--fastq-list-sample-id <Sample_ID_2>
```

**Tumor-Normal Pairs Input**

If using fastq\_lists or tumor\_fastq\_lists comprising of multiple samples (RGSMs) in somatic mode, you can use a loop to iterate through the two lists to create tumor-normal pairs for testing. Create a \*.txt file with the RGSM of each normal sample to be tested (one per line), and then create a separate \*.txt file with the RGSM of the tumor samples to be tested. Make sure that the tumor sample RGSM is listed in the same order as the corresponding normal samples and to include a blank line after the last sample.

You can use the following example script to perform testing in somatic mode. Each iteration takes one entry from the tumor samples list and one entry from the normal samples list (from top to bottom) to create a tumor-normal pair as input for the DRAGEN run.

```
#!/bin/bash

HT="/staging/HT/"
tumor_fastq_list="/staging/inputs/tumor_fastq_list.csv"
normal_fastq_list="/staging/inputs/normal_fastq_list.csv"

tumor_samples_list="/staging/inputs/tumor_samples_list.txt"
normal_samples_list="/staging/inputs/normal_samples_list.txt"

while read -u 3 -r tumor_RGSM && read -u 4 -r normal_RGSM; do
output_dir="/staging/results/${tumor_RGSM}_${normal_RGSM}"
mkdir -p ${output_dir}

dragen \
-r ${HT} \
--tumor-fastq-list ${tumor_fastq_list} \
--tumor-fastq-list-sample-id ${tumor_RGSM} \
--fastq-list ${normal_fastq_list} \
--fastq-list-sample-id ${normal_RGSM} \
--output-directory ${output_dir} \
--output-file-prefix ${tumor_RGSM}_${normal_RGSM}
done 3<${tumor_samples_list} 4<${normal_samples_list}

```

```

Sample fastq_list.csv content:

RGPL,RGID,RGSM,RGLB,Lane,Read1File,Read2File
DRAGEN_RGPL,DRAGEN_RGID_N1.1,normal-1,ILLUMINA,1,/staging/inputs/normal-1_S1_L001_R1_001.fastq.gz,/staging/inputs/normal-1_S1_L001_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_N1.2,normal-1,ILLUMINA,2,/staging/inputs/normal-1_S1_L002_R1_001.fastq.gz,/staging/inputs/normal-1_S1_L002_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_N2.1,normal-2,ILLUMINA,1,/staging/inputs/normal-2_S1_L001_R1_001.fastq.gz,/staging/inputs/normal-2_S1_L001_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_N2.2,normal-2,ILLUMINA,2,/staging/inputs/normal-2_S1_L002_R1_001.fastq.gz,/staging/inputs/normal-2_S1_L002_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_N3.1,normal-3,ILLUMINA,1,/staging/inputs/normal-3_S1_L001_R1_001.fastq.gz,/staging/inputs/normal-3_S1_L001_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_N3.2,normal-3,ILLUMINA,2,/staging/inputs/normal-3_S1_L002_R1_001.fastq.gz,/staging/inputs/normal-3_S1_L002_R2_001.fastq.gz

```

The following are examples of the FASTQ lists and samples lists used as input for the script.

```
Sample tumor_fastq_list.csv content:

RGPL,RGID,RGSM,RGLB,Lane,Read1File,Read2File
DRAGEN_RGPL,DRAGEN_RGID_T1.1,tumor-1,ILLUMINA,1,/staging/inputs/tumor-1_S1_L001_R1_001.fastq.gz,/staging/inputs/tumor-1_S1_L001_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_T1.2,tumor-1,ILLUMINA,2,/staging/inputs/tumor-1_S1_L002_R1_001.fastq.gz,/staging/inputs/tumor-1_S1_L002_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_T2.1,tumor-2,ILLUMINA,1,/staging/inputs/tumor-2_S1_L001_R1_001.fastq.gz,/staging/inputs/tumor-2_S1_L001_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_T2.2,tumor-2,ILLUMINA,2,/staging/inputs/tumor-2_S1_L002_R1_001.fastq.gz,/staging/inputs/tumor-2_S1_L002_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_T3.1,tumor-3,ILLUMINA,1,/staging/inputs/tumor-3_S1_L001_R1_001.fastq.gz,/staging/inputs/tumor-3_S1_L001_R2_001.fastq.gz
DRAGEN_RGPL,DRAGEN_RGID_T3.2,tumor-3,ILLUMINA,2,/staging/inputs/tumor-3_S1_L002_R1_001.fastq.gz,/staging/inputs/tumor-3_S1_L002_R2_001.fastq.gz

```

```
Sample normal_samples_list content

normal-1
normal-2
normal-3

```

```
Sample tumor_samples_list content

tumor-1
tumor-2
tumor-3

```

#### FASTQ ORA Input Files

You can use the same options as the other FASTQ input file types for ORA files. To use the ORA file, replace the FASTQ file name with the ORA file name and specify the ORA reference directory using `--ora-reference`.

See ORA Compression and Decompression for more information on ORA reference files.

The following command represents paired-end in two matched ORA FASTQ files (-1 and -2 options).

```
dragen -r <REF_DIR> -1 <fastq.ora1> -2 <fastq.ora2> \
--ora-reference <ORADATA_DIR> \
--output-directory <OUT_DIR> --output-file-prefix <OUT_PREFIX> \
--RGID <RGID> --RGSM <RGSM>
```

#### BAM Input Files

BAM files can be used as input to the mapper/aligner. By default, `--enable-map-align` is true. When a BAM file input is provided with map/align enabled, DRAGEN ignores any alignment or duplicate marking information contained in the input file, reads are re-mapped and the new alignments are fed downstream to the variant callers. Any existing flags in the input BAM are erased when reads are re-mapped. BAM re-mapping is supported for multiple BAM inputs at a time, such as in paired tumor-normal input to somatic variant calling. Outputting the re-mapped BAM(s) can be enabled by setting `--enable-map-align-output=true`.

Alternatively, existing alignments in the BAM file can be used as input to the variant callers by setting the `--enable-map-align` option to false.

If the input file contains paired-end reads, it is important to specify that the input data should be sorted so that pairs can be processed together. Other pipelines would require you to re-sort the input data set by read name. DRAGEN vastly increases the speed of this operation by pairing the input reads, and sending them on to the mapper/aligner when pairs are identified. Use the `--pair-by-name` option to enable or disable this feature (the default is true).

Specify single-ended input in one BAM file with the (`-b`) and `--pair-by-name=false` options, as follows:

```
dragen -r <ref_dir> -b <bam> --output-directory <out_dir> \
--output-file-prefix <out_prefix> --pair-by-name false
```

Specify paired-end input in one BAM file with the (`-b`) and `\--pair-by-name=true` options, as follows:

```
dragen -r <ref_dir> -b <bam> --output-directory <out_dir> \
--output-file-prefix <out_prefix> --pair-by-name true
```

#### CRAM Input

You can use CRAM files as input to the DRAGEN mapper/aligner and variant caller. The DRAGEN functionality available when using CRAM input is the same as when using BAM input. Supported CRAM input file formats are v3.0 and v3.1.

By default, the CRAM compressor and decompressor uses the DRAGEN reference specified with the `--ref-dir` option. CRAM compression is reference based, and the reference used for compression is not part of the CRAM file. Therefore, the CRAM input file must have been created with the same reference than what is provided to DRAGEN for the analysis.

DRAGEN supports the re-alignment of a CRAM input that was created with a different reference in one step. Re-aligning a CRAM file that was created with a different reference requires use of the `--cram-reference` option. This option will make the CRAM decompressor use the specified reference.

* `--cram-reference` can be either a fasta file, or a DRAGEN hash table folder.
* If pointing to a fasta file, the fasta .fai index file must be present next to the fasta file
* CRAM output will always be compressed using the `--ref-dir` reference

*Example: CRAM was created with hg19, re-analysis with hg38*

```
dragen -r <ref_dir HG38> --cram-input <cram> --output-directory <out_dir> \
--output-file-prefix <out_prefix> --cram-reference <ref_dir HG19>
```

```
dragen -r <ref_dir HG38> --cram-input <cram> --output-directory <out_dir> \
--output-file-prefix <out_prefix> --cram-reference <hg19.fa>
```

The following options are used for providing a CRAM input to either mapper/aligner or variant caller:

* `--cram-input`--The name and path for the CRAM file
* `--cram-input`--One usage example is paired-end input in a single CRAM file. In addition, set the `--pair-by-name option` to true.

```
dragen -r <ref_dir> --cram-input <cram> --output-directory <out_dir> \
--output-file-prefix <out_prefix> --pair-by-name true
```

**Multiple BAM or CRAM Input Files**

To provide multiple BAM input files, you can use the `--bam-list <csv file name>` option to specify the name of a CSV file containing the list of BAM files. For example:

```
dragen -r <ref_dir> --bam-list <CSV_FILE> \
--output-directory <OUT_DIR> --output-file-prefix <OUT_PREFIX>
```

To provide multiple CRAM input files, you can use the `--cram-list <csv file name>` option.

**BAM or CRAM CSV Input File Format**

The first line of the CSV file specifies the header containing the title for each column and each subsequent line is a data line. All lines in the CSV file must contain the same number of comma-separated values and should not contain white space or any other extraneous characters.

An example BAM CSV file:

```
BamFile
/path/to/bam/one
/path/to/bam/two
```

Column titles are case sensitive. The following column titles are required:

* BamFile -- path to BAM file

Please note that only the "BamFile" column is supported as this time. Extra fields may be specified in the CSV file but they will not be processed by DRAGEN.

CRAM CSV input follows the same format above, with "CramFile" as the column title instead.

Restrictions and Limitations:

DRAGEN bam-list and cram-list are intended to mirror manually merging BAM or CRAM files via a utility such as samtools or MergeSamFiles (Picard). As a result, using bam-list or cram-list is analogous to having a single merged BAM or CRAM input file. Please note that some callers (i.e. DRAGEN variant calling) are unable to process a bam-list or cram-list that is composed of input files containing multiple samples.

In the case where identical read group IDs appear across multiple files and you want to treat them as distinct read groups, you can use the `--prepend-filename-to-rgid=true` option to distinguish between read groups.

If enabled, the resulting output BAM or CRAM file will contain all read groups from the input BAM or CRAM files passed in the CSV list file.

**Tumor-Normal Pairs Input**

You can also use `--tumor-bam-list <csv file name>` or `--tumor-cram-list <csv file name>` when running with tumor-only or tumor-normal inputs to DRAGEN. The CSV file has the same format as the options described above.

#### BCL Input Files

BCL is the output format of Illumina sequencing systems. Under limited circumstances, DRAGEN can read directly from BCL for map-align operations, saving the time needed for conversion to FASTQ.

DRAGEN can read directly from BCL in the following circumstances:

* Only one lane is input as part of a run (specified on the command-line).
* The lane has only a single sample specified in the SampleSheet.csv file. When converting BCL to FASTQ is required, DRAGEN provides a BCL to FASTQ converter (see DRAGEN BCL Data Conversion).

The following example command is for BCL input with only one lane of input:

```
dragen --bcl-input-dir <BCL_ROOT> --bcl-only-lane <num> -r <ref_dir> \
--output-directory <out_dir> --output-file-prefix <out_prefix>
```

For additional BCL conversion options, see Input File Types.

#### Handling of N bases

One of the techniques that DRAGEN uses to optimize handling sequences can lead to the overwriting the base quality score assigned to N base calls.

When you use the `--fastq-n-quality` and `--fastq-offset` options, the base quality scores are overwritten with a fixed base quality. The default values for these options are 2 and 33 to match the Illumina minimum quality of 35 (ASCII character ‘#’).

#### Read Names for Paired-End Reads

By a common convention, read names can include suffixes, such as `/1` or `/2`), which indicate the end of a pair the read represents. For BAM input using the `--pair-by-name` option, DRAGEN ignores these suffixes to find matching pair names. By default, DRAGEN uses the forward slash character as the delimiter for these suffixes and ignores the `/1` and `/2` when comparing names. By default, DRAGEN strips these suffixes from the original read names.

DRAGEN has the following options to control how suffixes are used:

* To change the delimiter character, for suffixes, use the `--pair-suffix-delimiter` option. Valid values for this option include forward-slash (/), dot (.), and colon (:).
* To preserve the entire name, including the suffixes, set `--strip-input-qname-suffixes` to false.
* To append a new set of suffixes to all read names, set `--append-read-index-to-name` to true. The delimiter is determined by the `--pair-suffix-delimiter` option. By default, the delimiter is a slash, so `/1` and `/2` are added to the names.

#### Gene Annotation Input Files

When processing RNA-Seq data, you can supply a gene annotations file by using the `--annotation-file` option. Providing this file improves the accuracy of the mapping and aligning stage (see \[Input Files]{.underline}). The file should conform to the GTF/GFF format specification and should list annotated transcripts that match the reference genome being mapped against. The similar GFF3 format is currently not supported, due to inconsistent contig naming between GENCODE and Ensembl. See the RNA user guide section for more details on potential issues and workarounds.

DRAGEN can take the SJ.out.tab file (see \[SJ.out.tab]{.underline}) as an annotations file to help guide the aligner in a two-pass mode of operation.

### Networked Streaming

#### AWS S3, Azure Blob Storage, and AWS Presigned URL Input Streaming

DRAGEN can stream input files directly from an AWS S3 bucket, Azure Blob storage account, or by using AWS presigned URLs (presigned URLs are not supported for Azure Blob storage at this time). With streaming, input files are not required to be downloaded locally prior to being processed. The files are streamed over the network directly into the DRAGEN processor.

Input streaming is most beneficial for large input files. DRAGEN supports input streaming for BAMs and compressed FASTQ files. For FASTQ files, input streaming can be used in all the configurations, including single-end FASTQs, paired-end FASTQs, and FASTQ lists.

Input streaming is supported for the following use cases:

* Mapping/aligning of FASTQ and BAM.
* Germline and somatic small variant calling from BAM (without remapping).
* Configuration files, including your [DRAGEN Cloud Credentials](/dragen-v4.5/reference/licensing/cloud_licensing#providing-credentials-via-license-credentials-file).

For other file types that are significantly smaller in size, download them locally before running the analysis.

**Streaming FASTQ Input Using AWS S3**

```
dragen -f \
  -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  -1 s3://s3-bucket-name/path/to/object_1.fastq.gz \
  -2 s3://s3-bucket-name/path/to/object_2.fastq.gz \
  --RGID object_ID \
  --RGSM sample_name \
  --output-directory /staging/examples/ \
  --output-file-prefix streaming
```

**Streaming FASTQ Input Using Azure Blob Storage Account**

```
AZ_ACCOUNT_NAME="storage-account-name" AZ_ACCOUNT_KEY="<account-key>" dragen -f \
  -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  -1 https://storage-account-name.blob.core.windows.net/path/to/object_1.fastq.gz \
  -2 https://storage-account-name.blob.core.windows.net/path/to/object_2.fastq.gz \
  --RGID object_ID \
  --RGSM sample_name \
  --output-directory /staging/examples/ \
  --output-file-prefix streaming
```

**Streaming FASTQ Input Using Presigned URLs (for AWS only)**

```
dragen -f \
  -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  -1 https://bucket-name.amazonaws.com/path/to/object_1.fastq.gz?querystring \
  -2 https://bucket-name.amazonaws.com/path/to/object_2.fastq.gz?querystring \
  --RGID object_ID \
  --RGSM sample_name \
  --output-directory /staging/examples/ \
  --output-file-prefix streaming
```

**Streaming BAM Input Using AWS S3**

```
dragen -f \
  -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  -b s3://s3-bucket-name/path/to/object_1.bam \
  --output-directory /staging/examples/ \
  --output-file-prefix streaming
```

**Streaming BAM Input Using Azure Blob Storage Account**

```
AZ_ACCOUNT_NAME="storage-account-name" AZ_ACCOUNT_KEY="<account-key>" dragen -f \
  -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  -b https://storage-account-name.blob.core.windows.net/path/to/object_1.bam \
  --output-directory /staging/examples/ \
  --output-file-prefix streaming
```

**Streaming BAM Input Using Presigned URLs (for AWS only)**

```
dragen -f \
  -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  -b https://bucket-name.amazonaws.com/path/to/object_1.bam?querystring \
  --output-directory /staging/examples/ \
  --output-file-prefix streaming
```

#### Azure Blob Storage Output Streaming

DRAGEN can stream its output to an Azure Blob Storage Account Container. Output streaming is beneficial for large output files and for sharing results.

**Streaming output to Azure Blob Storage Account**

```
AZ_ACCOUNT_NAME="storage-account-name" AZ_ACCOUNT_KEY="<account-key>" dragen -f \
  -r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  -1 SRA056922.fastq \
  --RGID object_ID \
  --RGSM sample_name \
  --output-directory https://storage-account-name.blob.core.windows.net/path/to/output \
  --intermediate-results-dir /staging/examples \
  --output-file-prefix streaming
```

#### Security and Permissions

To stream input files from cloud storage or write output to Azure Blob Storage, you must have permission to access the remote files.

**AWS S3 (input streaming only)**

S3 requires AWS authentication and credentials. The authentication should already be set up on the instance you are running, for example, via IAM policies. S3 output streaming is not supported.

**Azure Blob Storage Account**

Azure requires authentication and environment variables. DRAGEN supports two cases: (1) Using managed identities and (2) Storage account access keys.

To use managed identities you must run DRAGEN on an Azure instance. The instance must have `Contributor` permissions (read/write) on the Storage Account it wants to read and write to. If the instance has a single managed identity, only the `AZ_ACCOUNT_NAME=<azure-storage-account-name>` environment variable is required. For multiple managed identities, you must also provide the `AZR_IDENT_CLIENT_ID=<client-id>` environment variable, with the client id of the identity that can access your storage bucket. This can be found on the Azure Portal.

With storage account access keys, DRAGEN can write to an Azure bucket both on and off Azure instances. For this use case, find the [Storage Account Access Key](https://docs.microsoft.com/en-us/azure/storage/common/storage-account-keys-manage?toc=%2Fazure%2Fstorage%2Fblobs%2Ftoc.json\&tabs=azure-portal) and set the environment variables `AZ_ACCOUNT_NAME=<azure-storage-account-name>` and `AZ_ACCOUNT_KEY=<account-key>`.

**Presigned URL (AWS only)**

An AWS presigned URL most likely has a query string attached to it, which provides the authentication credentials or necessary tokens to grant permission to the S3 bucket (e.g., `https://bucket-name.amazonaws.com/path/to/folder?querystring`). Currently, streaming input to DRAGEN Azure presigned URLs is not supported.

### Sample Sex

Use the `--sample-sex` command line option to control the sex karyotype input used in downstream components, such as variant callers. If a sample sex karyotype input is not specified using the command line, the sex karyotype is automatically determined. The sex karyotype input is converted to a reference sex karyotype for use in variant calling. Other components might support sex karyotype input. Refer to the corresponding section for the component you are using.

The `--sample-sex` option supports the following values. Values are not case-sensitive.

* `none`: No sex karyotype input. Components use a default reference sex karyotype.
* `auto`: The sex karyotype is estimated by the Ploidy Estimator. If using CNV calling, sex karyotype is determined using a separate sex estimation module. If DRAGEN cannot estimate the sex karyotype, then components do not have a sex karyotype input. This behavior is then the same as `none`. `auto` is the default value.
* `female`: Sex karyotype input is XX.
* `male`: Sex karyotype input is XY.

The following example command lines use `--sample-sex` to specify the sex karyotype.

```
--sample-sex FEMALE

--sample-sex MALE

--sample-sex NONE
```

If the value is `none`, `female`, or `male`, the Ploidy Estimator could still run and produce output, but variant callers will not use any estimated sex karyotype that is different than the sex karyotype provided via the command-line.

The sex karyotype input is converted to the reference sex karyotype for the different components as follows. See the relevant component section for more information on how `--sample-sex` is used.

#### Reference Sex Karyotype

| Sex Karyotype Input | CNV Caller | DRAGEN-STR | Ploidy Caller | Small Variant Caller | SV Caller |
| ------------------- | ---------: | ---------: | ------------: | -------------------: | --------: |
| X0                  |         XX |         XY |            XX |                 XXYY |        XX |
| XX                  |         XX |         XX |            XX |                   XX |        XX |
| XXX                 |         XX |         XX |            XX |                 XXYY |        XX |
| XXXX                |         XX |         XX |            XX |                 XXYY |        XX |
| XXXXX               |         XX |         XX |            XX |                 XXYY |        XX |
| XY                  |         XY |         XY |            XY |                   XY |        XY |
| XXY                 |         XY |         XX |            XY |                 XXYY |        XY |
| XXXY                |         XY |         XX |            XY |                 XXYY |        XY |
| XXXXY               |         XY |         XX |            XY |                 XXYY |        XY |
| XYY                 |         XY |         XY |            XY |                 XXYY |        XY |
| XXYY                |         XY |         XX |            XY |                 XXYY |        XY |
| XXXYY               |         XY |         XX |            XY |                 XXYY |        XY |
| XYYY                |         XY |         XY |            XY |                 XXYY |        XY |
| XXYYY               |         XY |         XX |            XY |                 XXYY |        XY |
| XYYYY               |         XY |         XY |            XY |                 XXYY |        XY |
| None                |      XX/XY |         XX |         XX/XY |                 XXYY |      XXYY |

* For sex karyotype input of None, CNV/Ploidy Caller independently check the coverage ratio of X and Y to determine the reference sex karyotype. Detection of minimal Y coverage will yield XY, otherwise XX.

### Preservation or Stripping of BQSR Tags

The Picard Base Quality Score Recalibration (BQSR) tool produces output BAM files that include tags BI and BD. BQSR calculates these tags relative to the exact sequence for a read. If a BAM file with BI and BD tags is used as input to mapper/aligner with hard clipping enabled, the BI and/or BD tags can become invalid.

The recommendation is to strip these tags when using BAM files as input. To remove the BI and BD tags, set the `--preserve-bqsr-tags` option to false. If you preserve the tags, DRAGEN warns you to disable hard clipping.

### Read Group Options

DRAGEN assumes that all the reads in a given FASTQ belong to the same read group. DRAGEN creates a single @RG read group descriptor in the header of the output BAM file, with the ability to specify the following standard BAM attributes:

| Attribute | Argument | Description                                                                                                                                          |
| --------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| ID        | `--RGID` | Read group identifier. If you include any of the read group parameters, RGID is required. It is the value written into each output BAM record.       |
| LB        | `--RGLB` | Library.                                                                                                                                             |
| PL        | `--RGPL` | Platform/technology used to produce the reads. The BAM standard allows for values CAPILLARY, LS454, ILLUMINA, SOLID, HELICOS, IONTORRENT and PACBIO. |
| PU        | `--RGPU` | Platform unit, eg, flowcell-barcode.lane.                                                                                                            |
| SM        | `--RGSM` | Sample.                                                                                                                                              |
| CN        | `--RGCN` | Name of the sequencing center that produced the read.                                                                                                |
| DS        | `--RGDS` | Description.                                                                                                                                         |
| DT        | `--RGDT` | Date the run was produced.                                                                                                                           |
| PI        | `--RGPI` | Predicted mean insert size.                                                                                                                          |

If any of these arguments are present, DRAGEN adds an RG tag to all the output records to indicate that they are members of a read group. The following example shows a command line that includes read group parameters:

```
dragen --RGID 1 --RGCN Broad --RGLB Solexa-135852 \
--RGPL Illumina --RGPU 1 --RGSM NA12878 \
-r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
-1 SRA056922.fastq --output-directory /staging/tmp/ \
--output-file-prefix rg_example
```

When using the `--fastq-list` option to input multiple read groups, BAM tags (and others) are specified for each read group by adding columns to the `fastq_list.csv` file. Each column heading consists of four capital letters and each begins with 'RG'. For each column, each read group's values for that column are propagated to the output BAM file in an identically named tag.

### License Options

To suppress the license status message at the end of the run, use the `--lic-no-print` option. The following shows an example of the license status message:

```
LICENSE_MSG| =====================================================
LICENSE_MSG| License report
LICENSE_MSG|   Genome status [ACxxxxxxxxxxx] : used 1263.9 Gbases
since 2018-Feb-15 (1263886160894 bases, unlimited)
LICENSE_MSG|   Genome  bases [ACxxxxxxxxxxx] : 202000000
LICENSE_MSG|   Genome  bases [total]         : 202000000
```

## Autogenerated MD5SUM for BAM and CRAM Output Files

An MD5SUM file is generated automatically for BAM and CRAM output files. The MD5SUM file has the same name as the output file, with an .md5sum extension appended (eg, whole\_genome\_run\_123.bam.md5sum). The MD5SUM file is a single-line text file that contains the md5sum of the output file, which exactly matches the output of the Linux md5sum command.

The MD5SUM calculation is performed as the output file is written, so there is no measurable performance impact (compared to the Linux md5sum command, which can take several minutes for a 30x BAM).

## Configuration Files

Command line options can be stored in a configuration file. The location of the default configuration file is `<INSTALL_PATH>/config/dragen-user-defaults.cfg`. You can override this file by using the `--config-file (-c)` option to specify a different file. The configuration file used for a given run supplies the default settings for that run, any of which can be overridden by command line options.

The recommended approach is to use the dragen-user-defaults.cfg file as a template to create default settings for different use cases. Copy dragen-user-defaults.cfg, rename the copy, then modify the new file for the specific use-case. Best practice is to put options that rarely change into the configuration file and to specify options that vary from run to run on the command line.

## Licensing

DRAGEN utilizes quota based licensing for a majority of features. More information can be found in the [Licensing Reference Section](/dragen-v4.5/reference/licensing).


# DRAGEN Secondary Analysis

The DRAGEN secondary analysis software utilizes a highly reconfigurable Field Programmable Gate Array (FPGA) card and is available on a preconfigured DRAGEN server that can be seamlessly integrated into bioinformatics workflows. The platform can be loaded with highly optimized algorithms for many different NGS secondary analysis pipelines, including the following:

* Whole genome
* Exome
* RNA-Seq
* Methylome
* Cancer

All user interaction is accomplished via DRAGEN software that runs on the host server and manages all communication with the FPGA card. This user guide summarizes the technical aspects of the system and provides detailed information for all DRAGEN command line options. If you are working with DRAGEN for the first time, Illumina recommends that you first read the *Getting Started section*, which provides a short introduction to DRAGEN, including running a test of the server, generating a reference genome, and running example commands.

## DNA Pipeline

DRAGEN DNA Pipeline

![](/files/CvECfz8twOTCMr3rvleI)

The DRAGEN DNA Pipeline massively accelerates the secondary analysis of NGS data. For example, the time taken to process an entire human genome at 30x coverage is reduced from approximately 10 hours (using the current industry standard, BWA-MEM+GATK-HC software) to approximately 20 minutes. Time scales linearly with coverage depth.

These pipelines harness the tremendous power of the DRAGEN server and include highly optimized algorithms for mapping, aligning, sorting, duplicate marking, and haplotype variant calling. They also use platform features such as hardware-accelerated compression and optimized BCL conversion, together with the full set of platform tools.

Unlike all other secondary analysis methods, DRAGEN DNA Applications do not reduce accuracy to achieve speed improvements. Accuracy for both SNPs and INDELs is improved over that of BWA-MEM+GATK-HC in side-by-side comparisons.

In addition to haplotype variant calling, the pipeline supports calling of copy number and structural variants as well as detection of repeat expansions.

## RNA Pipeline

DRAGEN secondary anaylsis includes an RNA-seq (splicing-aware) aligner, as well as RNA-specific analysis components for gene expression quantification and gene fusion detection.

![](/files/iuMoCxH1VWCvdCiOBHla)

The DRAGEN RNA Pipeline shares many components with the DNA Pipeline. Mapping of short seed sequences from RNA-Seq reads is performed similarly to mapping DNA reads. In addition, splice junctions (the joining of noncontiguous exons in RNA transcripts) near the mapped seeds are detected and incorporated into the full read alignments.

DRAGEN secondary analysis uses hardware accelerated algorithms to map and align RNA-Seq--based reads faster and more accurately than popular software tools. For instance, it can align 100 million paired-end RNA-Seq--based reads in about three minutes. With simulated benchmark RNA-Seq data sets, its splice junction sensitivity and specificity are unsurpassed.

## Methylation Pipeline

The DRAGEN Methylation Pipeline provides support for automating the processing of bisulfite sequencing data to generate a BAM with the tags required for methylation analysis and reports detailing the locations with methylated cytosines.


# DRAGEN Apps


# DRAGEN Germline Enrichment ICA App

DRAGEN Germline Enrichment is an accurate and efficient end-to-end (FASTQ to VCF) secondary analysis solution for whole exome and targeted panel NGS data; it can be used for both germline and somatic variant calling. This app takes input files in FASTQ, ORA, BAM, and CRAM format. Files may be decompressed, go through Map/Align/Sort, and go through variant calling.

Before launching the app in [ICA](https://ica.illumina.com) the corresponding bundle containing the pipeline must be linked to your project. Follow the steps on the [ICA help site](https://help.ica.illumina.com/home/h-bundles#linking-an-existing-bundle-to-a-project) to link the DRAGEN bundle corresponding to the version and region you want to use (e.g. 'DRAGEN 4.4').

## In-run PON support for CNV and Targeted Caller

Version 4.4.4 of the DRAGEN Germline Enrichment App introduced support for automatic generation of an in-run PON for both CNV calling and the Targeted Caller. An in-run PON is required for Targeted Caller from WES data (see [Targeted Caller | Exome calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#exome-calling-using-in-run-pon)), and is the preferred replacement for a pre-built PON for CNV calling (see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals)).

Use the following steps to run the DRAGEN Germline Enrichment App using an in-run PON:

1. In your project go to 'Pipelines' and search for 'DRAGEN\_Germline\_Enrichment'. Click on the DRAGEN Germline Enrichment pipeline corresponding to the version you want to use.
2. Click 'Start Analysis'
3. Under 'General' section enter the name for the analysis under the 'User reference' field.
4. Complete the 'Pricing' section.
5. Select FASTQs/BAMs/CRAMs for all samples in the sequencing run. At least 5 samples are required for CNV and at least 30 samples are required for Targeted Caller.
6. Click '+' in the 'Reference' field and select the Illumina DRAGEN References version corresponding to the DRAGEN version being used. The correct reference version can be found on the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). For example, for DRAGEN 4.4 select 'Illumina DRAGEN v11 References'.
7. For human samples, select one of the supported pangenome reference TAR files containing 'graph' in the name and click 'Add'. Note that Targeted Caller only supports hg38, hg19 and hs37d5.
8. Click '+' in the 'Target BED File' field and select 'Illumina Enrichment BEDs' and then select the reference build selected previously.
9. Select the bed file corresponding to the library prep kit used. If the Illumina CS/PGx Custom Enrichment Research Panel was used, select 'Illumina\_Custom\_Enrichment\_Panel\_v2\_DRAGEN\_Targeted\_Calling.*\<hg38|hg19|hs37d5>*.bed'
10. Do not select files for the 'CNV Panel of Normals' or 'CNV Combined Counts' fields. An in-run PON will be built instead.
11. In the 'Settings | Variant Calling Options' section select the 'true' radio button for the 'Enable CNV calling' field.
12. Select the 'true' radio button for the 'CNV Enable In-Run Panel of Normals' field.
13. If the Illumina CS/PGx Custom Enrichment Research Panel was used go to the 'Settings | Additional Options' section and select the 'true' radio button for the 'Enable Targeted Calling' field.
14. In the field 'Excluded Sample for In-Run Panel of Normals' list the name of any samples from the sequencing run that should not be included in the PON. Examples of samples that can be listed here include:
    1. Samples that have not met QC requirements.
    2. Related samples from a large pedigree consisting of more than \~6% of the total number of PON samples (e.g. more than quad for PON of size 50).
    3. Samples that are known to not be copy-neutral in certain target regions of interest and would otherwise make up a large portion of the total PON samples.
15. The 'Samples Per Node' field defaults to 5. This value can be increased to reduce the cost overhead for spinning up more nodes, but will increase the total runtime. Decreasing this value will increase the cost overhead, but decrease the total runtime.


# DRAGEN Enrichment BSSH App

DRAGEN Enrichment is an accurate and efficient end-to-end (FASTQ to VCF) secondary analysis solution for whole exome and targeted panel NGS data; it can be used for both germline and somatic variant calling. This app takes input files in FASTQ, BAM, and CRAM formats. Files may be decompressed, go through Map/Align/Sort, and go through variant calling.

## In-run PON support for Germline CNV and Targeted Caller

Version 4.4.4 of the DRAGEN Enrichment BSSH App introduced support for automatic generation of an in-run PON for both CNV calling and the Targeted Caller. An in-run PON is required for Targeted Caller from WES data (see [Targeted Caller | Exome calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#exome-calling-using-in-run-pon)), and is the preferred replacement for a pre-built PON for CNV calling (see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals)).

Use the following steps to run the DRAGEN Enrichment BSSH App on Germline WES data using an in-run PON:

1. Go to version 4.4.4 or later of the [DRAGEN Enrichment BSSH App](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-enrichment.html) on BaseSpace and click 'Launch Application'.
2. Set the Analysis Name and output Project.
3. Select FASTQs/BAMs/CRAMs for all samples in the sequencing run. At least 5 samples are required for CNV and at least 30 samples are required for Targeted Caller. Input file types cannot be mixed.
4. Make sure 'Variant Caller Mode' is 'Germline' and select a reference. For human samples, select one of the supported pangenome references. Note that Targeted Caller only supports hg38, hg19 and hs37d5.
5. Select the 'Targeted Regions' corresponding to the enrichment protocol for the sequencing run. For Targeted Caller the 'Illumina CS/PGx Custom Enrichment Research Panel' must be used.
6. Expand the 'CNV' section and check the box 'Enable CNV'. For data generated using the Illumina CS/PGx Custom Enrichment Research Panel, check the box 'Enable Targeted Calling'. 'Enable CNV' must be checked before checking the 'Enable Targeted Calling'.
7. Check the box 'Enable In-Run Panel of Normals'.
8. In the field 'In-Run Panel of Normals Excluded Samples' list any samples from the sequencing run that should not be included in the PON. Examples of samples that can be listed here include:
   1. Samples that have not met QC requirements.
   2. Related samples from a large pedigree consisting of more than \~6% of the total number of PON samples (e.g. more than quad for PON of size 50).
   3. Samples that are known to not be copy-neutral in certain target regions of interest and would otherwise make up a large portion of the total PON samples.
9. Expand the 'Advanced Settings' section and set the 'Samples Per Node' field to 5. This value can be increased to reduce the cost overhead for spinning up more nodes, but will increase the total runtime. Decreasing this value will increase the cost overhead, but decrease the total runtime.
10. Click 'Launch Application'.


# DRAGEN Germline Enrichment from BCLs BSSH App

[DRAGEN Germline Enrichment from BCLs](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-germline-enrichment-from-bcls.html) provides accurate and efficient end-to-end (BCL to VCF) germline variant calling of whole exome and targeted panel NGS data. This app takes sequencing output in BCL format, converts them to FASTQs and then performs mapping and variant calling.

## In-run PON support for Germline CNV and Targeted Caller

An in-run PON is required for Targeted Caller from WES data (see [Targeted Caller | Exome calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#exome-calling-using-in-run-pon)), and is the preferred replacement for a pre-built PON for CNV calling (see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals)). Given samples from a single sequencing run, the DRAGEN Germline Enrichment from BCLs app will generate the required PON files and use them to analyze each sample in the run. The analysis workflow may be started from a planned run using the Run Planning tool in BSSH or from an existing sequencing run by launching the DRAGEN Germline Enrichment from BCLs app on BSSH. The DRAGEN Germline Enrichment from BCls app does not allow any samples from the sequencing run to be excluded. As a workaround, the DRAGEN BCL Convert Application can be used through Run Planning tool to generate FASTQs and the FASTQs can be used with either the [DRAGEN Germline Enrichment ICA app](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-germline-enrichment-ica-app) or the [DRAGEN Enrichment BSSH app](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-enrichment-bssh-app).

## DRAGEN Germline Enrichment from BCls through Run Planning Tool

A new run can be configured in BaseSpace Sequence Hub using the Run Planning tool. Follow the steps for your sequencing system described on the [Basespace Sequence Hub Help Center | Plan Runs](https://help.basespace.illumina.com/sequence/plan-runs) page.

During the configuration stage, use the following steps to configure the run for DRAGEN Germline Enrichment:

1. In the 'Application' field select the desired version under 'DRAGEN Germline Enrichment'
2. Select the appropriate 'Library Prep Kit' for the samples being sequenced. If the Illumina CS/PGx Custom Enrichment Research Panel was used with Exome 2.5 then select 'Illumina DNA Prep with Exome 2.5 Enrichment'.
3. Select the 'Index Adapter Kit' for the samples being sequenced.
4. Click 'Next'.
5. Complete the [Analysis Settings](#analysis-settings) section.
6. Click 'Next'.
7. Click 'Save as Planned'.

## DRAGEN Germline Enrichment from BCLs from existing run

The DRAGEN Germline Enrichment from BCLs app can run from an existing sequencing run containing BCLs using the steps below.

1. Go to the [BaseSpace Sequence Hub Dashboard](https://basespace.illumina.com/dashboard) and click on the 'Apps' tab.
2. In the search bar enter 'DRAGEN Germline Enrichment From BCLs' and click on the app.
3. Make sure the version selected is 4.4.4 or newer and click 'Launch Application'
4. Select the sequencing run containing BCLs
5. If the Sample Sheet in the selected sequencing run is invalid or outdated there may be validation errors. In this case an 'Override Sample Sheet' must be selected. The 'Override Sample Sheet' can be exported from a run created using the [Run Planning tool](#dragen-germline-enrichment-from-bcls-from-existing-run) as described above. In the final step of the Run Planning tool, instead of clicking 'Save as Planned', click 'Export' to generate the Override Sample Sheet. Note that the Sample Sheet must first be uploaded to a project on BSSH before it can be selected in the app.
6. Click 'Edit' in the 'Configuration: DRAGEN Germline Enrichment' section.
7. Complete the [Analysis Settings](#analysis-settings) section.
8. Click 'Next'.
9. Click 'Requeue' to start the app.

### Analysis Settings

Follow the steps below to configure the analysis settings:

1. Select the 'Reference TAR'. For human samples, select one of the supported pangenome references. Note that Targeted Caller only supports hg38, hg19 and hs37d5.
2. Select the 'Panel BED'. If the Illumina CS/PGx Custom Enrichment Research Panel was used, select 'Illumina\_Custom\_Enrichment\_Panel\_v2\_DRAGEN\_Targeted\_Calling.*\<hg38|hg19|hs37d5>*.bed'
3. Under 'DRAGEN copy number and targeted calling settings' select the 'In-Run' radio button for the 'CNV Baseline Source' field.
4. If the Illumina CS/PGx Custom Enrichment Research Panel was used select the 'Yes' radio button for the 'Enable targeted paralog calling' field.


# DRAGEN Recipes

### Overview

The following sub-pages contain recommended command line options for specific DRAGEN pipelines. For an overview of DRAGEN command line parsing, also see [Multicaller Workflows](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/multi-caller)

#### Germline Pipelines

* [DNA Germline Amplicon](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-amplicon)
* [DNA Germline Panel](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-panel)
* [DNA Germline WES](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-wes)
* [DNA Germline WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-wgs)
* [5 Base DNA Germline Panel](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-germline-panel)
* [5 Base DNA Germline WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-germline-wgs)

#### Germline with UMI Pipelines

* [DNA Germline Panel UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-panel-umi)
* [DNA Germline WES UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-wes-umi)
* [5 Base DNA Germline Panel UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-germline-panel-umi)

#### RNA and scRNA Pipelines

* [Illumina scRNA](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/illumina-scrna)
* [Illumina scRNA Perturb seq](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/illumina-scrna-perturb-seq)
* [Other scRNA prep](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/other-scrna-prep)
* [RNA Amplicon](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/rna-amplicon)
* [RNA Panel](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/rna-panel)
* [RNA WTS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/rna-wts)

#### Somatic Pipelines

* [DNA Somatic Tumor-Normal Heme WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-heme-wgs)
* [DNA Somatic Tumor-Normal MRD](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-mrd)
* [DNA Somatic Tumor-Normal Solid Panel](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-solid-panel)
* [DNA Somatic Tumor-Normal Solid WES](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-solid-wes)
* [DNA Somatic Tumor-Normal Solid WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-solid-wgs)
* [DNA Somatic Tumor-Only Heme WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-heme-wgs)
* [DNA Somatic Tumor-Only Solid Amplicon](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-solid-amplicon)
* [DNA Somatic Tumor-Only Solid Panel](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-solid-panel)
* [DNA Somatic Tumor-Only Solid WES](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-solid-wes)
* [DNA Somatic Tumor-Only Solid WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-solid-wgs)
* [DNA Somatic Tumor-Only ctDNA Amplicon](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-ctdna-amplicon)
* [5 Base DNA Somatic Tumor-Normal Solid Panel](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-somatic-tumor-normal-solid-panel)
* [5 Base DNA Somatic Tumor-Normal Solid WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-somatic-tumor-normal-solid-wgs)
* [5 Base DNA Somatic Tumor-Only Solid Panel](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-somatic-tumor-only-solid-panel)
* [5 Base DNA Somatic Tumor-Only Solid WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-somatic-tumor-only-solid-wgs)

#### Somatic with UMI Pipelines

* [DNA Somatic Tumor-Normal Solid Panel UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-solid-panel-umi)
* [DNA Somatic Tumor-Normal Solid WES UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-solid-wes-umi)
* [DNA Somatic Tumor-Normal Solid WGS UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-normal-solid-wgs-umi)
* [DNA Somatic Tumor-Only Solid Panel UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-solid-panel-umi)
* [DNA Somatic Tumor-Only Solid WES UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-solid-wes-umi)
* [DNA Somatic Tumor-Only Solid WGS UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-solid-wgs-umi)
* [DNA Somatic Tumor-Only ctDNA Panel UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-somatic-tumor-only-ctdna-panel-umi)
* [5 Base DNA Somatic Tumor-Only Solid Panel UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-somatic-tumor-only-solid-panel-umi)
* [5 Base DNA Somatic Tumor-Only ctDNA Panel UMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/5-base-dna-somatic-tumor-only-ctdna-panel-umi)

#### TruPath Pipelines

* [Illumina TruPath Genome WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/illumina-trupath-genome-wgs)


# DNA Germline Amplicon

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking false        #default=false 
# Amplicon 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED          #Optional. Auto-generated based on amplicon target bed. 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# SV 
--enable-sv true 
# CNV 
--enable-cnv true 
--cnv-combined-counts $PATH             #CNV PON. Required for amplicon CNV calling on CASE samples. 
--cnv-target-bed $PATH                  #Optional. Auto-generated based on amplicon target bed. 
--cnv-filter-qual $NUM                  #CNV filter quality. Adjust CNV filter quality thresholds according to the user’s validation study. 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                     |
| -------------------------------- | ----------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).  |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false). |

### Duplicate Marking

| Option                             | Description                                                                                                                                                                                               |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking false` | The Amplicon Pipeline disables duplicate marking. In amplicon assays, fragments originate from a limited number of unique start and end positions, making conventional duplicate detection inappropriate. |

### SNV

DRAGEN amplicon does not employ machine learning based variant recalibration (DRAGEN-ML).

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### CNV

Amplicon CNV requires PON input. In PON mode, the DRAGEN CNV Pipeline is broken down into two distinct stages. The target counts stage is performed on each sample (case and normals), to bin the alignments. The normalization and call detection stage is then performed with the case sample against the panel of normals to determine the events.

| Option                                 | Description                                                                                                                                                                                                                                                                                                       |
| -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. By default, `bed` is used for standard panels and `hslm` for Pillar panels with a pre-built PON.                                                                                                                                                           |
| `--amplicon-cnv-use-default-pon false` | We recommend including in-run normal samples—matched in sample type and library preparation—in the same sequencing run to serve as the PON. If generating a custom PON is not feasible, for Pillar panels, the pre-packaged panel-specific PON can be used as a fallback. To enable this, set the option to true. |
| `--cnv-segmentation-bed $PATH`         | You can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. If bed segmentation mode is used, the segmentation bed is auto-generated from amplicon target bed by default                                                                                                |
| `--cnv-filter-qual $NUM`               | QUAL value at which to hard filter CNV VCF. You can adjust CNV filter quality thresholds according to the your validation study                                                                                                                                                                                   |

### In-run PON

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

**Step 1. Generate CNV target counts of individual samples from the sequencing run.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.


# DNA Germline Panel

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# SV 
--enable-sv true 
--sv-exome true 
--sv-locus-node-target-file $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON. See 'In-run PON' section below. 
# HLA genotyper 
--enable-hla true 
--hla-as-filter-min-threshold 29.0      #panel specific setting 
--hla-as-filter-ratio-threshold 0.85    #panel specific setting 
--hla-exome true                        #If panel contains only coding regions.  Sets HLA output to 3 field resolution (coding sequence only) 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true`     | By default, DRAGEN marks duplicate reads and exclude them from variant calling.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--enable-positional-collapsing true` | Alternative to `--enable-duplicate-marking=true`. Instead of discarding duplicate reads, DRAGEN can optionally perform positional collapsing, merging them into higher-quality consensus reads. This is beneficial for small panels without UMIs and coverage between 300X and 1000X. However, it's slower than standard duplicate marking and less effective on samples with coverage lower than 300X. For very high coverage (1000X+), avoid it due to potential read collisions. For high-sensitivity panels with 1000X+ coverage, consider using UMIs. |

### SNV

DRAGEN SNV VC employs machine learning based variant recalibration (DRAGEN-ML). It processes read and other contextual evidence to remove false positives, recover false negatives and reduce zygosity errors. No additional setup is required. DRAGEN-ML is enabled by default as needed, when running the germline SNV VC on hg19 or hg38.

Note that we do not recommend changing the default QUAL thresholds of 3 for DRAGEN-ML and 10 for DRAGEN without ML. These values differ from each other because DRAGEN-ML improves the calibration of QUAL scores, leading to a change in the scoring range.

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |
| `--vc-emit-ref-confidence GVCF`             | To enable gVCF output.                                                                                                                       |
| `--vc-enable-vcf-output`                    | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                | Description                                                                                                                                               |
| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true` | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`   | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`        | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |

### In-run PON

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

**Step 1. Generate CNV target counts of individual samples from the sequencing run.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.


# DNA Germline Panel UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# SV 
--enable-sv true 
--sv-exome true 
--sv-locus-node-target-file $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON. See 'In-run PON' section below. 
# HLA genotyper 
--enable-hla true 
--hla-as-filter-min-threshold 29.0      #panel specific setting 
--hla-as-filter-ratio-threshold 0.85    #panel specific setting 
--hla-exome true                        #If panel contains only coding regions.  Sets HLA output to 3 field resolution (coding sequence only) 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

DRAGEN SNV VC employs machine learning based variant recalibration (DRAGEN-ML). It processes read and other contextual evidence to remove false positives, recover false negatives and reduce zygosity errors. No additional setup is required. DRAGEN-ML is enabled by default as needed, when running the germline SNV VC on hg19 or hg38.

Note that we do not recommend changing the default QUAL thresholds of 3 for DRAGEN-ML and 10 for DRAGEN without ML. These values differ from each other because DRAGEN-ML improves the calibration of QUAL scores, leading to a change in the scoring range.

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |
| `--vc-emit-ref-confidence GVCF`             | To enable gVCF output.                                                                                                                       |
| `--vc-enable-vcf-output`                    | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                | Description                                                                                                                                               |
| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true` | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`   | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`        | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |

### In-run PON

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

**Step 1. Generate CNV target counts of individual samples from the sequencing run.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.


# DNA Germline WES

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# SV 
--enable-sv true 
--sv-exome true 
--sv-target-bed $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF  #needed for AOH detection 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON. See 'In-run PON' section below. 
# HLA genotyper 
--enable-hla true 
--hla-exome true                        #Set HLA output to 3 field resolution (coding sequence only) 
# Targeted caller (only if using the Illumina CS/PGx Custom Enrichment Research Panel) 
--enable-targeted true 
--targeted-pon $PATH                    #Targeted PON. See 'In-run PON' section below. 
--targeted-systematic-noise $PATH       #Targeted systematic noise file 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### SNV

DRAGEN SNV VC employs machine learning based variant recalibration (DRAGEN-ML). It processes read and other contextual evidence to remove false positives, recover false negatives and reduce zygosity errors. No additional setup is required. DRAGEN-ML is enabled by default as needed, when running the germline SNV VC on hg19 or hg38.

Note that we do not recommend changing the default QUAL thresholds of 3 for DRAGEN-ML and 10 for DRAGEN without ML. These values differ from each other because DRAGEN-ML improves the calibration of QUAL scores, leading to a change in the scoring range.

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |
| `--vc-emit-ref-confidence GVCF`             | To enable gVCF output.                                                                                                                       |
| `--vc-enable-vcf-output`                    | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                               |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF to enable AOH detection through allele-specific CN calling (ASCN).                                                           |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default true).                                                                                                     |
| `--cnv-enable-mosaic-calling true`       | Enable MOSAIC-calling mode (default true).                                                                                                                |

### In-run PON

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

For Targeted Caller PON requirements and generation options see [Targeted Caller | Exome calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#exome-calling-using-in-run-pon).

CNV and Targeted Caller require separate PON files, but the intermediate counts files can be generated in the same DRAGEN command line invocation. Follow the steps below to generate the CNV and Targeted Caller PON files. Note that Targeted Caller is only supported with the Illumina CS/PGx Custom Enrichment Research Panel.

**Step 1. Generate CNV target counts and Targeted exome counts of individual samples from the sequencing run.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
# Targeted caller (only if using the Illumina CS/PGx Custom Enrichment Research Panel) 
--targeted-generate-exome-counts true 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

**Step 3. Targeted Caller PON file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--targeted-pon-counts-list $TARGETED_PON_COUNTS_LIST 
```

`$TARGETED_PON_COUNTS_LIST` is a text file with one line for each path to a Targeted Caller exome counts file generated in step 1 (`<output-file-prefix>.targeted.exome.counts.json.gz`). Individual exome counts files are merged into a single `<output-file-prefix>.targeted.pon.json.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN Targeted Caller using the `--targeted-pon` option.

### Targeted Caller

A systematic noise file corresponding to one of the pre-built pangenome references can be downloaded from the \[DRAGEN Software Support Site page]<https://support.illumina.com/sequencing/sequencing\\_software/dragen-bio-it-platform/product\\_files.html>).


# DNA Germline WES UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# SV 
--enable-sv true 
--sv-exome true 
--sv-target-bed $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF  #needed for AOH detection 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON. See 'In-run PON' section below. 
# HLA genotyper 
--enable-hla true 
--hla-exome true                        #Set HLA output to 3 field resolution (coding sequence only) 
# Targeted caller (only if using the Illumina CS/PGx Custom Enrichment Research Panel) 
--enable-targeted true 
--targeted-pon $PATH                    #Targeted PON. See 'In-run PON' section below. 
--targeted-systematic-noise $PATH       #Targeted systematic noise file 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

DRAGEN SNV VC employs machine learning based variant recalibration (DRAGEN-ML). It processes read and other contextual evidence to remove false positives, recover false negatives and reduce zygosity errors. No additional setup is required. DRAGEN-ML is enabled by default as needed, when running the germline SNV VC on hg19 or hg38.

Note that we do not recommend changing the default QUAL thresholds of 3 for DRAGEN-ML and 10 for DRAGEN without ML. These values differ from each other because DRAGEN-ML improves the calibration of QUAL scores, leading to a change in the scoring range.

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |
| `--vc-emit-ref-confidence GVCF`             | To enable gVCF output.                                                                                                                       |
| `--vc-enable-vcf-output`                    | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                               |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF to enable AOH detection through allele-specific CN calling (ASCN).                                                           |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default true).                                                                                                     |
| `--cnv-enable-mosaic-calling true`       | Enable MOSAIC-calling mode (default true).                                                                                                                |

### In-run PON

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

For Targeted Caller PON requirements and generation options see [Targeted Caller | Exome calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#exome-calling-using-in-run-pon).

CNV and Targeted Caller require separate PON files, but the intermediate counts files can be generated in the same DRAGEN command line invocation. Follow the steps below to generate the CNV and Targeted Caller PON files. Note that Targeted Caller is only supported with the Illumina CS/PGx Custom Enrichment Research Panel.

**Step 1. Generate CNV target counts and Targeted exome counts of individual samples from the sequencing run.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
# Targeted caller (only if using the Illumina CS/PGx Custom Enrichment Research Panel) 
--targeted-generate-exome-counts true 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

**Step 3. Targeted Caller PON file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--targeted-pon-counts-list $TARGETED_PON_COUNTS_LIST 
```

`$TARGETED_PON_COUNTS_LIST` is a text file with one line for each path to a Targeted Caller exome counts file generated in step 1 (`<output-file-prefix>.targeted.exome.counts.json.gz`). Individual exome counts files are merged into a single `<output-file-prefix>.targeted.pon.json.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN Targeted Caller using the `--targeted-pon` option.

### Targeted Caller

A systematic noise file corresponding to one of the pre-built pangenome references can be downloaded from the \[DRAGEN Software Support Site page]<https://support.illumina.com/sequencing/sequencing\\_software/dragen-bio-it-platform/product\\_files.html>).


# DNA Germline WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# SV 
--enable-sv true 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF  #needed for AOH detection 
--cnv-enable-self-normalization true 
# HLA genotyper 
--enable-hla true 
# Targeted caller 
--enable-targeted true 
# Star allele 
--enable-star-allele true 
# PGX 
--enable-pgx true                       #PGX 
# Short tandem repeats 
--repeat-genotype-enable true 
# Multi-Region Joint Detection (MRJD) 
--enable-mrjd true 
--mrjd-enable-high-sensitivity-mode true 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### SNV

DRAGEN SNV VC employs machine learning based variant recalibration (DRAGEN-ML). It processes read and other contextual evidence to remove false positives, recover false negatives and reduce zygosity errors. No additional setup is required. DRAGEN-ML is enabled by default as needed, when running the germline SNV VC on hg19 or hg38.

Note that we do not recommend changing the default QUAL thresholds of 3 for DRAGEN-ML and 10 for DRAGEN without ML. These values differ from each other because DRAGEN-ML improves the calibration of QUAL scores, leading to a change in the scoring range.

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |
| `--vc-emit-ref-confidence GVCF`             | To enable gVCF output.                                                                                                                       |
| `--vc-enable-vcf-output`                    | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                               |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF to enable AOH detection through allele-specific CN calling (ASCN).                                                           |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default true).                                                                                                     |
| `--cnv-enable-mosaic-calling true`       | Enable MOSAIC-calling mode (default true).                                                                                                                |

### Multi-Region Joint Detection (MRJD)

| Option                                | Description                                                                                                                                                                                                                     |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-mrjd`                       | If set to true, MRJD is enabled for the DRAGEN pipeline.                                                                                                                                                                        |
| `--mrjd-enable-high-sensitivity-mode` | If set to true, MRJD high sensitivity mode is enabled for the DRAGEN pipeline. See the MRJD section in the user guide for information on variant types reported in MRJD default mode and high-sensitivity mode (default=false). |

For futher details refer to [MRJD](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/multi-region-joint-detection).


# DNA Somatic Tumor-Normal Heme WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Recommended 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--heme-sv true 
--sv-systematic-noise $PATH             #Optional 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# DUX4 
--enable-dux4-caller true 
# CNV 
--heme-cnv true 
--cnv-use-somatic-vc-baf true 
--cnv-enable-self-normalization true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# TMB 
--enable-tmb true 
# HLA genotyper 
--enable-hla true 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #recommended 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                 | Description                                                                                                                                               |
| -------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-enable-cyto-output true`        | Enable Cytogenetics-compatible output (default false).                                                                                                    |
| `--heme-cnv true`                      | Configures DRAGEN to use CNV settings for HEME.                                                                                                           |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                     |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |
| `--heme-sv true`                                  | Configures DRAGEN to use SV settings for Liquid Tumors (e.g., AML/MLL).                  |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

### DUX4

| Option                          | Description                                                                                               |
| ------------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--dux4-skip-sanity-check true` | Bypass the requirements checks if the input datasets don't comply with parameters listed in prerequisites |

For more information, see [DUX4-rearrangement Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/dux4-rearrangement-caller).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files are recommended for WGS workflows. Prebuilt WGS noise files are available.

#### Prebuilt

Prebuilt WGS SV systematic noise files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WGS noise files                            | Description      |
| --------------------------------------------------- | ---------------- |
| `WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz`      | For WGS, FF/FFPE |
| `IDPF_WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz` | For WGS, HEME    |

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Normal MRD

> **For conceptual background and pipeline overview, see** [**DRAGEN MRD Pipeline Overview**](/dragen-v4.5/product-guides/dragen-v4.5/mrd)

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

## Step 0: Fastq generation

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--output-directory $OUTPUT_DIR 
--sample-sheet $SAMPLE_SHEET 
--bcl-input-directory $RUN_FOLDER 
--bcl-conversion-only true 
--strict-mode true 
# if using ora compression (.fastq.ora) rather than gzip (.fastq.gz) 
--ora-reference $ORA_REFERENCE 
--fastq-compression-format dragen 
```

BCL conversion is optional if FASTQ data already exists. If starting from BCL files, this step must be completed before running the MRD pipeline to ensure sample-specific FASTQs are available as input.

## Step 1: Read alignment and targeted variant calling

### Step 1A: Read alignment and targeted germline variant calling (FFPE)

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--validate-pangenome-reference false    #required 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH 
--output-file-prefix $PREFIX 
--events-log-file $EVENTS_LOG_FILE 
--watchdog-active-timeout 1800 
--watchdog-idle-timeout 1800 
# Inputs (e.g. FQ list) 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional for BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-duplicate-marking true         #default=true 
--enable-targeted false 
--Aligner.hard-clips 7                  #for FFPE samples only, uses hard clipping for all alignment types 
# Small variant caller 
--enable-variant-caller true            #targeted germline calling for ~37K SNP sites with high population allele frequencies (typically close to 50% VAF) 
--vc-target-bed $COMMON_GERMLINE_TARGET_BED 
# QC 
--gc-metrics-enable true 
--qc-coverage-ignore-overlaps false     #de-duplicated conventional coverage is output rather than molecular coverage 
--qc-cross-cont-vcf $QC_CROSS_CONTAMINATION_VCF 
# ORA 
--ora-reference $ORA_REFERENCE 
```

### Step 1B: Read alignment and targeted germline variant calling (BC/Plasma)

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--validate-pangenome-reference false    #required 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH 
--output-file-prefix $PREFIX 
--events-log-file $EVENTS_LOG_FILE 
--watchdog-active-timeout 1800 
--watchdog-idle-timeout 1800 
# Inputs (e.g. FQ list) 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional for BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-duplicate-marking true         #default=true 
--enable-targeted false 
# Small variant caller 
--enable-variant-caller true            #targeted germline calling for ~37K SNP sites with high population allele frequencies (typically close to 50% VAF) 
--vc-target-bed $COMMON_GERMLINE_TARGET_BED 
# QC 
--gc-metrics-enable true 
--qc-coverage-ignore-overlaps false     #de-duplicated conventional coverage is output rather than molecular coverage 
--qc-cross-cont-vcf $QC_CROSS_CONTAMINATION_VCF 
# ORA 
--ora-reference $ORA_REFERENCE 
```

### Read alignment and targeted variant calling notes

For consistency, use the linear reference. For FFPE samples, use `--Aligner.hard-clips=7` to use hard clipping for all alignment types; omit this parameter for Buffy Coat or Plasma samples.

## Step 2: Fingerprint generation + QC

### Step 2A: Fingerprint generation and FFPE normal-aware contamination QC

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH 
--output-file-prefix $PREFIX 
--events-log-file $EVENTS_LOG_FILE 
# Inputs (BAM files from Step 1) 
--bam-input $BUFFY_COAT_BAM 
--tumor-bam-input $FFPE_BAM 
# M/A 
--enable-map-align false 
--enable-map-align-output false 
# Small variant caller 
--enable-variant-caller true 
--vc-enable-germline-tagging true 
--vc-target-bed $HIGH_CONFIDENCE_REGIONS 
--vc-systematic-noise $VC_SYSTEMATIC_NOISE 
--vc-germline-tagging-db-files $VC_GERMLINE_TAGGING_DB_FILES 
# FP 
--mrd-fingerprint true 
# QC 
--qc-somatic-contam-vcf $QC_SOMATIC_CONTAMINATION_VCF 
```

### Fingerprint generation notes

For the DRAGEN MRD pipeline (similar to all somatic runs) it is recommended to use the linear hashtable. DRAGEN hashtables can be downloaded from [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

It is recommended to use `--vc-target-bed $BED` or `--vc-excluded-regions-bed $BED` to limit fingerprint calls to high-confidence regions. Construct a BED file covering only easily mapped regions, excluding ALU or highly repetitive regions where recurring noise tends to be more frequent.

Use a systematic noise file to further reduce false positives. Prebuilt systematic noise BED files can be downloaded from [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html):

| Prebuilt WGS noise files                           | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |

For more information see [SNV systematic noise](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode#systematic-noise-filtering). To download germline annotation files, refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana).

### FFPE normal-aware contamination QC notes

It is recommended to use `--qc-somatic-contam-vcf` for FFPE normal-aware contamination detection. The somatic contamination VCF files are bundled with every DRAGEN installation at `/opt/dragen/<version>/resources/qc/`. For the hg38 reference, use `somatic_sample_cross_contamination_resource_hg38.vcf.gz`.

For samples not run in matched Tumor/Normal mode (i.e., BC and Plasma samples in Step 1B), a germline contamination VCF file can be passed to `--qc-cross-cont-vcf`. These files are also located at `/opt/dragen/<version>/resources/qc/`. For the hg38 reference, use `sample_cross_contamination_resource_hg38.vcf.gz`.

### Step 2B: FFPE/BC sample matching QC

```
  
/opt/dragen/$VERSION/bin/dragen 
--ref-dir $REF_DIR 
--output-directory $OUTPUT_DIR 
--output-file-prefix $SAMPLE_ID 
--events-log-file $EVENTS_LOG_FILE 
# Inputs 
--checkfingerprint-expected-vcf $BUFFY_COAT_TARGETED_GERMLINE_VCF 
--checkfingerprint-observed-vcf $FFPE_TARGETED_GERMLINE_VCF 
# QC 
--enable-checkfingerprint true 
```

## Step 3: MRD detection

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR 
--output-directory $OUTPUT_DIR 
--output-file-prefix $SAMPLE_ID 
--events-log-file $EVENTS_LOG_FILE 
# Inputs 
--bam-input $PLASMA_BAM 
--mrd-probes-file $FINGERPRINT_VCF 
# M/A 
--enable-map-align false 
--enable-map-align-output false 
--enable-sort false 
# MRD 
--enable-mrd true 
```

### MRD detection notes

The command line parameters that control MRD detect are:

| Parameter Name          | Description                                                                                           |
| ----------------------- | ----------------------------------------------------------------------------------------------------- |
| `--enable-mrd`          | Enables MRD detect. Default = "false".                                                                |
| `--mrd-probes-file`     | Path to the individual's tumor fingerprint VCF file                                                   |
| `--mrd-score-threshold` | Threshold used to determine the presence/absence of residual cancer DNA in the plasma. Default = 4.0. |

The MRD detect module generates an output summary file using the standard DRAGEN output directory and prefix: .mrd\_summary.json. The file is a valid JSON file that contains an array of JSON objects. DRAGEN supports running one sample at a time, so the array will be of length one.

The output JSON will include the following two fields of interest:

* Run\[1].TumorEstimate.illumina.eVAF
* Run\[1].TumorEstimate.illumina.score

The "eVAF" (estimated Variant Allele Frequency) is the allele-level fraction of detected variants in the plasma sample, reflecting ctDNA burden. See [Interpreting eVAF](/dragen-v4.5/product-guides/dragen-v4.5/mrd#interpreting-evaf) for details.

The "score" can be used to determine presence/absence of residual cancer DNA in the plasma. A higher score indicates that the presence of cancer DNA is more likely. The exact threshold score that is used to indicate a positive ctDNA status may depend on sample quality and coverage, and can be optimized for a specific pipeline. It is expected that this threshold will typically be between 4 - 7.

## Step 4: Plasma QC

### Step 4A: Plasma/BC sample matching QC

```
  
/opt/dragen/$VERSION/bin/dragen 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--validate-pangenome-reference false    #required 
--output-directory $OUTPUT 
--output-file-prefix $PREFIX 
--events-log-file $EVENTS_LOG_FILE 
# Inputs 
--checkfingerprint-expected-vcf $BUFFY_COAT_TARGETED_GERMLINE_VCF 
--checkfingerprint-observed-vcf $PLASMA_TARGETED_GERMLINE_VCF 
# QC 
--enable-checkfingerprint true 
```

### Step 4B: Plasma contamination QC

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR 
--output-directory $OUTPUT_DIR 
--output-file-prefix $SAMPLE_ID 
--events-log-file $EVENTS_LOG_FILE 
# Inputs 
--bam-input $PLASMA_BAM 
# M/A 
--enable-map-align false 
--enable-map-align-output false 
--enable-sort false 
# MRD 
--enable-mrd true 
--mrd-probes-file $COMMON_GERMLINE_VCF 
--mrd-blocklist $BUFFY_COAT_TARGETED_GERMLINE_VCF 
```

### Plasma QC notes

Similar to the MRD detection step, the output JSON will include the following field of interest:

```
 Run[1].TumorEstimate.illumina.eVAF 
```

The "eVAF" (estimated Variant Allele Frequency) is the observed allele fraction. For the plasma contamination QC workflow, the inferred foreign DNA fraction is approximately 2\*eVAF (since signal is measured at loci with population allele frequencies close to 50%). See [Interpreting eVAF in plasma contamination QC](/dragen-v4.5/product-guides/dragen-v4.5/mrd#interpreting-evaf-in-plasma-contamination-qc) and [Interpreting eVAF](/dragen-v4.5/product-guides/dragen-v4.5/mrd#interpreting-evaf) for details.


# DNA Somatic Tumor-Normal Solid Panel

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-enable-solid true 
--sv-locus-node-target-file $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# TMB 
--enable-tmb true 
# HLA genotyper 
--enable-hla true 
--hla-as-filter-min-threshold 29.0      #panel specific setting 
--hla-as-filter-ratio-threshold 0.85    #panel specific setting 
--hla-exome true                        #If panel contains only coding regions.  Sets HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #recommended 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true`     | By default, DRAGEN marks duplicate reads and exclude them from variant calling.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--enable-positional-collapsing true` | Alternative to `--enable-duplicate-marking=true`. Instead of discarding duplicate reads, DRAGEN can optionally perform positional collapsing, merging them into higher-quality consensus reads. This is beneficial for small panels without UMIs and coverage between 300X and 1000X. However, it's slower than standard duplicate marking and less effective on samples with coverage lower than 300X. For very high coverage (1000X+), avoid it due to potential read collisions. For high-sensitivity panels with 1000X+ coverage, consider using UMIs. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                 | Description                                                                                                                                                                                                                                                                      |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                                                                                                                                            |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-locus-node-target-file`                      | Specifies a BED file containing a set of regions to call. All SVs with **at least one** breakend within the specified regions will be called.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--sv-target-bed`                                  | (Optional) Specifies a BED file containing the set of regions to call. SVs with **both** breakends within the specified regions will be called. Use this option in place of `sv-locus-node-target-file` for higher specificity.                                                                                                                                                                                                                                                                                                                                                                                                   |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-emit-ref-confidence BP_RESOLUTION 
--vc-enable-vcf-output true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For panels we create GVCF files. Gather the full paths to the small variant caller hard filtered GVCFs (not VCFs) from step 1 and create an input file `${GVCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${GVCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Normal Solid Panel UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
--tumor-normal-has-umi STRING           #Sample(s) containing UMI ['tumor', 'both']. 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-enable-umi-solid true              #>= 1% VAF 
# SV 
--enable-sv true 
--sv-enable-solid true 
--sv-locus-node-target-file $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# TMB 
--enable-tmb true 
# HLA genotyper 
--enable-hla true 
--hla-as-filter-min-threshold 29.0      #panel specific setting 
--hla-as-filter-ratio-threshold 0.85    #panel specific setting 
--hla-exome true                        #If panel contains only coding regions.  Sets HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #recommended 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |
| `--tumor-normal-has-umi STRING`    | Specify if only the tumor, or if both the tumor and normal have UMIs. Options: 'both','tumor'.                                                                                                                                                                                                                                      |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                 | Description                                                                                                                                                                                                                                                                      |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                                                                                                                                            |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-locus-node-target-file`                      | Specifies a BED file containing a set of regions to call. All SVs with **at least one** breakend within the specified regions will be called.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--sv-target-bed`                                  | (Optional) Specifies a BED file containing the set of regions to call. SVs with **both** breakends within the specified regions will be called. Use this option in place of `sv-locus-node-target-file` for higher specificity.                                                                                                                                                                                                                                                                                                                                                                                                   |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-emit-ref-confidence BP_RESOLUTION 
--vc-enable-vcf-output true 
--vc-enable-umi-solid true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For panels we create GVCF files. Gather the full paths to the small variant caller hard filtered GVCFs (not VCFs) from step 1 and create an input file `${GVCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${GVCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Normal Solid WES

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-exome true 
--sv-target-bed $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-use-somatic-vc-baf true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# TMB 
--enable-tmb true 
# HLA genotyper 
--enable-hla true 
--hla-exome true                        #Set HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #recommended 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                 | Description                                                                                                                                                                                                                                                                      |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-enable-cyto-output true`        | Enable Cytogenetics-compatible output (default false).                                                                                                                                                                                                                           |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                                                                                                                                            |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Normal Solid WES UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
--tumor-normal-has-umi STRING           #Sample(s) containing UMI ['tumor', 'both']. 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-exome true 
--sv-target-bed $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-use-somatic-vc-baf true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# TMB 
--enable-tmb true 
# HLA genotyper 
--enable-hla true 
--hla-exome true                        #Set HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #recommended 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |
| `--tumor-normal-has-umi STRING`    | Specify if only the tumor, or if both the tumor and normal have UMIs. Options: 'both','tumor'.                                                                                                                                                                                                                                      |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                 | Description                                                                                                                                                                                                                                                                      |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-enable-cyto-output true`        | Enable Cytogenetics-compatible output (default false).                                                                                                                                                                                                                           |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                                                                                                                                            |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Normal Solid WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-systematic-noise $PATH             #Optional 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# CNV 
--enable-cnv true 
--cnv-use-somatic-vc-baf true 
--cnv-enable-self-normalization true 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# TMB 
--enable-tmb true 
# HLA genotyper 
--enable-hla true 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #recommended 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                 | Description                                                                                                                                               |
| -------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-enable-cyto-output true`        | Enable Cytogenetics-compatible output (default false).                                                                                                    |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                     |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files are recommended for WGS workflows. Prebuilt WGS noise files are available.

#### Prebuilt

Prebuilt WGS SV systematic noise files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WGS noise files                            | Description      |
| --------------------------------------------------- | ---------------- |
| `WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz`      | For WGS, FF/FFPE |
| `IDPF_WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz` | For WGS, HEME    |

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Normal Solid WGS UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
--tumor-normal-has-umi STRING           #Sample(s) containing UMI ['tumor', 'both']. 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-systematic-noise $PATH             #Optional 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# CNV 
--enable-cnv true 
--cnv-use-somatic-vc-baf true 
--cnv-enable-self-normalization true 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# TMB 
--enable-tmb true 
# HLA genotyper 
--enable-hla true 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #recommended 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |
| `--tumor-normal-has-umi STRING`    | Specify if only the tumor, or if both the tumor and normal have UMIs. Options: 'both','tumor'.                                                                                                                                                                                                                                      |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                 | Description                                                                                                                                               |
| -------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-enable-cyto-output true`        | Enable Cytogenetics-compatible output (default false).                                                                                                    |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                     |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--vc-detect-systematic-noise=true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files are recommended for WGS workflows. Prebuilt WGS noise files are available.

#### Prebuilt

Prebuilt WGS SV systematic noise files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WGS noise files                            | Description      |
| --------------------------------------------------- | ---------------- |
| `WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz`      | For WGS, FF/FFPE |
| `IDPF_WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz` | For WGS, HEME    |

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only Heme WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Required 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--heme-sv true 
--sv-systematic-noise $PATH             #Recommended 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# DUX4 
--enable-dux4-caller true 
# CNV 
--heme-cnv true 
--cnv-population-b-allele-vcf $POP_VCF 
--cnv-enable-self-normalization true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
--vc-germline-tag-hotspots false        #When germline tagging is enabled, disable it only for somatic hotspot variants 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 3. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### CNV

| Option                                   | Description                                                                                                                                               |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.    |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default false).                                                                                                    |
| `--heme-cnv true`                        | Configures DRAGEN to use CNV settings for HEME.                                                                                                           |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL)                      |
| ----------------------------------- | ----------------------------------------------------------------------- |
| `--heme-sv true`                    | Configures DRAGEN to use SV settings for Liquid Tumors (e.g., AML/MLL). |
| `--sv-min-scored-variant-size $INT` | 100000                                                                  |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

### DUX4

| Option                          | Description                                                                                               |
| ------------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--dux4-skip-sanity-check true` | Bypass the requirements checks if the input datasets don't comply with parameters listed in prerequisites |

For more information, see [DUX4-rearrangement Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/dux4-rearrangement-caller).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-enable-germline-tagging=true 
--vc-germline-tag-hotspots=false        #When germline tagging is enabled, disable it only for somatic hotspot variants 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files are recommended for WGS workflows. Prebuilt WGS noise files are available.

#### Prebuilt

Prebuilt WGS SV systematic noise files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WGS noise files                            | Description      |
| --------------------------------------------------- | ---------------- |
| `WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz`      | For WGS, FF/FFPE |
| `IDPF_WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz` | For WGS, HEME    |

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).


# DNA Somatic Tumor-Only Solid Amplicon

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# Amplicon 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
--enable-duplicate-marking false        #default=false 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED          #Optional. Auto-generated based on amplicon target bed. 
--vc-systematic-noise $PATH             #optional for SNV systematic noise. 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
# CNV 
--enable-cnv true 
--cnv-combined-counts $PATH             #CNV PON. Required for amplicon CNV calling on CASE samples. 
--cnv-target-bed $PATH                  #Optional. Auto-generated based on amplicon target bed. 
--cnv-filter-qual $NUM                  #CNV filter quality. Adjust CNV filter quality thresholds according to the user’s validation study. 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
--amplicon-enable-msi true 
```

## Notes and additional options

### Pillar Amplicon Specific Settings

To support the varied designs of amplicon panels and the specific requirements of different analysis types (e.g., SNV, CNV, SV, MSI, RNA fusion, RNA splice variants, and RNA 3'/5' imbalance ratio), panel-specific parameter settings have been integrated into the command-line options. Each supported Pillar panel has a dedicated option, and the details for these DNA panels are listed in the table below:

|            **Panel Name**           |     **Short Name**    | **Panel Code** | **Sample Type** | **Default variant caller enabled** |      **Command Line Options**     |
| :---------------------------------: | :-------------------: | :------------: | :-------------: | :--------------------------------: | :-------------------------------: |
|  oncoReveal BRCA1 & BRCA2 plus CNV  |        BRCA CNV       |      BR283     |       DNA       |              SNV, CNV              |     --amplicon-enable-dna-brca    |
|         oncoReveal Lymphoid         |        Lymphoid       |    P-LYM-01    |       DNA       |               SNV, SV              |   --amplicon-enable-dna-lymphoid  |
|       oncoReveal Essential MPN      |          MPN          |       MY7      |       DNA       |                 SNV                |     --amplicon-enable-dna-mpn     |
| oncoReveal Multi-Cancer v4 with CNV | Multi-Cancer with CNV |      HS341     |       DNA       |              SNV, CNV              | --amplicon-enable-dna-multicancer |
|          oncoReveal Myeloid         |        Myeloid        |      MY766     |       DNA       |               SNV, SV              |   --amplicon-enable-dna-myeloid   |
|       oncoReveal Nexus 21 Gene      |         Nexus         |    P-CMC-01    |       DNA       |               SNV, SV              |    --amplicon-enable-dna-nexus    |
|      oncoReveal Solid Tumor v2      |     Solid Tumor v2    |     P-ST-02    |       DNA       |                 SNV                |  --amplicon-enable-dna-solidtumor |

For more detail on the amplicon pipeline, please refer to [DRAGEN Amplicon Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-amplicon-pipeline)

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                     |
| -------------------------------- | ----------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).  |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false). |

### Amplicon post-alignment processing

| Option                                 | Description                                                                                                                                                   |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--amplicon-primer-length INT`         | If an alignment starts inside the primer region of the amplicon target, the alignment is assigned to the amplicon.                                            |
| `--amplicon-allow-partial-target true` | In order to detect deletion events that are close to the target boundaries, we now require only one of the reads to start in the primer region (Default=true) |

For more detail on the amplicon post-alignment processing, please refer to [DRAGEN Amplicon Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-amplicon-pipeline)

### Duplicate Marking

| Option                             | Description                                                                                                                                                                                               |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking false` | The Amplicon Pipeline disables duplicate marking. In amplicon assays, fragments originate from a limited number of unique start and end positions, making conventional duplicate detection inappropriate. |

### SNV

| Option                                      | Description                                                                                                                                                                                                   |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                                                                                  |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                  |
| `--vc-target-vaf $FLOAT`                    | The default is 0.03 (3%)                                                                                                                                                                                      |
| `--vc-af-call-threshold $FLOAT`             | If the AF filter is enabled using --vc-enable-af-filter=true, the option sets the allele frequency call threshold for nuclear chromosomes to emit a call in the VCF. The default value is 0.01.               |
| `--vc-af-filter-threshold $FLOAT`           | If the AF filter is enabled using --vc-enable-af-filter=true, the option sets the allele frequency filter threshold for nuclear chromosomes to mark emitted VCF calls as filtered. The default value is 0.05. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### CNV

Amplicon CNV requires PON input. In PON mode, the DRAGEN CNV Pipeline is broken down into two distinct stages. The target counts stage is performed on each sample (case and normals), to bin the alignments. The normalization and call detection stage is then performed with the case sample against the panel of normals to determine the events.

| Option                                 | Description                                                                                                                                                                                                                                                                                                       |
| -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. By default, `bed` is used for standard panels and `hslm` for Pillar panels with a pre-built PON.                                                                                                                                                           |
| `--amplicon-cnv-use-default-pon false` | We recommend including in-run normal samples—matched in sample type and library preparation—in the same sequencing run to serve as the PON. If generating a custom PON is not feasible, for Pillar panels, the pre-packaged panel-specific PON can be used as a fallback. To enable this, set the option to true. |
| `--cnv-segmentation-bed $PATH`         | You can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. If bed segmentation mode is used, the segmentation bed is auto-generated from amplicon target bed by default                                                                                                |
| `--cnv-filter-qual $NUM`               | QUAL value at which to hard filter CNV VCF. You can adjust CNV filter quality thresholds according to the your validation study                                                                                                                                                                                   |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                     | Specifies a BED file containing the set of regions to call. Default as amplicon target bed.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| `--enable-variant-deduplication true` | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`        | Optional systematic noise BEDPE file containing the set of noisy paired regions.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

DRAGEN has pre-built systematic noise files for Pillar panels. These files are packaged directly with DRAGEN.

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only Solid Panel

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-enable-solid true 
--sv-locus-node-target-file $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# TMB 
--enable-tmb true 
--tmb-enable-proxi-filter true          #Optional for Tumor-Only 
# HLA genotyper 
--enable-hla true 
--hla-as-filter-min-threshold 29.0      #panel specific setting 
--hla-as-filter-ratio-threshold 0.85    #panel specific setting 
--hla-exome true                        #If panel contains only coding regions.  Sets HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true`     | By default, DRAGEN marks duplicate reads and exclude them from variant calling.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--enable-positional-collapsing true` | Alternative to `--enable-duplicate-marking=true`. Instead of discarding duplicate reads, DRAGEN can optionally perform positional collapsing, merging them into higher-quality consensus reads. This is beneficial for small panels without UMIs and coverage between 300X and 1000X. However, it's slower than standard duplicate marking and less effective on samples with coverage lower than 300X. For very high coverage (1000X+), avoid it due to potential read collisions. For high-sensitivity panels with 1000X+ coverage, consider using UMIs. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 3. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                                                                                                                                                      |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.                                                                                                                           |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

The Tumor-Normal pipeline is more effective than the Tumor-Only pipeline at removing or tagging germline variants. The Tumor-Only may subsequently report somewhat elevated TMB values. The TMB proxi filter is an optional setting on top of the regular database germline filter. It will aggressively filter additional germline variants based on allele frequencies.

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-locus-node-target-file`                      | Specifies a BED file containing a set of regions to call. All SVs with **at least one** breakend within the specified regions will be called.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--sv-target-bed`                                  | (Optional) Specifies a BED file containing the set of regions to call. SVs with **both** breakends within the specified regions will be called. Use this option in place of `sv-locus-node-target-file` for higher specificity.                                                                                                                                                                                                                                                                                                                                                                                                   |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-emit-ref-confidence BP_RESOLUTION 
--vc-enable-vcf-output true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For panels we create GVCF files. Gather the full paths to the small variant caller hard filtered GVCFs (not VCFs) from step 1 and create an input file `${GVCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${GVCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only Solid Panel UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-enable-umi-solid true              #>= 1% VAF 
# SV 
--enable-sv true 
--sv-enable-solid true 
--sv-locus-node-target-file $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# TMB 
--enable-tmb true 
--tmb-enable-proxi-filter true          #Optional for Tumor-Only 
# HLA genotyper 
--enable-hla true 
--hla-as-filter-min-threshold 29.0      #panel specific setting 
--hla-as-filter-ratio-threshold 0.85    #panel specific setting 
--hla-exome true                        #If panel contains only coding regions.  Sets HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 4. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                                                                                                                                                      |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.                                                                                                                           |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

The Tumor-Normal pipeline is more effective than the Tumor-Only pipeline at removing or tagging germline variants. The Tumor-Only may subsequently report somewhat elevated TMB values. The TMB proxi filter is an optional setting on top of the regular database germline filter. It will aggressively filter additional germline variants based on allele frequencies.

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-locus-node-target-file`                      | Specifies a BED file containing a set of regions to call. All SVs with **at least one** breakend within the specified regions will be called.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--sv-target-bed`                                  | (Optional) Specifies a BED file containing the set of regions to call. SVs with **both** breakends within the specified regions will be called. Use this option in place of `sv-locus-node-target-file` for higher specificity.                                                                                                                                                                                                                                                                                                                                                                                                   |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-emit-ref-confidence BP_RESOLUTION 
--vc-enable-vcf-output true 
--vc-enable-umi-solid true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For panels we create GVCF files. Gather the full paths to the small variant caller hard filtered GVCFs (not VCFs) from step 1 and create an input file `${GVCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${GVCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only Solid WES

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-exome true 
--sv-target-bed $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# TMB 
--enable-tmb true 
--tmb-enable-proxi-filter true          #Optional for Tumor-Only 
# HLA genotyper 
--enable-hla true 
--hla-exome true                        #Set HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 3. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                                                                                                                                                      |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.                                                                                                                           |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default false).                                                                                                                                                                                                                           |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

The Tumor-Normal pipeline is more effective than the Tumor-Only pipeline at removing or tagging germline variants. The Tumor-Only may subsequently report somewhat elevated TMB values. The TMB proxi filter is an optional setting on top of the regular database germline filter. It will aggressively filter additional germline variants based on allele frequencies.

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only Solid WES UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-exome true 
--sv-target-bed $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# TMB 
--enable-tmb true 
--tmb-enable-proxi-filter true          #Optional for Tumor-Only 
# HLA genotyper 
--enable-hla true 
--hla-exome true                        #Set HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 4. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                                                                                                                                                      |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.                                                                                                                           |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default false).                                                                                                                                                                                                                           |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

The Tumor-Normal pipeline is more effective than the Tumor-Only pipeline at removing or tagging germline variants. The Tumor-Only may subsequently report somewhat elevated TMB values. The TMB proxi filter is an optional setting on top of the regular database germline filter. It will aggressively filter additional germline variants based on allele frequencies.

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only Solid WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-systematic-noise $PATH             #Recommended 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF 
--cnv-enable-self-normalization true 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# TMB 
--enable-tmb true 
--tmb-enable-proxi-filter true          #Optional for Tumor-Only 
# HLA genotyper 
--enable-hla true 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 3. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                               |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.    |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default false).                                                                                                    |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

The Tumor-Normal pipeline is more effective than the Tumor-Only pipeline at removing or tagging germline variants. The Tumor-Only may subsequently report somewhat elevated TMB values. The TMB proxi filter is an optional setting on top of the regular database germline filter. It will aggressively filter additional germline variants based on allele frequencies.

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files are recommended for WGS workflows. Prebuilt WGS noise files are available.

#### Prebuilt

Prebuilt WGS SV systematic noise files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WGS noise files                            | Description      |
| --------------------------------------------------- | ---------------- |
| `WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz`      | For WGS, FF/FFPE |
| `IDPF_WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz` | For WGS, HEME    |

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only Solid WGS UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-target-vaf $NUM                    #Default = 0.03 (>= 3% VAF) 
# SV 
--enable-sv true 
--sv-systematic-noise $PATH             #Recommended 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF 
--cnv-enable-self-normalization true 
# HRD Scoring 
--enable-hrd true                       #requires CNV 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# TMB 
--enable-tmb true 
--tmb-enable-proxi-filter true          #Optional for Tumor-Only 
# HLA genotyper 
--enable-hla true 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 4. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                               |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.    |
| `--cnv-enable-cyto-output true`          | Enable Cytogenetics-compatible output (default false).                                                                                                    |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

The Tumor-Normal pipeline is more effective than the Tumor-Only pipeline at removing or tagging germline variants. The Tumor-Only may subsequently report somewhat elevated TMB values. The TMB proxi filter is an optional setting on top of the regular database germline filter. It will aggressively filter additional germline variants based on allele frequencies.

| Option                              | Description                                                                                                                                                               |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`         | Variant minimum allele frequency for usable variants (default=0.05)                                                                                                       |
| `--vc-callability-tumor-thresh INT` | Required read coverage to use a site (default=50).                                                                                                                        |
| `--tmb-enable-proxi-filter BOOL`    | Use variant vaf information to increase germline filtering. Recommended for TO, but not for TN. May be overly aggressive at tagging variants as germline (default=false). |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--vc-detect-systematic-noise=true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files are recommended for WGS workflows. Prebuilt WGS noise files are available.

#### Prebuilt

Prebuilt WGS SV systematic noise files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WGS noise files                            | Description      |
| --------------------------------------------------- | ---------------- |
| `WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz`      | For WGS, FF/FFPE |
| `IDPF_WGS_hg38_v3.0.0_systematic_noise.sv.bedpe.gz` | For WGS, HEME    |

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only ctDNA Amplicon

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# Amplicon 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
--enable-duplicate-marking false        #default=false 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED          #Optional. Auto-generated based on amplicon target bed. 
--vc-systematic-noise $PATH             #optional for SNV systematic noise. 
--vc-target-vaf $NUM                    #Default = 0.001 (>= 0.1% VAF) 
# SV 
--enable-sv true 
# CNV 
--enable-cnv true 
--cnv-combined-counts $PATH             #CNV PON. Required for amplicon CNV calling on CASE samples. 
--cnv-target-bed $PATH                  #Optional. Auto-generated based on amplicon target bed. 
--cnv-filter-qual $NUM                  #CNV filter quality. Adjust CNV filter quality thresholds according to the user’s validation study. 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
--amplicon-enable-msi true 
```

## Notes and additional options

### Pillar Amplicon Specific Settings

To support the varied designs of amplicon panels and the specific requirements of different analysis types (e.g., SNV, CNV, SV, MSI, RNA fusion, RNA splice variants, and RNA 3'/5' imbalance ratio), panel-specific parameter settings have been integrated into the command-line options. Each supported Pillar panel has a dedicated option, and the details for these DNA panels are listed in the table below:

|      **Panel Name**      | **Short Name** | **Panel Code** | **Sample Type** | **Default variant caller enabled** |      **Command Line Options**     |
| :----------------------: | :------------: | :------------: | :-------------: | :--------------------------------: | :-------------------------------: |
|    oncoReveal Core LBx   |    Core LBx    |    P-LBX-01    |      cfDNA      |            SNV, CNV, MSI           |    --amplicon-enable-cfdna-core   |
| oncoReveal Essential LBx |  Essential LBx |    P-LBX-04    |      cfDNA      |            SNV, CNV, MSI           | --amplicon-enable-cfdna-essential |

For more detail on the amplicon pipeline, please refer to [DRAGEN Amplicon Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-amplicon-pipeline)

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                     |
| -------------------------------- | ----------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).  |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false). |

### Amplicon post-alignment processing

| Option                                 | Description                                                                                                                                                   |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--amplicon-primer-length INT`         | If an alignment starts inside the primer region of the amplicon target, the alignment is assigned to the amplicon.                                            |
| `--amplicon-allow-partial-target true` | In order to detect deletion events that are close to the target boundaries, we now require only one of the reads to start in the primer region (Default=true) |

For more detail on the amplicon post-alignment processing, please refer to [DRAGEN Amplicon Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-amplicon-pipeline)

### Duplicate Marking

| Option                             | Description                                                                                                                                                                                               |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking false` | The Amplicon Pipeline disables duplicate marking. In amplicon assays, fragments originate from a limited number of unique start and end positions, making conventional duplicate detection inappropriate. |

### SNV

| Option                                      | Description                                                                                                                                                                                                         |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                                                                                        |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                        |
| `--vc-target-vaf $FLOAT`                    | For ctDNA, the default is 0.001 (0.1%).                                                                                                                                                                             |
| `--vc-af-call-threshold $FLOAT`             | If the AF filter is enabled using --vc-enable-af-filter=true, the option sets the allele frequency call threshold for nuclear chromosomes to emit a call in the VCF. For ctDNA, the default is 0.001.               |
| `--vc-af-filter-threshold $FLOAT`           | If the AF filter is enabled using --vc-enable-af-filter=true, the option sets the allele frequency filter threshold for nuclear chromosomes to mark emitted VCF calls as filtered. For ctDNA, the default is 0.003. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### CNV

Amplicon CNV requires PON input. In PON mode, the DRAGEN CNV Pipeline is broken down into two distinct stages. The target counts stage is performed on each sample (case and normals), to bin the alignments. The normalization and call detection stage is then performed with the case sample against the panel of normals to determine the events.

| Option                                 | Description                                                                                                                                                                                                                                                                                                       |
| -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. By default, `bed` is used for standard panels and `hslm` for Pillar panels with a pre-built PON.                                                                                                                                                           |
| `--amplicon-cnv-use-default-pon false` | We recommend including in-run normal samples—matched in sample type and library preparation—in the same sequencing run to serve as the PON. If generating a custom PON is not feasible, for Pillar panels, the pre-packaged panel-specific PON can be used as a fallback. To enable this, set the option to true. |
| `--cnv-segmentation-bed $PATH`         | You can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. If bed segmentation mode is used, the segmentation bed is auto-generated from amplicon target bed by default                                                                                                |
| `--cnv-filter-qual $NUM`               | QUAL value at which to hard filter CNV VCF. You can adjust CNV filter quality thresholds according to the your validation study                                                                                                                                                                                   |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                     | Specifies a BED file containing the set of regions to call. Default as amplicon target bed.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| `--enable-variant-deduplication true` | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`        | Optional systematic noise BEDPE file containing the set of noisy paired regions.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

DRAGEN has pre-built systematic noise files for Pillar panels. These files are packaged directly with DRAGEN.

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--vc-detect-systematic-noise=true 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For WES and WGS pipelines gather the full paths to the small variant hard filtered VCFs (not GVCFs) from step 1 and create a lines file `${VCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${VCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--enable-dna-amplicon true 
--amplicon-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# DNA Somatic Tumor-Only ctDNA Panel UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
--umi-source STRING                     #Default='qname' 
--umi-library-type STRING               #e.g. random-duplex 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-enable-umi-liquid true             #>= 0.1% VAF 
# SV 
--enable-sv true 
--sv-enable-liquid true 
--sv-locus-node-target-file $SV_TARGET_BED 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
--cnv-combined-counts $PATH             #CNV PON 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
# TMB 
--enable-tmb true 
--vc-callability-tumor-thresh 1000 
--tmb-vaf-threshold 0.002 
--tmb-enable-proxi-filter true          #Optional for Tumor-Only 
# HLA genotyper 
--enable-hla true 
--hla-as-filter-min-threshold 29.0      #panel specific setting 
--hla-as-filter-ratio-threshold 0.85    #panel specific setting 
--hla-exome true                        #If panel contains only coding regions.  Sets HLA output to 3 field resolution (coding sequence only) 
# Microsatellite Instability (MSI) 
--enable-msi true 
--msi-microsatellites-file $PATH 
--msi-ref-normal-input $PATH            #required 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-source STRING`              | Specify the input type for the UMI sequence. Options: `qname`, `fastq`, `bamtag`.                                                                                                                                                                                                                                                   |
| `--umi-library-type STRING`        | Set the batch option for different UMIs correction. Options: `random-duplex`, `random-simplex`, `nonrandom-duplex`.                                                                                                                                                                                                                 |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-emit-ref-confidence GVCF`                          | To enable gVCF output.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--vc-enable-vcf-output`                                 | To enable VCF file output during a gVCF run, set to true. The default value is false.                                                                                                                                                                                                                                                                                                                                                             |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 2. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### HLA

| Option                            | Description                                                                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-hla`                    | Enable HLA typer (this setting by default will only genotype class 1 genes)                                                     |
| `--hla-as-filter-min-threshold`   | Internal option to set min alignment score threshold. The default is 59 and works for WES and WGS. Set to 29 for panels.        |
| `--hla-as-filter-ratio-threshold` | Minimum Alignment score of a read mate to be considered. The default is 0.67 and works for WES and WES. Set to 0.85 for panels. |
| `--hla-enable-class-2`            | Extend genotyping to HLA class 2 genes (default=true).                                                                          |
| `--hla-exome`                     | Output HLA alleles at 3-field resolution (default 4-field resolution) for panels targeting only coding sequences or WES.        |

### CNV

| Option                                   | Description                                                                                                                                                                                                                                                                      |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`    | Enable or disable GC bias correction when generating target counts.                                                                                                                                                                                                              |
| `--cnv-segmentation-mode $SEG_MODE`      | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis.                                                                                                                        |
| `--cnv-segmentation-bed $PATH`           | Specify a segmentation bed file to add pre-defined segments to be called. If you are using somatic targeted panels with a set of genes supplied with the capture kit, then you can bypass segmentation by specifying a cnv-segmentation-bed and using cnv-segmentation-mode=bed. |
| `--cnv-population-b-allele-vcf $POP_VCF` | Specify a population SNP VCF. This option can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.                                                                                                                           |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### TMB

The Tumor-Normal pipeline is more effective than the Tumor-Only pipeline at removing or tagging germline variants. The Tumor-Only may subsequently report somewhat elevated TMB values. The TMB proxi filter is an optional setting on top of the regular database germline filter. It will aggressively filter additional germline variants based on allele frequencies.

| Option                          | Description                                                                                 |
| ------------------------------- | ------------------------------------------------------------------------------------------- |
| `--tmb-vaf-threshold FLOAT`     | Variant minimum allele frequency for usable variants. Default=0.05. Set to 0.002 for ctDNA. |
| `--vc-callability-tumor-thresh` | Required read coverage to use a site. Default=50. Set to 1000 for ctDNA.                    |

See the user guide: [TMB Germline Variants](https://help.dragen.illumina.com/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/pages/gNnLZydItf4mjzOt19Ly#id-4.-support-for-germline-variants).

### MSI

Microsatellite sites and PON files can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For more information on MSI calling, see [MSI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-locus-node-target-file`                      | Specifies a BED file containing a set of regions to call. All SVs with **at least one** breakend within the specified regions will be called.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--sv-target-bed`                                  | (Optional) Specifies a BED file containing the set of regions to call. SVs with **both** breakends within the specified regions will be called. Use this option in place of `sv-locus-node-target-file` for higher specificity.                                                                                                                                                                                                                                                                                                                                                                                                   |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. For high sensitivity applications, including panels or clinical WES/WGS assays, it is recommended to create your own systematic noise file as described under Custom.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |

#### Custom

This section describes how to generate systematic noise files from phenotypically normal (non-tumor) samples to optimize the performance of a specific assay. For best accuracy, the normal samples should ideally closely match the sequencer, sample type, library prep, and coverage of the tumor samples of interest. It is typically recommended to use 30-70 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on each of approximately 30-70 normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--vc-detect-systematic-noise=true 
--vc-target-bed $VC_TARGET_BED          #Region assessed in assay 
--vc-target-bed-padding 500 
--vc-emit-ref-confidence BP_RESOLUTION 
--vc-enable-vcf-output true 
--vc-enable-umi-liquid true 
--vc-enable-germline-tagging=true 
--variant-annotation-data $NIRVANA_PATH 
--intermediate-results-dir $PATH 
--output-directory $PATH 
--output-file-prefix $STRING 
```

For panels we create GVCF files. Gather the full paths to the small variant caller hard filtered GVCFs (not VCFs) from step 1 and create an input file `${GVCF_LIST}` by specifying 1 file per line.

**Step 2. Generate the final noise file.**

This step generates a bed file containing mean and max noise estimates per position. This can be used directly during variant calling (argument --vc-systematic-noise). The distribution of noise per position can also be plotted to identify particularly noisy positions that could be troubleshooted (e.g. modify assay settings or DRAGEN settings) or blocklisted

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--build-sys-noise-vcfs-list ${GVCF_LIST} 
```

The SNV systematic noise files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### SV Systematic Noise

SV systematic noise files have not been tested with WES, enrichment and amplicon panels. It is considered an experimental mode for these assays.

#### Custom

Custom systematic noise files can be generated for WES, Panels or Amplicon. For best accuracy the normal samples should ideally closely match the sequencer, sample type, library prep and coverage of the tumor samples of interest. It is typically recommended to use 30 - 100 normals when building a noise file, but fewer can be used.

**Step 1. Run DRAGEN somatic tumor-only on normal samples with `--sv-detect-systematic-noise` set to true to generate VCF output per normal sample.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--umi-enable true 
--umi-source STRING                     #default='qname' 
--umi-library-type STRING               #see 'UMI' 
--sv-detect-systematic-noise true 
```

**Step 2. Build the BEDPE file using input VCFs from previous step.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--sv-build-systematic-noise-vcfs-list $VCF_LIST#one VCF per line. 
```

Systematic noise BEDPE files can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

### CNV Panel of Normals (PON)

For CNV PON requirements and generation options see [CNV Preprocessing | Panel of Normals](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#panel-of-normals).

If a matched normal is available it is recommended to include it in the PON.

**Step 1. Generate CNV target counts of individual normal samples.**

Any samples that should not be included in the final PON file can be excluded from this step. Any options used for CNV target counts generation (BED file, GC Bias Correction, etc.) should be matched when processing the case samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# CNV 
--enable-cnv true 
--cnv-target-bed $PATH 
```

**Step 2. CNV combined counts file generation.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--enable-cnv true 
--cnv-generate-combined-counts true 
--cnv-normals-list $CNV_NORMALS_LIST 
```

`$CNV_NORMALS_LIST` is a text file with one line for each path to a CNV target counts file generated in step 1 (either `<output-file-prefix>.target.counts.gz` or `<output-file-prefix>.target.counts.gc-corrected.gz`). Individual target counts files are merged into a single `<output-file-prefix>.combined.counts.txt.gz` PON file in the output directory. The PON file is used for each case sample run of DRAGEN CNV using the `--cnv-combined-counts` option.

### MSI baseline file (PON)

It is recommended to use a baseline file that matches the sample type (FF/FFPE), assay type (WGS/WES/Panel) and genome build (hg19/hg38) of the samples being analyzed. If matched normals are available, it is recommended to include them in the baselines.

For MSI baseline generation, see [Baseline microsatellite repeat distribution](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/biomarkers/biomarker-msi#baseline-microsatellite-repeat-distribution). It is required that non-tumor inputs are used for baseline generation.

**Step 1. Generate MSI baselines for individual normal samples.**

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# MSI 
--enable-msi true 
--msi-generate-baseline true 
--msi-microsatellites-file $MICROSATELLITE_LIST#see 'Product Files' for available files 
```

**Step 2. Combine MSI baselines to a single file.**

```
touch combined_msi_baseline.dist 
sed 's/$/ $SAMPLEID $OUTPUT/$PREFIX.microsat_normal.dist | grep -v '#' >> combined_msi_baseline.dist 
```

The combined MSI baseline files can then be used for each case sample run of DRAGEN MSI using `--msi-ref-normal-input combined_msi_baseline.dist`.


# Illumina TruPath Genome WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' 
--fastq-list-sample-id $STRING 
# Illumina TruPath Genome 
--enable-proximity true 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-phasing-min-fragments 0 
# SV 
--enable-sv true 
# CNV 
--enable-cnv true 
--cnv-enable-self-normalization true 
# Targeted caller 
--enable-targeted true 
# Short tandem repeats 
--repeat-genotype-enable true 
# Multi-Region Joint Detection (MRJD) 
--enable-mrjd true 
```

## Notes and additional options

### Hashtable

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

TruPath Genome Prep only supports fastq list and fastq input. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```


# Illumina scRNA

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-rna true 
--annotation-file $GTF                  #GTF or GFF3 format 
--enable-map-align true                 #required for RNA/scRNA 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# Illumina Single Cell Prep 
--scrna-enable-pipseq-mode true 
--single-cell-threshold ratio           #['fixed', 'ratio', inflection'] 
```

## Notes and additional options

### Hashtable

For DRAGEN RNA/scRNA runs, it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                    |
| -------------------------------- | -------------------------------------------------------------- |
| `--enable-map-align true`        | For RNA/scRNA pipelines, map-align should always be turned on. |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                |

### Illumina single-cell RNA Prep PIPseq options

PIPseq mode batch option automatically sets the barcode/BI source, the barcode and binning index positions and the barcode sequence list options.

By default the barcode/BI is obtained from read 1 and the transcript is obtained from read 2.

To change the barcode or binning index positions, use `--scrna-barcode-position` and `--scrna-umi-position`. These settings should be provided in the form `<startPos>_<endPos>` for each barcode. Connect multiple barcode sequence positions with a '+'.

For example, a library with the cell-barcode split into three blocks of 9 bp separated by fixed linker sequences and an 8 bp BI would be set using: `--scrna-barcode-position 0_8+21_29+43_51`, and `--scrna-umi-position 52_59`.

The following table lists some optional settings:

| Option                          | Description                                                                                                                                       |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--scrna-enable-pipseq-mode`    | Option to enable Illumina PIPseq mode.                                                                                                            |
| `--scrna-barcode-position`      | See example above or refer to [Illumina scRNA PIPseq](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina). |
| `--scrna-umi-position`          | See example above or refer to [Illumina scRNA PIPseq](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina). |
| `--single-cell-threshold`       | Cell filtering can be set to \['fixed', 'ratio', or 'inflection'].                                                                                |
| `--scrna-barcode-sequence-list` | A known barcode sequence list can be optionally provided.                                                                                         |
| `--umi-source`                  | Optionally override the default barcode/BI source, valid option include \['read1', 'read2', 'qname', 'fastq'].                                    |

For more details on Illumina single cell prep pipeline options, refer to the [Illumina scRNA PIPseq Pipeline User Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina)


# Illumina scRNA Perturb seq

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-rna true 
--annotation-file $GTF                  #GTF or GFF3 format 
--enable-map-align true                 #required for RNA/scRNA 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# Illumina Single Cell Perturb-seq 
--scrna-enable-pipseq-mode true 
--single-cell-threshold ratio           #['fixed', 'ratio', inflection'] 
--scrna-enable-pipseq-crispr-mode true 
--scrna-feature-barcode-groups $FEATURE_RGIDS 
--scrna-feature-barcode-reference $FEATURE_REF_CSV 
```

## Notes and additional options

### Hashtable

For DRAGEN RNA/scRNA runs, it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                    |
| -------------------------------- | -------------------------------------------------------------- |
| `--enable-map-align true`        | For RNA/scRNA pipelines, map-align should always be turned on. |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                |

### Illumina single-cell RNA Prep Perturb-seq options

PIPseq mode batch option automatically sets the barcode/BI source, the barcode and binning index positions and the barcode sequence list options.

By default the barcode/BI is obtained from read 1 and the transcript is obtained from read 2.

PIPseq CRISPR mode option configures DRAGEN to automatically process CRISPR feature reads.

CRISPR mode enables feature counting, transforming of CRISPR cell-barcodes to match transcript barcodes, guide RNA calling, and additional metrics.

CRISPR mode requires a feature barcode reference CSV file to be provided with the CRISPR 'hook' and 'grab' sequences and positions. Refer to [Illumina scRNA CRISPR Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#crispr-mode) for details.

CRISPR mode also requires the feature barcode groups to be provided, i.e. the RGIDs of the CRISPR input FASTQs (multiple read groups should be comma-separated).

To change the barcode or binning index positions, use `--scrna-barcode-position` and `--scrna-umi-position`. These settings should be provided in the form `<startPos>_<endPos>` for each barcode. Connect multiple barcode sequence positions with a '+'.

For example, a library with the cell-barcode split into three blocks of 9 bp separated by fixed linker sequences and an 8 bp BI would be set using: `--scrna-barcode-position 0_8+21_29+43_51`, and `--scrna-umi-position 52_59`.

The following table lists some optional settings:

| Option                              | Description                                                                                                                                                                                    |
| ----------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--scrna-enable-pipseq-mode`        | Option to enable Illumina PIPseq mode.                                                                                                                                                         |
| `--scrna-barcode-position`          | See example above or refer to [Illumina scRNA PIPseq](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina).                                              |
| `--scrna-umi-position`              | See example above or refer to [Illumina scRNA PIPseq](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina).                                              |
| `--scrna-enable-pipseq-crispr-mode` | Option to enable Illumina PIPseq CRISPR mode.                                                                                                                                                  |
| `--scrna-feature-barcode-groups`    | Comma-separated list of CRISPR read groups (RGID strings), required in PIPseq CRISPR mode.                                                                                                     |
| `--scrna-feature-barcode-reference` | CSV-file of feature barcode reference sequences. Refer to [Illumina scRNA CRISPR Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#crispr-mode). |
| `--single-cell-threshold`           | Cell filtering can be set to \['fixed', 'ratio', or 'inflection'].                                                                                                                             |
| `--scrna-barcode-sequence-list`     | A known barcode sequence list can be optionally provided.                                                                                                                                      |
| `--umi-source`                      | Optionally override the default barcode/BI source, valid option include \['read1', 'read2', 'qname', 'fastq'].                                                                                 |

For more details on Illumina single cell perturb-seq pipeline options, refer to the [Illumina scRNA CRISPR Mode User Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-illumina#crispr-mode)


# Other scRNA prep

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-rna true 
--annotation-file $GTF                  #GTF or GFF3 format 
--enable-map-align true                 #required for RNA/scRNA 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# Other Single Cell Prep 
--enable-single-cell-rna true 
--umi-source qname                      #default='qname' 
--scrna-barcode-position $BARCODE_POS   #see notes 
--scrna-umi-position $UMI_POS           #see notes 
--scrna-barcode-sequence-list $PATH 
--single-cell-threshold ratio           #['fixed', 'ratio', inflection'] 
--single-cell-threshold-filterby umi    #['umi', 'read'] 
```

## Notes and additional options

### Hashtable

For DRAGEN RNA/scRNA runs, it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                    |
| -------------------------------- | -------------------------------------------------------------- |
| `--enable-map-align true`        | For RNA/scRNA pipelines, map-align should always be turned on. |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                |

### Other single-cell RNA Prep options

To change the barcode or binning index positions, use `--scrna-barcode-position` and `--scrna-umi-position`. These settings should be provided in the form `<startPos>_<endPos>` for each barcode. Connect multiple barcode sequence positions with a '+'.

For example, a library with the cell-barcode split into three blocks of 9 bp separated by fixed linker sequences and an 8 bp BI would be set using: `--scrna-barcode-position 0_8+21_29+43_51`, and `--scrna-umi-position 52_59`.

The following table lists some optional settings:

| Option                          | Description                                                                                                                    |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| `--enable-single-cell-rna true` | Option to enable single-cell rna mode.                                                                                         |
| `--scrna-barcode-position`      | See example above or refer to [scRNA](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-other). |
| `--scrna-umi-position`          | See example above or refer to [scRNA](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-other). |
| `--single-cell-threshold`       | Cell filtering can be set to \['fixed', 'ratio', or 'inflection'].                                                             |
| `--scrna-barcode-sequence-list` | A known barcode sequence list can be optionally provided.                                                                      |
| `--umi-source`                  | Optionally override the default barcode/BI source, valid option include \['read1', 'read2', 'qname', 'fastq'].                 |

For more details on single-cell RNA options, refer to the [DRAGEN Single-Cell RNA User Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-single-cell-pipeline/dragen-scrna-other).


# RNA Amplicon

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-rna true 
--annotation-file $GTF                  #GTF or GFF3 format 
--enable-map-align true                 #required for RNA/scRNA 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# Amplicon 
--enable-rna-amplicon true 
--amplicon-target-bed $PATH 
--amplicon-enable-imbalance-ratio true  #optional for panels that support 3'/5' imbalance-ratio 
--enable-duplicate-marking false        #default=false 
# RNA Splice Variants 
--enable-rna-splice-variant true 
# RNA Gene Fusions 
--enable-rna-gene-fusion true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
```

## Notes and additional options

### Hashtable

For DRAGEN RNA/scRNA runs, it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### RNA Amplicon

To enable RNA amplicon, set:

* `--enable-rna-amplicon true`, and
* `--amplicon-target-bed $PATH`.

If RNA amplicon mode is enabled and the amplicon bed file already includes the gene name, then it is not required to set the enriched option (rna-enriched-regions or rna-enriched-genes), since DRAGEN will read the enriched genes names from the amplicon BED file (fifth column).

#### Pillar Specific Amplicon Settings

To support the varied designs of amplicon panels and the specific requirements of different analysis types (e.g., SNV, CNV, SV, MSI, RNA fusion, RNA splice variants, and RNA 3'/5' imbalance ratio), panel-specific parameter settings have been integrated into the command-line options. Each supported Pillar panel has a dedicated option, and the details for these RNA panels are listed in the table below:

|             **Panel Name**            |      **Short Name**      | **Panel Code** | **Sample Type** |             **Default variant caller enabled**            |      **Command Line Options**     |
| :-----------------------------------: | :----------------------: | :------------: | :-------------: | :-------------------------------------------------------: | :-------------------------------: |
|            oncoReveal Heme            |           Heme           |    P-HFU-01    |       RNA       |                         RNA fusion                        |     --amplicon-enable-rna-heme    |
|         oncoReveal Fusion LBx         |        Fusion LBx        |    P-LBX-03    |      cfRNA      |               RNA fusion, RNA splice-variant              | --amplicon-enable-cfrna-lbxfusion |
| oncoReveal Multi-Cancer RNA Fusion v2 | Multi-Cancer with Fusion |      SF-V2     |       RNA       | RNA fusion, RNA splice-variant, RNA 3'/5' imbalance-ratio | --amplicon-enable-rna-multicancer |

For more detail on the amplicon pipeline, please refer to [DRAGEN Amplicon Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-amplicon-pipeline)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                    |
| -------------------------------- | -------------------------------------------------------------- |
| `--enable-map-align true`        | For RNA/scRNA pipelines, map-align should always be turned on. |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                |

### Duplicate Marking

| Option                             | Description                                                                                                                                                                                               |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking false` | The Amplicon Pipeline disables duplicate marking. In amplicon assays, fragments originate from a limited number of unique start and end positions, making conventional duplicate detection inappropriate. |

### RNA Variant Calling

| Option                  | Description                                                                                                                                                                                               |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed $PATH` | Restrict the variants called to a target bed. A bed file specifying the gene-coding regions should be provided to avoid calling erroneous variants in unenriched or noncoding regions due to noisy reads. |

### RNA Splice

| Option                                    | Description                                                                                                                                                                                                                                                                                                                       |
| ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--rna-splice-variant-normals $PATH`      | Optional list of normal splice variants that will be used to filter false positive calls. The file should be a tab separated file with the following first four columns: (1) contig name, (2) first base of the splice junction (1-based), (3) last base of the splice junction (1-based), (4) strand (0: undefined, 1: +, 2: -). |
| `--rna-splice-variant-knowns $PATH`       | File with a list of expected splice junctions, these will be passed                                                                                                                                                                                                                                                               |
| `--rna-splice-variant-fusion-genes $PATH` | List of hotspot genes that may contain spliced fusions                                                                                                                                                                                                                                                                            |

### RNA Fusion

| Option                            | Description                                                                                                              |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--rna-gf-restrict-genes`         | Ignore genes that are not protein coding for gene fusions (Default=true)                                                 |
| `--rna-gf-enable-post-filters`    | Enable stringent post-filtering of RNA gene fusion candidates by quality flags to reduce false positives (Default=false) |
| `--rna-gf-output-fusion-sequence` | Output assembled gene fusion sequence in fusion\_candidates.final file (Default=true)                                    |
| `--rna-gf-report-intronic`        | Report fusion calls with intronic breakpoints (Default=true)                                                             |
| `--rna-gf-report-antisense`       | Report fusion calls with antisense and intronic breakpoints (Default=false)                                              |
| `--rna-gf-report-intergenic`      | Report fusion calls with intergenic, antisense and intronic breakpoints (Default=false)                                  |
| `--rna-gf-report-read-through`    | Report read-through fusion calls (Default=false)                                                                         |
| `--rna-gf-ptd-genes`              | List of gene names where we allow PTD/ITD associated self fusions                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)


# RNA Panel

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-rna true 
--annotation-file $GTF                  #GTF or GFF3 format 
--enable-map-align true                 #required for RNA/scRNA 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--rna-enriched-genes $PATH              #for RNA panels 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# RNA Quantification 
--enable-rna-quantification true 
--rna-library-type A                    #see 'RNA Quant' 
--rna-quantification-gc-bias true 
# RNA Splice Variants 
--enable-rna-splice-variant true 
# RNA Gene Fusions 
--enable-rna-gene-fusion true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
```

## Notes and additional options

### Hashtable

For DRAGEN RNA/scRNA runs, it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                    |
| -------------------------------- | -------------------------------------------------------------- |
| `--enable-map-align true`        | For RNA/scRNA pipelines, map-align should always be turned on. |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                |

### Duplicate Marking

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true`     | By default, DRAGEN marks duplicate reads and exclude them from variant calling.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--enable-positional-collapsing true` | Alternative to `--enable-duplicate-marking=true`. Instead of discarding duplicate reads, DRAGEN can optionally perform positional collapsing, merging them into higher-quality consensus reads. This is beneficial for small panels without UMIs and coverage between 300X and 1000X. However, it's slower than standard duplicate marking and less effective on samples with coverage lower than 300X. For very high coverage (1000X+), avoid it due to potential read collisions. For high-sensitivity panels with 1000X+ coverage, consider using UMIs. |

### RNA Variant Calling

| Option                  | Description                                                                                                                                                                                               |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed $PATH` | Restrict the variants called to a target bed. A bed file specifying the gene-coding regions should be provided to avoid calling erroneous variants in unenriched or noncoding regions due to noisy reads. |

### RNA Quant

| Option               | Description                                                                                                                                                              |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `--rna-library-type` | Set the library according to the read orientations. Set to 'A' to auto detect the correct read orientation. Alternatively select 'IU', 'ISR', 'ISF', 'U', 'SR', or 'SF'. |

### RNA Splice

| Option                                    | Description                                                                                                                                                                                                                                                                                                                       |
| ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--rna-splice-variant-normals $PATH`      | Optional list of normal splice variants that will be used to filter false positive calls. The file should be a tab separated file with the following first four columns: (1) contig name, (2) first base of the splice junction (1-based), (3) last base of the splice junction (1-based), (4) strand (0: undefined, 1: +, 2: -). |
| `--rna-splice-variant-knowns $PATH`       | File with a list of expected splice junctions, these will be passed                                                                                                                                                                                                                                                               |
| `--rna-splice-variant-fusion-genes $PATH` | List of hotspot genes that may contain spliced fusions                                                                                                                                                                                                                                                                            |

### RNA Fusion

| Option                            | Description                                                                                                              |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--rna-gf-restrict-genes`         | Ignore genes that are not protein coding for gene fusions (Default=true)                                                 |
| `--rna-gf-enable-post-filters`    | Enable stringent post-filtering of RNA gene fusion candidates by quality flags to reduce false positives (Default=false) |
| `--rna-gf-output-fusion-sequence` | Output assembled gene fusion sequence in fusion\_candidates.final file (Default=true)                                    |
| `--rna-gf-report-intronic`        | Report fusion calls with intronic breakpoints (Default=true)                                                             |
| `--rna-gf-report-antisense`       | Report fusion calls with antisense and intronic breakpoints (Default=false)                                              |
| `--rna-gf-report-intergenic`      | Report fusion calls with intergenic, antisense and intronic breakpoints (Default=false)                                  |
| `--rna-gf-report-read-through`    | Report read-through fusion calls (Default=false)                                                                         |
| `--rna-gf-ptd-genes`              | List of gene names where we allow PTD/ITD associated self fusions                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)


# RNA WTS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-rna true 
--annotation-file $GTF                  #GTF or GFF3 format 
--enable-map-align true                 #required for RNA/scRNA 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# RNA Quantification 
--enable-rna-quantification true 
--rna-library-type A                    #see 'RNA Quant' 
--rna-quantification-gc-bias true 
# RNA Splice Variants 
--enable-rna-splice-variant true 
# RNA Gene Fusions 
--enable-rna-gene-fusion true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
```

## Notes and additional options

### Hashtable

For DRAGEN RNA/scRNA runs, it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                    |
| -------------------------------- | -------------------------------------------------------------- |
| `--enable-map-align true`        | For RNA/scRNA pipelines, map-align should always be turned on. |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### RNA Variant Calling

| Option                  | Description                                                                                                                                                                                               |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed $PATH` | Restrict the variants called to a target bed. A bed file specifying the gene-coding regions should be provided to avoid calling erroneous variants in unenriched or noncoding regions due to noisy reads. |

### RNA Quant

| Option               | Description                                                                                                                                                              |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `--rna-library-type` | Set the library according to the read orientations. Set to 'A' to auto detect the correct read orientation. Alternatively select 'IU', 'ISR', 'ISF', 'U', 'SR', or 'SF'. |

### RNA Splice

| Option                                    | Description                                                                                                                                                                                                                                                                                                                       |
| ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--rna-splice-variant-normals $PATH`      | Optional list of normal splice variants that will be used to filter false positive calls. The file should be a tab separated file with the following first four columns: (1) contig name, (2) first base of the splice junction (1-based), (3) last base of the splice junction (1-based), (4) strand (0: undefined, 1: +, 2: -). |
| `--rna-splice-variant-knowns $PATH`       | File with a list of expected splice junctions, these will be passed                                                                                                                                                                                                                                                               |
| `--rna-splice-variant-fusion-genes $PATH` | List of hotspot genes that may contain spliced fusions                                                                                                                                                                                                                                                                            |

### RNA Fusion

| Option                            | Description                                                                                                              |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--rna-gf-restrict-genes`         | Ignore genes that are not protein coding for gene fusions (Default=true)                                                 |
| `--rna-gf-enable-post-filters`    | Enable stringent post-filtering of RNA gene fusion candidates by quality flags to reduce false positives (Default=false) |
| `--rna-gf-output-fusion-sequence` | Output assembled gene fusion sequence in fusion\_candidates.final file (Default=true)                                    |
| `--rna-gf-report-intronic`        | Report fusion calls with intronic breakpoints (Default=true)                                                             |
| `--rna-gf-report-antisense`       | Report fusion calls with antisense and intronic breakpoints (Default=false)                                              |
| `--rna-gf-report-intergenic`      | Report fusion calls with intergenic, antisense and intronic breakpoints (Default=false)                                  |
| `--rna-gf-report-read-through`    | Report read-through fusion calls (Default=false)                                                                         |
| `--rna-gf-ptd-genes`              | List of gene names where we allow PTD/ITD associated self fusions                                                        |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)


# 5 Base DNA Germline Panel

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true`     | By default, DRAGEN marks duplicate reads and exclude them from variant calling.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--enable-positional-collapsing true` | Alternative to `--enable-duplicate-marking=true`. Instead of discarding duplicate reads, DRAGEN can optionally perform positional collapsing, merging them into higher-quality consensus reads. This is beneficial for small panels without UMIs and coverage between 300X and 1000X. However, it's slower than standard duplicate marking and less effective on samples with coverage lower than 300X. For very high coverage (1000X+), avoid it due to potential read collisions. For high-sensitivity panels with 1000X+ coverage, consider using UMIs. |

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)


# 5 Base DNA Germline Panel UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)


# 5 Base DNA Germline WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN pangenome hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
# SV 
--enable-sv true 
# CNV 
--enable-cnv true 
--cnv-enable-self-normalization true 
```

## Notes and additional options

### Hashtable

For DRAGEN germline runs, it is recommended to use the pangenome hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--bam-input $PATH 
```

CRAM Input

```
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                      | Description                                                                                                                                  |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                           | Limit variant calling to region of interest.                                                                                                 |
| `--vc-combine-phased-variants-distance INT` | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2) |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### CNV

| Option                                | Description                                                                                                                                               |
| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true` | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`   | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`        | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |


# 5 Base DNA Somatic Tumor-Normal Solid Panel

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true`     | By default, DRAGEN marks duplicate reads and exclude them from variant calling.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--enable-positional-collapsing true` | Alternative to `--enable-duplicate-marking=true`. Instead of discarding duplicate reads, DRAGEN can optionally perform positional collapsing, merging them into higher-quality consensus reads. This is beneficial for small panels without UMIs and coverage between 300X and 1000X. However, it's slower than standard duplicate marking and less effective on samples with coverage lower than 300X. For very high coverage (1000X+), avoid it due to potential read collisions. For high-sensitivity panels with 1000X+ coverage, consider using UMIs. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. These files should also be used in 5-Base workflows. The 5-Base workflows have not been tested with custom noise files.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |


# 5 Base DNA Somatic Tumor-Normal Solid WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH                      #see 'Input Options' for FQ, BAM or CRAM 
--fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Recommended 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
# SV 
--enable-sv true 
--sv-systematic-noise $PATH             #Optional 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# CNV 
--enable-cnv true 
--cnv-use-somatic-vc-baf true 
--cnv-enable-self-normalization true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--enable-variant-annotation true 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
--fastq-list $PATH 
--fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
--fastq-file1 $PATH 
--fastq-file2 $PATH 
--RGSM $STRING 
--RGID $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
--bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
--cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | In the TN pipeline this must be set to false for BAM/CRAM input.                                     |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-enable-liquid-tumor-mode true`                     | Tumor-in-normal contamination. Only use if there is some tumor leakage in the normal control.                                                                                                                                                                                                                                                                                                                                                     |
| `--vc-override-tumor-pcr-params-with-normal false`       | Mixed sample preparation. Only use if the tumor and normal samples exhibit different PCR (indel) noise patterns, e.g., due to using different sample preparation.                                                                                                                                                                                                                                                                                 |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 17.5. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 10. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### CNV

| Option                                 | Description                                                                                                                                               |
| -------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true`  | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`    | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`         | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |
| `--cnv-normal-cnv-vcf $CNV_NORMAL_VCF` | Specify germline CNVs from the matched normal sample.                                                                                                     |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                                            | Recommended Value for Liquid Tumors (e.g. AML/MLL)                                       |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| `--sv-enable-liquid-tumor-mode true`              | DRAGEN can account for Tumor-in-Normal (TiN) contamination by running liquid tumor mode. |
| `--sv-tin-contam-tolerance $TIN_CONTAM_TOLERANCE` | Set the Tumor-in-Normal (TiN) contamination tolerance level.                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. These files should also be used in 5-Base workflows. The 5-Base workflows have not been tested with custom noise files.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |


# 5 Base DNA Somatic Tumor-Only Solid Panel

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                                | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true`     | By default, DRAGEN marks duplicate reads and exclude them from variant calling.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--enable-positional-collapsing true` | Alternative to `--enable-duplicate-marking=true`. Instead of discarding duplicate reads, DRAGEN can optionally perform positional collapsing, merging them into higher-quality consensus reads. This is beneficial for small panels without UMIs and coverage between 300X and 1000X. However, it's slower than standard duplicate marking and less effective on samples with coverage lower than 300X. For very high coverage (1000X+), avoid it due to potential read collisions. For high-sensitivity panels with 1000X+ coverage, consider using UMIs. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 3. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. These files should also be used in 5-Base workflows. The 5-Base workflows have not been tested with custom noise files.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |


# 5 Base DNA Somatic Tumor-Only Solid Panel UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
--vc-enable-umi-solid true              #>= 1% VAF 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 4. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. These files should also be used in 5-Base workflows. The 5-Base workflows have not been tested with custom noise files.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |


# 5 Base DNA Somatic Tumor-Only Solid WGS

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments. This recipe includes the recommended commands for solid samples. These settings support fresh frozen samples, as well as some optional settings for FFPE samples.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
--enable-duplicate-marking true         #default=true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-systematic-noise $PATH             #Required 
--vc-excluded-regions-bed $BED          #FFPE: optionally mask ALUs 
# SV 
--enable-sv true 
--sv-systematic-noise $PATH             #Recommended 
--enable-oncovirus-detection true       #Optional 
--oncovirus-detection-db $PATH          #Optional 
# CNV 
--enable-cnv true 
--cnv-population-b-allele-vcf $POP_VCF 
--cnv-enable-self-normalization true 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Duplicate Marking

| Option                            | Description                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------- |
| `--enable-duplicate-marking true` | By default, DRAGEN marks duplicate reads and exclude them from variant calling. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.).                                                                                                                                                                                                                                                                           |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 3. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

Clinical applications require maximum confidence in variant calls to prevent false positive diagnoses. High-specificity mode reduces false positives with some loss in sensitivity.

| High Specificity Option      | Description                                                                                               |
| ---------------------------- | --------------------------------------------------------------------------------------------------------- |
| `--vc-high-specificity true` | Apply aggressive filters in the small variant caller to boost specificity, with some loss in sensitivity. |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### CNV

| Option                                | Description                                                                                                                                               |
| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gcbias-correction true` | Enable or disable GC bias correction when generating target counts.                                                                                       |
| `--cnv-segmentation-mode $SEG_MODE`   | Option to override the default segmentation algorithm. Defaults include `slm` for germline WGS, `aslm` for somatic WGS, and `hslm` for targeted analysis. |
| `--cnv-segmentation-bed $PATH`        | Specify a segmentation bed file to add pre-defined segments to be called.                                                                                 |

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

### SV

| Option                                             | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--sv-target-bed`                                  | Specifies a BED file containing the set of regions to call. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sv-exclusion-bed`                               | (Optional) Specifies a BED file containing the set of regions to exclude for the SV calling. Optionally gzip or bgzip format.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `--enable-variant-deduplication true`              | Relevant when both SV and SNV callers are enabled in somatic workflows. Can increase sensitivity and prevent the occurrence of replicated variants within genes such as FLT3 and KMT2A. Filter all small indels in the structural variant VCF that appear and are passing in the small variant VCF. DRAGEN will create a new VCF that contains variants in SV VCF that are not matching a variant from SNV VCF file. The new deduplicated SV VCF file will have the same prefix passed by `--output-file-prefix` followed by `sv.small_indel_dedup`. DRAGEN normalizes variants by trimming and left shifting by up to 500 bases. |
| `--sv-systematic-noise $BEDPE`                     | Systematic noise BEDPE file containing the set of noisy paired regions (optionally gzip or bzip compressed). Optional for WGS Tumor-Normal, but strongly recommended for WGS Tumor-Only. Has not been validated in WES, enrichment and amplicon panels.                                                                                                                                                                                                                                                                                                                                                                           |
| `--sv-somatic-ins-tandup-hotspot-regions-bed $BED` | Specify a custom BED of ITD hotspot regions to increase sensitivity for calling ITDs in somatic variant analysis. The default file includes FLT3, ARHGEF7, KMT2A, and UBTF exonic regions with some padding on both sides (300 bps)                                                                                                                                                                                                                                                                                                                                                                                               |
| `--sv-min-candidate-variant-size`                  | Run SV caller and report all SVs/indels at or above this size. The default value is set to 10.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--sv-min-scored-variant-size`                     | After candidate identification, only score and report SVs/indels at or above this size. The default value is set to 50. This parameter doesn't affect the somatic hotspot region.                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--enable-oncovirus-detection`                     | Set to enable detection of oncoviral integration. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--oncovirus-detection-db`                         | Specifies the directory containing oncovirus detection resource files. See [Oncovirus Detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/oncovirus-detection) for more information.                                                                                                                                                                                                                                                                                                                                                                                                                           |

| Option                              | Recommended Value for Liquid Tumors (e.g. AML/MLL) |
| ----------------------------------- | -------------------------------------------------- |
| `--sv-min-scored-variant-size $INT` | 100000                                             |

For more information, see [Structural Variant Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. These files should also be used in 5-Base workflows. The 5-Base workflows have not been tested with custom noise files.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |


# 5 Base DNA Somatic Tumor-Only ctDNA Panel UMI

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis. For clarity, some default parameters are explicitly included and annotated with comments.

```
  
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--ref-dir $REF_DIR                      #path to DRAGEN linear hashtable 
--output-directory $OUTPUT 
--intermediate-results-dir $PATH        #e.g. SSD /staging 
--output-file-prefix $PREFIX 
# Inputs 
--tumor-fastq-list $PATH                #see 'Input Options' for FQ, BAM or CRAM 
--tumor-fastq-list-sample-id $STRING 
# Mapper 
--enable-map-align true                 #optional with BAM/CRAM input 
--enable-map-align-output true          #optionally save the output BAM 
--enable-sort true                      #default=true 
# UMI 
--umi-enable true 
# 5-Base 
--methylation-conversion illumina 
--methylation-generate-cytosine-report true 
--methylation-compress-cx-report true 
# Small variant caller 
--enable-variant-caller true 
--vc-target-bed $VC_TARGET_BED 
--vc-systematic-noise $PATH             #Required 
--vc-enable-umi-liquid true             #>= 0.1% VAF 
# Annotation 
--variant-annotation-data $NIRVANA_PATH 
--vc-enable-germline-tagging true 
```

## Notes and additional options

### Hashtable

For DRAGEN somatic runs it is recommended to use the linear hashtable.

See: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Input options

DRAGEN input sources include: fastq list, fastq, bam, or cram. For BCL input, first create FASTQs using [BCL conversion](/dragen-v4.5/product-guides/dragen-v4.5/bcl-conversion).

FQ list Input

```
--tumor-fastq-list $PATH 
--tumor-fastq-list-sample-id $STRING 
```

FQ Input

```
--tumor-fastq1 $PATH 
--tumor-fastq2 $PATH 
--RGSM-tumor $STRING 
--RGID-tumor $STRING 
```

BAM Input

```
--tumor-bam-input $PATH 
```

CRAM Input

```
--tumor-cram-input $PATH 
```

### Mapping and Aligning

| Option                           | Description                                                                                          |
| -------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--enable-map-align true`        | Optionally disable map & align (default=true).                                                       |
| `--enable-map-align-output true` | Optionally save the output BAM (default=false).                                                      |
| `--Aligner.clip-pe-overhang 2`   | Clean up any unwanted UMI indexes. Only use when reads contain UMIs, but UMI collapsing was not run. |

### Fractional (Raw Reads) Downsampling

DRAGEN can subsample a random, fractional percentage of reads from an input file using the fractional downsampler. You can use downsampling to subsample data sets in order to simulate different amounts of sequencing. DRAGEN randomly subsamples reads from primary analysis without any modification (e.g. no trimming, no filtering, etc.).

Downsampling may be useful to reduce runtime on very deep samples. For Tumor-Normal analyses it is also recommended to use a normal sample with coverage that is less than the tumor sample. If the matched normal has deeper coverage than the tumor sample, then the fractional samples may be used to reduce coverage on the normal sample.

| Option                             | Description                                                                                                 |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--enable-fractional-down-sampler` | Set to true to enable fractional downsampling. The default value is false.                                  |
| `--down-sampler-normal-subsample`  | Specify the fraction of reads to keep as a subsample of normal input data. The default value is 1.0 (100%). |
| `--down-sampler-tumor-subsample`   | Specify the fraction of reads to keep as a subsample of tumor input data. The default value is 1.0 (100%).  |
| `--down-sampler-random-seed`       | Specify the random seed for different runs of the same input data. The default value is 42.                 |

### UMI

| Option                             | Description                                                                                                                                                                                                                                                                                                                         |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--umi-nonrandom-whitelist $PATH`  | If UMI is nonrandom, either a whitelist or correction table is required. The whitelist includes a valid UMI sequence per line.                                                                                                                                                                                                      |
| `--umi-correction-table $PATH`     | If UMI is nonrandom, either a whitelist or correction table is required. The correction table defaults to the table used by TruSight Oncology: \<INSTALL\_PATH>/resources/umi/umi\_correction\_table.txt.gz.                                                                                                                        |
| `--umi-min-supporting-reads INT`   | Specify the number of matching UMI input reads required to generate a consensus read. Any family with insufficient supporting reads is discarded. Most pipelines perform better with a setting of 1. A setting of 2 may potentially be relevant for samples with ultra deep coverage (e.g. ctDNA). (Default since DRAGEN V4.5 is 1) |
| `--umi-metrics-interval-file $BED` | Target region in BED format.                                                                                                                                                                                                                                                                                                        |
| `--umi-emit-multiplicity both`     | Set the consensus sequence type to output. DRAGEN UMI allows collapsing duplex sequences from the two strands of the original molecules. For more information, see [Merge Duplex UMIs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#merge-duplex-umis).                                 |
| `--umi-start-mask-length INT`      | Number of additional bases to ignore from start of read. The default is 0. To reduce FP optionally set to 1.                                                                                                                                                                                                                        |
| `--umi-end-mask-length INT`        | Number of additional bases to ignore from end of read. The default is 0. To reduce FP optionally set to 3.                                                                                                                                                                                                                          |

For more information see: [UMI Options](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/unique-molecular-identifiers#umi-options).

### 5-Base Methylation

| Option                                        | Description                                                                                                                                                                                                                       |
| --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--methylation-conversion STRING`             | Library conversion for methylation analysis. Options: `none`, `c_t`, `mc_t`, `illumina` (default=none).                                                                                                                           |
| `--methylation-protocol STRING`               | Library protocol for methylation analysis. Options: `none`, `directional`, `non-directional`, `directional-complement`, `pbat`. The default value for `methylation-conversion=illumina` is `directional`, otherwise it is `none`. |
| `--methylation-mapq-threshold INT`            | Only reads with MAPQ greater or equal than the threshold will be included in methyl-seq analysis (default=0).                                                                                                                     |
| `--methylation-generate-mbias-report true`    | Whether to generate a per-sequencer-cycle methylation bias report (default=true).                                                                                                                                                 |
| `--mbias-report-include-overlaps`             | Calculate methylation stats for overlapping bases between mates (default=false).                                                                                                                                                  |
| `--methylation-generate-cytosine-report true` | Whether to generate a genome-wide cytosine methylation CX\_report file (default=false).                                                                                                                                           |
| `--methylation-compress-cx-report true`       | Set to true to enable compression of the CX\_report (default=true).                                                                                                                                                               |
| `--methylation-keep-ref-cytosine true`        | Set to true to keep all reference cytosines in the CX\_report file, even if they don't appear in the input reads (default=false).                                                                                                 |
| `--enable-cpg-methylated-mapping true`        | Enable methylated mapping with base conversions restricted to CpG context (default=true). When false, runs DRAGEN Methylation 3-base map/align instead.                                                                           |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in VCF files (default=c).                                                                                                                                             |
| `--methylation-report-to-vcf`                 | Specify methylation type (none, cg, or c) which is reported in gVCF files (default=cg).                                                                                                                                           |

For more information see: [5-Base Pipeline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-5-base/dragen-5base-pipeline).

### SNV

| Option                                                   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--vc-target-bed`                                        | Limit variant calling to region of interest.                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--vc-combine-phased-variants-distance INT`              | Maximum distance in base pairs (BP) over which phased variants will be combined. Set to 0 to disable. Valid range is \[0; 15] BP (Default=2)                                                                                                                                                                                                                                                                                                      |
| `--vc-systematic-noise $PATH`                            | Systematic noise file. This filter is recommended for removing systematic noise observed in normal samples (i.e. systematic alignment errors, sequencing errors, etc.). When working with panels it is recommended that a custom systematic noise file be created for each assay.                                                                                                                                                                 |
| `--vc-somatic-hotspots $PATH`                            | DRAGEN has a default set of hotspot variants (positions and alleles) where it will assign an increased prior probability. Use this option to override with a custom hotspots file.                                                                                                                                                                                                                                                                |
| `--vc-sq-filter-threshold $NUM`                          | Threshold for sensitivity-specificity tradeoff using SQ score. The pipeline specific default threshold is 2. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                                                                                                   |
| `--vc-systematic-noise-filter-threshold $INT`            | Threshold for sensitivity-specificity tradeoff using AQ score for non-hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 60. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                        |
| `--vc-systematic-noise-filter-threshold-in-hotspot $INT` | Threshold for sensitivity-specificity tradeoff using AQ score for hotspot variants. This is only used when supplying a systematic noise file. Pipeline specific default value = 20. Raise this value to improve specificity at the cost of sensitivity, or lower it to improve sensitivity at the cost of specificity.                                                                                                                            |
| `--vc-excluded-regions-bed $BED`                         | Hard filter variants that overlap with this region. ALU regions comprise approximately 11% of the genome, and are often exceptionally noisy regions in FFPE samples. Optionally filter out ALU regions using the DRAGEN excluded regions filter. ALU bed files can be downloaded as part of the Bed File Collection: [Bed File Collection](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) |

High-coverage sequencing panels allow for the detection of low-frequency alleles. DRAGEN supports 3 main settings for improved sensitivity on low VAF variant calls.

| High Sensitivity Option       | Description                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--vc-target-vaf FLOAT`       | The default is 0.03 (3%). Set to e.g. 0.01 to improve SNV sensitivity on 1% VAF variants (assuming sufficient coverage). |
| `--vc-enable-umi-solid true`  | Optimized for 1% and higher VAFs on UMI (or read position collapsed) samples with approx 300-1000X coverage.             |
| `--vc-enable-umi-liquid true` | Optimized for 0.1% and higher VAFs on UMI samples with 1000X or higher coverage as expected in liquid biopsies.          |

For more detail on the small variant caller in somatic mode please refer to [Somatic Mode](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/somatic-mode)

### Annotation

For instructions on how to download the Nirvana annotation database, please refer to [Nirvana](/dragen-v4.5/product-guides/dragen-v4.5/nirvana)

## Resource Files

DRAGEN requires resource files for components such as SNV, SV, and CNV. The following notes provide references for downloading these files or generating them for custom workflows or assays.

### SNV Systematic Noise

Systematic noise files are considered essential in Tumor-Only workflows. It is also recommended for Tumor-Normals workflows.

DRAGEN has pre-built systematic noise files for WGS, WES and for Pillar Amplicons. These files should also be used in 5-Base workflows. The 5-Base workflows have not been tested with custom noise files.

#### Prebuilt

Prebuilt systematic noise BED files (WES and WGS) can be downloaded here: [Product Files](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

| Prebuilt WES/WGS noise files                       | Description              |
| -------------------------------------------------- | ------------------------ |
| `WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WGS FF               |
| `FFPE_WGS_hg38_v2.0.0_systematic_noise.snv.bed.gz` | For WGS FFPE (only hg38) |
| `WES_hg38_v2.0.0_systematic_noise.snv.bed.gz`      | For WES FF and FFPE      |


# DRAGEN WGS Germline Accuracy

This section contains comprehensive accuracy metrics for DRAGEN's germline variant calling.

## Available components

* Standard SBS
  * [Small Variants](/dragen-v4.5/product-guides/dragen-v4.5/germline/small_variants_accuracy) — SNPs and indels detection accuracy across HG001-HG007 samples
  * [Multi-Region Joint Detection](/dragen-v4.5/product-guides/dragen-v4.5/germline/mrjd_accuracy) — MRJD performance in high-homology paralogous regions
  * [Copy Number Variation](/dragen-v4.5/product-guides/dragen-v4.5/germline/copy_number_variation_accuracy) — CNV calling performance metrics
  * [HLA Genotyping](/dragen-v4.5/product-guides/dragen-v4.5/germline/hla_genotyping_accuracy) — HLA Class I and II genotyping accuracy
  * [Targeted Callers](/dragen-v4.5/product-guides/dragen-v4.5/germline/targeted_callers_accuracy) — Performance for specialized variant callers
  * [Ploidy Estimator](/dragen-v4.5/product-guides/dragen-v4.5/germline/ploidy_estimator_accuracy) — Ploidy estimation accuracy
  * [Structural Variants](/dragen-v4.5/product-guides/dragen-v4.5/germline/structural_variants_accuracy) — SV and MEI detection metrics
  * [Tandem Repeats](/dragen-v4.5/product-guides/dragen-v4.5/germline/tandem_repeats_accuracy) — STR expansion detection and classification accuracy

## Reference

**DRAGEN Version**: 4.5.4 | **Reference Genome**: GRCh38


# Small Variants

Generated on **2026-05-07**

## Small variants

The DRAGEN Germline Small Variant Caller takes mapped and aligned DNA reads as input and calls SNPs and indels through a combination of column-wise detection and local de novo assembly of haplotypes. For more information, refer to the [DRAGEN user guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling).

**DRAGEN**: DRAGEN 4.4.6 vs 4.5.4 | **Truthset**: GIAB/NIST high-confidence, GIAB/NIST v4.2.1 | **Reference**: GRCh38

Results reported for HG001-HG007 sequenced with NSX 1.3 with 10B flow cell downsampled at 35x raw coverage.

<details>

<summary>SNPs table (click to expand)</summary>

| Sample | Version | Recall | Precision | F1-score | FN   | FP   |
| ------ | ------- | ------ | --------- | -------- | ---- | ---- |
| HG001  | 4.4.6   | 1.000  | 0.999     | 0.999    | 1336 | 2634 |
| HG001  | 4.5.4   | 1.000  | 1.000     | 1.000    | 721  | 1658 |
| HG002  | 4.4.6   | 0.999  | 1.000     | 0.999    | 2896 | 1749 |
| HG002  | 4.5.4   | 0.999  | 1.000     | 1.000    | 2198 | 1026 |
| HG003  | 4.4.6   | 0.999  | 0.999     | 0.999    | 3511 | 3005 |
| HG003  | 4.5.4   | 0.999  | 0.999     | 0.999    | 2936 | 2338 |
| HG004  | 4.4.6   | 0.999  | 0.999     | 0.999    | 3748 | 2584 |
| HG004  | 4.5.4   | 0.999  | 1.000     | 0.999    | 2963 | 1803 |
| HG005  | 4.4.6   | 0.999  | 0.999     | 0.999    | 3023 | 2047 |
| HG005  | 4.5.4   | 0.999  | 1.000     | 0.999    | 2413 | 1369 |
| HG006  | 4.4.6   | 0.999  | 0.999     | 0.999    | 3103 | 2408 |
| HG006  | 4.5.4   | 0.999  | 1.000     | 0.999    | 2484 | 1699 |
| HG007  | 4.4.6   | 0.999  | 0.999     | 0.999    | 3467 | 2327 |
| HG007  | 4.5.4   | 0.999  | 1.000     | 0.999    | 2955 | 1602 |

</details>

<details>

<summary>Indels table (click to expand)</summary>

| Sample | Version | Recall | Precision | F1-score | FN   | FP   |
| ------ | ------- | ------ | --------- | -------- | ---- | ---- |
| HG001  | 4.4.6   | 0.996  | 0.997     | 0.997    | 1867 | 1437 |
| HG001  | 4.5.4   | 0.996  | 0.997     | 0.997    | 1849 | 1231 |
| HG002  | 4.4.6   | 0.996  | 0.997     | 0.997    | 2101 | 1421 |
| HG002  | 4.5.4   | 0.996  | 0.998     | 0.997    | 1998 | 1229 |
| HG003  | 4.4.6   | 0.996  | 0.997     | 0.997    | 1852 | 1420 |
| HG003  | 4.5.4   | 0.997  | 0.998     | 0.997    | 1789 | 1164 |
| HG004  | 4.4.6   | 0.996  | 0.997     | 0.997    | 1898 | 1384 |
| HG004  | 4.5.4   | 0.996  | 0.998     | 0.997    | 1877 | 1226 |
| HG005  | 4.4.6   | 0.998  | 0.998     | 0.998    | 897  | 670  |
| HG005  | 4.5.4   | 0.998  | 0.999     | 0.998    | 810  | 554  |
| HG006  | 4.4.6   | 0.997  | 0.998     | 0.998    | 1147 | 838  |
| HG006  | 4.5.4   | 0.997  | 0.998     | 0.998    | 1116 | 741  |
| HG007  | 4.4.6   | 0.997  | 0.998     | 0.998    | 1246 | 868  |
| HG007  | 4.5.4   | 0.997  | 0.998     | 0.998    | 1161 | 762  |

</details>

![HG001-HG007, NSX 10B v1.3, 35x raw coverage](/files/mkj2KXTcsFqXVXCy2tuK)

**HG002 T2TQ100 performance**

**DRAGEN**: 4.4.6 vs 4.5.4 | **Truthset**: T2TQ100-v1.1-v0.019 | **Reference**: GRCh38

<details>

<summary>HG002 T2TQ100-v1.1-v0.019 small variants comparison (click to expand)</summary>

| Version | SNPs FN | SNPs FP | SNPs FP+FN | Indels FN | Indels FP | Indels FP+FN |
| ------- | ------- | ------- | ---------- | --------- | --------- | ------------ |
| 4.4.6   | 21571   | 8752    | 30323      | 40208     | 23681     | 63889        |
| 4.5.4   | 16490   | 5764    | 22254      | 36945     | 19661     | 56606        |

</details>

![HG002 T2TQ100-v1.1-v0.019 FP+FN comparison](/files/6KgGeK6sSvR1mEZmwd2r)

**A legacy of continuous innovation and improving accuracy in each version.**

Historical HG002 benchmarked releases using the HG002 small variants NIST v4.2.1 truthset show FP+FN dropping from 41,244 in DRAGEN 3.4.5 to 4,364 in DRAGEN 4.5.4.

<details>

<summary>Historical version trend table (click to expand)</summary>

| Version | FP+FN  |
| ------- | ------ |
| 3.4.5   | 41,244 |
| 3.5.7   | 40,056 |
| 3.6.3   | 39,046 |
| 3.7.5   | 21,479 |
| 3.8.4   | 20,402 |
| 3.9.3   | 19,052 |
| 3.10.4  | 12,851 |
| 4.0.3   | 12,815 |
| 4.2.4   | 11,163 |
| 4.3.3   | 6,615  |
| 4.4.6   | 5,955  |
| 4.5.4   | 4,364  |

</details>

![WGS small variants FP+FN improvement across DRAGEN versions](/files/t5di3F41BncBGjunNngL)

**Mosaic detection**

Non-cancer post-zygotic mosaic variants have typical allele fraction (AFs) lower than 50% and therefore more challenging to find with the default small variant caller that has been optimized to detect germline variants with typical AFs of 0%, 50% or 100%. The Mosaic variant caller has been designed to address this gap by training a dedicated machine learning model trained on lowAF variants. For more information, refer to the [DRAGEN user guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/mosaic-detection).

**DRAGEN**: DRAGEN 4.5.4 vs 4.4.6 | **Truthset**: HG002 NIST Mosaic v1.0 | **Reference**: GRCh38 Results reported for HG002 sample [fastqs](https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/data/AshkenazimTrio/HG002_NA24385_son/NIST_HiSeq_HG002_Homogeneity-10953946/HG002_HiSeq300x_fastq/) sequenced with HiSeq2500 at 300x available from NIST.

<details>

<summary>HG002 mosaic small variants comparison (click to expand)</summary>

| Version | SNPs FN | SNPs FP | SNPs TP | SNPs FP+FN | Indels FN | Indels FP | Indels TP | Indels FP+FN |
| ------- | ------- | ------- | ------- | ---------- | --------- | --------- | --------- | ------------ |
| 4.4.6   | 0       | 318     | 85      | 318        | 0         | 196       | 0         | 196          |
| 4.5.4   | 0       | 311     | 85      | 311        | 0         | 185       | 0         | 185          |

</details>

![HG002 Mosaic FP comparison](/files/FpdSLHjroFgW6ICAT8IJ)


# Copy Number Variation

Generated on **2026-05-13**

## Copy Number Variation

The DRAGEN Copy Number Variant (CNV) Pipeline can call CNV events using next-generation sequencing (NGS) data. This pipeline supports multiple applications in a single interface via the DRAGEN Host Software, including processing of whole-genome sequencing (WGS) data and whole-exome sequencing (WES) data. For more information, refer to the [DRAGEN user guide for germline CNV calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline).

**DRAGEN**: DRAGEN 4.5.4 | **Truthset**: HG002 NIST CNV v0.6 | **Reference**: GRCh38

<details>

<summary>Standard WGS CNV table (click to expand)</summary>

| Subtype   | Recall | Precision | F1-score |
| --------- | ------ | --------- | -------- |
| 1kb-10kb  | 0.968  | 0.977     | 0.973    |
| 10kb-50kb | 0.950  | 0.952     | 0.951    |
| >50kb     | 1.000  | 1.000     | 1.000    |

</details>

![CNV for Standard WGS](/files/i8sjnGzBiIdmYJh4oI3I)

### Cytogenetics Modality

Conventional cytogenetics methodologies typically focus on larger alterations than the ones provided by NGS analyses. The Cytogenetics modality for the CNV caller allows the user to visualize variants at different resolutions, aiming at providing a more flexible workspace for different use cases. For more information, refer to [DRAGEN manual](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline#cytogenetics-modality).

<details>

<summary>Cytogenetics CNV summary table (click to expand)</summary>

| Group        | TP  | FN | Recall |
| ------------ | --- | -- | ------ |
| DEL 25kb-1Mb | 15  | 1  | 0.938  |
| DEL >=1Mb    | 46  | 1  | 0.979  |
| DUP 50kb-1Mb | 10  | 0  | 1.000  |
| DUP >=1Mb    | 29  | 1  | 0.967  |
| AOH >=500kb  | 43  | 0  | 1.000  |
| TOTAL        | 143 | 3  | 0.979  |

</details>

![Cytogenetics CNV Recall by Group](/files/09QKbE4RRxh5X260Crko)


# Structural Variants

Generated on **2026-05-07**

## Structural Variants

The DRAGEN Structural Variant (SV) caller identifies large insertions, deletions and genomic arrangements from mapped DNA reads. For details, refer to the [DRAGEN user guide for structural variant calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling).

**DRAGEN**: DRAGEN 4.5.4 | **Truthset**: HG002 T2TQ100-v1.1-v0.018 on Truvari v4.2.2 | **Reference**: GRCh38

<details>

<summary>SV WGS historical comparison table (click to expand)</summary>

| Method       | Recall | Precision | F1-score |
| ------------ | ------ | --------- | -------- |
| DRAGEN 4.5.4 | 0.880  | 0.969     | 0.922    |
| DRAGEN 4.4   | 0.807  | 0.963     | 0.878    |
| DRAGEN 4.3   | 0.511  | 0.953     | 0.665    |

</details>

**Comparison with previous DRAGEN versions**

![SV WGS accuracy comparison across DRAGEN versions](/files/H7MjuwTSYx2TqG9M9Otb)

**Mobile Element Insertions**

The general purpose SV routine can detect Mobile Element Insertions (MEIs) with assembled inserted sequences like other regular insertions. If missed by the general purpose SV routine, MEIs will be rescued by the MEI specific routine based on similarity between assembled contigs and known sequences in the MEI catalog.

Recall for **HG002 NIST v0.6 MEI** (Wittyer v0.5.4.1):

<details>

<summary>SV MEI historical comparison table (click to expand)</summary>

| Method       | Recall |
| ------------ | ------ |
| DRAGEN 4.5.4 | 0.985  |
| DRAGEN 4.4   | 0.906  |
| DRAGEN 4.3   | 0.885  |

</details>

**Comparison with previous DRAGEN versions**

![SV MEI recall trend across DRAGEN versions](/files/zwTlJmWWgraqLrcCWqoe)


# HLA Genotyping

Generated on **2026-05-07**

## HLA Genotyping

DRAGEN includes a dedicated genotyper for genotyping the Human Leukocyte Antigen (HLA) genes. This typer is capable of matching to established HLA alleles from the IMGT Database as well as calling novel alleles. For WGS data or panels covering the entire HLA genes, DRAGEN can output HLA types at full resolution. For more information, refer to [DRAGEN manual](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/hla-typing).

Assessment based on 47 samples with long read truth calls. 3- and 4-field resolution for HLA Class I accuracy is not supported in DRAGEN 4.4.

<details>

<summary>HLA Class I Accuracy Table</summary>

**2-field comparison**

| Gene | Field resolution | DRAGEN version | Accuracy | 2-field improvement (4.5.4 - 4.4) |
| ---- | ---------------- | -------------- | -------- | --------------------------------- |
| A    | 2-field          | 4.4            | 98.94%   |                                   |
| B    | 2-field          | 4.4            | 96.81%   |                                   |
| C    | 2-field          | 4.4            | 98.94%   |                                   |
| All  | 2-field          | 4.4            | 98.23%   |                                   |
| A    | 2-field          | 4.5.4          | 98.94%   | +0.00 pp                          |
| B    | 2-field          | 4.5.4          | 98.94%   | +2.13 pp                          |
| C    | 2-field          | 4.5.4          | 100.00%  | +1.06 pp                          |
| All  | 2-field          | 4.5.4          | 99.29%   | +1.06 pp                          |

**Other field resolutions**

| Gene | Field resolution | DRAGEN version | Accuracy |
| ---- | ---------------- | -------------- | -------- |
| A    | 3-field          | 4.5.4          | 98.94%   |
| B    | 3-field          | 4.5.4          | 96.81%   |
| C    | 3-field          | 4.5.4          | 100.00%  |
| All  | 3-field          | 4.5.4          | 98.58%   |
| A    | 4-field          | 4.5.4          | 94.68%   |
| B    | 4-field          | 4.5.4          | 92.55%   |
| C    | 4-field          | 4.5.4          | 93.62%   |
| All  | 4-field          | 4.5.4          | 93.62%   |

</details>

![HLA Class I Accuracy](/files/YwEgzYaUgwLn57MDgr56)

<details>

<summary>HLA Class II Accuracy Table</summary>

| Gene | Field resolution | DRAGEN version | Accuracy |
| ---- | ---------------- | -------------- | -------- |
| DPA1 | 2-field          | 4.5.4          | 94.05%   |
| DPB1 | 2-field          | 4.5.4          | 98.81%   |
| DQA1 | 2-field          | 4.5.4          | 99.40%   |
| DQB1 | 2-field          | 4.5.4          | 100.00%  |
| DRB1 | 2-field          | 4.5.4          | 93.45%   |
| All  | 2-field          | 4.5.4          | 97.14%   |
| DPA1 | 3-field          | 4.5.4          | 91.67%   |
| DPB1 | 3-field          | 4.5.4          | 98.21%   |
| DQA1 | 3-field          | 4.5.4          | 98.81%   |
| DQB1 | 3-field          | 4.5.4          | 92.26%   |
| DRB1 | 3-field          | 4.5.4          | 93.45%   |
| All  | 3-field          | 4.5.4          | 94.88%   |
| DPA1 | 4-field          | 4.5.4          | 87.50%   |
| DPB1 | 4-field          | 4.5.4          | 93.45%   |
| DQA1 | 4-field          | 4.5.4          | 89.29%   |
| DQB1 | 4-field          | 4.5.4          | 83.33%   |
| DRB1 | 4-field          | 4.5.4          | 75.00%   |
| All  | 4-field          | 4.5.4          | 85.71%   |

</details>

![HLA Class II Accuracy](/files/ZkoljKAVK74GYX9lc69S)


# Targeted Callers

Generated on **2026-05-07**

## Targeted Callers

Repetitive regions in the human genome pose a challenge for general variant calling approaches which typically cannot make use of potentially misplaced MAPQ0 reads. Furthermore, high sequence homology of some genes with a pseudogene paralog can lead to a wide variety of common structural variants (SVs) in the population, requiring specialized targeted calling approaches. DRAGEN supports targeted calling for a number of genes/targets as described in subsequent target-specific sections. For more information, refer to the [DRAGEN user guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller).

**DRAGEN**: DRAGEN 4.5.4 | **Assay**: WGS

**Performance summary**

**Targeted caller concordance summary (WGS)**

| Target               | Concordant | Total | Concordance |
| -------------------- | ---------- | ----- | ----------- |
| CYP2D6               | 211        | 212   | 99.5%       |
| SMN1                 | 247        | 250   | 98.8%       |
| SMN2                 | 244        | 250   | 97.6%       |
| GBA                  | 24         | 24    | 100.0%      |
| CYP2B6 (GeT-RM)      | 70         | 76    | 92.1%       |
| CYP2B6 (PacBio HiFi) | 123        | 125   | 98.4%       |
| CYP21A2              | 82         | 84    | 97.6%       |
| HBA1/2               | 244        | 247   | 98.8%       |
| RHD/RHCE             | 41         | 42    | 97.6%       |

**CYP2D6**

Concordant across 211 out of 212 (99.5%) WGS samples with calls from GeT-RM and long read sequencing data.

**CYP2D6 concordance by star allele type**

| Star Allele Type  | Number of Samples | Concordance     |
| ----------------- | ----------------- | --------------- |
| without SV        | 90                | 89 (98.9%)      |
| deletion          | 15                | 15 (100%)       |
| duplication       | 12                | 12 (100%)       |
| fusion/conversion | 26                | 26 (100%)       |
| **Total**         | **143**           | **142 (99.3%)** |

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/cyp2d6-calling) | [Cyrius: accurate CYP2D6 genotyping using WGS data](https://www.illumina.com/science/genomics-research/articles/cyrius-cyp2d6-genotyping-wgs-data.html)

**SMN1/2**

Concordant across 247/250 (98.8%) WGS samples for SMN1 and 244/250 (97.6%) WGS samples for SMN2, compared against digital PCR and MLPA.

**SMN1 concordance by copy number**

| SMN1 CN   | Total Samples | Concordant      |
| --------- | ------------- | --------------- |
| 0         | 42            | 42 (100%)       |
| 1         | 23            | 23 (100%)       |
| 2         | 96            | 94 (97.9%)      |
| 3         | 68            | 68 (100%)       |
| 4         | 20            | 20 (100%)       |
| 6         | 1             | 0 (0%)          |
| **Total** | **250**       | **247 (98.8%)** |

**SMN2 concordance by copy number**

| SMN2 CN   | Total Samples | Concordant      |
| --------- | ------------- | --------------- |
| 0         | 26            | 26 (100%)       |
| 1         | 73            | 71 (97.3%)      |
| 2         | 92            | 89 (96.7%)      |
| 3         | 54            | 54 (100%)       |
| 4         | 5             | 4 (80%)         |
| **Total** | **250**       | **244 (97.6%)** |

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/smn-calling) | [SMA carrier screening from WGS data](https://www.illumina.com/science/genomics-research/articles/sma-carrier-screening-wgs-data.html)

**GBA**

Concordant across 24/24 (100%) WGS samples compared against digital PCR and long read sequencing data.

**GBA concordance by variant type**

| Copy Number Change | Recombinant Variant | Total Samples | Concordant    |
| ------------------ | ------------------- | ------------- | ------------- |
| Gain               | None                | 10            | 10 (100%)     |
| Loss               | None                | 2             | 2 (100%)      |
| Loss               | RecNciI             | 1             | 1 (100%)      |
| Loss               | A495P               | 1             | 1 (100%)      |
| Neutral            | L483P               | 7             | 7 (100%)      |
| Neutral            | c.1263del           | 1             | 1 (100%)      |
| Neutral            | c.1263del+RecTL     | 2             | 2 (100%)      |
| **Total**          |                     | **24**        | **24 (100%)** |

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/gba-calling) | [GBA gene for Gaucher and Parkinson disease research](https://www.illumina.com/science/genomics-research/articles/using-dragen-for-gaucher-and-parkinson-disease-research--resolvi.html)

**CYP2B6**

Concordance against calls from GeT-RM and long read sequencing data. Concordance is based on metabolizer status when multiple genotypes are reported. Two discordant samples in the PacBio comparison are due to a novel star allele (\*U1) resulting in multiple possible genotypes.

**CYP2B6 concordance by reference source**

| Star Allele Genotype Source | Number of Samples | DRAGEN Concordance |
| --------------------------- | ----------------- | ------------------ |
| GeT-RM                      | 76                | 70 (92.1%)         |
| PacBio HiFi reads           | 125               | 123 (98.4%)        |

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/cyp2b6-calling)

**CYP21A2**

Concordant across 82/84 (97.6%) WGS samples across three benchmark sets with calls from long-range PCR, MLPA and long read sequencing.

**CYP21A2 concordance by benchmark set**

| Benchmark Set | Total Samples | Concordant     |
| ------------- | ------------- | -------------- |
| Internal      | 14            | 13 (92.9%)     |
| 1KGP          | 66            | 65 (98.5%)     |
| Coriell       | 4             | 4 (100%)       |
| **Total**     | **84**        | **82 (97.6%)** |

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/cyp21a2-calling) | [CYP21A2 blog](https://www.illumina.com/science/genomics-research/articles/CYP21A2.html)

**HBA1/2**

Concordant across 244/247 (98.8%) WGS samples compared against long read sequencing data.

**HBA1/2 WGS concordance by genotype**

| HBA Genotype | Total Samples | Concordant      |
| ------------ | ------------- | --------------- |
| aa/aa        | 202           | 199 (98.5%)     |
| -α3.7/aa     | 31            | 31 (100%)       |
| --/aa        | 4             | 4 (100%)        |
| -α3.7/-α3.7  | 4             | 4 (100%)        |
| αααα3.7/aa   | 3             | 3 (100%)        |
| -α4.2/aa     | 2             | 2 (100%)        |
| αααα4.2/aa   | 1             | 1 (100%)        |
| **Total**    | **247**       | **244 (98.8%)** |

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/hba-calling) | [HBA targeted caller](https://www.illumina.com/science/genomics-research/articles/HBA-targeted-caller.html)

**LPA**

LPA copy number is assessed via allele-specific KIV-2 repeat quantification. Accuracy was evaluated by trio/duo-based inheritance analysis and comparison against Bionano optical mapping.

![LPA benchmark: KIV-2 copy number concordance with Bionano and trio inheritance analysis](/files/j2fgig1YVvfJwW7yigaE)

**Figure**: **(A)** Trio-based phased KIV-2 copy number comparison in 120 offspring vs best-matching parental call. **(B)** Duo-based comparison in 153 offspring. **(C)** Total copy number vs Bionano optical mapping (n=145). **(D)** Allelic copy number vs Bionano optical mapping (n=145). Pearson's r and P-value shown per panel.

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/lpa-calling) | [KIV-2 copy number variants for cardiovascular disease risk](https://www.illumina.com/science/genomics-research/articles/using-whole-genome-sequencing-to-evaluate-copy-number-variants-o.html)

**RHD/RHCE**

Concordant across 41/42 (97.6%) WGS samples compared against long read sequencing data.

**RHD/RHCE concordance by gene conversion status**

| RHCE\*CE-D(2)-CE Gene Conversion | Total Samples | Concordant     |
| -------------------------------- | ------------- | -------------- |
| Not present                      | 24            | 24 (100%)      |
| Heterozygous                     | 10            | 9 (90%)        |
| Homozygous                       | 8             | 8 (100%)       |
| **Total**                        | **42**        | **41 (97.6%)** |

**References:** [DRAGEN Product Guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller/rh-calling)

**Star Allele Caller**

The Star Allele Caller identifies the genotypes and metabolism status of the following PGx genes that are included in FDA's PGx recommendations or have CPIC Level A designation: CACNA1S, CFTR, CYP2C19, CYP2C9, CYP3A4, CYP3A5, CYP4F2, IFNL3, RYR1, NUDT15, SLCO1B1, TPMT, UGT1A1, VKORC1, DPYD, G6PD, MT-RNR1, BCHE, ABCG2, NAT2, F5 and UGT2B17. For more information, refer to the [DRAGEN user guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/star-allele-caller).

| Sample source      | # samples | Compare against                   | Genes tested                                                                                             | Concordance\* |
| ------------------ | --------- | --------------------------------- | -------------------------------------------------------------------------------------------------------- | ------------- |
| 1K genomes project | 3201      | PharmCAT (from DRAGEN gVCF)       | DPYD, CACNA1S, UGT1A1, TPMT, CYP3A5, CFTR, CYP2C19, CYP2C9, SLCO1B1, NUDT15, VKORC1, CYP4F2, RYR1, IFNL3 | 100%          |
| Coriell            | 96        | PharmCAT (from DRAGEN gVCF)       | DPYD, CACNA1S, UGT1A1, TPMT, CYP3A5, CFTR, CYP2C19, CYP2C9, SLCO1B1, NUDT15, VKORC1, CYP4F2, RYR1, IFNL3 | 100%          |
| Coriell            | 96        | GeT-RM (using orthogonal methods) | CYP2C19, CYP2C9, CYP3A5, CYP4F2, TPMT                                                                    | 94%\*         |
| 1K genomes project | 100       | Aldy (from BWA/GATK BAM)          | G6PD, CFTR, IFNL3, TPMT, VKORC1                                                                          | 100%          |

\* *Truth set from 2016; 6% calls from outdated star allele definition.*


# Multi-Region Joint Detection

Generated on **2026-05-11**

## Multi-Region Joint Detection

DRAGEN Multi-region Joint Detection (MRJD) is a de novo germline small variant caller for paralogous regions. In DRAGEN v4.3, MRJD covers regions that include six clinically relevant genes: NEB, TTN, SMN1/2, PMS2, STRC, and IKBKG. MRJD is compatible with hg38, hg19 and GRCh37 reference genome. The table below includes hg38 region coordinates covered by MRJD. For more information, refer to the [DRAGEN user guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/multi-region-joint-detection).

<details>

<summary>hg38 region coordinates covered by MRJD (click to expand)</summary>

| Chromosome | Start     | End       | Description       |
| ---------- | --------- | --------- | ----------------- |
| chr2       | 151578759 | 151588523 | NEB exon 98-105   |
| chr2       | 151589318 | 151599076 | NEB exon 90-97    |
| chr2       | 151599871 | 151609628 | NEB exon 82-89    |
| chr2       | 178653238 | 178654995 | TTN exon 172-180  |
| chr2       | 178657498 | 178659255 | TTN exon 181-189  |
| chr2       | 178661759 | 178663516 | TTN exon 190-198  |
| chr5       | 70049522  | 70077596  | SMN2              |
| chr5       | 70924940  | 70953013  | SMN1              |
| chr7       | 5970924   | 5980896   | PMS2 exon 13-15   |
| chr7       | 5980968   | 5987689   | PMS2 exon 11-12   |
| chr7       | 6737007   | 6743712   | PMS2CL exon 2-3   |
| chr7       | 6743880   | 6753867   | PMS2CL exon 4-6   |
| chr15      | 43599563  | 43602630  | STRC exon 24-29   |
| chr15      | 43602982  | 43611000  | STRC exon 14-23   |
| chr15      | 43611040  | 43618800  | STRC exon 1-13    |
| chr15      | 43699379  | 43702452  | STRCP1 exon 23-28 |
| chr15      | 43702488  | 43710472  | STRCP1 exon 13-22 |
| chr15      | 43710502  | 43718262  | STRCP1 exon 1-12  |
| chrX       | 154555884 | 154565047 | IKBKG exon 3-10   |
| chrX       | 154639390 | 154648553 | IKBKGP1           |

</details>

Sample description: 147 cell line samples from the Illumina Polaris 1 diversity panel with orthogonal small variant calls derived from long-range PCR data; RTG Tools ploidy-squash mode was used for benchmarking.

**MRJD performance summary**

| Comparison                                                                 | SNV recall                                               | INDEL recall                                             | Additional notes                                           |
| -------------------------------------------------------------------------- | -------------------------------------------------------- | -------------------------------------------------------- | ---------------------------------------------------------- |
| MRJD high sensitivity mode\* vs long-range PCR orthogonal truth (Figure 3) | 99.7%                                                    | 97.1%                                                    | Aggregated recall in PMS2 high-homology region             |
| MRJD high sensitivity mode\* vs long-read-based approach                   | >99.7%                                                   | >99.7%                                                   | Independent 147-sample concordance analysis                |
| Spurious call rate analysis                                                | -                                                        | -                                                        | <0.7% spurious call rate                                   |
| Non-cell-line validation                                                   | Detected all expected clinically relevant small variants | Detected all expected clinically relevant small variants | 22 non-cell-line samples (Broad Clinical Labs + Tempus AI) |

\*For details on MRJD calling in high sensitivity mode refer to [DRAGEN user guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/multi-region-joint-detection#high-sensitivity-mode)

![MRJD Figure 3 performance summary](https://assets.illumina.com/content/dam/illumina-marketing/images/genomics-research/articles/pms2-small-variant-detection/PMS2-small-variant-detection-figure3.png)


# Ploidy Estimator

Generated on **2026-05-07**

## Ploidy Estimator

The Ploidy Estimator uses reads from the mapper/aligner to calculate the sequencing depth of coverage for each autosome and allosome in the human genome. The sex karyotype of the sample is then estimated using the ratios of the median sex chromosome coverages to the median autosomal coverage. The sex karyotype is estimated based on the range the ratios fall in. If the ratios are outside all expected ranges, then the Ploidy Estimator does not determine a sex karyotype. 100% recall and precision were achieved across all chromosomes in 21 samples with various triploid and sex chromosome aneuploidies as well as 2 normal samples. Accuracy was assessed using PASS calls from the ploidy.vcf. For more see [DRAGEN documentation](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/ploidy-calling/ploidy-estimator).

**Example: HG00142 trisomy event**

HG00142 is reported in Polaris whole chromosome events as an example with chromosome-level copy number gains consistent with trisomy.

<details>

<summary>HG00142 chromosome-level ploidy metrics (click to expand)</summary>

| CHROM | median | mad | z-score |
| ----- | ------ | --- | ------- |
| chr11 | 130.4  | 9.6 | 169.4   |
| chr13 | 129.6  | 9.7 | 144.9   |

</details>

![HG00142 whole chromosome trisomy example](https://raw.githubusercontent.com/Illumina/Polaris/master/cohorts/1000_genomes/plots/chromosome_level_events_trisomies_3_2.png)

Source: Polaris whole chromosome trisomy events, sample HG00142.


# Tandem Repeats

Short tandem repeats (STRs) are regions of the genome consisting of repetitions of short DNA segments called repeat units. STRs can expand to lengths beyond the normal range and cause mutations called repeat expansions. For more information, refer to the [DRAGEN user guide](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/repeat-expansions).

DRAGEN-STR ships with two STR catalogs. The *default* catalog contains a restricted number of well-studied loci whose expansion is linked with various diseases. The *expanded* catalog contains \~174,000 highly-polymorphic STR loci located in and around genes, and is more suitable for whole-genome explorations.

## STR size genotyping accuracy in the expanded catalog

Accuracy on the expanded catalog loci was evaluated against the GIAB tandem repeat benchmark (v1.0) (1), using Truvari (2) for variant comparison.

**DRAGEN**: DRAGEN 4.5.4 | **Truthset**: HG002 GIABTR v1.0 | **Reference**: GRCh38

<details>

<summary>Standard WGS STR table (click to expand)</summary>

| Subtype | Recall | Precision | F1-score | FN   | FP  |
| ------- | ------ | --------- | -------- | ---- | --- |
| STR     | 0.9617 | 0.9871    | 0.9742   | 2007 | 664 |

</details>

![STR for Standard WGS](/files/u2qsHsasArrdVOd4n79B)

## Classification of samples with known pathogenic STR expansions

We sequenced 40 Coriell cell lines with known STR expansions and compared DRAGEN's STR classification against orthogonal validation methods. The swimplot shows long-allele size distributions in repeat-count units. Dots are colored by [STRipy database](https://www.stripy.org/database) range classification (normal/intermediate/pathogenic) according to each locus thresholds and orthogonally validated size prediction. Shaded regions and classification thresholds from Ibañez K. et al., Lancet Neurology 21, 234–245 (2022).

<details>

<summary>STR classification metrics table (click to expand)</summary>

| Subtype                             | Recall | Precision | F1-score |
| ----------------------------------- | ------ | --------- | -------- |
| STR Classification Accuracy Metrics | 0.9824 | 1.0000    | 0.9911   |

</details>

STR lengths distribution across 20 loci for 40 Coriell samples with pathogenic STRs. Shaded regions show normal, intermediate, and pathogenic ranges (STRipy database, motif-length scaled). Dots are colored by classification.

![STR size distribution, 20 loci, bp, filled intermediate](/files/q6WyQWO4BOcp3VU4pvAA)

#### References

1. <https://www.nature.com/articles/s41587-024-02225-z>
2. <https://link.springer.com/article/10.1186/s13059-022-02840-6>


# DRAGEN Reference Support

DRAGEN supports the construction of reference hash tables for both human and non-human reference genomes. The reference autodetect feature of DRAGEN is able to recognize the reference hash tables build on the four Human reference genomes: hg19 (`hg19`), GRCh37/hs37d5 (`hs37d5`), GRCh38/hs38d1(`hg38`), and T2T-CHM13v2.0 (`chm13`).

DRAGEN supports pangenome reference hash tables which extend the reference genomes with alternative variant paths from a sample cohort used to construct the pangenome reference. A pangenome-based reference improves the mapping accuracy of Illumina reads in the “Difficult-to-Map Regions” of the genome and the downstream variant calling.

Pre-built human references are available for download at [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

The pangenome is the recommended reference for germline human analyses. The accuracy achieved with pangenome references are highlighted in the plot below.

![](/files/qMMmfpxS6t19T2w2HyvN)

In the following tables we summarize the reference support for each DRAGEN component and the recommended reference type for each component.

| Pipeline                   | hg19      | hs37d5    | hg38 | chm13     | Recommended (Human) | Recommended (Non-Human) |
| -------------------------- | --------- | --------- | ---- | --------- | ------------------- | ----------------------- |
| **Germline**               | Yes       | Yes       | Yes  | *Note1\** | Pangenome           | Linear                  |
| **Somatic**                | Yes       | Yes       | Yes  | *Note2\** | Linear              | Linear                  |
| **RNA**                    | Yes       | Yes       | Yes  | *Note1\** | Linear              | Linear                  |
| **Methyl 5-base Germline** | Yes       | Yes       | Yes  | No        | Pangenome           | Linear                  |
| **Methyl 5-base Somatic**  | Yes       | Yes       | Yes  | No        | Linear              | Linear                  |
| **Methyl TruSeq**          | Yes       | Yes       | Yes  | No        | Linear              | Linear                  |
| **scRNA**                  | Yes       | Yes       | Yes  | *Note1\** | Linear              | Linear                  |
| **TruPath**                | *Note3\** | *Note3\** | Yes  | No        | Pangenome           | Linear                  |
| **Annotation**             | Yes       | Yes       | Yes  | No        | Pangenome           | Linear                  |

*Note1\* DRAGEN™ supports the component execution; however, the component's accuracy has not been established. Validated only for SNV. Accuracy not validated for CNV, SV, Joint Genotyping, HLA, gVCFGenotyper, RNA and scRNA. Not supported for STR, Targeted Callers, MRJD*

*Note2\* DRAGEN™ supports the component execution; however, the component's accuracy has not been established.*

*Note3\* Experimental use only. The component's functionality and accuracy has not been established with this reference.*

### Component availability per pipeline

| Pipeline    | Components                                                                                                  |
| ----------- | ----------------------------------------------------------------------------------------------------------- |
| Germline    | SNV, CNV, SV, STR, Targeted Callers, MRJD, RNA, De Novo, Joint Genotyping, Biomarkers (HLA), gVCF genotyper |
| Somatic     | SNV, UMI SNV, CNV, SV                                                                                       |
| Methylation | 5-base, TruSeq DNA Methyl, TruSeq Methyl Capture                                                            |
| Single cell | RNA, ATAC                                                                                                   |
| TruPath     | SNV, CNV, SV, STR, Targeted Callers, MRJD                                                                   |
| Annotation  | Nirvana                                                                                                     |

By default, DRAGEN will error out if a linear reference is provided when running a component for which a pangenome reference is recommended as listed in the above table. If you are sure that a linear reference is desired, the error can be suppressed by setting `--validate-pangenome-reference=false`.

See [Prepare a Reference Genome](/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support/prepare-a-reference-genome) for how to build a custom reference genome.


# Prepare a Reference Genome

Before a reference genome can be used with DRAGEN, it must be converted from FASTA format into a custom binary format for use with the DRAGEN hardware. The options used in this preprocessing step offer tradeoffs between performance and mapping quality.

Pre-built DRAGEN reference genomes are available for download in the Illumina customer portal. If you find that performance and mapping quality with these are adequate, there is a good chance that you can simply work with these supplied reference genomes. Depending on your read lengths and other particular aspects of your application, you may be able to improve mapping quality and/or performance by tuning the reference preprocessing options.

## Hash Table Background

The DRAGEN mapper extracts many overlapping seeds (subsequences or K-mers) from each read, and looks up those seeds in a hash table residing in memory on its PCIe card, to identify locations in the reference genome where the seeds match. Hash tables are ideal for extremely fast lookups of exact matches. The DRAGEN hash table must be constructed from a chosen reference genome using the `--build-hash-table option`, which extracts many overlapping seeds from the reference genome, populates them into records in the hash table, and saves the hash table as a binary file.

### Automatic Reference Detection

DRAGEN will attempt to detect the provided reference in order to automatically apply recommended resources and settings. There are four human references that DRAGEN can detect: hg38, hg19, hs37d5, and chm13v2. DRAGEN is able to detect references that contain a subset of the primary contigs from one of these references, as long as the names and lengths of the detected contigs are consistent with the names and lengths from the standarad assemblies of these references.

In detail, automatic reference detection operates as follows:

We define a primary contig of a human genome to be an autosome (1-22) or sex chromosome (X,Y). Let F be the input fasta. For each reference genome R in hg38, hg19, hs37d5, and chm13v2, DRAGEN checks if there are any contigs in F that have the same name and length as a primary contig in R, and that there are no contigs in F that have the same name as a contig in R, but with different length. If these conditions hold for exactly one of hg38, hg19, hs37d5, and chm13v2, then that reference is detected and resources may be applied automatically.

The DRAGEN hash table builder will automatically apply decoy contigs and mask bed files to detected reference. Other pipelines may also apply automatic resources. For example variant callers may apply machine learning models and target bed files.

#### Naming Conventions

In order for DRAGEN to correctly detect the provided reference, it is important to use the standard naming conventions for each of the four human assemblies that DRAGEN detects:

| Assembly            | Autosome and Sex Chromosome Names |
| ------------------- | --------------------------------- |
| hg38, hg19, chm13v2 | chr1-chr22, chrX, chrY            |
| hs37d5              | 1-22, X, Y                        |

### Reference Seed Interval

The size of the DRAGEN hash table is proportionate to the number of seeds populated from the reference genome. The default is to populate a seed starting at every position in the reference genome, ie, roughly 3 billion seeds from a human genome. This default requires at least 32 GB of memory on the DRAGEN PCIe board.

To operate on larger, nonhuman genomes or to reduce hash table congestion, it is possible to populate less than all reference seeds using the `--ht-ref-seed-interval` option to specify an average reference interval. The default interval for 100% population is `--ht-ref-seed-interval 1`, and 50% population is specified with `--ht-ref-seed-interval 2`. The population interval does not need to be an integer. For example, `--ht-ref-seed-interval 1.2` indicates 83.3% population, with mostly 1-base and some 2-base intervals to achieve a 1.2 base interval on average.

### Hash Table Occupancy

It is characteristic of hash tables that they are allocated a certain size, but always retain some empty records, so they are less than 100% occupied. A healthy amount of empty space is important for quick access to the DRAGEN hash table. Approximately 90% occupancy is a good upper bound. Empty space is important because records are pseudo-randomly placed in the hash table, resulting in an abnormally high number of records in some places. These congested regions can get quite large as the percentage of empty space approaches zero, and queries by the DRAGEN mapper for some seeds can become increasingly slow.

### Hash Table / Seed Length

The hash table is populated with reference seeds of a single common length. This primary seed length is controlled with the `--ht-seed-len` option, which defaults to 21.

The longest primary seed supported is 27 bases when the table is 8 GB to 31.5 GB in size. Generally, longer seeds are better for run time performance, and shorter seeds are better for mapping quality (success rate and accuracy). A longer seed is more likely to be unique in the reference genome, facilitating fast mapping without needing to check many alternative locations. But a longer seed is also more likely to overlap a deviation from the reference (variant or sequencing error), which prevents successful mapping by an exact match of that seed (although another seed from the read may still map), and there are fewer long seed positions available in each read.

Longer seeds are more appropriate for longer reads, because there are more seed positions available to avoid deviations.

Seed Length Recommendations

| Value for `--ht-seed-len` | Read Length           |
| ------------------------- | --------------------- |
| 21                        | 100 bp to 150 bp      |
| 17 to 19                  | shorter reads (36 bp) |
| 27                        | 250+ bp               |

### Hash Table / Seed Extensions

Due to repetitive sequences, some seeds of any given length match many locations in the reference genome. DRAGEN uses a unique mechanism called seed extension to successfully map such high-frequency seeds. When the software determines that a primary seed occurs at many reference locations, it extends the seed by some number of bases at both ends, to some greater length that is more unique in the reference.

For example, a 21-base primary seed may be extended by 7 bases at each end to a 35-base extended seed. A 21-base primary seed may match 100 places in the reference. But 35-base extensions of these 100 seed positions may divide into 40 groups of 1-3 identical 35-base seeds. Iterative seed extensions are also supported, and are automatically generated when a large set of identical primary seeds contains various subsets that are best resolved by different extension lengths.

The maximum extended seed length, by default equal to the primary seed length plus 128, can be controlled with the `--ht-max-ext-seed-len` option. For example, for short reads, it is advisable to set the maximum extended seed shorter than the read length, because extensions longer than the whole read can never match.

It is also possible to tune how aggressively seeds are extended using the following options (advanced usage):

`--ht-cost-coeff-seed-len`

`--ht-cost-coeff-seed-freq`

`--ht-cost-penalty`

`--ht-cost-penalty-incr`

There is a tradeoff between extension length and hit frequency. Faster mapping can be achieved using longer seed extensions to reduce seed hit frequencies, or more accurate mapping can be achieved by avoiding seed extensions or keeping extensions short, while tolerating the higher hit frequencies that result. Shorter extensions can benefit mapping quality both by fitting seeds better between SNPs, and by finding more candidate mapping locations at which to score alignments. The default extension settings along with default seed frequency settings, lean aggressively toward mapping accuracy, with relatively short seed extensions and high hit frequencies.

The defaults for the seed frequency options are as follows:

| Option                      | Default |
| --------------------------- | ------- |
| `--ht-cost-coeff-seed-len`  | 1       |
| `--ht-cost-coeff-seed-freq` | 0.5     |
| `--ht-cost-penalty`         | 0       |
| `--ht-cost-penalty-incr`    | 0.7     |
| `--ht-max-seed-freq`        | 16      |
| `--ht-target-seed-freq`     | 4       |

### Seed Frequency Limit and Target

One primary or extended seed can match multiple places in the reference genome. All such matches are populated into the hash table, and retrieved when the DRAGEN mapper looks up a corresponding seed extracted from a read. The multiple reference positions are then considered and compared to generate aligned mapper output. However, the DRAGEN software enforces a limit on the number of matches, or frequency, of each seed, which is controlled with the `--ht-max-seed-freq option`. By default, the frequency limit is 16. In practice, when the software encounters a seed with higher frequency, it extends it to a sufficiently long secondary seed that the frequency of any particular extended seed pattern falls within the limit. However, if a maximum seed extension would still exceed the limit, the seed is rejected, and not populated into the hash table. Instead, a single High Frequency record is populated.

This seed frequency limit does not tend to impact DRAGEN mapping quality notably, for two reasons. First, because seeds are rejected only when extension fails, only extremely high-frequency primary seeds, typically with many thousands of matches are rejected. Such seeds are not very useful for mapping. Second, there are other seed positions to check in a given read. If another seed position is unique enough to return one or more matches, the read can still be properly mapped. However, if all seed positions were rejected as high frequency, often this means that the entire read matches similarly well in many reference positions, so even if the read were mapped it would be an arbitrary choice, with very low or zero MAPQ.

Thus, the default frequency limit of 16 for `--ht-max-seed-freq` works well. However, it may be decreased or increased, up to a maximum of 256. A higher frequency limit tends to marginally increase the number of reads mapped (especially for short reads), but commonly the additional mapped reads have very low or zero MAPQ. This also tends to slow down DRAGEN mapping, because correspondingly large numbers of possible mappings are occasionally considered.

In addition to a frequency limit, a target seed frequency can be specified with `--ht-target-seed-freq` option. This target frequency is used when extensions are generated for high frequency primary seeds. Extension lengths are chosen with a preference toward extended seed frequencies near the target. The default of 4 for `--ht-target-seed-freq` means that the software is biased toward generating shorter seed extensions than necessary to map seeds uniquely.

### References with ALT contigs

When building a reference hash table from a fasta with ALT contigs, it may be desired to mask certain regions of high similarity, or to establish a liftover realtionships between primary and alternate contigs. The recommended approach is masking, as described in the Map-Align section. When hg19 or hg38 alt contigs are detected, the hash table builder will require a liftover file or a bed file to mask the alt contigs. If non are provided, a mask bed file from `<INSTALL_PATH>/resources/ht_builder/` will be used automaticaly.

### Masked References

DRAGEN has adopted a masked approach to handle native reference ALT contigs, where strategic regions are masked to increased accuracy. The hash table builder will build the mapper hash table as if the regions that were specified in the argument for `ht-mask-bed` were masked with N's. The hash table builder will only allow setting one of `ht-mask-bed` or `ht-alt-liftover`. Each line in the bed file is expected to contain a contig name, start position (0-based), and end position (1-based), seperated by a single tab or space. Lines that start with # are ignored by the hash table builder to allow commenting. Any line with a contig name that is not found in the input fasta is skipped and logged to the DRAGEN log file. Likewise, lines that describe empty intervals are skipped. If all lines are skipped this way, the hash table builder will issue an error and abort, unless the mask bed file was automatically applied (see Automatic masking). The hash table builder will always issue an error and abort if an interval described in the BED file is outside of the range of the corresponding contig in the fasta. Lines that are not skipped are written to a file called mask.bed that will be present in the hash table output directory, and whose digest will appear in hash\_table.cfg. This file is used when a reference is loaded to the FPGA card to dynamically mask reference.bin.

### Automatic masking

When running from a fasta for which hg38 or hg19 is detected (See Automatic Reference Detection), and no argument for `ht-mask-bed` or `ht-alt-liftover` was provided, the hash table builder will automatically apply the corresponding bed file for the detected reference from `<INSTALL_PATH>/resources/ht_builder/`. Note that the hash table builder will identify alt contigs by name. So when running from an input fasta that contains alt contig with standard names but modified base content, it is recommended to suppress automatic masking by setting `ht-suppress-mask=true` or by passing a custom mask bed file to `ht-mask-bed`.

### Handling Decoy Contigs

The behavior of DRAGEN with respect to the handling of decoy contigs in the reference has changed since version 2.6.

Starting with DRAGEN 3.x, DRAGEN's hash table builder automatically detects the absence of the decoy contigs from the reference and adds it to the FASTA file, prior to building the hash table. The decoys file is found at `<INSTALL_PATH>/resources/ht_builder/hs_decoys.fa.gz`. If the reference is missing the decoy contigs, then the reads which map to the decoy contigs are artificially marked as unmapped in the output BAM (because the original reference does not have the decoy contig). This results in an artificially lower mapping rate, however, the accuracy of variant calling is improved thanks to removing false positive caused by decoy reads.

Illumina recommends using this feature by default. However, you can to set the `--ht-suppress-decoys` option to true to suppress adding these decoys to the hash table.

The table below describes the difference in behavior between older DRAGEN versions (2.6 and earlier) and DRAGEN 3.x versions with respect to the handling of decoy contigs in the hash table builder:

| DRAGEN Behavior                                         | DRAGEN 2.6 and earlier versions                                                                                                                                                                                                | DRAGEN 3.0 and later versions                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Reference does not include the decoy contigs (eg, hg19) | <p>Decoy reads mismap elsewhere in the genome due to the lack of contigs in the reference.<br><br>Artificially higher mapping rate.<br><br>False positive calls in noisy regions to which the decoy contigs are mismapped.</p> | <p>DRAGEN automatically detects the absence of the decoy contig from the reference and adds it to the FASTA file.<br><br>Artificially lower mapping rate because decoy reads which map to the decoy contigs are artificially marked as unmapped in the output BAM (because the original reference does not have the decoy contig).<br><br>False positive calls are avoided thanks to adding the decoy contigs under the hood. Therefore this helps variant calling.</p> |
| Reference includes the decoy contigs (eg, hs37d5)       | <p>Decoy reads map to the decoy contigs.<br><br>High mapping rate.<br><br>No false positive calls caused by decoy reads because decoy reads map to the right place</p>                                                         | <p>Decoy reads map to the decoy contigs.<br><br>High mapping rate.<br><br>No false positive calls caused by decoy reads because decoy reads map to the right place</p>                                                                                                                                                                                                                                                                                                  |

## Prepare a Pangenome Reference

DRAGEN analysis is capable of mapping on a pangenome hash table. The pangenome hash table introduces alternate graph paths to the linear reference hash table to represent more broadly the allelic diversity of the population over the whole genome or in specific regions defined in a bed file. Gain on accuracy from this methodology has been described in scientific blogs available on the [Illumina Genomics Research Hub site](https://www.illumina.com/science/genomics-research.html). Mutigenome hash tables for CHM13\_v2, hg38, hg19 and hs37d5 assemblies are available on the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

See [DRAGEN Multigenome Mapper](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/dna-map-align#dragen-multigenome-mapper) for information on the multigenome mapping method.

It is possible to build a custom pangenome reference in order to:

* customize the released pangenome hash table with custom bed files or hash table builder options. A set of bed files are available in the resource files on the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).
* generate a population-specific-pangenome hash table from pangenome msVCF generated from the BSSH app.
* generate a human or non-human pangenome hash table from customer-provided msVCF.

The input files required are a single multi-sample VCF file containing the set of population variants, and optionally bed files restricting graph to some region. The generated files, including hash\_table.cmp and associated files in the specified output directory, can then be used as the reference hash table for the DRAGEN mapper. DRAGEN software supports the tool on human reference with files available on the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). For non-human, the user provides the required resource files.

### Usage

To enable the pangenome hash table builder, example command usage is :

`dragen --build-hash-table true (required) --ht-graph-msvcf-file <path to a multi-sampple VCF file (required for pangenome reference) --ht-reference <reference.fasta> (required) --ht-graph-extra-kmer-bed < graph.bed> (optional) --ht-mask-bed <mask.bed> (optional) --ht-graph-exclusion-bed <exclusion bed> (optional) --output-directory <DIR> (required) [options]`

### Inputs

#### Set of population variants, in a multi-sample VCF (msVCF)

The custom pangenome hash table builder tool uses a set of population variants provided by the user to generate a pangenome hash table. The variants must be specified in VCF format, in a single multi-sample VCF (msVCF) file containing the variants for a set of individuals. This multi-sample VCF file must have specific formatting described below.

#### Specific msVCF input formatting

The custom pangenome hash table builder tool only supports msVCF file input respecting the format described below:

* msVCF compliant with 4.2 VCF format specification
* with variants positionally sorted in the same contig order as the main FASTA reference genome provided in --ht-reference
* records shall include diploid or haploid GT calls
* supports multi-allelic variants merged in multi-line or separated in multiple lines
* with the following FILTER codes, non-PASS records are ignored:
  * \##FILTER=\<ID=PASS,Description="All filters passed">
* with the following FORMAT field :
  * \##FORMAT=\<ID=GT,Number=1,Type=String,Description="Genotype">
* for better results, we recommend variants to be left-aligned.
* maximum number of recommended samples in the msVCF is 256. Higher number may lead to very high memory usage at hash table creation.

> Note: INFO/FORMAT subfields must be defined in the header. Events with undefined subfields are ignored.

To build a high-performance custom genome it is highly recommended to use long read sequencing data. We recommend using external tools such as Whatshap (<https://github.com/whatshap/whatshap>) to generate phased input. DRAGEN analysis leverages the phasing information to reconstruct population haplotypes.

#### Reference genome

A reference genome in FASTA format must be provided. Reference genomes are available to download from the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

> Note: the reference genome provided as input must be the same as the one used to generate the input phased msVCF. If the msVCF contains variants from regions not present in the fasta file, the pangenome reference builder will stop with an error.

#### Exclusion bed file (optional)

This bed file is used to filter out regions of the msVCF file. Variants that fall within intervals defined in the "Graph exclusion bed" file will be ignored and not used in any part of the pangenome reference builder. The result will be the same as if the input msVCF did not contain any variants in the regions defined in the exclusion bed. The file is optional, by default every variants in the msVCF file will be used. Exclusion bed files are available to download from [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

A custom exclusion bed file can also be provided given the following format: tab delimited with first three columns being: contig name, start position, end position. Any line with a contig name that is not found in the input FASTA is skipped. Any lines that describe empty intervals are skipped.

> Note: records of the exclusion bed file provided must be from the same build as the reference genome used to build the pangenome reference.

#### Extra kmer bed file (optional)

This file is used to define regions in the genome where extra seeds will be indexed in the hash table. By default, only seed extracted from the primary reference will be extracted and saved in the reference hash table for mapping. This option will additionally generate seeds from population variants in the defined regions. It is recommended to include the expected difficult regions in this bed file. Extra-kmer-bed files are available to download from [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) for the human hg38, hg19, hs37d5, and chm13 references.

An Extra-kmer-bed bed file can also be provided given the following format: tab delimited with first three columns being: contig name, start position, end position. Any line with a contig name that is not found in the input FASTA is skipped. Any lines that describe empty intervals are skipped.

> Note: records of the Extra-kmer-bed file provided must be from the same build as the reference genome used to build the graph reference.

#### Mask bed file (recommended)

A mask bed file must be provided in order to mask certain regions of high similarity between primary and alternate contigs present in the main genome FASTA. Mask bed files are available to download from the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

A custom mask bed file can also be provided given the following format: tab delimited with first three columns being: contig name, start position, end position. Any line with a contig name that is not found in the input FASTA is skipped. Any lines that describe empty intervals are skipped.

> Note: records of the mask bed file provided must be from the same build as the reference genome used to build the graph reference.

### Command line options

| Option                    | Required             | Description                                                              |
| ------------------------- | -------------------- | ------------------------------------------------------------------------ |
| --build-hash-table        | Yes                  | Set to true                                                              |
| --ht-graph-msvcf-file     | Yes                  | Path to the multi-sample VCF file containing population variants         |
| --ht-reference            | Yes                  | Path to the reference genome FASTA file.                                 |
| --ht-graph-extra-kmer-bed | No                   | Path to the extra kmer bed file                                          |
| --ht-mask-bed             | No (but recommended) | Path to the mask bed file                                                |
| --ht-graph-exclusion-bed  | No                   | Path to the exclusion bed file                                           |
| --output-directory        | Yes                  | Specify the directory where all related hash table files will be written |

> Note: The custom graph reference hash table end to end pipeline will return an error if options --ht-alt-liftover or --ht-allow-mask-and-liftover are specified.

### Output

The hash table builder generates the following outputs:

| File                   | Description                                                                                                                                                                                                                                                                                                                                                                            |
| ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| reference.bin          | The reference sequences, encoded in 4 bits per base. Four-bit codes are used, so the size in bytes is roughly half the reference genome size. In between reference sequences, N are trimmed and padding is automatically inserted. For example, hg19 has 3,137,161,264 bases in 93 sequences. This is encoded in 1,526,285,312 bytes = 1.46 GB, where 1 GB means 1 GiB or 2^30^ bytes. |
| hash\_table.cmp        | Compressed hash table. The hash table is decompressed and used by the DRAGEN mapper to look up primary seeds with length specified by the `--ht-seed-len` option and extended seeds of various lengths.                                                                                                                                                                                |
| hash\_table.cfg        | A list of parameters and attributes for the generated hash table, in a text format. This file provides key information about the reference genome and hash table.                                                                                                                                                                                                                      |
| hash\_table.cfg.bin    | A binary version of hash\_table.cfg used to configure the DRAGEN hardware.                                                                                                                                                                                                                                                                                                             |
| hash\_table\_stats.txt | A text file listing extensive internal statistics on the constructed hash including the hash table occupancy percentages. This table is for information purposes. It is not used by other tools.                                                                                                                                                                                       |
| mask.bed               | Present only for masked hash tables. A tab delimeted bed file that describes the masked regions. Contains all lines from the input bed file that are not comment lines, lines that describe empty intervals, or lines with contig names that were not found in the input fasta.                                                                                                        |

## Prepare a linear Reference

### Usage

Use the `--build-hash-table` option to transform a reference FASTA into the hash table for DRAGEN mapping. It takes as input a FASTA file (multiple reference sequences being concatenated) and a preexisting output directory. Build command usage is as follows:

```
dragen --build-hash-table true [options] --ht-reference
<reference.fasta> --output-directory <outdir>
```

### Input

The `--ht-reference` and `--output-directory` options are required for building a hash table. The `--ht‑reference` option specifies the path to the reference FASTA file, while `--output-directory` specifies a preexisting directory where the hash table output files are written. Illumina recommends organizing various hash table builds into different folders. As a best practice, folder names should include any nondefault parameter settings used to generate the contained hash table. The sequence names in the reference FASTA file must be unique.

### Command line options

| Option             | Required             | Description                                                                                                                                                   |
| ------------------ | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --build-hash-table | Yes                  | Set to true                                                                                                                                                   |
| --ht-reference     | Yes                  | Path to the reference genome FASTA file.                                                                                                                      |
| --ht-mask-bed      | No (but recommended) | Path to the mask bed file. If not provided, the DRAGEN software automatically applies BED files for hg38 and hg19 from `<INSTALL_PATH>/resources/ht_builder`. |
| --output-directory | Yes                  | Specify the directory where all related hash table files will be written                                                                                      |

#### Liftover Based ALT-Aware Hash Tables

While masking is the recommended approach to dealing with ALT contigs, DRAGEN also supports a liftover based method. To enable liftover based ALT-aware mapping in DRAGEN, build the hash table with a liftover file by using the `--ht-alt-liftover` option. The hash table builder classifies each reference sequence as primary or alternate based on the liftover file, and packs primaries before alternates in reference.bin. SAM liftover files for hg38DH and hg19 are in the `<INSTALL_PATH>/resources/ht_builder` folder.

**Custom Liftover Files**

Custom liftover files can be used in place of those provided with DRAGEN. Liftover files must be SAM format, but no SAM header is required. SEQ and QUAL fields can be omitted ('\*'). Each alignment record should have an alternate haplotype reference sequence name as QNAME, indicating the RNAME and POS of its liftover alignment in a destination (normally primary assembly) reference sequence.

Reverse-complemented alignments are indicated by bit 0x10 in FLAG. Records flagged unmapped (0x4) or secondary (0x100) are ignored. The CIGAR may include hard or soft clipping, leaving parts of the ALT contig unaligned.

A single reference sequence cannot serve as both an ALT contig (appearing in QNAME) and a liftover destination (appearing in RNAME). Multiple ALT contigs can align to the same primary assembly location. Multiple alignments can also be provided for a single ALT contig (extras optionally be flagged 0x800 supplementary), such as to align one portion forward and another portion reverse-complemented. However, each base of the ALT contig only receives one liftover image, according to the first alignment record with an M CIGAR operation covering that base.

SAM records with QNAME missing from the reference genome are ignored, so that the same liftover file may be used for various reference subsets, but an error occurs if any alignment has its QNAME present but its RNAME absent.

## Options for advanced users

### Primary Seed Length

The `--ht-seed-len` option specifies the initial length in nucleotides of seeds from the reference genome to populate into the hash table. At run time, the mapper extracts seeds of this same length from each read, and looks for exact matches (unless seed editing is enabled) in the hash table.

The maximum primary seed length is a function of hash table size. The limit is k=27 for table sizes from 16 GB to 64 GB, covering typical sizes for whole human genome, or k=26 for sizes from 4 GB to 16 GB.

The minimum primary seed length depends mainly on the reference genome size and complexity. It needs to be long enough to resolve most reference positions uniquely. For whole human genome references, hash table construction typically fails with k < 16. The lower bound may be smaller for shorter genomes, or higher for less complex (more repetitive) genomes. The uniqueness threshold of `--ht-seed-len 16` for the 3.1Gbp human genome can be understood intuitively because log4(3.1 G) ≈ 16, so it requires at least 16 choices from 4 nucleotides to distinguish 3.1 G reference positions.

#### Accuracy Considerations

For read mapping to succeed, at least one primary seed must match exactly (or with a single SNP when edited seeds are used). Shorter seeds are more likely to map successfully to the reference, because they are less likely to overlap variants or sequencing errors, and because more of them fit in each read. So for mapping accuracy, shorter seeds are mainly better.

However, very short seeds can sometimes reduce mapping accuracy. Very short seeds often map to multiple reference positions, and lead the mapper to consider more false mapping locations. Due to imperfect modeling of mutations and errors by Smith-Waterman alignment scoring and other heuristics, occasionally these noise matches may be reported. Run time quality filters such as `--Aligner.aln_min_score` can control the accuracy issues with very short seeds.

#### Speed Considerations

Shorter seeds tend to slow down mapping, because they map to more reference locations, resulting in more work such as Smith-Waterman alignments to determine the best result. This effect is most pronounced when primary seed length approaches the reference genome's uniqueness threshold, eg, K=16 for whole human genome.

#### Application Considerations

**Read Length**---Generally, shorter seeds are appropriate for shorter reads, and longer seeds for longer reads. Within a short read, a few mismatch positions (variants or sequencing errors) can chop the read into only short segments matching the reference, so that only a short seed can fit between the differences and match the reference exactly. For example, in a 36 bp read, just one SNP in the middle can block seeds longer than 18 bp from matching the reference. By contrast, in a 250 bp read, it takes 15 SNPs to exceed a 0.01% chance of blocking even 27 bp seeds.

**Paired Ends**---The use of paired end reads can make longer seeds yield good mapping accuracy. DRAGEN uses paired end information to improve mapping accuracy, including with rescue scans that search the expected reference window when only one mate has seeds mapping to a given reference region. Thus, paired end reads have essentially twice the opportunity for an exact matching seed to find their correct alignments.

**Variant or Error Rate**---When read differences from the reference are more frequent, shorter seeds may be required to fit between the difference positions in a given read and match the reference exactly.

**Mapping Percentage Requirement**---If the application requires a high percentage of reads to be mapped somewhere (even at low MAPQ), short seeds may be helpful. Some reads that do not match the reference well anywhere are more likely to map using short seeds to find partial matches to the reference.

### Maximum Seed Length

The `--ht-max-ext-seed-len` option limits the length of extended seeds populated into the hash table. Primary seeds (length specified by `--ht-seed-len`) that match many reference positions can be extended to achieve more unique matching, which may be required to map seeds within the maximum hit frequency (`--ht-max-seed-freq`).

Given a primary seed length k, the maximum seed length can be configured between k and k+128. The default is the upper bound, k+128.

#### When to Limit Seed Extension

The `--ht-max-ext-seed-len` option is recommended for short reads, eg, less than 50 bp. In such cases, it is helpful to limit seed extension to the read length minus a small margin, such as 1-4 bp. For example, with 36 bp reads, setting `--ht-max-ext-seed-len` to 35 might be appropriate. This ensures that the hash table builder does not plan a seed extension longer than the read causing seed extension and mapping to fail at run time, for seeds that could have fit within the read with shorter extensions.

While seed extension can be similarly limited for longer reads, eg, setting `--ht-max-ext-seed-len` to 99 for 100 bp reads, there is little utility in this because seeds are extended conservatively in any event. Even with the default k+128 limit, individual seeds are only extended to the lengths required to fit under the maximum hit frequency (`--ht-max-seed-freq`), and at most a few bases longer to approach the target hit frequency (`‑‑ht‑target-seed-freq`), or to avoid taking too many incremental extension steps.

### Maximum Hit Frequency

The `--ht-max-seed-freq` option sets a firm limit on the number of seed hits (reference genome locations) that can be populated for any primary or extended seed. If a given primary seed maps to more reference positions than this limit, it must be extended long enough that the extended seeds subdivide into smaller groups of identical seeds under the limit. If, even at the maximum extended seed length (`--ht-max-ext-seed-len`), a group of identical reference seeds is larger than this limit, their reference positions are not populated into the hash table. Instead, a single High Frequency record is populated.

The maximum hit frequency can be configured from 1 to 256. However, if this value is too low, hash table construction can fail because too many seed extensions are needed. The practical minimum for a whole human genome reference, other options being default, is 8.

#### Accuracy Considerations

Generally, a higher maximum hit frequency leads to more successful mapping. There are two reasons for this. First, a higher limit rejects fewer reference positions that cannot map under it. Second, a higher limit allows seed extensions to be shorter, improving the odds of exact seed matching without overlapping variants or sequencing errors.

However, as with very short seeds, allowing high hit counts can sometimes hurt mapping accuracy. Most of the seed hits in a large group are not to the true mapping location, and occasionally one of these noise hits may be reported due to imperfect scoring models. Also, the mapper limits the total number of reference positions it considers, and allowing very high hit counts can potentially crowd out the actual best match from consideration.

#### Speed Considerations

Higher maximum hit frequencies slow down read mapping, because seed mapping finds more reference locations, resulting in more work, such as Smith-Waterman alignments, to determine the best result.

### Pangenome Reference

The DRAGEN Software enables the user to build a custom pangenome hash table from a set of population variants. The population variants are specified in a single multi-sample VCF file.

* `--ht-graph-msvcf-file: Input file containing list of population variants, in multi-sample VCF format.`

This replaces the previous options that were previously used to build a graph Reference that are now deprecated.

List of deprecated options :

* `--ht-pop-alt-contigs: Population based alternate contigs FASTA.`
* `--ht-pop-alt-liftover: Liftover SAM file of population alternate contigs.`
* `--ht-pop-snps: Population based SNPs VCF`

### ALT-Contigs

The following options control building hash tables from references with ALT-contigs. See References with ALT contigs for more information.

* `--ht-mask-bed`: Set a custom BED file that defines which regions to mask. If not provided, the DRAGEN software automatically applies BED files for hg38 and hg19 from `<INSTALL_PATH>/resources/ht_builder`.
* `--ht-alt-liftover`: Set a liftover file to build a liftover based ALT-aware hash table. SAM liftover files for hg38DH and hg19 are provided in `<INSTALL_PATH>/resources/ht_builder`.
* `--ht-allow-mask-and-liftover`: Allow the use of both `--ht-mask-bed` and `--ht-alt-liftover` together.
* `--ht-suppress-mask`: Suppress automatic detection of the default mask bed files when building the hash table.

### Decoy Contigs

* `--ht-decoys` The DRAGEN software automatically detects the use of hg19 and hg38 references and adds decoys to the hash table when they are not found in the FASTA file. Use the `--ht-decoys` option to specify the path to a decoys file. The default is `<INSTALL_PATH>/resources/ht_builder/hs_decoys.fa.gz`.
* `--ht-suppress-decoys`: Suppress automatic detection of the default decoys file when building the hash table.

### Processing Options

* `--ht-num-threads` The `--ht-num-threads` option determines the maximum number of worker CPU threads that are used to speed up hash table construction. The default for this option is 8, with a maximum of 32 threads allowed. If your server supports execution of more threads, it is recommended that you use the maximum. For example, the DRAGEN servers contain 24 cores that have hyperthreading enabled, so a value of 32 should be used. When using a higher value, adjust `--ht-max-table-chunks` needs to be adjusted as well. The servers have 128 GB of memory available.
* `--ht-max-table-chunks` The `--ht-max-table-chunks` option controls the memory footprint during hash table construction by limiting the number of \~1 GB hash table chunks that reside in memory simultaneously. Each additional chunk consumes roughly twice its size (\~2 GB) in system memory during construction. The hash table is divided into power-of-two independent chunks, of a fixed chunk size, X, which depends on the hash table size, in the range 0.5 GB < X ≤ 1 GB. For example, a 24 GB hash table contains 32 independent 0.75 GB chunks that can be constructed by parallel threads with enough memory and a 16 GB hash table contains 16 independent 1 GB chunks. The default is `--ht-max-table-chunks` equal to `--ht-num-threads`, but with a minimum default `--ht-max-table-chunks` of 8. It makes sense to have these two options match, because building one hash table chunk requires one chunk space in memory and one thread to work on it. Nevertheless, there are build-speed advantages to raising `--ht-max-table-chunks` higher than `--ht-num-threads`, or to raising `--ht-num-threads` higher than `--ht-max-table-chunks`.

### Size Options

* `--ht-mem-limit` Memory Limit. The `--ht-mem-limit` option controls the generated hash table size by specifying the DRAGEN card memory available for both the hash table and the encoded reference genome. The `‑‑ht‑mem-limit` option defaults to 32 GB when the reference genome approaches WHG size, or to a generous size for smaller references. Normally there is little reason to override these defaults.
* `--ht-size` Hash Table Size. This option specifies the hash table size to generate, rather than calculating an appropriate table size from the reference genome size and the available memory (option `--ht-mem-limit`). Using default table sizing is recommended and using `--ht-mem-limit` is the next best choice.

### Seed Population Options

* `--ht-ref-seed-interval` Seed Interval. The `--ht-ref-seed-interval` option defines the step size between positions of seeds in the reference genome populated into the hash table. An interval of 1 (default) means that every seed position is populated, 2 means 50% of positions are populated, etc. Noninteger values are supported, eg, 2.5 yields 40% populated. Seeds from a whole human reference are easily 100% populated with 32 GB memory on DRAGEN boards. If a substantially larger reference genome is used, change this option.
* `--ht-soft-seed-freq-cap` and `--ht-max-dec-factor` Soft Frequency Cap and Maximum Decimation Factor for Seed Thinning. Seed thinning is an experimental technique to improve mapping performance in high-frequency regions. When primary seeds have higher frequency than the cap indicated by the `--ht-soft-seed-freq-cap option`, only a fraction of seed positions are populated to stay under the cap. The `--ht-max-dec-factor` option specifies a maximum factor by which seeds can be thinned. For example, `--ht-max-dec-factor 3` retains at least 1/3 of the original seeds. `--ht-max-dec-factor 1` disables any thinning. Seeds are decimated in careful patterns to prevent leaving any long gaps unpopulated. The idea is that seed thinning can achieve mapped seed coverage in high frequency reference regions where the maximum hit frequency would otherwise have been exceeded. Seed thinning can also keep seed extensions shorter, which is also good for successful mapping. Based on testing to date, seed thinning has not proven to be superior to other accuracy optimization methods.
* `--ht-rand-hit-hifreq` and `--ht-rand-hit-extend` Random Sample Hit with HIFREQ Record and EXTEND Record. Whenever a HIFREQ or EXTEND record is populated into the hash table, it stands in place of a large set of reference hits for a certain seed. Optionally, the hash table builder can choose a random representative of that set, and populate that HIT record alongside the HIFREQ or EXTEND record. Random sample hits provide alternative alignments that are very useful in estimating MAPQ accurately for the alignments that are reported. They are never used outside of this context for reporting alignment positions, because that would result in biased coverage of locations that happened to be selected during hash table construction. To include a sample hit, set `--ht-rand-hit-hifreq` to 1. The `--ht-rand-hit-extend` option is a minimum pre-extension hit count to include a sample hit, or zero to disable. Modifying these options is not recommended.

### Seed Extension Control

DRAGEN seed extension is dynamic, applied as needed for particular K-mers that map to too many reference locations. Seeds are incrementally extended in steps of 2--14 bases (always even) from a primary seed length to a fully extended length. The bases are appended symmetrically in each extension step, determining the next extension increment if any.

There is a potentially complex seed extension tree associated with each high frequency primary seed. Each full tree is generated during hash table construction and a path from the root is traced by iterative extension steps during seed mapping. The hash table builder employs a dynamic programming algorithm to search the space of all possible seed extension trees for an optimal one, using a cost function that balances mapping accuracy and speed. The following options define that cost function:

* `--ht-target-seed-freq` Target Hit Frequency. The `--ht-target-seed-freq` option defines the ideal number of hits per seed for which seed extension should aim. Higher values lead to fewer and shorter final seed extensions, because shorter seeds tend to match more reference positions.
* `--ht-cost-coeff-seed-len` Cost Coefficient for Seed Length The `--ht-cost-coeff-seed-len` option assigns the cost component for each base by which a seed is extended. Additional bases are considered a cost because longer seeds risk overlapping variants or sequencing errors and losing their correct mappings. Higher values lead to shorter final seed extensions.
* `--ht-cost-coeff-seed-freq` Cost Coefficient for Hit Frequency. The `--ht-cost-coeff-seed-freq` option assigns the cost component for the difference between the target hit frequency and the number of hits populated for a single seed. Higher values result primarily in high-frequency seeds being extended further to bring their frequencies down toward the target.
* `--ht-cost-penalty` Cost Penalty for Seed Extension. The `--ht-cost-penalty` option assigns a flat cost for extending beyond the primary seed length. A higher value results in fewer seeds being extended at all. Current testing shows that zero (0) is appropriate for this parameter.
* `--ht-cost-penalty-incr` Cost Increment for Extension Step. The `--ht-cost-penalty-incr` option assigns a recurring cost for each incremental seed extension step taken from primary to final extended seed length. More steps are considered a higher cost because extending in many small steps requires more hash table space for intermediate EXTEND records, and takes substantially more run time to execute the extensions. A higher value results in seed extension trees with fewer nodes, reaching from the root primary seed length to leaf extended seed lengths in fewer, larger steps.

## Pipeline Specific Hash Tables

### RNA-Seq

When building a hash table, DRAGEN configures the options for DNA analysis by default. To run RNA-Seq data, you must build an RNA-Seq hash table by setting `--ht-build-rna-hashtable` to true. If running RNA-Seq alignment, use the original `--output-directory` instead of the automatically generated subdirectory.

### CNV

If using the CNV pipeline, set `--ht-build-cnv-hashtable` to true. The command generates an additional Kmer hash map that is used in the CNV algorithm. Illumina recommends to always use the `--ht-build-cnv-hashtable` option, so you can perform CNV calling with the same hash table used for mapping and aligning.

### Methylation

To run the methylation pipeline, you must build a methylation-specific hash table. DRAGEN can build a single-pass or legacy multi-pass methylation hash table. Methylation runs using a single-pass hash table are completed faster than the legacy multipass hash tables. Single-pass hash tables are recommended for building methylation tables and running analyses.

| Hash Table Type | Hash Table Commands                                               |
| --------------- | ----------------------------------------------------------------- |
| single-pass     | `--ht-methylated-combined=true` `--ht-seed-len 27`                |
| multi-pass      | `--ht-methylated=true` `--ht-seed-len 27` `--ht-max-seed-freq 16` |

#### Single-pass

The following is an example of a single-pass hash table build. The example generates a combined hash table in your reference index folder under the methyl\_converted subdirectory.

`dragen --build-hash-table true \ --output-directory $REFDIR \ --ht-reference $FASTA \ --ht-num-threads 40 \ --ht-methylated-combined=true \ --ht-seed-len 27`

#### Multipass

Multi-pass methylation mapping requires building two special hash tables with reference bases converted from C to T in one table and G to A in the other table. The conversions are performed automatically when using the `--ht-methylated` command line option. The converted hash tables are generated in two subdirectories under the folder specified using the `--output-directory` command line option. The subdirectories are named CT\_converted and GA\_converted, corresponding with the base conversions. When using the hash tables for methylated alignment runs, make sure to refer to the `--output-directory` folder, not the subdirectories.

The base conversions remove a significant amount of information from the hash tables. You might need to use different hash table parameters than in a conventional hash table build. The following options are recommended for building hash tables for mammalian species.

`dragen --build-hash-table=true --output-directory $REFDIR --ht-reference $FASTA --ht-max-seed-freq 16 --ht-seed-len 27 --ht-num-threads 40 --ht-methylated=true`

## HLA

To run the HLA caller, an HLA-specific anchored reference hash table must be built. Set `--ht-build-hla-hashtable` to true. The command will create a `anchored_hla` subdirectory inside the `--output-directory`. The HLA-specific reference subdirectory can be built at the same time as the primary reference construction.

An HLA resource file is packaged with DRAGEN and located at the following path after installation: `<INSTALL_PATH>/resources/hla/HLA_resource.v1.fasta.gz`. This file is used by default when building the HLA-specific anchored hash table. A custom file can be specified with `--ht-hla-reference`. See the HLA section for more information [Using Custom HLA Reference Files](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/hla-typing#using-custom-hla-reference-files)


# DRAGEN DNA Pipeline

The DRAGEN DNA Pipeline accelerates the secondary analysis of NGS data by harnessing the tremendous power available on the DRAGEN Platform. The pipeline includes highly optimized algorithms for mapping, aligning, sorting, duplicate marking, and haplotype variant calling. In addition to haplotype variant calling, the pipeline supports calling of copy number and structural variants as well as detection of repeat expansions and targeted calls.

![](/files/hANbA0wBQnUV6lC3Swyf)


# DNA Mapping

### DNA Mapping

#### Seed Density Option

The *seed-density* option controls how many (normally overlapping) primary seeds from each read the mapper looks up in its hash table for exact matches. The maximum density value of 1.0 generates a seed starting at every position in the read, ie, (L-K+1) K-base seeds from an L-base read.

Seed density must be between 0.0 and 1.0. Internally, an available seed pattern equal or close to the requested density is selected. The sparsest pattern is one seed per 32 positions, or density 0.03125.

* **Accuracy Considerations**--Generally, denser seed lookup patterns improve mapping accuracy. However, for modestly long reads (eg, 50 bp+) and low sequencer error rates, there is little to be gained beyond the default 50% seed lookup density.
* **Speed Considerations**--Denser seed lookup patterns generally slow down mapping, and sparser seed patterns speed it up. However, when the seed mapping stage can run faster than the aligning stage, a sparser seed pattern does not make the mapper much faster.

**Relationship to Reference Seed Interval**

Functionally, a denser or sparser seed lookup pattern has an impact very similar to a shorter or longer reference seed interval (build hash table option `--ht-ref-seed-interval`). Populating 100% of reference seed positions and looking up 50% of read seed positions has the same effect as populating 50% of reference seed positions and looking up 100% of read seed positions. Either way, the expected density of seed hits is 50%.

More generally, the expected density of seed hits is the product of the reference seed density (the inverse of the reference seed interval) and the seed lookup density. For example, if 50% of reference seeds are populated and 33.3% (1/3) of read seed positions are looked up, then the expected seed hit density should be 16.7% (1/6).

DRAGEN automatically adjusts its precise seed lookup pattern to ensure it does not systematically miss the seed positions populated from the reference. For example, the mapper does not look up seeds matching only odd positions in the reference when only even positions are populated in the hash table, even if the reference seed interval is 2 and seed-density is 0.5.

#### Map Orientations Option

The `--Mapper.map-orientations` option is used in mapping reads for bisulfite methylation analysis. It is set automatically based on the value set for `‑‑methylation-protocol`.

The `--Mapper.map-orientations` option can restrict the orientation of read mapping to only forward in the reference genome, or only reverse-complemented. The valid values for `--map-orientations` are as follows.

* 0--Either orientation (default)
* 1--Only forward mapping
* 2--Only reverse-complemented mapping

If mapping orientations are restricted and paired end reads are used, the expected pair orientation can only be FR, not FF or RF.

#### Seed-Editing Options

Although DRAGEN primarily maps reads by finding exact reference matches to short seeds, it can also map seeds differing from the reference by one nucleotide by also looking up single-SNP edited seeds. Seed editing is usually not necessary with longer reads (100 bp+), because longer reads have a high probability of containing at least one exact seed match. This is especially true when paired ends are used, because a seed match from either mate can successfully align the pair. But seed editing can, for example, be useful to increase mapping accuracy for short single-ended reads, with some cost in increased mapping time. The following options control seed editing:

Seed Editing Options

| Command-Line Option Name    | Configuration File Option Name |
| --------------------------- | ------------------------------ |
| `--Mappper.seed-density`    | seed-density                   |
| `-Mapper.edit-mode`         | edit-mode                      |
| `--Mapper.edit-seed-num`    | edit-seed-num                  |
| `--Mapper.edit-read-len`    | edit-read-len                  |
| `--Mapper.edit-chain-limit` | edit-chain-limit               |

**edit-mode and edit-chain-limit**

The edit-mode and edit-chain-limit options control when seed editing is used. The following four edit-mode values are available:

| Mode | Description              |
| ---- | ------------------------ |
| 0    | No editing (default)     |
| 1    | Chain length test        |
| 2    | Paired chain length test |
| 3    | Full seed editing        |

Edit mode 0 requires all seeds to match exactly. Mode 3 is the most expensive because every seed that fails to match the reference exactly is edited. Modes 1 and 2 employ heuristics to look up edited seeds only for reads most likely to be salvaged to accurate mapping.

The main heuristic in edit modes 1 and 2 is a seed chain length test. Exact seeds are mapped to the reference in a first pass over a given read, and the matching seeds are grouped into chains of similarly aligning seeds. If the longest seed chain (in the read) exceeds a threshold edit-chain-limit, the read is judged not to require seed editing, because there is already a promising mapping position.

Edit mode 1 triggers seed editing for a given read using the seed chain length test. If no seed chain exceeds `edit-chain-limit` (including if no exact seeds match), then a second seed mapping pass is attempted using edited seeds. Edit mode 2 further optimizes the heuristic for paired-end reads. If either mate has an exact seed chain longer than `edit-chain-limit`, then seed editing is disabled for the pair, because a rescue scan is likely to recover the mate alignment based on seed matches from one read. Edit mode 2 is the same as mode 1 for single-ended reads.

**edit-seed-num and edit-read-len**

For edit modes 1 and 2, when the heuristic triggers seed editing, these options control how many seed positions are edited in the second pass over the read. Although exact seed mapping can use a densely overlapping seed pattern, such as seeds starting at 50% or 100% of read positions, most of the value of seed editing can be obtained by editing a much sparser pattern of seeds, even a nonoverlapping pattern. Generally, if a user application can afford to spend some additional amount of mapping time on seed editing, a greater increase in mapping accuracy can be obtained for the same time cost by editing seeds in sparse patterns for a large number of reads, than by editing seeds in dense patterns for a small number of reads.

Whenever seed editing is triggered, these two options request edit-seed-num seed editing positions, distributed evenly over the first edit-read-len bases of the read. For example, with 21-base seeds, edit-seed-num=6 and edit-read-len=100, edited seeds can begin at offsets {0, 16, 32, 48, 64, 80} from the 5' end, consecutive seeds overlapping by 5 bases. Because sequencing technologies often yield better base qualities nearer the (5') beginning of each read, this can focus seed editing where it is most likely to succeed. When a particular read is shorter than `edit-read-len`, fewer seeds are edited.

Seed editing is more expensive when the reference seed interval (build hash table option ‑-ht‑ref-seed-interval) is greater than 1. For edit modes 1 and 2, additional seed editing positions are automatically generated to avoid missing the populated reference seed positions. For edit mode 3, the time cost can increase dramatically because query seeds matching unpopulated reference positions typically miss and trigger editing.

### DNA Aligning

#### Smith-Waterman Alignment Scoring Settings

The first stage of mapping is to generate seeds from the read and look for exact matches in the reference genome. These results are then refined by running full Smith-Waterman alignments on the locations with the highest density of seed matches. This well-documented algorithm works by comparing each position of the read against all the candidate positions of the reference. These comparisons correspond to a matrix of potential alignments between read and reference. For each of these candidate alignment positions, Smith-Waterman generates scores that are used to evaluate whether the best alignment passing through that matrix cell reaches it by a nucleotide match or mismatch (diagonal movement), a deletion (horizontal movement), or an insertion (vertical movement). A match between read and reference provides a bonus, on the score, and a mismatch or indel imposes a penalty. The overall highest scoring path through the matrix is the alignment chosen.

The specific values chosen for scores in this algorithm indicate how to balance, for an alignment with multiple possible interpretations, the possibility of an indel as opposed to one or more SNPs, or the preference for an alignment without clipping. The default DRAGEN scoring values are reasonable for aligning moderate length reads to a whole human reference genome for variant calling applications. But any set of Smith-Waterman scoring parameters represents an imprecise model of genomic mutation and sequencing errors, and differently tuned alignment scoring values can be more appropriate for some applications.

The following alignment options control Smith-Waterman Alignment:

| Command-Line Option Name    | Configuration File Option Name |
| --------------------------- | ------------------------------ |
| `--Aligner.global`          | `global`                       |
| `--Aligner.match-score`     | `match-score`                  |
| `--Aligner.match-n-score`   | `match-n-score`                |
| `--Aligner.mismatch-pen`    | `mismatch-pen`                 |
| `--Aligner.gap-open-pen`    | `gap-open-pen`                 |
| `--Aligner.gap-ext-pen`     | `gap-ext-pen`                  |
| `--Aligner.unclip-score`    | `unclip-score`                 |
| `--Aligner.no-unclip-score` | `no-unclip-score`              |
| `--Aligner.aln-min-score`   | `aln-min-score`                |
| `--Aligner.min-score-coeff` | `min-score-coeff`              |

* **global** The `global` option (value can be 0 or 1) controls whether alignment is forced to be end-to-end in the read. When set to 1, alignments are always end-to-end, as in the Needleman-Wunsch global alignment algorithm (although not end-to-end in the reference), and alignment scores can be positive or negative. When set to 0, alignments can be clipped at either or both ends of the read, as in the Smith-Waterman local alignment algorithm, and alignment scores are nonnegative. Generally, `global=0` is preferred for longer reads, so significant read segments after a break of some kind (large indel, structural variant, chimeric read, and so forth) can be clipped without severely decreasing the alignment score. Setting global=1 might not have the desired effect with longer reads because insertions at or near the ends of a read can function as pseudoclipping. Also, with global=0, multiple (chimeric) alignments can be reported when various portions of a read match widely separated reference positions. Using `global=1` is sometimes preferable with short reads, which are unlikely to overlap structural breaks, unable to support chimeric alignments, and are suspected of incorrect mapping if they cannot align well end-to-end. Consider using the unclip-score option, or increasing it, instead ofsetting global=1, to make a soft preference for unclipped alignments.
* **match-score** The `match-score` option specifies the score for a read nucleotide matching a reference nucleotide (A, C, G, or T), or matching a reference 2–3 nucleotide IUPAC-IUB code. Its value is an unsigned integer, from 0 to 15. match\_score=0 can only be used when global=1. A higher match score results in longer alignments, and fewer long insertions.
* **match-n-score** The `match-n-score` option specifies the score for an aligned position where the read position and/or the reference position is an N code. This option is a signed integer, from -16 to 15.
* **mismatch-pen** The `mismatch-pen` option is the penalty (negative score) for a read nucleotide mismatching any reference nucleotide or IUPAC-IUB code, except N. This option is an unsigned integer, from 0 to 63. A higher mismatch penalty results in alignments with more insertions, deletions, and clipping to avoid SNPs.
* **gap-open-pen** The `gap-open-pen` option is the penalty (negative score) for opening a gap (ie, an insertion or deletion). This value is only for a 0-base gap. It is always added to the gap length times gap-ext-pen. This option is an unsigned integer, from 0 to 127. A higher gap open penalty causes fewer insertions and deletions of any length in alignment CIGARs, with clipping or alignment through SNPs used instead.
* **gap-ext-pen** The `gap-ext-pen` option is the penalty (negative score) for extending a gap (ie, an insertion or deletion) by one base. This option is an unsigned integer, from 0 to 15. A higher gap extension penalty causes fewer long insertions and deletions in alignment CIGARs, with short indels, clipping, or alignment through SNPs used instead.
* **unclip-score** The `unclip-score` option is the score bonus for an alignment reaching the beginning or end of the read. An end-to-end alignment receives twice this bonus. This option is an unsigned integer, from 0 to 127. A higher unclipped bonus causes alignment to reach the beginning and/or end of a read more often, where this can be done without too many SNPs or indels. A nonzero unclip-score is useful when global=0 to make a soft preference for unclipped alignments. Unclipped bonuses have little effect on alignments when global=1, because end-to-end alignments are forced anyway (although 2 × unclip-score does add to every alignment score unless no-unclip-score = 1). Note that, especially with longer reads, setting unclip-score much higher than gap-open-pen can have the undesirable effect of insertions at or near one end of a read being utilized as pseudoclipping, as happens with global=1
* **no-unclip-score** The `no-unclip-score` option can be 0 or 1. The default is 1. When no-unclip-score is set to 1, any unclipped bonus (unclip-score) contributing to an alignment is removed from the alignment score before further processing, such as comparison with aln-min-score, comparison with other alignment scores, and reporting in AS or XS tags. However, the unclipped bonus still affects the best-scoring alignment found by Smith-Waterman alignment to a given reference segment, biasing toward unclipped alignments When unclip-score > 0 causes a Smith-Waterman local alignment to extend out to one or both ends of the read, the alignment score stays the same or increases if no-unclip-score=0, whereas it stays the same or decreases if no-unclip-score=1. The default, no-unclip-score=1, is recommended when global=1, because every alignment is end-to-end, and there is no need to add the same bonus to every alignment. When changing no-unclip-score, consider whether aln-min-score should be adjusted. When no-unclip-score=0, unclipped bonuses are included in alignment scores compared to the aln-min-score floor, so the subset of alignments filtered out by aln-min-score can change significantly with no-unclip-score.
* **aln-min-score** The `aln-min-score` option specifies a minimum acceptable alignment score. Any alignment results scoring lower are discarded. Increasing or decreasing aln-min-score can reduce or increase the percentage of reads mapped. This option is a signed integer (negative alignment scores are possible with global=0). aln-min-score also affects MAPQ estimates. The primary contributor to MAPQ calculation is the difference between the best and second-best alignment scores. A read's best alignment score is saved in the AS SAM tag, and the second-best score (if available) is saved in the XS tag. aln-min-score serves as the suboptimal alignment score if nothing higher was found except the best score. Therefore, increasing aln-min-score can decrease reported MAPQ for some low-scoring alignments. You can use the min-score-coeff option to adjust aln-min-score as a function of read length.
* **min-score-coeff** The `min-score-coeff` option makes adjustments to `aln-min-score` per read base. When using the `min-score-coeff` and `aln-min-score` options together, you can define the minimum alignment score for each read as an affine function of read lengths. The minimum score for an N-base read is calculated as follows: `(min-score-coeff)\*N+(aln-min-score)` The `min-score-coeff` option is an integer ranging from –64 to 63.999. If the value is 0, then the minimum alignment score is fixed at aln-min-score for all read length. You can use positive values for `min-score-coeff` to allow shorter reads to match with lower alignment scores, but require longer reads to achieve higher scores.

#### Paired-End Options

DRAGEN can process paired-end data passed via a pair of FASTQ files or in a single interleaved FASTQ file. The hardware maps the two ends separately, and then determines a set of alignments that seem most likely to form a pair in the expected orientation and having roughly the expected insert size. The alignments for the two ends are evaluated for the quality of their pairing, with larger penalties for insert sizes far from the expected size. The following options control processing of paired-end data:

* **Reorientation** The `pe-orientation` option specifies the expected paired-end orientation. Only pairs with this orientation can be flagged as proper pairs. Valid values are as follows:
  * 0--FR (default)
  * 1--RF
  * 2--FF
* **unpaired-pen** For paired end reads, best mapping positions are determined jointly for each pair, according to the largest pair score found, considering the various combinations of alignments for each mate. A pair score is the sum of the two alignment scores minus a pairing penalty, which estimates the unlikelihood of insert lengths further from the mean insert than this aligned pair. The `unpaired-pen` option specifies how much alignment pair scores should be penalized when the two alignments are not in properly paired position or orientation. This option also serves as the maximum pairing penalty for properly paired alignments with extreme insert lengths. The unpaired-pen option is specified in Phred scale, according to its potential impact on MAPQ. Internally, it is scaled into alignment score space based on Smith-Waterman scoring parameters.
* **pe-max-penalty**

The `pe-max-penalty` option limits how much the estimated MAPQ for one read can increase because its mate aligned nearby. A paired alignment is never assigned MAPQ higher than the MAPQ that it would have received mapping single-ended, plus this value. By default, pe-max-penalty = mapq-max = 255, effectively disabling this limit. The key difference between `unpaired-pen` and `pe-max-penalty` is that `unpaired-pen` affects calculated pair scores and thus which alignments are selected and pe-max-penalty affects only reported MAPQ for paired alignments.

#### Mean Insert Size Detection

When working with paired-end data, DRAGEN must choose among the highest-quality alignments for the two ends to try to choose likely pairs. To make this choice, DRAGEN uses a skew normal insert model to evaluate the likelihood that a pair of alignments constitutes a pair. This model is based on the observation that common library preparation methods have insert-size distributions that are sometimes close to normal, but also sometimes clearly asymmetric, often skewing toward longer insert sizes. The skew normal insert model is used only for the DNA mode.

If you know the statistics of your library prep for an input file (and the file consists of a single read group), you can specify the characteristics of the insert-length distribution: mean, standard deviation, shape (or skewness) and three quartiles. These characteristics can be specified with the `Aligner.pe-stat-mean-insert`, `Aligner.pe-stat-stddev-insert`, `Aligner.pe-stat-shape-insert`, `Aligner.pe-stat-quartiles-insert`, and `Aligner.pe-stat-mean-read-len` options. However, it is typically preferable to allow DRAGEN to detect these characteristics automatically.

Dragen automatically samples the insert-length distribution. When the software starts execution, it runs a sample of up to 2,000,000 pairs through the aligner, calculates the distribution, and then uses the resulting statistics for evaluating all pairs in the input set.

The DRAGEN host software reports the statistics in its stdout log in a report, as follows:

```
Initial paired-end statistics detected for read group RGID, based on 39042 high quality pairs for FR orientation
        Quartiles (25 50 75) = 398 409 420
        Mean = 410.192
        Standard deviation = 14.1254
        NOTE: DRAGEN's insert estimates include corrections for clipping (so they are not identical to TLEN)

        Skew-normal insert distribution applied:
          Position (xi) = 424.084
          Scale (omega) = 19.8719
          Shape (alpha) = -1.88125

        To rerun with identical insert stats, specify:
          --Aligner.pe-stat-mean-insert=424.084
          --Aligner.pe-stat-stddev-insert=19.8719
          --Aligner.pe-stat-shape-insert=-1.88125
          --Aligner.pe-stat-quartiles-insert="398 409 420"
          --Aligner.pe-stat-mean-read-len=101
```

Note that the `Mean`, `Standard deviation` and `Quartiles` reported above are the sample mean, standard deviation and quartiles calculated from the initial sample of up to 2,000,000 pairs, assuming a normal distribution. The sample mean and standard deviation are used to fit the parameters of a skew-normal distribution. A skew-normal distribution is defined by starting with an underlying normal distribution (whose mean we call `position` or `xi` and standard deviation we call `scale` or `omega`) and folding a varying portion of the probability mass from one side of the mean (e.g., left side) to the other (e.g., right) side. The portion folded varies smoothly, from 0% at the original mean, approaching 100% from the left tail to the right tail. A `shape` parameter which we call `alpha` controls how rapidly the folded fraction increases, and at `alpha=0` there is no folding and the distribution remains normal.

In the standard output, we also include the command line options needed to reproduce the DRAGEN run with the same insert stat settings. Note that when specifying stats on the command line, the skew-normal `xi` value should be used for `Aligner.pe-stat-mean-insert`. The `omega` value should be used for `Aligner.pe-stat-stddev-insert`, and the `alpha` value should be used for `Aligner.pe-stat-shape-insert`. If `Aligner-pe-stat-shape-insert` is not specified on the command line, a default value of 0 is assumed.

The insert length distribution for each sample is written to fragment\_length\_hist.csv. Each sample starts with the following lines

```
 #Sample: sample name
 FragmentLength,Count
```

These lines are followed by the histogram for the first \~2M read pairs for DNA (\~100K read pairs for RNA). The histogram counts are aggregated across all read groups sharing the same sample id (`RGSM` field).

When the number of sample pairs is very small, there is not enough information to characterize the distribution with high confidence. In this case, DRAGEN applies default statistics that specify a very wide insert distribution, which tends to admit pairs of alignments as proper pairs, even if they may lie tens of thousands of bases apart. In this situation, DRAGEN outputs a message, as follows:

```
WARNING: Less than 28 high quality pairs found - standard deviation is
calculated from the small samples formula
```

The small samples formula calculates standard deviation as follows:

```
 if samples < 3 then                                                     
      standard deviation = 10000                                          
 else if samples < 28 then                                               
    standard deviation = 25 * (standard deviation + 1) / (samples - 2) 
 end if                                                                   
                                                                          
 if standard deviation < 12 then                                         
      standard deviation = 12                                             
 end if                                                                   
```

The default model is "standard deviation = 10000". If the first 2M reads are unmapped or if all pairs are improper pairs, then the standard deviation is set to 10000 and the mean and quartiles are set to 0. Note that the minimum value for standard deviation is 12, which is independent of the number of samples. Also, in the DNA mode when we have fewer than 1000 high quality alignments we revert to the normal distribution based insert model, because of insufficient number of samples to accurately estimate the parameters of the skew normal distribution.

For RNA-Seq data, the insert size distribution is not normal due to pairs containing introns. The DRAGEN software estimates the distribution using a kernel density estimator to fit a long tail to the samples. This estimate leads to a more accurate mean and standard deviation for RNA-Seq data and proper pairing.

DRAGEN writes detected paired-end stats into a tab-delimited log file in the output directory called .insert-stats.tab. This file contains the statistical distribution of detected insert sizes for each read group, including quartiles, mean, standard deviation, shape, minimum, and maximum. The information matches the standard-out report above. Additionally, the log file includes the minimum and maximum insert limits that DRAGEN applied for rescue scans. Note that the reported mean and standard deviation in this tab-limited log file are the `xi` and `omega` parameters of the skew-normal distribution.

#### Rescue Scans

For paired-end reads, where a seed hit is found for one mate but not the other, rescue scans hunt for missing mate alignments within a rescue radius of the mean insert length. Normally, the DRAGEN host software sets the rescue radius to 2.5 standard deviations of the empirical insert distribution. But in cases where the insert standard deviation is large compared to the read length, the rescue radius is restricted to limit mapping slowdowns. In this case, a warning message is displayed, as follows:

```
Rescue radius = 220
     Effective rescue sigmas = 0.5
            WARNING: Default rescue sigmas value of 2.5 was overridden by host software!
            The user may wish to set rescue sigmas value explicitly with --Aligner.rescue-sigmas
```

Although the user can ignore this warning, or specify an intermediate rescue radius to maintain mapping speed, it is recommended to use 2.5 sigmas for the rescue radius to maintain mapping sensitivity. To disable rescue scanning, set max-rescues to 0.

#### Output Options

DRAGEN can track multiple independent alignments for each read. These alignments include the optimal (primary) one, as well as those mapping different subsegments of the read, (chimeric/supplementary), and sub-optimal (secondary) mappings of the read to different areas of the reference.

For DNA alignment by default, DRAGEN can emit one primary alignment for each read, up to three chimeric alignments (Aligner.supp-aligns=3), and no secondary alignments (Aligner.sec-aligns=0). The maximum user-specified value for supp-aligns or sec-aligns is 4095.

You can use the following configuration options to control how many of each type of alignment to include in DRAGEN output.

* **mapq-max** The `mapq-max` option specifies a ceiling on the estimated MAPQ that can be reported for any alignment, from 0 to 255. If the calculated MAPQ is higher, this value is reported instead. The default is 60.
* **supp-aligns**, **sec-aligns** The `supp-aligns` and `sec-aligns` options restrict the maximum number of supplementary (ie, chimeric and SAM FLAG 0x800) alignments and secondary (ie, suboptimal and SAM FLAG 0x100) alignments, respectively, that can be reported for each read. A maximum of 4095 supplementary alignments and 4095 secondary alignments can be reported for any read, in addition to a primary alignment. High settings for these two options impact speed so it is advisable to increase only as needed.
* **sec-phred-delta** The `sec-phred-delta` option controls which secondary alignments are emitted based on the alignment score relative to the primary reported alignment. Only secondary alignments with likelihood within this Phred value of the primary are reported.
* **sec-aligns-hard** The `sec-aligns-hard` option suppresses the output of all secondary alignments if there are more secondary alignments than can be emitted. Set sec-aligns-hard to 1 to force the read to be unmapped when not all secondary alignments can be output.
* **supp-as-sec** When the `supp-as-sec` option is set to 1, then supplementary (chimeric) alignments are reported with SAM FLAG 0x100 instead of 0x800. The default is 0. The supp-as-sec option provides compatibility with tools that do not support FLAG 0x800.
* **hard-clips** The hard-clips option is used as a field of 3 bits, with values ranging from 0 to 7. The bits specify alignments, as follows:
  * Bit 0--primary alignments
  * Bit 1--supplementary alignments
  * Bit 2--secondary alignments

Each bit determines whether local alignments of that type are reported with hard clipping (1) or soft clipping (0). The default is 6, meaning primary alignments use soft clipping and supplementary and secondary alignments use hard clipping.

### Mapping with ALT-contigs

The GRCh38 human reference contains many more alternate haplotypes (ALT contigs) than previous versions of the reference. Generally, including ALT contigs in the mapping reference improves mapping and variant calling specificity, because misalignments are eliminated for reads matching an ALT contig but scoring poorly against the primary assembly. However, mapping with GRCh38's ALT contigs without special treatment can substantially degrade variant calling sensitivity in corresponding regions, because many reads align equally well to an ALT contig and to the corresponding position in the primary assembly.

#### Masked Based ALT-awareness

The recomeneded and default approach for dealing with ALT-contigs in DRAGEN is masking regions of ALT contigs of high similarity to their corresponding primary contig. This approach is more accurate than liftover based ALT-awarness because there are many places where the "correct" or most useful liftover between a long ALT haplotype and the primary assembly is ambiguous. Incorrect liftover can produce dense clusters of mismapped reads and false variant calls. The base masking approach has the benefits of using ALT contigs without the negative consequences.

Masked hash tables are built from a standard hg18 or hg38 FASTA that contains ALT contigs. The hash table builder will automatically mask regions of the ALT contigs with Ns.

#### Liftover Based ALT-awarness

With liftover based ALT-awareness, the mapper and aligner are aware of the liftover relationship between ALT contig positions and corresponding primary assembly positions. Seed matches within ALT contigs are used to obtain corresponding primary assembly alignments, even if the latter score poorly. Liftover groups are formed, each containing a primary assembly alignment candidate, and zero or more ALT alignment candidates that lift to the same location. Each liftover group is scored according to its best-matching alignments, taking properly paired alignments into account. The winning liftover group provides its primary assembly representative as the primary output alignment, with MAPQ calculated based on the score difference to the second-best liftover group. Emitting primary alignments within the primary assembly maintains normal aligned coverage and facilitates variant calling there. If the --Aligner.en-alt-hap-aln option is set to 1 and --Aligner.supp-aligns is greater than 0, then corresponding alternate haplotype alignments can also be output, flagged as supplementary alignments.

DRAGEN requires ALT-Aware hash tables for any hg19 or GRCh38 reference where ALT contigs are detected. To disable this requirement in DRAGEN, set the --ht-alt-aware-validate option to false.

The following is a comparison of alternative options for dealing with alternate haplotypes.

* Mapping without ALT contigs in the reference:
  * False-positive variant calls result when reads matching an alternate haplotype misalign somewhere else.
  * Poor mapping and variant calling sensitivity where reads matching an ALT contig differ greatly from the primary assembly.
* Mapping with ALT contigs but no ALT awareness:
  * False-positive variant calls from misaligned reads matching ALT contigs are eliminated.
  * Low or zero aligned coverage in primary assembly regions covered by alternate haplotypes, due to some reads mapping to ALT contigs.
  * Low or zero MAPQ in regions covered by alternate haplotypes, where they are similar or identical to the primary assembly.
  * Variant calling sensitivity is dramatically reduced throughout regions covered by alternate haplotypes.
* Mapping with ALT contigs and ALT awareness:
  * False-positive variant calls from misaligned reads matching ALT contigs are eliminated.
  * Normal aligned coverage in regions covered by alternate haplotypes because primary alignments are to the primary assembly.
  * Normal MAPQs are assigned because alignment candidates in alternative haplotypes are not considered in competition.
  * Good mapping and variant calling sensitivity where reads matching an ALT contig differ greatly from the primary assembly.

### DRAGEN Multigenome Mapper

The Multigenome Mapper in DRAGEN significantly improves the accuracy of mapping Illumina reads, particularly in challenging regions such as segmental duplications and other difficult to map regions. This advanced method leverages population haplotypes from pangenome references to incorporate additional variant information, constructing alternative haplotype paths that improve reads mapping. By offering these alternate paths, the Multigenome Mapper enables reads containing population-specific variants to align directly to their most likely genomic locations, reducing mapping ambiguity. This improved mapping also results in improved variant calling accuracy.

When given a set of population variants (VCF) or haplotypes, the pangenome reference modification is categorized in the following types:

* Alternate contigs represent population haplotypes. Alt-contigs can have a single variant or a combination of nearby phased variants.
* Ambiguous codes (IUPAC codes) to represent SNPs. To improve alignment, it edits the reference FASTA with isolated population SNPs.
* Haplotype database. An additional haplotype database is built and used to augment the reference FASTA with population variants. A multigenome based mapper algorithm is used to score read alignment according to the variants in this database.

The DRAGEN pangenome hashtables are available to download from the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).


# Read Trimming

DRAGEN can remove artifacts from reads using hardware accelerated read trimming. Hardware accelerated read trimming is available on U200 and cloud systems, as part of the DRAGEN mapper and adds no additional run time. DRAGEN provides multiple independent trimming filters that target different types of artifacts or use cases. You can enable and configure the artifacts or use cases independently to tailor the read-trimming to your analysis. Read trimming uses two different modes, hard-trimming and soft-trimming.

To enable hard-trimming mode, use `--read-trimmers`. In hard-trimming mode, potential artifacts are removed from input reads. Reads that are trimmed to fewer than 20 bases are filtered and replaced with a placeholder read that uses 10 N bases. DRAGEN assigns the filtered reads a 0x200 flag set.

DRAGEN contains a novel lossless soft-trimming mode. In soft-trimming mode, reads are mapped as though they had been trimmed, but no bases are removed. To enable the trimmer in soft mode, use `--soft-read-trimmers`.

Soft-trimming suppresses systematic mismapping of reads that contain trimmable artifacts, without actually losing the trimmed bases in aligned output. Soft-trimming prevents reads with trimmable artifacts, such as Poly-G artifacts, from being mapped to reference G homopolymers, or prevents adapter sequences from being mapped to the matching reference loci. Soft-trimming might map reads to different positions in the reference than they would have been if not using soft-trimming. When using soft-trimmed, DRAGEN does not filter reads and does not map reads with bases that would have been trimmed entirely.

Soft-trimming for Poly-G artifacts is enabled by default on supported systems.

## Read Trimming Tools

### Fixed-Length Trimming

Fixed-length trimming removes a fixed number of bases from the 5' end of each read. If you are analyzing sequencing data from an amplicon of fixed size and expect the read-length to consistently exceed the length of quality sequence data, you can use the expected number in fixed-length trimming.

### Poly-G Trimming

Poly-G artifacts appear on two-channel sequencing systems when the dark base G is called after synthesis has terminated. As a result, DRAGEN calls several erroneous high-confidence G bases on the ends of affected reads. For contaminated samples, many affected reads can be mapped to reference regions with high G content. The affected reads can cause problems for processing downstream.

### Quality Trimming

Base quality can degrade over the length of a read toward the 5' end and separate from any artifacts from early termination of synthesis. The lower quality bases can affect mapping and alignment results, and might lead to incorrect variant or methylation calls downstream. The quality trimming tool calculates a rolling average of the base quality inward from the 5' end and removes the minimum number of bases, so the average number of bases is above the threshold specified using `--trim-min-quality`.

### Adapter Trimming

Problems during library preparation, or libraries with smaller inserts can result in the synthesis of high quality reads containing sequence from the adapters used. If not removed before analysis, noninsert bases can reduce mapping efficiency and downstream accuracy. The adapter trimming tool uses the adapter sequences from the input FASTA file, and then removes all hits greater than a specified size. Adapter trimming allows for a 10% mismatch. For 3' adapters, trimming is from the first matching adapter base to the end of the read. For 5' adapters, trimming is from the first (3') matching adapter base to the beginning (5') of the read.

### Ambiguous Base Trimming

If quality trimming is not feasible due to reduced yield or other limitations, an alternative option is to remove only explicitly ambiguous bases from the ends of read. If enabled the ambiguous base trimmer applies a simple exact-match search to both ends of all processed reads, regardless of mate-pair status.

### Minimum Length Trimming

You can maximize trimmer sensitivity, by using the minimum length trimming tool to remove a fixed number of bases from each read after the trimmer tools above have run. For example, if you would like to remove 5 bp from each read, a 7 bp adapter hit could be missed if five of the bases are removed first. To mitigate this issue, DRAGEN provides an optional minimum trim-length filter.

### Maximum Length Trimming

If using libraries of fixed-size inserts, such as small PCR amplicons, it is more convenient to specify a length that all reads should be trimmed to rather than the number of bases to remove. You can use the maximum length trimming tool.

### PolyA Tail Trimming

If using RNA libraries, reads overlapping the poly-A tail of the transcripts may contain long poly-A/poly-T sequences at the end of the reads which may result in incorrect alignment. The poly-A trimmer mitigates this by trimming the poly-A tail from the end of the read. See additional description in [RNA alignment](/dragen-v4.5/product-guides/dragen-v4.5/dragen-rna-pipeline/rna-alignment#polya-trimming) section.

## Read Trimming Metrics

The trimmer generates a metrics file titled `\<output prefix\>.trimmer_metrics.csv`. Metrics are available on an aggregate level over all input data. The metrics units are in reads or bases.

* **Total input reads** Total number of reads in the input files.
* **Total input bases** Total number of bases in the input reads.
* **Total input bases R1** Total number of bases in R1 reads.
* **Total input bases R2** Total number of bases in R2 reads.
* **Average input read length** Total number of input bases divided by the number of input reads.
* **Total trimmed reads** Total number of reads trimmed by at least one base, not including soft-trimming.
* **Total trimmed bases** Total number of bases trimmed, not including soft-trimming.
* **Average bases trimmed per read** The number of trimmed bases divided by the number of input reads.
* **Average bases trimmed per trimmed read** The number of trimmed bases divided by the number of trimmed reads.
* **Remaining poly-G K-mers R1 3prime** The number of R1 3' read ends that contain likely Poly-G artifacts after trimming.
* **Remaining poly-G K-mers R2 3prime** The number of R2 3' read ends that contain likely Poly-G artifacts after trimming.
* **Total filtered reads** The number of reads that were filtered out during trimming.
* **Reads filtered for minimum read length R1** The number of R1 reads that were filtered due to being trimmed below the minimum read length.
* **Reads filtered for minimum read length R2** The number of R2 reads that were filtered due to being trimmed below the minimum read length.
* **\<Trimmer tool> trimmed reads** The number of reads with at least one base trimmed by TRIMMER. DRAGEN reports the metric for both R1 and R2 mates and the filtering status (unfiltered or filtered) of the trimmed read. The metric includes reads that were trimmed during soft-trimming. Each trimming tool above produces the metric.
* **\<Trimmer tool> trimmed bases** The number of bases trimmed by TRIMMER. The metric is produced for both R1 and R2 mates and the filtering status (unfiltered or filtered) of the trimmed read. The metric includes bases from reads that were trimmed during soft trimming. Each trimming tool above produces the metric.

## Read Trimming Settings

### Read trimmer

| Option                 | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `--read-trimmers`      | <p>To enable trimming filters in hard-trimming mode, set to a comma-separated list of the trimmer tools you would like to use (in the order of execution). To disable trimming, set to <code>none</code>. During mapping, artifacts are removed from all reads. The following are valid trimmer names:</p><ul><li><code>fixed-len</code>—Fixed-length trimming</li><li><code>polyg</code>—Poly-G trimming</li><li><code>quality</code>—Quality trimming</li><li><code>adapter</code>—Adapter trimming</li><li><code>n</code>—Ambiguous base trimming</li><li><code>min-len</code>—Minimum length trimming</li><li><code>cut-end</code>—Maximum length trimming</li><li><code>polya</code>—RNA Poly-A tail trimming. See additional description in <a href="/pages/a02FlLtIzfQthUzBdkCL#polya-trimming">RNA alignment</a> section</li><li><code>bisulfite</code>—Bisulfite trimming</li></ul><p>Read trimming is disabled by default (default: "none").</p>                                                                                 |
| `--soft-read-trimmers` | <p>To enable trimming filters in soft-trimming mode, set to a comma-separated list of the trimmer tools you would like to use (in the order of execution). To disable soft trimming, set to <code>none</code>. During mapping, reads are aligned as if trimmed, and bases are not removed from the reads. The following are the valid trimmer names.</p><ul><li><code>fixed-len</code>—Fixed-length trimming</li><li><code>polyg</code>—Poly-G trimming</li><li><code>quality</code>—Quality trimming</li><li><code>adapter</code>—Adapter trimming</li><li><code>n</code>—Ambiguous base trimming</li><li><code>min-len</code>—Minimum length trimming</li><li><code>cut-end</code>—Maximum length trimming</li><li><code>polya</code>—RNA Poly-A tail trimming. See additional description in <a href="/pages/a02FlLtIzfQthUzBdkCL#polya-trimming">RNA alignment</a> section</li><li><code>bisulfite</code>—Bisulfite trimming</li></ul><p>Soft-trimming is enabled for the <code>polyg</code> filter by default (default: "polyg").</p> |
| `--trimming-only`      | Disables mapping and alignment to run read-trimming only.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |

### Filtering after the trimmer's execution

| Option                    | Description                                                                                                                                                                                                 |
| ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--trim-min-length`       | Specify a minimum read length allowed after the trimmer execution. DRAGEN filters any reads with a length less than the value after all read-trimming steps are completed (default: 20).                    |
| `--trim-min-len-read1`    | Specify a minimum read length allowed for read1 after the trimmer execution. DRAGEN filters any reads with a length of read1 less than the value after all read-trimming steps are completed (default: 20). |
| `--trim-min-len-read2`    | Specify a minimum read length allowed for read2 after the trimmer execution. DRAGEN filters any reads with a length of read2 less than the value after all read-trimming steps are completed (default: 20). |
| `--trim-filter-dummy-len` | Specify the number of N bases in dummy reads that replace filtered reads (default: 10).                                                                                                                     |
| `--trim-filter-set-flag`  | If enabled, dummy reads will have their 0x200 SAM flag set (default: true).                                                                                                                                 |

### Fixed-length trimming

| Option             | Description                                                                     |
| ------------------ | ------------------------------------------------------------------------------- |
| `--trim-r1-5prime` | Specify a fixed number of bases to trim from the 5' end of Read 1 (default: 0). |
| `--trim-r1-3prime` | Specify a fixed number of bases to trim from the 3' end of Read 1 (default: 0). |
| `--trim-r2-5prime` | Specify a fixed number of bases to trim from the 5' end of Read 2 (default: 0). |
| `--trim-r2-3prime` | Specify a fixed number of bases to trim from the 3' end of Read 2 (default: 0). |

### Quality trimming

| Option                     | Description                                                                                                   |
| -------------------------- | ------------------------------------------------------------------------------------------------------------- |
| `--trim-min-quality`       | Specify the minimum read quality. DRAGEN trims bases from the 3' end of reads with a quality below the value. |
| `--trim-quality-r1-5prime` | Specify the quality cutoff below which to trim from the 5' end of read 1.                                     |
| `--trim-quality-r1-3prime` | Specify the quality cutoff below which to trim from the 3' end of read 1.                                     |
| `--trim-quality-r2-5prime` | Specify the quality cutoff below which to trim from the 5' end of read 2.                                     |
| `--trim-quality-r2-3prime` | Specify the quality cutoff below which to trim from the 3' end of read 2.                                     |

### Adapter trimming

| Option                      | Description                                                                                                                                                                                                  |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `--trim-adapter-read1`      | Specify the FASTA file that contains adapter sequences to trim from the 3' end of Read 1.                                                                                                                    |
| `--trim-adapter-read2`      | Specify the FASTA file that contains adapter sequences to trim from the 3' end of Read 2.                                                                                                                    |
| `--trim-adapter-r1-5prime`  | Specify the FASTA file that contains adapter sequences to trim from the 5' end of Read 1. NB: the sequences should be in reverse order (with respect to their appearance in the FASTQ) but not complemented. |
| `--trim-adapter-r2-5prime`  | Specify the FASTA file that contains adapter sequences to trim from the 5' end of Read 2. NB: the sequences should be in reverse order (with respect to their appearance in the FASTQ) but not complemented. |
| `--trim-adapter-stringency` | Specify the minimum number of adapter bases required for trimming (default: 4).                                                                                                                              |

### Bisulfite trimming

| Option                    | Description                                                                                                     |
| ------------------------- | --------------------------------------------------------------------------------------------------------------- |
| `--trim-bisulfite-ends`   | Enable both 5-Prime and 3-Prime bisulfite trimming.                                                             |
| `--trim-bisulfite-5prime` | If a 3' adapter was trimmed, trim an additional 2bp from the 3' end, unless the 5' end matches 'CAA' or 'CGA'". |
| `--trim-bisulfite-3prime` | If the 5' end matches 'CAA' or 'CGA', trim the first two of these 5' bases.                                     |

### Minimum-length trimming

| Option                 | Description                                                                         |
| ---------------------- | ----------------------------------------------------------------------------------- |
| `--trim-min-r1-5prime` | Specify the minimum number of bases to trim from the 5' end of Read 1 (default: 0). |
| `--trim-min-r1-3prime` | Specify the minimum number of bases to trim from the 3' end of Read 1 (default: 0). |
| `--trim-min-r2-5prime` | Specify the minimum number of bases to trim from the 5' end of Read 2 (default: 0). |
| `--trim-min-r2-3prime` | Specify the minimum number of bases to trim from the 3' end of Read 2 (default: 0). |

### Maximum-length trimming

| Option                 | Description                                                                               |
| ---------------------- | ----------------------------------------------------------------------------------------- |
| `--trim-max-length`    | Specify the maximum number of bases that can be trimmed from the sequences of both reads. |
| `--trim-max-len-read1` | Specify the maximum number of bases that can be trimmed from the sequences of read1.      |
| `--trim-max-len-read2` | Specify the maximum number of bases that can be trimmed from the sequences of read2.      |

### PolyA trimming

| Option                  | Description                                                              |
| ----------------------- | ------------------------------------------------------------------------ |
| `--trim-polya-min-trim` | The minimum number of poly-As required for polya trimming (default: 20). |

### PolyG trimming

| Option                              | Description                                                                                |
| ----------------------------------- | ------------------------------------------------------------------------------------------ |
| `--trim-polyg-kmer-len`             | How many bases to check at each read end for poly-G artifact detection (default: 25).      |
| `--trim-polyg-kmer-non-g`           | The maximum number of non-G bases in the K-mer for poly-G artifact detection (default: 2). |
| `--trim-polyg-g-score-r1-5prime`    | The score for G bases on the 5' end of read 1 (default: 0).                                |
| `--trim-polyg-g-score-r1-3prime`    | The score for G bases on the 3' end of read 1 (default: 15).                               |
| `--trim-polyg-g-score-r2-5prime`    | The score for G bases on the 5' end of read 2 (default: 0).                                |
| `--trim-polyg-g-score-r2-3prime`    | The score for G bases on the 3' end of read 2 (default: 15).                               |
| `--trim-polyg-min-trim-r1-5prime`   | The minimum number of G's to trim from the 5' end of read 1 (default: 6).                  |
| `--trim-polyg-min-trim-r1-3prime`   | The minimum number of G's to trim from the 3' end of read 1 (default: 6).                  |
| `--trim-polyg-min-trim-r2-5prime`   | The minimum number of G's to trim from the 5' end of read 2 (default: 6).                  |
| `--trim-polyg-min-trim-r2-3prime`   | The minimum number of G's to trim from the 3' end of read 2 (default: 6).                  |
| `--trim-polyg-early-exit-threshold` | The signed score threshold for poly-G trimming to exit early (default: -500).              |

### PolyX trimming

| Option                            | Description                                                                                 |
| --------------------------------- | ------------------------------------------------------------------------------------------- |
| `--trim-polyx-bases-r1-5prime`    | The bases to trim for polyX trimming from the 5' end of read 1 (default: empty string "" ). |
| `--trim-polyx-bases-r1-3prime`    | The bases to trim for polyX trimming from the 3' end of read 1 (default: empty string "" ). |
| `--trim-polyx-bases-r2-5prime`    | The bases to trim for polyX trimming from the 5' end of read 2 (default: empty string "" ). |
| `--trim-polyx-bases-r2-3prime`    | The bases to trim for polyX trimming from the 3' end of read 2 (default: empty string "" ). |
| `--trim-polyx-min-trim-r1-5prime` | The minimum number of X's to trim from the 5' end of read 1 (default: 20).                  |
| `--trim-polyx-min-trim-r1-3prime` | The minimum number of X's to trim from the 3' end of read 1 (default: 20).                  |
| `--trim-polyx-min-trim-r2-5prime` | The minimum number of X's to trim from the 5' end of read 2 (default: 20).                  |
| `--trim-polyx-min-trim-r2-3prime` | The minimum number of X's to trim from the 3' end of read 2 (default: 20).                  |


# Sorting and Duplicate Marking

### Sorting

The map/align system produces a BAM file sorted by reference sequence and position by default. Creating this BAM file typically eliminates the requirement to run samtools sort or any equivalent postprocessing command. The `‑‑enable-sort option` can be used to enable or disable creation of the BAM file, as follows:

* To enable, set to true.
* To disable, set to false.

On the reference hardware system, running with sort enabled increases run time for a 30x full genome by about 6--7 minutes.

### Duplicate Marking

Marking or removing duplicate aligned reads is a common best practice in whole-genome sequencing. Not doing so can bias variant calling and lead to incorrect results.

The DRAGEN system can mark or remove duplicate reads, and produces a BAM file with duplicates marked in the FLAG field, or with duplicates entirely removed.

In testing, enabling duplicate marking adds minimal run time over and above the time required to produce the sorted BAM file. The additional time is approximately 1--2 minutes for a 30x whole human genome, which is a huge improvement over the long run times of open source tools.

#### The Duplicate Marking Algorithm

The DRAGEN duplicate-marking algorithm is modeled on the Picard toolkit's MarkDuplicates feature. All the aligned reads are grouped into subsets in which all the members of each subset are potential duplicates.

For two pairs to be duplicates, they must have the following:

* Identical alignment coordinates (position adjusted for soft- or hard-clips from the CIGAR) at both ends.
* Identical orientations (direction of the two ends, with the left-most coordinate being first).

In addition, an unpaired read may be marked as a duplicate if it has identical coordinate and orientation with either end of any other read, whether paired or not.

Unmapped read pairs are never marked as duplicates.

When DRAGEN has identified a group of duplicates, it picks one as the best of the group, and marks the others with the BAM duplicate flag (0x400, or decimal 1024). For this comparison, duplicates are scored based on the average sequence Phred quality. Pairs receive the sum of the scores of both ends, while unpaired reads get the score of the one mapped end. The idea of this score is to try, all other things being equal, to preserve the reads with the highest-quality base calls.

If two reads (or pairs) have exactly matching quality scores, DRAGEN breaks the tie by choosing the pair with the higher alignment score. If there are multiple pairs that also tie on this attribute, then DRAGEN chooses a winner arbitrarily.

The score for an unpaired read R is the average Phred quality score per base, calculated as follows:

![](/files/9PA2K8GfkC7RurOR8JYR)

Where R is a BAM record, QUAL is its array of Phred quality scores, and dedup-min-qual is a DRAGEN configuration option with default value of 15. For a pair, the score is the sum of the scores for the two ends.

This score is stored as a one-byte number, with values rounded down to the nearest one-quarter. This rounding may lead to different duplicate marks from those chosen by Picard, but because the reads were very close in quality this has negligible impact on variant calling results.

#### Duplicate Marking Limitations

The limitations to DRAGEN duplicate marking implementation are as follows:

* When there are two duplicate reads or pairs with very close Phred sequence quality scores, DRAGEN might choose a different winner from that chosen by Picard. These differences have negligible impact on variant calling results.
* If using a single FASTQ file as input, DRAGEN accepts only a single library ID as a command-line argument (RGLB). For this reason, the FASTQ inputs to the system must be already separated by library ID. Library ID cannot be used as a criterion for distinguishing non-duplicates.
* DRAGEN does not distinguish between optical and PCR duplicates.

#### Duplicate Marking Settings

The following options can be used to configure duplicate marking in DRAGEN:

* `--enable-duplicate-marking` Set to true to enable duplicate marking. When `\--enable-duplicate-marking is enabled`, the output is sorted, regardless of the value of the enable-sort option.
* `--remove-duplicates` Set to true to suppress the output of duplicate records. If set to false, set the 0x400 flag in the FLAG field of duplicate BAM records. When --remove-duplicates is enabled, then enable- duplicate-marking is forced to enabled as well.
* `--dedup-min-qual`\
  Specifies the Phred quality score below which a base should be excluded from the quality score calculation used for choosing among duplicate reads.


# Small Variant Calling

The DRAGEN Small Variant Caller is a high-speed haplotype caller implemented with a hybrid of hardware and software. The caller performs localized *de novo* assembly in regions of interest to generate candidate haplotypes, and then performs read likelihood calculations using a hidden Markov model (HMM).

Variant calling is disabled by default. To enable variant calling, set the `--enable-variant-caller` option to true. The VCF header is annotated with `##source=DRAGEN_SNV` to indicate the file is generated by the DRAGEN SNV pipeline.

## The Variant Caller Algorithm

The DRAGEN Small Variant Caller performs the following steps:

**Active Region Identification**---Identifies areas where multiple reads disagree with the reference are identified, and selects windows around them (active regions) for processing.

**Localized Haplotype Assembly**--- Assembles all overlapping reads in each active region into a de Bruijn graph (DBG). A DBG is a directed graph based on overlapping K-mers (length K subsequences) in each read or multiple reads. When all reads are identical, the DBG is linear. Where there are differences, the graph forms bubbles of multiple paths that diverge and rejoin. If the local sequence is too repetitive and K is too small, cycles can form, which invalidate the graph. DRAGEN uses K=10 and 25 as the default values. If those values produce an invalid graph, then additional values of K=35, 45, 55, 65 are tried until a cycle-free graph is obtained. From this cycle-free DBG, DRAGEN extracts every possible path to produce a complete list of candidate haplotypes, ie, hypotheses for what the true DNA sequence might be on at least one strand. In addition to graph assembly, haplotypes are also generated via columnwise detection, with candidate variant events identified directly from BAM alignments. Columnwise detection is enabled by default in all small variant calling pipelines and is supplementary to the DBG, but is especially useful in highly repetitive regions where DBG assembly of reads is more likely to fail.

**Haplotype Alignment**---Uses the Smith-Waterman algorithm to align each extracted haplotype to the reference genome. The alignments determine what variations from the reference are present.

**Read Likelihood Calculation**---Tests each read against each haplotype to estimate a probability of observing the read assuming the haplotype was the true original DNA sampled. This calculation is performed by evaluating a pair hidden Markov model (HMM), which accounts for the various possible ways the haplotype might have been modified by PCR or sequencing errors into the read observed. The HMM evaluation uses a dynamic programming method to calculate the total probability of any series of Markov state transitions arriving at the observed read.

**Genotyping**---Forms the possible diploid combinations of variant events from the candidate haplotypes and, for each combination, calculates the conditional probability of observing the entire read pileup. Calculations use the constituent probabilities of observing each read, given each haplotype from the pair HMM evaluation. These calculations feed into the Bayesian formula to calculate the likelihood that each genotype is the genotype of the sample being analyzed, given the entire read pileup observed. Genotypes with maximum likelihood are reported.

## Read filtering and reporting of vcf DP fields

In most pipelines, DRAGEN reports two types of depth counts, both of which may differ from the information in the BAM pileup due to various filtering steps that are applied throughout variant calling. Briefly:

* **Unfiltered depth** is the number of reads (fragment-based) covering the position, downstream of any read collapsing, deduplication, downsampling and read disqualification, but upstream of informative read determination. Overlapping mate pairs are counted as single reads. When overlapping mate pairs are present, this may cause an apparent discrepancy between the reported depth and the pileup as viewed in a browser such as IGV. To resolve this, use the "View as pairs" option in IGV. Unfiltered depth is reported as INFO/DP, except in the case of gVCF homref calls, where it is reported as FORMAT/DP.
* **Informative depth** is the number of reads (fragment-based) actually used to make the calling decision, where badly mated reads and uninformative reads (reads that could not be assigned to a specific allele) have been excluded. Informative depth is reported as FORMAT/DP, except in the case of gvcf homref calls, where it is not reported. The FORMAT/AD and FORMAT/AF fields are based on informative depth.

The following figure summarizes the different filtering steps in more detail.

![](/files/5Hg4rjYU1S7Hvv0vsXR9)

* Filter 1 acts on the reads present in the BAM input (in UMI pipelines, these are the collapsed reads produced by the read collapsing step, not the raw reads) and filters out the following reads:
  * Duplicate reads.
  * Soft-clipped bases. DRAGEN filters out soft-clipped bases only when calculating coverage reports.
  * **\[Somatic]** Reads with MAPQ=0.
  * **\[Somatic]** Reads with MAPQ < vc-min-tumor-read-qual, where vc-min-tumor-read-qual >1. By default, germline runs with machine learning (ML) enabled consider all reads, even those with MAPQ 0, resulting in increased sensitivity. MAPQ read filtering is controlled by `--vc-min-read-qual` for germline and tumor/normal (T/N) runs, but it does not affect tumor-only (T/O) runs. In contrast, `--vc-min-tumor-read-qual` controls filtering for tumor samples in T/N and T/O runs and has no effect on normal-only samples.
* Filter 2 trims bases with BQ < 10 and filters out the following reads:
  * Unmapped reads.
  * Secondary reads.
  * Reads with bad cigars.
* Filter 3 occurs after downsampling and HMM. Filter 3 filters out the following reads:
  * Disqualified reads. Reads are disqualified if their HMM score is below a threshold.
* Filter 4 occurs after the genotyper runs. The genotyper adds annotation information to the FORMAT field. Filter 4 filters out the following reads:
  * Badly mated reads. A badly mated read is a read where the pair is mapped to two different reference contigs.
  * Non-informative reads. For example, if the HMM scores of the read against two different haplotypes are almost equal, the read is filtered out because it does not provide enough information to distinguish which of the two haplotypes are more likely.

## Mosaic Calling

Since DRAGEN 4.3 the mosaic small variant caller runs downstream of the germline small variant caller. Non-cancer post-zygotic mosaic variants with typical AF lower than 50% detected by the mosaic caller are reported in the output VCF file with a `MOSAIC` INFO flag. As default, `MOSAIC` tagged variants with `AF` smaller than 20% are filtered with the `MosaicLowAF` filter. To further enhance sensitivity in WGS samples, if the median depth of the sample detected by the ploidy estimator (`cvg`) exceeds `100x`, a default of `20/cvg` threshold will be applied if the coverage is between `100x` and `200x`, and 10% if the coverage exceeds `200x`. On WES data the default value of the threshold is 10% regardless of the estimated coverage. This is likely to have a greater impact on exome data, which typically has higher coverage. Exome users looking to control the number of low allele frequency (AF) mosaic variants can set the option `--vc-mosaic-af-filter-threshold` to 0.2 to override the dynamic coverage-based thresholding.

See [Mosaic detection](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/mosaic-detection) for further details on the mosaic small variant caller and the mosaic detection mode and a comparison with previous versions of DRAGEN features, namely DRAGEN 4.2, DRAGEN 4.3, and DRAGEN 4.4.

## Variant Caller Options

The following options control the variant caller stage of the DRAGEN host software.

* `--enable-variant-caller`

  Set `--enable-variant-caller` to true to enable the variant caller stage for the DRAGEN pipeline.
* `--vc-target-bed`

  **\[Optional]** Restricts processing of the small variant caller, target BED related coverage, and callability metrics to regions specified in a BED file. The BED file is a text file containing at least three tab-delimited columns. The first three columns are chromosome, start position, and end position. The positions are zero-based. For example:

```
# header information
chr11 0 246920
chr11 255660 255661
```

If the reference span of the variant overlaps with any of the regions in the target BED, then the variant is output. If the reference span does not overlap, the variant is not output. For SNPs and Insertions, the reference span is 1 bp. For deletions, the reference span is the length of the deletion.

* `--vc-target-bed-padding`

  **\[Optional]** Pads all target BED regions with the specified value. For example, if a BED region is 1:1000–2000 and the specified padding value is 100, the result is equivalent to using a BED region of 1:900–2100 and a padding value of 0. Any padding added to --vc-target-bed-padding is used by the small variant caller and by the target bed coverage/callability reports. The default padding is 0.
* `--vc-target-coverage`

  Specifies the target coverage for downsampling. The default value is 500 for germline mode and 1000 for somatic mode.
* `--vc-remove-all-soft-clips`

  Set to true to ignore soft-clipped bases during the haploytype assembly step.
* `--vc-decoy-contigs`

  Specifies a comma-separated list of contigs to skip during variant calling. This option can be set in the configuration file.
* `--vc-enable-decoy-contigs`

  Set to true to enable variant calls on the decoy contigs. The default value is false.
* `--vc-enable-phasing`

  Enable variants to be phased when possible. The default value is true.
* `--vc-combine-phased-variants-distance`

  Set the maximum distance in base pairs between phased variants to be combined. The default value is 0, which disables the option. If the user wants to enable the combining of phased variants the recommended value of the distance is 15 base pairs. The valid range is \[0; 15].
* `--vc-enable-mosaic-detection`

  Set to true to enable DRAGEN mosaic detection. Set to false to disable DRAGEN mosaic detection.
* `--vc-mosaic-af-filter-threshold`

  Set the allele frequency threshold for the application of the `MosaicLowAF` filter to mosaic calls. All `MOSAIC` tagged variants with `AF` smaller than the `AF` threshold are filtered with the `MosaicLowAF` filter. The default mosaic `AF` filter threshold is set to `0.2` if the median depth of the sample detected by the ploidy caller is `<= 100x` and `0.1` if the detected depth is `> 100x`.

## Downsampling Options for Small Variant Calling

You can use the following options for downsampling reads in the small variant calling pipeline.

| Option                           | Description                                                                                                                                                                       |
| -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --vc-target-coverage             | Specifies the maximum number of reads covering any given position.                                                                                                                |
| --vc-max-reads-per-active-region | Specifies the maximum number of reads covering a given active region.                                                                                                             |
| --vc-max-reads-per-raw-region    | Specifies the maximum number of reads covering a given raw region.                                                                                                                |
| --vc-min-reads-per-start-pos     | Specifies the minimum number of reads with a start position overlapping any given position.                                                                                       |
| --high-coverage-support-mode     | Applies the high coverage mode down-sample options if set to true. Enabling this option is recommended for targeted panels with coverage over 1000x, but will slow down run time. |

For mitochondrial small variant calling, the downsampling options can be set separately because the mitochondrial contig contains a higher depth than the rest of the contigs in a WGS data set. The following are the downsampling options for the mitochondrial contig.

* `--vc-target-coverage-mito`
* `--vc-max-reads-per-active-region-mito`
* `--vc-max-reads-per-raw-region-mito` The target coverage and max/min reads in raw/active region options are not directly related and could be triggered independently.

The following are the default downsampling values for each small variant calling mode.

| Mode          | Downsampling Option                     | Default Value |
| ------------- | --------------------------------------- | ------------- |
| Germline      | `--vc-target-coverage`                  | 500           |
| Germline      | `--vc-max-reads-per-active-region`      | 10000         |
| Germline      | `--vc-max-reads-per-raw-region`         | 30000         |
| Somatic       | `--vc-target-coverage`                  | 1000          |
| Somatic       | `--vc-max-reads-per-active-region`      | 10000         |
| Somatic       | `--vc-max-reads-per-raw-region`         | 30000         |
| High Coverage | `--vc-target-coverage`                  | 100000        |
| High Coverage | `--vc-max-reads-per-active-region`      | 200000        |
| High Coverage | `--vc-max-reads-per-raw-region`         | 200000        |
| Mitochondrial | `--vc-target-coverage-mito`             | 40000         |
| Mitochondrial | `--vc-max-reads-per-active-region-mito` | 200000        |
| Mitochondrial | `--vc-max-reads-per-raw-region-mito`    | 200000        |

The target coverage downsampling step runs first and is meant to limit the total coverage at a given position. This step is approximate and the coverage after downsampling at a given position could be a bit higher than the threshold due to the `--vc-min-reads-per-start-pos` setting.

If the number of reads at any position with same start position is equal to or lower than the `--vc-min-reads-per-start-pos`, that position is skipped for downsampling to make sure that there is always at least a minimum number of reads (set to `--vc-min-reads-per-start-pos`, default value is 10) at any start position.

The next downsampling step is to apply the `--vc-max-reads-per-raw-region` and `--vc-max-reads-per-active-region` limits. These options are used to limit the total number of reads in an entire region using a leveling downsampling method.

This downsampling mechanism scans each start position from the start boundary of the region and discards one read from that position, then moves on to the next position, until the total number of reads falls below the threshold. It can potentially take several passes across the entire region for the total number of reads in the entire region to fall below the threshold. After the threshold is met, the downsampling step is stopped regardless of which position was considered last in the region.

When downsampling occurs, the choice of which reads to keep or remove is random. However, the random number generator is seeded to a default value to make sure that the generator produces the same set of values in each run. This ensures reproducible results, which means there is no run to run variation when using the same input data.

## Small Variant Caller Output

By default, the DRAGEN small variant caller outputs a VCF file (`<output-file-prefix>.hard-filtered.vcf.gz`) in VCF 4.2 format containing both filtered and PASSing variant records.

### Variant Representation

DRAGEN outputs variants in the VCF file following variant normalization conventions described here <https://genome.sph.umich.edu/wiki/Variant_Normalization>. The normalization of a variant representation in VCF consists of two parts: parsimony and left alignment pertaining to the nature of a variant's length and position respectively.

* Parsimony means representing a variant in as few nucleotides as possible without reducing the length of any allele to 0.
* Left aligning a variant means shifting the start position of that variant to the left till it is no longer possible to do so.
* A variant is normalized if and only if it is parsimonious and left aligned

Additional notes on variant representation in the DRAGEN VCF:

* Reference-trimming of alleles: A single padding reference base is used to represent insertions and deletions (i.e. the reference base preceding the insertion or deletion is included).
* Allele decomposition: By default, phased variants are represented as contiguous individual variant records in the VCF. When phasing can be determined, the FORMAT/GT is phased and the FORMAT/PS contains the coordinate position of the first variant in the set of phased variants; this information indicates which variants have occurred on the same haplotype. DRAGEN offers functionality to merge phased variant records into a single VCF record; please see the [Combine Phased Variants](#combine-phased-variants) section for details.

In some cases, such as complex variants in repetitive regions, some variants cannot be normalized (i.e. converted into a standard representation) or represented uniquely. To counteract this problem, when comparing two VCFs (e.g. a DRAGEN VCF against a truth set VCF), it is recommended to use the RTG vcfeval tool which performs variant comparisons using a haplotype-aware approach. RTG vcfeval has been adopted as the standard VCF comparison tool by GA4GH and PrecisionFDA, as described in [Best Practices for Benchmarking Germline Small Variant Calls in Human Genomes](https://www.biorxiv.org/content/biorxiv/early/2018/02/23/270157.full.pdf).

### Multi-allelic Variants and Overlapping Variants

A multiallelic site is a specific locus in a genome that contains three or more observed alleles, counting the reference as one, and therefore allowing for two or more variant alleles. Multi-allelic calls are output in a single variant record in the VCF as follows:

```
chr1 2656216 . A T,C 107.65 PASS
AC=1,1;AF=0.500,0.500;AN=2;DP=12;FS=0.000;MQ=28.95;QD=8.97;SOR=3.056;FractionInformativeReads=0.750
GT:AD:AF:DP:GQ:PL:GL:GP:PRI:SB:MB
1/2:0,5,4:0.556,0.444:9:15:177,144,46,122,0,72:-17.704,-14.420,-4.626,-12.220,0.000,-7.244:1.076e+02,1.096e+02,1.465e+01,8.758e+01,1.520e-01,4.082e+01:0.00,34.77,37.77,34.77,69.54,37.77:0,0,1,8:0,0,4,5
```

Two indels are considered as multi-allelic if they share the same reference base preceding the indel. For example:

```
chr1 7392258 . C CT,CTTT 234.76 PASS
AC=1,1;AF=0.500,0.500;AN=2;DP=44;FS=0.000;MQ=199.22;QD=5.34;SOR=2.226;FractionInformativeReads=0.659
GT:AD:AF:DP:GQ:PL:GL:GP:PRI:SB:MB
1/2:0,15,14:0.517,0.483:29:50:245,256,55,190,0,55:-24.476,-25.634,-5.492,-18.976,0.000,-5.500:2.348e+02,2.513e+02,5.292e+01,1.848e+02,4.401e-05,5.300e+01:0.00,5.00,8.00,5.00,10.00,8.00:0,0,7,22:0,0,17,12
```

DRAGEN employs joint detection of overlapping variants, a feature designed to detect overlapping SNP and INDEL variants and output them in a single VCF record represented as a multi-allelic genotype. However, if a SNP overlaps an INDEL but the SNP does not align with the reference base preceding the indel, the SNP and INDEL are represented as two different variant records, as shown in the example below.

```
chr1 1029628 . C CGT 49.88 PASS
AC=1;AF=0.500;AN=2;DP=37;FS=7.791;MQ=105.32;MQRankSum=-1.315;QD=1.35;ReadPosRankSum=1.423;SOR=1.510;FractionInformativeReads=0.892;R2\_5P\_bias=-19.742
GT:AD:AF:DP:GQ:PL:GL:GP:PRI:SB:MB:PS
0|1:17,16:0.485:33:48:81,0,50:-8.088,0.000,-5.000:4.988e+01,6.653e-05,5.300e+01:0.00,31.00,34.00:10,7,5,11:11,6,9,7:1029628

chr1 1029629 . A G 50.00 PASS
AC=1;AF=0.500;AN=2;DP=37;FS=1.289;MQ=105.32;MQRankSum=-0.659;QD=1.35;ReadPosRankSum=-0.199;SOR=0.604;FractionInformativeReads=1.000;R2\_5P\_bias=-24.923
GT:AD:AF:DP:GQ:PL:GL:GP:PRI:SB:MB:PS
0|1:16,21:0.568:37:48:85,0,49:-8.477,0.000,-4.934:5.000e+01,6.886e-05,5.234e+01:0.00,34.77,37.77:9,7,10,11:10,6,13,8:1029628
```

### QUAL, QD, and GQ Formulation

In single sample VCF and gVCF, the QUAL follows the definition of the VCF specification. For more information on the VCF specification, see the most current VCF documentation available on samtools/hts-specs GitHub repository.

* QUAL is the Phred-scaled probability that the site has no variant and is computed as:

  ```
  QUAL = -10\*log10 (posterior genotype probability of a
  homozygous-reference genotype (GT=0/0))
  ```

  That is, QUAL = GP (GT=0/0), where GP = posterior genotype probability in Phred scale. QUAL = 20 means there is 99% probability that there is a variant at the site. The GP values are also given in Phred-scaled in the VCF file.
* GQ for non-homref calls is the Phred-scaled probability that the call is incorrect. GQ=-10\*log10(p), where p is the probability that the call is incorrect. GQ=-10\*log10(sum(10.^(-GP(i)/10))) where the sum is taken over the GT that did not win. GQ of 3 indicates a 50 percent chance that the call is incorrect, and GQ of 20 indicates a 1 percent chance that the call is incorrect.
* In gvcf mode, the evidence in favor of homozygous reference calls is also assessed. However, the posterior probability is not of interest in this case (with zero evidence, e.g. due to zero coverage, the strong prior in favor of homref would yield a strong posterior in favor of homref), so the value of GQ for homref calls reflects the evidence directly, defined using the likelihood ratio between the likelihoods for the homref hypothesis and the strongest competing variant hypothesis: 10\*log10\[P(D|homref)/P(D|variant)] where D represents the pileup data.
* QD is the QUAL normalized by the read depth, DP.

| Metric                | QUAL                                                                                                                               | GQ (non-homref)                                                                              | GQ (homref)                                                | QD                       |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | ------------------------ |
| **Description**       | Probability that the site has no variant                                                                                           | Probability that the call is incorrect                                                       | Evidence supporting homref call                            | Qual normalized by depth |
| **Formulation**       | QUAL = GP(GT=0/0)                                                                                                                  | GQ =-10\*log10(p)                                                                            | GQ = 10\*log10\[P(D\|homref)/P(D\|variant)]                | QUAL/DP                  |
| **Scale**             | Unsigned Phred                                                                                                                     | Unsigned Phred                                                                               | Signed Phred                                               | Unsigned Phred           |
| **Numerical example** | <p>QUAL=20: 1 % chance that there is no variant at the site.<br>Qual=50: 1 in 1e5 chance that there is no variant at the site.</p> | <p>GQ=3, 50% that the call is incorrect.<br>GQ=20, 1% change that the call is incorrect.</p> | <p>GQ=0: no evidence.<br>GQ>0: evidence favors homref.</p> |                          |

The QUAL scores generated by DRAGEN differ significantly from those of GATK, as DRAGEN's algorithms for small variant detection provide more realistic scores. This improvement stems from two key factors:

* Correlated Errors: DRAGEN accounts for real-world correlated errors, unlike GATK, which assumes errors are uncorrelated, leading to inflated QUAL scores in GATK.
* Machine Learning (ML): DRAGEN-ML further recalibrates QUAL scores, making them more accurate than DRAGEN without ML. With ML enabled, QUAL scores tend to not exceed 75, compared to GATK, where they can exceed 1000. Consequently, DRAGEN-ML uses a lower QUAL filtering threshold (3.0103) compared to DRAGEN without ML (10.41 for SNP and 7.83 for Indel).

![](/files/43u5NziZfCFJUWoap8JP)

Our recommendation is to use the default filtering thresholds in DRAGEN: QUAL threshold of 3.0103 with ML enabled.

### gVCF Output

A genomic VCF (gVCF) file contains information on variants and positions determined to be homozygous to the reference genome. For homozygous regions, the gVCF file includes statistics that indicate how well reads support the absence of variants or alternative alleles. The gVCF file includes an artificial `<NON_REF>` allele. Reads that do not support the reference or any variants are assigned the `<NON_REF>` allele. DRAGEN uses these reads to determine if the position can be called as a homozygous reference, as opposed to remaining uncalled. The resulting score represents the Phred-scaled level of confidence in a homozygous reference call. In germline mode, the score is `FORMAT/GQ` and in somatic mode the score is `FORMAT/SQ`.

The following options are available to enable and control gVCF output.

* `--vc-emit-ref-confidence`

  To enable gVCF output, set to `GVCF`. By default, contiguous runs of homozygous reference calls with similar scores are collapsed into blocks (hom-ref blocks). Hom-ref blocks save disk space and processing time of downstream analysis tools. DRAGEN recommends using the default mode.

  To produce unbanded output, set `--vc-emit-ref-confidence` to `BP_RESOLUTION`.
* `--vc-enable-vcf-output`

  To enable VCF file output during a gVCF run, set to true. The default value is false.
* `--vc-gvcf-bands`

  If using the default `--vc-emit-ref-confidence gvcf` (banded mode), DRAGEN collapses gVCF records with a similar GQ or SQ score. By default, the cutoffs are `1 10 20 30 40 60 80` for germline and `1 3 10 20 50 80` for somatic. For example, to define the bands \[0, 10), \[10, 50), and ≥ 50 use `--vc-gvcf-bands 10 50`.
* `--vc-compact-gvcf`

  This option, when used for germline in conjunction with `--vc-emit-ref-confidence gvcf`, produces a much smaller gVCF output file than the default. It can be used when the gVCF is destined for ingestion into gVCF Genotyper, offering further savings on disk space and gVCF Genotyper runtime compared to the default. This option implies `--vc-gvcf-bands 0 1 10 20 30` and additionally omits certain metrics that are not used by gVCF Genotyper. Note that files generated using this option will be rejected by the Pedigree Caller.

Not all entries in the gVCF are contiguous. The file might contain gaps that are not covered by either a variant line or a hom-ref block. The gaps correspond to regions that are not callable. A region is not callable if there is not at least one read mapped to the region with a MAPQ score above zero.

In germline mode, the thresholds for calling are lower for gVCFs than for VCFs. The gVCF output could show a different number of variants than a VCF run for the same sample. There is likely a different number of biallelic and multiallelic calls because gVCF mode includes all possible alleles at a locus, rather than only the two most likely alleles. This means that a biallelic call in the VCF can be output as a multiallelic call in the gVCF. The genotype in the gVCF still points to the two most likely alleles, so the variant call remains the same.

The following are example gVCF records that include a hom-ref block call and a variant call.

```
1 39224 . C <NON_REF> . PASS END=39260
GT:AD:DP:GQ:MIN_DP:PL:SPL:ICNT
0/0:2,0:2:3:1:0,3,37:0,3,37:3,0

1 39261 . T C,<NON_REF> 15.59 PASS
DP=3;MQ=12.73;MQRankSum=0.736;ReadPosRankSum=0.736;FractionInformativeReads=1.000
GT:AD:AF:DP:F1R2:F2R1:GQ:PL:SPL:ICNT:GP:PRI:SB:MB
0/1:1,2,0:0.667,0.000:3:1,0,0:0,2,0:5:49,0,1,69,7,75:66,0,8:1,0:1.5592e+01,1.5915e+00,5.5412e+00,7.0100e+01,4.3330e+01,8.0068e+01:0.00,34.77,37.77,34.77,69.54,37.77:0,1,0,2:0,1,2,0
```

In single sample gVCF, FORMAT/DP reported at a HomRef position is the median DP in the band and AD is the corresponding value, so sum of AD will be DP even in a homref band. The minimum is also computed and printed as MIN\_DP for the band.

## Phasing and Phased Variants

DRAGEN supports output of phased variant records in both the germline and the somatic VCF and gVCF files. When two or more variants are phased together, the phasing information is encoded in a sample-level annotation, FORMAT/PS. FORMAT/PS identifies which set the phased variant is in. The value in the field in an integer representing the position of the first phased variant in the set. All records in the same contig with matching PS values belong to the same phase set.

```
##FORMAT=<ID=PS,Number=1,Type=Integer,Description="Physical phasing
ID information, where each unique ID within a given sample (but not
across samples) connects records within a phasing group">
```

The following is an example of a DRAGEN single sample gVCF, where two SNPs are phased together.

```
chr1 1947645 . C T,<NON_REF> 48.44 PASS
DP=35;MQ=250.00;MQRankSum=4.983;ReadPosRankSum=3.217;FractionInformativeReads=1.000;R2_5P_bias=0.000
GT:AD:AF:DP:F1R2:F2R1:GQ:PL:SPL:ICNT:GP:PRI:SB:MB:PS
0|1:20,15,0:0.429:35:9,7,0:11,8,0:47:83,0,50,572,758,622:255,0,255:19,0:4.844e+01,8.387e-05,5.300e+01,4.500e+02,4.500e+02,4.500e+02:0.00,34.77,37.77,34.77,69.54,37.77:11,9,10,5:12,8,8,7:1947645

chr1 1947648 . G A,<NON_REF> 50.00 PASS
DP=36;MQ=250.00;MQRankSum=5.078;ReadPosRankSum=2.563;FractionInformativeReads=1.000;R2_5P_bias=0.000
GT:AD:AF:DP:F1R2:F2R1:GQ:PL:SPL:ICNT:GP:PRI:SB:MB:PS
1|0:16,20,0:0.556:36:8,9,0:8,11,0:48:85,0,49,734,613,698:255,0,255:16,0:5.000e+01,7.067e-05,5.204e+01,4.500e+02,4.500e+02,4.500e+02:0.00,34.77,37.77,34.77,69.54,37.77:10,6,11,9:8,8,12,8:1947645
```

During the genotyping step, all haplotypes and all variants are considered over an active region. For each pair of variants, if both variants occur on all of the same haplotypes or if either is a homozygous variant, they are phased together. If the variants only occur on different haplotypes, they are phased opposite to each other. If any heterozygous variants are present on some of the same haplotypes but not others, phasing is aborted and no phasing information is output for the active region.

### Combine Phased Variants

The DRAGEN small variant caller supports combining multiple nearby phased variant records into a single VCF record. When the functionality is enabled, the caller will output both multi-nucleotide variants (MNVs; multiple phased SNVs combined into a single VCF record) and complex indels (multiple phased insertions, deletions, and/or SNVs combined into a single VCF record) in the VCF. FGT‑tagged variants are excluded and skipped during MNV combination.

For example, assuming reference at position `chr2 115035` is `A`, the following two phased SNVs can be combined into an MNV.

```
chr2 115034 . G C GT:PS 0|1:115034
chr2 115036 . C T GT:PS 0|1:115034
```

The phased SNVs are combined as follows.

```
chr2 115034 . GAC CAT GT:PS 0|1:115034
```

The following two phased indels can be combined into a complex indel.

```
chr2 61569261 . GCACA G     GT:PS 0|1:61569261
chr2 61569266 . C     CGTGG GT:PS 0|1:61569261
```

The phased indels are combined as follows.

```
chr2 61569263 . ACAC GTGG GT:PS 0|1:61569261
```

Individual variant records existing on the same haplotypes are deemed to be in phase and will be merged if they are within a configurable distance threshold of one another. For each consecutive pair of phased variants in a phase set, the variants will be combined if the difference between their POS fields does not exceed the threshold. For deletions, the number of deleted bases is taken into account and subtracted from the POS difference between the deletion and downstream phased variant when calculating the distance between calls. Please note that variant records without a PS tag may be merged into MNVs and complex indels together with calls having PS tags due to algorithmic differences between variant phasing and variant merging.

In the somatic pipeline, combining phased variants is enabled by default, consistent with HGVS guidelines. In the germline pipeline, the functionality can be enabled using the command line options detailed below.

#### Command line options for merging phased variants

* `--vc-combine-phased-variants-distance` Specifies the maximum distance over which phased insertions, deletions, and SNVs will be combined into an MNV or complex indels. This threshold applies to any phased variant cluster containing SNVs, indels, or a mixture of both. The default is `2` for somatic and `0` for germline (disabled).
* `--vc-combine-phased-variants-distance-snvs-only` Specifies the maximum distance over which phased SNVs (but not complex indels) will be combined into an MNV. This value applies exclusively to phased variant groups composed entirely of SNVs, and defaults to the value specified for complex indels `vc-combine-phased-variants-distance` above.

For both options, a value of 0 disables merging. When either option is enabled with a value \[1, 15], all phased variants within the specified distance are merged into an MNV for `vc-combine-phased-variants-distance-snvs-only`, or into an MNV / complex indel for`vc-combine-phased-variants-distance`.

* `--vc-mnv-emit-component-calls` Specifies whether or not to emit the individual component variant records along with the merged variant records. When set to `true`, all component calls making up an MNV or complex indel will be emitted in the VCF along with the merged variant record. The default is `true` for somatic and `false` (disabled) for germline.
* `--vc-combine-phased-variants-max-vaf-delta` Specifies the threshold for filtering MNV component variant calls when the events comprising to the MNV have different allele frequencies. The default value is 0.1, which means that any SNV or INDEL with an AF that is more than 0.1 greater than the MNV AF shall be emitted as a PASSing call, while the remaining components shall be emitted with the 'mnv\_component' FILTER flag. Only applicable when `vc-combine-phased-variants-distance` is greater than 0 and `vc-mnv-emit-component-calls` is true. (Default=0.1)

DRAGEN can output all component SNVs and/or INDELs that make up a merged MNV or complex indel along with the merged call itself. Merged calls and their component calls can be identified and linked to one another by a common value in the INFO.MNVTAG field. This behavior is default in somatic mode and can be enabled in germline mode by setting `--vc-mnv-emit-component-calls=true`. When `vc-mnv-emit-component-calls` is enabled and DRAGEN reports an MNV or complex indel call, the component calls that make up the merged call are filtered with the `mnv_component` filter flag unless the difference in VAF between the component call and merged call is greater than the value of `vc-combine-phased-variants-max-vaf-delta` (default: 0.1). This avoids component calls being doubly represented in the VCF output. In the case where VAF difference between a given component call and merged call exceeds the threshold value of `vc-combine-phased-variants-max-vaf-delta`, that is considered evidence for the component call existing both as a standalone variant and as part of the MNV or complex indel and the component call will be emitted as a PASSing VCF record. For example, in the following MNV group, there are two component SNVs making up the MNV. The MNV call is emitted as a PASSing call while one component SNV with AF equal to that of the MNV is filtered with the `mnv_component` FILTER flag and the other component SNV with VAF greater than that of the MNV by more than 0.1 is emitted as a PASSing call.

```
chr1    1771073 .       T       C       .       mnv_component
DP=65;MQ=250.00;FractionInformativeReads=0.892;SoftClipRatio=0.00;MNVTAG=chr1:1771073_TAT->CAC
GT:SQ:AD:AF:F1R2:F2R1:DP:SB:MB  0/1:42.70:17,41:0.7069:10,18:7,23:58:10,7,13,28:9,8,22,19

chr1    1771073 .       TAT     CAC     .       PASS
DP=65;MQ=250.00;FractionInformativeReads=0.892;SoftClipRatio=0.00;MNVTAG=chr1:1771073_TAT->CAC;GermlineStatus=Germline_DB
GT:SQ:AD:AF:F1R2:F2R1:DP:SB:MB  0/1:42.70:17,41:0.7069:10,18:7,23:58:10,7,13,28:9,8,22,19

chr1    1771075 .       T       C       .       PASS
DP=67;MQ=250.00;FractionInformativeReads=0.881;SoftClipRatio=0.00;STR;RU=AC;RPA=7;MNVTAG=chr1:1771073_TAT->CAC;GermlineStatus=Germline_DB
GT:SQ:AD:AF:F1R2:F2R1:DP:SB:MB  1/1:62.99:0,59:1.0000:0,29:0,30:59:0,0,24,35:0,0,32,27
```

DRAGEN supports phasing of the genotypes listed in the below table. Only the first row in the table is relevant to somatic, since the somatic pipeline only emits 0/1 and 0|1 genotypes. MNV calls can still be phased with other variant calls that fell outside the phased variants distance.

| GT variant 1 | GT variant 2 | GT MNV | Relevant Pipeline    | Supported in DRAGEN |
| ------------ | ------------ | ------ | -------------------- | ------------------- |
| 0\|1         | 0\|1         | 0/1    | Germline and Somatic | Yes in 4.0          |
| 0/1          | 1/1          | 1/2    | Germline             | No                  |
| 0/1          | 1/2          | 1/2    | Germline             | No                  |
| 1/1          | 1/1          | 1/1    | Germline             | Yes in 4.2          |

Examples of diploid haplotypes where phasing is supported:

```
-------------------------------------------------------------- H0 ( REF ) 
-----------------x---------------------------y---------------- H1
```

```
-----------------x---------------------------y---------------- H1
-----------------x---------------------------y-----------------H1
```

Examples of diploid haplotypes where phasing is not supported:

```
-----------------x---------------------------y---------------- H1
---------------------------------------------y---------------- H2
```

```
-----------------x----------------------------y--------------- H1
----------------------------------------------z--------------- H2
```

## Ploidy Support

The small variant caller currently only supports either ploidy 1 or 2 on all contigs within the reference except for the mitochondrial contig, which uses a continuous allele frequency approach (see Mitochondrial Calling). The selection of ploidy 1 or 2 for all other contigs is determined as follows.

* If `--sample-sex` is not specified on the command line, the Ploidy Estimator determines the sex. If the Ploidy Estimator cannot determine the sex karyotype or detects sex chromosome aneuploidy, all contigs are processed with ploidy 2.
* If `--sample-sex` is specified on the command line, contigs are processed as follows.
  * For female samples, DRAGEN processes all contigs with ploidy 2, and marks variant calls on chrY with a filter PloidyConflict.
  * For male samples, DRAGEN processes all contigs with ploidy 2, except for the sex chromosomes. DRAGEN processes chrX with ploidy 1, except in the PAR regions, where it is processed with ploidy 2. chrY is processed with ploidy 1 throughout.

For male samples in germline calling mode, DRAGEN calls potential mosaic variants in non-PAR regions of sex chromosomes. A variant is called as mosaic when the allele frequency (FORMAT/AF) is below 75% or if multiple alt alleles are called, suggesting incompatibility with the haploid assumption. The GT field for bi-allelic mosaic variants is "0/1", denoting a mixture of reference and alt alleles, as opposed to the regular GT of "1" for haploid variants. The GT field for multi-allelic mosaic variants is "1/2" in VCF. You can disable the calling of mosaic variants by setting `--vc-enable-sex-chr-diploid` to false.

An example germline VCF record of a mosaic variant in a haploid region: chrX 18622368 . C T 48.84 PASS AC=1;AF=0.500;AN=2;DP=22;FS=4.154;MQ=248.02;MQRankSum=3.272;QD=2.27;ReadPosRankSum=2.671;SOR=1.546;FractionInformativeReads=1.000;`MOSAIC` GT:AD:AF:DP:F1R2:F2R1:GQ:PL:GP:PRI:SB:MB 0/1:9,13:`0.5909`:22:1,8:8,5:48:84,0,51:4.8837e+01,7.4031e-05,5.4007e+01:0.00,34.77,37.77:5,4,4,9:3,6,5,8

DRAGEN detects sex chromosomes by the naming convention, either X/Y or chrX/chrY. No other naming convention is supported.

| --sample-sex   | Ploidy Estimation | Sample Sex in Small VC |
| -------------- | ----------------- | ---------------------- |
| male           | Not relevant      | Male                   |
| female         | Not relevant      | Female                 |
| none           | Not relevant      | None                   |
| auto (default) | XY                | Male                   |
| auto (default) | XX                | Female                 |
| auto (default) | Everything else   | None                   |

## Overlapping Mates in the Small Variant Calling

Instead of treating overlapping mates as independent evidence for a given event, DRAGEN handles overlapping mates in both the germline and somatic pipelines as follows.

* When the two overlapping mates agree with each other on the allele with the highest HMM score, the genotyper uses the mate with the greatest difference between the highest and the second highest HMM score. The HMM score of the other mate becomes zero.
* When the two overlapping mates disagree, the genotyper sums the HMM score between the two mates, assigns the combined score to the mate that agrees with the combined result, and changes the HMM score of the other mate to zero.
* The base qualities of overlapping mates are no longer adjusted.

## Mitochondrial Calling

Typically, there are approximately 100 mitochondria in each mammalian cell. Each mitochondrion harbors 2–10 copies of mitochondrial DNA (mtDNA). For example, if 20% of the chrM copies have a variant, then the allele frequency (AF) is 20%. This is also referred to as continuous allele frequency. The expectation is that the AF of variants on chrM is anywhere between 0% and 100%.

DRAGEN processes chrM through a continuous AF pipeline, which is similar to the somatic variant calling pipeline. In this case, a single ALT allele is considered and the AF is estimated. The estimated AF can be anywhere between 0% and 100%. Default variant AF thresholds are applied to mitochondrial variant calling.

* `--vc-enable-af-filter-mito`\
  Whether to enable the allele frequency for mitochondrial variant calling. The default is true.
* `--vc-af-call-threshold-mito`\
  Set the threshold for emitting calls in the VCF. The default is 0.01.
* `--vc-af-filter-threshold-mito`\
  Set the threshold to mark emitted vcf call as filtered. The default is 0.02.

QUAL and GQ are not output in the chrM variant records. Instead, the confidence score is FORMAT/SQ, which gives the Phred-scaled confidence that a variant is present at a given locus. A call is made if FORMAT/SQ> vc-sq-call-threshold (default = 3.0). For details on mitochondrial downsampling options, see [Downsampling Options for Small Variant Calling](#downsampling-options-for-small-variant-calling).

```
##FORMAT=<ID=SQ,Number=A,Type=Float,Description="Somatic quality">
```

The following filters can be applied to mitochondrial variant calls.

* `--vc-sq-call-threshold`\
  Set the SQ threshold for emitting calls in the VCF. The default is 0.1.
* `--vc-sq-filter-threshold`\
  Set the SQ threshold to mark emitted VCF calls as filtered. The default is 3.0
* `--vc-enable-triallelic-filter` Enables the multiallelic filter. The default value is false.

If FORMAT/SQ < vc-sq-call-threshold, the variant is not emitted in the VCF. If FORMAT/SQ > vc-sq-call-threshold but FORMAT/SQ < vc-sq-filter-threshold, the variant is emitted in the VCF but FILTER=weak\_evidence.

If FORMAT/SQ> vc-sq-call-threshold, FORMAT/SQ > vc-sq-filter-threshold, and no other filters are triggered, the variant is output in the VCF and FILTER=PASS.

The following are example VCF records on the chrM. The examples show one call with very high AF and another with low AF. In both cases FORMAT/SQ > vc-sq-call-threshold. FORMAT/SQ is also > vc-sq-filter-threshold, so the FILTER annotation is PASS.

```
chrM    513     .       GCA     G       .       PASS    DP=4937;MQ=235.28;FractionInformativeReads=0.883
GT:SQ:AD:AF:F1R2:F2R1:DP:SB:MB  1/1:95.46:33,4327:0.992:7,1081:26,3246:4360:31,2,2371,1956:10,23,2811,1516

chrM    7028    .       C       T       .       PASS    DP=8868;MQ=60.19;FractionInformativeReads=0.993 
GT:SQ:AD:AF:F1R2:F2R1:DP:SB:MB  0/1:21.48:8622,181:0.021:4190,82:4432,99:8803:4344,4278,94,87:5032,3590,101,80
```

### FORMAT/GT

For homref calls (e.g. in NON\_REF regions of gVCF output) the FORMAT/GT is hard-coded to 0/0. The FORMAT/AF yields an estimate on the variant allele frequency, which ranges anywhere within \[0,1]. For variant calls with FORMAT/AF < 95%, the FORMAT/GT is set to 0/1. For variants with very high allele frequencies (FORMAT/AF ≥ 95%), the FORMAT/GT is set to 1/1.

The following is an example of a variant record on chrM in a trio joint VCF. The variant was detected in the second sample with a confidence score that passed the filter threshold. In the first and third samples GT=0/0, which indicates a tentative hom-ref call (ie, that position for the sample is in a NON\_REF region over which no variant was detected with sufficient confidence), but the weak\_evidence filter tag indicates that this call is made with low confidence.

```
chrM 2623 . A G . PASS DP=18772;MQ=111.77 GT:AD:AF:DP:FT:SQ:F1R2:F2R1 0/0:6841,7:0.001:4334:weak_evidence:0:.:. 0/1:6736,2053:0.234:8789:PASS:21.32:3394,1060:3342,993 0/0:6086,9:0.001:5613:weak_evidence:0:.:.
```

## Personalized Germline Small Variant Calling

We leverage the new pangenome reference and multi-genome mapper output to compute a personalized 2-haplotype reference for the input sample.

The computed 2-haplotype reference is used to impute variants, adjust priors probabilities for genotypes in the variant caller, create a new personalized machine learning model and significantly boosts accuracy of variant calling. False negatives are reduced by adjusting genotype priors based on imputed phased variants in the computed haplotypes. False positives are reduced by limiting the impact of noise from other population haplotypes.

To enable personalized variant calling, including the personalized machine learning model, set `--enable-personalization` to true (default: false). This outputs two files in the output directory: `.personal_haplotypes.tsv.gz` and `.personal.vcf.gz`.

`.personal_haplotypes.tsv.gz` describes the personalized 2-haplotype reference. Each line contains the following fields: `#CHROM START END HAPS`. By default each line represents a 4kbp bin of the reference genome (indicated by the `CHROM`, `START` and `END` fields). For each 4kbp bin, the `HAPS` field denotes the pair of ancestral haplotypes (from the pangenome reference panel) that are inferred for the sample.

`.personal.vcf.gz` describes the variants imputed for the sample. Each variant is annotated along with genotype (`GT`), posterior probabilities (`PGP`, Personalized Genotype Posterior) and the inferred best haplotype pair (`HAPS`). The variant quality score (`QUAL`) is computed as `-10 * log10(probability that the imputed genotype is incorrect)` and is capped at 999.

The `.personal.vcf.gz` is useful when running in split mode and is beneficial to save along with the BAM/CRAM output. To enable personalized variant calling and machine learning in split-run scenarios, simply provide the personal variant VCF (`.personal.vcf.gz`) along with the BAM/CRAM input (`--enable-personalization true --vc-pg-variants <$OUT_PREFIX.personal.vcf.gz>`).

Note that personalization is only available for the germline small variant caller (WGS and WES) when used with a pangenome reference.

## Joint Detection of Overlapping Variants

When variants at multiple loci in a single active region are detected jointly, genotyping can benefit. DRAGEN combines loci into a joint detection region if the following conditions are met:

* Loci have alleles that overlap each other.
* Loci are in the STR region or less than 10 bases apart of the STR region.
* Loci are less than 10 bases apart of each other.

Joint detection generates a haplotype list where all possible combinations of the alleles in the joint detection regions are represented. This calculation leads to a larger number of haplotypes. During genotyping, joint detection calculates the likelihoods that each haplotype pair is the truth, given the observed read pileup. Genotype likelihoods are calculated as the sum of the likelihoods of haplotype pairs that support the alleles in the genotypes. Genotypes with maximum likelihood are reported.

Joint detection is enabled by default. To disable joint detection, set `--vc-enable-joint-detection` to false.

## Modeling of Correlated Errors Across Reads

DRAGEN has two algorithms that model correlated errors across reads in a given pileup.

### Foreign Read Detection

Foreign read detection (FRD) detects mismapped reads. FRD modifies the probability calculation to account for the possibility that a subset of the reads were mismapped. Instead of assuming that mapping errors occur independently per read, FRD estimates the probability that a burst of reads is mismapped, by incorporating such evidence as MAPQ and skewed AF.

Mapping errors typically occur in bursts, but treating mapping errors as independent error events per read can result in high confidence scores in spite of low MAPQ and/or skewed AF. One possible strategy to mitigate overestimation of confidence scores is to include a threshold on the minimum MAPQ used in the calculation. However, this strategy can discard evidence and result in false positives.

FRD extends the legacy genotyping algorithm by incorporating an additional hypothesis that reads in the pileup might be foreign reads (ie, their true location is elsewhere in the reference genome). The algorithm exploits multiple properties (skewed allele frequency and low MAPQ) and incorporates this evidence into the probability calculation.

Sensitivity is improved by rescuing FN, correcting genotypes, and enabling lowering of the MAPQ threshold for incoming reads into the variant caller. Specificity is improved by removing FP and correcting genotypes.

### Base Quality Dropoff

The base quality drop off (BQD) algorithm detects systematic and correlated base call errors caused by the sequencing system. BQD exploits certain properties of those errors (strand bias, position of the error in the read, base quality) to estimate the probability that the alleles are the result of a systematic error event rather than a true variant.

Bursts of errors that occur at a specific locus have distinct characteristics differentiating them from true variants. The base quality drop off (BQD) algorithm is a detection mechanism that exploits certain properties of those errors (strand bias, position of the error in the read, low mean base quality over said subset of reads at the locus of interest) and incorporates them into the probability calculation.


# ROH Caller

#### ROH Caller

Regions of homozygosity (ROH) are detected as part of the small variant caller. The caller detects and outputs the runs of homozygosity from whole genome calls on autosomal human chromosomes. Sex chromosomes are ignored unless the sample sex karyotype is XX, as specified on the command line or determined by the Ploidy Estimator. ROH output allows downstream tools to screen for and predict consanguinity between the parents of the proband subject.

A region is defined as consecutive variant calls on the chromosome with no large gap in between these variants. In other words, regions are broken by chromosome or by large gaps with no SNV calls. The gap size is set to 3 Mbases.

An alternative algorithm to detect ROH regions is provided on the Germline WGS ASCN caller, see [relevant section for potential differences between the two approaches](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline#comparison-with-roh-caller).

**ROH Algorithm**

The ROH algorithm runs on the small variant calls. The algorithm excludes variants with multiallelic sites, indels, complex variants, non-PASS filtered calls, and homozygous reference sites. The variant calls are then filtered further using a block list BED, and finally depth filtering is applied after the block list filter. The default value for the fraction of filtered calls is 0.2, which filters the calls with the highest 10% and lowest 10% in DP values. The algorithm then uses the resulting calls to find regions.

The ROH algorithm first finds seed regions that contain at least 50 consecutive homozygous SNV calls with no heterozygous SNV or gaps of 500,000 bases between the variants. The regions can be extended using a scoring system that functions as follows.

* Score increases with every additional homozygous variant (0.025) and decreases with a large penalty (1-0.025) for every heterozygous SNV. This provides some tolerance of presence of heterozygous SNV in the region.
* Each region expands on both ends until the regions reach the end of a chromosome, a gap of 500,000 bases between SNVs occurs, or the score becomes too low (0).

Overlapping regions are merged into a single region. Regions can be merged across gaps of 500,000 bases between SNVs if a single region would have been called from the beginning of the first region to the end of the second region without the gap. There is no maximum size for regions, but regions always end at chromosome boundaries.

**ROH Options**

* `--vc-enable-roh` Set to true to enable the ROH caller. The ROH caller is enabled by default for human autosomes only. Set to false to disable.
* `--vc-roh-blacklist-bed` If provided, the ROH caller ignores variants that are contained in any region in the block list BED file. DRAGEN distributes block list files for all popular human genomes and automatically selects a block list to match the genome in use, unless this option is used to select a file.
* `--vc-roh-enable-depth-filter` Enable depth filter for ROH. The depth filter is enabled by default. Set to false to disable.
* `--vc-roh-min-seed-size` The minimum number of consecutive homozygous SNPs to form a ROH (Default=50)
* `--vc-roh-max-gap-length` The maximum gap length (bp) between two homozygous SNPs to be included in the same region (Default=500000)
* `--vc-roh-error-rate` The rate of genotyping errors (Default=0.025)

**ROH Output**

The ROH caller produces an ROH output file named `<output-file-prefix>.roh.bed` in which each row represents one region of homozygosity. The BED file contains the following columns:

`Chromosome Start End Score #Homozygous #Heterozygous`

* Score is a function of the number of homozygous and heterozygous variants, where each homozygous variant increases the score by 0.025, and each heterozygous variant reduces the score by 0.975.
* Start and end positions are a 0-based, half-open interval.
* \#Homozygous is number of homozygous variants in the region.
* \#Heterozygous is number of heterozygous variants in the region. The caller also produces a metrics file named `<output-file-prefix>.roh_metrics.csv` that lists the number of large ROH and percentage of SNPs in large ROH (>3 MB).

#### Concordance with PLINK

The table below demonstrates how the PLINK options can be tuned to behave similarly to the DRAGEN ROH caller default settings (see column DRAGEN default). We observed that PLINK ROH calls (see column PLINK default) in default settings are more conservative compared to DRAGEN default settings. By default, PLINK reports ROH regions of size 1MB or larger (see PLINK option --homozyg-kb ) with at least 100 homozygous SNPs (see PLINK option --homozyg-snp) while DRAGEN ROH caller reports smaller regions with at least 50 homozygous SNPs (see DRAGEN ROH Algorithm section). In addition, PLINK by default allows for only 1 heterozygous SNP per scanning window (specified by PLINK option --homozyg-window-het) while DRAGEN uses a soft score threshold penalty without setting an upper bound on the allowed number of heterozygous SNPs (see DRAGEN ROH Algorithm section). The PLINK ROH calls are largely comparable to the DRAGEN ROH calls after relaxing the default PLINK settings, shown in column PLINK tuned. Prior to PLINK ROH calling, the input DRAGEN hard-filtered VCF files are filtered as per the instructions in DRAGEN ROH Algorithm section.

| PLINK option               | PLINK default | PLINK tuned | DRAGEN default                                                                    | PLINK Definitions                                                                                                                      |
| -------------------------- | ------------- | ----------- | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
| --homozyg-density          | 50            | 50          |                                                                                   | Minimum required density to call a ROH (1 SNP in 50 kb), can be increased to relax the per SNP density.                                |
| --homozyg-gap              | 1000          | 1000        | 3000                                                                              | Maximal interval between two homozygous SNPs in a ROH (in kb)                                                                          |
| --homozyg-kb               | 1000          | 500         | All sizes reported                                                                | Minimal length of reported ROH (in kb)                                                                                                 |
| --homozyg-snp              | 100           | 50          | 50                                                                                | Minimal number of homozygous SNPs in the reported ROH                                                                                  |
| --homozyg-window-het       | 1             | 2           | Soft score threshold (1-0.025) penalty for a het SNP and 0.025 gain for a hom SNP | Maximum number of heterozygous SNPs allowed in a scanning window                                                                       |
| --homozyg-window-missing   | 5             | 5           |                                                                                   | Number of missing calls allowed in a scanning window                                                                                   |
| --homozyg-window-snp       | 50            | 50          |                                                                                   | Variants in a scanning window                                                                                                          |
| --homozyg-window-threshold | 0.05          | 0.05        |                                                                                   | For a SNP to be eligible for inclusion in a ROH, the hit rate/overlap of all scanning windows containing the SNP must be at least 0.05 |


# B-Allele Frequency Output

B-Allele frequency (BAF) output is enabled by default in germline and somatic VCF and gVCF runs.

The BAF value is calculated as either `AF` or `(1 - AF)`, where

* `AF = (alt_count / (ref_count + alt_count))`
* `BAF = 1 - AF`, only when ref base < alt base, order of priority for bases is `A < T < G < C < N`.

The B-allele frequency values are often plotted to visually inspect the spread away from a perfectly diploid heterozygous call (BAF=50%). This plot is more easily interpreted if it is symmetric about the BAF=50% line. To ensure the symmetry, a heuristic must be used to determine when `BAF = AF` or `BAF = 1-AF`. This definition of B-Allele Frequency is based on the definition that is used for bead arrays, as most users are accustomed to that implementation. Here, the choice of the B allele is based on the color of dye attached to each nucleotide. A and T get one color, G and C get the other color. The bead array implementation has much more complex rule for tie-breaking between A and T or G and C that involves top and bottom strands. This is unnecessary and so the simpler hierarchical approach of using a priority for the nucleotides `A<T<G<C<N` is used.

For each small variant VCF entry with exactly one SNP alternate allele, the output contains a corresponding entry in the BAF output file.

* `<NON_REF>` lines are excluded
  * ForceGT variants (as marked by the "FGT" tag in the INFO field) are not included in the output, unless the variant also contains the "NML" tag in the INFO field.
  * Variants where the ref\_count and alt\_count are both zero are not included in the output.

## BAF Options

* `--vc-enable-baf` Enable or disable B-allele frequency output. Enabled by default.

## BAF Output

The BF generates are BigWig-compressed files, named `<output-file-prefix>.baf.bw` and `<output-file-prefix>.hard-filtered.baf.bw`. The hard-filtered file only contains entries for variants that pass the filters defined in the VCF (ie, PASS entries).

Each entry contains the following information: `Chromosome Start End BAF`

Where:

* Chromosome is a string matching a reference contig.
* Start and end values are zero-based, half open intervals.
* BAF is a floating point value.


# Somatic Mode

### Somatic Mode

The DRAGEN Somatic Pipeline allows ultrarapid analysis of Next-Generation Sequencing (NGS) data to identify cancer-associated mutations in somatic chromosomes. DRAGEN calls SNVs and indels from both matched tumor-normal pairs and tumor-only samples using a probability model that considers the possibility of somatic variants, germline variants, and various systematic noise artifacts. The model is informed by sample-specific nucleotide and indel noise patterns that are estimated from the data at runtime. When considering somatic variants, DRAGEN does not make any ploidy assumptions, which enables detection of low-frequency alleles. For loci with coverage up to 100x in the tumor sample, DRAGEN can detect variant allele frequencies down to approximately 5%. This limit scales with increasing depth on a per-locus basis. It is recommended to provide DRAGEN with a systematic noise file that contains position- and allele-specific noise frequencies as estimated from a panel of normal samples (see below); DRAGEN uses this noise file to filter calls that can be explained as resulting from position- and allele-specific noise. After multiple filtering steps, the output is generated as a VCF file. Variants that fail the filtering steps are kept in the output VCF. The variants include a FILTER annotation that indicates which filtering steps have failed.

For the tumor-normal pipeline, both samples are analyzed jointly. DRAGEN assumes that germline variants and systematic noise artifacts are shared by both samples, whereas somatic variants are present only in the tumor sample. Only somatic variants are reported. To detect systematic noise artifacts, DRAGEN recommends that the coverage in the normal sample be at least half of the coverage in the tumor sample.

The tumor-only pipeline produces output that contains both germline and somatic variants and can be further analyzed to identify tumor mutations. The caller does not attempt to distinguish between them: filtering out common germline variants as reported in databases is currently the most reliable way to remove germline variants. The tumor-only pipeline provides a germline tagging feature and requires this feature to be explicitly enabled or disabled. When germline tagging is enabled, a variant annotation data directory must be passed in via the command line; DRAGEN will then tag variants that are common in the gnomAD, 1000 Genomes, or ABraOM databases as germline so they can be filtered out if desired (see details below in [Germline Tagging in the Tumor-Only Pipeline](#germline-tagging-in-the-tumor-only-pipeline)). The tumor-only pipeline also requires the presence of a systematic noise file by default. To run without germline tagging and/or systematic noise files, these options need to be disabled explicitly.

#### Variant Scoring

DRAGEN uses a Bayesian approach to compute the posterior probability that a somatic variant is present and reports this as a phred-scale quantity, "somatic quality" (SQ):

`##FORMAT=<ID=SQ,Number=1,Type=Float,Description="Somatic quality">`

DRAGEN scores variants by computing likelihoods for several hypotheses and noise processes, taking into account many factors such as: the numbers of alt-supporting and ref-supporting reads in the tumor and normal samples (and hence the alt allele frequencies in both samples); mapping qualities and how these are distributed across the reads in the tumor and normal pileups; basecall qualities; forward vs reverse strand support; sample-wide estimates of insertion and deletion error probabilities as functions of repeat period, repeat length, and indel length; sample-wide estimates of nucleotide error biases; whether there are nearby co-phased events; and whether the positions and alleles in question are known somatic hotspots or associated with sequence-specific error patterns. You can use SQ as the primary metric to describe the confidence with which the caller made a somatic call. SQ is reported as a format field for the tumor sample (exception: for homozygous reference calls in gvcf mode it is instead a likelihood ratio, analogous to homref GQ as described in the germline section). Variants with SQ score below the SQ filter threshold are filtered out using the `weak_evidence` tag. To trade off sensitivity against specificity, adjust the SQ filter threshold. Lower thresholds produce a more sensitive caller and higher thresholds produce a more conservative caller. If performing tumor-normal analysis, the SQ field for the normal sample contains the Phred-scaled posterior probability that a putative call is a germline variant. The somatic caller does not test for diploid genotype candidates and does not output GQ or QUAL values.

If tumor SQ > `vc-sq-call-threshold` (default is 3 for tumor-normal and 0.1 for tumor-only), then FORMAT/GT is hard-coded to 0/1 for the tumor sample and 0/0 for the normal sample (if present), and the tumor-sample FORMAT/AF yields an estimate of the somatic variant allele frequency, which ranges anywhere within \[0,1].

* If the value for `vc-sq-filter-threshold` is lower than `vc-sq-call-threshold`, the filter threshold value is used instead of the call threshold value.
* If tumor SQ < `vc-sq-call-threshold`, the variant is not emitted in the VCF.
* If tumor SQ > `vc-sq-call-threshold` but tumor SQ < `vc-sq-filter-threshold`, the variant is emitted in the VCF, but FILTER=weak\_evidence.
* If tumor SQ > `vc-sq-call-threshold` and tumor SQ > `vc-sq-filter-threshold`, the variant is emitted in the VCF and FILTER=PASS (unless the variant is filtered by a different filter).
* The default vc-sq-filter-threshold is 17.5 for tumor-normal and 3.0 for tumor-only analysis.
* When the `vc-sq-filter-threshold` argument is specified, the threshold will be applied to all variant types. To set the SQ threshold separately for SNVs and indels, use `vc-sq-snv-filter-threshold` and `vc-sq-indel-filter-threshold`.

The following is an example somatic T/N VCF record. Tumor SQ > `vc-sq-call-threshold` but tumor SQ < `vc-sq-filter-threshold`, so the FILTER is marked as weak\_evidence.

```
chr2 593701 . G A . weak_evidence
DP=97;MQ=48.74;SQ=3.86;NLOD=9.83;FractionInformativeReads=1.000
GT:SQ:AF:F1R2:F2R1:DP:SB:MB 0/0:9.83:33,0:0.000:14,0:19,0:33
0/1:3.86:61,3:0.047:29,2:32,1:64:35,26,0,3:39,22,1,2
```

The clustered-events penalty is an exception to the above rule for emitting variants. By default, the clustered-events penalty replaces the (obsolete) clustered-events filter. Instead of applying a hard filter when too many events are clustered together, DRAGEN applies a penalty to the SQ scores of co-phased clustered events. Clustered events with weak evidence are no longer called, but clustered events with strong evidence can still be called. This is equivalent to lowering the prior probability of observing clustered co-phased variants. The penalty is applied after the decision to emit variants, so that penalized variants still appear in the VCF if their unpenalized score is high enough. Variants that are combined into an MNV via the `--combine-phased-variants-distance` option are treated as a single variant for the purposes of the penalty. The penalty will not be applied to somatic hotspot variants. To disable the clustered-events penalty, set `--vc-clustered-event-penalty=0`.

#### Somatic Mode Options

To run DRAGEN somatic small variant calling, enable the variant caller with `--enable-variant-caller=true` and pass in tumor, and optionally, matched normal inputs via the command line. FASTQ (both gzipped and Ora-compressed), FASTQ list, BAM and CRAM inputs are all supported input types. For all input types, reads will be aligned by the DRAGEN map/align module and resulting alignments fed into the caller by default. For BAM and CRAM inputs, you can bypass map/align and use existing alignments as variant caller input by setting `--enable-map-align=false`.

Please see the DRAGEN Recipe sections for recommended command lines in typical workflows. The following command line options are typically used for somatic small-variant calling:

* `--tumor-fastq1 and --tumor-fastq2`

  Inputs a pair of FASTQ files into the mapper aligner and somatic variant caller. You can use these options with OTHER FASTQ options to run in tumor-normal mode. For example:

  ```
  dragen -f -r  /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
  --tumor-fastq1 <TUMOR_FASTQ1> \
  --tumor-fastq2 <TUMOR_FASTQ2> \
  --RGID-tumor <RG0-tumor> ---RGSM-tumor <SM0-tumor> \
  -1 <NORMAL_FASTQ1> \
  -2 <NORMAL_FASTQ2> \
  --RGID <RG0> --RGSM <SM0> \
  --enable-variant-caller true \
  --output-directory /staging/examples/ \
  --output-file-prefix SRA056922_30x_e10_50M 
  ```
* `--tumor-fastq-list`

Inputs a list of FASTQ files into the mapper aligner and somatic variant caller. You can use these options with other FASTQ options to run in tumor-normal mode. For example:

```
dragen -f \
-r /staging/human/reference/hg19/hg19.fa.k_21.f_16.m_149 \
--tumor-fastq-list <TUMOR_FASTQ_LIST> \
--fastq-list <NORMAL_FASTQ_LIST> \
--enable-variant-caller true \
--output-directory /staging/examples/ \
--output-file-prefix SRA056922_30x_e10_50M
```

* `--tumor-bam-input` and `--tumor-cram-input` Inputs a mapped BAM or CRAM file into the somatic variant caller. You can use these options with other BAM/CRAM options to run in tumor-normal mode. When the mapper is enabled (default), reads from the input BAM/CRAM files are re-mapped and updated alignments are sent to the caller (supported for both tumor-normal and tumor-only BAM/CRAM input). When the mapper is disabled (`--enable-map-align=false`), the existing BAM/CRAM alignments will be used in the caller.
* `--vc-sq-call-threshold`, `--vc-sq-filter-threshold`, `--vc-sq-snv-filter-threshold`, and `--vc-sq-indel-filter-threshold` These options control the thresholds for emitting calls in the VCF and applying the `weak_evidence` filter tag (see above).
* `--vc-target-vaf` This option allows the user to adjust the allele frequencies of haplotypes that will be considered by the caller as potentially appearing in the sample. It is not a hard threshold, but the variant caller will aim to detect variants with allele frequencies larger than this setting. In the case of tumor-normal runs, the frequency is measured with respect to the full set of reads (tumor and normal combined). The default threshold of 0.03 was selected to be as low as possible without incurring an excessive false positive cost; a lower setting may increase sensitivity for low-frequency variants, but may increase false positives and runtime; a higher setting may reduce false positives. Setting the vc-target-vaf to 0 will result in all haplotypes with at least two supporting reads being taken into consideration.
* `--vc-somatic-hotspots`, `--vc-use-somatic-hotspots`, and `--vc-hotspot-log10-prior-boost` DRAGEN uses a hotspot VCF to indicate somatic mutations that are expected with increased frequency. The default hotspot file (automatically selected from `<INSTALL_PATH>/resources/hotspots/somatic_hotspots_*` based on the reference) is mostly based on the Memorial Sloan Kettering Cancer Center (MSKCC) published hotspots and positions in COSMIC with population allele counts (AC) >= 50. It is somewhat conservative and boosts only a few thousand positions. You can specify a custom hotspot file via the `--vc-somatic-hotspots` option (note: input VCF records must be sorted in the same order as contigs in the selected reference) or disable the hotspots feature with `vc-use-somatic-hotspots=false`. The effect of the hotspot file is that the prior probability for hotspot variants is boosted by a factor, up to a maximum prior of 0.5. An SNV is considered to match a hotspot variant only if the allele in question is identical, whereas insertions or deletions are considered to match any insertion/deletion allele respectively. You can use `vc-hotspotlog10-prior-boost` to control the size of the adjustment. The default value is 4 (log10 scale) corresponding to an increase of 40 phred, and reducing this value will result in a smaller adjustment.
* `vc-systematic-noise` This option allows the user to specify the systematic noise file. To run without a systematic noise file (not recommended), specify `vc-systematic-noise=NONE`.
* `--vc-combine-phased-variants-distance` This option is the same as in the germline variant caller (see "Combine Phased Variants" in the germline small-variant caller section).
* `vc-skip-germline-tagging=true` This option disables the germline tagging feature in the tumor-only pipeline (not recommended).
* `--vc-callability-tumor-thresh` Specifies the callability threshold for tumor samples. The somatic callable regions report includes all regions with tumor coverage above the tumor threshold. The default value is 50. For more information on the somatic callable regions report, see [Somatic Callable Regions Report](/dragen-v4.5/product-guides/dragen-v4.5/qc-metrics-reporting#somatic-callable-regions-report).
* `--vc-callability-normal-thresh` Specifies the callability threshold for normal samples, if present. If applicable, the somatic callable regions report includes all regions with normal coverage above the normal threshold. The default value is 5. For more information on the somatic callable regions report, see [Somatic Callable Regions Report](/dragen-v4.5/product-guides/dragen-v4.5/qc-metrics-reporting#somatic-callable-regions-report).
* `--vc-excluded-regions-bed` Optional excluded regions BED file specifying where variants will be hard-filtered. Useful, e.g., to exclude ALU regions that tend to be especially noisy in FFPE samples.
* `--vc-call-hotspots-in-excluded-regions` Do not apply excluded regions filter to hotspot variants (Default=false).

**High-Specificity Mode**

The default configuration provides balanced sensitivity and specificity, and the sensitivity-specificity tradeoff can be adjusted via the SQ filter threshold as described above. However, for applications requiring especially high specificity, DRAGEN provides an umbrella option that enables maximum-specificity filtering by selecting a high SQ threshold while also enabling multiple extra filter options. This aggressive filtering improves precision at the cost of some sensitivity:

```
--vc-high-specificity=true
```

The high-specificity mode applies the following settings:

| Parameter                                    | Default               | High-Specificity Setting | Purpose                                                               |
| -------------------------------------------- | --------------------- | ------------------------ | --------------------------------------------------------------------- |
| `--vc-systematic-noise-filter-threshold`     | 10 (T/N), 60 (T/O)    | 60                       | More aggressive filtering of position-/allele-specific noise patterns |
| `--vc-output-variant-read-position`          | false                 | true                     | Output variant read position information in VCF                       |
| `--vc-enable-trimer-context`                 | false                 | true                     | Enable trimer context for error rate estimation                       |
| `--vc-enable-soft-clip-ratio-filter`         | false                 | true                     | Enable filtering based on soft-clipping patterns                      |
| `--vc-soft-clip-ratio-filter-threshold`      | 0.1                   | 0.05                     | More aggressive filtering of variants with high soft-clip ratios      |
| `--vc-override-tumor-pcr-params-with-normal` | true                  | false                    | Use separate PCR parameters for tumor (not normal)                    |
| `--vc-clustered-event-penalty`               | 4.0 (T/N), 7.0 (T/O)  | 7                        | Increase penalty applied to clustered co-phased variants              |
| `--vc-sq-filter-threshold`                   | 17.5 (T/N), 3.0 (T/O) | 23                       | Higher confidence threshold for PASS calls                            |
| `--vc-enable-af-filter`                      | false                 | true                     | Enable allele frequency filtering                                     |
| `--vc-af-filter-threshold`                   | 0.1                   | 0.05                     | More aggressive filtering of low-frequency variants                   |
| `--vc-use-somatic-hotspots`                  | true                  | false                    | Disable somatic hotspot boosting                                      |

#### Tumor-in-normal contamination and liquid tumor mode

In a tumor-normal analysis, DRAGEN accounts for tumor-in-normal (TiN) contamination by running liquid tumor mode. Liquid tumor mode is disabled by default, but we recommend enabling it with `--vc-enable-liquid-tumor-mode=true` if TiN contamination is expected. When liquid tumor mode is enabled, DRAGEN is able to call variants in the presence of TiN contamination up to a specified maximum tolerance level (default: 0.15). If using the default maximum contamination TiN tolerance, somatic variants are expected to be observed in the normal sample with allele frequencies up to 15% of the corresponding allele in the tumor sample. `vc-tin-contam-tolerance` enables liquid tumor mode and allows you to set the maximum contamination TiN tolerance.

Liquid tumor mode is not equivalent to liquid biopsy. Liquid tumors in liquid tumor mode refer to hematological cancers, such as leukemia. For liquid tumors, it is not feasible to use blood as a normal control because the tumor is present in the blood. Skin or saliva is typically used as the normal sample. However, skin and saliva samples can still contain blood cells, so that the matched normal control sample contains some traces of the tumor sample and somatic variants are observed at low frequencies in the normal sample. If the contamination is not accounted for, it can severely impact sensitivity by suppressing true somatic variants.

Liquid tumor mode typically uses a library that is WGS or WES with medium depth for example (100x T/ 40xN), and the lowest VAF detected for these types of depths is \~5%. Liquid biopsy typically uses a targeted gene panel (eg 500 genes), with very high raw depth, and uses UMI indexing (collapsing down to a depth of >2000x) to enable sensitivity at VAF down to 0.1 % in some cases (the limit of detection will vary depending on coverage and data quality).

#### Mixing tumor and normal samples from different sequencing protocols

If using different sequencing systems or different library preparation methods for tumor and normal samples, we recommend setting `--vc-override-tumor-pcr-params-with-normal=false`. In tumor-normal mode, DRAGEN estimates a set of PCR error parameters separately for each of the tumor and normal samples. By default, DRAGEN ignores the tumor-sample parameters and uses normal-sample parameters for analysis of both samples. This default prevents overestimation of tumor-sample error rates that can occur if the somatic variant rate is high.

**Allele frequency and related settings**

There is no hard limit on the allele frequencies at which DRAGEN can report calls, but there are a number of points in the pipeline where low allele frequency can affect calling. The `vc-target-vaf` setting affects the threshold used to detect candidate haplotypes during localized haplotype assembly, but does not affect variant scoring. Once a candidate haplotype is detected, all putative variants appearing in the haplotype are scored and calls scoring above the SQ call threshold are emitted regardless of the allele frequency or the number of supporting reads.

The probability calculation in the somatic caller assesses variant and noise hypotheses at fixed allele frequencies defined by a discrete grid (by default at coverages <200: 0, 0.05, 0.1, ... 1.0). This means that the calculation will assess variants with allele frequencies below 0.05 as if the true frequency is equal to 0.05; this strategy does not preclude such variants from being called but may result in lower scores compared to if the true frequency had been considered. At positions with higher coverage, DRAGEN adds extra grid points as in the table below in order to consider hypotheses involving lower allele frequencies and effectively achieve a lower limit of detection (LOD), with the lowest VAF halving every time the coverage doubles:

| Coverage | Lowest AF |
| -------- | --------- |
| 0-199    | 0.05      |
| 200-399  | 0.025     |
| 400-799  | 0.0125    |
| ...      | ...       |

If calls below a certain VAF are not of interest, you can use `--vc-enable-af-filter` (see Post Somatic Calling Filtering below) to apply a hard filter on VAF.

#### Sample-specific NTD Error Bias Estimation

DRAGEN can compensate for oxidation and deamination artifacts that might exist upstream of the sequencing system, and are common in FFPE samples. DRAGEN does this by estimating nucleotide mutation biases on a per sample basis, taking account of read orientation. During variant calling, DRAGEN then corrects for nucleotide substitution biases by combining the estimated parameters with the basecall quality scores, thus modifying the nucleotide error rates used by the hidden Markov model.

Nucleotide (NTD) Error Bias Estimation is on by default and recommended as a replacement for the orientation bias filter. Both methods take account of strand-specific biases (systematic differences between F1R2 and F2R1 reads). In addition, NTD error estimation accounts for non-strand-specific biases such as sample-wide elevation of a certain snv type, e.g. C->T or any other transition or transversion. This is done by collecting counts (sampled from across the genome, and counted per read orientation) of reads supporting each specific nucleotide subsitution (C->T, G->A, etc.). The estimated rate of each substitution is [written to a metrics file named "\*.allele-transition-noise-metrics.csv"](/dragen-v4.5/product-guides/dragen-v4.5/qc-metrics-reporting#somatic-allele-transition-noise-metrics-file). NTD error estimation can also capture these biases in a trinucleotide context, e.g. in the case of C->T it will break down the counts as ACA->ATA, CCA->CTA, GCA->CTA, TCA->TTA, etc.

This feature can be disabled by specifying `--vc-enable-unequal-ntd-errors=false` or set to auto-detect by specifying `--vc-enable-unequal-ntd-errors=auto`. In auto-detect mode, DRAGEN will run the estimation but then disable the use of the estimated parameters if it determines that the sample does not exhibit nucleotide error bias. When the feature is enabled, DRAGEN will by default estimate a smaller set of parameters in a monomer context. To estimate a larger set of parameters in a trimer context (recommended on sufficiently large panels when coverage is above 1000X), specify `--vc-enable-trimer-context=true`.

To specify the regions from which to estimate nucleotide substitution biases, use `--vc-snp-error-cal-bed`. Alternatively, if `--vc-target-bed` is used to specify the target regions for variant calling, and the total bed regions are sufficiently small (maximum 4 megabases), `--vc-snp-error-cal-bed` can be omitted and DRAGEN will use the target bed file for bias estimation. Otherwise, DRAGEN will use a default bed file selected to match the reference, and covering a mixture of coding and non-coding regions.

DRAGEN requires a panel size of at least 150kbp to correctly estimate nucleotide mutation biases when using trimer context, or at least 10kbp when using monomer context. If this requirement is not met for trimer context, DRAGEN falls back on the monomer model, and if it is not met for monomer context, DRAGEN turns the bias estimation feature off.

#### Unique Molecular Identifier (UMI) Support in Small Panels

DRAGEN includes two ultra-sensitive, UMI-aware variant calling pipelines designed for UMI-collapsed reads. These pipelines leverage the improved read and base quality typical of simplex- and duplex-collapsed data.

Both modes are disabled by default. They are not recommended for WGS and generally not for WES, as the increased sensitivity substantially extends runtime.

The pipelines assume input from FASTQs with UMIs (`--enable-umi true`) or from pre-collapsed BAMs.

Enable them with:

* `--vc-enable-umi-solid` Optimized for solid tumors with post-collapsed coverage of \~200–1000X and allele frequencies ≥5%.
* `--vc-enable-umi-liquid` Optimized for liquid biopsy, detecting low-VAF somatic variants in cell-free DNA from plasma. Typically assumes >1000X post-collapsed coverage and allele frequencies ≥0.1%. (Note: distinct from liquid tumor mode.)

If your UMI-collapsed reads do not meet the recommended post-collapsed coverage depths for the options listed above, we recommend you run with default settings.

If a third-party tool is used to produce the collapsed reads, then configure the tool so that the base call quality scores quantify the error produced by the sequencing system only. DRAGEN uses Sample-specific NTD Error Bias Estimation (see above) to account for errors upstream of the sequencing system, so such errors should not be included in base call quality scores.

#### gVCF Output

You can output a gVCF file for tumor-only data sets. A gVCF file reports information on every position of the input genome, including homozygous reference (homref) positions, i.e. positions where no alt allele (either germline or somatic) is present. DRAGEN creates a new \<NON\_REF> allele, to which reads that do not support the reference allele or any reported variant allele are assigned. In tumors, variants could exist at arbitrarily low allele frequencies and be undetectable. Thus, a somatic homref call cannot guarantee that no somatic variant at any allele frequency exists at the position. Instead, DRAGEN considers a position to be a homozygous reference if there are no somatic variants with an allele frequency at or above the limit of detection (LOD). Whereas the SQ score for an ordinary alt allele is a phred-scale posterior probability, the SQ score for the \<NON\_REF> allele is a phred-scale ratio between the likelihood of a homref call and the likelihood of a variant call with allele frequency at the LOD (if an alt allele is also reported, the \<NON\_REF> SQ score is capped at the complement of the posterior probability for the alt allele). If the LOD value is lowered, fewer homref calls are made. If the LOD value is increased, more homref calls are made.

By default the LOD is set to 5%, but you can enter a different value using the `--vc-gvcf-homref-lod` option.

#### Analysis of Low-Quality FFPE Samples

An increased number of FP variants can be observed in FFPE samples that are highly degraded. For samples with a DNA integrity number (DIN) less than 2, it is recommended to enable the low-quality FFPE mode with `--vc-enable-low-qual-ffpe-mode=true`. This option is designed to reduce the number of FPs in two ways:

1. Reducing the spurious signal created by PCR duplicates. This is accomplished by collapsing reads based on alignment position and always requiring at least 3 supporting reads to call a variant.
2. Filtering out variant calls where the alt allele support is low quality:
   1. Alt-supporting read positions are skewed to the ends of reads compared to ref-supporting positions: the distance of the event position to the nearest end of the alignment (excluding soft clips) is collected for each read. A score is calculated as a phred-scaled P value from a one-sided Mann Whitney test, and the difference between the alt and ref median values is also calculated. If the score exceeds 25 and the median difference exceeds 2, the call is filtered with "alt\_closer\_to\_end".
   2. Alt-supporting reads have increased mutations compared to ref-supporting reads: the edit distances are calculated for each read. In alt-supporting reads, the edit distance is adjusted to exclude the variant being considered. A score is calculated as a phred-scaled P value from a one-sided Mann Whitney test, and the difference between the alt and ref mean values is also calculated. If the score exceeds 25 and the mean difference exceeds 1.0, the call is filtered with "alt\_higher\_edit\_distance".

#### Post Somatic Calling Filtering

DRAGEN can add a number of filters by populating the FILTER column in the vcf. The output is provided in the `<output-file-prefix>.hard-filtered.vcf.gz` output file. When machine learning is enabled, calls are filtered solely on the ML-recalibrated SQ scores and no other hard filtering is applied.

**Options**

The following options are available for post somatic calling filtering:

* `--vc-sq-call-threshold`

  Governs which calls are emitted in the VCF; only calls with SQ greater than vc-sq-call-threshold will be emitted. The default is 3.0 for tumor-normal and 0.1 for tumor-only. If the value for `vc-sq-filter-threshold` is lower than `vc-sq-call-threshold`, the filter threshold value is used instead of the call threshold value.
* `--vc-sq-filter-threshold`

  Sets the SQ filter threshold for both SNVs and indels. Calls with SQ below threshold are filtered with the `weak_evidence` flag. The default is 17.5 for tumor-normal and 3.0 for tumor-only. Incompatible with `--vc-sq-snv-filter-threshold` and `--vc-sq-indel-filter-threshold` options.
* `--vc-sq-snv-filter-threshold`

  Sets the SQ filter threshold independently for SNVs only. Incompatible with `--vc-sq-filter-threshold` option.
* `--vc-sq-indel-filter-threshold`

  Sets the SQ filter threshold independently for indels only. Incompatible with `--vc-sq-filter-threshold` option.
* `--vc-enable-triallelic-filter`

  Enables the multiallelic filter. The default is true. This filter will not be applied to somatic hotspot variants.
* `--vc-enable-non-primary-allelic-filter`

  Similar to the triallelic filter, but filters less aggressively. Keep the allele per multiallelic position with highest alt AD, and only filter the rest (Default=false). This filter will not be applied to somatic hotspot variants. Cannot be enabled when the triallelic filter is also on.
* `--vc-enable-af-filter`

  Enables the allele frequency filter for nuclear chromosomes. The default value is false. When set to true, the VCF excludes variants with allele frequencies below the AF call threshold or variants with an allele frequency below the AF filter threshold and tagged with low AF filter tag. The default AF call threshold is 1% and the default AF filter threshold is 5%. To change the threshold values, use the `vc-af-call-threshold` and `vc-af-filter-threshold` command-line options. Please use `vc-enable-af-filter-mito` and corresponding threshold options for mitochondrial allele frequency filtering.
* `--vc-enable-non-homref-normal-filter`

  Enables the non-homref normal filter. The default value is true. When set to true, the VCF filters out variants if the normal sample genotype is not a homozygous reference.
* `--vc-enable-vaf-ratio-filter`

  Adds one condition to be filtered out by the alt\_allele\_in\_normal filter. The default value is false. When set to true, the VCF filters out variants if the normal sample AF is greater than 20% of tumor sample AF.
* `--vc-depth-filter-threshold`

  Filters all somatic variants (alt or homref) with a depth below this threshold. The default value is 0 (no filtering).
* `vc-homref-depth-filter-threshold`

  In gvcf mode, filters all somatic homref variants with a depth below this threshold. The default value is 3.
* `vc-depth-annotation-threshold`

  Filters all non-PASS somatic alt variants with a depth below this threshold. The default value is 0 (no filtering).
* `vc-min-supporting-read-count` Filters all variant calls with fewer ALT-supporting reads than the threshold value. (Default = 3)
* `vc-enable-soft-clip-ratio-filter` Enables filtering on ratio of soft-clipped bases supporting variant calls. (Default = true)
* `vc-soft-clip-ratio-filter-threshold` Filter threshold for ratio of soft-clipped bases supporting a variant; all variants with INFO:softCLipRatio will be filtered. (Default=0.8)
* `vc-soft-clip-ratio-filter-vaf-cutoff` Limits application of soft\_clip\_ratio filter to only variants with VAF below cutoff. (Default=0.2)
* `vc-enable-indel-repeat-length-filter` Enables filtering of indels called in long homopolymers or STRs. (Default=true)
* `vc-indel-repeat-length-filter-threshold` Filter indel calls if they are in a homopolymer or STR with repeat length greater than this threshold. (Default=25)
* `vc-max-indel-length` Maximum allowed indel length to be called by somatic small VC. Set < 0 to disable this filter. (Default=50)
* `vc-enable-non-homref-normal-filter` Enables the non-homref-normal filter for somatic mode, filtering all calls where the normal sample genotype is not homozygous reference. Only applies to T/N mode. (Default=true)

**Filters**

| Somatic Mode                | Filter ID                      | Description                                                                                                                                                                                                                                                                                                                                                     |
| --------------------------- | ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Tumor-Only & Tumor-Normal   | weak\_evidence                 | Variant does not meet likelihood threshold. The likelihood ratio for SQ tumor-normal is < 17.5 or < 3.0 for SQ tumor-only.                                                                                                                                                                                                                                      |
| Tumor-Only & Tumor-Normal   | multiallelic                   | Site filtered if there are two or more ALT alleles at this location in the tumor. Not applied to somatic hotspot variants.                                                                                                                                                                                                                                      |
| Tumor-Only & Tumor-Normal   | base\_quality                  | Median base quality of ALT reads at this locus is < 20.                                                                                                                                                                                                                                                                                                         |
| Tumor-Only & Tumor-Normal   | mapping\_quality               | Median mapping quality of ALT reads at this locus is < 20 (tumor-normal) or < 30 (tumor-only).                                                                                                                                                                                                                                                                  |
| Tumor-Only & Tumor-Normal   | fragment\_length               | Absolute difference between the median fragment length of alt reads and median fragment length of ref reads at a given locus > 10000.                                                                                                                                                                                                                           |
| Tumor-Only & Tumor-Normal   | read\_position                 | Median of distances between the start and end of read and a given locus < 5 (the variant is too close to edge of all the reads). To output variant read position to the INFO field, use `--vc-output-variant-read-position=true`.                                                                                                                               |
| Tumor-Only & Tumor-Normal   | low\_af                        | Allele frequency is below the threshold specified with `--vc-af-filter-threshold` (default is 5%). Enabled only when using `--vc-enable-af-filter=true`.                                                                                                                                                                                                        |
| Tumor-Only & Tumor-Normal   | systematic\_noise              | If AQ score is < 10 (default) for tumor-normal or < 60 (default) for tumor-only, the site is filtered.                                                                                                                                                                                                                                                          |
| Tumor-Only & Tumor-Normal   | low\_frac\_info\_reads         | The fraction of informative reads (denominator excludes filtered\_out reads) is below the threshold. The default threshold value is 0.5. This filter is bypassed for long indels where ratio of indel length / estimated read length exceeds 0.4.                                                                                                               |
| Tumor-Only & Tumor-Normal   | filtered\_reads                | More than 50% of reads have been filtered out.                                                                                                                                                                                                                                                                                                                  |
| Tumor-Only & Tumor-Normal   | long\_indel                    | Indel length is greater than value of vc-max-indel-length (default = 50bp).                                                                                                                                                                                                                                                                                     |
| Tumor-Only & Tumor-Normal   | low\_depth                     | The site was filtered because the number of reads is too low. The filter is off by default.                                                                                                                                                                                                                                                                     |
| Tumor-Only & Tumor-Normal   | low\_tlen                      | The site was filtered because the fraction of low TLEN ALT supporting reads is above a threshold. The default threshold is 0.4. Reads with TLEN smaller than -2.25 (default) standard deviations from the mean are considered to be low TLEN. This filter is not applied for reads sampled from tight insert distributions i.e., stddev / mean < 0.1 (default). |
| Tumor-Only and Tumor-Normal | no\_reliable\_supporting\_read | No reliable supporting read was found in the tumor sample. A reliable supporting read is a read supporting the alt allele with mapping quality ≥ 40, fragment length ≤ 10,000, base call quality ≥ 25, and distance from start/end of read ≥ 5.                                                                                                                 |
| Tumor-Only & Tumor-Normal   | too\_few\_supporting\_reads    | Variant is supported by < vc-min-supporting-read-count (default = 3) reads in the tumor sample. This filter is not applied in UMI-aware pipelines.                                                                                                                                                                                                              |
| Tumor-Only & Tumor-Normal   | soft\_clip\_ratio              | Variant filtered due to ratio of soft-clipped bases supporting the variant greater than vc-soft-clip-ratio-filter-threshold (default = 0.8) and VAF less than vc-soft-clip-ratio-filter-vaf-cutoff (default = 0.2).                                                                                                                                             |
| Tumor-Only & Tumor-Normal   | long\_repeat                   | Variant call filtered because it is in a homopolymer or STR with repeat length greater than vc-indel-repeat-length-filter-threshold (default=25). Only applied to indels.                                                                                                                                                                                       |
| Tumor-Normal                | noisy\_normal                  | More than three alleles are observed in the normal sample at allele frequency above 9.9%.                                                                                                                                                                                                                                                                       |
| Tumor-Normal                | alt\_allele\_in\_normal        | ALT allele frequency in the normal sample is above 0.2 plus the maximum contamination tolerance. For solid tumor mode, the value is 0. For liquid tumor mode, the default value is 0.15. See `vc-enable-vaf-ratio-filter` for optional conditions.                                                                                                              |
| Tumor-Normal                | non\_homref\_normal            | Normal sample genotype is not a homozygous reference.                                                                                                                                                                                                                                                                                                           |

#### Systematic Noise Filtering

The DRAGEN systematic noise filter significantly improves somatic variant calling precision, especially in tumor-only mode. DRAGEN enforces its use in the tumor-only pipeline by refusing to start a run without a noise file (this option can explicitly be disabled). This filter removes noise that consistently appears at specific locations in the reference genome. This noise can arise from:

* Mis-mapping in low-complexity regions: Repetitive sequences with low information content can lead to reads mapping to incorrect locations.
* PCR noise in homopolymer regions: Regions with long stretches of the same nucleotide (e.g., AAAAA) can introduce errors during PCR amplification.

To determine whether a variant should be filtered, the systematic noise filter compares the observed variant's allele frequency (AF) to the noise level at the matching locus in the systematic noise file. Variants are filtered if their AF is not statistically sufficiently higher than the recorded noise.

Note that the systematic noise filter specifically aims to remove noise, not germline variants; however, it may inadvertently filter some germline variants. For this reason, it is not ideal to evaluate the systematic noise file on germline admixture datasets.

Newer versions of the systematic noise filter will include allele-specific information along with two columns for noise frequency: one for the "mean" noise and one for the "max" noise. During a VC run, DRAGEN will automatically detect the input sample type as either WGS or WES/panel and will apply the optimal noise values based on sample type and run context. For WGS data, the "max" noise is used by default; for WES/panel data or whenever UMI is enabled, the "mean" noise is used.

WES and WGS prebuilt systematic noise files are available for download (see below).

Custom panels will require custom noise files. It is recommended to use normal samples sequenced on the same instrument type and using the same library prep. Building your own noise file is especially helpful for clean UMI samples that tend to have less noise than WGS/WES samples. To generate a noise file it is recommended to use approximately 30-70 normal samples, although fewer normal samples (1-10) can still be used to generate useful noise files.

The systematic noise filter is used in the DRAGEN tumor-only or tumor-normal pipeline by adding `--vc-systematic-noise NOISE_FILE_PATH`.

| Option                                            | Description                                                                                                                                                                                                                                                                                                 |
| ------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --vc-systematic-noise                             | Specifies a systematic noise BED file. If a somatic variant does not pass the AQ threshold, the variant is marked as 'systematic\_noise' in the FILTER column of the output VCF.                                                                                                                            |
| --vc-systematic-noise-filter-threshold            | Set the AQ threshold. Higher values filter more aggressively. By default the threshold value is 10 for tumor-normal and 60 for tumor-only. The valid range spans 0-100. For tumor-normal runs the threshold may be set higher (e.g. to 60) to improve specificity at the possible cost of some sensitivity. |
| --vc-systematic-noise-filter-threshold-in-hotspot | Set the AQ threshold to use in hotspot regions, where one may want to filter less aggressively than in the rest of the genome. By default, the threshold value is 10 for tumor-normal and 20 for tumor-only.                                                                                                |
| --vc-allele-specific-systematic-noise             | Apply systematic noise in an allele-specific manner when allele information is available. This setting is ignored for v1.x.x noise files (Default=true))                                                                                                                                                    |

#### Prebuilt Systematic Noise BED Files

Prebuilt systematic noise files can be downloaded here: [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

Somatic Systematic Noise Baseline Collection v2.0.0 noise files include allele specific information to better preserve sensitivity with systematic noise filtering enabled. Each v2.0.0 noise file includes both "mean" and "max" noise in separate columns, with the appropriate noise applied automatically based on auto-detected input type and run context.

The latest noise files (v2.0.0) contain more columns than earlier noise files and are therefore incompatible with versions of DRAGEN prior to v4.3. Older noise files are still supported in the current version of DRAGEN; however, the older noise files lack allele specific information and noise filtering will be applied by position only as was the default in v4.2 and earlier versions of DRAGEN.

The default WES and WGS noise files were generated using a combination of Nextera and TruSeq samples (with and without PCR). There are also hg38 WGS HEME and FFPE specific noise files. For details please refer to [SNV Systematic Noise Files](https://help.dragen.illumina.com/reference/resource-files/prebuilt-baseline-files).

### Custom Systematic Noise Files

The BaseSpace Sequence Hub DRAGEN Baseline Builder App or the DRAGEN Systematic Noise File Builder Pipeline on ICA can be used to build systematic noise files in the cloud.

For example command lines on how to build a custom noise file, please refer to the respective DRAGEN recipes: [DRAGEN Recipes](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes).

| Option                                   | Description                                                                                                                                                                                                                                                                    |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| --build-sys-noise-vcfs-list              | Text file containing the paths of normal VCFs. Specify the full VCF file paths. List one file per line.                                                                                                                                                                        |
| --build-sys-noise-germline-vaf-threshold | Variant calls with VAF higher than this threshold will be considered germline and will not contribute to the noise estimate. This option is disabled by default by setting the threshold to 1. (Default 1)                                                                     |
| --build-sys-noise-use-germline-tag       | This option will ensure that variants tagged by `vc-enable-germline-tagging=true` will not be counted as noise. (Default true)                                                                                                                                                 |
| --build-sys-noise-min-sample-cov         | Min coverage at a site for a sample to be used towards noise estimation. At low coverages estimated allele frequencies become less reliable. Accurate AF estimation is imporant for germline variant detection, and also for noise detection when using MAX noise. (Default 5) |
| --build-sys-noise-min-supporting-samples | Min number of samples with noise at a position in order for a position to be considered systematic-noise (Default 1).                                                                                                                                                          |

### Germline Tagging in the Tumor-Only Pipeline

When enabling DRAGEN for tumor-only somatic calling, potential germline variants can be tagged in the INFO field with 'GermlineStatus' using population databases. Current databases include 1KG, both exome and genome sequencing data from gnomAD. The following options are available for this feature:

* `--vc-enable-germline-tagging` Enable germline tagging. The default is 'false'. In a tumor-only analysis, this option must either be set 'true' (recommended) or germline tagging must be explicitly disabled with `--vc-skip-germline-tagging=true` (not recommended). Once the `vc-enable-germline-tagging` option is set to 'true', it will require the user to pass in a variant annotation data directory as follows:
  * `--variant-annotation-data` Nirvana annotation database (Downloadable at <https://support.illumina.com/content/dam/illumina-support/help/Illumina\\_DRAGEN\\_Bio\\_IT\\_Platform\\_v3\\_7\\_1000000141465/Content/SW/Informatics/Dragen/Nirvana\\_DownloadData\\_fDG.htm>)

Additional options to control how to define germline variants.

* `--germline-tagging-abraom-threshold` The minimum alternative allele count for the ABraOM database for a variant to be defined as germline (default=50).
* `--germline-tagging-db-threshold` The minimum alternative allele count across population databases for a variant to be defined as germline (default=50).
* `--germline-tagging-pop-af-threshold` The minimum population allele frequency for a variant to be defined as germline. Once specified, this will override the input from --germline-tagging-db-threshold.

```
1    11301714        .       A       G       .       PASS    
DP=3626;MQ=249.61;FractionInformativeReads=0.974;AQ=100.00;GermlineStatus=Germline_DB   
GT:SQ:AD:AF:F1R2:F2R1:DP:SB:MB  0/1:64.73:1772,1758:0.498:872,901:900,857:3530:846,926,843,915:894,878,874,884
```

The total number of variants tagged as germline and somatic in the VCF are [written to a metrics file named "\*.vc\_germline\_tagging\_metrics.csv"](/dragen-v4.5/product-guides/dragen-v4.5/qc-metrics-reporting#germline-tagging-counts-in-tumor-only-pipeline).


# Somatic ML for Small VC (Beta)

DRAGEN Somatic Tumor-Normal small variant calling has an optional workflow which employs machine learning-based variant recalibration. Variant calling accuracy is improved using powerful yet efficient machine learning techniques that augment the variant caller. A supervised machine learning method was developed to build a model that processes read and other contextual evidence to remove false positives and recover false negatives, for both SNVs and INDELs.

DRAGEN Somatic ML can be applied to WGS or WES samples. It also supports FF and FFPE sample types. **DRAGEN Somatic ML should not be used when running with mutational signatures analysis or with HRD or TMB biomarkers enabled in the DRAGEN run. See** [**Somatic ML limitations below**](#limitations)**.**

## Setup

DRAGEN Somatic ML is enabled using `--vc-ml-enable-recalibration true`. DRAGEN Somatic ML runs concurrently with DRAGEN Somatic SNV VC.

## Inputs

DRAGEN Somatic ML requires a run with BAM, CRAM or FASTQ input, since the machine learning model extracts information from the read pile-up. Recalibration of existing VCF files is not supported.

It is recommended that a pre-built systematic noise file (v2) should be supplied to run Somatic ML for optimal performance. They are available at [DRAGEN Software Support Sitepage](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

## Outputs

DRAGEN Somatic ML recalibrates all quality scores, changing the value of the SQ field in the output VCF/GVCF.

DRAGEN Somatic ML PHRED scores (SQ) are better calibrated than and differ significantly from those with ML disabled and, as a consequence, SQ scores are lower. For this reason, the default SQ filtering thresholds are much lower when DRAGEN Somatic ML is enabled (5 and 3 for SNVs and INDELs respectively, compared to 17.5 for both SNVs and INDELs when DRAGEN Somatic ML is disabled).

The following variant types are recalibrated:

* Autosomes and sex chromosomes
* ForceGT calls
* Non-primary contigs

The output SQ of DRAGEN Somatic ML is empirically more accurately calibrated than DRAGEN Somatic SNV VC without ML. Note that a small number of variant calls may have degraded accuracy with ML enabled compared to VC without ML.

## Limitations

The following features should not be run together with DRAGEN Somatic ML:

* Mutational signatures analysis
* TMB biomarker
* HRD biomarker

These use cases are untested and not fully supported in v4.5, but will be supported in future DRAGEN releases.


# Pedigree Analysis

DRAGEN supports pedigree-based and population-based germline variant joint analysis for multiple samples. A pedigree-based analysis deals with samples from the same species which are related to each other. A population-based analysis compares samples of the same species which are unrelated to each other. You can find more information about the population-based analysis in the [iterative gVCF Genotyper](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/iterative-gvcf-genotyper) section.

Joint analysis requires a gVCF file for each sample. To create a gVCF file, run the germline small variant caller with the `--vc-emit-ref-confidence gVCF` option.

The gVCF file contains information on the variant positions and positions determined to be homozygous to the reference genome. For homozygous regions, the gVCF file includes statistics that indicate how well reads support the absence of variants or alternative alleles. Contiguous homozygous runs of bases with similar levels of confidence are grouped into blocks, referred to as hom-ref blocks. Not all entries in the gVCF are contiguous. A reference might contain gaps that are not covered by either variant line or a hom-ref block. Gaps correspond to regions that are not callable. A region is not callable if there is not at least one read mapped to the region with a MAPQ score above zero.

## Command-Line Options

### Pedigree Mode Options

The following parameters are available.

* `--enable-joint-genotyping`--To run the Joint Genotyper, set to true.
* `--pedigree-file`--Specify the path to a pedigree file that describes the relationship between samples. It is possible to run JointGenotyper without a pedigree file on unrelated samples on versions prior to DRAGEN v3.10. It is not recommended for gVCF variant calls on DRAGEN v3.10 or later.
* `--variant` or `--variant-list`--Specify the gVCF input to the workflow. The pedigree caller can read input gVCF files from an AWS S3 bucket, Azure storage BLOB, or pre-signed URL.
* `--enable-multi-sample-gvcf=true`--Write output as a multisample gVCF file that includes both variant records and hom-ref block information.
* `--output-directory`--The output directory. This is required.
* `--output-file-prefix`--The prefix used to label all output files. This is required.
* `-r`--The directory where the hash table resides.

The output of the joint genotyper depends on the order of input gVCF files passed on the command line using `--variant` or `--variant-list`. It is recommended to use the same input order when re-analyzing gVCFs to ensure the output is consistent with previous runs.

### Small Variant De Novo Calling Options

Use the common [Pedigree Mode Options](#pedigree-mode-options), plus the following options for *de novo* small variant calling.

* `--qc-snp-denovo-quality-threshold`--Specify the minimum DQ value for a SNP to be considered *de novo*. The default is 0.05 if ML recalibration is off, 0.0017 if ML recalibration is on.
* `--qc-indel-denovo-quality-threshold`--Specify the minimum DQ value for an indel to be considered *de novo*. The default is 0.4 if ML recalibration is off, 0.04 if ML recalibration is on.

## &#x20;Joint Analysis Output Format

There are two available joint analysis output files:

* **Multisample VCF**--A VCF file containing a column with genotype information for each of the input samples according to the input variants.
* **Multisample gVCF**--A gVCF file augmenting the content of a multisample VCF file, similar to how a gVCF file augments a VCF file for a single sample. In between variant sites, the multisample gVCF contains statistics that describe the level of confidence that each sample is homozygous to the reference genome. Multisample gVCF is a convenient format for combining the results from a pedigree or small cohort into a single file. If using a large number of samples, fluctuation in coverage or variation in any of the input samples creates a new hom-ref block, which causes a highly fragmented block structure and a large output file that can be slow to create.

The multisample gVCF output is only available in the pedigree-based analysis.

The following example shows a single line from a multi-sample VCF where one sample has a variant, and the other two samples are in a gVCF gap. Gaps are represented by "./.:.:".

```
1 605262 . G A 13.41 DRAGENHardQUAL
AC=2;AF=1.000;AN=2;DP=2;FS=0.000;MQ=14.00;QD=6.70;SOR=0.693
GT:AD:AF:DP:GQ:FT:F1R2:F2R1:PL:GP ./.:.:.:.:.:LowDepth
1/1:0,2:1.000:2:4:PASS:0,0:0,2:50,6,0:1.383e+01,4.943e+00,1.951e+00
./.:.:.:.:.:LowDepth
```

### &#x20;Hom-ref Blocks FORMAT Fields

In hom-ref blocks, the following FORMAT fields are calculated uniquely.

* **FORMAT/DP**--In a single sample gVCF, the FORMAT/DP reported at a hom-ref position is the median DP in that band. In a multisample gVCF, the FORMAT/DP reported at a hom-ref position is the MIN\_DP from hom-ref calls.
* **FORMAT/AD**--In single sample gVCF, values represent the position in the band where DP=median DP. In the multisample gVCF, AD values at hom-ref positions are copied from the single sample gVCF.
* **FORMAT/AF**--Values are based on FORMAT/AD.
* **FORMAT/PL**--Values represent the Phred likelihoods per genotype hypothesis. For hom-ref blocks, each value in FORMAT/PL represents the minimum value across all positions within the band.
* **FORMAT/SPL** and **FORMAT/ICNT**--Parameters reported in the gVCF records, including both hom-ref blocks and variant records. The parameters are used to compute the confidence score of a variant being de novo in the proband of a trio. For SNP, FORMAT/PL and FORMAT/SPL are both used as input to the De Novo Caller. FORMAT/PL represents Phred likelihoods obtained from the genotyper, if the genotyper is called. FORMAT/SPL represents Phred likelihoods obtained from column-wise estimation, pregraph. Each value in FORMAT/SPL represents the minimum across all positions within the band. For INDEL, the PL value is computed in the joint pedigree calling step based on the FORMAT/ICNT reported in the gVCF file. FORMAT/ICNT consists of two values. The first value is the number of reads with no indels at the position, and the second value is the number of reads with indels at the position. Each value in FORMAT/ICNT represents the maximum of the value across all positions within the band.

In the following example hom-ref block, ICNT provides information on whether each sample contains an Indel at the position of interest. If the proband contains an indel at the position and the ICNT of the parents does not indicate any read supporting an indel, then the confidence score is high for the proband to have an indel *de novo* call at the position.

```
chr1 10288 . C <NON_REF> . PASS END=10290
GT:AD:DP:GQ:MIN_DP:PL:SPL:ICNT
0/0:131,4:135:69:132:0,69,1035:0,125,255:23,1

chr1 10291 . C
T,<NON_REF> 38.45 PASS
DP=100;MQ=24.72;MQRankSum=0.733;ReadPosRankSum=4.112;FractionInformativeReads=0.600;R2_5P_bias=0.000
GT:AD:AF:DP:F1R2:F2R1:GQ:PL:SPL:ICNT:GP:PRI:SB:MB
0/1:28,32,0:0.533,0.000:60:20,21,0:8,11,0:15:73,0,12,307,157,464:255,0,255:23,10:3.8452e+01,1.3151e-01,1.5275e+01,3.0757e+02,1.9173e+02,4.5000e+02:0.00,34.77,37.77,34.77,69.54,37.77:4,24,7,25:8,20,14,18
```

SPL and ICNT values are specific to DRAGEN. The GATK variant caller does not output SPL and ICNT values.

In a single sample gVCF, FORMAT/DP reported at a hom-ref position is the median DP in the band. The minimum is also computed and printed as MIN\_DP for the band.

In the multisample gVCF, MIN\_DP from hom-ref calls is printed as FORMAT/DP, and AD is just copied from the gVCF. Therefore, at a hom-ref position in the multi-sample gVCF output, the DP is not necessarily going to be the sum of AD.

## &#x20;Pedigree Mode

Use pedigree mode to jointly analyze samples from related individuals and to perform *de novo* calling.

To invoke pedigree mode, set the `--enable-joint-genotyping` option to true. Use the `--pedigree-file` option to specify the path to a pedigree file that describes the relationship between samples.

The pedigree file must be a tab-delimited text file with the file name ending in the .ped extension. The following information is required.

| Column Header  | Description                                                                                                             |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- |
| Family\_ID     | The pedigree identifier.                                                                                                |
| Individual\_ID | The ID of the individual.                                                                                               |
| Paternal\_ID   | The ID of the individual's father. If the founder, the value is 0.                                                      |
| Maternal\_ID   | The ID of the individual's mother. If the founder, the value is 0.                                                      |
| Sex            | The sex of the sample. If male, the value is 1. If female, the value is 2.                                              |
| Phenotype      | The genetic data of the sample. If unknown, the value is 0. If unaffected, the value is 1. If affected, the value is 2. |

The following is an example of an input pedigree file.

```
#Family_ID Individual_ID Paternal_ID Maternal_ID Sex Phenotype
FAM001 NA12877_Father 0 0 1 1
FAM001 NA12878_Mother 0 0 2 1
FAM001 NA12882_Proband NA12877_Father NA12878_Mother 2 2
FAM001 NA12883_Proband NA12877_Father NA12878_Mother 1 0
```

### De Novo Calling

The De Novo Caller identifies all the trios within the pedigree and generates a de novo score for each child. The De Novo Caller supports multiple trios within a single pedigree as long as they are all part of the same family. If multiple trios from different families need to be analyzed, each family will have to be processed in separate JointGenotyper runs. As the memory requirements and run time increases with each trio, it may be advantageous to split multiple trios within the same family into separate runs as well.

Pedigree Mode supports de novo calling for small, structural, and copy number variants.

Pedigree Mode is run in multiple steps. The following is an example workflow for a trio using FASTQ input.

1. Run single sample alignment and variant calling to generate per sample output using the following inputs for Pedigree Mode.
   * gVCF files for the Small Variant Caller.
   * \*.tn.tsv files for the Copy Number Caller.
   * BAM files for the Structural Variant Caller.
2. Run Pedigree Mode for Small Variant Caller. For more information, see [Small Variant De Novo Calling](#small-variant-de-novo-calling).
3. Run Pedigree Mode for Copy Number Caller. For more information, see [Multisample CNV Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline#multisample-germline-cnv-calling).
4. Run Pedigree Mode for Structural Variant Caller. For more information, see [Structural Variant De Novo Quality Scoring](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/sv-calling/sv-denovo-quality-scoring#structural-variant-de-novo-quality-scoring).
5. Run De Novo Small Variant Filtering. For more information, see [De Novo Small Variant Filtering](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/filtering#de-novo-small-variant-filtering).

### Small Variant De Novo Calling

The Small Variant De Novo Caller considers a trio of samples at a time. The samples are related via a pedigree file. The Small Variant De Novo Caller determines all positions that have a Mendelian conflict based on the genotype from the individual sample gVCFs. Sex chromosomes in males are treated as haploid apart from the PAR regions, which are treated as diploid.

Each of those positions is then processed through the Pedigree Caller to compute a joint posterior probability matrix for the possible genotypes. The probabilities are used to determine whether the proband has a *de novo* variant with a DQ confidence score. All three subjects are assumed to have an independent error probability.

At positions where the original genotype from the gVCFs shows a double Mendelian conflict (eg, 0/0+0/0->1/1 or 1/1+1/1->0/0), the genotypes of the trio samples can be adjusted to the highest joint posterior probability that has at least one Mendelian conflict.

The DQ formula is DQ = -10log10(1 - Pdenovo).

Pdenovo is the sum of all indexes in the joint posterior probability matrix with one or more Mendelian conflicts.

In the GT overwrite step, it is possible for the GT of the parents to be overwritten. In the case of multiple trios, the GT of the parents is based on the last trio processed. The trios are processed in the order they are listed in the pedigree file. DRAGEN currently does not add an annotation in the VCF in cases where the GT was overwritten.

The multisample VCF file is annotated with FORMAT/DQ and FORMAT/DN fields to output a VCF file that represents a *de novo* quality score and an associated *de novo* call. The DN field in the VCF is used to indicate the *de novo* status for each segment.

The following are the possible values:

* **Inherited**--The called trio genotype is consistent with Mendelian inheritance.
* **LowDQ**--The called trio genotype is inconsistent with Mendelian inheritance and DQ is less than the *de novo* quality threshold.
* **DeNovo**--The called trio genotype is inconsistent with Mendelian inheritance and DQ is greater than or equal to the *de novo* quality threshold.

The following is an example VCF line for a trio:

```
1	16355525	.	G	A	34.46	PASS	AC=1;AF=0.167;AN=6;DP=45;FS=6.69;MQ=108.04;MQRankSum=-0.156;QD=2.46;ReadPosRankSum=0;SOR=0.016	GT:AD:AF:DP:GQ:FT:F1R2:F2R1:PL:GP:PP:DPL:DN:DQ	0/1:11,3:0.214:14:39:PASS:8,2:3,1:74,0,47:39.454,0.00053613,49.99:0,1,104:74,0,47:DeNovo:0.67375	0/0:18,0:0:16:48:PASS:.:.:0,48,605:.:0,12,224:0,48,255:.:.	0/0:14,0:0:14:42:PASS:.:.:0,42,490:.:0,5,223:0,42,255:.:.
```


# De Novo Small Variant Filtering

The filtering step identifies de novo variants calls of the joint calling workflow in regions with ploidy changes. Since de novo calling can have reduced specificity in regions where at least one of the pedigree members shows non-diploid genotypes, the de novo variant filtering marks relevant variants and thus can improve specificity of the call set.

Based on the structural and copy number variant calls of the pedigree, the FORMAT/DN field in the proband is changed from the original DeNovo value to DeNovoSV or DeNovoCNV if the de novo variant overlaps with a ploidy-changing SV or CNV, respectively. All other variant details remain unchanged, and all variants of the input VCF will also be present in the filtered output VCF. Structural or copy number variants which result in no change of ploidy, such as inversions, are not considered in the filtering. As an example, a de novo SNV calls in the input VCF

```
chr1 234710899 . T C 44.74 PASS
AC=1;AF=0.167;AN=6;DP=73;FS=4.720;MQ=250.00;MQRankSum=5.310;QD=1.15;ReadPosRankSum=1.366;SOR=0.251
GT:AD:AF:DP:GQ:FT:F1R2:F2R1:PL:GL:GP:PP:DQ:DN
0/1:21,18:0.462:39:48:PASS:14,10:7,8:84,0,50:-8.427,0,-5:4.950e+01,7.041e-05,5.300e+01:15,0,120:3.2280e-01:DeNovo
0/0:13,0:0.000:11:30:PASS:.:.:0,30,450:.:.:10,0,227
0/0:25,0:0.000:22:60:PASS:.:.:0,60,899:.:.:0,33,227
```

Overlapping with an SV duplication in the proband, mother or father would be represented in the filtered output VCF as follows:

```
chr1 234710899 . T C 44.74 PASS
AC=1;AF=0.167;AN=6;DP=73;FS=4.720;MQ=250.00;MQRankSum=5.310;QD=1.15;ReadPosRankSum=1.366;SOR=0.251
GT:AD:AF:DP:GQ:FT:F1R2:F2R1:PL:GL:GP:PP:DQ:DN
0/1:21,18:0.462:39:48:PASS:14,10:7,8:84,0,50:-8.427,0,-5:4.950e+01,7.041e-05,5.300e+01:15,0,120:3.2280e-01:DeNovoSV
0/0:13,0:0.000:11:30:PASS:.:.:0,30,450:.:.:10,0,227
0/0:25,0:0.000:22:60:PASS:.:.:0,60,899:.:.:0,33,227
```

The following is an example command line for running the de novo filtering, based on the files returned by the joint calling workflows:

```
dragen \
--dn-enable-denovo-filtering true \
--dn-input-joint-vcf <JOINT_SMALL_VARIANT_VCF> \
--dn-output-joint-vcf <OUTPUT_VCF> \
--dn-sv-vcf <JOINT_SV_VCF> \
--dn-cnv-vcf <JOINT_CNV_VCF> \ 
--enable-map-align false
```

## De Novo Small Variant Filtering Options

The following options are used for de novo variant filtering:

* `--dn-input-vcf`---Joint small variant VCF from the de novo calling step to be filtered.
* `--dn-output-vcf`---File location to which the filtered VCF should be written. If not specified, the input VCF is overwritten.
* `--dn-sv-vcf`---Joint structural variant VCF from the SV calling step. If omitted, checks with overlapping structural variants are skipped.
* `--dn-cnv-vcf`--- Joint structural variant VCF from the CNV calling step. If omitted, checks with overlapping copy number variants are skipped.

## Germline Small Variant Hard Filtering

DRAGEN provides post-VCF variant filtering based on annotations present in the VCF records. Default and non-default variant hard filtering are described below. However, due to the nature of DRAGEN's algorithms, which incorporate the hypothesis of correlated errors from within the core of variant caller, the pipeline has improved capabilities in distinguishing the true variants from noise, and therefore the dependency on post-VCF filtering is substantially reduced. For this reason, the default post-VCF filtering in DRAGEN is very simple.

### Default Small Variant Hard Filtering

The default filters in the germline pipeline are as follows:

* \##FILTER=\<ID=DRAGENSnpHardQUAL,Description="Set if true:QUAL < 10.41 (3.0103 when ML recalibration is enabled)">
* \##FILTER=\<ID=DRAGENIndelHardQUAL,Description="Set if true:QUAL < 7.83 (3.0103 when ML recalibration is enabled)">
* \##FILTER=\<ID=MosaicHardQUAL,Description="Set if true:QUAL < 3.0103">
* \##FILTER=\<ID=MosaicLowAF,Description="Set if true:AF < 0.2 (0.1 if depth > 100x)">
* \##FILTER=\<ID=LowDepth,Description="Set if true:DP <= 1">
* \##FILTER=\<ID=PloidyConflict,Description="Genotype call from variant caller not consistent with chromosome ploidy">
* DRAGENSnpHardQUAL and DRAGENIndelHardQUAL: For all contigs other than the mitochondrial contig, the default hard filtering consists of thresholding the QUAL value only. A different default QUAL threshold value is applied to SNP and INDEL
* MosaicHardQUAL and MosaicLowAF: For all `MOSAIC` tagged variants, the default hard filtering consists of thresholding the QUAL value and the `FORMAT:AF` value.
* LowDepth: This filter is applied to all variants calls with INFO/DP <= 1
* PloidyConflict: This filter is applied to all variant calls on chrY of a female subject, if female is specified on the DRAGEN command line, of if female is detected by the ploidy estimator.

For the mitochondrial contig, DRAGEN processes it through a continuous AF pipeline, which is similar to the somatic variant calling pipeline. Please refer to [Mitochondrial Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling#mitochondrial-calling) for the filtering details.

### Non-Default Small Variant Hard Filtering

DRAGEN supports basic filtering of variant calls as described in the VCF standard. You can apply any number of filters with the `--vc-hard-filter` option, which takes a semicolon-delimited list of expressions, as follows:

```
<filter ID>:<snp|indel|all>:<list of criteria>,
```

where the list of criteria is itself a list of expressions, delimited by the || (OR) operator in this format:

```
<annotation ID> <comparison operator> <value>
```

The meaning of these expression elements is as follows:

* **filterID**---The name of the filter, which is entered in the FILTER column of the VCF file for calls that are filtered by that expression. The space character should be avoided.
* **snp/indel/all/homref/mosaic**---The subset of variant calls to which the expression should be applied.
* **annotation ID**---The variant call record annotation for which values should be checked for the filter. Supported annotations include FS, MQ, MQRankSum, QD, and ReadPosRankSum.
* **comparison operator**---The numeric comparison operator to use for comparing to the specified filter value. Supported operators include <, ≤, =, ≠, ≥, and >. For example, the following expression would mark with the label "SNPFilter" any SNPs with FS < 2.1 or with MQ < 100, and would mark with "indelFilter" any records with FS < 2.2 or with MQ < 110:

```
--vc-hard-filter="SNPFilter:snp:FS < 2.1 || MQ < 100; indelFilter:indel:FS < 2.2 || MQ < 110"
```

This example is for illustration purposes only and is NOT recommended for use with DRAGEN V3 output. Illumina recommends using the default hard filters. The only supported operation for combining value comparisons is OR, and there is no support for arithmetic combinations of multiple annotations. More complex expressions may be supported in the future.

The `--vc-hard-filter` value corresponding to the default hard filters are `'DRAGENSnpHardQUAL:snp: QUAL < 3.0103; DRAGENIndelHardQUAL:indel: QUAL < 3.0103; LowDepth:all: DP <= 1; MosaicHardQUAL:mosaic: QUAL < 3.0103; MosaicLowAF:mosaic: AF < 0.2’` if the sample's depth is less than or equals to 100x, otherwise they are `'DRAGENSnpHardQUAL:snp: QUAL < 3.0103; DRAGENIndelHardQUAL:indel: QUAL < 3.0103; LowDepth:all: DP <= 1; MosaicHardQUAL:mosaic: QUAL < 3.0103; MosaicLowAF:mosaic: AF < 0.1’`. We strongly suggest to not specify the `--vc-hard-filter` unless necessary.

## Orientation Bias Filter

The orientation bias filter is designed to reduce noise typically associated with the following:

* Pre-adapter artifacts introduced during genomic library preparation (eg, a combination of heat, shearing, and metal contaminates can result in the 8-oxoguanine base pairing with either cytosine or adenine, ultimately leading to G→T transversion mutations during PCR amplification), or
* FFPE (formalin-fixed paraffin-embedded) artifact. FFPE artifacts stem from formaldehyde deamination of cytosines, which results in C to T transition mutations. The orientation bias filter can only be used on somatic pipelines. To enable the filter, set the `--vc-enable-orientation-bias-filter` option to true. The default is false.

The artifact type to be filtered can be specified with the `--vc-orientation-bias-filter-artifacts` option. The default is C/T,G/T, which correspond to OxoG and FFPE artifacts. Valid values include C/T, or G/T, or C/T,G/T,C/A.

An artifact (or an artifact and its reverse compliment) cannot be listed twice. For example, C/T,G/A is not valid, because C→G and T→A are reverse compliments.

The orientation bias filter adds the following information:

* \##FORMAT=\<ID=F1R2,Number=R,Type=Integer,Description="Count of reads in F1R2 pair orientation supporting each allele">
* \##FORMAT=\<ID=F2R1,Number=R,Type=Integer,Description="Count of reads in F2R1 pair orientation supporting each allele">
* \##FORMAT=\<ID=OBC,Number=1,Type=String,Description="Orientation Bias Filter base context">
* \##FORMAT=\<ID=OBPa,Number=1,Type=String,Description="Orientation Bias prior for artifact">
* \##FORMAT=\<ID=OBParc,Number=1,Type=String,Description="Orientation Bias prior for reverse compliment artifact">
* \##FORMAT=\<ID=OBPsnp,Number=1,Type=String,Description="Orientation Bias prior for real variant">

Please note that the OBF filter runs as a standalone process after DRAGEN is complete. The VC metrics that are computed as part of DRAGEN SNV caller will not be updated and will not reflect the additional variants that are filtered in this stage.


# Autogenerated MD5SUM for VCF Files

An MD5SUM file is generated automatically for VCF output files. This file is in the same output directory and has the same name as the VCF output file, but with an .md5sum extension appended. For example, whole\_genome\_run\_123.vcf.md5sum. The MD5SUM files is a single-line text file that contains the md5sum of the VCF output file. This md5sum exactly matches the output of the Linux md5sum command.


# Force Genotyping

DRAGEN supports force genotyping (ForceGT) for small variant calling. Use `--vc-forcegt-vcf` to specify a VCF file containing variants to force genotype. The input list of small variants can be a \*.vcf or \*.vcf.gz file.

## Supported Modes

* **Germline**: Supported. When using joint genotyping with the `--vc-forcegt-vcf` option, the output joint VCF contains only variants tagged with `FGT`. Without this option, FGT-tagged variants are skipped.
* **Somatic**: Supported in both Tumor-Only (T/O) and Tumor-Normal (T/N) modes.

## Input Requirements

DRAGEN supports only a single ForceGT VCF input file. The input VCF must:

* Be a valid VCF 4.2 file (minimum 8 tab-delimited columns, sorted by contig and position).
* The header must list the same contig names as the reference used for variant calling. All variants must refer to one of these contig names.
* Contain normalized variants (parsimonious and left-aligned).
* Not contain multinucleotide or complex variants (e.g., `AT → C`). These are variants that require more than one substitution / insertion / deletion to go from REF allele to ALT allele and are ignored.
* Not contain deletions longer than 50bp — these are filtered out.
* Duplicate entries (same POS, REF, ALT) are ignored.

**Example of normalization:**

```
# Wrong (not parsimonious):
chrX  153592402  GC  GCG

# Correct (parsimonious):
chrX  153592403  C   CG
```

A nonnormalized variant will cause undefined behavour in DRAGEN.

## Output Behavior

The output VCF contains both regular variant calls and ForceGT variants. Each variant is tagged in the INFO field to indicate its origin:

| Scenario                                 | INFO Tag  |
| ---------------------------------------- | --------- |
| Regular call only (not in ForceGT input) | *(none)*  |
| ForceGT only (not called by pipeline)    | `FGT`     |
| Both regular and ForceGT (germline)      | `FGT;NML` |
| Both regular and ForceGT (somatic)       | `FGT;SOM` |

**Notes:**

* `NML` (normal): Indicates the variant was independently called by the pipeline in germline mode AND present in the ForceGT input.
* `SOM` (somatic): Indicates the variant was independently called by the pipeline in somatic mode AND present in the ForceGT input.
* `NML` and `SOM` **only** appear paired with `FGT`, never alone

**FILTER and INFO field behavior:**

* If a ForceGT variant matches a regular call with the same POS, REF, ALT, it inherits all FILTER and INFO fields from the regular call.
* If a ForceGT variant is at a novel site (no regular call), FILTER and INFO fields are calculated independently for that variant.

## Genotype Reporting

All variants in the ForceGT input VCF are genotyped and included in the output with the following GT values:

| Condition                            | Germline GT        | Somatic T/N GT | Somatic T/O GT |
| ------------------------------------ | ------------------ | -------------- | -------------- |
| No coverage at position              | `./.`              | `./.`          | `./.`          |
| Coverage but no ALT-supporting reads | `0/0`              | `0/0`          | `0/0`          |
| Coverage with ALT-supporting reads   | `0/1`, `1/1`, etc. | `0/1`          | `0/1` or `1/1` |

## ForceGT and Multiallelic Sites

In somatic mode, `--vc-split-multiallelic-calls` is enabled by default, which outputs multiallelic variants on separate lines. **It is not recommended to disable this option.**

ForceGT variants are combined into a single output line with regular calls only when they have an exact match (same POS, REF, and ALT). Otherwise, a separate ForceGT call is emitted.

**Example 1: ForceGT variant differs from regular call**

Both variants are output on separate lines:

```
chrX  100  .  G  C  .  PASS  .           ...  # Regular call (no tags)
chrX  100  .  G  A  .  PASS  FGT         ...  # ForceGT variant (different ALT)
```

**Example 2: ForceGT variant matches regular call exactly**

Combined into a single line with both tags:

```
# Germline mode:
chrX  100  .  G  A  .  PASS  FGT;NML     ...  # Called by pipeline AND in ForceGT input

# Somatic mode:
chrX  100  .  G  A  .  PASS  FGT;SOM     ...  # Called by pipeline AND in ForceGT input
```

**Example 3: Multiallelic site with partial ForceGT overlap**

If the pipeline calls a multiallelic site (e.g., G→A and G→T) and ForceGT input contains only G→A:

```
chrX  100  .  G  A  .  PASS  FGT;SOM     ...  # Matches ForceGT
chrX  100  .  G  T  .  PASS  .           ...  # Regular call only (no tags)
```

## Target BED Filtering

If a target BED file is provided via `--vc-target-bed`, only ForceGT variants overlapping the BED regions are included in the output.


# Machine Learning for Variant Calling

DRAGEN secondary analysis employs machine learning based variant recalibration (DRAGEN-ML) for germline SNV VC. Variant calling accuracy is improved using powerful yet efficient machine learning techniques that augment the variant caller, by exploiting more of the available read and context information that does not easily integrate into the Bayesian processing used by the haplotype variant caller. A supervised machine learning method was developed using truth from the PrecisionFDA v4.2.1 sets to build a model that processes read and other contextual evidence to remove false positives, recover false negatives and reduce zygosity errors, for both SNVs and INDELs.

## Setup

No additional setup is required. ML model files for human references are packaged with the DRAGEN installer. After installation, the files are present at `<INSTALL_PATH>/resources/ml_model`. DRAGEN-ML is enabled by default as needed, when running the germline SNV VC.

## Inputs

DRAGEN-ML requires a run with BAM or FASTQ input, since the machine learning model extracts information from the read pile-up. DRAGEN-ML runs concurrently with DRAGEN SNV VC. DRAGEN-ML can be applied to WGS or WES samples. Re-calibration of existing VCF files is not supported.

When DRAGEN-ML is enabled, the variant caller will use reads with MAPQ=0 for processing. All variant calls made in regions that have only MAPQ=0 reads (homologous regions) are identified by a `HML` flag in the VCF INFO field.

## Outputs

DRAGEN-ML recalibrates all quality scores, changing the values of the QUAL and GQ fields in the output VCF/GVCF.

* DRAGEN-ML also updates PL and GP in the output VCF/GVCF.
* The genotypes (GT field) of some variants may be changed by ML e.g., 0/1 to 1/1 or vice versa.
* DRAGEN-ML PHRED scores (e.g. QUAL) are better calibrated than and differ significantly from those with ML disabled and, as a consequence, QUAL scores tend to not exceed 75 and are capped at 60. By comparison, QUAL scores with ML disabled can exceed 1000. For this reason, the QUAL filtering threshold for both SNP and Indel is set to 3.0103 when DRAGEN-ML is enabled, compared to 10.41 of SNP and 7.83 of Indel for DRAGEN-VC when DRAGEN-ML is disabled.

The following variants types are recalibrated:

* Biallelic and multiallelic variants
* Autosomes and sex chromosomes, including haploid positions
* Force GT calls
* Non primary contigs

## Accuracy Improvements

DRAGEN-ML typically removes 30-50% of SNP FPs, with smaller gains on INDELS. FN counts are reduced by 10% or more. The output QUAL/GQ of DRAGEN-ML is empirically more accurately calibrated than DRAGEN SNV VC without ML. There are significant gains in accuracy statistics across the entire genome with ML enabled. Note that a small number of variant calls may have degraded accuracy with ML enabled compared to VC without ML.

## Run time

DRAGEN-ML adds about 10% to the run time compared to runs without ML.


# Evidence BAM

### Overview

The DRAGEN small variant caller is a haplotype-based caller which performs local assembly of all reads in an active region into a de Bruijn graph (DBG). The assembly process uses all the read bases including the soft-clip bases of reads. The soft-clip bases provide evidence for the presence of variants, specifically longer insertions and deletions which are not present in the read cigar and hence cannot be directly viewed in IGV.

The assembly and realignment step (using pair-HMM) performed by variant caller aims to correct mapping errors made by the original aligner and improves the overall variant caller accuracy. Using the evidence BAM, we can view how the variant caller sees the read evidence and how the reads have been realigned making it a very useful debugging tool.

By default, the evidence BAM contains only a subset of regions processed by the small variant caller. Only regions which have candidate indel variants and some percentage of soft-clip reads in the pileup are realgned and output in the evidence BAM. This is done to reduce the run-time overhead needed to generate the evidence BAM.

Please note this feature is only available in DNA germline and somatic modes, and not supported by the RNA variant caller.

### Outputs

The output of the VC Evidence BAM feature will match the output format that the customer has selected using --output-format option. The default format is bam.

* A bam/cram/sam file with the suffix `_evidence.bam/cram/sam` and the corresponding index file. The evidence BAM can be enabled along with the regular BAM output from the Map-Align step. When multiple BAM are passed as inputs to the variant caller, for e.g., in Tumor-Normal calling, then they will be combined in the evidence BAM output and tagged with appropriate read groups.
* A bed file with regions that were realigned and output in VC Evidence BAM with suffix ".realigned-regions.bed".

### Features

The evidence BAM consists of realigned reads, badly mated reads and reads that are disqualified by the variant caller based on the read likelihood scores.

* Disqualified and Badly Mated reads

  Reads that are badly-mated (when the read and its mate are mapped to different chromosmes) are tagged with a BM tag (integer) and reads that are disqualified (based on read likelihoods) are tagged with the DQ tag (integer). These reads are filtered out by the genotyper in the variant caller.
* Graph Haplotypes

  When enabling graph haplotypes output using `--vc-evidence-bam-output-haplotypes`, all the haplotypes constructed by the de Bruijn graph are output in the evidence BAM as single reads covering the entire active region. The reads and haplotypes are tagged with different read groups which makes it easily distinguishable in IGV. In IGV, we can use “Color Alignments By” or “Group Alignments By” > read group to separate out the reads from the haplotypes. The haplotypes are tagged with read group `EvidenceHaplotype` and the reads are part of the `EvidenceRead_Normal/Tumor` read group.

  The haplotypes are named as Haplotype 1, Haplotype 2 and so on and have an additional ‘HC’ tag (integer). The realigned reads also have an HC tag which encodes which haplotype best matches the read based on the likelihood calculation. Only reads which are supported by a single unique haplotype have the HC tag, reads which match more than one haplotype well do not have an HC tag. The use of this tag is primarily intended to enable highlighting of reads in IGV. Go to "Color Alignments By > Tag" and enter "HC" to view which reads are uniquely supported by a certain graph haplotypes.

### Command Line Arguments

| Name                                     | Description                                                                                | Default Value |
| ---------------------------------------- | ------------------------------------------------------------------------------------------ | ------------- |
| `vc-output-evidence-bam`                 | Enable evidence BAM output                                                                 | False         |
| `vc-evidence-bam-output-haplotypes`      | Output graph haplotypes in evidence BAM                                                    | False         |
| `vc-evidence-bam-clipped-read-threshold` | Percentage of clipped reads in active region to enable evidence BAM output for that region | 10%           |
| `vc-evidence-bam-force-output`           | Force evidence BAM output for all active regions                                           | False         |


# Mosaic Detection

The default mode of the small variant caller has been optimized to detect germline variants with typical AFs of 0%, 50% or 100%. Non-cancer post-zygotic mosaic variants have typical allele fraction (AFs) lower than 50% and therefore more challenging to find with the default small variant caller. To improve sensitivity of low AF calls, a new machine learning (ML) model trained using read and context evidence from low AF calls is used. This allows the model to identify variants down to approximately 5% AF on 35x WGS and 3% AF on 300x WGS. The mosaic ML model is applied to all calls that are rejected by the germline model and variants detected with the mosaic detection are identified by a `MOSAIC` flag in the VCF INFO field.

![](/files/PHJ6Rmrx49nDmbkYXM8o)

`MOSAIC` tagged variants with `QUAL` smaller than `3.0103` are filtered with the `MosaicHardQUAL` filter.

We provide an optional `MosaicLowAF` filtering option to filter `MOSAIC` tagged variants with `AF` smaller than the `AF` threshold. The threshold for this filter can be set with the `--vc-mosaic-af-filter-threshold` option. The default value for the mosaic `AF` filter threshold is set to `0.2` if the median depth of the sample detected by the ploidy caller is `<= 100x` and `0.1` if the detected depth is `>100x`. All `MOSAIC` variants can be reported, regardless of their `AF`, with the `--vc-mosaic-af-filter-threshold=0` option. For haploid regions, we apply separate handling that uses a threshold of `0.75` for assigning `MOSAIC` tags. Please refer to the [Ploidy Support](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling#ploidy-support) section for more details.

Furthermore, the output of `MOSAIC` tagged calls can be restricted using an optional target BED provided with the `--vc-mosaic-target-bed` option.

## High Sensitivity Mosaic detection (Alpha)

The DRAGEN high sensitivity mosaic detection is designed to lower the limit of detection in high-depth datasets to approximately 1.5% AF and it can be enabled with `--vc-enable-high-sensitivity-mosaic-detection=true`. We suggest enabling the high hensitivity mosaic detection when interested in finding very low AF variants in high-depth samples (>=200x). Enabling high-sensitivity mosaic detection in DRAGEN will increase runtime by approximately 20–25%.

**Note:** Internal parameters of DRAGEN are changed from the default settings when the high sensitivity mosaic detection is enabled. This directly affects both the germline and mosaic small variant calling pipelines. Therefore, a direct comparison of the VCF files generated could highlight differences in the variant calls scores, as well as variants private to one VCF (even among non `MOSAIC` tagged calls).

## Command line options

* `--vc-enable-mosaic-detection`

  Set to true to enable mosaic detection. Set to false to disable mosaic detection.
* `--vc-mosaic-af-filter-threshold`

  Set the allele fraction threshold for the application of the `MosaicLowAF` filter to mosaic calls. All `MOSAIC` tagged variants with `AF` smaller than the `AF` threshold are filtered with the `MosaicLowAF` filter. The default mosaic `AF` filter threshold is set to `0.1` for WES samples, and for `WGS` sample it is `0.2` if `cvg <= 100x`, `20/cvg` if `100x < cvg <= 200x`, and `0.1` if `cvg > 200x` where `cvg` is the median depth of the sample detected by the ploidy caller, i.e., `min( max( 20/cvg, 0.1), 0.2)`.
* `--vc-mosaic-qual-filter-threshold`

  Set the `QUAL` threshold for the application of the `MosaicHardQUAL` filter to mosaic calls. All `MOSAIC` tagged variants with `QUAL` smaller than the threshold `QUAL` are filtered with the `MosaicHardQUAL` filter. The default mosaic `QUAL` filter threshold is set to `3.0103`.
* `--vc-mosaic-target-bed`

  Optional target BED file to restrict the output of `MOSAIC` tagged variant calls only in the specified regions.

## Comparison with DRAGEN Mosaic modes across releases

Small variant calling features comparison between default germline small variant caller and mosaic detection mode in DRAGEN 4.2, DRAGEN 4.3 and DRAGEN 4.4

| Tag        | Name                                    | Command line options                                                                             | QUAL threshold | MAPQ0 | Mosaic Detection | Mosaic AF filter threshold                                                                                                                                                                                  |
| ---------- | --------------------------------------- | ------------------------------------------------------------------------------------------------ | -------------- | ----- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 4.2        | DRAGEN 4.2 default Small Variant Caller | --enable-variant-caller=true                                                                     | 3              | No    | No               | N/A                                                                                                                                                                                                         |
| 4.2 HSM    | DRAGEN 4.2 High Sensitivity Mode        | --enable-variant-caller=true --vc-enable-high-sensitivity-mode=true                              | 0.4            | Yes   | Yes (Alpha)      | N/A                                                                                                                                                                                                         |
| 4.3        | DRAGEN 4.3 default Small Variant Caller | --enable-variant-caller=true                                                                     | 3              | Yes   | Yes (Full)       | 20%                                                                                                                                                                                                         |
| 4.3 Mosaic | DRAGEN 4.3 Mosaic Detection Mode        | --enable-variant-caller=true --vc-enable-mosaic-detection=true                                   | 0.4            | Yes   | Yes (Full)       | 0%                                                                                                                                                                                                          |
| 4.4        | DRAGEN 4.4 default Small Variant Caller | --enable-variant-caller=true                                                                     | 3.0103         | Yes   | Yes (Full)       | 20% if WGS and depth <=100x, 10% otherwise                                                                                                                                                                  |
| 4.4        | DRAGEN 4.4 default Small Variant Caller | --enable-variant-caller=true --vc-enable-mosaic-detection=true                                   | 3.0103         | Yes   | Yes (Full)       | 20% if WGS and depth <=100x, 10% otherwise                                                                                                                                                                  |
| 4.4 Mosaic | DRAGEN 4.4 Mosaic Detection Mode        | --enable-variant-caller=true --vc-enable-mosaic-detection=true --vc-mosaic-af-filter-threshold=0 | 3.0103         | Yes   | Yes (Full)       | 0%                                                                                                                                                                                                          |
| 4.5        | DRAGEN 4.5 default Small Variant Caller | --enable-variant-caller=true                                                                     | 3.0103         | Yes   | Yes (Full)       | <p>10% if WES,<br>20% if <code>cvg <= 100x</code>, <code>20/cvg</code> if <code>100x < cvg <= 200x</code>, and 10% if <code>cvg > 200x</code> if <a data-footnote-ref href="#user-content-fn-1">WGS</a></p> |
| 4.5        | DRAGEN 4.5 default Small Variant Caller | --enable-variant-caller=true --vc-enable-mosaic-detection=true                                   | 3.0103         | Yes   | Yes (Full)       | <p>10% if WES,<br>20% if <code>cvg <= 100x</code>, <code>20/cvg</code> if <code>100x < cvg <= 200x</code>, and 10% if <code>cvg > 200x</code> if <a data-footnote-ref href="#user-content-fn-1">WGS</a></p> |
| 4.5 Mosaic | DRAGEN 4.5 Mosaic Detection Mode        | --enable-variant-caller=true --vc-enable-mosaic-detection=true --vc-mosaic-af-filter-threshold=0 | 3.0103         | Yes   | Yes (Full)       | 0%                                                                                                                                                                                                          |

## Comparison with DRAGEN 4.3 command line behavior

The command line behavior of `--vc-enable-mosaic-detection=true` has changed between DRAGEN 4.3 and DRAGEN 4.4. In DRAGEN 4.3, explicitly setting `--vc-enable-mosaic-detection=true` caused two changes to default behavior:

1. The hard filter `QUAL` threshold for both SNPs and INDELs is lowered to `0.4` in this mode to allow low AF calls to be set as `PASS` in the FILTER field.
2. The mosaic `AF` filter threshold is set to `0.0`.

In DRAGEN 4.4, to improve the user experience, both these changes have been decoupled from the `--vc-enable-mosaic-detection=true` option, therefore the user can report all `MOSAIC` variants, regardless of their `AF`, with the `--vc-mosaic-af-filter-threshold=0` option.

[^1]: `cvg` is the median depth of the sample detected by the ploidy caller.


# VCF Imputation

The VCF imputation software can infer multi-allelic SNP and INDEL variants from low-coverage sequencing samples by packaging the GLIMPSE software (2020, Olivier Delaneau & Simone Rubinacci). The DRAGEN implementation of the GLIMPSE software allows for scalability of variant imputation:

* with an end-to-end pipeline where the 3 phases of the GLIMPSE software (Chunk, Phase and Ligate) get executed in a single command, on one chromosome or on multiple chromosomes
* with accceleration supported with Advanced Vector Extensions (AVX)

The DRAGEN VCF imputation software infers variants on autosomes and chromosome X of haploid and diploid species.

Upon completion, the software generates imputed variants based on a reference panel, a genetic map, and input samples provided. The DRAGEN secondary analysis software supports VCF imputation on human data and provides a reference panel and a genetic map for the hg38 reference build accessible on the [DRAGEN Software Support site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

For data other than human data (reference build hg38) the user needs to provide its own reference panel and genetic map. A custom reference panel can be built with the DRAGEN Population Haplotyping software.

Notes:

* The output is in biallelic format, one line per ALT.
* The VCF imputation software only supports input sample data generated with the DRAGEN secondary analysis software.

The following is an example of commands to impute SNPs on a single autosome chromosome:

```
dragen 
--enable-imputation true 
--imputation-ref-panel-dir <REF_PANEL_DIR>
--imputation-ref-panel-prefix <IRPv2.1> 
--imputation-chunk-input-region <chr22> 
--imputation-phase-input-list <VCF_to_be_imputed.txt> 
--imputation-genome-map-dir <MAP_DIR> 
--output-directory <OUT_DIR>
--output-file-prefix <OUT_PREFIX>
```

The following is an example of commands to impute SNPs and INDELs on all human chromosomes (autosomes and chromosome X):

```
dragen 
--enable-imputation true 
--imputation-ref-panel-dir <REF_PANEL_DIR>
--imputation-ref-panel-prefix <IRPv2.1> 
--imputation-chunk-input-region-list <chr_list.txt> 
--imputation-phase-input-list <VCF_to_be_imputed.txt> 
--imputation-genome-map-dir <MAP_DIR> 
--imputation-phase-sample-type-list <path to sample type file>
--imputation-phase-impute-reference-only-variants true
--output-directory <OUT_DIR>
--output-file-prefix <OUT_PREFIX>
```

## Inputs

### Sample Input

The imputation software infers multi-allelic SNP and INDEL variants from low-coverage sequencing samples that are provided by the user. To maximize the accuracy of the imputed variant per sample, the software leverages the information from all provided samples.

The sample(s) to be imputed must have the following format:

* VCF, multi-sample VCF, BCF or multi-sample BCF (zip or unzipped). gVCF is not supported
* Must contain GL (Genotype Likelihoods) or PL (phred-scaled genotype likelihoods) information

To achieve more accurate results, it is recommended to use input VCF generated with the force genotyping capability of the DRAGEN secondary analysis software so that it contains all the positions that are present in the reference panel. A file to be used as input of the force genotyping run of the DRAGEN variant caller, with all sites present in the IRP reference panel (built from human reference genome hg38) is provided in the Imputation files accessible in the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). When running the force genotype option (of DRAGEN variant caller) for imputation, it is recommended to disable the machine learning software (--vc-ml-enable-recalibration=false).

#### Recommendation when imputing SNPs and INDELs

To impute SNPs and INDELs and get the best accuracy on INDELs, it is recommended:

* to force genotype the input VCF with a SNPs-only sites.vcf file using DRAGEN argument --vc-forcegt-vcf. This SNPs-only sites.vcf file contains only the SNPs sites present in the reference panel. A SNPs-only VCF file is also available in the IRP reference panel (built from human reference genome hg38) in the Imputation files accessible in the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).
* and to set the command `--imputation-phase-impute-reference-only-variants` to true.

### Reference Panel

A per-chromosome reference panel in BCF or VCF format that lists all the imputation positions in the targeted regions along with the corresponding haplotypes must be provided. A reference panel (with prefix IRPv{x}) is available in the Imputation files accessible in the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). IRPv2.1 is a multi-allelic SNP, INDELs reference panel containing 3202 samples from the 1000 Genomes Project, which have been variant called using DRAGEN 4.0 against hg38.

Notes: IRPv1.x does nor support chrX, IRPv2.x supports chrX (chrY and chrM are not supported)

A custom reference panel can be built with the DRAGEN Population Haplotyping software. When providing a custom reference panel ensure the chromosome of mixed ploidy chromosome is divided into the PAR and non-PAR regions that exist, and the basename matches the subregions names defined in the JSON config file. The format should be `<PREFIX>.basename`. Examples: IRPv2.0.chrX\_par1, IRPv2.0.chrX\_par2, and IRPv2.0.chrX\_nonpar.

### Genetic Map

A genetic map per chromosome is required to obtain the imputed variants. You can use your own genetic map computed from the recombination rate of the species and its reference genome, or use a prebuilt genetic map corresponding to the human hg38 reference genome. A prebuilt map is available as part of the Imputation files, accessible at the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). DRAGEN does not generate custom genetic map files. The genetic map should follow the format:

* `<chromosome name>.gmap.gz`
* 3 columns: position, chromosome number, distance (cM)
* compliant with the reference genome used to generate the sample input

### JSON config file

This config file allows the proper handling of haploid/diploid chromosomes. This file is present in the same directory of the input reference panel with PREFIX and is available in the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). It must follow the naming convention: {$DIR}/{$PREFIX}.config.json. When the config file is not present in the directory, the software assumes that the imputation is done on all diploid chromosomes.

In the IRP reference panel folder available on DRAGEN support page, the JSON config file corresponds to human data. The user can edit this file if imputation is done on another species.

#### Example of JSON config file

For imputing VCF on human data with typename “M” for Male and “F” for Female (“M” and “F” are the values used in the sample type file):

```
{
  "regions": { 
    "chrX" : [ "chrX_par1", "chrX_nonpar", "chrX_par2" ] 
  },
  "ploidy" : {
     "chrX_nonpar" : { "M": 1, "F": 2},
     "default"     : { "M": 2, "F": 2}
  }
 }
```

#### Instructions to make a custom JSON configuration file:

The JSON config file is made of two fields as defined in the table below

| Fields  | Required/Optional                                                                                                                                                                                       | Purpose                                                                                                                        | Type                                                                                                                                                                                                                            |
| ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| regions | Required only when a chromosome of mixed ploidy is present in the Reference Panel folder                                                                                                                | Define contig name and subregion name of mixed ploidy chromosome                                                               | Dictionary in the form: contigname\_of\_mixed\_ploidy :\[contigname\_of\_mixed\_ploidy"\_par1", contigname\_of\_mixed\_ploidy"\_par2", contigname\_of\_mixed\_ploidy"\_nonpar1", contig\_name\_of\_mixed\_ploidy"\_nonpar2"...] |
| ploidy  | <ul><li>“default” is a required name</li><li><br></li><li>contigname\_of\_mixed\_ploidy\_"nonpar" is required only when a chromosome of mixed ploidy is present in the Reference Panel folder</li></ul> | <p>Define:<br></p><ul><li>ploidy behavior when different from “default”</li><li><br></li><li>default ploidy behavior</li></ul> | <p>Dictionary in the form: contigname\_of\_mixed\_ploidy\_"nonpar": { typename1 : 1, typename2 : 2}<br>"default"   : { "typename1": 2, "typename 2": 2}<br><br>typename is used in the Sample Type file input</p>               |

Note: ensure the subregion names match the genetic map name. Example: if "chrX\_nonpar" is defined in the "region" field of the JSON config file, then the genetic map corresponding to chromosome X non PAR region in the Reference Panel folder must be named "chrX\_nonpar".gmap.gz.

### Sample type file

The sample type file is required when haplotyping is performed on non-PAR regions of mixed ploidy chromosomes to define the typename of each sample.

The sample type file is a txt file with the following format

* 2 columns, tabs or space delimited
* First column: list of all sample names present in the input sample
* Second column: typename value for each sample. This typename value should match the typenames used in the JSON config file.

## Outputs

The VCF imputation software generates several outputs:

* The imputed variant file with concatenated imputed variants: one single VCF or msVCF file for all specified regions/chromosomes with name `<prefix>.impute.vcf.gz`
* The intermediate files:
  * chunk regions to be passed along to the internal Phase step with name `<prefix>.impute.chunk.out.txt`
  * imputed variants per chunks identified: VCF or msVCF depending on the input sample format with name `<prefix>_chr_start-end.impute.phase.vcf.gz`
  * text file with path to all the `<prefix>_chr_start-end.impute.phase.vcf.gz` generated with name `<prefix>.impute.phase.out.txt`

Note: while the imputation software can impute multi-allelic positions, the output is in biallelic format, one line per ALT. The bcftools software can be used to post-collapse all ALT in one line with the command: `bcftools norm -m +snps`

## Command Line Options

| Option                                            | Type   | Required                                                                                 | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| ------------------------------------------------- | ------ | ---------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --enable-imputation                               | NA     | Yes                                                                                      | Set to `true` to enable vcf imputation pipeline                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| --imputation-ref-panel-dir                        | STRING | Yes                                                                                      | Directory containing per-chromosome reference panel VCF/BCF format and optionally the JSON config file                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| --imputation-ref-panel-prefix                     | STRING | Yes                                                                                      | Prefix for reference panel files and the JSON config file                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| --imputation-genome-map-dir                       | STRING | Yes                                                                                      | Directory containing per-chromosome genome map files                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| --imputation-chunk-input-region                   | STRING | Yes for single region                                                                    | Target region, usually a full chromosome (e.g. chr20:1000000-2000000 or chr20).                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| --imputation-chunk-input-region-list              | STRING | Yes for list of regions                                                                  | Text file listing chromosomes or regions to be processed, one chromosome/region per line.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| --imputation-phase-input                          | STRING | Yes for single VCF file                                                                  | Sample input file with VCF/BCF format. Single VCF or multi-sample VCF                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| --imputation-phase-input-list                     | STRING | Yes for multiple VCF files                                                               | Text file listing sample input in VCF/BCF format, one input file per line                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| --imputation-phase-sample-type                    | STRING | Yes when imputing on a non PAR region of mixed ploidy chromosome AND a single VCF file   | Define typename of the VCF file imputed. The typename must match one of the two typenames defined in the JSON config file                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| --imputation-phase-sample-type-list               | STRING | Yes when imputing on a non PAR region of mixed ploidy chromosome AND a list of VCF files | Path to the Sample Type file                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| --output-directory                                | STRING | Yes                                                                                      | Output directory                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| --output-file-prefix                              | STRING | Yes                                                                                      | Output files prefix                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| --imputation-phase-threads                        | INT    | No                                                                                       | Specify the number of threads to use. Default is the number of system threads                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| --imputation-phase-filter-input-sample-in-ref     | NA     | No                                                                                       | Default is `true`: if sample ID matches between reference panel and sample input, then the corresponding samples are ignored from the reference panel to avoid imputation against itself. To be turned to `false` if all samples from the reference panel should be kept regardless of their presence in the sample input.                                                                                                                                                                                                                                                                                                                            |
| --imputation-phase-impute-reference-only-variants | STRING | No                                                                                       | <p>Default is <code>false</code>. If set to <code>true</code>, allows imputation at variants only present in the reference panel. The use of this option is intended only to allow imputation at sporadic missing variants. If the number of missing variants is non-sporadic, please re-run the genotype likelihood computation at all reference variants and avoid using this option, since data from the reads should be used.<br>When imputing INDELS and SNPs positons, it is recommended to input samples that have been variant called using <code>--vc-forcegt-vcf</code> with SNPs-only sites.vcf file AND to turn this command to true.</p> |
| --imputation-phase-input-independently            | STRING | No                                                                                       | Default is `false`. If set to `true`, allows to treat each sample input independently without using them in the reference panel calculation                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |

Note: with this end-to-end implementation of the GLIMPSE software, the parameters window\_size and buffer\_size are respectively set to 2 Mb and 200 kb.


# Multi-Region Joint Detection

DRAGEN Multi-region Joint Detection (MRJD) is a de novo germline small variant caller for paralogous regions. MRJD currently covers regions that include six clinically relevant genes: *NEB*, *TTN*, *SMN1/2*, *PMS2*, *STRC*, and *IKBKG*. MRJD is compatible with hg38, hg19 and GRCh37 reference genome. The table below includes hg38 region coordinates covered by MRJD.

| Chromosome | Start     | End       | Description         |
| ---------- | --------- | --------- | ------------------- |
| chr2       | 151578759 | 151588523 | *NEB* exon 98-105   |
| chr2       | 151589318 | 151599076 | *NEB* exon 90-97    |
| chr2       | 151599871 | 151609628 | *NEB* exon 82-89    |
| chr2       | 178653238 | 178654995 | *TTN* exon 172-180  |
| chr2       | 178657498 | 178659255 | *TTN* exon 181-189  |
| chr2       | 178661759 | 178663516 | *TTN* exon 190-198  |
| chr5       | 70049522  | 70077596  | *SMN2*              |
| chr5       | 70924940  | 70953013  | *SMN1*              |
| chr7       | 5970924   | 5980896   | *PMS2* exon 13-15   |
| chr7       | 5980968   | 5987689   | *PMS2* exon 11-12   |
| chr7       | 6737007   | 6743712   | *PMS2CL* exon 2-3   |
| chr7       | 6743880   | 6753867   | *PMS2CL* exon 4-6   |
| chr15      | 43599563  | 43602630  | *STRC* exon 24-29   |
| chr15      | 43602982  | 43611000  | *STRC* exon 14-23   |
| chr15      | 43611040  | 43618800  | *STRC* exon 1-13    |
| chr15      | 43699379  | 43702452  | *STRCP1* exon 23-28 |
| chr15      | 43702488  | 43710472  | *STRCP1* exon 13-22 |
| chr15      | 43710502  | 43718262  | *STRCP1* exon 1-12  |
| chrX       | 154555884 | 154565047 | *IKBKG* exon 3-10   |
| chrX       | 154639390 | 154648553 | *IKBKGP1*           |

## MRJD method

MRJD is a variant calling method that is designed to detect de novo germline small variants in paralogous regions of the genome. A conventional variant caller relies on the read aligner to determine which reads likely originated from a given location. This method works well when the region of interest does not resemble any other region of the genome over the span of a single read (or a pair of reads for paired-end sequencing). However, a significant fraction of the human genome does not meet this criterion. At least 5% of the human genome consists of segmental duplications. Many regions of the genome have near-identical copies elsewhere, and as a result, the true source location of a read might be subject to considerable uncertainty. If a group of reads is mapped with low confidence, a conventional variant caller might ignore the reads, even though they contain useful information. If a read is mismapped (i.e., the primary alignment is not the true source of the read), it can result in variant detection errors.

MRJD is designed in an attempt to tackle the complexities raised by segmental duplication regions. Instead of considering each region in isolation, MRJD considers all locations from which a group of reads may have originated and attempts to detect the underlying sequences jointly across all paralogous regions in the sample of interest.

Below is a diagram showing the general workflow of MRJD in a pair of paralogous regions. MRJD takes primary alignments in all paralogous regions, regardless of mapping quality, builds and places all copies in a pair of paralogous regions based on reads and prior knowledge, call small variants based on the placed copies, and output final genotypes.

![Workflow](/files/WZVmrKYKadZ0YQOlkYXD)

Figure 1. MRJD Caller workflow.

## Two modes of the MRJD Caller

There are two modes of the DRAGEN MRJD Caller, default mode and high sensitivity mode. Here are details on the differences between the two modes.

### Default mode

With `--enable-mrjd=true`, the MRJD Caller will report the following two types of variants:

1. **Uniquely placed variants**, which means the variant is found and placed in one of the paralogous regions without ambiguity. These variants will be labeled as **"UNIQUELY\_PLACED"** in the VCF INFO field.
2. **Region-ambiguous variants**. In this case, the aggregated genotype contains a variant allele with high confidence, but MRJD Caller is unable to place the variant allele in one of the paralogous regions with high confidence. The MRJD Caller will report the variant allele in all paralogous regions. These variants will be labeled as **"REGION\_AMBIGUOUS"** in the VCF INFO field.

### High Sensitivity mode

With both `--enable-mrjd=true` and `--mrjd-enable-high-sensitivity-mode=true`, the MRJD Caller reports the same variants as from the default mode, plus two other types of variants.

1. **Positions where the reference alleles in all paralogous regions are not the same.** It is well established that gene conversion, including reciprocal crossover, is a common event between paralogous regions (such as *PMS2* and *PMS2CL*). When reciprocal crossover event occurs, the prior model, without nearby information on phasing, might end up placing the converted haplotype in the source region instead of the destination region, resulting in no variant. The high sensitivity mode compensates for this event by reporting the variant in corresponding positions in all paralogous regions. These variants will be labeled as **"MRJD\_HS;REF\_DIFF\_SITE"** in the VCF INFO field.
2. **Variants that have been placed uniquely in one of the paralogous regions and no variant in the corresponding position in the other region.** The high sensitivity mode reports the variant in the rest of the paralogous regions. This is to compensate the fact that sometimes the prior knowledge that is used to help place the variant is not sufficient or is estimated incorrectly. In those cases, the variant allele still exists but is placed in the wrong paralog region. Therefore, reporting the variant in the other paralogous regions can help maximize sensitivity even with the lack of prior. These variants will be labeled as **"MRJD\_HS;ALT\_LOCATION"** in the VCF INFO field.

## Running DRAGEN MRJD

The MRJD Caller is disabled by default and requires WGS data aligned to a human reference genome build 38, 19, or GRCh37.

Here is the list of options related to MRJD.

* `--enable-mrjd` If set to true, MRJD is enabled for the DRAGEN pipeline.
* `--mrjd-enable-high-sensitivity-mode` If set to true, MRJD high sensitivity mode is enabled for the DRAGEN pipeline. See previous section on what variant types are reported in MRJD default mode and high sensitivity mode (default = ‘false’).

The following command-line example uses FASTQ input and runs MRJD Caller with high sensitivity mode:

```sh
dragen \
  -r <REF> \
  -1 <FQ1> \
  -2 <FQ2> \
  --RGID <RG> --RGSM <SM> \
  --output-dir <OUTPUT> \
  --output-file-prefix <PREFIX> \
  --enable-map-align=true \
  --enable-map-align-output=true \
  --enable-sort=true \
  --enable-duplicate-marking=true \
  --enable-mrjd true \
  --mrjd-enable-high-sensitivity-mode true
```

The following command-line example uses BAM input that has already been aligned and runs MRJD Caller with high sensitivity mode:

```sh
dragen \
  -r <REF> \
  -b <BAM> \
  --output-dir <OUTPUT> \
  --output-file-prefix <PREFIX> \
  --enable-map-align=false \
  --enable-mrjd true \
  --mrjd-enable-high-sensitivity-mode true
```

### Example WGS workflow that includes both DRAGEN Small Variant Caller and MRJD

Starting from DRAGEN v4.4, MRJD can run together with the DRAGEN Small Variant Caller in the same DRAGEN run. Here are the example command lines to run DNA Mapping using FASTQ files as input, followed by Small Variant Calling and MRJD.

```sh
# run DNA Mapping, Small Variant Calling, and MRJD
dragen \
  -r <HASH_TABLE> \
  -1 <FQ1> \
  -2 <FQ2> \
  --RGID <RG> --RGSM <SM> \
  --output-dir <OUTPUT_DIRECTORY> \
  --output-file-prefix <PREFIX> \
  --enable-map-align true \
  --enable-map-align-output true \
  --enable-sort true \
  --enable-duplicate-marking true \
  --enable-variant-caller true \
  --enable-mrjd true \
  --mrjd-enable-high-sensitivity-mode true
```

Using this workflow, two VCF files will be created (`<sample>.hard-filtered.vcf.gz` by DRAGEN Small Variant Caller and `<sample>.mrjd.hard-filtered.vcf.gz` by DRAGEN MRJD). To help user get a single VCF file for downstream analysis, we prepared an utility tool that replaces the DRAGEN Small Variant Caller output in the homology region of the six medically relevant and challenging genes with MRJD caller output. The tool also annotates the calls made by MRJD (with "MRJD" tag in the INFO column). Please refer to the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html) to download the utility tool.

## Output format

The MRJD Caller generates a `<sample>.mrjd.hard-filtered.vcf.gz` file in the output directory. The output file is a compressed VCFv4.2 formatted file that contains the VCF representation of the small variants from the identified genotype.

### Uniquely placed call

The following are example output format for uniquely placed variant. The `DRAGENHardQual` filter is applied to the records if the variant has a QUAL < 3.00.

![MRJD\_unique](/files/1hgSnUaXrVodwicFsyN6)

Figure 2. VCF output format example for uniquely placed call.

### Non-uniquely-placed call

For variants that are not uniquely placed, including region-ambiguous variants from default mode, and all variants from high sensitivity mode, the MRJD Caller will also report variants under diploid genotype format, which can be interpreted the same way as uniquely placed variant (the genotype is region-specific instead of being an aggregate across all regions). Under this format, the QUAL represents phred-scaled quality score for the assertion made in ALT (i.e. `−10log10 prob(GT==0/0)`). Note that the QUAL score will be equal to or less than 3 (if the QUAL > 3, then the call should be uniquely placed).

The QUAL, GT, GQ and PL will be reported similarly to the DRAGEN germline VC. To avoid losing information about the aggregated genotype across paralogous regions, the MRJD Caller reports genotype, phred-scaled quality score, and the phred-scaled genotype likelihoods for aggregated genotype using JGT, JQL, and JPL in the FORMAT column.

![MRJD\_region\_ambiguous](/files/nXM5DwtRwkIegBPScz5U)

Figure 3. VCF output format example for region-ambiguous call.


# Small Variant Calling Metrics

The QC metrics are printed to the standard output. In addition CSV files are written to the run output directory:

* \<output prefix>.vc\_metrics.csv

Metrics are reported for each sample in multi sample VCF and gVCF files and output in a csv file with the file name ending in "vc\_metrics.csv". Based on the run case, metrics are reported either as standard VARIANT CALLER or JOINT CALLER. Metrics are reported both for PREFILTER (including all variant calls regardless of FILTER field) and POSTFILTER (including only PASS calls) categories.

Panel of Normals (PON) and COSMIC filtered variants are counted as PASS variants in the POSTFILTER VCF metrics. These PASS variants can cause higher than expected variant counts in the POSTFILTER VCF metrics

* **Number of samples**---Number of samples in the population/ joint VCF.
* **Reads Processed**---The number of reads used for variant calling, excluding any duplicate marked reads and reads falling outside of the target region.
* **Total**---The total number of variants (SNPs + MNPs + indels).
* **Biallelic**---Number of sites in a genome that contains two observed alleles. The reference is counted as one allele, which allows for one variant allele.
* **Multiallelic**---Number of sites in the VCF that contain three or more observed alleles. The reference is counted as one, which allows for two or more variant alleles.
* **SNPs**---A variant is counted as an SNP when the reference, allele 1, and allele 2 are all length 1.
* **Insertions (Hom)**---Number of variants that contains homozygous insertions.
* **Insertions (Het)**---Number of variants where both alleles are insertions, but not homozygous.
* **Deletions (Het)**---Number of variants that contains homozygous deletions.
* **INDELS (Het)**---Number of variants where genotypes are either \[insertion+deletion], \[insertion+SNP], or \[deletion+SNP].
* **De Novo SNPs**---*De novo* marked SNPs with DQ greater than the threshold. Set the `--qc-snp-denovo-quality-threshold` option to the required threshold. The default is 0.05 if ML recalibration is off, 0.0017 if ML recalibration is on.
* **De Novo INDELs**---*De novo* marked indels with DQ values greater than the threshold. This DQ threshold can be specified by setting the `--qc-indel-denovo-quality-threshold` option to the required DQ threshold. The default is 0.4 if ML recalibration is off, 0.04 if ML recalibration is on.
* **De Novo MNPs**---*De novo* marked SNPs with DQ greater than the threshold. Set the `--qc-snp-denovo-quality-threshold` to the required threshold. The default is 0.05 if ML recalibration is off, 0.0017 if ML recalibration is on.
* **(Chr X SNPs)/(Chr Y SNPs) ratio in the genome (or the target region)** ---Number of SNPs in chromosome X (or in the intersection of chromosome X with the target region) divided by the number of SNPs in chromosome Y (or in the intersection of chromosome Y with the target region). If there was no alignment to either chromosome X or chromosome Y, this metric shows as NA.
* **SNP Transitions**---Number of transitions, the interchange of two purines (A<->G) or two pyrimidines (C<->T), across the genome.
* **SNP Transversions**---Number of transversions, the interchange of purine and pyrimidine bases, across the genome.
* **Ti/Tv ratio**---Ratio of transitions to transitions.
* **SNP Transitions (non-centromeric)**---Number of transitions across the genome, excluding centromeres.
* **SNP Transversions (non-centromeric)**---Number of transversions across the genome, excluding centromeres.
* **Ti/Tv ratio (non-centromeric)**---Ratio of transitions to transitions in non-centromeric regions.
* **Heterozygous**---Number of heterozygous variants.
* **Homozygous**---Number of homozygous variants.
* **Het/Hom ratio**---Heterozygous/ homozygous ratio.
* **In dbSNP**---Number of variants detected that are present in the dbSNP reference file. If no dbSNP file is provided via the `--bsnp` option, then both the In dbSNP and Novel metrics show as NA.
* **Novel**---Total number of variants minus number of variants in dbSNP.
* **Percent Callability**---Available in germline and somatic modes with gVCF output. The percentage of non-N reference positions having a PASSing genotype call. Multiallelic variants are not counted. Deletions are counted for all the deleted reference positions only for homozygous calls. Only autosomes and chromosomes X, Y, and M are considered. To produce this metric for non-human references, set --qc-callability-autosome-contigs to specify the autosome contig names. Optionally, --qc-callability-xym-contigs allows setting X, Y and M contig names.
* **Percent Autosome Callability**---Only autosomes are considered. To produce this metric for non-human references, set --qc-callability-autosome-contigs to specify the autosome contig names.
* **Percent QC Region Callability in Region i (i is equivalent to regions 1,2, or 3)**---Available if callability for custom regions is requested via the `--qc-coverage-region-i` option and the callability output is specified with `--qc-coverage-reports-i`. All contigs are considered. Setting --qc-callability-autosome-contigs enables outputting this metric for non-human references.

### Per Contig Het/Hom Ratio

When the germline small variant caller is executed, DRAGEN calculates a het/hom ratio per contig.

The het/hom ratio values can be used as an indication of whole chromosome uniparental disomy (UPD). UPD of certain chromosomes are associated with genetic syndromes known as imprinting disorders. Whole chromosome UPD have het/hom ratios close to 0.0. Ranges vary, but are usually between 1.0–2.0. The het/hom ratios should be interpreted in the context of the specific assay.

DRAGEN reports the ratios for both PREFILTER (including all variant calls regardless of FILTER field) and POSTFILTER (including only PASS calls) categories. The metrics are output to the `.vc_hethom_ratio_metrics.csv` file.

The file contains the following values for each primary contig processed.

* Contig
* Number of heterozygous variants
* Number of homozygous variants
* Het/Hom ratio

The following example shows a section of the metrics.

```
VARIANT CALLER POSTFILTER,HG04070,1 Heterozygous,185733
VARIANT CALLER POSTFILTER,HG04070,1 Homozygous,182928
VARIANT CALLER POSTFILTER,HG04070,1 Het/Hom ratio,1.015
VARIANT CALLER POSTFILTER,HG04070,2 Heterozygous,203946
VARIANT CALLER POSTFILTER,HG04070,2 Homozygous,174294
VARIANT CALLER POSTFILTER,HG04070,2 Het/Hom ratio,1.170
VARIANT CALLER POSTFILTER,HG04070,3 Heterozygous,192861
VARIANT CALLER POSTFILTER,HG04070,3 Homozygous,130087
VARIANT CALLER POSTFILTER,HG04070,3 Het/Hom ratio,1.483
VARIANT CALLER POSTFILTER,HG04070,4 Heterozygous,178389
VARIANT CALLER POSTFILTER,HG04070,4 Homozygous,157062
VARIANT CALLER POSTFILTER,HG04070,4 Het/Hom ratio,1.136
```


# Copy Number Variant Calling

The DRAGEN Copy Number Variant (CNV) Pipeline detects copy number aberrations and regions with loss of heterozygosity (LOH) from next-generation sequencing (NGS) data. It supports germline and somatic workflows for both whole-genome sequencing (WGS) and whole-exome sequencing (WES) in a single interface via the DRAGEN Host Software.

## Choose Your Workflow

Use the table below to identify the right pipeline for your use case and jump directly to its documentation.

| Sample type | Data type | Input samples  | Documentation                                                                                                 |
| ----------- | --------- | -------------- | ------------------------------------------------------------------------------------------------------------- |
| Germline    | WGS / WES | Single / Multi | [Germline CNV Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline) |
| Somatic     | WGS / WES | Single         | [Somatic CNV Calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-somatic)   |

Example commandlines are provided under [DRAGEN Recipes](https://github.com/illumina-swi/dragen-docs/blob/release/4.5-prod/product-guides/dragen-v4.5/user-guide/dragen-recipes/README.md).

Visit [Reference](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference) for more details on the CNV component.

## Before You Begin

Before running the CNV pipeline, ensure the following prerequisites are in place:

1. **CNV-enabled reference hashtable** — The hashtable must be built with `--ht-build-cnv-hashtable true`. This generates an additional k-mer uniqueness map used to correct mappability biases.
2. **Aligned BAM or CRAM input** — The pipeline accepts pre-aligned reads. If you are starting from FASTQ, first run map/align or use streaming alignments.
3. **Panel of normals (required for WES)** — WES normalization requires a panel of normals. WGS can use self-normalization instead.

## Pipeline Overview

The DRAGEN CNV Pipeline processes the input signal through the following stages:

**1. Target Counts** — Read counts and other signals are extracted from alignments and binned into target intervals.

![](/files/yWJoqeE43R0mksjg2i5R)

**2. Normalization** — The case sample is normalized against a panel of normals or against the estimated normal ploidy. Systematic biases (e.g., GC bias) are corrected to amplify event-level signals.

![](/files/OEh48qVpaaNPyjFojzE0)

**3. Segmentation** — The normalized signal is segmented using one of the available segmentation algorithms.

![](/files/vbtuo277XMiH0CO2plHd)

**4. CNV Calling** — Events are called from the segments, scored, and emitted in the output VCF.

![](/files/H2VMcuAyAlZaHbcQbBtP)


# Germline

## Overview

DRAGEN provides germline copy number variant (CNV) calling workflows that detect copy number aberrations and regions with absence of heterozygosity (AOH) in whole genome sequencing (WGS) and whole exome sequencing (WES) data. The CNV workflows leverage both depth of coverage and B-allele frequencies (BAFs) to provide comprehensive detection of:

* Copy number gains (duplications) and losses (deletions)
* Copy-neutral loss of heterozygosity (CNLOH)
* Whole-arm and whole-chromosome aneuploidies (via Cytogenetics modality)
* Mosaic alterations (WGS only, enabled by default)
* Minor allele copy number estimation

For applications that do not require allele-specific information, mosaic alterations and whole-arm/-chromosome aneuploidies, our legacy depth-only workflow is also available. See [Depth-Only Workflow](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy) for details.

## Workflow

The germline CNV workflow follows this processing pipeline:

![](/files/hCL7QIQITxJkh6xDXkAk)

The pipeline consists of the following modules:

1. **Target Counts** — Binning of read counts and other signals from alignments
2. **B-Allele Counts** — Extraction of allelic read counts
3. **Bias Correction** — Correction of GC bias and other systematic biases
4. **Normalization** — Detection of normal ploidy levels and normalization
5. **Segmentation** — Breakpoint detection via segmentation of normalized depth and BAF signals
6. **ASCN Calling** — Integration of depth and BAF segments to determine copy number states and allele-specific information

## Example Command Lines

### WGS

Note: add `--cnv-stop-after-intrinsic-corrections=true` if interested only in target counts generation + bias correction.

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--bam-input <BAM> \
--cnv-population-b-allele-vcf <POP_SNP_VCF>
```

### WES

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--bam-input <BAM> \
--cnv-target-bed <CNV_TARGET_BED> \
--cnv-normals-list <PANEL_OF_NORMALS> \
--cnv-population-b-allele-vcf <POP_SNP_VCF>
```

Alternatively, you can use a pre-combined panel of normals file:

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--bam-input <BAM> \
--cnv-target-bed <CNV_TARGET_BED> \
--cnv-combined-counts <CNV_PANEL_OF_NORMALS> \
--cnv-population-b-allele-vcf <POP_SNP_VCF>
```

## Required Options

| Option       | Description                           |
| ------------ | ------------------------------------- |
| --enable-cnv | Enable CNV processing (set to `true`) |

### Input

| Option                        | Description                                                                       |
| ----------------------------- | --------------------------------------------------------------------------------- |
| --fastq-file1, --fastq-file2  | FASTQ input files (requires `--enable-map-align true`)                            |
| --bam-input                   | BAM input file                                                                    |
| --cram-input                  | CRAM input file                                                                   |
| --ref-dir                     | DRAGEN reference genome hashtable directory                                       |
| --enable-map-align            | Enable mapper and aligner module                                                  |
| --cnv-population-b-allele-vcf | Population SNP catalog for BAF estimation                                         |
| --cnv-target-bed              | BED file defining exome capture regions (only for WES)                            |
| --sample-sex                  | Sample sex (e.g., `male`, `female`). If not specified, sex is estimated from data |

You can download a suitable population SNP catalog (Resource file "CNV Population SNP VCF") for your associated reference at [this page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html)

### Segmentation

The default segmentation mode depends on the sample type. Germline WGS samples use SLM by default. Germline WES samples use HSLM by default. The segmentation mode can be set explicitly with `--cnv-segmentation-mode (SLM|HSLM)`. See [Segmentation](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#segmentation) in the CNV reference section for a description of SLM and its HSLM variant.

| Option          | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --cnv-slm-eta   | Probability that the segmenter changes to any other state than the current state going from the current target to the next target. This could also be expressed as the probability that the true depth for adjacent targets is different for reasons that simple counting noise does not adequately explain. Likewise, the stay-in-state probability is (1.0 - eta). The default value is 4e-5, the range is (0.0, 1.0) excluding endpoints. Decreasing this value results in longer segments and reduced fragmentation; increasing produces shorter segments with more fragmentation. |
| --cnv-slm-omega | Scaling parameter modulating the relative weight between experimental and biological variance. The default is 0.3, the range is (0.0, 1.0) excluding endpoints. In general, decreasing this value produces longer segments with less fragmentation; increasing produces shorter segments with more fragmentation.                                                                                                                                                                                                                                                                      |

The following options apply to the HSLM segmentation method, which is only used in germline WES.

| Option            | Description                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --cnv-slm-stepeta | Distance normalization parameter. The default is 10000. Modifies the effective eta based on the genomic distance between consecutive target intervals. This can progressively relax stay-in-state, or "stickiness", of the segmenter as adjacent targets become farther apart, making the method adaptive to unequal spacing. Decreasing produces shorter segments with more fragmentation; increasing produces longer segments with less fragmentation. |
| --cnv-slm-fw      | Minimum number of depth bins or targets required for a segment to be retained. This is an internal hard filter at the segmentation stage. The default is 0, which disables it. This is largely vestigial; use of this option is not recommended.                                                                                                                                                                                                         |

The following options are documented here in proximity to segmentation options because of their direct relevance to each other. Once provisional calls for copy number (CN) and minor copy number (MCN) have been made on the resulting segments from the segmentation stage, adjacent segments with the same CN and MCN are joined together to form one single segment. This is continued until no two adjacent segments satisfy the merging criteria. Segment merging is a critical step which compensates for over-segmentation or over-fragmentation happening at the segmentation stage. However, segment merging cannot split segments apart, so it cannot compensate in the other direction. **Thus, segmentation can afford to produce a degree of over-segmentation, but there is no compensatory mechanism for under-segmentation.** The following options control segment merging in germline analyses and do not depend on segmentation method or the segmentation options in use.

| Option                | Description                                                                                                                                                                                                                            |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --cnv-merge-distance  | Maximum gap in base pairs between two adjacent segments that still allows them to be merged. The default is 1000 for germline WGS. For WES the default is effectively unlimited, since target intervals are inherently non-contiguous. |
| --cnv-merge-threshold | Maximum difference in segment mean (linear copy ratio) between two adjacent segments that still allows them to be merged. The default is 0.2 for germline WGS and 0.4 for germline WES.                                                |

Setting `--cnv-merge-threshold` to zero disables segment merging entirely. This is not recommended.

### Normalization

The following options are mutually exclusive:

| Option                          | Description                                                                |
| ------------------------------- | -------------------------------------------------------------------------- |
| --cnv-normals-list              | Text file containing paths to reference target counts files (one per line) |
| --cnv-normals-file              | Individual normal counts file (use multiple times for multiple files)      |
| --cnv-combined-counts           | Combined panel of normals file (`.combined.counts.txt.gz`)                 |
| --cnv-enable-self-normalization | Use self-normalization for sample normalization (only available for WGS)   |

### Output

| Option               | Description                               |
| -------------------- | ----------------------------------------- |
| --output-directory   | Output directory for all results          |
| --output-file-prefix | Prefix prepended to all output file names |

### Workflow Configuration

| Option                                 | Description                                                            | Default | WGS | WES |
| -------------------------------------- | ---------------------------------------------------------------------- | ------- | :-: | :-: |
| --cnv-enable-mosaic-calling            | Enable detection of mosaic alterations                                 | `true`  |  ✓  |     |
| --cnv-enable-cyto-output               | Enable cytogenetics-compatible output VCF                              | `true`  |  ✓  |  ✓  |
| --cnv-enable-legacy-vcf-format         | Use VCF v4.2 format instead of VCF v4.4                                | `false` |  ✓  |  ✓  |
| --cnv-stop-after-intrinsic-corrections | Stop processing after generating target counts and GC-corrected counts | `false` |  ✓  |  ✓  |

**Note:** Mosaic calling is available for WES but not recommended (disabled by default) due to lack of extensive validation.

### Output Filtering

| Option                        | Description                                                | Default        |
| ----------------------------- | ---------------------------------------------------------- | -------------- |
| --cnv-enable-ref-calls        | Emit copy-neutral (REF) calls in output VCF                | `true` for WGS |
| --cnv-filter-length           | Minimum event length (bp) for PASS calls                   | `10000`        |
| --cnv-exclude-bed             | BED file specifying intervals to exclude from analysis     | Not set        |
| --cnv-exclude-bed-min-overlap | Minimum overlap fraction for exclusion                     | `0.5`          |
| --cnv-post-vcf-target-bed     | BED file used to only emit calls overlapping BED intervals | Not set        |

## Output Files

The germline CNV workflow generates the following output files:

| File                           | Description                                | Format           |
| ------------------------------ | ------------------------------------------ | ---------------- |
| .target.counts.gz              | Raw target counts before bias correction   | gzipped TSV      |
| .target.counts.gc-corrected.gz | GC-bias corrected target counts            | gzipped TSV      |
| .tn.tsv.gz                     | Tangent-normalized coverage signal         | gzipped TSV      |
| .ballele.counts.gz             | B-allele counts at population SNP sites    | gzipped TSV      |
| .baf.bedgraph.gz               | B-allele frequency in bedgraph format      | gzipped bedGraph |
| .seg                           | Segmentation results (depth and BAF)       | TSV              |
| .cnv.vcf.gz                    | Primary CNV calls (VCF v4.4 by default)    | gzipped VCF      |
| .cyto.vcf.gz                   | Cytogenetics-compatible calls (if enabled) | gzipped VCF      |
| .cnv\_metrics.csv              | Summary metrics including predicted sex    | CSV              |
| .cnv.gff3                      | Variant calls in GFF format                | GFF              |
| .tn.bw                         | Tangent-normalized signal track            | BigWig           |

### Target Counts Output

`<prefix>.target.counts.gz`

Compressed tab-delimited file containing the number of read counts per target interval. This is the raw signal as extracted from the alignments of the BAM or CRAM file. The format is identical for both the case sample and any panel of normals samples. There is also a bigWig representation of a `target.counts.diploid` file, which is normalized to the normal ploidy level of 2 instead of raw counts.

Columns:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. Count of alignments in this interval
6. Count of improperly paired alignments in this interval

Header lines starting with `#` contain the DRAGEN version, command line, and other meta information.

Example:

```
#TARGET COUNTS FILE
##DRAGENVersion=<VERSION_INFO>
##DRAGENCommandLine=<CommandLineOptions>
...
contig  start  stop   name                <SampleName> improper_pairs
1       565480 565959 target-wgs-1-565480 7          6
1       566837 567182 target-wgs-1-566837 9          0
1       713984 714455 target-wgs-1-713984 34         4
1       721116 721593 target-wgs-1-721116 47         1
1       724219 724547 target-wgs-1-724219 24         21
```

`<prefix>.target.counts.gc-corrected.gz`

Contains GC-corrected read counts per target interval. The format is equivalent to the `*.target.counts.gz` file:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. GC-corrected read counts in this interval
6. Count of improperly paired alignments in this interval

Example:

```
#GC CORRECTED FILE
##DRAGENVersion=<VERSION_INFO>
##DRAGENCommandLine=<CommandLineOptions>
...
contig  start   stop    name    <SampleName> improper_pairs
chr1    818022  819840  target-wgs-chr1-818022:819840   1071.353133     6
chr1    819840  821337  target-wgs-chr1-819840:821337   1051.014997     19
chr1    821337  822485  target-wgs-chr1-821337:822485   1098.6502       10
chr1    822485  824431  target-wgs-chr1-822485:824431   1117.28308      7
```

For more information, see [Target Counts File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#target-counts-file) and [GC Bias Correction](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#gc-bias-correction).

### Normalized Coverage Output

`<prefix>.tn.tsv.gz`

Contains the normalized signal of the case sample per target interval, i.e., the log2-transformed copy ratio signal. A strong signal deviation from 0.0 indicates a potential for a CNV event. The format is equivalent to the `*.target.counts.gz` file:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. Log2-transformed copy ratio in this interval
6. Count of improperly paired alignments in this interval

Header lines are also included that start with `#`. In some cases, the normalization counts could be patched internally with intervals from other processes, such as the SegDups extension. In such cases, patches are indicated (sorted in order of application) with header lines starting with `#patch`:

```
#patch 1 = <normalized_counts_patch_1_filename>
#patch 2 = <normalized_counts_patch_2_filename>
...
```

and the original (unpatched) `*.tn.tsv.gz` is renamed as `*.tn.unpatched.tsv.gz`. Note: this file is reported in output for inspection, but most use cases will use the (patched) `*.tn.tsv.gz` file downstream of normalization.

An example of a `*.tn.tsv.gz` file is shown below.

```
#title = Normalized coverage profile
#sex = UNDETERMINED
contig  start   stop    name    <SampleName> improper_pairs
chr1    818022  819840  target-wgs-chr1-818022:819840   -0.18479358083014644    6
chr1    819840  821337  target-wgs-chr1-819840:821337   -0.21244441644669046    19
chr1    821337  822485  target-wgs-chr1-821337:822485   -0.14849555308041734    10
chr1    822485  824431  target-wgs-chr1-822485:824431   -0.12423291178926463    7
chr1    830446  832304  target-wgs-chr1-830446:832304   -0.1438261733656668     1
```

For more information, see [Normalization](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#normalization).

### B-Allele Counts

In germline ASCN runs, B-allele counts are calculated at bi-allelic sites taken from a collection of high-frequency SNVs in the population. Each B-allele site consists of a reference allele and a variant allele, and the number of reads in the sample supporting each of these alleles is counted.

B-allele counts are written both to gzipped tsv file `*.ballele.counts.gz` and gzipped bedgraph file `*.baf.bedgraph.gz`.

`<prefix>.ballele.counts.gz`

Columns:

1. Contig identifier
2. Start, BED-style (zero-based inclusive) start position of the reference allele
3. Stop, BED-style (one-based inclusive) stop position of the reference allele
4. Base sequence for the reference allele
5. Base sequence for the first allele being counted
6. Base sequence for the second allele being counted
7. The number of qualified reads containing a sequence matching the first allele
8. The number of qualified reads containing a sequence matching the second allele
9. Population frequency for the first allele
10. Population frequency for the second allele

Example:

```
contig  start   stop    refAllele       allele1 allele2 allele1Count    allele2Count    allele1AF       allele2AF
chr1    51478   51479   T       T       A       4       2       0.6747  0.3253
chr1    82733   82734   T       T       C       111     36      0.79346 0.20654
chr1    83083   83084   T       T       A       0       0       0.1538  0.8462
chr1    86330   86331   A       A       G       9       9       0.87384 0.12616
chr1    88315   88316   G       G       A       0       0       0.8926  0.1074
```

`<prefix>.baf.bedgraph.gz`

B-allele frequency in bedgraph format. Allele count ratios are calculated by sorting alleles according to base priority {A, T, G, C} (descending), producing frequencies deterministically distributed above and below 0.5. This provides easy visualization in IGV of significant BAF changes between neighboring segments.

Example:

```
chr1    11021   11022   0.333333
chr1    14463   14464   0.755102
chr1    16494   16495   0.317708
chr1    38741   38742   0.5
chr1    39014   39015   0.44186
```

### Segmentation Results

`<prefix>.seg`

Contains the segments produced by the segmentation algorithm. The `Segment_Mean` value of a segment is the ratio of the mean of that segment to the whole-sample median, without log transformation (linear copy-ratio). A strong signal deviation from 1.0 indicates a potential for a CNV event.

The file has the following columns:

1. Sample name
2. Contig identified
3. Start position
4. End position
5. Number of intervals in the segment
6. Linear copy-ratio of the segment

An example of a `*.seg` file is shown below.

```
Sample  Chromosome      Start   End     Num_Probes      Segment_Mean
<SampleName> chr1    818022  1117426 224     0.82500341336435279
<SampleName> chr1    1117426 4063702 2438    0.91726081432236528
<SampleName> chr1    4063702 4067591 3       0.38861386123247205
<SampleName> chr1    4067591 7705829 3302    0.93021316913709917
<SampleName> chr1    7705829 9357003 1405    0.98147825043799442
<SampleName> chr1    9357003 9377365 19      0.50269670724395654
<SampleName> chr1    9377365 12859821        2905    1.0684818476332989
```

`<prefix>.baf.seg`

In addition to segmentation of target counts, some workflows perform segmentation of B-allele loci. The output file has suffix `*.baf.seg` and it has the same format of the `*.seg` file with two modifications. First, the `Segment_Mean` value is the mean over B-allele loci of the smaller observed allele fraction. Second, there is an additional column:

7. `BAF_SLM_STATE`: Integer between 0 and 10, indicating bins of minor-allele fraction (low to high), or `.` when the BAF data are too variable to estimate a minor-allele fraction

An example of BAF segmentation output file is shown below:

```
Sample  Chromosome      Start   End     Num_Probes      Segment_Mean    BAF_SLM_STATE
<SampleName> chr1    820348  1104646 194     0.29301737166888697     6
<SampleName> chr1    1105091 1533754 444     0.26185904799069076     5
<SampleName> chr1    1533810 1534166 9       0.41958837071702065     8
<SampleName> chr1    1534217 9356793 6689    0.26034515815016335     5
<SampleName> chr1    9358304 9376529 27      0.46450553586280602     10
```

### VCF Output

`<prefix>.cnv.vcf.gz`

The CNV VCF file follows the standard VCF format [v4.4](https://samtools.github.io/hts-specs/VCFv4.4.pdf). The VCF header is annotated with `##source=<DRAGEN_SOURCE>`, where `<DRAGEN_SOURCE>` identifies the caller which produced the VCF, e.g.:

* `DRAGEN_ASCN`: CNV caller
* `DRAGEN_ASCN_SV`: CNV caller + SV support
* `DRAGEN_CNV`: [legacy depth-only CNV caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy) (note: for legacy reasons this caller uses VCF version [v4.2](https://samtools.github.io/hts-specs/VCFv4.2.pdf))

Due to the nature of how CNV events are represented, not all fields are applicable. In general, if more information is available about an event, then the information is annotated. To include copy neutral (REF) calls, set `--cnv-enable-ref-calls` to true. AOH/LOH events are not available in the legacy depth-only caller.

#### Example Records

```bash
# Example REF call
chr1    819841  DRAGEN:REF:chr1:819841-6103865  N       .       1000    PASS
  END=6103865;REFLEN=5284025
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/0:2:1:1000:1000:2.00155:1.000775:1.000775:129.1:0.5:4544:10920:66,10:0.00368019

# Example copy-neutral LOH call
chr1    6104347 DRAGEN:CNLOH:chr1:6104348-6727324       N       <LOH>   1000    PASS
  END=6727324;REFLEN=622977;SVLEN=622977;LOHTYPE=AOH;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  1/1:2:0:1000:1000:1.9876:0.001988:0.993798:128.2:0.001:528:916:10,12:0.00766703

# Example GAIN call
chr1    16715826        DRAGEN:GAIN:chr1:16715827-16949283      N       <DUP>   744     PASS
  END=16949283;REFLEN=233457;SVLEN=233457;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:3:1:1000:99:3.08217:1.134239:1.541085:198.8:0.368:49:26:20,14:0.0384615

# Example GAIN LOH call
chr15   20212550        DRAGEN:GAINLOH:chr15:20212551-20421468  N       <LOH>   390     PASS
  END=20421468;REFLEN=208918;SVLEN=208918;LOHTYPE=AOH;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  1/1:6:0:1:1:5.90559:0.000000:2.952793:380.91:0:76:1:9,8:0

# Example LOSS call
chr1    25274774        DRAGEN:LOSS:chr1:25274775-25331683      N       <DEL>   226     PASS
  END=25331683;REFLEN=56909;SVLEN=56909;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:1:0:1000:1000:1.01085:0.000000:0.505426:65.2:0:7:10:5,1:0

# Example MOSAIC GAIN call
chr17   26781858        DRAGEN:GAIN:chr17:26781859-26940176  N  <DUP>   70      PASS
  END=26940176;CIPOS=-6985,1424;CIEND=-1519,1732;REFLEN=158318;MOSAIC;SVLEN=158318;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:MF:SM:SD:MAF:BC:AS:PE
  ./1:3:.:0:.:2.267890:.:0.268000:1.133945:123.6:.:78:0:17,118

# Example MOSAIC LOSS call
chr17   21884022        DRAGEN:LOSS:chr17:21884023-21988202  N  <DEL>   1000    PASS
  END=21988202;CIPOS=-1254,1259;CIEND=-1574,1412;REFLEN=104180;MOSAIC;SVLEN=104180;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:MF:SM:SD:MAF:BC:AS:PE
  0/1:1:0:0:1000:1.135780:.:0.864000:0.567890:61.9:.:79:0:80,348
```

#### Header

The following is an example of some of the header lines that are specific to CNV:

```
##fileformat=VCFv4.4
##ModelSource=DEPTH+BAF
##DiploidCoverage=371.000000
##OverallPloidy=1.998571
##HomozygosityIndex=0.001064
##OutlierBafFraction=0.024958
...
```

The following header lines are specific to the germline WGS ASCN caller:

| ID                 | Description                                                                                                                                                                                                                                                                                 |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| ModelSource        | The primary basis on which the final model was chosen. Value: `DEPTH+BAF`.                                                                                                                                                                                                                  |
| DiploidCoverage    | Expected read count for a target bin in a diploid region.                                                                                                                                                                                                                                   |
| OverallPloidy      | Length-weighted average of copy number for PASS events.                                                                                                                                                                                                                                     |
| OutlierBafFraction | A QC metric that measures the fraction of b-allele frequencies that are incompatible with the segment the BAFs belong to. High values might indicate substantial cross-sample contamination, or a different source of a mosaic genome, such as bone marrow transplantation. Range: \[0, 1]. |
| HomozygosityIndex  | Autosomal AOH/LOH percentage, considering only PASS AOH/LOH ≥ 2 Mb (default). Used as a proxy for consanguinity. A custom minimum size can be set through `--cnv-min-length-homozygosity-index`. The Cyto VCF (`*.cyto.vcf.gz`) also provides resolution-specific homozygosity indexes.     |

#### Records

All coordinates in the VCF are 1-based.

| ID    | Description                                                                                                                                                                                                                                                                                                                                                                                                                           |
| ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| CHROM | The chromosome (or contig) on which the copy number variant occurs.                                                                                                                                                                                                                                                                                                                                                                   |
| POS   | Start position of the variant. If any of the ALT alleles is a symbolic allele (e.g., `<DEL>`), POS denotes the coordinate of the base preceding the polymorphism.                                                                                                                                                                                                                                                                     |
| ID    | Encodes the event type and coordinates of the event (1-based, inclusive). Event types include `GAIN`, `LOSS`, `REF`, `CNLOH`, and `GAINLOH`.                                                                                                                                                                                                                                                                                          |
| REF   | Contains `N` for all CNV events.                                                                                                                                                                                                                                                                                                                                                                                                      |
| ALT   | Specifies the type of CNV event: `<DEL>`, `<DUP>`, or `<LOH>`. REF calls have ALT `.`. With `--cnv-enable-legacy-vcf-format` (VCF v4.2), the `ALT` field contains `<DEL>,<DUP>` in place of `<LOH>` for AOH/LOH events.                                                                                                                                                                                                               |
| QUAL  | Estimated quality score used in hard filtering. Note: different workflows provide different QUAL score distributions - it is recommended to compare QUAL scores only within results from the same workflow (e.g., it is incorrect to compare QUAL scores between the CNV caller and the [legacy (depth-only) CNV caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy)). |

#### FILTER

The FILTER column contains `PASS` if the CNV event passes all filters, otherwise the column contains the name of the failed filter. Default values are defined in the header line for each available FILTER.

| ID               | Description                                                                                                                                                                                                                                                |
| ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| binCount         | CNV events with a bin count lower than a threshold.                                                                                                                                                                                                        |
| chromArmBinCount | A whole-arm alteration call is based on a minimal portion (default 500 intervals) of the entire arm (e.g., in acrocentric chromosomes, where the short arm is mainly consisting of poor mappability regions, that are ignored during copy-number calling). |
| cnvLength        | The length of the CNV is lower than a threshold.                                                                                                                                                                                                           |
| cnvMosaicLength  | A MOSAIC call below a certain length has been filtered as candidate FP.                                                                                                                                                                                    |
| cnvQual          | The QUAL of the CNV is lower than a threshold.                                                                                                                                                                                                             |
| mosaicFraction   | The mosaic fraction of a CNV is below a defined threshold (`--cnv-filter-mosaic-fraction`). This filter is applied only to small CNVs with lengths shorter than the specified size threshold (`--cnv-filter-mosaic-fraction-max-length`, default: 200000). |

#### INFO

The INFO column contains information representing the event.

| ID      | Description                                                                                                                                                                                                                    |
| ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| REFLEN  | Length of the event.                                                                                                                                                                                                           |
| SVLEN   | Length of the event. Only present for non-REF records. Note: in VCF v4.2 format (enabled with `--cnv-enable-legacy-vcf-format`), `SVLEN` is a signed representation of `REFLEN` (e.g., a negative value indicates a deletion). |
| SVTYPE  | Always `CNV`. Only present for non-REF records.                                                                                                                                                                                |
| END     | End position of the event (1-based, inclusive).                                                                                                                                                                                |
| LOHTYPE | Type of loss of heterozygosity. Possible values: `AOH` (Absence of Heterozygosity).                                                                                                                                            |
| MOSAIC  | Tag identifying mosaic calls (if mosaic calling is enabled).                                                                                                                                                                   |
| CIPOS   | Confidence interval around the nominal `POS`.                                                                                                                                                                                  |
| CIEND   | Confidence interval around the nominal `END`.                                                                                                                                                                                  |

If using a segment BED file, then the segment identifier is carried over from the input to `SEGID` field.

When matching CNV with SV output, additional INFO annotations are added.

#### FORMAT

The common FORMAT fields are described in the header:

| ID   | Description                                                                                                                                                                                                |
| ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| GT   | Genotype                                                                                                                                                                                                   |
| SM   | Linear copy ratio of the segment mean                                                                                                                                                                      |
| CN   | Estimated total copy number of sample                                                                                                                                                                      |
| BC   | Number of read count bins                                                                                                                                                                                  |
| PE   | Number of improperly paired end reads at start and stop breakpoints                                                                                                                                        |
| AS   | Number of allelic read count sites                                                                                                                                                                         |
| CNF  | Floating point estimate of copy number                                                                                                                                                                     |
| CNQ  | Exact total copy number Q-score                                                                                                                                                                            |
| MAF  | Estimate for the minor allele frequency                                                                                                                                                                    |
| MCN  | Estimated minor-haplotype copy number                                                                                                                                                                      |
| MCNF | Floating point estimate of minor-haplotype copy number                                                                                                                                                     |
| MCNQ | Minor copy number Q-score                                                                                                                                                                                  |
| MF   | Mosaic fraction estimate (for MOSAIC calls)                                                                                                                                                                |
| OBF  | Per-segment Outlier BAF Fraction. Percentage of BAF counts which are considered "outlier" with respect to the chosen segment call. Higher values might indicate segments where BAF counts are problematic. |
| SD   | Best estimate of segment's bias-corrected read count                                                                                                                                                       |

For more information, see [CNV VCF](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#cnv-vcf-file).

### Cytogenetics Output

`<prefix>.cyto.vcf.gz`

The Cytogenetics modality output has a similar format to the standard CNV VCF (`*.cnv.vcf.gz`). A list of differences is indicated below:

* Records can have the `INFO/RES` field. In such case, such field indicate the resolution(s) associated with the record.
* Records can have the `INFO/SEGID` field. In such case, such field can either indicate custom predefined segments indicated in input by the user (similar to the standard CNV VCF), or Cytogenetics-specific predefined segments which are typically whole-arm/-chromosome segments automatically injected during the caller execution. In the latter case, the annotation field indicates the ID or name for the arm or chromosome.
* The VCF header is annotated with `##source=DRAGEN_CYTO` to indicate the file is generated by the Cytogenetics modality.

**Note:** The Cyto VCF also provides resolution-specific homozygosity indexes (i.e., computed on each specific resolution's callset). The default minimum size considered is the same as the main `HomozygosityIndex`, and for each resolution in output, there will be an additional header line on the Cyto VCF indicating the resulting metric, e.g., `##HomozygosityIndex(25k)=0.001015`.

### CNV Metrics Output

`<prefix>.cnv_metrics.csv`

DRAGEN CNV outputs metrics in CSV format. The following metrics are reported:

**Sex Genotyper**

| Metric           | Description                                                                               |
| ---------------- | ----------------------------------------------------------------------------------------- |
| Estimated sex    | Estimated sex of the case sample (and panel of normals samples if applicable).            |
| Confidence score | Range: \[0.0, 1.0]. If the sample sex is specified via `--sample-sex`, this value is 0.0. |

DRAGEN Sex Genotyper requires a minimum of 300 target intervals to confidently determine sex genotype; if the panel covers fewer intervals on the sex chromosomes, genotyping will fail and an undetermined genotype is returned. Users may lower this requirement by setting `--cnv-sex-genotyper-num-interval-requirement` to a smaller value, at the risk of increased false genotype calls.

**CNV Summary**

* Bases in reference genome in use
* Average alignment coverage over genome - The average alignment coverage over the genome is calculated by dividing the total number of bases from processed alignment records (excluding those filtered by the Target Counts stage in DRAGEN CNV) by the genome length. Alignment records are filtered taking into consideration duplicate marking status (if available), MAPQ, and mapping status.
* Number of alignment records processed
  * Number of filtered records (total)
  * Number of filtered records (due to duplicates)
  * Number of filtered records (due to MAPQ)
  * Number of filtered records (due to being unmapped)
* PMAD - Pairwise Median Absolute Deviation measures the variation in read coverage between adjacent bins. It measures variability due to various factors, such as DNA degradation, extraction, amplification or library preparation. Higher values indicate noisier sample data. PMAD is calculated as following:
  * Define a vector v\[i] as normalized counts of i-th interval in log scale, and d\[i] as pairwise differences of consecutive normalized counts between i and i+1 intervals, i.e. d\[i] = (v\[i] - v\[i+1])
  * PMAD is median absolute deviation of d, i.e. PMAD = Median(|d\[i]-Median(d)|)
* Coverage MAD - Median absolute deviation of normalized case counts. Higher values indicate noisier sample data.
* Median Bin Count - Median of raw counts normalized by interval size.
* Number of target intervals
* Number of normal samples
* Number of segments
* Number of amplifications - Note: GAINLOH events (ALT=LOH and CN > 2) are also included here
* Number of deletions
* Number of CNLOHs (Copy-Neutral LOHs)
* Number of PASS amplifications - Note: GAINLOH events (ALT=LOH and CN > 2) are also included here
* Number of PASS deletions
* Number of PASS CNLOHs (Copy-Neutral LOHs)
* Post-Normalization Bin Count Sigma - Standard deviation of post-PoN-normalization median-normalized coverage values.

Coverage MAD and Median Bin Count are only printed for WES germline/somatic CNV. Post-Normalization Bin Count Sigma is only printed when PoN normalization has been applied.

Example (not all metrics are shown):

```
SEX GENOTYPER,,<SampleName>,FEMALE,0.0000
CNV SUMMARY,,Bases in reference genome,3217346917
CNV SUMMARY,,Average alignment coverage over genome,34.5
CNV SUMMARY,,PMAD,0.031799
CNV SUMMARY,,Number of target intervals,2873465
CNV SUMMARY,,Number of segments,1247
CNV SUMMARY,,Number of amplifications,87
CNV SUMMARY,,Number of deletions,54
CNV SUMMARY,,Number of CNLOHs,12
CNV SUMMARY,,Number of PASS amplifications,65
CNV SUMMARY,,Number of PASS deletions,38
CNV SUMMARY,,Number of PASS CNLOHs,8
```

For more information, see [CNV Metrics](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#cnv-metrics-file).

### Track Files (IGV)

To generate additional equivalent bigWig and gff files, set the `--cnv-enable-tracks` option to true. These files can be loaded into IGV along with other tracks that are available, such as RefSeq genes. Using these tracks alongside publicly available tracks allows for easier interpretation of calls. DRAGEN autogenerates IGV session XML file if tracks are generated by DRAGEN CNV. The `*.cnv.igv_session.xml` can be loaded directly into IGV for analysis.

The following IGV tracks are automatically populated in the output IGV session file:

| Track File            | Description                                                                                                                                                                                                            | Recommended View   |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| \*.target.counts.bw   | BigWig representation of target counts bins. Values are GC-corrected if GC correction was performed.                                                                                                                   | Barchart or points |
| \*.improper\_pairs.bw | BigWig representation of improper pairs counts.                                                                                                                                                                        | Barchart           |
| \*.tn.bw              | BigWig representation of the tangent normalized signal.                                                                                                                                                                | Points             |
| \*.seg.bw             | BigWig representation of the segments.                                                                                                                                                                                 | Points             |
| \*.baf.seg.bw         | BigWig representation of BAF segments (if available).                                                                                                                                                                  | Points             |
| \*.baf.bedgraph.gz    | BED graph representation of B-allele frequency (if available).                                                                                                                                                         | Points             |
| \*.cnv.gff3           | GFF3 representation of CNV events: DEL=blue, DUP=red, filtered=light gray, REF=green (if enabled), AOH/LOH=magenta. An example is shown below (different workflows may output different attributes on the 9th column). | —                  |

Example GFF3 output:

```
##gff-version 3
chr1    DRAGEN  LOSS    12779193        12859821        30      .       .       Alt=DEL;LinearCopyRatio=0.576;CopyNumber=1;Genotype=0/1;Qual=30;Filter=PASS;Start=12779192;Stop=12859821;Length=80629;BinCount=24;ImproperPairsCount=16,7;color=#0000FF;
chr1    DRAGEN  REF     13106280        13122338        19      .       .       Alt=REF;LinearCopyRatio=1.05981;CopyNumber=2;Genotype=./.;Qual=19;Filter=PASS;Start=13106279;Stop=13122338;Length=16059;BinCount=8;ImproperPairsCount=3,1;color=#00FF00;
chr1    DRAGEN  GAIN    13225213        13247040        66      .       .       Alt=DUP;LinearCopyRatio=2.016;CopyNumber=4;Genotype=./1;Qual=66;Filter=PASS;Start=13225212;Stop=13247040;Length=21828;BinCount=9;ImproperPairsCount=7,5;color=#FF0000;
```

#### IGV Session

![](/files/WGi22zHxhS0kedSW46he)

File extension: `*.igv_session.xml`

The IGV session XML file is prepopulated with track files generated by DRAGEN. The session file loads the reference genome that best matches the standard reference genomes in an IGV installation, by comparing the name of the `--ref-dir` specified on the command-line. Standard UCSC human reference genomes are autodetected, but any variations from the standard reference genomes might not be autodetected. To edit the genome detection, alter the `genome` attribute in the `Session` element to the reference genome you would like for analysis before loading into IGV. The reference identifier used by IGV might differ from the actual name of the genome. The following is an example edited session file.

```
<?xml version="1.0" encoding="utf-8"?>
<Session genome="b37" hasGeneTrack="false" hasSequenceTrack="true" version="8">
    <Resources>
        <Resource path="example.cnv.gff3"/>
        <Resource path="example.cnv.excluded_intervals.bed.gz"/>
        <Resource path="example.target.counts.bw"/>
        <Resource path="example.improper.pairs.bw"/>
        <Resource path="example.tn.bw"/>
        <Resource path="example.seg.bw"/>
    </Resources>
    <Panel height="500" width="1200" name="DataPanel">
        ...
    </Panel>
</Session>
```

Note that depending on the IGV version installed, it may come prepackaged with different flavors of GRCh37. The reference naming conventions have changed so a user may have to edit the `genome` field in the XML file directly. For example, IGV has traditionally packaged a `b37` reference genome, but may also include a `1kg_v37` or a `1kg_b37+decoy`, which will appear on the IGV user interface as "1kg, b37" or "1kg, b37+decoy" respectively.

You can determine what the correct encoding of a reference genome by going to `File > Save Session...` and then inspecting the generated igv\_session.xml file.

When the Cytogenetics Modality is enabled, DRAGEN CNV produces an additional IGV session xml `*.cyto.igv_session.xml` shown below.

![](/files/4f11wpwXtvo98GriV38t)

## Advanced Topics

### Cytogenetics Modality

Conventional cytogenetics methodologies typically focus on larger alterations than the ones provided by NGS analyses. The Cytogenetics modality for the CNV caller allows the user to visualize variants or alterations at different resolutions, aiming at providing a more flexible workspace for different use cases.

It is enabled with `--cnv-enable-cyto-output` (default true for germline workflows). Not available for somatic WES workflows.

From the same sample, and during the same run, the Cytogenetics modality starts from the high resolution results (before smoothing) provided in the standard output CNV VCF. The output callset then undergoes multiple rounds of smoothing, going progressively from finer resolution to coarser resolution calls (larger alterations). Each round of smoothing produces a smoothed callset which is set aside and becomes the starting point for callsets with higher degree of smoothing.

![](/files/5gBu7zidRsCoAwQCYe6z)

At the end of the smoothing procedure, the Cytogenetics modality produces several outputs, e.g.:

* Multiple GFF3 files, one for each round of smoothing (extension `*cyto.<resolution_ID>.gff3`).
* A single VCF file, with extension `*.cyto.vcf.gz`. This file contains all callsets identified through the smoothing iterations, where the iteration identifier is stored on the `INFO/RES` field. Identical alterations across resolutions are deduplicated. In such case, the `INFO/RES` field will contain a comma-separated list of resolution identifiers.
  * Some resolutions will be based on depth of coverage only (no BAF). Their `INFO/RES` value will reflect the original callset used as a starting point, with added suffix `_depth`. E.g., for depth-only calls derived from resolution `1M`, the new callset will have resolution ID `1M_depth`. Note: calls made at different resolutions or with different information (depth+BAF versus depth-only) may occasionally conflict. For instance, in a region that is AOH that also has a mosaic DEL, the region may be reported as AOH for the depth+BAF calling but may be reported as (mosaic) DEL for the depth-only track. The event type with the strongest evidence will be output for each resolution.
  * An additional callset which does not conform to the ones above (no `INFO/RES` field) is the one containing whole-arm/-chromosome aneuploidies. For this callset, all reported records have the chromosome name or arm name in the `INFO/SEGID` field. Entries for this callset will not be present on any GFF3 file. For more details see the section on whole-chromosome aneuploidies below.
* A single IGV session file, with extension `*.cyto.igv_session.xml`, which provides a convenient way to load the multiple GFF3 files and other typical tracks found on the standard `*.cnv.igv_session.xml`. Below an example screenshot of one of such IGV sessions:
  * The first 5 tracks provide the DRAGEN CNV calls (Blue/DEL, Green/REF, Magenta/AOH, Red/DUP) at decreasing degree of resolution (from high to low, top to bottom).
  * The remaining tracks are similar to the standard `*cnv.igv_session.xml` run, e.g.: poor mappability regions, target counts coverage, improper pairs, B-allele frequency, etc.

![](/files/4f11wpwXtvo98GriV38t)

Below, an example set of calls from the `*.cyto.vcf.gz` output file (note additional `INFO/RES` annotation with respect to `*.cnv.vcf.gz` output file):

```
# Example REF call
chr1    819841  DRAGEN:REF:chr1:819841-6103865  N       .       1000    PASS
  END=6103865;REFLEN=5284025;RES=25k,500k,50k
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/0:2:1:1000:1000:2.00155:1.000775:1.000775:129.1:0.5:4544:10920:66,10:0.00368019

# Example GAIN call
chr1    16605768        DRAGEN:GAIN:chr1:16605769-16645359      N       <DUP>   427     PASS
  END=16645359;REFLEN=39591;RES=25k;SVLEN=39591;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE
  ./1:6:.:1:.:6.27065:.:3.135326:404.457:.:23:0:6,11

# Example LOSS call
chr1    25274774        DRAGEN:LOSS:chr1:25274775-25331683      N       <DEL>   226     PASS
  END=25331683;REFLEN=56909;RES=25k,50k;SVLEN=56909;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:1:0:1000:1000:1.01085:0.000000:0.505426:65.2:0:7:10:5,1:0
```

**Selection of appropriate resolution**

Since the most-informative resolution may vary depending on circumstances (event sizes, distance between calls, presence of smaller calls causing fragmentation, etc), no one-size-fits-all recommendation can work for all cases. However, some practical recommendations to consider are the following:

* Each resolution `INFO/RES` ID identifies the *minimum size* for alterations to be considered PASS.
* If only minimal call smoothing is necessary, resolution 25k can provide a good balance and provide calls in size ranges compatible with Chromosomal Microarray (CMA).
* When comparing against technologies such as karyotyping, resolution 1M may be the more appropriate to reduce call fragmentation.

Note: if the use case under consideration is not impacted by call fragmentation, it is typically recommended to use the `*.cnv.vcf.gz` or `*.cnv_sv.vcf.gz` output results (instead of the ones in `*.cyto.vcf.gz`), to take full advantage of the superior detail of NGS.

**Additional options**

| Option                                          | Description                                                                                    |
| ----------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| --cnv-cyto-keep-resolutions=\<resolution\_list> | Comma-separated list of resolutions to output (currently supported: 25k,50k,500k,1M,1M\_depth) |

**Whole-chromosome Aneuploidy Detection**

For some use cases, it is sometimes necessary to inspect a sample at arm or whole-chromosome level. Typically this would require the use of an additional caller, together with the standard CNV caller with automated segment detection. On the same run, the Cytogenetics modality provides such set of calls within the same VCF file (with extension `*.cyto.vcf.gz`).

```
chr21  12000000   DRAGEN:GAIN:chr21:12000001-46709983  N   <DUP>  1000  PASS
  END=46709983;REFLEN=34709983;SEGID=chr21q;SVLEN=34709983;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:3:1:1000:1000:3.00155:1.002518:1.500775:193.6:0.334:29570:66224:0,0:0.0016016

chrX   1        DRAGEN:LOSS:chrX:2-156040895     N     <DEL>  1000  PASS
  END=156040895;REFLEN=156040894;SEGID=chrX;SVLEN=156040894;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:1:0:1000:1000:0.996364:0.000996:0.498182:82.2:0.001:122580:144548:0,0:0.00995089
```

In the example above, two calls derived from such callset. The segment ID annotation (`INFO/SEGID`) provides the name for the segment call under consideration (i.e., for this example, q-arm of chromosome 21 and the entire chromosome X). REF calls are not displayed by default unless required explicitly by the user (i.e., with `--cnv-enable-ref-calls true`. Note: this will enable REF calls for both CNV and CYTO VCF files).

Note: acrocentric chromosomes (13, 14, 15, 21, and 22) have short arms characterized by repetitive regions. These regions create mappability issues and they are typically excluded from analysis. Thus, calling short arm alterations for these chromosomes is challenging, being based on a small percentage of total arm's length. To avoid false positive calls (in this case, indicating an alteration on the full short arm with evidence only coming from a minimal portion of it), the algorithm has a hard threshold (default 500 intervals) on the minimum number of intervals required when calling whole-arm alterations. When the chromosome arm call does not satisfy this threshold, the call is filtered with `FILTER` `chromArmBinCount`. The default can be changed with option `cnv-filter-chrom-arm-bin-count`.

### MOSAIC fraction estimation

For MOSAIC alterations, DRAGEN attempts inference of the mosaic fraction (`MF`), that is, the percentage of cells showing the alteration.

After copy number calling, the call's mosaic fraction is preliminarily estimated from the total and minor-allele copy-number (`CN`, `MCN`) and floating point estimates (`CNF`, `MCNF`). For example: in the case of `CN=4`, `CNF=4.48`, the `CN` of the population without the alteration is considered $$CN=5$$, and then the mosaic fraction preliminary estimate is $$MF=1-0.48=0.52$$.

The call observed `MAF` is then cross-checked with the expected `MAF`:

$$MAF=\frac{M\_1n\_1(1-q)+M\_2n\_2q}{n\_1(1-q)+n\_2q}$$

\* Note: this algorithm assumes only 2 cells populations (population 2: with the MOSAIC alteration called with `CN` and `MCN`, population 1: all remaining cells).

where:

* $$MAF$$ is the expected `MAF` of the mixture
* $$n\_1$$ and $$n\_2$$ represent the expected `CN` of the 2 cell populations
* $$M\_1$$ and $$M\_2$$ represent the expected `MAF` of the 2 cell populations
* $$q$$ denotes the mosaic fraction (aka `MF`, fraction of the 2nd cell population)

If the observed `MAF` is consistent with the expected `MAF` (considering a 5% tolerance on the `MF` value), the `MF` value is returned. Otherwise, the algorithm investigates alternative ($$n\_1$$, $$q$$) configurations that are compatible with $$n\_2$$ and the copy-number floating point estimate (`CNF`). If at least one alternative passes the expected `MAF` compatibility check, the updated $$q$$ is returned in the `MF` field. In all other cases, `MF=.`.

### Low-pass WGS support

The germline WGS caller supports reliable detection of CNVs from low-pass WGS data. Low-pass WGS is a highly cost-effective approach for CNV detection, providing genome-wide resolution at substantially lower cost than standard WGS or WES:

* Cost-effective CNV detection at low sequencing depth (1× to 10×)
* Comparable performance to WGS for cytogenetic-scale events (>1 Mb)
* Detects CNVs down to a few hundred kilobases
* Supports whole-chromosome aneuploidy and mosaic events

#### CNV Detection Capabilities

* **Variant types**:
  * Deletions
  * Duplications
* **Resolution tiers**:
  * Cytogenetic (coarse): ≥ 1 Mb
  * CNV (fine): 200 kb – 1 Mb
* **Minimum event size**:
  * 200 kb hard filter
* **B-allele frequency (BAF)**:
  * Not estimated in low-pass mode

#### Output Files

| Output File | Resolution           | Size Range    |
| ----------- | -------------------- | ------------- |
| cyto.vcf.gz | Coarse (cytogenetic) | ≥ 1 Mb        |
| cnv.vcf.gz  | Fine                 | 200 kb – 1 Mb |

#### Command-Line Usage

Enable low-pass CNV calling using the `--cnv-enable-lowpass=true` option:

```bash
dragen \
  -1 sample_R1.fastq.gz \
  -2 sample_R2.fastq.gz \
  --RGID RGID \
  --RGSM RGSM \
  --enable-map-align=true \
  --enable-map-align-output=true \
  --ref-dir=<REFERENCE> \
  --output-file-prefix=dragen \
  --output-directory=<OUTPUT_DIR> \
  --cnv-enable-lowpass=true
```

#### Example records

**CNV**

```
chr4       123918579       DRAGEN:LOSS:chr4:123918580-124314854       N       <DEL>   190     PASS
  END=124314854;CIPOS=-59095,63223;CIEND=-54689,56222;REFLEN=396275;SVLEN=396275;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:MF:SM:SD:MAF:BC:AS:PE
  0/1:1:0:1000:190:0.987000:.:.:0.493500:197.4:.:7:0:7,10
```

**Cytogenetics**

```
chrX       1       DRAGEN:GAIN:chrX:2-156040895       N       <DUP>   1000    PASS
  END=156040895;REFLEN=156040894;SEGID=chrX;SVLEN=156040894;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:MF:SM:SD:MAF:BC:AS:PE       1:2:0:1000:.:1.942077:.:.:1.942077:355.4:.:2472:0:0,0
```

**Mosaic events**

```
chr8       1       DRAGEN:GAIN:chr8:2-145138636       N       <DUP>   1000    PASS
  END=145138636;REFLEN=145138635;SEGID=chr8;MOSAIC;SVLEN=145138635;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:MF:SM:SD:MAF:BC:AS:PE
  ./1:3:.:0:.:2.272202:.:0.272000:1.136101:314.7:.:2530:0:0,0
```

#### Hard Filter Options

Low-pass CNV calling applies filters based on CNV length and bin count to reduce noise associated with low sequencing coverage.

| Option                 | Default | Description                         |
| ---------------------- | ------- | ----------------------------------- |
| --cnv-filter-length    | 200 kb  | Minimum CNV length for a PASS call. |
| --cnv-filter-bin-count | 4       | Minimum bin count for a PASS call.  |

### CNV with SV Support

The DRAGEN CNV caller leverages depth/BAF as its primary signal for calling copy number variants. CNV alone poses challenges for calling events that are less than 10kbp. The sensitivity of CNVs at lengths less than 10kbp can be improved by leveraging junction signals from the DRAGEN structural variant caller.

When both the DRAGEN CNV and SV caller are executed in a single invocation, then an additional integration step is done at the end of a DRAGEN run to improve the CNV calls. This feature is enabled automatically when DRAGEN detects a germline WGS analysis.

The SV/CNV Integration module takes in DEL and DUP calls from the output data structures of the germline CNV and SV callers, identifies putative matches, updates annotations, filters, scores, and outputs the refined records in CNV VCF. By leveraging junction signals from the SV caller and depth/BAF signals from the CNV caller, this approach allows for sensitive CNV detection down to 1kbp while also improving recall and precision across length scales. This is achieved by rescuing previously low quality calls if evidence is found from both callers, and also by adjusting CNV breakends to the more accurate SV breakends. The matching algorithm takes into account the proximity of the events as well as the transition states at the breakends, among other things.

#### Example command lines

The following is an example command line for running a germline WGS analysis for both CNV and SV.

```
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--bam-input <BAM> \
--enable-map-align false \
--enable-cnv true \
--cnv-population-b-allele-vcf <POP_SNP_VCF> \
--enable-sv true
```

Other optional CNV or SV parameters can also be added.

Note: There is a high sensitivity mode that can be enabled with `--sv-cnv-enable-high-sensitivity-mode=true`. This option is experimental and will disable many filters in the processing chain to allow for more SV+CNV calls to pass. It is recommended that users apply their own training and downstream filters when using this option.

#### VCF Output

CNV calls with SV support are output in the CNV VCF (`*.cnv.vcf.gz`). The VCF header includes all header information from the individual CNV and SV callers, with some header lines deduplicated and additional header lines added from SV/CNV integration. For details on the individual caller header lines, please refer to the CNV and SV sections of the user guide. In cases where users want to obtain a separate CNV/SV VCF file while keeping the original CNV and SV VCF outputs, they can specify `--sv-cnv-output-as-cnv-vcf=false`. CNV calls with SV support are then output in a separate CNV/SV VCF with the `*.cnv_sv.vcf.gz` extension. In this case, the original CNV and SV VCF files prior to integration are also available in the DRAGEN output directory, as described elsewhere.

Newly added header lines from SV/CNV integration are described in the following table.

| Header Field        | Number | Type    | Description                                                                                                         |
| ------------------- | ------ | ------- | ------------------------------------------------------------------------------------------------------------------- |
| END\_LEFT\_BND\_OF  | 1      | String  | ID of CNV whose left end is matched to the end of SV                                                                |
| END\_RIGHT\_BND\_OF | 1      | String  | ID of CNV whose right end is matched to the end of SV                                                               |
| LEFT\_BND           | 1      | String  | ID of SV that matches the left end of CNV record                                                                    |
| LEFT\_BND\_OF       | 1      | String  | ID of CNV whose left end is matched to SV                                                                           |
| MatchSv             | 1      | Integer | ID of original SV that was merged with CNV record                                                                   |
| OrigCnvEnd          | 1      | Integer | Coordinate of original CNV end                                                                                      |
| OrigCnvPos          | 1      | Integer | Coordinate of original CNV pos                                                                                      |
| RIGHT\_BND          | 1      | String  | ID of SV that matches the right end of CNV record                                                                   |
| RIGHT\_BND\_OF      | 1      | String  | ID of CNV whose right end is matched to SV                                                                          |
| SVCLAIM             | A      | String  | Claim made by the structural variant call. Valid values are D, J, DJ for abundance, adjacency and both respectively |

Records that can be matched or rescued will have annotations indicating the breakpoint linkage between a CNV and SV record. If a complete match is found, then the `MatchSv` annotation will be present in the record, indicating the SV record's `ID` field for this CNV record. In this case, BND notations refer to the merged record ID itself rather than the SV before merging. Furthermore, the use of the `SVCLAIM` field will indicate if the record has evidence arising from depth/BAF signal `D`, or junction signals `J`, or both `DJ`.

Because of the mixing of standalone SV records and CNV records, the FORMAT field may have different annotations. For details on the CNV or SV specific annotations, please refer to the individual CNV and SV user guide sections.

Records that can be matched or rescued will have FILTER set to PASS. The original FILTERs are retained for records that were not matched or rescued. For example, the `cnvLength` FILTER will still be applied to standalone CNV records (those with `SVCLAIM=D`).

Example records are shown below.

```
# Merged record, note presence of SVCLAIM=DJ and MatchSv
chr1    24478046        DRAGEN:LOSS:chr1:24478047-24480950      N       <DEL>   1000    PASS
  END=24480950;CIPOS=0,9;CIEND=0,9;REFLEN=2904;SVLEN=2904;SVTYPE=DEL;LEFT_BND=DRAGEN:LOSS:chr1:24478047-24480950;OrigCnvPos=24477572;RIGHT_BND=DRAGEN:LOSS:chr1:24478047-24480950;OrigCnvEnd=24481505;SVCLAIM=DJ;MatchSv=DRAGEN:DEL:3301:0:1:0:0:0;CIGAR=1M2904D;HOMLEN=9;HOMSEQ=CCACCACGC;RIGHT_BND_OF=DRAGEN:REF:chr1:21465759-24478046;LEFT_BND_OF=DRAGEN:LOSS:chr1:24478047-24480950;END_RIGHT_BND_OF=DRAGEN:LOSS:chr1:24478047-24480950;END_LEFT_BND_OF=DRAGEN:REF:chr1:24480951-25405351   GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:MF:SM:SD:MAF:BC:AS:PE:OBF:GQ:PL:PR:SR:SB:FS:MLQS:VF:VF1:VAF1:VF2:VAF2       1/1:1:0:15:15:1.212644:0.369856:.:0.606322:105.5:0.305:3:10:54,59:0:11:999,14,0:59,45:28,16:15,13,6,10:4.442:.:71,54:36,54:0.600000:35,54:0.606742
 
# CNV record that did not match, note presence of SVCLAIM=D
chr1    169257422       DRAGEN:LOSS:chr1:169257423-169273108    N       <DEL>   277     PASS
  END=169273108;CIPOS=-7669,1182;CIEND=-1265,10511;REFLEN=15686;SVLEN=15686;SVTYPE=CNV;SVCLAIM=D  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:MF:SM:SD:MAF:BC:AS:PE:OBF   0/1:1:0:1000:1000:1.067816:0.000000:.:0.533908:92.9:0:14:38:1,5:0
  
# SV record that did not match, note presence of SVCLAIM=J
chr1    16744928        DRAGEN:LOSS:chr1:16744929-16746692      N       <DEL>   999     PASS
  END=16746692;SVTYPE=DEL;SVLEN=1764;CIGAR=1M1764D;CIPOS=0,24;CIEND=0,24;HOMLEN=24;HOMSEQ=GCCAACATGGTGAAACCCTGTCTC;SVCLAIM=J GT:GQ:PL:PR:SR:SB:FS:MLQS:VF:VF1:VAF1:VF2:VAF2  1/1:4:999,6,0:22,21:49,33:22,27,18,15:3.011:.:63,43:38,43:0.530864:25,43:0.632353
```

### Coverage Uniformity

The DRAGEN CNV pipeline provides a measure of the quality of the data for a sample. If using the WGS self-normalization method, the additional `CoverageUniformity` metric is present in the VCF header. The CNV pipeline assumes that post-normalization target counts are independently and identically distributed (IID). Coverage in most high-quality WGS samples is uniform enough for the CNV caller to produce accurate calls, but some samples violate the IID assumption. Issues during library preparation or sample contamination can lead to several extreme outliers and/or waviness of target counts, which can result in a large number of false positive CNV calls. The `CoverageUniformity` metric quantifies the degree of local coverage correlation in the sample to help identify poor-quality samples.

A larger value for this metric means the coverage in a sample is less uniform, which indicates that the sample has more nonrandom noise, and could be considered poor quality. The CoverageUniformity metric depends on factors other than sample quality, such as the `cnv-interval-width` setting and sample mean coverage. It is recommended to use this score to compare the quality of samples from similar mean coverage and the same command line options. Because of this, DRAGEN CNV only provides the metric and does not take any action based on it.

### Call Smoothing

The segmentation stage might produce adjacent or nearby segments that are assigned the same copy number and have similar depth and BAF data. This segmentation can result in a region with consistent true copy number being fragmented into several pieces. The fragmentation might be undesirable for downstream use of copy number estimates. Also, for some uses, it can be preferable to smooth short segments that would be assigned different copy numbers whether due to a true copy number change or an artifact. To reduce undesirable fragmentation, initial segments can be merged during a postcalling segment smoothing step.

After initial calling, segments shorter than the specified value of `--cnv-filter-length` are deemed negligible. Among the remaining nonnegligible segments, successive pairs are evaluated for merging. The caller combines two successive segments that are within `--cnv-merge-distance` of one another and have the same CN and MCN assignments, along with any intervening negligible segments into a single segment that is recalled and rescored. If the merged segment receives the same CN and MCN as its constituent nonneglible pieces with a sufficiently high-quality score, the original segments are replaced with the merged segment. The merged segment might be further merged with other initial or merged segments to either side. Merging proceeds until all segment pairs that meet the criteria are merged.

### QUAL Model

QUAL estimation is based on a model associated with the most likely diploid coverage estimated from depth of coverage and B-allele frequency.

Given such diploid coverage, for each segment, the algorithm calls the most likely copy number state (complete with total copy number CN, and minor allele copy number MCN).

The probability of the REF state is used in input to the scoring algorithm which outputs the QUAL value (a PHRED score capped at 1000). The QUAL value is the PHRED score where the probability of error is the probability of REF when an alteration is called, or the probability of having a non-REF call when the segment should be called REF.

Note: this is different from how QUAL is computed in the legacy [depth-only caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy).

## Comparison with ROH caller

Both the [ROH caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/small-variant-calling/roh-caller) and the germline CNV caller can detect runs-of-homozygosity (ROH) regions.

The two algorithms underlying the two different approaches might occasionally disagree. The differences are due to the following:

* The ROH caller requires minor-allele frequency to be \~0. In contrast, the germline CNV caller will assign to each segment its most likely copy-number state. This can include MOSAIC alterations, not available in the ROH caller.
* The ROH caller is dependent on the small variant caller, and only uses the SNPs that it calls. In contrast, the germline CNV caller works with a catalog of SNPs from population variation studies, such as 1000 Genomes.
* The ROH caller uses a blacklist bed file to filter certain sites and reduce call fragmentation. In contrast, the germline CNV caller does not need to filter any site but provides an alternative smoothing algorithm to reduce call fragmentation, which is agnostic on the sample under consideration.
* The ROH caller identifies ROH regions but does not provide the total copy number of the region under consideration. In contrast, the germline CNV caller also reports the copy number for the region (which could be different from reference ploidy).

## Limitations

The following features (available in the [depth-only workflow](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy)) are not yet supported:

* Multisample/Pedigree mode

## Multisample Germline CNV Calling

Multisample Germline CNV calling is possible starting from tangent normalized counts files (`*.tn.tsv.gz`) specified with the `--cnv-input` option (one per sample). Multisample CNV analysis benefits from using joint segmentation to increase the sensitivity of detection of copy number variable segments. For each copy number variable segment identified, the copy number genotype of each sample is emitted in a single VCF entry to facilitate annotation and interpretation.

Multisample Germline CNV analysis is supported for [legacy (depth-only) WGS and WES workflows](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy).

### Example command lines

The following is an example command line for running a trio analysis:

```
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-cnv true \
--cnv-input <FATHER_TN_TSV> \
--cnv-input <MOTHER_TN_TSV> \
--cnv-input <PROBAND_TN_TSV> \
--pedigree-file <PEDIGREE_FILE>
```

### De Novo CNV Calling Options

Make sure all input samples have gone through the same single sample workflow and have identical intervals. If the samples are WES inputs, then you must generate the samples using the same panel of normals, and the autosomal intervals for all samples must match.

The following options are used in DeNovo CNV calling:

| Option                    | Description                                                                                                                      |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| --cnv-input               | Input tangent-normalized signal files (`*.tn.tsv.gz`) from single sample runs. Can be specified multiple times, once per sample. |
| --cnv-filter-de-novo-qual | Phred-scaled threshold for calling an event as de novo in the proband. Default: `0.125`.                                         |
| --pedigree-file           | Pedigree file specifying the relationship between input samples.                                                                 |

### Joint Segmentation

First, CNV calling is performed on each sample independently. Joint segmentation then uses the copy number variable segments from each single sample analysis to derive a set of joint copy number variable segments. This set of joint segments is determined simply by taking the union of all breakpoints from the copy number variable segments of all samples. This results in the splitting of any partially overlapping segments across different samples. For example:

![](/files/FZz3VkDBHAJK1lcxA8gD)

Following joint segmentation, copy number calling is again performed independently on each sample using the joint segments. Segments can be merged as with the single sample analysis, but each joint segment is emitted in the multisample VCF as a single entry. The quality score (`QS` in the VCF) from the sample's merged segment, if applicable, is used for filtering the call. Sample calls are filtered using the sample's FT field in the multisample VCF. The `QUAL` column of the multisample VCF is always missing (ie, "."). The `FILTER` column of the multisample VCF is `SampleFT` if none of the sample's `FT` fields are `PASS`, and `PASS` if any of the sample's `FT` fields are `PASS`.

Note, however, that when a single segment in one sample overlaps multiple segments in another sample, the larger segment annotation is replicated across multiple records, e.g. (only relevant VCF fields are printed below):

```
DRAGEN:REF:chr22:21917617-22385563	GT:SM:CN:BC:PE:QS:FT	./.:1.01773:2:867:0,0:62:PASS	./.:1.00693:2:379:0,0:61:PASS
DRAGEN:LOSS:chr22:22385564-22549952	GT:SM:CN:BC:PE:QS:FT	./.:1.01773:2:867:0,0:62:PASS	0/1:0.695867:1:135:0,0:7:cnvQual
DRAGEN:LOSS:chr22:22549953-23041393	GT:SM:CN:BC:PE:QS:FT	./.:1.01773:2:867:0,0:62:PASS	0/1:0.614398:1:341:0,0:40:PASS
DRAGEN:LOSS:chr22:23041394-23055519	GT:SM:CN:BC:PE:QS:FT	./.:1.01773:2:867:0,0:62:PASS	0/1:0.31226:1:141:0,0:52:PASS
DRAGEN:LOSS:chr22:23055520-23198595	GT:SM:CN:BC:PE:QS:FT	0/1:0.57652:1:168:0,0:41:PASS	0/1:0.31226:1:141:0,0:52:PASS
DRAGEN:LOSS:chr22:23198596-23241095	GT:SM:CN:BC:PE:QS:FT	0/1:0.57652:1:168:0,0:41:PASS	1/1:0.128:0:39:0,0:42:PASS
```

The previous can be visualized as:

![](/files/MEyPDktJOTAyP7Hr1JbA)

### De Novo Calling Stage

A de novo event is defined as the existence of a genotype at a particular locus in a proband's genome that did not result from standard Mendelian inheritance from the parents. The de novo calling stage identifies putative de novo events in the proband of each trio of a multisample analysis. In some cases, these putative de novo events may be real, but they can also arise from sequencing or analysis artifacts. Consequently, a de novo quality score is assigned to each putative de novo event and used to filter out low-quality de novo events. Trios are specified by specifying a .ped file with the `--pedigree-file` option. Multiple trios can be specified (eg, quad analysis), and all valid trios will be processed.

For each joint segment in a trio, the de novo caller determines if there is a Mendelian inheritance conflict for the called copy number genotypes. The CNV caller does not identify the copy number for each allele of a given diploid segment, which means assumptions are made about the possible allelic composition of the parent genotypes.

The assumption is that the copy number 0 allele is not present for diploid regions of a parent's genome (sex dependent) when the assigned copy number is 2 or greater. This results in simplifications, as follows:

| Parent Copy Number Genotype | Possible Copy Number Alleles | Assumed Possible Copy Number Alleles |
| --------------------------- | ---------------------------- | ------------------------------------ |
| 2                           | 0/2, 1/1                     | 1/1                                  |
| 3                           | 0/3, 1/2                     | 1/2                                  |
| 4                           | 0/4, 1/3, 2/2                | 1/3, 2/2                             |
| N                           | x/(N-x) for x <= N/2         | x/(N-x) for 1 <= x <= N/2            |

The following are examples of consistent and inconsistent copy number genotypes for diploid regions using these assumptions:

| Mother Copy Number | Father Copy Number | Proband Copy Number | Mendelian Consistent? |
| ------------------ | ------------------ | ------------------- | --------------------- |
| 2                  | 2                  | 2                   | Yes                   |
| 2                  | 2                  | 1                   | No                    |
| 3                  | 2                  | 4                   | No                    |
| 3                  | 2                  | 2                   | Yes                   |
| 2                  | 0                  | 2                   | No                    |

If a joint segment has a Mendelian inheritance conflict, a Phred-scaled de novo quality score (`DQ` field in the VCF) is calculated using the likelihoods for each copy number state (see Quality Scoring section) of each sample in the trio, combined with a prior for the trio genotypes:

$$DQ = -10\log \left( \frac{1-\sum\_C{p(CN\_m|data) \cdot p(CN\_f|data) \cdot p(CN\_p|data) \cdot p(CN\_m,CN\_f,CN\_p)}}{\sum\_G{p(CN\_m|data) \cdot p(CN\_f|data) \cdot p(CN\_p|data) \cdot p(CN\_m,CN\_f,CN\_p)}} \right)$$

Where:

* $$G$$ is the set of all genotypes
* $$C$$ is the set of conflicting genotypes
* $$CN\_m$$ is the Mother copy number
* $$CN\_f$$ is the Father copy number
* $$CN\_p$$ is the Proband copy number
* $$p(CN\_m,CN\_f,CN\_p)$$ is the prior for the trio genotype

The `DN` field in the VCF is used to indicate the de novo status for each segment. Possible values are:

* `Inherited` - the called trio genotype is consistent with Mendelian inheritance
* `LowDQ` - the called trio genotype is inconsistent with Mendelian inheritance and DQ is less than the de novo quality threshold (default 0.125)
* `DeNovo` - the called trio genotype is inconsistent with Mendelian inheritance and `DQ` is greater than or equal to the de novo quality threshold (default 0.125)

### Multisample CNV VCF Output

The records in a multisample CNV VCF differ slightly from the single sample case. The major differences are as follows:

The per-record entries are broken down into the segments among the union of all the input samples breakpoints, which means there are more entries in the overall VCF.

The `QUAL` column is not used and its value is ".". The per-sample quality is carried over into the `SAMPLE` columns with the `QS` tag.

The `FILTER` column indicates `PASS` if any of the individual `SAMPLE` columns `PASS`. Otherwise, it indicates `SampleFT`.

The per-sample annotations are carried over from their originating calls. The single sample filters are applied at the sample level and are emitted in the `FT` annotation.

Additionally, if a valid pedigree is used, then de novo calling is performed, which adds the following two annotations to the proband sample.

```
##FORMAT=<ID=DQ,Number=1,Type=Float,Description="De novo quality">
##FORMAT=<ID=DN,Number=1,Type=String,Description="Possible values are `Inherited', 'DeNovo' or 'LowDQ'. Threshold for a passing de novo call is DQ > 0.125000">
```

While the VCF contains many entries, due to the joint segmentation stage, the number of de novo events can be found by extracting entries that have a `DN` and `DQ` annotation. These records are also extracted and are converted to GFF3 in the de novo calling case.

### Chromosome X and Y behavior

The sample sex from the single sample analysis (either estimated or overriden using the `--sample-sex` option) can be overriden by specifying the sex in the input pedigree file (i.e. 1 for male and 2 for female). To use the sample sex from the single sample analysis an unknown sex can be specified in the pedigree file using the value 0 (rather than 1 for male or 2 for female).

Note that when all samples in the pedigree are female, then no calls on chrY will be emitted for any sample. When the pedigree includes at least one male sample, only the male samples will have genotype info reported in the VCF for chrY and any VCF entries on chrY will have a "missing" Genotype column (i.e. ".") for all corresponding female samples in the pedigree.


# Germline Depth-Only

The DRAGEN CNV pipeline also supports a legacy depth-only mode that analyzes read depth coverage to identify copy number variations without using B-allele frequency information. This workflow follows the same processing steps as the ASCN workflow (binning, bias correction, normalization, segmentation, and calling) but operates solely on coverage depth signals.

## Key Differences from ASCN Workflow

The depth-only workflow differs from the germline ASCN workflow in the following ways:

* **No B-allele analysis**: Does not require `--cnv-population-b-allele-vcf` or generate B-allele count files
* **No allele-specific information**: Cannot distinguish between copy-neutral LOH and diploid regions, or determine minor allele copy numbers
* **Limited output**: Does not produce `*.ballele.counts.gz`, `*.baf.bedgraph.gz`, `*.baf.seg`, or cytogenetics VCF files
* **VCF format**: Uses VCF v4.2 instead of v4.4 by default
* **No mosaic detection**: Mosaic alteration calling is not available
* **Different quality metrics**: QUAL scores and filtering thresholds differ from ASCN workflows

## QUAL

Quality scores are computed using a probabilistic model that uses a mixture of heavy tailed probability distributions (one per integer copy number) with a weighting for event length. Noise variance is estimated. The output VCF contains a Phred-scaled metric that measures confidence in called amplification (CN > 2 for diploid locus), deletion (CN < 2 for diploid locus), or copy neutral (CN=2 for diploid locus) events.

The scoring algorithm also calculates exact copy-number quality scores that are inputs to the DeNovo CNV detection pipeline.

## Example Command Lines

**WGS with self-normalization:**

```bash
dragen \
  -r <HASHTABLE> \
  --output-directory <OUTPUT> \
  --output-file-prefix <SAMPLE> \
  --enable-map-align false \
  --enable-cnv true \
  --bam-input <BAM> \
  --cnv-enable-self-normalization true \
  --sample-sex <SEX>
```

**WES with panel of normals:**

```bash
dragen \
  -r <HASHTABLE> \
  --output-directory <OUTPUT> \
  --output-file-prefix <SAMPLE> \
  --enable-map-align false \
  --bam-input <BAM> \
  --sample-sex <SEX> \
  --enable-cnv true \
  --cnv-target-bed <CNV_TARGET_BED> \
  --cnv-combined-counts <CNV_PANEL_OF_NORMALS>
```

**Note:** To explicitly enable depth-only mode when using DRAGEN v4.5 or later, omit the `--cnv-population-b-allele-vcf` option.

## Depth-Only Specific VCF Annotations

The depth-only (non-ASCN) workflow produces VCF files in [VCF v4.2](https://samtools.github.io/hts-specs/VCFv4.2.pdf) format with the following specific annotations.

### Header

The VCF header is annotated with `##source=DRAGEN_CNV` to indicate the file is generated by the legacy depth-only CNV caller.

### Available FILTERs

The depth-only workflow applies the following filters to identify potentially false positive CNV calls:

* `cnvBinSupportRatio` - For CNVs greater than 80kb, indicates the percent span of supporting target intervals is lower than a threshold.
* `cnvCopyRatio` - Indicates that the segment mean of the CNV is not far enough from copy neutral.
* `cnvLength` - Indicates that the length of the CNV is lower than a threshold.
* `cnvLikelihoodRatio` - Indicates a log10 likelihood ratio of ALT to REF is less than a threshold.
* `cnvQual` - Indicates that the QUAL of the CNV is lower than a threshold.
* `dinucQual` - Applied based on the percentage of bases in a segment that belong to a two-base set (GC, CT, or AC), determined by individual occurrences. A CNV call is filtered out if any of these percentages fall outside typical ranges, indicating a likely false positive.
* `highCN` - Indicates a CNV call with implausible copy number (>6). Note: this filter is only applied in WGS depth-only workflows.

### INFO Fields

Depth-only workflows include the following INFO fields:

| ID    | Description                         |
| ----- | ----------------------------------- |
| `GCP` | Percentage of bases that are G or C |
| `CTP` | Percentage of bases that are C or T |
| `ACP` | Percentage of bases that are A or C |

These dinucleotide composition metrics are used by the `dinucQual` filter to identify false positive calls based on unusual base composition patterns.

### FORMAT Fields

The depth-only workflow provides the following common FORMAT fields:

| ID   | Description                                                         |
| ---- | ------------------------------------------------------------------- |
| `GT` | Genotype                                                            |
| `SM` | Linear copy ratio of the segment mean                               |
| `CN` | Estimated copy number                                               |
| `BC` | Number of bins in the region                                        |
| `PE` | Number of improperly paired end reads at start and stop breakpoints |

Germline depth-only CNV also includes:

| ID   | Description                          |
| ---- | ------------------------------------ |
| `LR` | Log10 likelihood ratio of ALT to REF |

The likelihood ratio is computed using the same probabilistic model described in [QUAL](#qual). To calculate the likelihood ratio, bins within each segment are evaluated on a per-bin basis. For each bin, the probabilities of amplification (CN > ploidy), deletion (CN < ploidy), and copy-neutral (CN = ploidy) states are calculated. The log10 likelihood ratios are then computed as `log10(p(amplification) / p(neutral))` for amplification, and `log10(p(deletion) / p(neutral))` for deletion. The segment-level likelihood ratio is obtained by summing the per-bin log10 likelihood ratios across all bins in the segment. For copy-neutral calls, the likelihood ratio is set to `0`.

Note that depth-only workflows do not provide allele-specific annotations such as `MCN`, `MCNQ`, `MCNF`, `MAF`, `AS`, or `OBF`, which are exclusive to ASCN workflows.

### Segmentation Output Files

The depth-only workflow generates additional segmentation output files:

* `*.seg.called` — Identical to `*.seg` with an additional column indicating the initial call (`+` for duplication or `-` for deletion)
* `*.seg.called.merged` — Similar to `*.seg.called` with potentially merged segments, plus additional columns for QUAL, FILTER, copy number assignment, ploidy, and improper pairs count

## VCF Format Differences

Different workflows or modalities can use a different version of the VCF specs:

* CNV caller: [VCF v4.4](https://samtools.github.io/hts-specs/VCFv4.4.pdf)
* Legacy (depth-only) CNV caller: [VCF v4.2](https://samtools.github.io/hts-specs/VCFv4.2.pdf)

**General**

| Field        | VCF v4.2             | VCF v4.4        |
| ------------ | -------------------- | --------------- |
| `INFO/SVLEN` | Positive or Negative | Always Positive |

**Absence/Loss of Heterozygosity (AOH/LOH)** - only present when emitting ASCN calls with legacy VCF v4.2

| Field       | VCF v4.2      | VCF v4.4 |
| ----------- | ------------- | -------- |
| `ALT`       | `<DEL>,<DUP>` | `<LOH>`  |
| `FORMAT/GT` | `1/2`         | `1/1`    |


# Somatic

## Overview

DRAGEN provides somatic copy number variant (CNV) calling workflows that detect copy number aberrations and regions with loss of heterozygosity (LOH) in whole genome sequencing (WGS) and whole exome sequencing (WES) data. The CNV workflows leverage both depth of coverage and B-allele frequencies (BAFs) to provide comprehensive detection of:

* Copy number gains (duplications) and losses (deletions)
* Copy-neutral loss of heterozygosity (CNLOH)
* Subclonal alterations (WGS only, enabled by default)
* Minor allele copy number estimation

## Workflow

The DRAGEN somatic CNV workflow follows this processing pipeline: ![](/files/EixZ9C7uQKuUdfIaMdKm)

The pipeline consists of the following modules:

1. **Target Counts** — Binning of read counts and other signals from alignments
2. **B-Allele Counts** — Extraction of allelic read counts
3. **Bias Correction** — Correction of GC bias and other systematic biases
4. **Normalization** — Detection of normal ploidy levels and normalization
5. **Segmentation** — Breakpoint detection via segmentation of normalized depth and BAF signals
6. **Allele Specific Copy Number (ASCN) Calling** — Integration of depth and BAF segments to determine copy number states and allele-specific information

**B-Allele Frequency Inputs**: The pipeline supports multiple input options for estimating B‑allele frequencies (BAF), depending on the availability of a matched normal sample.

* Matched normal already processed
  * If the matched normal sample has been processed with the germline small variant caller, the resulting VCF file can be provided directly.
* Matched normal not yet processed
  * If the matched normal has not been processed, the user may provide raw reads or aligned reads and enable concurrent execution of the germline small variant caller.
  * In this case, DRAGEN CNV consumes the small variant caller output to estimate B‑allele frequencies from germline SNVs.
* No matched normal available
  * A population SNV VCF may be provided.
  * DRAGEN estimates B‑allele frequencies using variants from the population SNV VCF.

For WES, population SNVs are intersected with the regions defined in cnv-target-bed. The target BED file must contain the same target intervals used to generate the PON.

**Depth-Only Workflow (Legacy)**: For applications that require only fold-change detection without purity/ploidy model estimation, a legacy depth-only workflow is also available for WES and targeted panels. See [Depth-Only Workflow](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy) for details.

## Example Command Lines

### WGS — Tumor-Normal (concurrent SNV caller)

If the matched normal has not been pre-processed, you can run the somatic SNV caller concurrently with CNV, which feeds germline heterozygous sites directly to the CNV caller:

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--tumor-bam-input <TUMOR_BAM> \
--bam-input <NORMAL_BAM> \
--enable-variant-caller true \
--cnv-use-somatic-vc-baf true
```

To additionally enable [Germline-aware Mode](#germline-aware-mode) and [VAF-aware Mode](#vaf-aware-mode), add the following flags:

```bash
--cnv-normal-cnv-vcf <CNV_NORMAL_VCF>   # germline-aware mode
# VAF-aware mode is enabled by default for tumor/matched-normal runs with --enable-variant-caller true
```

### WGS — Tumor-Only (population SNP VCF)

If no matched normal is available, run in tumor-only mode using a population SNP catalog:

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--tumor-bam-input <TUMOR_BAM> \
--cnv-population-b-allele-vcf <POP_SNP_VCF>
```

### WES — Tumor-Normal (concurrent SNV caller)

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--tumor-bam-input <TUMOR_BAM> \
--bam-input <NORMAL_BAM> \
--enable-cnv true \
--enable-variant-caller true \
--cnv-use-somatic-vc-baf true \
--cnv-normals-list <PANEL_OF_NORMALS> \
--cnv-target-bed <TARGET_BED> \
--vc-target-bed <TARGET_BED>
```

### WES — Tumor-Only (population SNP VCF)

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--tumor-bam-input <TUMOR_BAM> \
--cnv-population-b-allele-vcf <POP_SNP_VCF> \
--cnv-normals-list <PANEL_OF_NORMALS> \
--cnv-target-bed <TARGET_BED>
```

## Required Options

| Option         | Description                           |
| -------------- | ------------------------------------- |
| `--enable-cnv` | Enable CNV processing (set to `true`) |

### Input Options

**DNA inputs**

| Option                             | Description                                            |
| ---------------------------------- | ------------------------------------------------------ |
| `--tumor-fastq1`, `--tumor-fastq2` | FASTQ input files (requires `--enable-map-align true`) |
| `--tumor-bam-input`                | BAM input file                                         |
| `--tumor-cram-input`               | CRAM input file                                        |

**B-Allele inputs**

| Option                          | Description                                                                                                              |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `--cnv-normal-b-allele-vcf`     | Specify a matched normal SNV VCF.                                                                                        |
| `--cnv-population-b-allele-vcf` | Specify a population SNP catalog.                                                                                        |
| `--cnv-use-somatic-vc-baf`      | If running in tumor-normal mode with the SNV caller enabled, use this option to specify the germline heterozygous sites. |

For more information on specifying b-allele loci, see [Specification of B-Allele Loci](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#b-allele-counts).

**PON inputs**

| Option                  | Description                                                                                                                                                        |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `--cnv-target-bed`      | BED file defining exome capture regions (only for WES)                                                                                                             |
| `--cnv-normals-file`    | Specify individual normal counts file (target.counts.gz or target.counts.gc-corrected.gz) for PON. You can use this option multiple times, one time for each file. |
| `--cnv-normals-list`    | Specify text file that contains paths to the list of reference target counts files to be used as a panel of normals (new line separated).                          |
| `--cnv-combined-counts` | Specify combined PON file (.combined.counts.txt.gz).                                                                                                               |

**Other inputs**

| Option               | Description                                                                        |
| -------------------- | ---------------------------------------------------------------------------------- |
| `--ref-dir`          | DRAGEN reference genome hashtable directory                                        |
| `--enable-map-align` | Enable mapper and aligner module                                                   |
| `--sample-sex`       | Sample sex (e.g., `male`, `female`). If not specified, sex is estimated from data. |

#### Pop SNP download

Population VCF files can be downloaded from link below:

| Reference                            | Size  | Download                                                                                                                                 |
| ------------------------------------ | ----- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| hg38 CNV Population SNP VCF v1.0     | 1.8GB | [Download](https://webdata.illumina.com/downloads/software/dragen/resource-files/misc/hg38_1000G_phase1.snps.high_confidence.vcf.gz)     |
| hg19 CNV Population SNP VCF v1.0     | 1.8GB | [Download](https://webdata.illumina.com/downloads/software/dragen/resource-files/misc/hg19_1000G_phase1.snps.high_confidence.vcf.gz)     |
| hs37d5 CNV Population SNP VCF v1.0   | 1.8GB | [Download](https://webdata.illumina.com/downloads/software/dragen/resource-files/misc/hs37d5_1000G_phase1.snps.high_confidence.vcf.gz)   |
| CHM13-v2 CNV Population SNP VCF v1.0 | 4.0GB | [Download](https://webdata.illumina.com/downloads/software/dragen/resource-files/misc/chm13_v2_1000G_phase1.snps.high_confidence.vcf.gz) |

> Download links in the table opens an external Illumina download page: <https://support.illumina.com/sequencing/sequencing\\_software/dragen-bio-it-platform/product\\_files.html>

### Output Options

| Option                     | Description                                                                        |
| -------------------------- | ---------------------------------------------------------------------------------- |
| `--output-directory`       | Output directory for all results                                                   |
| `--output-file-prefix`     | Prefix prepended to all output file names                                          |
| `--cnv-enable-cyto-output` | Enable cytogenetics-compatible output VCF (default false) - only available for WGS |

### Target Counting Options

| Option                              | Description                                                                                                                                                                                                                                                                                                                                                                |
| ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-counts-method`               | Specifies the counting method for an alignment to be counted in a target bin. Values are midpoint, start, or overlap. The default value is overlap when using the panel of normals approach, which means if an alignment overlaps any part of the target bin, the alignment is counted for that bin. In the self-normalization mode, the default counting method is start. |
| `--cnv-min-mapq`                    | Specifies the minimum MAPQ for an alignment to be counted during target counts generation. The default value is 3 for self-normalization and 20 otherwise. When generating counts for panel of normals, all MAPQ0 alignments are counted.                                                                                                                                  |
| `--cnv-target-bed`                  | Specifies a properly formatted BED file that indicates the target intervals to sample coverage over. For use in WES analysis.                                                                                                                                                                                                                                              |
| `--cnv-interval-width`              | Specifies the width of the sampling interval for CNV processing. This option controls the effective window size. The default is 1000 for WGS analysis and 500 for WES analysis.                                                                                                                                                                                            |
| `--cnv-skip-contig-list`            | Specifies a comma-separated list of contig identifiers to skip when generating intervals for WGS analysis. The default contigs that are skipped, if not specified, are `chrM,MT,m,chrm`.                                                                                                                                                                                   |
| `--cnv-filter-duplicate-alignments` | Filter duplicate marked alignments during target counts if option is set to `true`. The default setting is `true` unless map/align is enabled and duplicate marking is disabled.                                                                                                                                                                                           |

Note that `--cnv-filter-duplicate-alignments` is only available with duplicate marking option set to true. For more information, see [Filter Duplicate Alignments](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#filter-duplicate-alignments)

For more information of target counting method description, see [Target Counts](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#target-counts)

### GC Bias Correction Options

| Option                            | Description                                                                                                                                                       |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-gc-bias-correction` | Enable or disable GC bias correction when generating target counts. The default is true.                                                                          |
| `--cnv-enable-gcbias-smoothing`   | Enable or disable smoothing the GC bias correction across adjacent GC bins with an exponential kernel. The default is true.                                       |
| `--cnv-num-gc-bins`               | Specifies the number of bins for GC bias correction. Each bin represents the GC content percentage. Allowed values are 10, 20, 25, 50, or 100. The default is 25. |

For more information, see [GC Bias Correction](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#gc-bias-correction)

### Normalization Options

A Panel of normals (PON) is used to provide the reference baseline for copy number variants. PON is required for WES, while WGS can use either self (recommended) or PON normalization.

**Self-normalization option**

| Option                            | Description                                                                                                 |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `--cnv-enable-self-normalization` | Enable/disable self normalization mode, which does not require a panel of normals (only available for WGS). |

**PON normalization options**

| Option                                       | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `--cnv-extreme-percentile`                   | Specifies the extreme median percentile value at which to filter out samples. The default is 2.5.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| `--cnv-max-percent-zero-samples`             | Specifies the number of zero coverage samples allowed for a target. If the target exceeds the specified threshold, then the target is filtered out. The default value is 5%. The option is sensitive to the number of normal samples being used. Make sure you adjust the threshold accordingly. If your panel of normals size is small and the threshold not adjusted, the option could filter out targets that were not intended to be.                                                                                                                                                        |
| `--cnv-max-percent-zero-targets`             | Specifies the number of zero coverage targets allowed for a sample. If sample exceeds the specified threshold, then the sample is filtered out. The default value is 2.5%. The option is sensitive to the total number of target intervals. Make sure you adjust the threshold accordingly. If the capture kit has a small number of probes and the threshold not adjusted, the option could filter out targets that were not intended to be.                                                                                                                                                    |
| `--cnv-target-factor-threshold`              | Specifies the bottom percentile of panel of normals medians to filter out useable targets. The default is 1% for whole genome processing and 5% for targeted sequencing processing.                                                                                                                                                                                                                                                                                                                                                                                                              |
| `--cnv-truncate-threshold`                   | Specifies a percentage threshold for truncating extreme outliers. The default is 0.1%.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| `--cnv-enable-gender-matched-pon`            | Enable/disable gender matched PON normalization. If enabled, DRAGEN uses matched gender PON for sex chromosome normalization. Sex chromosome intervals are filtered if PON has no matched gender sample. The default value is true.                                                                                                                                                                                                                                                                                                                                                              |
| `--cnv-enable-cross-gender-adjustments-chrX` | Enable normalization on chrX by adjusting coverage of PON samples according to the expected number of copies of chrX in male and female samples. If the case sample is male, coverage of female PON samples is scaled down by a factor of 2 on chrX. If the case sample is female, coverage of male PON samples is scaled up by a factor of 2 on chrX. If no male PON samples are available, chrY intervals will be filtered. This feature is only supported for germline enrichment runs. The default value is false; if set to true, then `--cnv-enable-gender-matched-pon` must also be true. |

DRAGEN will select PON normalization if PON is provided. For more information, see [normalization](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#normalization)

### Segmentation Options

The segmentation method for both WGS and WES somatic workflows, and in both tumor-normal and tumor-only configurations, is a variant of shifting level models (SLM) called adaptive shifting level models, or ASLM. This can be overridden with the option `--cnv-segmentation-mode` (see [segmentation](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#segmentation)), but is not recommended.

| Option           | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| --cnv-slm-eta    | Probability that the segmenter changes to any other state than the current state going from the current target to the next target. This could also be expressed as the probability that the true depth for adjacent targets is different for reasons that simple counting noise does not adequately explain. Likewise, the stay-in-state probability is (1.0 - eta). The effective default value is 3e-3, the range is (0.0, 1.0) excluding endpoints. Decreasing this value results in longer segments and reduced fragmentation; increasing produces shorter segments with more fragmentation. |
| --cnv-slm-bafeta | Similar to above, but between adjacent B allele sites. The default value is 1e-7 for somatic WES, 1e-12 for somatic WGS tumor-normal, and 1e-20 for somatic WGS tumor-only. The range is (0.0, 1.0) excluding endpoints. Decreasing this value results in longer segments and reduced fragmentation; increasing produces shorter segments with more fragmentation. However, see below for the limited purpose of segmentation on B allele frequencies.                                                                                                                                           |

The B allele segmentation is performed separately and independently of the depth segmentation. It is a crude segmentation to find the segments which have a balanced B allele frequency, indicating both parental haplotypes are present at equal copy number. A subset of these B allele balanced segments, subject to some additional criteria, are then used to identify a common variance parameter for the depth domain. The ordinary [SLM method](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#shifting-level-models-segmentation) used for depth-based segmentation is then extended to have a state-dependent emission variance computed from the common variance and scaled by the state mean. **The B allele segmentation is not directly used after this, but it plays a critical role in determining the parameterization of the depth-based segmentation.** However, there is no analogous parameter in ASLM to `--cnv-slm-omega` in SLM and HSLM as described for germline analyses.

The following options are documented here in proximity to segmentation options because of their direct relevance to each other. Once provisional calls for copy number (CN) and minor copy number (MCN) have been made on the resulting segments from the segmentation stage, given the selected [purity/ploidy model](#purity-ploidy-model-selection-options), adjacent segments with the same CN and MCN are joined to form a single segment. This is continued until no two adjacent segments satisfy the merging criteria. [Segment merging](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#call-smoothing) is a critical step which compensates for over-segmentation or over-fragmentation happening at the segmentation stage. However, segment merging cannot split segments apart, so it cannot compensate in the other direction. **Thus, segmentation can afford to produce a degree of over-segmentation, but there is no compensatory mechanism for under-segmentation.** These options control segment merging in somatic analyses and do not depend on the segmentation option settings.

| Option                | Description                                                                                                                                                                                                                                                                                  |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| --cnv-merge-distance  | Maximum gap in base pairs between two adjacent segments that still allows them to be merged. The default is 10000 for somatic WGS, meaning segments must be within 10 kb of each other. For WES, the default is effectively unlimited, since target intervals are inherently non-contiguous. |
| --cnv-merge-threshold | Maximum difference in segment mean (linear copy ratio) between two adjacent segments that still allows them to be merged. The default is 0.025 for somatic WGS and 0.4 for somatic WES.                                                                                                      |

Setting `--cnv-merge-threshold` to zero disables segment merging entirely. This is not recommended.

You can specify additional CBS options

| Option                | Description                                                                                                                           |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-cbs-alpha`     | Specifies the significance level for the test to accept change points. The default is 0.01.                                           |
| `--cnv-cbs-eta`       | Specifies the Type I error rate of the sequential boundary for early stopping when using the permutation method. The default is 0.05. |
| `--cnv-cbs-kmax`      | Specifies maximum width of smaller segment for permutation. The default is 25.                                                        |
| `--cnv-cbs-min-width` | Specifies the minimum number of markers for a changed segment. The default is 2.                                                      |
| `--cnv-cbs-nmin`      | Specifies the minimum length of data for maximum statistic approximation. The default is 200.                                         |
| `--cnv-cbs-nperm`     | Specifies the number of permutations used for p-value computation. The default is 10000.                                              |
| `--cnv-cbs-trim`      | Specifies the proportion of data to be trimmed for variance calculations. The default is 0.025.                                       |

For more information, see [segmentation](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#segmentation)

### Purity ploidy model selection options

| Option                                    | Description                                                                                                                                                                 |
| ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-use-somatic-vc-vaf`                | Use the variant allele frequencies (VAFs) from the somatic SNVs to help select the tumor model for the sample. For more information, see [VAF-aware Mode](#vaf-aware-mode). |
| `--cnv-somatic-essential-genes-bed`       | BED file containing genes where the model should not predict HOMDEL                                                                                                         |
| `--cnv-somatic-enable-het-calling`        | Enable HET-calling mode for heterogeneous segments.                                                                                                                         |
| `--cnv-somatic-enable-lower-ploidy-limit` | Enable check on lower ploidy limit based on essential genes                                                                                                                 |
| `--cnv-normal-cnv-vcf`                    | Specify germline CNVs from the matched normal sample. For more information, see [Germline-aware Mode](#germline-aware-mode).                                                |
| `--cnv-somatic-min-purity`                | Specify minimum purity to consider                                                                                                                                          |
| `--cnv-somatic-max-purity`                | Specify maximum purity to consider                                                                                                                                          |
| `--cnv-ascn-min-ploidy`                   | Specify minimum ploidy to consider                                                                                                                                          |
| `--cnv-ascn-max-ploidy`                   | Specify maximum ploidy to consider                                                                                                                                          |

For more information, see [ASCN calling](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#allele-specific-copy-number-calling)

### Filtering Options

| Option                            | Description                                                                                    |
| --------------------------------- | ---------------------------------------------------------------------------------------------- |
| `--cnv-enable-ref-calls`          | Emit copy-neutral (REF) calls in output VCF (defaul`true` for WGS, `false` for WES)            |
| `--cnv-filter-qual`               | QUAL value at which to hard filter CNV VCF (default `40` for WGS/WES, `90` for WES depth-only) |
| `--cnv-filter-length`             | Minimum event length (bp) for PASS calls (default `10000` for WGS, `0` for WES)                |
| `--cnv-filter-del-mean`           | SM value used to hard filter DELs in CNV VCF (Somatic WGS)                                     |
| `--cnv-filter-dup-mean`           | SM value used to hard filter DUPs in CNV VCF (Somatic WGS)                                     |
| `--cnv-filter-cnloh-maf`          | MAF value used to hard filter CNLOHs in CNV VCF (Somatic WGS)                                  |
| `--cnv-somatic-filter-het-length` | Minimum event length to hard filter subclonal CNV VCF                                          |
| `--cnv-post-vcf-target-bed`       | BED file to keep only VCF entries overlapping with target regions                              |

If --cnv-post-vcf-target-bed is specified, VCF records that do not overlap the provided BED intervals are filtered out. This is a post‑processing hard filter applied only to the output VCF and does not affect any upstream workflow steps or CNV modeling.

### Other Options

| Option                                         | Description                                                                |
| ---------------------------------------------- | -------------------------------------------------------------------------- |
| `--cnv-enable-tracks`                          | Enables generation of IGV track files                                      |
| `--cnv-generate-pon-metric-file`               | Generate PON metric file for WES/targeted panel                            |
| `--cnv-exclude-bed`                            | BED file specifying intervals to exclude from analysis                     |
| `--cnv-exclude-bed-min-overlap`                | Minimum overlap fraction for exclusion (default `0.5`)                     |
| `--cnv-sex-genotyper-num-interval-requirement` | Number of sex contig interval requirements for sex genotyper (default:300) |

## CNV Output Files

The somatic CNV workflow generates the following output files:

| File                                           | Description                                         | Format           |
| ---------------------------------------------- | --------------------------------------------------- | ---------------- |
| `<prefix>.tumor.target.counts.gz`              | Raw target counts before bias correction            | gzipped TSV      |
| `<prefix>.tumor.target.counts.gc-corrected.gz` | GC-bias corrected target counts                     | gzipped TSV      |
| `<prefix>.tumor.ballele.counts.gz`             | B-allele counts at population SNP sites             | gzipped TSV      |
| `<prefix>.baf.bedgraph.gz`                     | B-allele frequency in bedgraph format               | gzipped bedGraph |
| `<prefix>.tn.tsv.gz`                           | Tangent-normalized coverage signal                  | gzipped TSV      |
| `<prefix>.cnv.excluded_intervals.bed.gz`       | List of target regions excluded                     | gzipped TSV      |
| `<prefix>.cnv.pon_metrics.tsv.gz`              | Coverage statistics of PON per interval             | gzipped TSV      |
| `<prefix>.cnv.pon_correlation.txt.gz`          | Correlation between CASE and PON                    | gzipped TSV      |
| `<prefix>.seg`                                 | Segmentation results (depth and BAF)                | TSV              |
| `<prefix>.cnv.purity.coverage.models.tsv`      | Model likelihood score for purity/ploidy estimation | TSV              |
| `<prefix>.cnv.vcf.gz`                          | Primary CNV calls (VCF v4.4 by default)             | gzipped VCF      |
| `<prefix>.cyto.vcf.gz`                         | Cytogenetics-compatible calls (if enabled)          | gzipped VCF      |
| `<prefix>.cnv_metrics.csv`                     | Summary metrics including predicted sex             | CSV              |
| `<prefix>.cnv.gff3`                            | Variant calls in GFF format                         | GFF              |
| `<prefix>.tn.bw`                               | Tangent-normalized signal track                     | BigWig           |

### Target Counts Output

`<prefix>.tumor.target.counts.gz`

Compressed tab-delimited file containing the number of read counts per target interval. This is the raw signal as extracted from the alignments of the BAM or CRAM file. The format is identical for both the case sample and any panel of normals samples. There is also a bigWig representation of a `target.counts.diploid` file, which is normalized to the normal ploidy level of 2 instead of raw counts.

Columns:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. Count of alignments in this interval
6. Count of improperly paired alignments in this interval

Header lines starting with `#` contain the DRAGEN version, command line, and other meta information.

Example:

```
#TARGET COUNTS FILE
##DRAGENVersion=<VERSION_INFO>
##DRAGENCommandLine=<CommandLineOptions>
#TargetCountOptions=<CNV_COUNTS_OPTIONS>
...
#Input target file: BED_FILENAME
contig  start   stop    name    WES_EA_N_1      improper_pairs
chr1    12080   12251   target-wes-chr1-12080:12251     662     0
chr1    12595   12802   target-wes-chr1-12595:12802     220     1
...
```

For more information, see [Target Counts File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#target-counts-file)

### GC-Corrected Counts Output

`<prefix>.tumor.target.counts.gc-corrected.gz`

Contains GC-corrected read counts per target interval. The format is equivalent to the `*.target.counts.gz` file:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. GC-corrected read counts in this interval
6. Count of improperly paired alignments in this interval

Example:

```
#GC CORRECTED FILE
##DRAGENVersion=<VERSION_INFO>
##DRAGENCommandLine=<CommandLineOptions>
#TargetCountOptions=<CNV_COUNTS_OPTIONS>
#Original input file: sample.target.counts.gz // raw counts filename
contig  start   stop    name    SampleName      improper_pairs
chr1    12080   12251   target-wes-chr1-12080:12251     981.529698      0
chr1    12595   12802   target-wes-chr1-12595:12802     50.05497673     1
chr1    13163   13658   target-wes-chr1-13163:13658     1086.20189      4
...
```

For more information, see [GC bias correction](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#gc-bias-correction)

### B-Allele Counts

In somatic ASCN runs, B-allele counts are calculated at sites in the tumor sample where the normal sample is likely to be heterozygous. When analyzed in conjunction with a matched normal sample, the sites are those that are called as heterozygous SNVs in the normal sample. When analyzed in tumor-only mode, sites are selected from a population collection (similar to germline ASCN runs). Each B-allele site consists of a reference allele and a variant allele, and the number of reads in the sample supporting each of these alleles is counted.

B-allele counts are written both to gzipped tsv file `*.ballele.counts.gz` and gzipped bedgraph file `*.baf.bedgraph.gz`.

`<prefix>.ballele.counts.gz`

Columns:

1. Contig identifier
2. Start, BED-style (zero-based inclusive) start position of the reference allele
3. Stop, BED-style (one-based inclusive) stop position of the reference allele
4. Base sequence for the reference allele
5. Base sequence for the first allele being counted
6. Base sequence for the second allele being counted
7. The number of qualified reads containing a sequence matching the first allele
8. The number of qualified reads containing a sequence matching the second allele

Additionally, in the case of B-allele sites from a population VCF, the following two additional columns are added after the columns listed above:

9. Population frequency for the first allele
10. Population frequency for the second allele

Example:

```
contig  start   stop    refAllele       allele1 allele2 allele1Count    allele2Count    allele1AF       allele2AF
chr1    51478   51479   T       T       A       4       2       0.6747  0.3253
chr1    82733   82734   T       T       C       111     36      0.79346 0.20654
chr1    83083   83084   T       T       A       0       0       0.1538  0.8462
chr1    86330   86331   A       A       G       9       9       0.87384 0.12616
chr1    88315   88316   G       G       A       0       0       0.8926  0.1074
```

### B-Allele Counts BED Graph

`<prefix>.baf.bedgraph.gz`

B-allele frequency in bedgraph format. Allele count ratios are calculated by sorting alleles according to base priority {A, T, G, C} (descending), producing frequencies deterministically distributed above and below 0.5. This provides easy visualization in IGV of significant BAF changes between neighboring segments.

Example:

```
chr1    11021   11022   0.333333
chr1    14463   14464   0.755102
chr1    16494   16495   0.317708
chr1    38741   38742   0.5
chr1    39014   39015   0.44186
```

### Normalization Output

`<prefix>.tn.tsv.gz`

Contains the normalized signal of the case sample per target interval, i.e., the log2-transformed copy ratio signal. A strong signal deviation from 0.0 indicates a potential for a CNV event. The format is equivalent to the `*.target.counts.gz` file:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. Log2-transformed copy ratio in this interval
6. Count of improperly paired alignments in this interval

Header lines are also included that start with `#`. In some cases, the normalization counts could be patched internally with intervals from other processes, such as the SegDups extension. In such cases, patches are indicated (sorted in order of application) with header lines starting with `#patch`:

```
#patch 1 = <normalized_counts_patch_1_filename>
#patch 2 = <normalized_counts_patch_2_filename>
...
```

and the original (unpatched) `*.tn.tsv.gz` is renamed as `*.tn.unpatched.tsv.gz`. Note: this file is reported in output for inspection, but most use cases will use the (patched) `*.tn.tsv.gz` file downstream of normalization.

An example of a `*.tn.tsv.gz` file is shown below.

```
#title = Tangent normalized coverage profile
#sex = MALE
contig  start   stop    name    SampleName      improper_pairs
chr1    12080   12251   target-wes-chr1-12080:12251     -0.3025426810360819     0
chr1    12595   12802   target-wes-chr1-12595:12802     -0.10691600293612752    0
chr1    13163   13658   target-wes-chr1-13163:13658     -0.55258557719170587    6
...
```

For more information, see [Normalization](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#normalization)

### Excluded Intervals Output

`<prefix>.cnv.excluded_intervals.bed.gz`

To improve accuracy, the DRAGEN CNV Pipeline excludes genomic intervals if one or more of the target intervals failed at least one quality requirement. The excluded intervals are reported to \*.cnv.excluded\_intervals.bed.gz file. The file has a bed format, identifies the regions of the genome that are not callable for CNV analysis and describes the reason intervals were excluded in the fourth column. The following are the possible reasons for exclusion.

Example:

```
chr1    258648  258852  PON_TARGET_FACTOR_THRESHOLD
...
chrX    151717091       151717377       EXCLUDE_BED
chrY    348335  348455  PON_UNMATCHED_GENDER
...
```

* 4th column provides reason for excluded intervals

For more information, see [Excluded Intervals File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#excluded-intervals-file)

### PON Metrics Output

`<prefix>.cnv.pon_metrics.tsv.gz`

The DRAGEN CNV Pipeline generates the PON Metrics File (.cnv.pon\_metrics.tsv.gz) if a Panel of Normals is provided and --cnv-generate-pon-metric-file is set to true. If PON size is less than 2, then an empty file will be generated.

The PON Metric File includes basic statistics of the coverage profile for each interval. To remove sample coverage bias, DRAGEN applies sample median normalization, and then computes the metrics.

Example:

```
contig  start   stop    name    mean    std     normalizedStd min     25%     50%     75%     max     intervalSize    gcContents
1       12098   12178   target-wes-1-12098:12178/1      3.6259044560802365      0.46661435469856077      0.1286890927079175     2.7961783439490446      3.2573018790849675      3.7105263157894739      4.0162683823529415      4.3298969072164946      80      0.49382716049382713
1       12178   12258   target-wes-1-12178:12258/2      5.0685579775753595      0.70638315915955963      0.13936570564740217     3.9044585987261144      4.5225944682508761      5.067708333333333       5.5778115844038769      6.3277777777777775      80      0.46913580246913578
1       12553   12637   target-wes-1-12553:12637/1      4.6990858287992054      0.62537786269786677      0.13308500535681309     3.7417218543046356      4.0305632538350444      5.0382165605095546      5.2151580459770113      5.5773195876288657      84      0.6705882352941176
...
```

For more information, see [PON Metrics File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#pon-metrics-file)

### PON Correlation Output

`<prefix>.cnv.pon_correlation.txt.gz`

The DRAGEN CNV Pipeline generates the PON Correlation File (.cnv.pon\_correlation.txt.gz) if a Panel of Normals is provided. The PON Correlation File includes correlation between CASE sample and each PON sample.

Example:

```
Correlation of case sample CASE_SAMPLE_NAME
  PON1: 0.9786
  PON2: 0.9868
  PON3: 0.9912
  ...
```

For more information, see [PON Correlation File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#pon-correlation-file)

### PON Combined Counts Output

`<prefix>.combined.counts.txt.gz`

If PON samples are provided by `--cnv-normals-file` or `--cnv-normals-list`, then CNV generate single PON file for later uses by `--cnv-combined-counts` option.

Example:

```
#COMBINED COUNTS FILE
##DRAGENVersion=<VERSION_INFO>
##DRAGENCommandLine=<CommandLineOptions>
#TargetCountOptions=<CNV_COUNTS_OPTIONS>
contig  start   stop    name    PON1  PON2  PON3  PON4  PON5  PON6  PON7  PON8  PON9
chr1    69411   69541   target-wes-chr1-69411:69541/1   0       1.8140869319999999      0       0       3.6301639140000002      0       3.6322517749999998      2.7239599229999998      0
chr1    69541   69670   target-wes-chr1-69541:69670/2   0       1.732555405     0       0       3.4641199290000002      0       3.4667661330000001      3.4653864840000002      0
chr1    785931  786282  target-wes-chr1-785931:786282/1 41.683179699999997      37.341243050000003      52.024789929999997      59.030795980000001      53.800898459999999      50.370179270000001      43.404515570000001      43.349519780000001      46.891320659999998
chr1    817466  817596  target-wes-chr1-817466:817596/1 1.9608427310000001      0       0.98101929590000003     1.9595376019999999      0       0.9799059067    1.957367598     4.9118208430000001      0.97657295369999997
chr1    826645  826950  target-wes-chr1-826645:826950/1 67.020948630000007      66.92125953     76.833943899999994      64.963436049999999      37.414978750000003      90.559157929999998      61.079245899999997      78.87869293     69.023087360000005
...
```

### Segmentation Results

`<prefix>.seg`

Contains the segments produced by the segmentation algorithm. The `Segment_Mean` value of a segment is the ratio of the mean of that segment to the whole-sample median, without log transformation (linear copy-ratio). A strong signal deviation from 1.0 indicates a potential for a CNV event.

The file has the following columns:

1. Sample name
2. Contig identified
3. Start position
4. End position
5. Number of intervals in the segment
6. Linear copy-ratio of the segment

An example of a `*.seg` file is shown below.

```
Sample  Chromosome      Start   End     Num_Probes      Segment_Mean
<SampleName> chr1    818022  1117426 224     0.82500341336435279
<SampleName> chr1    1117426 4063702 2438    0.91726081432236528
<SampleName> chr1    4063702 4067591 3       0.38861386123247205
<SampleName> chr1    4067591 7705829 3302    0.93021316913709917
<SampleName> chr1    7705829 9357003 1405    0.98147825043799442
<SampleName> chr1    9357003 9377365 19      0.50269670724395654
<SampleName> chr1    9377365 12859821        2905    1.0684818476332989
```

### BAF Segmentation Output

`<prefix>.baf.seg`

In addition to segmentation of target counts, some workflows perform segmentation of B-allele loci. The output file has suffix `*.baf.seg` and it has the same format of the `*.seg` file with two modifications. First, the `Segment_Mean` value is the mean over B-allele loci of the smaller observed allele fraction. Second, there is an additional column:

7. `BAF_SLM_STATE`: Integer between 0 and 10, indicating bins of minor-allele fraction (low to high), or `.` when the BAF data are too variable to estimate a minor-allele fraction

An example of BAF segmentation output file is shown below:

```
Sample  Chromosome      Start   End     Num_Probes      Segment_Mean    BAF_SLM_STATE
<SampleName> chr1    820348  1104646 194     0.29301737166888697     6
<SampleName> chr1    1105091 1533754 444     0.26185904799069076     5
<SampleName> chr1    1533810 1534166 9       0.41958837071702065     8
<SampleName> chr1    1534217 9356793 6689    0.26034515815016335     5
<SampleName> chr1    9358304 9376529 27      0.46450553586280602     10
```

### Purity/Coverage Models Output

`<prefix>.cnv.purity.coverage.models.tsv`

Contains the tested purity and diploid-coverage models along with their log-likelihood scores. Each row corresponds to a candidate model evaluated by the ASCN caller during model selection.

Columns:

1. Model purity (Cellularity) — fraction of cells in the sample due to tumor \[0, 1]
2. Model diploid coverage — expected read count for a target bin in a diploid region
3. Model log-likelihood — log-likelihood score for this purity/coverage hypothesis
4. Approximate ploidy - approximate sample ploidy estimation before CNV calling, derived from the sample mean coverage
5. Failed constraints - model search constraints that were not satisfied by the model

The model with the highest log-likelihood is selected as the best estimate of tumor purity and ploidy. The selected purity is reported as `EstimatedTumorPurity` in the VCF header.

Example:

```
#Purity Coverage        logL    ApproxPloidy    FailedConstraints
0.05    400     -19966231.8629  5.99    MIN_MHD,HOM_DEL,USER_MIN_PURITY
0.05    401     -19952483.5915  5.88    MIN_MHD,HOM_DEL,USER_MIN_PURITY
0.10    402     -19938276.3649  5.77    
...
```

### VCF Output

`<prefix>.cnv.vcf.gz`

The CNV VCF file follows the standard VCF format [v4.4](https://samtools.github.io/hts-specs/VCFv4.4.pdf). The VCF header is annotated with `##source=<DRAGEN_SOURCE>`, where `<DRAGEN_SOURCE>` identifies the caller which produced the VCF, e.g.:

* `DRAGEN_ASCN`: CNV caller
* `DRAGEN_ASCN_SV`: CNV caller + SV support
* `DRAGEN_CNV`: [legacy depth-only CNV caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy) (note: for legacy reasons this caller uses VCF version [v4.2](https://samtools.github.io/hts-specs/VCFv4.2.pdf))

Due to the nature of how CNV events are represented, not all fields are applicable. In general, if more information is available about an event, then the information is annotated. To include copy neutral (REF) calls, set `--cnv-enable-ref-calls` to true. AOH/LOH events are not available in the legacy depth-only caller.

#### Example Records

```bash
# Example REF call
chr1    819841  DRAGEN:REF:chr1:819841-6103865  N       .       1000    PASS
  END=6103865;REFLEN=5284025
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/0:2:1:1000:1000:2.00155:1.000775:1.000775:129.1:0.5:4544:10920:66,10:0.00368019

# Example copy-neutral LOH call
chr1    6104347 DRAGEN:CNLOH:chr1:6104348-6727324       N       <LOH>   1000    PASS
  END=6727324;REFLEN=622977;SVLEN=622977;LOHTYPE=AOH;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  1/1:2:0:1000:1000:1.9876:0.001988:0.993798:128.2:0.001:528:916:10,12:0.00766703

# Example GAIN call
chr1    16715826        DRAGEN:GAIN:chr1:16715827-16949283      N       <DUP>   744     PASS
  END=16949283;REFLEN=233457;SVLEN=233457;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:3:1:1000:99:3.08217:1.134239:1.541085:198.8:0.368:49:26:20,14:0.0384615

# Example GAIN LOH call
chr15   20212550        DRAGEN:GAINLOH:chr15:20212551-20421468  N       <LOH>   390     PASS
  END=20421468;REFLEN=208918;SVLEN=208918;LOHTYPE=AOH;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  1/1:6:0:1:1:5.90559:0.000000:2.952793:380.91:0:76:1:9,8:0

# Example LOSS call
chr1    25274774        DRAGEN:LOSS:chr1:25274775-25331683      N       <DEL>   226     PASS
  END=25331683;REFLEN=56909;SVLEN=56909;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:1:0:1000:1000:1.01085:0.000000:0.505426:65.2:0:7:10:5,1:0
```

#### Header

The VCF header includes somatic-specific fields in addition to the common CNV header lines:

```
##fileformat=VCFv4.4
##ModelSource=DEPTH+BAF
##EstimatedTumorPurity=0.72
##DiploidCoverage=384.000000
##OverallPloidy=2.103412
##OutlierBafFraction=0.031287
##AlternativeModelDedup=0.72,192
##AlternativeModelDup=0.72,768
...
```

| ID                                        | Description                                                                                                                                                                      |
| ----------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| ModelSource                               | basis on which the final tumor model was chosen (e.g., `DEPTH+BAF`, `DEPTH+BAF_DOUBLED`, `VAF`, `SAMPLE_MEDIAN`).                                                                |
| EstimatedTumorPurity                      | fraction of cells in the sample due to tumor. Range: \[0, 1] or `NA` if a confident model could not be determined.                                                               |
| DiploidCoverage                           | expected read count for a target bin in a diploid region.                                                                                                                        |
| OverallPloidy                             | length-weighted average of copy number for PASS events in the tumor fraction.                                                                                                    |
| OutlierBafFraction                        | fraction of B-allele frequencies incompatible with their segment call. High values may indicate a mismatched normal, cross-sample contamination, or bone marrow transplantation. |
| AlternativeModelDedup/AlternativeModelDup | alternative models corresponding to one fewer or one more whole-genome duplication, given as `(purity, diploid_coverage)`. Useful for manual investigation.                      |

#### Records

All coordinates in the VCF are 1-based.

| ID    | Description                                                                                                                                                                                                                                                                                                                                                                                                                           |
| ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| CHROM | The chromosome (or contig) on which the copy number variant occurs.                                                                                                                                                                                                                                                                                                                                                                   |
| POS   | Start position of the variant. If any of the ALT alleles is a symbolic allele (e.g., `<DEL>`), POS denotes the coordinate of the base preceding the polymorphism.                                                                                                                                                                                                                                                                     |
| ID    | Encodes the event type and coordinates of the event (1-based, inclusive). Event types include `GAIN`, `LOSS`, `REF`, `CNLOH`, and `GAINLOH`.                                                                                                                                                                                                                                                                                          |
| REF   | Contains `N` for all CNV events.                                                                                                                                                                                                                                                                                                                                                                                                      |
| ALT   | Specifies the type of CNV event: `<DEL>`, `<DUP>`, or `<LOH>`. REF calls have ALT `.`. With `--cnv-enable-legacy-vcf-format` (VCF v4.2), the `ALT` field contains `<DEL>,<DUP>` in place of `<LOH>` for AOH/LOH events.                                                                                                                                                                                                               |
| QUAL  | Estimated quality score used in hard filtering. Note: different workflows provide different QUAL score distributions - it is recommended to compare QUAL scores only within results from the same workflow (e.g., it is incorrect to compare QUAL scores between the CNV caller and the [legacy (depth-only) CNV caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline/cnv-germline-legacy)). |

#### FILTER

The FILTER column contains `PASS` if the CNV event passes all filters, otherwise the column contains the name of the failed filter. Default values are defined in the header line for each available FILTER.

| ID        | Description                                         |
| --------- | --------------------------------------------------- |
| binCount  | CNV events with a bin count lower than a threshold. |
| cnvLength | The length of the CNV is lower than a threshold.    |
| cnvQual   | The QUAL of the CNV is lower than a threshold.      |

#### INFO

The INFO column contains information representing the event.

| ID      | Description                                                                                                                                                                                                                    |
| ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| REFLEN  | Length of the event.                                                                                                                                                                                                           |
| SVLEN   | Length of the event. Only present for non-REF records. Note: in VCF v4.2 format (enabled with `--cnv-enable-legacy-vcf-format`), `SVLEN` is a signed representation of `REFLEN` (e.g., a negative value indicates a deletion). |
| SVTYPE  | Always `CNV`. Only present for non-REF records.                                                                                                                                                                                |
| END     | End position of the event (1-based, inclusive).                                                                                                                                                                                |
| LOHTYPE | Type of loss of heterozygosity. Possible values: `CNLOH` (Copy-Neutral LOH), `GAINLOH` (LOH with copy number gain).                                                                                                            |
| HET     | Tag identifying subclonal (heterogeneous) calls, present when `--cnv-somatic-enable-het-calling` is set                                                                                                                        |
| CIPOS   | Confidence interval around the nominal `POS`.                                                                                                                                                                                  |
| CIEND   | Confidence interval around the nominal `END`.                                                                                                                                                                                  |

The meaning of the SVLEN, SVTYPE, END, CIPOS, and CIEND fields match their [VCF v4.2](https://samtools.github.io/hts-specs/VCFv4.2.pdf) definitions.

If using a segment BED file, then the segment identifier is carried over from the input to `SEGID` field.

When [Germline-aware Mode](#germline-aware-mode) is enabled, DRAGEN annotates somatic VCF entries with:

| ID   | Description                                                          |
| ---- | -------------------------------------------------------------------- |
| NCN  | Germline copy number from the matched normal sample.                 |
| SCND | Somatic copy number difference relative to the germline copy number. |

When matching CNV with SV output, additional INFO annotations are added.

#### FORMAT

The common FORMAT fields are described in the header:

| ID   | Description                                                                                                                                                                                                |
| ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| GT   | Genotype                                                                                                                                                                                                   |
| SM   | Linear copy ratio of the segment mean                                                                                                                                                                      |
| CN   | Estimated total copy number of tumor fraction                                                                                                                                                              |
| BC   | Number of read count bins                                                                                                                                                                                  |
| PE   | Number of improperly paired end reads at start and stop breakpoints                                                                                                                                        |
| AS   | Number of allelic read count sites                                                                                                                                                                         |
| CNF  | Floating point estimate of copy number                                                                                                                                                                     |
| CNQ  | Exact total copy number Q-score                                                                                                                                                                            |
| MAF  | Estimate for the minor allele frequency                                                                                                                                                                    |
| MCN  | Estimated minor-haplotype copy number                                                                                                                                                                      |
| MCNF | Floating point estimate of minor-haplotype copy number                                                                                                                                                     |
| MCNQ | Minor copy number Q-score                                                                                                                                                                                  |
| MF   | Mosaic fraction estimate (for MOSAIC calls)                                                                                                                                                                |
| OBF  | Per-segment Outlier BAF Fraction. Percentage of BAF counts which are considered "outlier" with respect to the chosen segment call. Higher values might indicate segments where BAF counts are problematic. |
| SD   | Best estimate of segment's bias-corrected read count                                                                                                                                                       |

For more information, see [CNV VCF](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#cnv-vcf-file).

### Cytogenetics Output

`<prefix>.cyto.vcf.gz`

The Cytogenetics modality output has a similar format to the standard CNV VCF (`*.cnv.vcf.gz`). A list of differences is indicated below:

* Records can have the `INFO/RES` field. In such case, such field indicate the resolution(s) associated with the record.
* Records can have the `INFO/SEGID` field. In such case, such field can either indicate custom predefined segments indicated in input by the user (similar to the standard CNV VCF), or Cytogenetics-specific predefined segments which are typically whole-arm/-chromosome segments automatically injected during the caller execution. In the latter case, the annotation field indicates the ID or name for the arm or chromosome.
* The VCF header is annotated with `##source=DRAGEN_CYTO` to indicate the file is generated by the Cytogenetics modality.

**Note:** The Cyto VCF also provides resolution-specific homozygosity indexes (i.e., computed on each specific resolution's callset). The default minimum size considered is the same as the main `HomozygosityIndex`, and for each resolution in output, there will be an additional header line on the Cyto VCF indicating the resulting metric, e.g., `##HomozygosityIndex(25k)=0.001015`.

### CNV Metrics Output

`<prefix>.cnv_metrics.csv`

The following metrics are reported:

**Sex Genotyper**

| Metric           | Description                                                                               |
| ---------------- | ----------------------------------------------------------------------------------------- |
| Estimated sex    | Estimated sex of the case sample (and panel of normals samples if applicable).            |
| Confidence score | Range: \[0.0, 1.0]. If the sample sex is specified via `--sample-sex`, this value is 0.0. |

DRAGEN Sex Genotyper requires a minimum of 300 target intervals to confidently determine sex genotype; if the panel covers fewer intervals on the sex chromosomes, genotyping will fail and an undetermined genotype is returned. Users may lower this requirement by setting `--cnv-sex-genotyper-num-interval-requirement` to a smaller value, at the risk of increased false genotype calls.

**CNV Summary**

* Bases in reference genome in use
* Average alignment coverage over genome - The average alignment coverage over the genome is calculated by dividing the total number of bases from processed alignment records (excluding those filtered by the Target Counts stage in DRAGEN CNV) by the genome length. Alignment records are filtered taking into consideration duplicate marking status (if available), MAPQ, and mapping status.
* Number of alignment records processed
  * Number of filtered records (total)
  * Number of filtered records (due to duplicates)
  * Number of filtered records (due to MAPQ)
  * Number of filtered records (due to being unmapped)
* PMAD - Pairwise Median Absolute Deviation measures the variation in read coverage between adjacent bins. It measures variability due to various factors, such as DNA degradation, extraction, amplification or library preparation. Higher values indicate noisier sample data. PMAD is calculated as following:
  * Define a vector v\[i] as normalized counts of i-th interval in log scale, and d\[i] as pairwise differences of consecutive normalized counts between i and i+1 intervals, i.e. d\[i] = (v\[i] - v\[i+1])
  * PMAD is median absolute deviation of d, i.e. PMAD = Median(|d\[i]-Median(d)|)
* Coverage MAD - Median absolute deviation of normalized case counts. Higher values indicate noisier sample data.
* Median Bin Count - Median of raw counts normalized by interval size.
* Number of target intervals
* Number of normal samples
* Number of segments
* Number of amplifications - Note: GAINLOH events (ALT=LOH and CN > 2) are also included here
* Number of deletions
* Number of CNLOHs (Copy-Neutral LOHs)
* Number of PASS amplifications - Note: GAINLOH events (ALT=LOH and CN > 2) are also included here
* Number of PASS deletions
* Number of PASS CNLOHs (Copy-Neutral LOHs)
* Post-Normalization Bin Count Sigma - Standard deviation of post-PoN-normalization median-normalized coverage values.

Coverage MAD and Median Bin Count are only printed for WES germline/somatic CNV. Post-Normalization Bin Count Sigma is only printed when PoN normalization has been applied.

Example:

```
SEX GENOTYPER,,sample,UNDETERMINED,0.0000
SEX GENOTYPER,,v1r1_normal_60,UNDETERMINED,0.0000
...
CNV SUMMARY,,Bases in reference genome,3217346917
CNV SUMMARY,,OutlierBafFraction,0.049278
CNV SUMMARY,,beta-binomial overdispersion M,184.400000
CNV SUMMARY,,PMAD,0.067799
CNV SUMMARY,,Coverage MAD,0.06750
CNV SUMMARY,,Median Bin Count,1.80
...
```

For more information, see [CNV Metrics](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-reference#cnv-metrics-file)

### Track Files (IGV)

To generate additional equivalent bigWig and gff files, set the `--cnv-enable-tracks` option to true. These files can be loaded into IGV along with other tracks that are available, such as RefSeq genes. Using these tracks alongside publicly available tracks allows for easier interpretation of calls. DRAGEN autogenerates IGV session XML file if tracks are generated by DRAGEN CNV. The `*.cnv.igv_session.xml` can be loaded directly into IGV for analysis.

The following IGV tracks are automatically populated in the output IGV session file:

| Track File            | Description                                                                                                                                                                                                            | Recommended View   |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| `*.target.counts.bw`  | BigWig representation of target counts bins. Values are GC-corrected if GC correction was performed.                                                                                                                   | Barchart or points |
| `*.improper_pairs.bw` | BigWig representation of improper pairs counts.                                                                                                                                                                        | Barchart           |
| `*.tn.bw`             | BigWig representation of the tangent normalized signal.                                                                                                                                                                | Points             |
| `*.seg.bw`            | BigWig representation of the segments.                                                                                                                                                                                 | Points             |
| `*.baf.seg.bw`        | BigWig representation of BAF segments (if available).                                                                                                                                                                  | Points             |
| `*.baf.bedgraph.gz`   | BED graph representation of B-allele frequency (if available).                                                                                                                                                         | Points             |
| `*.cnv.gff3`          | GFF3 representation of CNV events: DEL=blue, DUP=red, filtered=light gray, REF=green (if enabled), AOH/LOH=magenta. An example is shown below (different workflows may output different attributes on the 9th column). | —                  |

Example GFF3 output:

```
##gff-version 3
chr1    DRAGEN  LOSS    12779193        12859821        30      .       .       Alt=DEL;LinearCopyRatio=0.576;CopyNumber=1;Genotype=0/1;Qual=30;Filter=PASS;Start=12779192;Stop=12859821;Length=80629;BinCount=24;ImproperPairsCount=16,7;color=#0000FF;
chr1    DRAGEN  REF     13106280        13122338        19      .       .       Alt=REF;LinearCopyRatio=1.05981;CopyNumber=2;Genotype=./.;Qual=19;Filter=PASS;Start=13106279;Stop=13122338;Length=16059;BinCount=8;ImproperPairsCount=3,1;color=#00FF00;
chr1    DRAGEN  GAIN    13225213        13247040        66      .       .       Alt=DUP;LinearCopyRatio=2.016;CopyNumber=4;Genotype=./1;Qual=66;Filter=PASS;Start=13225212;Stop=13247040;Length=21828;BinCount=9;ImproperPairsCount=7,5;color=#FF0000;
```

#### IGV Session

![](/files/WGi22zHxhS0kedSW46he)

File extension: `*.igv_session.xml`

The IGV session XML file is prepopulated with track files generated by DRAGEN. The session file loads the reference genome that best matches the standard reference genomes in an IGV installation, by comparing the name of the `--ref-dir` specified on the command-line. Standard UCSC human reference genomes are autodetected, but any variations from the standard reference genomes might not be autodetected. To edit the genome detection, alter the `genome` attribute in the `Session` element to the reference genome you would like for analysis before loading into IGV. The reference identifier used by IGV might differ from the actual name of the genome. The following is an example edited session file.

```
<?xml version="1.0" encoding="utf-8"?>
<Session genome="b37" hasGeneTrack="false" hasSequenceTrack="true" version="8">
    <Resources>
        <Resource path="example.cnv.gff3"/>
        <Resource path="example.cnv.excluded_intervals.bed.gz"/>
        <Resource path="example.target.counts.bw"/>
        <Resource path="example.improper.pairs.bw"/>
        <Resource path="example.tn.bw"/>
        <Resource path="example.seg.bw"/>
    </Resources>
    <Panel height="500" width="1200" name="DataPanel">
        ...
    </Panel>
</Session>
```

Note that depending on the IGV version installed, it may come prepackaged with different flavors of GRCh37. The reference naming conventions have changed so a user may have to edit the `genome` field in the XML file directly. For example, IGV has traditionally packaged a `b37` reference genome, but may also include a `1kg_v37` or a `1kg_b37+decoy`, which will appear on the IGV user interface as "1kg, b37" or "1kg, b37+decoy" respectively.

You can determine what the correct encoding of a reference genome by going to `File > Save Session...` and then inspecting the generated igv\_session.xml file.

![](/files/4f11wpwXtvo98GriV38t)

## Germline-aware Mode

To specify germline CNVs from a matched normal sample, use `--cnv-normal-cnv-vcf`. When specified, CNV records marked as `PASS` in the normal sample are used during tumor-sample segmentation to make sure that confident germline CNV boundaries are also boundaries in the somatic output. Segments with germline copy number changes that are relative to reference ploidy are excluded from somatic model selection. During somatic copy number calling and scoring, the germline copy number is used to modify the expected depth contribution from the normal contamination fraction of the tumor sample. The process leads to more accurate assignment of somatic copy number in regions of germline CNV. DRAGEN then annotates the somatic WGS CNV VCF entries with germline copy number (`NCN`) and the somatic copy number difference relative to germline (`SCND`) for the segments that have germline CNVs.

**Example:**

```bash
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--tumor-bam-input <TUMOR_BAM> \
--bam-input <NORMAL_BAM> \
--enable-variant-caller true \
--cnv-use-somatic-vc-baf true \
--cnv-normal-cnv-vcf <CNV_NORMAL_VCF>
```

## VAF-aware Mode

If both the small variant caller and the CNV caller are enabled in a tumor-matched normal run, somatic SNV variant allele frequencies (VAFs) can inform the purity and ploidy model selection. VAF-based modeling is particularly useful when a tumor has limited copy number variation and/or CNVs are mostly subclonal (e.g., many liquid tumors), preventing the depth+BAF signal from reaching a clear model.

VAF information can also help determine the presence or absence of a whole-genome duplication even in clonal tumors with clear CNVs.

For tumor/matched-normal runs with `--enable-variant-caller true`, VAF-based modeling is enabled by default. To disable it, set `--cnv-use-somatic-vc-vaf false`.

## Advanced Topics

### Cytogenetics Modality

Conventional cytogenetics methodologies typically focus on larger alterations than the ones provided by NGS analyses. The Cytogenetics modality for the CNV caller allows the user to visualize variants at different resolutions, aiming at providing a more flexible workspace for different use cases. For more information, refer to [DRAGEN manual](https://help.dragen.illumina.com/dragen-v4.4/product-guide/dragen-v4.5/user-guide/dragen-dna-pipeline/cnv-calling/cnv-germline.md).

It is enabled with `--cnv-enable-cyto-output` (default true for germline workflows). Not available for somatic WES workflows.

From the same sample, and during the same run, the Cytogenetics modality starts from the high resolution results (before smoothing) provided in the standard output CNV VCF. The output callset then undergoes multiple rounds of smoothing, going progressively from finer resolution to coarser resolution calls (larger alterations). Each round of smoothing produces a smoothed callset which is set aside and becomes the starting point for callsets with higher degree of smoothing.

![](/files/5gBu7zidRsCoAwQCYe6z)

At the end of the smoothing procedure, the Cytogenetics modality produces several outputs, e.g.:

* Multiple GFF3 files, one for each round of smoothing (extension `*cyto.<resolution_ID>.gff3`).
* A single VCF file, with extension `*.cyto.vcf.gz`. This file contains all callsets identified through the smoothing iterations, where the iteration identifier is stored on the `INFO/RES` field. Identical alterations across resolutions are deduplicated. In such case, the `INFO/RES` field will contain a comma-separated list of resolution identifiers.
  * Some resolutions will be based on depth of coverage only (no BAF). Their `INFO/RES` value will reflect the original callset used as a starting point, with added suffix `_depth`. E.g., for depth-only calls derived from resolution `1M`, the new callset will have resolution ID `1M_depth`. Note: calls made at different resolutions or with different information (depth+BAF versus depth-only) may occasionally conflict. For instance, in a region that is AOH that also has a mosaic DEL, the region may be reported as AOH for the depth+BAF calling but may be reported as (mosaic) DEL for the depth-only track. The event type with the strongest evidence will be output for each resolution.
  * An additional callset which does not conform to the ones above (no `INFO/RES` field) is the one containing whole-arm/-chromosome aneuploidies. For this callset, all reported records have the chromosome name or arm name in the `INFO/SEGID` field. Entries for this callset will not be present on any GFF3 file. For more details see the section on whole-chromosome aneuploidies below.
* A single IGV session file, with extension `*.cyto.igv_session.xml`, which provides a convenient way to load the multiple GFF3 files and other typical tracks found on the standard `*.cnv.igv_session.xml`. Below an example screenshot of one of such IGV sessions:
  * The first 5 tracks provide the DRAGEN CNV calls (Blue/DEL, Green/REF, Magenta/AOH, Red/DUP) at decreasing degree of resolution (from high to low, top to bottom).
  * The remaining tracks are similar to the standard `*cnv.igv_session.xml` run, e.g.: poor mappability regions, target counts coverage, improper pairs, B-allele frequency, etc.

![](/files/4f11wpwXtvo98GriV38t)

Below, an example set of calls from the `*.cyto.vcf.gz` output file (note additional `INFO/RES` annotation with respect to `*.cnv.vcf.gz` output file):

```
# Example REF call
chr1    819841  DRAGEN:REF:chr1:819841-6103865  N       .       1000    PASS
  END=6103865;REFLEN=5284025;RES=25k,500k,50k
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/0:2:1:1000:1000:2.00155:1.000775:1.000775:129.1:0.5:4544:10920:66,10:0.00368019

# Example GAIN call
chr1    16605768        DRAGEN:GAIN:chr1:16605769-16645359      N       <DUP>   427     PASS
  END=16645359;REFLEN=39591;RES=25k;SVLEN=39591;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE
  ./1:6:.:1:.:6.27065:.:3.135326:404.457:.:23:0:6,11

# Example LOSS call
chr1    25274774        DRAGEN:LOSS:chr1:25274775-25331683      N       <DEL>   226     PASS
  END=25331683;REFLEN=56909;RES=25k,50k;SVLEN=56909;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:1:0:1000:1000:1.01085:0.000000:0.505426:65.2:0:7:10:5,1:0
```

**Selection of appropriate resolution**

Since the most-informative resolution may vary depending on circumstances (event sizes, distance between calls, presence of smaller calls causing fragmentation, etc), no one-size-fits-all recommendation can work for all cases. However, some practical recommendations to consider are the following:

* Each resolution `INFO/RES` ID identifies the *minimum size* for alterations to be considered PASS.
* If only minimal call smoothing is necessary, resolution 25k can provide a good balance and provide calls in size ranges compatible with Chromosomal Microarray (CMA).
* When comparing against technologies such as karyotyping, resolution 1M may be the more appropriate to reduce call fragmentation.

Note: if the use case under consideration is not impacted by call fragmentation, it is typically recommended to use the `*.cnv.vcf.gz` or `*.cnv_sv.vcf.gz` output results (instead of the ones in `*.cyto.vcf.gz`), to take full advantage of the superior detail of NGS.

**Additional options**

| Option                                          | Description                                                                                    |
| ----------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| --cnv-cyto-keep-resolutions=\<resolution\_list> | Comma-separated list of resolutions to output (currently supported: 25k,50k,500k,1M,1M\_depth) |

**Whole-chromosome Aneuploidy Detection**

For some use cases, it is sometimes necessary to inspect a sample at arm or whole-chromosome level. Typically this would require the use of an additional caller, together with the standard CNV caller with automated segment detection. On the same run, the Cytogenetics modality provides such set of calls within the same VCF file (with extension `*.cyto.vcf.gz`).

```
chr21  12000000   DRAGEN:GAIN:chr21:12000001-46709983  N   <DUP>  1000  PASS
  END=46709983;REFLEN=34709983;SEGID=chr21q;SVLEN=34709983;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:3:1:1000:1000:3.00155:1.002518:1.500775:193.6:0.334:29570:66224:0,0:0.0016016

chrX   1        DRAGEN:LOSS:chrX:2-156040895     N     <DEL>  1000  PASS
  END=156040895;REFLEN=156040894;SEGID=chrX;SVLEN=156040894;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF
  0/1:1:0:1000:1000:0.996364:0.000996:0.498182:82.2:0.001:122580:144548:0,0:0.00995089
```

In the example above, two calls derived from such callset. The segment ID annotation (`INFO/SEGID`) provides the name for the segment call under consideration (i.e., for this example, q-arm of chromosome 21 and the entire chromosome X). REF calls are not displayed by default unless required explicitly by the user (i.e., with `--cnv-enable-ref-calls true`. Note: this will enable REF calls for both CNV and CYTO VCF files).

Note: acrocentric chromosomes (13, 14, 15, 21, and 22) have short arms characterized by repetitive regions. These regions create mappability issues and they are typically excluded from analysis. Thus, calling short arm alterations for these chromosomes is challenging, being based on a small percentage of total arm's length. To avoid false positive calls (in this case, indicating an alteration on the full short arm with evidence only coming from a minimal portion of it), the algorithm has a hard threshold (default 500 intervals) on the minimum number of intervals required when calling whole-arm alterations. When the chromosome arm call does not satisfy this threshold, the call is filtered with `FILTER` `chromArmBinCount`. The default can be changed with option `cnv-filter-chrom-arm-bin-count`.

### Joint SV/CNV calling

Somatic joint calling performs copy number segment matching against all SVs with the starts and ends being matched independently.

Somatic joint calling is not enabled by default and must be enabled with `--enable-cnv-sv-somatic true`.

To ensure copy number neutral SVs have matching copy number segments, whenever `--enable-cnv-sv-somatic` is enabled, `--cnv-enable-ref-calls` is automatically enabled as well.

The following steps are performed:

* SV calling is performed.
* The SV call set is filtered to only PASS SV records.
* For each SV, the breakpoint(s) at which a copy number transition would occur, if it were base-pair consistent with the SV, are obtained.
* CNV segmentation is performed to obtain CNV breakpoints.
* If `--cnv-enable-sv-forced-segmentation` is enabled, SV breakpoints are added to the CNV breakpoints. Segments are generated from the combined CNV and SV breakpoints.
  * If a matching CNV breakpoint is found, the CNV breakpoint is adjusted to the SV breakpoint rather than adding a new breakpoint.
  * If a matching CNV breakpoint is not found, the SV breakpoint is added. CNV segments are therefore split at the internal SV breakpoints.
* CNV calling is performed on the segments.
* Adjacent CNV segments in which the END/CIEND of the left segment overlaps the POS/CIPOS of the right segment are adjusted to remove the gap.
* CNV segment start and end are independently matched to SV breakends based on POS/CIPOS and END/CIEND respectively. When there are multiple matching SVs, the inner-most position is matched.
* If a segmentation gap is created due to SV matching, short CNV segments filling the gaps between SVs are created. Short CNV segments CN is set to the CN of the containing pre-adjusted segment.
* SV `<DEL>`/`<DUP>` records that correspond to a single CNV `<DEL>`/`<DUP>` record are merged into a single VCF record. As with germline joint CNV+SV calling, these VCF record contains both the SV and CNV INFO and FORMAT fields.
* The joint call set is written to the `.cnv_sv.vcf.gz` output file. `cnv.vcf.gz` and `.sv.vcf.gz` outputs are unaffected.

When `--cnv-enable-sv-forced-segmentation` is enabled, the somatic joint CNV+SV call set forms a [breakpoint graph](https://en.wikipedia.org/wiki/Sequence_graph).

#### Example command lines

```
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--bam-input <NORMALBAM> \
--tumor-bam-input <TUMORBAM> \
--enable-map-align false \
--enable-cnv true \
--enable-sv true \
--enable-cnv-sv-somatic true \
--cnv-enable-sv-forced-segmentation true

```

#### Joint SV/CNV VCF Output

The original CNV and SV VCF output files, prior to integration, are available for users in the DRAGEN output directory, as described elsewhere. Additionally, there is an enhanced CNV VCF available with the `*.cnv_sv.vcf.gz` extension. The VCF header lines in the `*.cnv_sv.vcf.gz` mostly correspond to a concatenation of the individual header lines from the CNV and SV VCFs, with a few lines deduplicated and some new ones added. For details on the legacy header lines, please refer to the individual CNV and SV user guide sections.

Newly added header lines are described in the following table.

| Header Field        | Number | Type    | Description                                                                                                          |
| ------------------- | ------ | ------- | -------------------------------------------------------------------------------------------------------------------- |
| END\_LEFT\_BND\_OF  | 1      | String  | ID of CNV whose left end is matched to the end of SV                                                                 |
| END\_RIGHT\_BND\_OF | 1      | String  | ID of CNV whose right end is matched to the end of SV                                                                |
| LEFT\_BND           | 1      | String  | ID of SV that matches the left end of CNV record                                                                     |
| LEFT\_BND\_OF       | 1      | String  | ID of CNV whose left end is matched to SV                                                                            |
| MatchSv             | 1      | Integer | ID of original SV that was merged with CNV record                                                                    |
| OrigCnvEnd          | 1      | Integer | Coordinate of original CNV END                                                                                       |
| OrigCnvPos          | 1      | Integer | Coordinate of original CNV POS                                                                                       |
| RIGHT\_BND          | 1      | String  | ID of SV that matches the right end of CNV record                                                                    |
| RIGHT\_BND\_OF      | 1      | String  | ID of CNV whose right end is matched to SV                                                                           |
| SVCLAIM             | A      | String  | Claim made by the structural variant call. Valid values are D, J, DJ for: abundance, adjacency and both respectively |

Records that can be matched or rescued will have annotations indicating the breakpoint linkage between a CNV and SV record. If a complete match is found, then the `MatchSv` annotation will be present in the record, indicating the SV record's `ID` field for this CNV record. In this case, BND notations refer to the merged record ID itself rather than the SV before merging. Furthermore, the use of the `SVCLAIM` field will indicate if the record has evidence arising from depth signal `D`, or junction signals `J`, or both `DJ`.

Because of the mixing of standalone SV records and CNV records, the FORMAT field may have different annotations. For details on the CNV or SV specific annotations, please refer to the individual CNV and SV user guide sections.

Records that can be matched or rescued will have FILTER set to PASS. The original FILTERs are retained for records that were not matched or rescued. For example, the `cnvLength` FILTER will still be applied to standalone CNV records (those with `SVCLAIM=D`).

Example records are shown below.

```
# Merged record, note presence of SVCLAIM=DJ and MatchSv
chr1    9357666 DRAGEN:LOSS:chr1:9357667-9377061        N       <DEL>   1000    PASS    END=9377061;REFLEN=19395;SVLEN=19395;SVTYPE=DEL;LEFT_BND=DRAGEN:LOSS:chr1:9357667-9377061;OrigCnvPos=9357666;CIPOS=0,2;RIGHT_BND=DRAGEN:LOSS:chr1:9357667-9377061;OrigCnvEnd=9377061;CIEND=0,2;SVCLAIM=DJ;MatchSv=DRAGEN:DEL:1268:0:1:0:0:0;HOMLEN=2;HOMSEQ=TC;SOMATIC;SOMATICSCORE=444.26;LCF;RIGHT_BND_OF=DRAGEN:GAINLOH:chr1:4066343-9357666;LEFT_BND_OF=DRAGEN:LOSS:chr1:9357667-9377061;END_RIGHT_BND_OF=DRAGEN:LOSS:chr1:9357667-9377061;END_LEFT_BND_OF=DRAGEN:GAINLOH:chr1:9377062-9495567      GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:PR:SR:VF:VF1:VAF1:VF2:VAF2       1/1:0:0:1000:585:0.007000:.:0.003500:0.7:.:19:0:95,103:0,100:0,49:0,119:0,119:1.000000:0,119:1.000000
 
# CNV record that did not match, note presence of SVCLAIM=D
chr1    143540109       DRAGEN:GAIN:chr1:143540110-143751543    N       <DUP>   1000    PASS    END=143751543;CIPOS=-269657,1792;CIEND=-1808,799863;REFLEN=211434;SVLEN=211434;SVTYPE=CNV;SVCLAIM=D     GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE:OBF       0/1:3:1:1000:1000:3.000000:1.134000:1.500000:300:0.378:139:119:24,19:0.0168067
  
# SV record that did not match, note presence of SVCLAIM=J
chr1    156918006       DRAGEN:DUP:TANDEM:15023:0:1:0:0:0       N       <DUP:TANDEM>    .       PASS    END=156930478;SVTYPE=DUP;SVLEN=12472;CIPOS=0,3;CIEND=0,3;HOMLEN=3;HOMSEQ=GTG;SOMATIC;SOMATICSCORE=174.11;LCF;RIGHT_BND_OF=DRAGEN:GAIN:chr1:156918005-156918006;LEFT_BND_OF=DRAGEN:GAIN:chr1:156918007-156930478;END_RIGHT_BND_OF=DRAGEN:GAIN:chr1:156918007-156930478;END_LEFT_BND_OF=DRAGEN:GAIN:chr1:156930479-157982548;SVCLAIM=J  PR:SR:VF:VF1:VAF1:VF2:VAF2:PSL  114,38:70,27:162,65:93,65:0.411392:69,65:0.485075:DRAGEN_BND_15023_1_1_2_4_0_0
```


# Reference

## Reference

This reference provides detailed documentation of the DRAGEN CNV pipeline covering contents below.

* [Input](#input)
* [Preprocessing](#preprocessing)
* [Segmentation](#segmentation)
* [Allele-Specific Copy Number Calling](#allele-specific-copy-number-calling)
* [Output Files](#output-files)

## Input

The DRAGEN CNV pipeline supports multiple input formats. To run the DRAGEN CNV pipeline directly with FASTQ input without generating a BAM or CRAM file, see [Streaming Alignments](#streaming-alignments) for instructions on streaming alignment records directly from the DRAGEN map/align stage.

DRAGEN CNV also supports running from an already mapped and aligned BAM or CRAM file. If you have data that has not yet been mapped and aligned, see [Generate an Alignment File](#generate-an-alignment-file).

### Reference Hashtable

For the DRAGEN CNV pipeline, the hashtable must be generated with the `--ht-build-cnv-hashtable` option set to true, in addition to any other options required by other pipelines. When `--ht-build-cnv-hashtable` is true, DRAGEN generates an additional k-mer uniqueness map that the CNV algorithm uses to counteract mappability biases. You only need to generate the k-mer uniqueness map file one time per reference hashtable. The generation takes about 1.5 hours per whole human genome.

The reference hashtable is a pregenerated binary representation of the reference genome. For information on generating a hashtable, see [Prepare a Reference Genome](/dragen-v4.5/product-guides/dragen-v4.5/dragen-reference-support/prepare-a-reference-genome).

The following example command generates a hashtable.

```
dragen \
--build-hash-table true \
--ht-reference \<FASTA\> \
--ht-build-cnv-hashtable true \
--output-directory \<OUTPUT\> \
<OTHER HASHTABLE OPTIONS> \
```

### Generate an Alignment File

The following command-line examples show how to run the DRAGEN map/align pipeline depending on your input type. The map/align pipeline generates an alignment file in the form of a BAM or CRAM file that can then be used in the DRAGEN CNV Pipeline.

You need to generate alignment files for all samples that have not already been mapped and aligned, including any samples to be used as references for normalization. Each sample must have a unique sample identifier. Use the `--RGSM` option to specify the identifier. For BAM and CRAM input files, the sample identifier is taken from the file, so the `--RGSM` option is not required.

The following example command maps and aligns a FASTQ file:

```
dragen \
-r <HASHTABLE> \
-1 <FASTQ1> \
-2 <FASTQ2> \
--RGSM <SAMPLE> \
--RGID <RGID> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align true
```

The following example command maps and aligns an existing BAM file:

```
dragen \
-r <HASHTABLE> \
--bam-input <BAM> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align true
```

The following example command maps and aligns an existing CRAM file:

```
dragen \
-r <HASHTABLE> \
--cram-input <CRAM> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align true
```

### Streaming Alignments

DRAGEN can map and align FASTQ samples, and then directly stream them to downstream callers, such as the CNV Caller and the Haplotype Variant Caller. You can use this process to skip generation of a BAM or CRAM file, which bypasses the need to store additional files.

To stream alignments directly to the DRAGEN CNV pipeline, run the FASTQ sample through a regular DRAGEN map/align workflow, and then provide additional arguments to enable CNV. The following example command line maps and aligns a FASTQ file, and then sends the file to the Germline CNV WGS pipeline.

```
dragen \
-r <HASHTABLE> \
-1 <FASTQ1> \
-2 <FASTQ2> \
--RGSM <SAMPLE> \
--RGID <RGID> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align true \
--enable-cnv true \
--cnv-enable-self-normalization true
```

For information on running CNV concurrently with the Haplotype Variant Caller, see [Concurrent CNV and Small Variant Calling](#concurrent-cnv-and-small-variant-calling).

### Concurrent CNV and Small Variant Calling

DRAGEN can perform mapping and aligning of FASTQ samples, and then directly stream the data to downstream callers. If the input is a FASTQ sample, a single sample can run through both the CNV and the small VC. This triggers self-normalization by default.

Run the FASTQ sample through a regular DRAGEN map/align workflow, and then provide additional arguments to enable the CNV, VC, or both. The options that apply to CNV in the standalone workflows are also applicable here.

The following examples show different commands.

#### Map/Align FASTQ With CNV

```
dragen \
-r <HASHTABLE> \
-1 <FASTQ1> \
-2 <FASTQ2> \
--RGSM <SAMPLE> \
--RGID <RGID> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align true \
--enable-cnv true \
--cnv-enable-self-normalization true
```

#### Map/Align FASTQ With VC

```
dragen \
-r <HASHTABLE> \
-1 <FASTQ1> \
-2 <FASTQ2> \
--RGSM <SAMPLE> \
--RGID <RGID> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align true \
--enable-variant-caller true
```

#### Map/Align FASTQ With CNV and VC

```
dragen \
-r <HASHTABLE> \
-1 <FASTQ1> \
-2 <FASTQ2> \
--RGSM <SAMPLE> \
--RGID <RGID> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align true \
--enable-cnv true \
--cnv-enable-self-normalization true \
--enable-variant-caller true
```

#### BAM Input to CNV and VC

```
dragen \
-r <HASHTABLE> \
--bam-input <BAM> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--cnv-enable-self-normalization true \
--enable-variant-caller true
```

## Preprocessing

### Target Counts

The target counts stage is the first processing stage for the DRAGEN CNV pipeline. This stage bins the alignments into intervals. The primary analysis format for CNV processing is the target counts file, which contains the feature signals that are extracted from the alignments to be used in downstream processing. The binning strategy, interval sizes, and their boundaries are controlled by the target counts generation options, and the normalization technique used.

When working with whole genome sequence data, the intervals are autogenerated from the reference hashtable. Only the primary contigs from the reference hashtable are considered for binning. You can specify additional contigs to bypass with the `--cnv-skip-contig-list` option.

With whole exome sequence data, DRAGEN uses the target BED file supplied with the `--cnv-target-bed` option to determine the intervals for analysis. The target BED file should contain intervals that match those in the panel of normals file. If the intervals in the target BED file and the panel of normals file do not match, DRAGEN will use the target intervals from the panel of normals file.

The target counts stage generates a `*.target.counts.gz` file. You can use the file later in place of any BAM or CRAM by specifying the file with the `--cnv-input` or `--cnv-tumor-input` option for the normalization stage. The `*.target.counts.gz` file is an intermediate file for the DRAGEN CNV pipeline and should not be modified.

Further details are available in [Output Files - Target Counts File](#target-counts-file).

#### Whole Genome

If the samples are whole genome, then the effective target intervals width is specified with the `--cnv-interval-width` option. The higher the coverage of a sample, the higher the resolution that can be detected. This option is important when running with a panel of normals because all samples must have matching intervals. For self-normalization, the actual width of a given target interval might be larger than the specified value.

The default value for WGS is 1000 bp with a sample coverage of ≥ 30x.

| WGS Coverage per Sample | Recommended Resolution\* (bp) |
| ----------------------- | ----------------------------- |
| 5                       | 10000                         |
| 10                      | 5000                          |
| >= 30                   | 1000                          |

Using a `--cnv-interval-width` of less than 250 bp for WGS analysis can drastically increase runtime.

The intervals are autogenerated for every primary contig in the reference. Only references that have the UCSC or GRC convention are supported. For example, `chr1, chr2, chr3, ..., chrX, chrY` or `1, 2, 3, ..., X, Y`. You can specify a list of contigs to skip by using the `--cnv-skip-contig-list` option. This option takes a comma-separated list of contig identifiers. The contig identifiers must match the reference hashtable that you are using. By default, only the mitochondrial chromosomes are skipped. Non-primary contigs are never processed.

For example, to skip chromosome M, X, and Y, use the following option:

```
--cnv-skip-contig-list "chrM,chrX,chrY"
```

#### Whole Exome

If the samples are whole exome samples, supply a target BED file with the `--cnv-target-bed <TARGET_BED>` option. The intervals in the target BED file indicate regions where alignments are expected based on the target capture kit. The BED file intervals are further split into intervals of smaller size, depending on the value of `cnv-interval-width`.

To use a standard BED file, make sure that there is no header present in the file. In this case, all columns after the third column are ignored, similar to the operation of DRAGEN Variant Caller.

#### Target Counts Options

The following options control the generation of target counts.

* `--cnv-counts-method` --- Specifies the counting method for an alignment to be counted in a target bin. Values are midpoint, start, or overlap. The default value is overlap when using the panel of normals approach, which means if an alignment overlaps any part of the target bin, the alignment is counted for that bin. In the self-normalization mode, the default counting method is start.
* `--cnv-min-mapq` --- Specifies the minimum MAPQ for an alignment to be counted during target counts generation. The default value is 3 for self-normalization and 20 otherwise. When generating counts for panel of normals, all MAPQ0 alignments are counted.
* `--cnv-target-bed` --- Specifies a properly formatted BED file that indicates the target intervals to sample coverage over. For use in WES analysis.
* `--cnv-interval-width` --- Specifies the width of the sampling interval for CNV processing. This option controls the effective window size. The default is 1000 for WGS analysis and 500 for WES analysis.
* `--cnv-skip-contig-list` --- Specifies a comma-separated list of contig identifiers to skip when generating intervals for WGS analysis. The default contigs that are skipped, if not specified, are `chrM,MT,m,chrm`.
* `--cnv-filter-duplicate-alignments` --- Filter duplicate marked alignments during target counts if option is set to `true`. The default setting is `true` unless map/align is enabled and duplicate marking is disabled.

Target counts options are recorded in the header of each counts file, to facilitate review and validation of panel of normals. If PON counts are generated with different count options than CASE sample, then DRAGEN will return an option validation error.

#### Filter Duplicate Alignments

PCR duplicates are often considered as noise in coverage depth information. DRAGEN CNV has an option to include/exclude duplicate marked alignments: `--cnv-filter-duplicate-alignments` when counting alignments. This relies on the alignments having the duplicate-marked bit (0x400) in the SAM flag set correctly.

If `--enable-map-align=false`, then duplicate marking should be present in the input file (pre-aligned BAM/CRAM). If `--enable-map-align=true`, then `--enable-duplicate-marking=true` should be set.

Note that CNV will wait for duplicate marking from the Map/Aligner which may increase overall run time.

![](/files/tnGHzYTL5MXu5moFelQD)

| Input format | enable-map-align | Required option                                              |
| ------------ | ---------------- | ------------------------------------------------------------ |
| Fastq        | TRUE             | `--enable-map-align=true`, `--enable-duplicate-marking=true` |
| BAM          | TRUE             | `--enable-map-align=true`, `--enable-duplicate-marking=true` |
| BAM          | FALSE            | `--enable-map-align=false`                                   |

#### Target Counts Dropout Regions

In the WGS case where a BED file is not specified for a given reference, the same intervals should be generated each time. The intervals created take into account the mappability of the reference genome using a k-mer uniqueness map created during hashtable generation.

Due to ambiguity that may arise from non-unique genomic loci, only regions corresponding to unique k-mers are considered. A position in the reference genome is marked as a unique k-mer if the k-mer starting at that position does not show up anywhere else in the reference genome (or non-unique, otherwise). Furthermore, if the k-mer contains any bases other than A, C, T or G, it is marked as non-unique.

For WGS samples and in absence of a `cnv-target-bed` file, the target intervals are auto generated based on the pre-computed k-mer-uniqueness map for a given input reference hashtable, and the `cnv-interval-width` option, which defaults to 1000bp. The `cnv-interval-width` option determines the minimum number of unique k-mer positions required in the interval. There is an upper bound to the length of the interval: when the length of the interval is greater than double the size of `cnv-interval-width`, without reaching the required count of unique k-mer positions, the interval is discarded and the process starts again at the next genomic position. Regions that are discarded are denoted as "dropout" regions, and denoted with exclusion reason `NON_KMER_UNIQUE` in the `*.cnv.excluded_intervals.bed.gz` file.

A dropout region is a complex region that does not count alignments and results in an interval missing from the analysis. Dropout regions include centromeres, telomeres, and low complexity regions. If there is sufficient signal in the flanking regions, an event can still span these dropout regions, even if alignment counting does not occur in the regions. The event is handled by the segmentation stage.

Some of the excluded intervals can be rescued through the segmental duplication extension to the germline CNV workflow, as described in the following section.

#### Rescue of target counts in Segmental Duplications

The germline WGS CNV workflow can be extended to call copy number alterations in a curated subset of segmentally duplicated regions. Segmental duplications are large blocks of DNA ≥ 1kb, characterized by a high degree of sequence identity at nucleotide level (> 90%). This poses a challenge for traditional approaches, and such regions are usually excluded.

This extension complements the original germline WGS CNV workflow by using a tailored algorithm to compute the normalized coverage in such regions, which is then injected before the segmentation step and becomes part of the main CNV workflow in downstream steps. We recommend WGS data aligned to a supported human reference genome (currently only `hg38`) with at least 30x coverage. See below for additional requirements.

**Supported duplications**

The following pairs of genes defining Segmental Duplications are included:

|          |                   |
| -------- | ----------------- |
| CYP2A6   | CYP2A7            |
| FCGR3A   | FCGR3B            |
| RHD      | RHCE              |
| STRC     | STRCP1            |
| ACSM2A   | ACSM2B            |
| ACTR3B   | ACTR3C            |
| AQP12A   | AQP12B            |
| ASAH2    | ASAH2B            |
| CCDC74A  | CCDC74B           |
| CD177    | CD177p1           |
| CD8B     | CD8B2             |
| CFH1     | CFHR1             |
| CYP4A11  | CYP4A22           |
| DHX40    | DHX40P1           |
| EIF5AL1  | EIF5AP4           |
| FCGR2A   | FCGR2C            |
| FFAR3    | GPR42             |
| FOLH1    | FOLH1B            |
| FRMPD2   | FRMPD2B           |
| GPAT2    | GPAT2P1           |
| GSTT2B   | GSTT2             |
| DDT      | DDTL              |
| HCAR2    | HCAR3             |
| HSPA1A   | HSPA1B            |
| KRT81    | KRT86             |
| LGALS7   | LGALS7B           |
| MRPL45   | MRPL45P2          |
| MSTO1    | MSTO2p            |
| MUC20    | MUC20P1           |
| MZT2A    | MZT2B             |
| OTOA     | OTOAp1            |
| PDPR     | PDPR2P            |
| PIEZ02   | ENST00000591853.1 |
| ZP3      | POMZP3            |
| PRAMEF7  | PRAMEF8           |
| PROS1    | PROS2P            |
| RMND5A   | ANAPC1P2          |
| ROCK1    | ROCK1p1           |
| SERPINB3 | SERPINB4          |
| SYT3     | ZNF473CR          |
| TBC1D26  | TBC1D28           |
| TOP3B    | TOP3BP1           |
| TUBA3D   | TUBA3E            |
| ZNF443   | ZNF799            |

**Extension requirements**

This extension is enabled by default in the germline WGS CNV workflow. However, it requires:

* Normalization set to self-normalization (`--cnv-enable-self-normalization=true`).
* GC bias correction enabled (`--cnv-enable-gcbias-correction=true`).
* Counts method set to `start` (`--cnv-counts-method=start`).
* Interval width not greater than 10kb. However, we recommend using the `cnv-interval-width` default (1kb) for best performance.
* A supported reference genome builds in input (currently supported based on: `hg38`).

If necessary, the extension can be disabled through setting `--cnv-enable-segdups-extension` to false.

**Algorithm**

![](/files/2ZQUvyAmBaVhAQba1RVH)

* For each duplicated region, the extension collects all reads falling on top of the two homologous intervals of the pair, and it computes the normalized joint coverage (output to `*.cnv.segdups.joint_coverage.tsv.gz`).
* Through differentiating sites between the two homologous intervals, the extension computes the proportion of coverage to associate to the first and to the second interval (output to `*.cnv.segdups.site_ratios.tsv.gz`).
* Such proportion is used to redistribute the joint normalized coverage between the two homologous intervals.
* The rescued intervals are output to the `*.cnv.segdups.rescued_intervals.tsv.gz` file for inspection and they are automatically injected before the segmentation step.
  * During integration with the original intervals from the CNV caller, the rescued intervals are considered higher priority, thus replacing all original intervals that they overlap with.

See [Output Files - Segdups Extension Files](#segdups-extension-files) for a description of the extension output files.

### B-Allele Counts

In workflows supporting B-allele frequency (BAF), a source of heterozygous SNP sites is required to measure B-allele counts of the input sample. The following are the available modes, of which some are only available in somatic workflows.

| Option                        | Description                                                                                                                                                                                                                                                                                                                       |
| ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `cnv-population-b-allele-vcf` | Specify a population SNP VCF. This option is available for both the germline and the somatic workflows. In somatic, it can be used when a matched normal sample is not available and analysis must be performed in tumor-only mode.                                                                                               |
| `cnv-normal-b-allele-vcf`     | (Somatic-specific) Specify a matched normal SNV VCF. Use when a matched normal sample and the matched normal SNV VCF are available. To use this option, you must run the matched normal sample through the DRAGEN Germline workflow.                                                                                              |
| `cnv-use-somatic-vc-baf`      | (Somatic-specific) Set to `true` to enable DRAGEN to identify germline variants during a tumor/matched-normal run, rather than requiring a separate run on the normal sample. Use if and only if tumor and matched normal input are available. Also enable the Somatic SNV Caller via `enable-variant-caller` to use this option. |

To specify a population SNP VCF, use `--cnv-population-b-allele-vcf` option. To obtain a population SNP VCF, process an appropriate catalog of population variation, such as from dbSNP, the 1000 genome project, or other large cohort discovery efforts. A suitable example file for this parameter is "1000G\_phase1.snps.high\_confidence.vcf.gz" from the GATK resource bundle. Only high-frequency SNPs should be included. For example, include SNPs with minor allele population frequency ≥ 10% to limit run time impact and reduce artifacts. Specify the ALT allele frequency by adding `AF=<alt frequency>` to the `INFO` section of each record. Additional `INFO` fields might be present, but DRAGEN only parses and uses the AF field. Sites specified with `--cnv-population-b-allele-vcf` can be either heterozygous or homozygous in the germline genome from which the tumor genome derives

The following is an example valid population SNP record (note: it needs to be tab-delimited):

```
chr1  51479  .  T  A  1000  PASS  AF=0.3253
```

DRAGEN considers the following requirements when parsing records from the b-allele VCF:

* Only simple SNV sites.
* Records must be marked `PASS` in the `FILTER` field.
* If there are records with the same `CHROM` and `POS` values in the `VCF`, then DRAGEN uses the first record that occurs.

A suitable population B-allele VCF is provided for selected references at [this page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html).

#### Somatic-specific options

To specify a matched normal sample SNV VCF, use the `--cnv-normal-b-allele-vcf` option. The VCF file should come from processing the matched normal sample through the DRAGEN germline small variant caller with filters applied. Typically, this file name has a `*.hard-filtered.vcf.gz` extension. All records marked as PASS that are determined to be heterozygous in the normal sample are used to measure the b-allele counts of the tumor sample. You can also use equivalent gVCF file (`*.hard-filtered.gvcf.gz`), but the processing time is significantly longer due to the number of records, most of which are not heterozygous sites.

If a tumor sample and matched normal input are available, you can avoid having to separately process the matched normal with the DRAGEN germline pipeline by specifying `--cnv-use-somatic-vc-baf true`. If using this option, DRAGEN determines the germline heterozygous sites from the matched normal input and measures the b-allele counts of the tumor sample. The information is passed to the Somatic WGS CNV Caller to simplify the overall somatic workflow.

To enable `--cnv-use-somatic-vc-baf`, enter the following command line options.

* `--tumor-bam-input <TUMOR_BAM>`—Specify the tumor input
* `--bam-input <NORMAL_BAM>`—Specify the matched normal input
* `--enable-variant-caller true`—Enable the somatic SNV variant caller
* `--cnv-use-somatic-vc-baf true`—Enable somatic VC BAF

### GC Bias Correction

GC Biases measure the relationship between GC content and read coverage across a genome. Biases can occur in library prep, capture kits, sequencing system differences, and mapping. Biases can result in difficulties calling CNV events. The DRAGEN GC bias correction module attempts to correct these biases.

The GC bias correction module immediately follows the target counts stage and operates on the `*.target.counts.gz` file. GC bias correction generates a GC bias corrected version of the file, which has a `*.target.counts.gc-corrected.gz` extension in the file name. The GC bias corrected versions are recommended for any downstream processing when working with WGS data. For WES, if there are enough target regions, then the GC bias corrected counts can also be used. See [Output Files - Bias Correction file](#bias-correction-file) for further details on GC-corrected target counts files.

Typical whole-exome capture kits have over 200,000 targets spanning the regions of interest. If your BED file has fewer than 200,000 targets, or if the target regions are localized to a specific region in the genome (such that GC bias statistics might be skewed), then GC bias correction should be disabled.

The following options control the GC bias correction module.

* `--cnv-enable-gcbias-correction` --- Enable or disable GC bias correction when generating target counts. The default is true.
* `--cnv-enable-gcbias-smoothing` --- Enable or disable smoothing the GC bias correction across adjacent GC bins with an exponential kernel. The default is true.
* `--cnv-num-gc-bins` --- Specifies the number of bins for GC bias correction. Each bin represents the GC content percentage. Allowed values are 10, 20, 25, 50, or 100. The default is 25.

### Normalization

The DRAGEN CNV pipeline supports two normalization algorithms:

* Self-Normalization --- Estimates the autosomal diploid level from the sample under analysis to determine the baseline level to normalize by. Sex chromosomes and PAR regions are handled based on the sample sex.
* Panel of Normals --- A reference-based normalization algorithm that uses additional matched normal samples to determine a baseline level from which to call CNV events. The matched normal samples here means it has undergone the same library prep and sequencing workflow as the case sample.

Which algorithm to use depends on the available data and the application. Use the following guidelines to select the mode of normalization.

**Self-Normalization**

* Whole genome sequencing
* Single sample analysis
* Additional matched samples are not readily available
* Simpler workflow via a single invocation
* Only references with `chr1, chr2, chr3, ..., chrX, chrY` or `1, 2, 3, ..., X, Y` naming conventions are supported.

**Panel of Normals**

* Whole genome sequencing
* Whole exome sequencing
* Targeted panels, including somatic panels
* Additional matched samples are available
* Nonhuman samples

The table below shows supported normalization methods for CNV workflow:

| Platform | Germline   | Somatic (T/N) | Somatic (T/O) | Germline (depth-only) | Somatic (depth-only) |
| -------- | ---------- | ------------- | ------------- | --------------------- | -------------------- |
| **WGS**  | Self / PoN | Self / PoN    | Self / PoN    | Self / PoN            | No workflow          |
| **WES**  | PoN        | PoN           | PoN           | PoN                   | PoN                  |

"No workflow" indicates that no workflow exists for this configuration.

#### Self Normalization

The DRAGEN CNV pipeline provides the self-normalization mode that does not require a reference sample or a panel of normals. To enable this mode, set `--cnv-enable-self-normalization` to true. Self-normalization mode bypasses the need to run two stages and can save time. It uses the statistics within the case sample to determine the baseline from which to make a call.

Because self normalization uses the statistics within the case sample, this mode is not recommended for WES or targeted sequencing analysis due to the potential for insufficient data.

The self-normalization mode is the recommended approach for whole-genome sequencing single sample processing. The pipeline continues through to the segmentation and calling stage to produce the final called events.

```
dragen \
-r <HASHTABLE> \
--bam-input <BAM> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--cnv-enable-self-normalization true
```

If you are running from a FASTQ sample, then the default mode of operation is self-normalization.

When operating in self-normalization mode, the `--cnv-interval-width` option used during the target counts stage becomes the effective interval width based on the number of unique k-mer positions. You typically do not have to modify this option.

Self-normalization autogenerates the target intervals to use during the analysis based on the reference genome and is only compatible with standard human references or similar mammalian references (chr1, chr2, chr3, ..., chrX, chrY).

If the user wishes to attempt self-normalization mode on non-standard human references, an override can be set via `--cnv-bypass-contig-check=true`. Under this setting, the CNV caller will do a naive median normalization across all of the contigs within the reference genome. This feature is purely for experimental and research use only, and no claims or validation is made for the use of this feature.

#### Panel of Normals

The panel of normals mode uses a set of matched normal samples to determine the baseline level from which to call CNV events. Proper sample selection and preparation are critical for constructing an accurate and reliable CNV PON. High-quality germline samples—meeting stringent sequencing quality criteria such as a high percentage of bases over Q30, sufficient total read depth (yield), appropriate GC content, and minimal adapter contamination—must be used. Additionally, all samples should originate from the same sample type (e.g., FFPE, fresh-frozen) and be processed under identical experimental conditions, including the same library preparation kit, sequencing platform, and capture panel version. Even minor variations in hybridization efficiency or read depth distribution can introduce systematic artifacts, leading to inaccurate CNV calls.

Below are the key recommendations for preparing a high-quality PON:

* Sample Selection: Normal samples should be sourced from individuals without known chromosomal abnormalities to establish a clean and representative reference baseline. Additionally, normal samples should not be drawn from a cohort that is likely to be enriched for particular CNVs, or enriched for individuals affected by a particular disease or syndrome with a genetic component. Normal samples should ideally be unrelated to each other and to the case samples to be processed. No more than \~6% of the samples in the PON should be related to any case sample. For example, a 50-sample PON containing a quad (mother/father/sibling/proband) may be used to analyze each sample in the quad provided there are no additional related samples in the PON.
* Balanced sample sex: The normal sample set should include both male and female samples in similar numbers to ensure a well-represented reference baseline.
* Exclude Low-Quality Samples: Samples with unusually uneven target coverage, low sequencing depth, or high technical noise should be removed to minimize variability and ensure consistency in the PON.
* Standardized Library Preparation: All samples must be processed using the same library preparation protocol. Any deviations such as differences in hybridization efficiency, incubation time, or temperature can lead to inconsistent coverage patterns, increasing the likelihood of false positive CNV calls.
* Adequate Number of Reference Samples: A sufficient number (a minimum of 50 samples is recommended, though not mandatory) of high-quality reference samples is essential for reliable coverage estimation and robust CNV detection.

By following these guidelines, the PON can effectively minimize technical biases, improving the accuracy and reliability of CNV detection.

In PON mode, the DRAGEN CNV Pipeline is broken down into two distinct stages. The target counts stage is performed on each sample (case and normals), to bin the alignments. The normalization and call detection stage is then performed with the case sample against the panel of normals to determine the events.

CNV PONs can also be built in the cloud using the [DRAGEN Baseline Builder App on BaseSpace](https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/dragen-baseline-builder.html) or the DRAGEN Systematic Noise File Builder Pipeline on [ICA](https://www.illumina.com/products/by-type/informatics-products/connected-analytics.html).

**In-run PON for Germline Exome**

Some pre-built PONs are available for download from the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). When possible, however, it is recommended to utilize an in-run PON created from samples from the same sequencing run and library prep as the case samples. This will ensure that any biases that may have been introduced during library prep and/or sequencing will be properly normalized. If the samples in the sequencing run are sufficiently diverse and contain a large majority of copy neutral samples for each target region, then it is recommended that the PON consist of as many samples from the sequencing run as possible, but can be limited to 96 samples without significantly impacting the accuracy of coverage normalization. When the sequencing run is enriched with samples containing specific CNVs of interest, the in-run PON should be built from only those samples in the run without the enriched events (i.e. "normal samples"). A minimum of 50 such normal samples is recommended as with a pre-built PON. The table below summarizes the available options and high-level steps for running CNV using an in-run PON. CNV and Targeted Caller require separate PON files, but the intermediate counts files can be generated in the same DRAGEN command line invocation. For additional details click on the link for each option.

| Analysis option                                                                                                                                                                                             | Steps                                                                                                                                                                                                                                                          |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| [BSSH planned sequencing run (from BCLs)](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-germline-enrichment-from-bcls-bssh-app#dragen-germline-enrichment-from-bcls-through-run-planning-tool) | <ol><li>Create run using the Run Planning tool in BSSH</li><li>Start planned run in Control Software on instrument</li></ol>                                                                                                                                   |
| [BSSH existing sequencing run (from BCLs)](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-germline-enrichment-from-bcls-bssh-app#dragen-germline-enrichment-from-bcls-from-existing-run)        | <ol><li>Run DRAGEN Germline Enrichment from BCLs App</li></ol>                                                                                                                                                                                                 |
| [ICA from FASTQs/BAMs/CRAMs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-germline-enrichment-ica-app)                                                                                        | <ol><li>Run DRAGEN Germline Enrichment App</li></ol>                                                                                                                                                                                                           |
| [BSSH from FASTQs/BAMs/CRAMs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-enrichment-bssh-app)                                                                                               | <ol><li>Run DRAGEN Enrichment App</li></ol>                                                                                                                                                                                                                    |
| [Local or AMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-wes)                                                                                                                     | <ol><li>BCL to FASTQ conversion</li><li>Generate CNV target counts and Targeted Caller exome counts for each PON sample</li><li>Generate CNV combined counts PON file</li><li>Generate Targeted Caller PON file</li><li>Perform case sample analyses</li></ol> |

**Target Counts Stage**

Target counts should be generated for all normal samples used as a panel of normals. The case samples and all samples to be used as a panel of normals sample must have identical intervals and therefore should be generated with identical settings including reference version, target bed, counting methods, duplicate marking/filtering, filtering method/cutoff, etc. The target counts stage also performs GC Bias correction, if enabled. GC Bias correction is enabled by default, but can be disabled if desired.

The following examples are for WES processing, where a panel of normals is required.

The following is an example command for processing a BAM file.

```
dragen \
-r <HASHTABLE> \
--bam-input <BAM> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--cnv-target-bed <BED> \
```

The following is an example command for processing a CRAM file.

```
dragen \
-r <HASHTABLE> \
--cram-input <CRAM> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--cnv-target-bed <BED> \
```

The following example is for WGS processing, where a panel of normals is optional.

```
dragen \
-r <HASHTABLE> \
--bam-input <BAM> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--cnv-enable-self-normalization true \
```

#### Generating Panel of Normals (Combined Counts)

When running an analysis with a panel of normals (set of target counts), then a column wise concatenated version of the panel is output as a \*.combined.counts.txt.gz file. If the user wishes to generate this file without running the actual calling step, then this can be done by adding the `--cnv-generate-combined-counts=true` option to the command line. The individual panel of normals target counts file must be specified either via `--cnv-normals-file` (one per file) or `--cnv-normals-list` (single text file with paths to each sample).

The following is an example command line using a normals list:

```
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--cnv-normals-list <NORMALS_LIST> \
--enable-cnv true \
--cnv-generate-combined-counts true \
```

**Normalization and Call Detection Stage**

The next step in the CNV pipeline when using a panel of normals is to perform the normalization and to make the calls. This involves a separate execution of DRAGEN during which the normalization is performed and calls are made. This step requires the specification of a set of target counts files to be used for reference-based median normalization.

Ideally the panel of normals samples follow library prep and sequencing workflows that are identical to the workflows of the case sample under analysis. In order to be applicable to both male and female case samples, the panel of normals should include a balanced set of both male and female samples. DRAGEN automatically handles calling on sex chromosomes based on the predicted sex of each sample in the panel.

The presence of CNVs in the panel can result in artifactual calls in the test sample at locations where at least some of the panel samples have copy number changes. This leads to two considerations regarding construction of a panel.

Firstly, while it is not generally possible to select samples with no CNVs, panel samples should not be clearly aneuploid or contain large-scale somatic CNVs; further, if there is a region of particular interest, samples should be selected to be normal in that region.

Secondly, for optimal bias correction, a minimum of 50 samples is recommended as a panel. DRAGEN can run with smaller numbers of samples in the panel, down to even just a single sample, but smaller panels increase the likelihood of artifactual calls. Larger panels do not entirely prevent such issues, but they limit it to regions where non-reference copy numbers are common.

The following is an example of PON files, which uses a subset of the GC corrected files from the target counts stage.

```
/data/output/sample1.target.counts.gc-corrected.gz
/data/output/sample2.target.counts.gc-corrected.gz
/data/output/sample4.target.counts.gc-corrected.gz
/data/output/sample5.target.counts.gc-corrected.gz
/data/output/sample7.target.counts.gc-corrected.gz
/data/output/sample8.target.counts.gc-corrected.gz
...
```

DRAGEN accepts 3 different file formats for a Panel of Normals (PON).

| Option                  | Description                                                                                                                                                                                                                                                                                                                                                                                        |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--cnv-normals-file`    | Individual normal file. This option uses a single file name and can be specified multiple times.                                                                                                                                                                                                                                                                                                   |
| `--cnv-normals-list`    | List of normal files. A plain text file in which each line in the file contains a path pointing to a `*.target.counts.gz` or `*.target.counts.gc-corrected.gz` file generated from the target counts stage. Relative paths are supported if the paths are relative to the current working directory. Absolute paths are recommended in case the workflow is used later or shared with other users. |
| `--cnv-combined-counts` | PON file which combines all normal files in a single file. Combined counts file can be found from output folder of prior DRAGEN run with same panel of normals (`*.combined.counts.txt.gz` file). Some pre-packaged PON file directly downloaded from Illumina support site need to use this option.                                                                                               |

The CNV caller can also be started from the `*.target.counts.gz` (raw counts) or `*.target.counts.gc-corrected.gz` (GC corrected counts) files of the case sample, by specifying the selected file with the `--cnv-input` or `--cnv-tumor-input` option and the PON options described above. When selecting GC corrected counts the option `--cnv-enable-gcbias-correction` should be set to false to disable the GC-correction stage; GC-corrected inputs are not supported for somatic WGS analysis.

For example, the following command normalizes the case sample against the panel of normals.

```
dragen \
-r <HASHTABLE> \
--output-directory <OUTPUT> \
--output-file-prefix <SAMPLE> \
--enable-map-align false \
--enable-cnv true \
--cnv-input <CASE_COUNTS> \
--cnv-normals-list <NORMALS> \
--cnv-enable-gcbias-correction false
```

See [Output Files - Target Counts File](#target-counts-file) for a description of the target counts files.

#### Normalization Options

These options control the preconditioning of the panel of normals and the normalization of the case sample.

* `--cnv-enable-self-normalization` --- Enable/disable self normalization mode, which does not require a panel of normals.
* `--cnv-extreme-percentile` --- Specifies the extreme median percentile value at which to filter out samples. The default is 2.5.
* `--cnv-input` --- Specifies a target counts file for the case sample under analysis when using a panel of normals, for germline analysis (see --cnv-tumor-input for somatic analysis).
* `--cnv-normals-file` --- Specifies a target.counts.gz file to be used in the panel of normals. You can use this option multiple times, one time for each file.
* `--cnv-normals-list` --- Specifies a text file that contains paths to the list of reference target counts files to be used as a panel of normals. Absolute paths are recommended in case the workflow is used later or shared with other users. Relative paths are supported if the paths are relative to the current working directory.
* `--cnv-max-percent-zero-samples` --- Specifies the number of zero coverage samples allowed for a target. If the target exceeds the specified threshold, then the target is filtered out. The default value is 5%. The option is sensitive to the number of normal samples being used. Make sure you adjust the threshold accordingly. If your panel of normals size is small and the threshold not adjusted, the option could filter out targets that were not intended to be.
* `--cnv-max-percent-zero-targets` --- Specifies the number of zero coverage targets allowed for a sample. If sample exceeds the specified threshold, then the sample is filtered out. The default value is 2.5%. The option is sensitive to the total number of target intervals. Make sure you adjust the threshold accordingly. If the capture kit has a small number of probes and the threshold not adjusted, the option could filter out targets that were not intended to be.
* `--cnv-target-factor-threshold` --- Specifies the bottom percentile of panel of normals medians to filter out useable targets. The default is 1% for whole genome processing and 5% for targeted sequencing processing.
* `--cnv-tumor-input` --- Specifies a target counts file for the case sample under analysis when using a panel of normals, for somatic analysis (see --cnv-input for germline analysis).
* `--cnv-truncate-threshold` --- Specifies a percentage threshold for truncating extreme outliers. The default is 0.1%.
* `--cnv-enable-gender-matched-pon` --- Enable/disable gender matched PON normalization. If enabled, DRAGEN uses matched gender PON for sex chromosome normalization. Sex chromosome intervals are filtered if PON has no matched gender sample. The default value is true.
* `--cnv-enable-cross-gender-adjustments-chrX` --- Enable normalization on chrX by adjusting coverage of PON samples according to the expected number of copies of chrX in male and female samples. If the case sample is male, coverage of female PON samples is scaled down by a factor of 2 on chrX. If the case sample is female, coverage of male PON samples is scaled up by a factor of 2 on chrX. If no male PON samples are available, chrY intervals will be filtered. This feature is only supported for germline enrichment runs. The default value is false; if set to true, then `--cnv-enable-gender-matched-pon` must also be true.

### Exclude BED Filtering

You can input an exclude BED to the CNV caller to filter out regions from analysis. Inputting an exclude bed is useful if there are certain regions in the genome that are known to be problematic due to library prep, sequencing, or mapping issues. You can also exclude intervals that specify common CNVs to aid in downstream analysis. You can specify an exclude BED file using `--cnv-exclude-bed`. DRAGEN does not provide an exclude BED. The intervals to exclude should be formatted in standard three-column BED format.

The intervals in the exclude BED are compared with the original target counts intervals. If the overlap is greater than `--cnv-exclude-bed-min-overlap`, the target counts interval are excluded from analysis. The `*.target.counts.gz` file still includes the interval, so you can inspect the original read counts. The normalization stage removes intervals. The `*.tn.tsv.gz` file excludes the intervals removed during normalization.

An excluded interval does not guarantee that a CNV call does not span the interval. If there is sufficient data flanking the region, the segmentation stage along with any merging might still generate a call spanning the excluded interval. However, the call would not take read counts from excluded intervals into account. You can view explanations for excluded intervals in the `*.excluded_intervals.bed.gz` file. See [Output Files - Excluded interval files](#excluded-interval-files) for further details.

## Segmentation

After a case sample has been normalized, the sample goes through a segmentation stage. DRAGEN implements multiple segmentation algorithms, including the following algorithms:

* Circular Binary Segmentation (CBS)
* Shifting Level Models (SLM)

The SLM algorithm has three variants, SLM, Heterogeneous SLM (HSLM), and Adaptive SLM (ASLM). HSLM is for use in exome analysis and handles target capture kits that are not equally spaced. ASLM includes additional sample-specific estimation of technical variability of depth of coverage, as opposed to changes in copy number. The estimations are based on the median variance within fixed windows or a preliminary set of segments based on b-allele ratios. The ASLM algorithm mitigates over segmentation due to noisy or wavy samples; this is the default mode for somatic GWGS analysis.

By default, SLM is the segmentation algorithm for germline whole genome processing, ASLM is the algorithm for somatic whole genome processing, and HSLM is the algorithm for whole exome processing.

If you have specific regions of interests, you can also run with a `--cnv-segmentation-bed`. The option pre-defines the segments to estimate copy numbers region of interest listed in the bed file. See Targeted Segmentation (Segment BED) for more information.

* `--cnv-segmentation-mode` --- Specifies the segmentation algorithm to perform. The following values are available.
  * `bed` --- This option is not applicable to T/N and T/O of somatic WGS and somatic WES workflows
  * `cbs`
  * `slm` --- The default for germline WGS analysis.
  * `hslm` --- The default for germline WES analysis.
  * `aslm` --- The default for somatic analysis, either WGS or WES sample types.

### Circular Binary Segmentation

Circular Binary Segmentation is implemented directly in DRAGEN and is based on *A faster circular binary segmentation for the analysis of array CGH data¹* with enhancements to improve sensitivity for NGS data. The following options control Circular Binary Segmentation.

* `--cnv-cbs-alpha` --- Specifies the significance level for the test to accept change points. The default is 0.01.
* `--cnv-cbs-eta` --- Specifies the Type I error rate of the sequential boundary for early stopping when using the permutation method. The default is 0.05.
* `--cnv-cbs-kmax` --- Specifies maximum width of smaller segment for permutation. The default is 25.
* `--cnv-cbs-min-width` --- Specifies the minimum number of markers for a changed segment. The default is 2.
* `--cnv-cbs-nmin` --- Specifies the minimum length of data for maximum statistic approximation. The default is 200.
* `--cnv-cbs-nperm` --- Specifies the number of permutations used for p-value computation. The default is 10000.
* `--cnv-cbs-trim` --- Specifies the proportion of data to be trimmed for variance calculations. The default is 0.025.

¹Venkatraman ES, Olshen AB. A faster circular binary segmentation algorithm for the analysis of array CGH data. Bioinformatics. 2007;23(6):657-663. doi:10.1093/bioinformatics/btl646

### Shifting Level Models Segmentation

The Shifting Level Models (SLM) segmentation mode follows from the R implementation of *SLMSuite: a suite of algorithms for segmenting genomic profiles²*. The options relevant for SLM and HSLM mode are described in the [germline](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline#segmentation) workflow page. The options for ASLM are described in the [somatic](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-somatic#segmentation) workflow page.

²Orlandini V, Provenzano A, Giglio S, Magi A. SLMSuite: a suite of algorithms for segmenting genomic profiles. BMC Bioinformatics. 2017;18(1). doi:10.1186/s12859-017-1734-5

### User-Defined Segmentation (Segment BED)

DRAGEN CNV optionally accepts additional regions of interest by specifying a `--cnv-segmentation-bed` file. For example, the specified intervals might correspond to gene boundaries matched to the targeted assay. or might covers entire chromosome-arms. Intervals provided by `--cnv-segmentation-bed` will be appended to the CNV VCF with an INFO tag of `SEGID` provided by the name column of the input bed file.

The recommended format for the BED file includes four columns and a header. The four columns are `contig`, `start`, `stop`, and `name`. The name column represents the name of the region and must be unique within the BED file. The name is used in the output VCF and annotated as a segment identifier in the `INFO/SEGID` field. The following example file is in the recommended format with nominal use cases of gene level, arm level, and/or whole chromosome:

```
contig  start      stop       name
chr1    115245083  115261621  NRAS
chr1    204485504  204526342  MDM4
chr2    16075981   16090656   MYCN
chr2    29416087   30143527   ALK
chr3    12626010   12704516   RAF1
chr3    138374228  138478187  PIK3CB
chr3    178866307  178952154  PIK3CA
chr3    195776751  195806640  TFRC
chr16   29500000   30200000   16p11.2 // pathogenic arm-level CNV related to 16p11.2 Microdeletion Syndrome
chr17   15229779   15265326   PMP22   // pathogenic gene-level CNV related to Charcot-Marie-Tooth Disease Type 1A
chr21   1          46709983   chr21   // pathogenic whole chromosome CNV related to Down Syndrome
```

If a three-column BED file is used including the `contig`, `start`, and `stop` values, then segment identifiers are autogenerated from the coordinate fields.

You can mix user-defined segmentation with standard segmentation modes using the `--cnv-segmentation-mode` option (`cbs`, `slm`, `aslm`, `hslm`). For example:

```
dragen \
--cnv-segmentation-mode slm \
--cnv-segmentation-bed <BED_FILE> \
```

In this case, the CNV VCF output will include results from the selected segmentation method (`SLM` in this example), plus additional entries from the user-provided segmentation BED file. If variants are called REF, then they may be filtered based on option `--cnv-enable-ref-calls`. If you set `--cnv-segmentation-mode=bed`, the CNV VCF will include only the entries defined in the segmentation BED file.

If some segments in the --cnv-segmentation-bed file are not covered by any target intervals (from --cnv-target-bed), or if all overlapping target intervals are filtered out (e.g., due to k-mer uniqueness filtering), then the associated segments will not be output to the VCF.

Examples of output VCF entries are below:

```
# Example of REF (only shown if --cnv-enable-ref-calls=true)
chr1    1       DRAGEN:REF:chr1:115245083-115261621     N       .       136     PASS
  END=115261621;REFLEN=16539;SEGID=NRAS GT:SM:CN:BC:PE:LR
  ./.:1.00491:2:32:0,0:0.0

# Examples of DEL/DUP
chr1    204485503       DRAGEN:LOSS:chr1:204485504-204526342    N       <DEL>   55      PASS
  SVLEN=-40839;SVTYPE=CNV;END=204526342;REFLEN=40839;SEGID=MDM4;GCP=0.359;CTP=0.531;ACP=0.475
  GT:SM:CN:BC:PE:LR
  0/1:0.512937:1:64:23,17:122.6

chr2   16075980       DRAGEN:GAIN:chr2:16075981-16090656    N       <DUP>   76      PASS
  SVLEN=14676;SVTYPE=CNV;END=16090656;REFLEN=14676;SEGID=MYCN;GCP=0.401;CTP=0.498;ACP=0.503
  GT:SM:CN:BC:PE:LR
  ./1:1.82095:4:67:0,0:168.1
```

The table below shows CNV workflows supporting the `cnv-segmentation-bed` option:

|     | Germline | Somatic (T/N) | Somatic (T/O) | Germline (depth-only) | Somatic (depth-only) |
| --- | -------- | ------------- | ------------- | --------------------- | -------------------- |
| WGS | ✓        | ✓             | ✓             | ✓                     | No workflow          |
| WES | ✓        | ✓             | ✓             | ✓                     | ✓                    |

ASCN CNV requires `cnv-segmentation-mode` not equal to `bed` to calculate likelihood of purity/ploidy model from segments deriven by data. The table below shows CNV workflows supporting `cnv-segmentation-mode=bed` option:

|     | Germline | Somatic (T/N) | Somatic (T/O) | Germline (depth-only) | Somatic (depth-only) |
| --- | -------- | ------------- | ------------- | --------------------- | -------------------- |
| WGS | ✗        | ✗             | ✗             | ✓                     | No workflow          |
| WES | ✗        | ✗             | ✗             | ✓                     | ✓                    |

"No workflow" indicates that no workflow exists for this configuration.

## Allele-Specific Copy Number Calling

Selecting a diploid coverage level is a key component of an allele-specific copy number (ASCN) caller. In the somatic case, the caller also needs to identify the most likely tumor purity. DRAGEN CNV ASCN callers use a grid-search approach that evaluates many candidate models to attempt to fit the observed read and b-allele counts across all segments in the input sample. A log likelihood score is emitted for each candidate, and all scores are output (in `*.cnv.coverage.models.tsv` or `*.cnv.purity.coverage.models.tsv`, respectively for germline or somatic workflows). The caller chooses the model with the highest log likelihood and then computes several measures of model confidence based on the relative likelihood of the chosen model compared to alternative models.

Note: if BAF data is not sufficient it might be discarded during model estimation, leading to a model based on coverage depth only. In such case, the model will not be able to detect alterations that cannot be easily identified without BAF (e.g., whole-genome trisomy).

### Somatic-specific extensions

#### Default purity/ploidy model

If the confidence in the chosen model is low, the caller returns the default model with estimated tumor purity set to `NA`. This can be identified on the output VCF header lines:

```
##fileformat=VCFv4.4
##ModelSource=DEGENERATE_DIPLOID
##EstimatedTumorPurity=NA
##DiploidCoverage=337.000000
##OverallPloidy=2.008310
```

The default model provides an alternative methodology to identify large somatic alterations (length of at least 1 Mb): records are filtered by this model based on their segment mean value (`SM`) or, in the case of copy-neutral LOHs, by their minor allele frequency value (`MAF`). The threshold values for `SM` used by the caller are estimated automatically considering the variance of the sample, with larger `SM` thresholds for DUPs when the variance is higher. For `MAF` values, PASSing copy-neutral LOHs are called when the `MAF` is below a certain threshold. The user can use alternative threshold values through the `--cnv-filter-del-mean`, `--cnv-filter-dup-mean` and `--cnv-filter-cnloh-maf` parameters.

Finally, when the caller returns the default model, the fields regarding copy number states based on model estimation (i.e., `CN`, `CNF`, `CNQ`, `MCN`, `MCNF`, `MCNQ`) are omitted from the final VCF output. The following is a set of example calls from the final VCF output:

```
# LOSS - not PASSing since it is a short alteration (< 1Mb)
chr11   125205167       DRAGEN:LOSS:chr11:125205168-125215362   N       <DEL>   322     lengthDegenerate
  END=125215362;REFLEN=10195;SVLEN=10195;SVTYPE=CNV
  GT:SM:SD:MAF:BC:AS:PE:OBF
  0/1:0.506231:170.6:0.337:10:1:143,143:0

# REF - not PASSing (only GAIN/LOSS/LOH are output by the default model)
chr12   11415425        DRAGEN:REF:chr12:11415425-33369045      N       .       1000    segmentMean
  END=33369045;REFLEN=21953621
  GT:SM:SD:MAF:BC:AS:PE:OBF
  0/0:1.005638:338.9:0.5:19315:16563:13,86:0.00398647

# GAIN - PASSing
chr12   33369045        DRAGEN:GAIN:chr12:33369046-34655426     N       <DUP>   1000    PASS
  END=34655426;REFLEN=1286381;SVLEN=1286381;SVTYPE=CNV
  GT:SM:SD:MAF:BC:AS:PE:OBF
  0/1:1.487834:501.4:0.355:1045:723:86,89:0.00414938

# LOH - PASSing
chr13   18785987        DRAGEN:CNLOH:chr13:18785988-114351403   N       <LOH>   1000    PASS
  END=114351403;REFLEN=95565416;SVLEN=95565416;LOHTYPE=CNLOH;SVTYPE=CNV
  GT:SM:SD:MAF:BC:AS:PE:OBF
  1/1:0.996893:866.3:0.139:85604:167539:28,12:0.00247715
```

#### Grid search optimization informed by essential regions

In order to improve accuracy on the tumor ploidy model estimation, the somatic WGS CNV caller estimates whether the chosen model calls homozygous deletions on regions that are likely to reduce the overall fitness of cells, which are therefore deemed to be "essential" and under negative selection. In the current literature, recent efforts tried to map such cell-essential genes¹.

The check on essential regions is controlled with `--cnv-somatic-enable-lower-ploidy-limit`(default true). Default bedfiles describing the essential regions are provided for hg19, GRCh37, hs37d5, GRCh38, but a custom bedfile can also be provided in input through the `--cnv-somatic-essential-genes-bed=<BEDFILE_PATH>` parameter. In such case, the feature is automatically enabled. A custom essential regions bedfile needs to have the following format: 4-column, tab-separated, where the first 3 columns identify the coordinates of the essential region (chromosome, 0-based start, excluded end). The fourth column is the region id (string type). For the purpose of the algorithm, currently only the first 3 columns are used. However, the fourth might be helpful to investigate manually which regions drove the decisions on model plausibility made by the caller.

If the somatic WGS CNV caller does not find any overlap between any of the homozygous deletions and any of the essential regions, the model is considered plausible and the model optimization ends. Otherwise, when at least an overlap is found, the model is declared invalid and the model search is repeated on the subset of models that support at least one copy (CN = 1) for the essential region with the lowest coverage among the regions overlapping homozygous deletions.

¹E.g., in 2015 - <https://www.science.org/doi/10.1126/science.aac7041>

The following is an example taken from the output log where this feature is triggered, leading to an additional iteration of model fitting:

```
==================================================================
Checking for model plausibility
==================================================================
Call overlap found: chr18 65121873 80261473
- Call segment counts: 272.4
- Call length: 15139600
Model considered not plausible

==================================================================
Performing model fitting
==================================================================
Constraint on lowest non plausible HOMDEL coverage: 272.4
Performing grid search on 23867 models
```

and, an example where this feature is not triggered:

```
==================================================================
Checking for model plausibility
==================================================================
Plausible model identified
```

#### Rejection of models calling large portions of chromosome as CN0 (homozygous deletion)

Large chromosomal events are likely to negatively impact genome stability and cell viability. The option `--cnv-somatic-homdel-max-fraction` is the maximum allowed fraction for any chromosome that can be called as CN0 (default value: 0.7). If the number of bases on a chromosome are more than this fraction (over the total number of called bases), the weighted average coverage across all HOMDEL segments is taken as the coverage that needs to be at least CN1 for a model to be considered. Model fitting then restarts from the beginning with new constraints (and thus a reduced set of alternative models). This feature can be disabled by setting the parameter to `--cnv-somatic-homdel-max-fraction=1`, effectively allowing the total number of called bases on each chromosome to be CN0 without rejecting the model.

The following is an example taken from the output log where this feature is triggered:

```
==================================================================
Checking for model plausibility
==================================================================
Model considered not plausible
```

#### Constraining tumor purity

When a minimum and/or a maximum value of tumor purity for the input sample are known from additional evidence, it is possible to constrain the search of models based on either or both of these values. The available input options that can be provided are:

* `--cnv-somatic-min-purity` - float in \[0,1]
* `--cnv-somatic-max-purity` - float in \[min-purity,1]

#### Constraining sample ploidy

When a minimum and/or a maximum value of ploidy for the input sample are known from additional evidence, it is possible to constrain the search of models based on either or both of these values. The available input options that can be provided are:

* `--cnv-ascn-min-ploidy` - positive float between 0.5 and `--cnv-somatic-max-quartile-copy-number` (default 9)
* `--cnv-ascn-max-ploidy` - positive float between min ploidy and `--cnv-somatic-max-quartile-copy-number` (default 9)

Please note: the sample ploidy constraints are applied to a preliminary estimation of ploidy from sample parameters (**which might not be exactly respected in the final estimated ploidy in output**), equal to:

2 + (excess coverage with respect to diploid coverage) divided by (coverage of one copy)

For germline ASCN workflows, this is computed as:

$$2+\frac{m - c}{c/2}$$

while for somatic ASCN workflow, this expression is computed as:

$$2+\frac{m - c}{c\*p/2}$$

where:

* `m` is the mean coverage of the input sample
* `c` is the diploid coverage of the model under consideration
* `p` is the tumor purity (only for somatic workflows)

### Subclonal/Mosaic Calling Mode

DRAGEN uses a subclonal/mosaic calling mode for segments with a copy number that is estimated to be heterogeneous among different cells in the sample. Based on a statistical model, a segment is considered to be heterogeneous when the depths or BAF values in a segment are too far away from what is expected for the closest integer-copy number.

Note, in somatic this setting will only be honored when DRAGEN is able to identify a confident model. When a confident model cannot be identified, the caller will return a default model and this feature will always be disabled (see the [Default purity/ploidy model](#default-purityploidy-model) section for more details and nuances of this approach).

When a segment is considered as heterogeneous, the output for the segment is changed as follows.

* The MOSAIC (germline) or HET (somatic) tag is added to the INFO field for the segment.
* At least one of the CN and MCN values is given as a non-REF value. Specifically, the values are given as the integer values closest to CNF and MCNF. If the integer values would result in a REF call, then at least one of the CN and MCN values is adjusted to the closest non-REF value.
* The ID, ALT, and GT fields are set appropriately for the chosen CN and MCN.
* The QUAL score reflects confidence that the segment has nonreference copy number in at least a fraction of the sample.
* The CNQ and MCNQ values reflect confidence that the assigned CN and MCN values are true in all of the tumor cells, so at least one of the CNQ and MCNQ values is typically less than five.

To turn on this feature, specify either one of these options:

* `--cnv-enable-mosaic-calling=true` (for the germline ASCN workflow, default true)
* `--cnv-somatic-enable-het-calling=true` (for the somatic workflow, default false)

Note: this calling mode can be disabled for alterations smaller than N bases. It is recommended not to change default thresholds. If necessary, however, they can be changed with:

* `--cnv-filter-mosaic-length=N` (for the germline ASCN workflow, default 100000)
* `--cnv-somatic-filter-het-length=N` (for the somatic workflow, default 0)

The following is an example of MOSAIC GAIN call from the germline ASCN model (from `*.cnv.vcf.gz` output file):

```
chr2    89673674        DRAGEN:GAIN:chr2:89673675-89851643      N       <DUP>   1000    PASS
  END=89851643;REFLEN=177969;MOSAIC;SVLEN=177969;SVTYPE=CNV
  GT:CN:MCN:CNQ:MCNQ:CNF:MCNF:SM:SD:MAF:BC:AS:PE
  ./1:4:.:0:.:4.32403:.:2.162016:278.9:.:70:0:13,2
```

The assigned total CN is 4. However, inspecting the CNF annotation (CNF \~ 4.32), we can see that the segment above has a larger deviation from the diploid state with respect to the assigned integer CN state. This can support various hypotheses on the fraction of cells bearing the different CN states, for example:

* 32% of cells having CN5, 68% of cells having CN4

### Allele Specific Copy Number Examples

In addition to assigning total copy number based on depth, ASCN Callers make use of BAFs to call allele specific copy numbers. The following table provides examples for a DUP in a reference-diploid region:

| Total Copy Number (CN) | Minor Copy Number (MCN) | ASCN Scenario |
| ---------------------- | ----------------------- | ------------- |
| 4                      | 2                       | 2+2           |
| 4                      | 1                       | 3+1           |
| \*4                    | 0                       | 4+0           |

\*The entry represents a Absence or Loss of Heterozygosity (AOH/LOH) case. The total copy number is still considered a DUP, so the entry is annotated as `GAINLOH` to distinguish the value from Copy Neutral AOH/LOH (`CNLOH`), which would be annotated as 2+0.

### Call Smoothing

The segmentation stage might produce adjacent or nearby segments that are assigned the same copy number and have similar depth and BAF data. This segmentation can result in a region with consistent true copy number being fragmented into several pieces. The fragmentation might be undesirable for downstream use of copy number estimates. Also, for some uses, it can be preferable to smooth short segments that would be assigned different copy numbers whether due to a true copy number change or an artifact. To reduce undesirable fragmentation, initial segments can be merged during a postcalling segment smoothing step.

After initial calling, segments shorter than the specified value of `--cnv-filter-length` are deemed negligible. Among the remaining nonnegligible segments, successive pairs are evaluated for merging. On a trial basis, the caller combines two successive segments that are within `--cnv-merge-distance` (default value of 10000 for WGS Somatic CNV) of one another and have the same CN and MCN assignments, along with any intervening negligible segments into a single segment that is recalled and rescored. If the merged segment receives the same CN and MCN as its constituent nonneglible pieces with a sufficiently high-quality score, the original segments are replaced with the merged segment. The merged segment might be further merged with other initial or merged segments to either side. Merging proceeds until all segment pairs that meet the criteria are merged. Note: in somatic workflows, when the germline CN information is available, and two segments have different germline CN, they will not be merged.

### QUAL Model

The caller uses a model based on diploid coverage (and purity in somatic workflows) from depth of coverage and B-allele frequency.

Given the most likely diploid coverage (and purity in somatic workflows), for each segment, the algorithm calls the most likely copy number state (complete with total copy number CN, and minor allele copy number MCN).

The probability of the REF state is used in input to the scoring algorithm which outputs the QUAL value (a PHRED score capped at 1000). The QUAL value is the PHRED score where the probability of error is the probability of REF when an alteration is called, or the probability of having a non-REF call when the segment should be called REF.

Note: this is different from how QUAL is computed in (legacy) depth-only callers.

## Output Files

DRAGEN emits the calls in the standard VCF format. The VCF file includes only copy number gain and loss events. To include copy neutral (REF) calls in the output VCF, set `--cnv-enable-ref-calls` to true. AOH/LOH events are available in workflows where allele-specific copy number is available.

### CNV VCF File

File extension: `*.cnv.vcf.gz`

The CNV VCF file follows the standard VCF format. Due to the nature of how CNV events are represented versus how structural variants are represented, not all fields are applicable. In general, if more information is available about an event, then the information is annotated. Some fields in the DRAGEN CNV VCF are unique to CNVs. The VCF header is annotated with `##source=DRAGEN_CNV` to indicate the file is generated by the DRAGEN CNV pipeline.

#### VCF format differences between different callers

In the DRAGEN CNV component, two versions of the VCF specification are used for the `*.cnv.vcf.gz` file:

* For ASCN workflows, the format used is [VCF v4.4](https://samtools.github.io/hts-specs/VCFv4.4.pdf)
* For depth-only workflows (including multisample CNV calling), the format used is [VCF v4.2](https://samtools.github.io/hts-specs/VCFv4.2.pdf)

The differences between the two formats in output from DRAGEN are the following:

**General**

| Field        | VCF v4.2             | VCF v4.4        |
| ------------ | -------------------- | --------------- |
| `INFO/SVLEN` | Positive or Negative | Always Positive |

**Absence/Loss of Heterozygosity (AOH/LOH)**

| Field       | VCF v4.2      | VCF v4.4 |
| ----------- | ------------- | -------- |
| `ALT`       | `<DEL>,<DUP>` | `<LOH>`  |
| `FORMAT/GT` | `1/2`         | `1/1`    |

#### Header

The following is an example of some of the header lines that are specific to CNV:

```
##fileformat=VCFv4.2
##CoverageUniformity=0.402517
##contig=<ID=1,length=249250621>
##contig=<ID=2,length=243199373>
##contig=<ID=3,length=198022430>
##contig=<ID=4,length=191154276>
##contig=<ID=5,length=180915260>
...
##reference=file:///reference_genomes/Hsapiens/hs37d5/DRAGEN
##ALT=<ID=CNV,Description="Copy number variant region">
##ALT=<ID=DEL,Description="Deletion relative to the reference">
##ALT=<ID=DUP,Description="Region of elevated copy number relative to the reference">
##INFO=<ID=REFLEN,Number=1,Type=Integer,Description="Number of REF positions included in this record">
##INFO=<ID=SVLEN,Number=.,Type=Integer,Description="Difference in length between REF and ALT alleles">
##INFO=<ID=SVTYPE,Number=1,Type=String,Description="Type of structural variant">
##INFO=<ID=END,Number=1,Type=Integer,Description="End position of the variant described in this record">
##INFO=<ID=CIPOS,Number=2,Type=Integer,Description="Confidence interval around POS">
##INFO=<ID=CIEND,Number=2,Type=Integer,Description="Confidence interval around END">
##FILTER=<ID=cnvQual,Description="CNV with quality below <WORKFLOW-SPECIFIC DEFAULT VALUES>">
##FILTER=<ID=cnvCopyRatio,Description="CNV with copy ratio within +/- 0.2 of 1.0">
##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype">
##FORMAT=<ID=SM,Number=1,Type=Float,Description="Linear copy ratio of the segment mean">
##FORMAT=<ID=CN,Number=1,Type=Integer,Description="Estimated copy number">
##FORMAT=<ID=BC,Number=1,Type=Integer,Description="Number of bins in the region">
##FORMAT=<ID=PE,Number=2,Type=Integer,Description="Number of improperly paired end reads at start and stop breakpoints">
```

The following header lines are specific to the somatic CNV callers (WGS/WES) and the germline WGS CNV caller:

* `ModelSource` The primary basis on which the final model was chosen. The following values can be included:
  * `DEPTH+BAF`: Depth+BAF signal is used to determine model.
* `DiploidCoverage` Expected read count for a target bin in a diploid region. The numeric value is unlimited.
* `OverallPloidy` Length weighted average of copy number for PASS events (for the tumor fraction in somatic runs). The numeric value is unlimited.
* `OutlierBafFraction` A QC metric that measures the fraction of b-allele frequencies that are incompatible with the segment the BAFs belong to. High values might indicate a mismatched normal, substantial cross-sample contamination, or a different source of a mosaic genome, such as bone marrow transplantation. The range of this field is \[0, 1].
* `HomozygosityIndex` Autosomal AOH/LOH percentage, considering only PASS AOH/LOH greater or equal than a certain threshold. This metric can be used as a proxy for consanguinity in the germline WGS CNV caller. The default minimum size for PASS AOH/LOH to be considered is 2Mb, since it is often found that shorter ROHs "do not arise from inbreeding in recent generations and are common in all of the populations represented in the HGDP" (Kirin et al., 2010). However, a custom minimum size can be set through the option `--cnv-min-length-homozygosity-index`. Note: The Cyto VCF (`*.cyto.vcf.gz`) also provides resolution-specific homozygosity indexes (i.e., computed on each specific resolution's callset). The default minimum size considered is the same as the main `HomozygosityIndex`, and for each resolution in output, there will be an additional header line on the Cyto VCF indicating the resulting metric, e.g., `##HomozygosityIndex(25k)=0.001015`.

The following header lines are specific to the somatic CNV callers (WGS/WES):

* `ModelSource` can also have the following values (see section below for additional details):
  * `DEPTH+BAF_DOUBLED`: The initial depth+BAF model is duplicated based on VAF signal or excess segments at half the expected depth change.
  * `DEPTH+BAF_DEDUPLICATED`: The depth+BAF model is deduplicated based on VAF signal or insufficient segments supporting a duplication.
  * `DEPTH+BAF_WEAK`: Depth+BAF signal is used to determine tumor model, but this is associated with lower-confidence than `DEPTH+BAF`.
  * `VAF`: VAF signal is used to determine tumor model due to insufficient depth+BAF signal.
  * `SAMPLE_MEDIAN`: Sample is treated as high-purity diploid in absence of adequate signal from depth+BAF and VAF. Diploid coverage set to sample median.
  * `DEGENERATE_DIPLOID`: Sample is treated as high-purity diploid in absence of adequate signal from depth+BAF and VAF. The diploid coverage is set to lowest value observed in a substantial number of bases in segments with BAF=50%.
* `EstimatedTumorPurity` Estimated fraction of cells in the sample due to tumor. The range of this field is \[0, 1] or `NA` if a confident model could not be determined.
* `AlternativeModelDedup` An alternative to the best model corresponding to one less whole-genome duplication. The alternative is given as a pair of values (tumor purity, diploid coverage). This can be useful for manual investigation if the best model might involve a spurious genome duplication.
* `AlternativeModelDup` An alternative to the best model corresponding to one more whole-genome duplication. The alternative is given as a pair of values (tumor purity, diploid coverage). This can be useful for manual investigation where the best model might have missed a true genome duplication.

**Understanding `ModelSource` Annotation (Somatic only)**

The `ModelSource` indicates the type and strength of evidence used to determine the tumor purity and ploidy model for the sample. Possible values are listed in approximate order of decreasing evidence strength, with `DEPTH+BAF` variants representing the most robust determinations and the degenerate models representing the least confident scenarios.

* `DEPTH+BAF` represents the strongest evidence, where both sequencing depth (read coverage) and B-allele frequency (BAF) signals consistently support the chosen model, confirmed by variant allele frequency (VAF) data (if available).
* `DEPTH+BAF_DOUBLED` indicates that the initial depth and BAF model was adjusted upward by a whole-genome duplication factor, supported by either VAF evidence showing variants at the expected frequencies for a duplicated genome, or an excess of genomic segments with a closely matching state in the model with the WGD. Conversely, `DEPTH+BAF_DEDUPLICATED` means the model was adjusted downward by removing a whole-genome duplication, based on VAF data inconsistent with duplication or insufficient genomic segments supporting the higher ploidy hypothesis.
* `DEPTH+BAF_WEAK` reflects a scenario where depth and BAF signals provided the model, but with lower confidence than other `DEPTH+BAF` model sources. A model receiving this model source is found through depth and BAF, but either:
  * VAF data is available but the model is not concordant with VAF evidence.
  * Several regions of the genome have no closely matching state under the selected model.
* `VAF` indicates that variant allele frequency data from somatic mutations became the primary evidence source because depth and BAF signals were insufficient or conflicting.
* Finally, `DEGENERATE_DIPLOID` and `SAMPLE_MEDIAN` represent fallback models used when neither depth/BAF nor VAF provide adequate signal, and tumor purity cannot be reliably estimated. These assume the sample is high-purity diploid, with coverage set either to the lowest observed value in BAF-balanced regions (`DEGENERATE_DIPLOID`) or to the sample's median coverage (`SAMPLE_MEDIAN`).

#### Records

All coordinates in the VCF are 1-based.

**CHROM**

The CHROM column specifies the chromosome (or contig) on which the copy number variant being described occurs.

**POS**

The POS column is the start position of the variant. According to the VCF specification, if any of the ALT alleles is a symbolic allele, such as `<DEL>`, then the padding base is required and POS denotes the coordinate of the base preceding the polymorphism.

**ID**

The ID column is used to represent the event. The ID field encodes the event type and coordinates of the event (1-based, inclusive). In addition to representing `GAIN`, `LOSS` and `REF` events, in Somatic (WGS/WES) and Germline (WGS) CNV, the ID could include the Copy Neutral Loss/Absence of Heterozygosity (CNLOH) or Copy Number Gain with LOH (GAINLOH) events.

**REF**

The REF column contains an N for all CNV events.

**ALT**

The ALT column specifies the type of CNV event. Because the VCF contains only CNV events, only the `<DEL>`, `<DUP>` or `<LOH>` entries are used. If REF calls are emitted, their ALT will always be `.`. In workflows where allele-specific copy number (ASCN) is available, if the legacy DRAGEN VCF format (VCF v4.2) has been enabled with `--cnv-enable-legacy-vcf-format`, the `ALT` field will contain two alleles, `<DEL>,<DUP>`, in place of `<LOH>`, for AOH/LOH events.

**QUAL**

The QUAL column contains an estimated quality score for the CNV event, which is used in hard filtering. Each CNV workflow has different defaults and the value used can be found in the VCF header. Note: different workflows (e.g., germline WGS depth-only vs germline WGS) do not share the same underlying model and provide different QUAL score distributions. It is recommended to compare QUAL scores only within results from the same workflow. More details are available on [QUAL (depth-only)](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline#QUAL) and [QUAL (ASCN)](#qual-model).

**FILTER**

The FILTER column contains `PASS` if the CNV event passes all filters, otherwise the column contains the name of the failed filter. Default values are defined in the header line for each available FILTER.

| FILTER               | Germline WGS (depth-only) | Germline WGS | Germline WES | Somatic WGS | Somatic WES (depth-only) | Somatic WES |
| -------------------- | ------------------------- | ------------ | ------------ | ----------- | ------------------------ | ----------- |
| `binCount`           |                           | ✓            |              | ✓           |                          | ✓           |
| `chromArmBinCount`   |                           | ✓            | ✓            | ✓           |                          |             |
| `cnvBinSupportRatio` | ✓                         |              | ✓            |             | ✓                        |             |
| `cnvCopyRatio`       | ✓                         |              | ✓            |             | ✓                        |             |
| `cnvHetLength`       |                           |              |              | ✓           |                          | ✓           |
| `cnvLength`          | ✓                         | ✓            | ✓            | ✓           | ✓                        | ✓           |
| `cnvLikelihoodRatio` | ✓                         |              | ✓            |             |                          |             |
| `cnvMosaicLength`    |                           | ✓            | ✓            |             |                          |             |
| `cnvQual`            | ✓                         | ✓            | ✓            | ✓           | ✓                        | ✓           |
| `dinucQual`          | ✓                         |              | ✓            |             |                          |             |
| `mosaicFraction`     |                           | ✓            | ✓            |             |                          |             |
| `highCN`             | ✓                         |              |              |             |                          |             |
| `lengthDegenerate`   |                           |              |              | ✓           |                          | ✓           |
| `segmentMean`        |                           |              |              | ✓           |                          | ✓           |
| `SqQual`             |                           |              |              |             |                          | ✓           |

**FILTER description**

Available FILTERs:

* `binCount` - Filters CNV events with a bin count lower than a threshold.
* `cnvBinSupportRatio` which indicates, for CNVs greater than 80kb, the percent span of supporting target intervals is lower than a threshold.
* `cnvCopyRatio` which indicates that the segment mean of the CNV is not far enough from copy neutral.
* `cnvHetLength` which indicates that a HET call below a certain length has been filtered as candidate FP.
* `cnvLength` which indicates that the length of the CNV is lower than a threshold.
* `cnvLikelihoodRatio` indicates a log10 likelihood ratio of ALT to REF is less than a threshold.
* `cnvMosaicLength` which indicates that a MOSAIC call below a certain length has been filtered as candidate FP.
* `cnvQual` which indicates that the QUAL of the CNV is lower than a threshold.
* `chromArmBinCount` which indicates that a whole-arm alteration call is based on a minimal portion (default 500 intervals) of the entire arm (e.g., in acrocentric chromosomes, where the short arm is mainly consisting of poor mappability regions, that are ignored during copy-number calling).
* `dinucQual` is applied based on the percentage of bases in a segment that belong to a two-base set (GC, CT, or AC), determined by individual occurrences. A CNV call is filtered out if any of these percentages fall outside typical ranges, indicating a likely false positive.
* `mosaicFraction` which indicates that the mosaic fraction of a germline CNV is below a defined threshold (`--cnv-filter-mosaic-fraction`). This filter is applied only to small CNVs with lengths shorter than the specified size threshold (`--cnv-filter-mosaic-fraction-max-length`).
* `highCN` which indicates a CNV call with implausible copy number (>6).
* `lengthDegenerate` - Marks records as non-`PASS`ing based on each record's length (`REFLEN`) when the caller returns the default model. Segments having less than 1 Mb are assigned this filter when returning the default model.
* `segmentMean` - Marks records as non-`PASS`ing based on each record's segment mean (`SM`) when the caller returns the default model. Segments having insufficient `SM` in `DEL`s or `DUP`s are assigned this filter when returning the default model.
* `SqQual` - Marks records as non-`PASS`ing based on each record's somatic quality (SQ) when the caller returns the default model. Segments having insufficient SQ are assigned this FILTER when returning the default model. SQ is the somatic quality value which is a Phred scale score of p-value from 2-sample t-test comparing normalized counts of CASE vs PON.

**INFO**

The INFO column contains information representing the event.

* `REFLEN` indicates the length of the event.
* `SVLEN` indicates the length of the event and it is only present for non-REF records. Note: if the legacy DRAGEN VCF format (VCF v4.2) has been enabled with `--cnv-enable-legacy-vcf-format`, `SVLEN` is a signed representation of `REFLEN` (e.g., a negative value indicates a deletion).
* `SVTYPE` is always CNV and only present for non-REF records.
* `END` indicates the end position of the event (1-based, inclusive).

The legacy (depth-only) Germline CNV caller also includes the following INFO fields:

| ID  | Description                         |
| --- | ----------------------------------- |
| GCP | Percentage of bases that are G or C |
| CTP | Percentage of bases that are C or T |
| ACP | Percentage of bases that are A or C |

If using a segment BED file, then the segment identifier is carried over from the input to `SEGID` field.

In Germline WGS CNV the `MOSAIC` tag identifies mosaic calls. In Somatic CNV the `HET` tag identifies subclonal calls. See [Subclonal/Mosaic-Calling Mode](#subclonal-mosaic-calling-mode) for more details.

When matching CNV with SV output, additional INFO annotations are added. See [CNV With SV Support](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline#cnv-with-sv-support).

**FORMAT**

The common FORMAT fields are described in the header:

| ID   | Description                                                                                                                                                                                                |
| ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| GT   | Genotype                                                                                                                                                                                                   |
| SM   | Linear copy ratio of the segment mean                                                                                                                                                                      |
| CN   | Estimated copy number                                                                                                                                                                                      |
| BC   | Number of bins in the region                                                                                                                                                                               |
| PE   | Number of improperly paired end reads at start and stop breakpoints                                                                                                                                        |
| AS   | Number of allelic read count sites                                                                                                                                                                         |
| BC   | Number of read count bins                                                                                                                                                                                  |
| CN   | Estimated total copy number of sample (for the tumor fraction in Somatic callers). This field is not present if the model cannot be estimated with high confidence.                                        |
| CNF  | Floating point estimate of copy number (for the tumor fraction in Somatic callers). This field is not present if the model cannot be estimated with high confidence.                                       |
| CNQ  | Exact total copy number Q-score. This field is not present if the model cannot be estimated with high confidence.                                                                                          |
| MAF  | Estimate for the minor allele frequency                                                                                                                                                                    |
| MCN  | Estimated minor-haplotype copy number (for the tumor fraction in Somatic callers). This field is not present if the model cannot be estimated with high confidence.                                        |
| MCNF | Floating point estimate of minor-haplotype copy number (for the tumor fraction in Somatic callers). This field is not present if the model cannot be estimated with high confidence.                       |
| MCNQ | Minor copy number Q-score. This field is not present if the model cannot be estimated with high confidence.                                                                                                |
| OBF  | Per-segment Outlier BAF Fraction. Percentage of BAF counts which are considered "outlier" with respect to the chosen segment call. Higher values might indicate segments where BAF counts are problematic. |
| SD   | Best estimate of segment's bias-corrected read count                                                                                                                                                       |
| NCN  | Normal-sample copy number. The field is only present in somatic workflows with enabled germline-aware mode.                                                                                                |
| SCND | Difference between CN and NCN. The field is only present in somatic workflows with enabled germline-aware mode.                                                                                            |

Note: legacy depth-only callers (germline WGS/WES and somatic WES) do not include some of the FORMAT fields indicated above, due to limitation of the used legacy models. Germline WES (depth-only) also includes the following FORMAT fields:

| ID | Description                          |
| -- | ------------------------------------ |
| LR | Log10 likelihood ratio of ALT to REF |

**Note on genotype annotation in germline copy number calling (depth-only)**

Because germline copy number calling determines overall copy number rather than the copy number on each haplotype, the genotype type field contains missing values for diploid regions when CN is greater than or equal to 2. The following are examples of the GT field for various VCF entries:

| Diploid or Haploid? | ALT    | FORMAT:CN | FORMAT:GT |
| ------------------- | ------ | --------- | --------- |
| Diploid             | .      | 2         | ./.       |
| Diploid             | \<DUP> | >2        | ./1       |
| Diploid             | \<DEL> | 1         | 0/1       |
| Diploid             | \<DEL> | 0         | 1/1       |
| Haploid             | .      | 1         | 0         |
| Haploid             | \<DUP> | >1        | 1         |
| Haploid             | \<DEL> | 0         | 1         |

**Post VCF target BED**

The DRAGEN CNV pipeline can receive in input a target BED to only emit calls overlapping with the BED intervals. The post VCF target BED is provided through the `--cnv-post-vcf-target-bed` option.

#### Coverage Uniformity

The DRAGEN CNV pipeline provides a measure of the quality of the data for a sample. If using the WGS self-normalization method, the additional `CoverageUniformity` metric is present in the VCF header. The metric is only available for germline samples. The CNV pipeline assumes that post-normalization target counts are independently and identically distributed (IID). Coverage in most high-quality WGS samples is uniform enough for the CNV caller to produce accurate calls, but some samples violate the IID assumption. Issues during library preparation or sample contamination can lead to several extreme outliers and/or waviness of target counts, which can result in a large number of false positive CNV calls. The `CoverageUniformity` metric quantifies the degree of local coverage correlation in the sample to help identify poor-quality samples.

A larger value for this metric means the coverage in a sample is less uniform, which indicates that the sample has more nonrandom noise, and could be considered poor quality. The CoverageUniformity metric depends on factors other than sample quality, such as the `cnv-interval-width` setting and sample mean coverage. DRAGEN recommends using this score to compare the quality of samples from similar mean coverage and the same command line options. Because of this, DRAGEN CNV only provides the metric and does not take any action based on it.

**The CoverageUniformity metric is calculated as follows:**

1. Parse normalized counts from `tn.tsv.gz` for autosomes.
2. Split counts into disjoint tiled windows of size 100 bins.
3. Compute variance for all windows to measure observed distribution.
4. Shuffle normalized counts and compute variance for similar disjoint windows to create a background distribution.
5. Perform a two-sample Kolmogorov–Smirnov (KS) test between the observed and background variance distributions per chromosome.
6. Compute the final `CoverageUniformity` score as a weighted mean of KS test statistic across chromosomes, where weights correspond to chromosome sizes.

### CNV Metrics File

DRAGEN CNV outputs metrics in CSV format. The output follows the general convention for QC metrics reporting in DRAGEN. The CNV metrics are output to a file with a `*.cnv_metrics.csv` file extension. The following list summarizes the metrics that are output from a CNV run.

Sex Genotyper:

* Estimated sex of the case sample as well as that of all panel of normals samples are reported. For WGS workflows, the estimated sex karyotype will be reported and for non-WGS workflows the gender will be reported.
* Confidence score (ranging from 0.0 to 1.0). If the sample sex is specified, this metric is 0.0.

DRAGEN Sex Genotyper requires a minimum of 300 target intervals to confidently determine sex genotype; if the panel covers fewer intervals on the sex chromosomes, genotyping will fail and an undetermined genotype is returned. Users may lower this requirement by setting `--cnv-sex-genotyper-num-interval-requirement` to a smaller value, at the risk of increased false genotype calls.

CNV Summary:

* Bases in reference genome in use
* Average alignment coverage over genome - The average alignment coverage over the genome is calculated by dividing the total number of bases from processed alignment records (excluding those filtered by the Target Counts stage in DRAGEN CNV) by the genome length. Alignment records are filtered taking into consideration duplicate marking status (if available), MAPQ, and mapping status.
* Number of alignment records processed
  * Number of filtered records (total)
  * Number of filtered records (due to duplicates)
  * Number of filtered records (due to MAPQ)
  * Number of filtered records (due to being unmapped)
* PMAD - Pairwise Median Absolute Deviation measures the variation in read coverage between adjacent bins. It measures variability due to various factors, such as DNA degradation, extraction, amplification or library preparation. Higher values indicate noiser sample data. PMAD is calculated as following:
  * Define a vector v\[i] as normalized counts of i-th interval in log scale, and d\[i] as pairwise differences of consecutive normalized counts between i and i+1 intervals, i.e. d\[i] = (v\[i] - v\[i+1])
  * PMAD is median absolute deviation of d, i.e. PMAD = Median(|d\[i]-Median(d)|)
* Coverage MAD - Median absolute deviation of normalized case counts. Higher values indicate noiser sample data.
* Median Bin Count - Median of raw counts normalized by interval size.
* Number of target intervals
* Number of normal samples
* Number of segments
* Number of amplifications - Note: GAINLOH events (ALT=LOH and CN > 2) are also included here
* Number of deletions
* Number of CNLOHs (Copy-Neutral LOHs)
* Number of PASS amplifications - Note: GAINLOH events (ALT=LOH and CN > 2) are also included here
* Number of PASS deletions
* Number of PASS CNLOHs (Copy-Neutral LOHs)
* Post-Normalization Bin Count Sigma - Standard deviation of post-PoN-normalization median-normalized coverage values.

Coverage MAD and Median Bin Count are only printed for WES germline/somatic CNV. Post-Normalization Bin Count Sigma is only printed when PoN normalization has been applied.

### Intermediate and Visualization Files

Intermediate stages of the pipeline stages produce various intermediate output files. These files can be useful for visualization of the evidence or results from each stage, and may aid in fine-tuning options.

All files have a structure similar to a BED file with optional header line(s).

#### Target Counts File

The file `*.target.counts.gz` is a compressed tab-delimited text file that contains the number of read counts per target interval. This is the raw signal as extracted from the alignments of the BAM or CRAM file. The format is identical for both the case sample and any panel of normals samples. There is also a bigWig representation of a `target.counts.diploid` file, which is normalized to the normal ploidy level of 2 instead of raw counts.

It has the following columns:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. Count of alignments in this interval
6. Count of improperly paired alignments in this interval

Header lines are also included that start with `#`. They contain the DRAGEN version and command line that was used to generate the file, as well as other meta information.

An example of a `*.target.counts.gz` file is shown below.

```
#TARGET COUNTS FILE
##DRAGENVersion=<VERSION_INFO>
##DRAGENCommandLine=<CommandLineOptions>
...
contig  start  stop   name                <SampleName> improper_pairs
1       565480 565959 target-wgs-1-565480 7          6
1       566837 567182 target-wgs-1-566837 9          0
1       713984 714455 target-wgs-1-713984 34         4
1       721116 721593 target-wgs-1-721116 47         1
1       724219 724547 target-wgs-1-724219 24         21
1       725166 725544 target-wgs-1-725166 43         12
1       726381 726817 target-wgs-1-726381 47         14
1       753243 753655 target-wgs-1-753243 31         2
1       754322 754594 target-wgs-1-754322 27         0
1       754594 755052 target-wgs-1-754594 41         0
```

#### B-Allele counts

In germline runs, B-allele counts are calculated at bi-allelic sites taken from a collection of high-frequency SNVs in the population. In somatic runs, B-allele counts are calculated at sites in the tumor sample where the normal sample is likely to be heterozygous. When analyzed in conjunction with a matched normal sample, the sites are those that are called as heterozygous SNVs in the normal sample. When analyzed in tumor-only mode, sites are selected from a population collection (similar to germline runs). Each B-allele site consists of a reference allele and a variant allele, and the number of reads in the sample supporting each of these alleles is counted.

B-allele counts are written both to gzipped tsv file `*.ballele.counts.gz` and gzipped bedgraph file `*.baf.bedgraph.gz`.

**B-allele tsv**

The tsv file format is the following:

1. Contig identifier
2. Start, BED-style (zero-based inclusive) start position of the reference allele
3. Stop, BED-style (one-based inclusive) stop position of the reference allele
4. Base sequence for the reference allele
5. Base sequence for the first allele being counted
6. Base sequence for the second allele being counted
7. The number of qualified reads containing a sequence matching the first allele
8. The number of qualified reads containing a sequence matching the second allele

Additionally, in the case of B-allele sites from a population VCF, the following two additional columns are added after the columns listed above:

9. Population frequency for the first allele
10. Population frequency for the second allele

An example of B-allele counts file is provided below:

```
contig  start   stop    refAllele       allele1 allele2 allele1Count    allele2Count
chr1    11021   11022   G       G       A       4       2
chr1    14463   14464   A       A       T       111     36
chr1    16494   16495   G       G       C       122     262
chr1    38741   38742   C       C       T       9       9
chr1    39014   39015   A       A       C       38      48
chr1    39260   39261   T       T       C       199     143
chr1    48447   48448   C       C       T       8       15
chr1    48517   48518   A       A       G       13      15
chr1    91485   91486   G       G       C       1       4
chr1    91489   91490   A       A       G       1       3
chr1    98944   98945   C       C       T       46      114
```

**B-allele bedgraph file**

The bedgraph file format is similar to the BED format and it has the following columns:

1. Contig identifier
2. Start
3. Stop
4. Ratio of allele counts

The numerator and denominator of the ratio is determined by sorting the allele counts according to the priority of the corresponding bases. The order of the bases in descending priority is {A, T, G, C}.

When the priority of allele1 is higher than the priority of allele2, the output frequency is calculated by:

```
allele1Count / (allele1Count + allele2Count)
```

When the priority of allele2 is higher than the priority of allele1, the output frequency is calculated by:

```
allele2Count / (allele1Count + allele2Count)
```

By prioritizing the bases in this way, the output frequencies will be deterministically distributed in a roughly equal proportion above and below 0.5. When plotting these B-allele frequencies (e.g., in IGV), this gives an easy way to visually determine significant changes in b-allele frequency between neighboring segments of the genome. It also provides a similar visualization to that typically used for array data.

An example of the bedgraph file is shown below:

```
chr1    11021   11022   0.333333
chr1    14463   14464   0.755102
chr1    16494   16495   0.317708
chr1    38741   38742   0.5
chr1    39014   39015   0.44186
chr1    39260   39261   0.581871
chr1    48447   48448   0.652174
chr1    48517   48518   0.464286
```

#### Bias correction file

The file `*.target.counts.gc-corrected.gz` contains the number of GC-corrected read counts per target interval. The format is equivalent to the `*target.counts.gz` file:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. GC-corrected read counts in this interval
6. Count of improperly paired alignments in this interval

Header lines are also included that start with `#`. They contain the DRAGEN version and command line that was used to generate the file, as well as other meta information.

An example of a `*.target.counts.gc-corrected.gz` file is shown below.

```
#GC CORRECTED FILE
##DRAGENVersion=<VERSION_INFO>
##DRAGENCommandLine=<CommandLineOptions>
...
contig  start   stop    name    <SampleName> improper_pairs
chr1    818022  819840  target-wgs-chr1-818022:819840   1071.353133     6
chr1    819840  821337  target-wgs-chr1-819840:821337   1051.014997     19
chr1    821337  822485  target-wgs-chr1-821337:822485   1098.6502       10
chr1    822485  824431  target-wgs-chr1-822485:824431   1117.28308      7
chr1    830446  832304  target-wgs-chr1-830446:832304   1102.211816     1
chr1    832304  834311  target-wgs-chr1-832304:834311   1004.822683     5
chr1    836677  838659  target-wgs-chr1-836677:838659   1015.973037     7
chr1    841054  843056  target-wgs-chr1-841054:843056   1014.921403     3
```

#### Combined counts file

The file `*.combined.counts.txt.gz` is a column-wise concatenation of individual `*.target.counts.gz` and `*.target.counts.gc-corrected.gz` used to form the panel of normals.

#### Normalization file

The file `*.tn.tsv.gz` contains the normalized signal of the case sample, per target interval, i.e., the log2-transformed copy ratio signal. A strong signal deviation from 0.0 indicates a potential for a CNV event. The format is equivalent to the `*target.counts.gz` file:

1. Contig identifier
2. Start position
3. End position
4. Target interval name
5. Log2-transformed copy ratio in this interval
6. Count of improperly paired alignments in this interval

Header lines are also included that start with `#`. In some cases, the normalization counts could be patched internally with intervals from other processes, such as the SegDups extension. In such cases, patches are indicated (sorted in order of application) with header lines starting with `#patch`:

```
#patch 1 = <normalized_counts_patch_1_filename>
#patch 2 = <normalized_counts_patch_2_filename>
...
```

and the original (unpatched) `*.tn.tsv.gz` is renamed as `*.tn.unpatched.tsv.gz`. Note: this file is reported in output for inspection, but most use cases will use the (patched) `*.tn.tsv.gz` file downstream of normalization.

An example of a `*.tn.tsv.gz` file is shown below.

```
#title = Normalized coverage profile
#sex = UNDETERMINED
contig  start   stop    name    <SampleName> improper_pairs
chr1    818022  819840  target-wgs-chr1-818022:819840   -0.18479358083014644    6
chr1    819840  821337  target-wgs-chr1-819840:821337   -0.21244441644669046    19
chr1    821337  822485  target-wgs-chr1-821337:822485   -0.14849555308041734    10
chr1    822485  824431  target-wgs-chr1-822485:824431   -0.12423291178926463    7
chr1    830446  832304  target-wgs-chr1-830446:832304   -0.1438261733656668     1
chr1    832304  834311  target-wgs-chr1-832304:834311   -0.27728673450293895    5
chr1    836677  838659  target-wgs-chr1-836677:838659   -0.26136555699676262    7
```

#### Segmentation file

File extension: `*.seg`, `*.seg.called`, `*.seg.called.merged`

Files containing the segments produced by the segmentation algorithm. The `Segment_Mean` value of a segment is the ratio of the mean of that segment to the whole-sample median, without log transformation (linear copy-ratio). A strong signal deviation from 1.0 indicates a potential for a CNV event.

The `*.seg` file has the following columns:

1. Sample name
2. Contig identified
3. Start position
4. End position
5. Number of intervals in the segment
6. Linear copy-ratio of the segment

An example of a `*.seg` file is shown below.

```
Sample  Chromosome      Start   End     Num_Probes      Segment_Mean
<SampleName> chr1    818022  1117426 224     0.82500341336435279
<SampleName> chr1    1117426 4063702 2438    0.91726081432236528
<SampleName> chr1    4063702 4067591 3       0.38861386123247205
<SampleName> chr1    4067591 7705829 3302    0.93021316913709917
<SampleName> chr1    7705829 9357003 1405    0.98147825043799442
<SampleName> chr1    9357003 9377365 19      0.50269670724395654
<SampleName> chr1    9377365 12859821        2905    1.0684818476332989
```

**Germline-Specific (Depth-Only) Segmentation Output Files**

The `*.seg.called` file is identical to the `*.seg` file, with an additional column indicating the initial call for whether the segment is a duplication `+` or a deletion `-`.

The `*.seg.called.merged` file is identical to the `*.seg.called` file but with segments potentially merged when they meet internal merging criteria. In addition to the columns described above, this file has also the following columns:

8. QUAL
9. FILTER
10. Copy number assignment
11. Ploidy
12. Improper\_Pairs count

**B-allele segmentation file**

In addition to segmentation of target counts, some workflows perform segmentation of B-allele loci. The output file has suffix `*.baf.seg` and it has the same format of the `*.seg` file with two modifications. Firstly, the `Segment_Mean` value is the mean over B-allele loci of the smaller observed allele fraction. Secondly, there is an additional column:

7. `BAF_SLM_STATE`: Integer between 0 and 10, indicating bins of minor-allele fraction (low to high), or `.` when the BAF data are too variable to estimate a minor-allele fraction

An example of segmentation output file is shown below:

```
Sample  Chromosome      Start   End     Num_Probes      Segment_Mean    BAF_SLM_STATE
<SampleName> chr1    820348  1104646 194     0.29301737166888697     6
<SampleName> chr1    1105091 1533754 444     0.26185904799069076     5
<SampleName> chr1    1533810 1534166 9       0.41958837071702065     8
<SampleName> chr1    1534217 9356793 6689    0.26034515815016335     5
<SampleName> chr1    9358304 9376529 27      0.46450553586280602     10
<SampleName> chr1    9378480 12859495        1651    0.24172965924359388     5
```

#### Model identification file

In somatic callers the file `*.cnv.purity.coverage.models.tsv` describes the different tested models and their log-likelihood. It has columns:

1. Model purity (Cellularity)
2. Model diploid coverage
3. Model log-likelihood
4. Model ploidy
5. Failed constraints

An example is shown below:

```
#Purity Coverage        logL    ApproxPloidy    FailedConstraints
0.99    973     -19887608.9043  0.813   VAF_PEAKS
0.99    974     -19887600.1993  0.812   VAF_PEAKS
0.99    975     -19887591.592   0.811   VAF_PEAKS
1       115     -18380016.1561  6.98    VAF_PEAKS
1       116     -18384459.3513  6.91    VAF_PEAKS
1       117     -18390436.957   6.86    VAF_PEAKS
```

In the germline WGS caller the file `*.cnv.coverage.models.tsv` serves the same purpose. However, since germline analysis has no concept for tumor purity, the first column is set to the default value of 1.

#### Visualization files

To generate additional equivalent bigWig and gff files, set the `--cnv-enable-tracks` option to true. These files can be loaded into IGV along with other tracks that are available, such as RefSeq genes. Using these tracks alongside publicly available tracks allows for easier interpretation of calls. DRAGEN autogenerates IGV session XML file if tracks are generated by DRAGEN CNV. The `*.cnv.igv_session.xml` can be loaded directly into IGV for analysis.

The following IGV tracks are automatically populated in the output IGV session file:

* `*.target.counts.bw` --- Bigwig representation of the target counts bins. Setting the track view in IGV to barchart or points is recommended. Values are gc-corrected if gc-correction was performed.
* `*.improper_pairs.bw` --- BigWig representation of the improper pairs counts. Setting the track view in IGV to barchart is recommended.
* `*.tn.bw` --- BigWig representation of the tangent normalized signal. Setting the track view in IGV to points is recommended.
* `*.seg.bw` --- BigWig representation of the segments. Setting the track view in IGV to points is recommended.
* `*.baf.seg.bw` --- BigWig representation of the BAF segments (if available). Setting the track view in IGV to points is recommended.
* `*.baf.bedgraph.gz` --- BED graph representation of B-allele frequency (if available). Setting the track view in IGV to points is recommended.
* `*.cnv.gff3` --- GFF3 representation of the CNV events. DEL events show as blue and DUP events show as red. Filtered events are a light gray. If REF events are enabled, then they will show up as green. When the caller can call AOH/LOH events, they will show up as magenta. An example of DRAGEN CNV gff3 is shown below (different CNV workflows might output different attributes on the 9th column):

```
##gff-version 3
chr1    DRAGEN  LOSS    12779193        12859821        30      .       .       Alt=DEL;LinearCopyRatio=0.576;CopyNumber=1;Genotype=0/1;Qual=30;Filter=PASS;Start=12779192;Stop=12859821;Length=80629;BinCount=24;ImproperPairsCount=16,7;color=#0000FF;
chr1    DRAGEN  REF     13106280        13122338        19      .       .       Alt=REF;LinearCopyRatio=1.05981;CopyNumber=2;Genotype=./.;Qual=19;Filter=PASS;Start=13106279;Stop=13122338;Length=16059;BinCount=8;ImproperPairsCount=3,1;color=#00FF00;
chr1    DRAGEN  GAIN    13225213        13247040        66      .       .       Alt=DUP;LinearCopyRatio=2.016;CopyNumber=4;Genotype=./1;Qual=66;Filter=PASS;Start=13225212;Stop=13247040;Length=21828;BinCount=9;ImproperPairsCount=7,5;color=#FF0000;
```

For somatic WGS analyses, the following additional files are included in the IGV session xml:

* `*.tumor.baf.bedgraph.gz` --- Bedgraph representation of the B-allele frequencies. Setting the track view in IGV to points and windowing function to none is recommended.

**IGV Session**

![](/files/WGi22zHxhS0kedSW46he)

File extension: `*.igv_session.xml`

The IGV session XML file is prepopulated with track files generated by DRAGEN. The session file loads the reference genome that best matches the standard reference genomes in an IGV installation, by comparing the name of the `--ref-dir` specified on the command-line. Standard UCSC human reference genomes are autodetected, but any variations from the standard reference genomes might not be autodetected. To edit the genome detection, alter the `genome` attribute in the `Session` element to the reference genome you would like for analysis before loading into IGV. The reference identifier used by IGV might differ from the actual name of the genome. The following is an example edited session file.

```
<?xml version="1.0" encoding="utf-8"?>
<Session genome="b37" hasGeneTrack="false" hasSequenceTrack="true" version="8">
    <Resources>
        <Resource path="example.cnv.gff3"/>
        <Resource path="example.cnv.excluded_intervals.bed.gz"/>
        <Resource path="example.target.counts.bw"/>
        <Resource path="example.improper.pairs.bw"/>
        <Resource path="example.tn.bw"/>
        <Resource path="example.seg.bw"/>
    </Resources>
    <Panel height="500" width="1200" name="DataPanel">
        ...
    </Panel>
</Session>
```

Note that depending on the IGV version installed, it may come prepackaged with different flavors of GRCh37. The reference naming conventions have changed so a user may have to edit the `genome` field in the XML file directly. For example, IGV has traditionally packaged a `b37` reference genome, but may also include a `1kg_v37` or a `1kg_b37+decoy`, which will appear on the IGV user interface as "1kg, b37" or "1kg, b37+decoy" respectively.

You can determine what the correct encoding of a reference genome by going to `File > Save Session...` and then inspecting the generated igv\_session.xml file.

When the Cytogenetics Modality is enabled, DRAGEN CNV produces an additional IGV session xml `*.cyto.igv_session.xml` shown below. Please see [Cytogenetics Modality](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/cnv-overview/cnv-germline#cytogenetics-modality) for a description of the different tracks on this file.

![](/files/4f11wpwXtvo98GriV38t)

**Creating CNV coverage and BAF plots with third-party tools**

DRAGEN CNV outputs can be ingested using third-party libraries on most commonly used languages such as Python/R. The typically used files are:

* `*.target.counts.gz` or `*.target.counts.gc-corrected.gz`, containing the number of alignments, or corrected alignments, per interval. Used to display the coverage profile across all intervals.
* `*.tn.tsv.gz`, containing the log2-normalized copy ratio per interval.
* `*.baf.bedgraph.gz`, if BAF is available, containing the BAF for each considered site. Used to display the BAF profile across all sites.

In all previously specified files, the format is similar to BED, allowing them to be loaded as any other tab-separated files.

Using R, a good starting point is the [karyoploteR](https://bernatgel.github.io/karyoploter_tutorial//Examples/SNPArray/SNPArray.html) package. The main workflow involves reading the `*.target.counts.gz` file as an R dataframe, convert this to a GRanges object then plot the target intervals as points with the `karyoploteR` package. The same workflow can be used to plot the GC-corrected counts, the log2 normalized copy ratios and the BAF profiles.

![](/files/L1OP7DhtUWOIuNLkWW9M)

Using Python, the workflow is similar to R's but using Python's libraries such as [pandas](https://pandas.pydata.org/), to convert DRAGEN output files to dataframe, and [matplotlib](https://matplotlib.org/), to plot coverage and BAF profiles across the genome.

***

A similar workflow can be used to plot copy number calls (and minor copy number calls, if available) by using the `*.cnv.gff3` output file. Some examples of DRAGEN output GFF3 are shown below:

**Germline WGS**

```
chr1    DRAGEN  LOSS    12779193        12859821        30      .       .       Alt=DEL;LinearCopyRatio=0.576;CopyNumber=1;Genotype=0/1;Qual=30;Filter=PASS;Start=12779192;Stop=12859821;Length=80629;BinCount=24;ImproperPairsCount=16,7;color=#0000FF;
chr1    DRAGEN  REF     13106280        13122338        19      .       .       Alt=REF;LinearCopyRatio=1.05981;CopyNumber=2;Genotype=./.;Qual=19;Filter=PASS;Start=13106279;Stop=13122338;Length=16059;BinCount=8;ImproperPairsCount=3,1;color=#00FF00;
chr1    DRAGEN  GAIN    13225213        13247040        66      .       .       Alt=DUP;LinearCopyRatio=2.016;CopyNumber=4;Genotype=./1;Qual=66;Filter=PASS;Start=13225212;Stop=13247040;Length=21828;BinCount=9;ImproperPairsCount=7,5;color=#FF0000;
```

**Somatic WGS**

```
chr1    DRAGEN  GAIN    16605768        16949283        237     .       .       Start=16605769;Stop=16949283;Length=343515;Alt=<DUP>;Qual=237;Filter=PASS;Genotype=1/1;CopyNumber=4;MinorCopyNumber=2;CopyNumberQual=1;MinorCopyNumberQual=1;CopyNumberFloat=4.371887;MinorCopyNumberFloat=2.000000;BiasCorrectedReadCount=1182.6;MinorAlleleFrequency=0.5;BinCount=74;ImproperPairsCount=15,17;NumAllelicSites=223;color=#FF0000;
chr1    DRAGEN  CNLOH   16949283        23272950        1000    .       .       Start=16949284;Stop=23272950;Length=6323667;Alt=<LOH>;Qual=1000;Filter=PASS;Genotype=1/1;CopyNumber=2;MinorCopyNumber=0;CopyNumberQual=1000;MinorCopyNumberQual=1000;CopyNumberFloat=2.090573;MinorCopyNumberFloat=0.000000;BiasCorrectedReadCount=565.5;MinorAlleleFrequency=0;BinCount=5572;ImproperPairsCount=17,84;NumAllelicSites=2517;color=#FF00FF;
chr1    DRAGEN  LOSS    23272950        25394644        1000    .       .       Start=23272951;Stop=25394644;Length=2121694;Alt=<DEL>;Qual=1000;Filter=PASS;Genotype=0/1;CopyNumber=1;MinorCopyNumber=0;CopyNumberQual=1000;MinorCopyNumberQual=1000;CopyNumberFloat=1.069501;MinorCopyNumberFloat=0.000000;BiasCorrectedReadCount=289.3;MinorAlleleFrequency=0;BinCount=1718;ImproperPairsCount=84,5;NumAllelicSites=872;color=#0000FF;
```

From the output GFF3, the typical steps to follow are to parse each segment coordinates and the `CopyNumber` annotation (or any other annotation the user might want to plot), and to plot them using the libraries listed previously for coverage/BAF profiles (or any other library and language of user's choice).

### Excluded Intervals File

To improve accuracy, the DRAGEN CNV Pipeline excludes genomic intervals if one or more of the target intervals failed at least one quality requirement. The excluded intervals are reported to `*.cnv.excluded_intervals.bed.gz` file. The file has a bed format, identifies the regions of the genome that are not callable for CNV analysis and describes the reason intervals were excluded in the fourth column. The following are the possible reasons for exclusion.

| Exclusion Reason                 | Description                                                                     | Related DRAGEN Option                                                |
| -------------------------------- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------- |
| NON\_KMER\_UNIQUE                | Non-unique Kmer bases are larger than 50% of interval.                          | Not applicable. This reason only applies to self-normalization mode. |
| EXCLUDE\_BED                     | Interval overlaps with exclude BED larger than threshold.                       | `--cnv-exclude-bed-min-overlap`                                      |
| PON\_MAX\_PERCENT\_ZERO\_SAMPLES | Number of PON samples with 0 coverage is larger than threshold.                 | `--cnv-max-percent-zero-samples`                                     |
| PON\_TARGET\_FACTOR\_THRESHOLD   | Median coverage of interval is lower than threshold of overall median coverage. | `--cnv-target-factor-threshold`                                      |
| PON\_MISSING\_INTERVAL           | Target interval not found in PON.                                               | Not applicable                                                       |

An example of a `*.cnv.excluded_intervals.bed.gz` file is shown below:

```
chr1    0       818022  NON_KMER_UNIQUE
chr1    824431  830446  NON_KMER_UNIQUE
chr1    834311  836677  NON_KMER_UNIQUE
chr1    838659  841054  NON_KMER_UNIQUE
chr1    850451  853257  NON_KMER_UNIQUE
chr1    855442  860261  NON_KMER_UNIQUE
chr1    866189  868833  NON_KMER_UNIQUE
chr1    881779  884116  NON_KMER_UNIQUE
chr1    1016667 1018959 NON_KMER_UNIQUE
chr1    1075880 1079718 NON_KMER_UNIQUE
chr1    1137942 1140725 NON_KMER_UNIQUE
```

### Excluded Samples File

To improve accuracy, the DRAGEN CNV Pipeline excludes panel of normals samples if one or more of the samples failed at least one quality requirement. The excluded samples are reported to `*.cnv.excluded_samples.txt.gz` file. The file has a tsv (tab separated) format, identifies the excluded panel of normals samples and describes the reason. The following are the possible reasons for exclusion.

| Exclusion Reason                          | Description                                                  | Related DRAGEN Option                       |
| ----------------------------------------- | ------------------------------------------------------------ | ------------------------------------------- |
| PON\_SAMPLE\_NAME\_EQUAL\_TO\_CASE        | PON sample name is equal to case sample name                 | NA                                          |
| PON\_SAMPLE\_CORRELATION\_EQUAL\_TO\_CASE | PON sample counts are equal to case sample counts            | NA                                          |
| PON\_MAX\_PERCENT\_NAN\_SAMPLES           | number of nan values in sample is higher than threshold      | `--cnv-max-percent-nan-samples`(default=50) |
| MAX\_PERCENT\_ZERO\_TARGETS               | number of 0 target counts in sample is higher than threshold | `--cnv-max-percent-zero-targets`(default=5) |
| EXTREME\_PERCENTILE:UPPER                 | median coverage of sample is higher than threshold           | `--cnv-extreme-percentile`(default=2.5)     |
| EXTREME\_PERCENTILE:LOWER                 | median coverage of sample is lower than threshold            | `--cnv-extreme-percentile`(default=2.5)     |

An example of a `*.cnv.excluded_samples.txt.gz` file is shown below:

```
#name        reason                                 value        threshold
Sample1      MAX_PERCENT_ZERO_TARGETS               4776         418
Sample2      EXTREME_PERCENTILE:LOWER               0.000812534  0.20065
Sample3      EXTREME_PERCENTILE:UPPER               1.0003       1.00025
Sample4      PON_SAMPLE_NAME_EQUAL_TO_CASE          NA           NA
Sample5      PON_SAMPLE_CORRELATION_EQUAL_TO_CASE   NA           NA
```

The excluded samples output file may not exist if there are no excluded samples.

### Panel of Normals Files

#### PON Metrics File

The DRAGEN CNV Pipeline generates the PON Metrics File (`.cnv.pon_metrics.tsv.gz`) if a Panel of Normals is provided and `--cnv-generate-pon-metric-file` is set to `true`. If PON size is less than 2, then an empty file will be generated.

The PON Metric File includes basic statistics of the coverage profile for each interval. To remove sample coverage bias, DRAGEN applies sample median normalization, and then computes the following metrics:

| Column index | Column contents | Description                              |
| ------------ | --------------- | ---------------------------------------- |
| 1            | contig          | chromosome name                          |
| 2            | start           | genomic locus of interval start          |
| 3            | stop            | genomic locus of interval stop           |
| 4            | name            | interval name                            |
| 5            | mean            | average coverage depth                   |
| 6            | std             | standard deviation                       |
| 7            | normalizedStd   | normalized standard deviation (std/mean) |
| 8            | min             | minimum                                  |
| 9            | 25%             | 25 percentile                            |
| 10           | 50%             | median                                   |
| 11           | 75%             | 75 percentile                            |
| 12           | max             | maximum                                  |
| 13           | intervalSize    | interval size (stop-start)               |
| 14           | gcContents      | percent GC                               |

Example:

```
contig  start   stop    name    mean    std     normalizedStd min     25%     50%     75%     max     intervalSize    gcContents
1       12098   12178   target-wes-1-12098:12178/1      3.6259044560802365      0.46661435469856077      0.1286890927079175     2.7961783439490446      3.2573018790849675      3.7105263157894739      4.0162683823529415      4.3298969072164946      80      0.49382716049382713
1       12178   12258   target-wes-1-12178:12258/2      5.0685579775753595      0.70638315915955963      0.13936570564740217     3.9044585987261144      4.5225944682508761      5.067708333333333       5.5778115844038769      6.3277777777777775      80      0.46913580246913578
1       12553   12637   target-wes-1-12553:12637/1      4.6990858287992054      0.62537786269786677      0.13308500535681309     3.7417218543046356      4.0305632538350444      5.0382165605095546      5.2151580459770113      5.5773195876288657      84      0.6705882352941176
```

#### PON Correlation File

The DRAGEN CNV Pipeline generates the PON Correlation File (`.cnv.pon_correlation.txt.gz`) if a Panel of Normals is provided. The PON Correlation File includes correlation between CASE sample and each PON sample.

Example:

```
Correlation of case sample CASE_SAMPLE_NAME
  PON1: 0.9786
  PON2: 0.9868
  PON3: 0.9912
  ...
```

### SegDups Extension Files

The SegDups extension provides intermediate and final outputs. All intervals follow the bed format (0-based, start inclusive, end exclusive) and they are in tab-delimited text files (gzip compressed).

The final output has extension `.cnv.segdups.rescued_intervals.tsv.gz`, and contains the rescued target intervals which can then be injected before segmentation. It has columns:

1. Chromosome name
2. Start - 0-based inclusive
3. Stop - 0-based exclusive
4. Target interval name prefixed with "target-wgs-"
5. Sample Counts (in header, identifier taken from RGSM) - log2-scale normalized counts for each interval
6. Improper pairs - Kept for compatibility with CNV workflow, set to 0 for rescued intervals
7. Target region ID - ID of the target region (aka pair of rescued target intervals)

#### Intermediate files

The joint normalized coverage profile (log2-scale) for each region is provided in output to file `.cnv.segdups.joint_coverage.tsv.gz` with columns:

1. Target region ID
2. Joint normalized coverage (log2-scale) of the two intervals in the target region
3. Copy Number Float - estimate of joint copy number for the target region (e.g., CNF \~ 4)

The differentiating sites' data is provided in output to file `.cnv.segdups.site_ratios.tsv.gz` with columns:

1. Differentiating site name
2. Target (gene A) counts at site
3. Non-target (gene B) counts at site
4. Target ratio: gene A counts over total (i.e., gene A + gene B) counts at site


# Repeat Expansion Detection

Short tandem repeats (STRs) are regions of the genome consisting of repetitions of short DNA segments called repeat units. STRs can expand to lengths beyond the normal range and cause mutations called repeat expansions. Repeat expansions are responsible for many diseases, including Fragile X syndrome, amyotrophic lateral sclerosis, and Huntington's disease.

DRAGEN includes a repeat expansion detection tool for STRs, DRAGEN-STR. DRAGEN-STR Performs sequence-graph based realignment of reads that originate inside and around each target repeat. DRAGEN-STR then genotypes the length of the repeat in each allele based on these graph alignments.

DRAGEN-STR is designed for PCR-free whole genome samples. Repeats are only genotyped if the coverage at the locus is at least 10x, but a minimum of 30x is recommended. Sequencing reads must be paired-end with a minimum read length of 100 (2x100bp). DRAGEN-STR cannot be run on multiple FASTQ files that are assigned to different library IDs in the `fastq_list.csv` file.

DRAGEN-STR does not support somatic analysis.

> NOTE:
>
> DRAGEN STR is based on the ExpansionHunter tool. For more information about implementation details and performance assessment refer to these [publications](/dragen-v4.5/reference/citing-dragen#str-expansion-detection).

## Repeat Expansion Detection Options

To enable DRAGEN repeat expansion detection, the following command-line options are required.

* `--repeat-genotype-enable=true`
* `--repeat-genotype-specs=<path to specification file>`

You can use the `--sample-sex` option to specify the sex of the sample. The following options are optional.

* `--repeat-genotype-region-extension-length=<length of region around repeat to examine>` (default 1000 bp)
* `--repeat-genotype-min-baseq=<Minimum base quality for high confidence bases>` (default 20)

For more information on the specification file specified by `--repeat-genotype-specs` option, see [Repeat Expansion Specification Files](#repeat-expansion-specification-files).

The main output of repeat expansion detection is a VCF file that contains the variants found via this analysis.

## Repeat Expansion Specification Files

The repeat-specification (also called variant catalog) JSON file defines the repeat regions for DRAGEN-STR to analyze. Default repeat-specification for some pathogenic and polymorphic repeats are in the `<INSTALL_PATH>/resources/repeat-specs/` directory, based on the reference genome used with DRAGEN.

You can create specification files for new repeat regions by using one of the provided specification files as a template. See the [catalog documentation](https://github.com/Illumina/ExpansionHunter/blob/master/docs/04_VariantCatalogFiles.md) for details on the format.

`--repeat-genotype-specs` is required for DRAGEN-STR. If the option is not provided, DRAGEN attempts to autodetect the applicable catalog file from `<INSTALL_PATH>/resources/repeat-specs/` based on the reference provided.

Users can choose between any of the three default repeat-specification files packaged with DRAGEN using the command line option: `--repeat-genotype-use-catalog=<default|default_plus_smn|expanded>`. The `default` option includes \~60 repeats. The `default_plus_smn` option includes the SMN repeat in addition to all the repeats in the `default` catalog. The expanded catalog includes \~174K repeats, see [Covered Repeat Regions](#covered-repeat-regions). If `--repeat-genotype-use-catalog` is not specified on the command line, then the `default` catalog is used.

The repeat genotyping results will be incorrect if the selected reference genome is not compatible with the repeat specification file. When this occurs, many repeats may be marked as "LowDepth" in the VCF output file or estimated to have zero length. This can be further confirmed by visualizing read alignments with the [REViewer visualization tool](https://github.com/Illumina/REViewer).

## Covered Repeat Regions

The `default` variant catalog contains specifications on disease-causing repeats located in AFF2, AR, ARX\_1, ARX\_2, ATN1, ATXN1, ATXN10, ATXN2, ATXN3, ATXN7, ATXN8OS, BEAN1, C9ORF72, CACNA1A, CBL, CNBP, COMP, CSTB, DAB1, DIP2B, DMD, DMPK, EIF4A3, FMR1, FOXL2, FXN, GIPC1, GLS, HOXA13\_1, HOXA13\_2, HOXA13\_3, HOXD13, HTT, JPH3, LRP12, MARCHF6, NIPA1, NOP56, NOTCH2NLC, NUTM2B-AS1, PABPN1, PHOX2B, PPP2R2B, PRDM12, PRNP, RAPGEF2, RFC1, RUNX2, SAMD12, SOX3, STARD7, TBP, TBX1, TCF4, TNRC6A, VWA1, XYLT1, YEATS2, ZIC2 and ZIC3 genes. More information about disease-causing repeats can also be found [here](https://gnomad.broadinstitute.org/short-tandem-repeats?dataset=gnomad_r3).

For the `expanded` variant catalog, apart from the aforementioned disease-causing repeats, there are \~174K additional polymorphic repeats. They are initially detected using STR-Finder from the 1000 Genomes Project. After that, the candidate repeats are filtered out based on a customized quality control pipeline, see details [here](https://github.com/Illumina/RepeatCatalogs).

DRAGEN-STR can detect pathogenic expansions of FXN, ATXN3, ATN1, AR, DMPK, HTT, FMR1, ATXN1, C9ORF72 repeats with high accuracy (see [publications](/dragen-v4.5/reference/citing-dragen#str-expansion-detection)). The pathogenicity status of some repeats might depend on the presence of sequence interruptions or motif changes that DRAGEN-STR does not call. If you would like to visually inspect the relevant read alignments, you can use a Repeat Expansion Viewer third-party tool.

## Repeat Expansion Detection Output Files

### VCF Output File

The results of repeat genotyping are output as a separate VCF file, which provides the length of each allele at each callable repeat defined in the repeat-specification catalog file. The name is `<outputPrefix>.repeats.vcf` (\*.gz). The VCF output file lists with the following fields first.

Table 2 Core VCF Fields

| Field  | Description                                                                                                                         |
| ------ | ----------------------------------------------------------------------------------------------------------------------------------- |
| CHROM  | Chromosome identifier                                                                                                               |
| POS    | Position of the first base before the repeat region in the reference                                                                |
| ID     | Always `.`                                                                                                                          |
| REF    | The reference base at position POS                                                                                                  |
| ALT    | List of repeat alleles in format `<STRn>` . N is the number of repeat units. If REF, then `.`.                                      |
| QUAL   | Always `.`                                                                                                                          |
| FILTER | LowDepth filter is applied when the overall locus depth is below 10x or number of reads that span one or both breakends is below 5. |

Table 3 Additional INFO Fields

| Field | Description                                                     |
| ----- | --------------------------------------------------------------- |
| END   | Position of the last base of the repeat region in the reference |
| REF   | Number of repeat units spanned by the repeat in the reference   |
| RL    | Reference length in bp                                          |
| VARID | Variant ID from the variant catalog                             |
| RU    | Repeat unit in the reference orientation                        |
| REPID | Variant ID from the variant catalog                             |

Table 4 GENOTYPE (Per Sample) Fields

| Field | Description                                                                                                                                                                 |
| ----- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| GT    | Genotype                                                                                                                                                                    |
| SO    | Type of reads that support the allele. Values can be SPANNING, FLANKING, or INREPEAT. These values indicate if the reads span, flank, or are fully contained in the repeat. |
| REPCN | Number of repeat units spanned by the allele                                                                                                                                |
| REPCI | Confidence interval for REPCN                                                                                                                                               |
| ADSP  | Number of spanning reads consistent with the allele                                                                                                                         |
| ADFL  | Number of flanking reads consistent with the allele                                                                                                                         |
| ADIR  | Number of in-repeat reads consistent with the allele                                                                                                                        |
| LC    | Locus Coverage                                                                                                                                                              |

For example, the following VCF entry describes the ATXN1 repeat in a sample NA13537.

```
#CHROM  POS     ID      REF     ALT     QUAL    FILTER  INFO    FORMAT  NA13537
chr6	16327864	.	G	<STR33>,<STR58>	.	PASS	END=16327954;REF=30;RL=90;RU=TGC;VARID=ATXN1;REPID=ATXN1	GT:SO:REPCN:REPCI:ADSP:ADFL:ADIR:LC	1/2:SPANNING/INREPEAT:33/58:33-33/52-71:4/0:69/83:0/4:37.459459
```

In this example, the first allele spans 33 repeat units while the second allele spans 58 repeat units. The repeat unit is TGC (RU INFO field), so the sequence of the first allele is TGC x 33 and the sequence of the second allele is TGC x 58. The repeat spans 30 repeat units in the reference (REF INFO field).

The length of the short allele was estimated from spanning reads (SPANNING) while the length of the expanded allele was estimated from in-repeat reads (INREPEAT). The confidence interval for the size of the expanded allele is (52,71). There are 4 spanning and 69 flanking reads consistent with the repeat allele of size 33 that is 4 reads fully contain the repeat of size 33 and 69 flanking reads overlap at most 33 repeat units. There are 83 flanking and 4 in-repeat reads consistent with the repeat allele of size 58. The average coverage of this locus is 37.46x.

### &#x20;Additional Output Files

The sequence-graph alignments of reads in the targeted repeat regions are output in a BAM file. You can use a specialized GraphAlignmentViewer tool available on GitHub to visualize the alignments. Programs like Integrative Genomics Viewer (IGV) are not designed for displaying graph-aligned reads and cannot visualize these BAMs.

The BAMs store graph alignments in custom XG tags using the format `<LocusName>,<StartPosition>,<GraphCIGAR>`.

* **LocusName**---A locus identifier that matches the corresponding entry in the repeat expansion specification file.
* **StartPosition**---The starting alignment position of a read on the first graph node.
* **GraphCIGAR**---The alignment of a read against the graph starting from that position. GraphCIGAR consists of a sequence of graph node identifiers and linear CIGARS describing the alignment of the read to each node.

Quality scores in the BAM file are binary. High-scoring bases are assigned a score of 40, and low-scoring bases are assigned a score of 0.

## Motif analysis

Some STR loci have polymorphic motifs, meaning that the different repetitions of the motif have minor variations in their sequence. For example, an STR may contain some repetitions of AAAGG and some repetitions of AAGGG, all in the same haplotype.

In some cases, the motif composition affects the pathogenic threshold: a long expansion of one motif variant can be harmless while a short expansion of another can be pathogenic. The main example of this is the [RFC1 locus](https://pmc.ncbi.nlm.nih.gov/articles/PMC10689911/).

DRAGEN can estimate the global motif composition of STR loci and in some cases, if the STR expansion is heterozygous, also the per-allele composition, by leveraging reads that are only compatible with the estimated repeat size on one haplotype but not the other.

***

DRAGEN has two workflows to compute motif *fractions*:

Kmer-counting : Kmers extracted from the graph-alignment of each read to the locus graph are used to detect and count the motifs in the locus.

HMM labeling : The repetitive patterns in each read are labeled using an HMM model generated from a set of possible motifs passed on by the user.

***

There are advantages and disadvantages to each technique:

* Kmer counting does not need to know the set of motifs a priori, but it will only count kmers the same size as the pattern.
* HMM labeling needs the set of possible motifs to be known a priori, but it will perform much better when the set contains motifs of varying length.

**NOTE** If the set of motifs is known, it is always advisable to leverage the more accurate HMM labeling over the kmer-counting.

### Enabling motif analysis

Motif analysis is activated on a per-locus basis by adding a field to the respective catalog entry.

**NOTE** Either HMM or kmer motif analysis can be enabled, not both.

#### HMM labeling example

```yaml
{
     "LocusId": "RFC1",
     "LocusStructure": "(AARRG)*",
     "ReferenceRegion": "4:39348424-39348479",
     "VariantType": "Repeat",
     "MotifAnalysis": ["AAAAG", "AAAGG", "AAGGG", "AAGAG", "AACGG", "ACGGG", "ACAGG", "AAAGGG"], # for HMM labeling
}
```

#### Kmer-counting example

```yaml
{
     "LocusId": "RFC1",
     "LocusStructure": "(AARRG)*",
     "ReferenceRegion": "4:39348424-39348479",
     "VariantType": "Repeat",
     "MotifAnalysis": "kmer" # for kmer-counting
}
```

\======= **IMPORTANT** The HMM motif analysis is enabled by default only for the *RFC1* locus in built-in DRAGEN catalogs with the following motifs set:

```yaml
"MotifAnalysis": ["AAAAG", "AAAGG", "AAGGG", "AAGAG", "AACGG", "ACGGG", "ACAGG", "AAAGGG"]
```

### Output VCF

#### Additional `FORMAT` fields

When motif analysis is enabled on a locus, the following `FORMAT` fields will be present in the VCF record:

| Field    | Description                                                                                |
| -------- | ------------------------------------------------------------------------------------------ |
| `MOTIFS` | Set of high-quality motifs detected from graph-aligned reads                               |
| `MF`     | Fraction of quality weighted counts for each motif in `FORMAT:MOTIFS`                      |
| `AMF`    | Fraction of quality weighted counts for each motif in `FORMAT:MOTIFS` stratified by allele |

#### Example VCF record

```
:MOTIFS:MF:AMF AAAAG,AAAGG:0.15,0.75:0.454/0.189,1.000/0.000
```

#### The output in detail

To reduce false positives due to drops in base quality, DRAGEN has to detect a motif in high quality parts of the read to report it. When there are no high-quality motifs in the sample, the corresponding `MF` and `AMF` field will be `0` when using HMM labeling, or not reported in the VCF when kmer-counting.

If there is no genotype call, the `MF` and `AMF` fields will be empty. If a genotype call is present, but the genotype is homozygous, the `AMF` fields will be empty. Empty fields will be encoded with a dot (`.`).


# De Novo Repeat Expansion Detection

Short tandem repeats (STRs) are regions of the genome consisting of repetitions of short DNA segments called repeat units. STRs can expand to lengths beyond the normal range and cause mutations called repeat expansions. Repeat expansions are responsible for many diseases, including Fragile X syndrome, amyotrophic lateral sclerosis, and Huntington's disease.

STR profiling allows the discovery of expanded STR regions from paired-end reads across a cohort of samples. It is designed to work with PCR-free samples of 100-200bp paired-end reads at >30X coverage.

> **Note:**
>
> * STRs shorter than the read length are ignored; the program is appropriate only for detecting expansions that exceed the read length.
> * The location of each reported STR is approximate (up to about 500bp-1Kbp)
> * STRs are not genotyped; the program reports a depth-normalized count of reads originating inside each STR; this count can be used as a very approximate measure of the repeat length
> * To achieve best results all samples must be sequenced on the same instrument to similar coverage, have the same read and fragment lengths, and be subjected to the same computational pre-processing (e.g. reads must be aligned by the same aligner)

Briefly, the workflow can be separated in two distinct steps: profiling and analysis. In the profiling step, repetitive reads are found and used to infer the location of potential STR regions. The regions and the respective read counts are then saved in a "profile" on disk. The profiling step is run for each sample and the resulting profiles are merged into a single dataset for the analysis. In the analysis step the user needs to provide a table describing the experimental design to run either an outlier analysis which tests one sample against the rest or a case-control analysis where the samples are split in two groups.

## Running DRAGEN

The two steps of the workflow, profiling and analysis, are performed by two separate DRAGEN commands.

### Profiling

In the first step we compute the profiles which are going to be saved as ProtoBuf messages (`<out_prefix>.data`). The profile can be saved in a specific directory with the `--str-profiler-output-directory` flag. The sample name will be saved in the profile and can be specified at the profiling stage with the flag `--str-profiler-sample-name`. If not specified, the sample name in the RGSM field will be used instead.

DRAGEN has to be called once for each sample, for example with the command:

```bash
dragen \
  --enable-map-align=false \
  --output-directory=<out_dir>  \
  --output-file-prefix=<out_prefix>  \
  --bam-input <bam_input> \
  --enable-str-profiler=true \
  --str-profiler-sample-name=<optional_name> \            # if not set is == to RGSM
  --str-profiler-output-directory=<path_to_directory> \   # if not set is == to --output-directory
  -r=<dragen_ref>
```

After all the profiles are computed, they have to be divided in 'cases' and 'controls' directories. This can be achieved while computing the profiles by passing the directory with the `--str-profiler-output-directory` flag. The input can be a list of samples with the `--fastq-list` option. DRAGEN can take as input a list of FASTQ files and save each profile in the directory specified directory with `--str-profiler-output-directory`. A list of cases and a list of controls can be run in this manner.

Example command:

```bash
dragen \
  --enable-map-align=true \
  --output-directory=<out_dir>  \
  --output-file-prefix=<out_prefix>  \
  --fastq-list=<fastq_list> \
  --fastq-list-all-samples=true \                # necessary to run samples separately based on sample-id
  --enable-multi-sample=true \                   # necessary to run samples separately based on sample-id
  --enable-per-sample-map-align-output=true \    # can be set to false if map align output (BAM) is not needed
  --enable-str-profiler=true \
  --str-profiler-output-directory=<path_to_directory> \
  -r=<dragen_ref>
```

### Analysis

The analysis is performed with a separate DRAGEN command, which takes as input the path to the two directories.

Two analysis types can be specified:

* `outlier` = bootstraps the sampling distribution of the 95% quantile and then calculates the z-scores for the cases samples
* `casecontrol` = cases and controls counts are compared with a one-sided Wilcoxon rank-sum test and a Bonferroni correction is applied to the resulting p-values

Providing the `--str-profile-analysis flag` will trigger the analysis workflow. Example command:

```bash
dragen \
  --output-directory=<out_dir>  \
  --output-file-prefix=<out_prefix>  \
  --enable-str-profiler=true \
  --str-profiler-analysis=<outlier|casecontrol> \
  --str-profiler-cases-directory=<directory_with_cases_profiles> \
  --str-profiler-controls-directory=<directory_with_controls_profiles> \
  -r=<dragen_ref>
```

#### A note about bootstrapping

DRAGEN uses bootstrapping to approximate the distribution for the outlier analysis with a default number of iteration of 1000. This number can be adjusted with the flag `--str-profiler-resampling-rounds`. Increasing the number of resampling cycles will improve the accuracy of the approximation but also linearly increase the compute times.

DRAGEN will spread the computation across 48 threads by default, but the number can be adjusted on the command line with the flag `--str-profiler-threads`.

### Output

The output is composed of two tables, one for the "motif" level analysis and one for the "locus" level analysis which will be saved as `<output-prefix>.str_profiler_locus.tsv` and `<output-prefix>.str_profiler_motif.tsv` respectively. Below is a description of the locus analysis output. The motif table is the same as the locus table but without the ***contig***, ***start*** and ***end*** columns.

#### Outlier analysis (locus) output

| Column             | Description                                                      |
| ------------------ | ---------------------------------------------------------------- |
| contig             | Contig of the repeat region                                      |
| start              | Approximate start of the repeat                                  |
| end                | Approximate end of the repeat                                    |
| motif              | Inferred repeat motif                                            |
| top\_case\_zscore  | Top z-score of a case sample                                     |
| high\_case\_counts | Counts of case samples corresponding to z-score greater than 1.0 |
| counts             | Nonzero counts for all samples                                   |

#### Case-control analysis (locus) output

| Column       | Description                                                                                            |
| ------------ | ------------------------------------------------------------------------------------------------------ |
| contig       | Contig of the repeat region                                                                            |
| start        | Approximate start of the repeat                                                                        |
| end          | Approximate end of the repeat                                                                          |
| motif        | Inferred repeat motif                                                                                  |
| pvalue       | P-value from Wilcoxon rank-sum test                                                                    |
| bonf\_pvalue | P-value after Bonferroni correction                                                                    |
| counts       | Depth-normalized counts of anchored in-repeat reads for each sample (omitting samples with zero count) |


# Targeted Caller

Repetitive regions in the human genome pose a challenge for general variant calling approaches which typically cannot make use of potentially misplaced MAPQ0 reads. Furthermore, high sequence homology of some genes with a pseudogene paralog can lead to a wide variety of common structural variants (SVs) in the population, requiring specialized targeted calling approaches. DRAGEN supports targeted calling for a number of genes/targets as described in subsequent target-specific sections.

The targeted caller can be enabled using the command line option `--enable-targeted=true` or a subset of targets can be enabled by providing a space-separated list of target names. The supported target names for WGS are: `cyp2b6`, `cyp2d6`, `cyp21a2`, `gba`, `hba`, `lpa`, `rh`, and `smn`. For WGS TruPath data, only `lpa`, `hba`, and `smn` will run when the Targeted Caller is enabled, but a custom list of supported targets can be specified on the command line. The supported target names for WES are: `hba` and `smn`. For a list of all supported targeted caller options along with their default values, see [Targeted Caller Options](/dragen-v4.5/product-guides/dragen-v4.5/command-line-options#targeted-caller-options). The targeted caller produces a `<output-file-prefix>.targeted.json` file containing a summary of the variant caller results for each target. Additional detail of individual variant calls are reported in VCF format in the `<output-file-prefix>.targeted.vcf.gz` output file.

## Input Data

The targeted caller requires WES data or WGS data aligned to a human reference genome. WGS data should be at least 30x coverage as the caller may be less reliable at lower coverage. Human reference genome builds based on `hg19`, `hs37d5` (including `GRCh37`), or `hg38` are supported.

## Configuration files

The targeted caller utilizes several configuration files that are included in the `resources/targeted` directory of the DRAGEN install location. These files include information about the target regions, known variants, and known haplotypes for each target. It is possible to specify additional known variants by modifying these configuration files. Use the following steps to run DRAGEN with a custom set of targeted caller configuration files:

1. Copy the \<dragen\_install\_dir>/resources/targeted directory to a new location
2. Modify the configuration files in the new location as needed
3. Run DRAGEN with the additional command line option to specify the new targeted resources directory: `--targeted-resources-path /path/to/new/resources/targeted`

Note that modification of the targeted caller configuration files can cause unexpected results or errors and should be done with caution.

## Output Files

### Targeted JSON File

The targeted caller generates a `<output-file-prefix>.targeted.json` file in the output directory. The output file is a JSON formatted file containing the fields below.

| Fields in JSON           | Explanation                                                                                                                                          | Type and Possible Values | Present                       |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------ | ----------------------------- |
| sampleId                 | The sample name.                                                                                                                                     | string                   | always                        |
| softwareVersion          | The version of DRAGEN.                                                                                                                               | string                   | always                        |
| genomeBuild              | The reference genome build.                                                                                                                          | string                   | always                        |
| phenotypeDatabaseSources | Resources used for calling metabolism status (phenotype).                                                                                            | json array of strings    | CYP2B6 or CYP2D6 is enabled   |
| cyp2b6                   | The CYP2B6 caller fields.                                                                                                                            | dictionary               | CYP2B6 caller is enabled      |
| cyp2d6                   | The CYP2D6 caller fields.                                                                                                                            | dictionary               | CYP2D6 caller is enabled      |
| cyp21a2                  | The CYP21A2 caller fields.                                                                                                                           | dictionary               | CYP21A2 caller is enabled     |
| gba                      | The GBA caller fields.                                                                                                                               | dictionary               | GBA caller is enabled         |
| hba                      | The HBA caller fields.                                                                                                                               | dictionary               | HBA caller is enabled         |
| lpa                      | The LPA caller fields.                                                                                                                               | dictionary               | LPA caller is enabled         |
| rh                       | The RH caller fields.                                                                                                                                | dictionary               | RH caller is enabled          |
| smn                      | The SMN caller fields.                                                                                                                               | dictionary               | SMN caller is enabled         |
| hla                      | The HLA caller fields, see [HLA Typing](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/hla-typing#hla-output-files).                    | dictionary               | HLA caller is enabled         |
| locusAnnotations         | The Star Allele caller fields, see [Star Allele Caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/star-allele-caller#output-files) | dictionary               | Star Allele caller is enabled |

#### JSON content for recombinant variant detection

For cyp21a2 and gba, the fields below are included in the JSON output. Note that in the descriptions, target gene refers to CYP21A2 or GBA1 and nontarget gene refers to their respective pseudogene paralogs CYP21A1P or GBAP1. The `recombinantHaplotypes` field is limited to reporting only two phased haplotypes for the target gene. To see call sets containing all phased haplotypes for both target and nontarget genes, see the `phasedHaplotypes`->`depthMatchedHaplotypes`->`topHaplotypeSets` field.

| Fields in JSON              | Explanation                                                                                                                                                                                                                                                            | Type and Possible Values                                                                                                                                             |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| totalCopyNumber             | Total copy number of target and nontarget genes including hybrids                                                                                                                                                                                                      | nonnegative integer                                                                                                                                                  |
| deletionBreakpointInGene    | <ul><li>null (i.e. unknown) if totalCopyNumber > 3</li><li>true if CN <= 3 and a deletion-like recombinant variant haplotype is detected</li><li>false if CN <=3 and no deletion-like recombinant variant is detected</li></ul>                                        | true, false, null                                                                                                                                                    |
| recombinantHaplotypes       | List of detected haplotypes arising from nonallelic homologous recombination variant calling. Limited to two haplotypes. For all phased haplotypes in both target and nontarget genes, see the `phasedHaplotypes`->`depthMatchedHaplotypes`->`topHaplotypeSets` field. | Array of two strings. Each string consists of all associated allele IDs (if any) within the haplotype. Consecutive IDs in the same haplotype are separated by a '+'. |
| recombinantHaplotypesFilter | The filter status for the recombinant haplotypes call. Filter will be `RecombinantSiteDepthMismatch` if the reported haplotypes are incompatible with depth-based calls at all phasing sites within the haplotypes.                                                    | string (`PASS` or `RecombinantSiteDepthMismatch`)                                                                                                                    |
| phasedHaplotypes            | Summary of all detected phased haplotypes in target and nontarget gene regions                                                                                                                                                                                         | JSON object                                                                                                                                                          |

Note: A deletion-like recombinant variant haplotype (as opposed to a gene conversion-like recombinant variant haplotype) is defined as a haplotype with one or fewer switch sites (transitions from a target gene allele to a nontarget gene allele) after excluding some sites with common gene conversions in the nontarget gene.

The `phasedHaplotypes` json object will have the fields below.

| Fields in JSON         | Explanation                                                                                                                                                                               | Type and Possible Values                                                                                                                                                                    |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| targetAlleleDepths     | Array consisting of the most likely set of target allele copy numbers at each phasing site in the haplotypes.                                                                             | Array of nonnegative integers                                                                                                                                                               |
| targetAlleleDepthsQual | Phred-scaled quality for the set of depth calls in `targetAlleleDepths`                                                                                                                   | nonnegative decimal number                                                                                                                                                                  |
| rawHaplotypes          | The set of possible haplotypes that were phased from the read data.                                                                                                                       | Array of strings. Each character in a string corresponds to a phasing site in the haplotype (`T` for target gene allele, `N` for nontarget gene allele and `.` for unknown/unphased allele) |
| depthMatchedHaplotypes | Summary of sets of haplotypes that are consistent with depth-based calls at each site within the haplotypes. This field will not be present if no consistent set of haplotypes was found. | JSON object                                                                                                                                                                                 |

The `depthMatchedHaplotypes` json object, when present, will have the fields below.

| Fields in JSON           | Explanation                                                                                                                                                                                                                                                                                            | Type and Possible Values      |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------- |
| targetAlleleDepths       | Array consisting of the set of target allele copy numbers at each phasing site in the haplotypes. The set of target allele copy numbers are consistent with the haplotypes reported in `topHaplotypeSets`.                                                                                             | Array of nonnegative integers |
| targetAlleleDepthsQual   | Phred-scaled quality for the set of matched depth calls in `targetAlleleDepths`                                                                                                                                                                                                                        | nonnegative decimal number    |
| numMatchingHaplotypeSets | Total number of sets of haplotypes that matched the `targetAlleleDepths`                                                                                                                                                                                                                               | positive integer              |
| topHaplotypeSets         | Array of top matching sets of haplotypes. When available, at least 2 matching sets will be reported. If haplotype frequencies are available, the sets are prioritized by population prior. Any sets that are consistent with the haplotypes reported in the `recombinantHaplotypes` are also reported. | Array of JSON objects         |

Each haplotype set reported in the `topHaplotypeSets` array will have the fields below

| Fields in JSON      | Explanation                                                                                                                               | Type and Possible Values   |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | -------------------------- |
| populationPriorQual | Phred-scaled population prior for this set of haplotypes. Only reported if population prior information is available for the target gene. | nonnegative decimal number |
| haplotypes          | Array of JSON objects summarizing each unique haplotype within the set                                                                    | Array of JSON objects      |

Each haplotype reported in the `haplotypes` array will have the fields below

| Fields in JSON       | Explanation                                                                                                                                                                                                 | Type and Possible Values |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------ |
| recombinantAlleleIDs | The relevant identifiers for the variants in the haplotype. Multiple variant identifiers are delimited by `+`. `N/A` when the haplotype cannot be matched to a set of identifiers.                          | string                   |
| haplotype            | The sequence of alleles within the haplotype. Each character corresponds to a phasing site in the haplotype (`T` for target gene allele, `N` for nontarget gene allele and `.` for unknown/unphased allele) | string                   |
| copyNumber           | The number of copies of this haplotype in the set of haplotypes that match the `targetAlleleDepths`                                                                                                         | positive integer         |

#### JSON content for nonrecombinant variant detection

A `variants` field is used to report each nonrecombinant-like variant (i.e. not arising from nonallelic homologous recombination). Each variant will have the fields below.

| Fields in JSON   | Explanation                                      | Type and Possible Values         |
| ---------------- | ------------------------------------------------ | -------------------------------- |
| alleleId         | HGVS identifier of the variant allele            | string                           |
| alleleCopyNumber | Copy number of the allele in the called genotype | nonnegative integer              |
| genotypeQuality  | Phred-scaled quality for the called genotype     | nonnegative integer              |
| filter           | Filter for the called genotype                   | string. "PASS" when not filtered |

Recombinant-like and nonrecombinant-like variants are reported in VCF format. See [Targeted VCF File](#targeted-vcf-file) for details about how these variants are reported in VCF.

### Targeted VCF File

The targeted caller generates a `<output-file-prefix>.targeted.vcf.gz` file in the output directory. The output file is a `VCFv4.2` formatted file. The targets that have VCF output are: cyp2d6, cyp21a2, gba, hba, lpa, rh, and smn.

Small variants, structural variants, and copy number variants are reported in the same VCF file.

The `<output-file-prefix>.targeted.vcf.gz` file includes the following `source` header line:

```
##source=DRAGEN_TARGETED
```

For lpa, rh and smn targets, the `EVENT` and `EVENTTYPE` INFO fields are used to identify the called variants.

The `EVENT` and `EVENTTYPE` INFO fields are formally introduced in `VCFv4.4` to enable the representation of complex rearrangements. This is achieved using the `EVENT` field to group all the related VCF records together, and the `EVENTTYPE` to classify the event. The corresponding header lines are the following.

```
##INFO=<ID=EVENT,Number=A,Type=String,Description="Event name">
##INFO=<ID=EVENTTYPE,Number=A,Type=String,Description="Type of associated event">
```

However, the use of `EVENT` is not limited to complex rearrangements and can be used to associate nonsymbolic alleles, for example in cases of variant position ambiguity in high homology regions.

Since the `EVENTTYPE` values are implementation-defined, custom `EVENTTYPE` header lines are included to describe each `EVENTTYPE`.

```
##EVENTTYPE=<ID=GENE_CONVERSION,Description="Gene conversion event">
##EVENTTYPE=<ID=VARIANT_IN_HOMOLOGY_REGION,Description="Variant in homology region">
##EVENTTYPE=<ID=VNTR,Description="Variable number tandem repeat">
```

For cyp2d6, cyp21a2, gba, and hba targets, the `ALLELE_ID` INFO field is used to identify the called variant alleles.

```
##INFO=<ID=ALLELE_ID,Number=R,Type=String,Description="Identifier for each allele">
```

The missing value `.` is used when no identifier is available (e.g. a wild type allele) or applicable (e.g. allele index 0 for a structural variant record).

Additionally, the 'TargetedCaller' INFO field is used to indicate which targeted caller the current VCF record is generated from

```
##INFO=<ID=TargetedCaller,Number=1,Type=String,Description="Targeted Caller Name">
```

#### Nonrecombinant-like Variants In High Homology Regions

In the case of target variants in a high homology region, each variant is reported ambiguously at all corresponding homologous positions (i.e. in both the pseudogene and in the target gene). Additional analysis for these variants can be performed if absolute certainty that these variants are located in the target gene (e.g. in gba or cyp21a2) is required.

For lpa and smn the ploidy of the called genotype (`FORMAT/GT` field) corresponds to the combined copy number from all the homologous positions. For cyp21a2, gba and hba, this "joint" genotype from all the homologous positions is instead reported in a separate `FORMAT/JGT` field which is then collapsed into a diploid genotype and reported in the `FORMAT/GT` field. The following fields are reported for "joint" calls:

```
##INFO=<ID=JIDS,Number=.,Type=String,Description="IDs (from ID column) of calls associated with a joint genotype call in duplicated regions">
##FORMAT=<ID=JGT,Number=1,Type=String,Description="Joint genotype in duplicated regions">
##FORMAT=<ID=JGQ,Number=1,Type=Integer,Description="Quality of joint genotype in duplicated regions">
##FORMAT=<ID=JPL,Number=.,Type=Integer,Description="Normalized, Phred-scaled likelihoods for joint genotypes as defined in the VCF specification">
##FORMAT=<ID=JQL,Number=1,Type=Float,Description="Phred-scaled likelihood for homozygous reference joint genotype call">
##FORMAT=<ID=JVQL,Number=1,Type=Float,Description="Phred-scaled likelihood for nonvariant joint genotype call where overlapping deletion (*) ALT alleles are not considered to be variant.">
##FORMAT=<ID=JDP,Number=1,Type=Integer,Description="Total depth from all alleles in duplicated regions">
##FORMAT=<ID=JAD,Number=R,Type=Integer,Description="Total read depth for each allele in duplicated regions">
##FORMAT=<ID=JAF,Number=A,Type=Float,Description="Allele frequency for each alt allele in duplicated regions">
```

Note that the `FORMAT/GQ` and `FORMAT/JGQ` fields contain the unconditional genotype quality, unlike the VCF spec where `FORMAT/GQ` is defined as the genotype quality conditioned on the site being variant.

![High Homology Region Variant Example](/files/mIGL0nZn9npHlq1QsUuq)

In the depicted example there are two genes A and B that include a high homology region. The usual process to call variants in this regions is to make a joint pileup of the reads aligning in both genes A and B and call the variants using a model with a ploidy proportional to the total copy number of the regions. This generates divergent possible genotypes that are equally likely since the variant cannot be confidently placed in either gene A or gene B. For lpa and smn the variant would be reported as follows:

```
chr1 100 . A T . TargetedRepeatConflict EVENT=GeneA-B:50A>T;EVENTTYPE=VARIANT_IN_HOMOLOGY_REGION GT 0/0/0/1
chr1 200 . A T . TargetedRepeatConflict EVENT=GeneA-B:50A>T;EVENTTYPE=VARIANT_IN_HOMOLOGY_REGION GT 0/0/0/1
```

Given the unconventional ploidy of the `FORMAT/GT` field used in this representation, a `TargetedRepeatConflict` filter is applied to these records. The header line for the filter is the following.

```
##FILTER=<ID=TargetedRepeatConflict,Description="Set if call is in a targeted repeat region that cannot be placed">
```

For cyp21a2, gba and hba, a conventional diploid `FORMAT/GT` is reported and so no `TargetedRepeatConflict` filter is applied. Due to the ambiguity in placing target variants in high homology regions, the corresponding `QUAL` and `FORMAT/GQ` fields can be much lower than conventional small variant calls (i.e. Phred 3 for a single variant allele copy across two homologous diploid positions). Therefore, instead of filtering on `QUAL` and `FORMAT/GQ` for these records, the records are filtered based on the `FORMAT/JVQL` and `FORMAT/JGQ` fields:

```
##FILTER=<ID=TargetedLowJGQ,Description="Set if call has JGQ < 3.">
##FILTER=<ID=TargetedLowJVQL,Description="Set if call has JVQL < 3.00.">
```

Since the wild type alleles at homologous positions may be different from each other or different from the reference alleles, an additional filter is applied when only wild type alleles are detected across the homologous positions. This avoids making ambiguous variant calls when no target variant of interest is detected.

```
##FILTER=<ID=TargetedWT,Description="Region-ambiguous targeted call with GT containing only wild type alleles, ignoring any overlapping deletions.">
```

#### Rh Gene Conversion Events

In the case of an identified gene conversion even in rh, a small variant is reported at each differentiating site in the acceptor region.

![Gene Conversion Example](/files/ww3Q8vK5HcEUwWiz8vM4)

In the depicted example there are two genes A and B and gene A is the acceptor of a gene conversion from gene B (green box in the figure). Gene conversion are identified by observing variations in copy number at differentiating sites (blue and pink bars in the figure) in consecutive regions. Copy number variations between regions define the breakends of the gene conversion. An equivalent VCF representation for gene conversion would be using CNV and SV entries with breakends corresponding to the donor/acceptor regions, however, only the small variant representation is currently supported.

```
chr1 121 .   A T    . PASS EVENT=GC_AB;EVENTTYPE=GENE_CONVERSION; GT:PS 0|1:121
...
chr1 280 .   G A    . PASS EVENT=GC_AB;EVENTTYPE=GENE_CONVERSION; GT:PS 0|1:121
```

In the case of a detected gene conversion event, there may be differentiating sites with a genotype that is inconsistent with that gene conversion event. In these cases the `RecombinantConflict` filter is applied. The `RecombinantConflict` is defined by the following header line.

```
##FILTER=<ID=RecombinantConflict,Description="Set if call has a copy number that conflicts with a recombinant variant">
```

In the example, the resulting representation is as follows.

```
chr1 121 .   A T    . PASS EVENT=GC_AB;EVENTTYPE=GENE_CONVERSION; GT:PS 0|1:121
...
chr1 144 .   C T    . RecombinantConflict EVENT=GC_AB;EVENTTYPE=GENE_CONVERSION; GT:PS 1|1:121
chr1 153 .   A G    . RecombinantConflict EVENT=GC_AB;EVENTTYPE=GENE_CONVERSION; GT 0/0
...
chr1 280 .   G A    . PASS EVENT=GC_AB;EVENTTYPE=GENE_CONVERSION; GT:PS 0|1:121
```

#### Nonallelic Homologous Recombination

For cyp21a2 and gba, nonallelic homologous recombination can result in gene deletion or duplication in the case of reciprocal recombination or gene conversion in the case of nonreciprocal recombination. Both gene deletion and gene conversion can introduce loss-of-function variants and in both cases the targeted caller will report these variants in the target gene. In the case of gene deletion, the differentiating sites at the nontarget (i.e. pseudogene) positions will contain the overlapping deletion allele `*` while the differentiating sites in the target will contain any variant alleles. Although an equivalent VCF representation would be to simply report the deletion with a single structural variant VCF record, reporting small variant VCF records in the target gene allows for identification of the specific mutations that may occur in a gene transcript and matches well with annotation using HGVS nomenclature. Similarly, for gene conversions, variants are reported at differentiating sites in the target gene, rather than as pairs of structural variant breakends.

Calls at differentiating sites within the recombinant variant calling region will contain the same "joint" fields as are reported for nonrecombinant-like variants in high homology regions (see [Nonrecombinant-like Variants In High Homology Regions](#nonrecombinant-like-variants-in-high-homology-regions)). However, the collapsed diploid `FORMAT/GT` will be based on any detected recombination events. Because detected recombinant variants are placed in the target gene, these records are filtered differently than the ambiguously placed, nonrecombinant-like variants in high homology regions. The `INFO/Recombinant` flag is added to calls derived from recombinant variant calling to distinguish them from nonrecombinant-like variant calls in high homology regions. The `FORMAT/VQL` field is used to apply the `RecombinantLowVQL` filter for low quality recombinant variants and the `RecombinantREF` filter is applied when the collapsed diploid `FORMAT/GT` contains only reference alleles.

```
##FORMAT=<ID=VQL,Number=1,Type=Float,Description="Phred-scaled likelihood for nonvariant genotype call where overlapping deletion (*) ALT alleles are not considered to be variant.">
##FILTER=<ID=RecombinantLowVQL,Description="Region-ambiguous targeted call at recombinant site with VQL below 0.50.">
##FILTER=<ID=RecombinantREF,Description="Region-ambiguous targeted call at recombinant site with GT containing only reference alleles, ignoring any overlapping deletions.">
```

#### Overlapping Structural Variant Representation

The use of `GT=0` for symbolic structural variant alleles is formally disambiguated in `VCFv4.4`, specifying that *"GT=0 indicates the absence of any of the ALT symbolic structural variants defined in the record"*. With this convention we can report compound overlapping heterozygous structural variants.

![Overlapping Variants Representation Example](/files/l10tgmycAQ76vO4d1YeL)

In the hba genotype depicted above, two overlapping SVs can be represented as follows:

```
chr16	170262	.	G	<DEL>,<DUP>	.	.	END=174517;IMPRECISE;SVLEN=4255,4255;SVCLAIM=DJ,DJ;ALLELE_ID=.,-a4.2,aaa4.2	GT	0/2
chr16	173301	.	A	<DEL>,<DUP>	.	.	END=177104;IMPRECISE;SVLEN=3804,3804;SVCLAIM=DJ,DJ;ALLELE_ID=.,-a3.7,aaa3.7	GT	0/1
```

The relevant header lines for the VCF records above are:

```
##INFO=<ID=END,Number=1,Type=Integer,Description="End position of the variant described in this record">
##INFO=<ID=SVLEN,Number=A,Type=Integer,Description="Length of structural variant">
##INFO=<ID=SVCLAIM,Number=A,Type=String,Description="Claim made by the structural variant call. Valid values are D, J, DJ for abundance, adjacency and both respectively.">
##INFO=<ID=IMPRECISE,Number=0,Type=Flag,Description="Imprecise structural variation">
```

#### Variable Number Tandem Repeat Representation

![VNTR Example](/files/PHonDJpLq3bfaeDPylJM)

In the depicted example there is a Variable Number Tandem Repeat (VNTR) region composed of three repeat units in the reference. The `CN` INFO field is used to report the allele copy number, the `CN` FORMAT field to is used report the region total copy number given by the sum of the allele copy numbers, and the `REPCN` FORMAT field is used to report the repeat unit copy number equal to the allele copy number multiplied by the number of repeat units in the reference.

This VNTR can be represented as follows:

```
chr1 100 . A <DUP>,<DUP> . . END=400;EVENT=A;EVENTTYPE=VNTR;SVCLAIM=D;SVLEN=300;CN=2.6,4.3   GT:CN:REPCN 1|2:6.9:8|13
```

The `REPCN` and `CN` header lines are:

```
##FORMAT=<ID=REPCN,Number=1,Type=String,Description="Number of repeat units spanned by the allele">
##INFO=<ID=CN,Number=A,Type=Float,Description="Copy number of CNV / breakpoint">
##FORMAT=<ID=CN,Number=1,Type=Float,Description="Estimated copy number">
```

#### Additional Filters

For lpa, rh and smn, the `TargetedLowQual` filter is applied if the `QUAL` of a target variant is less than `3.00`.

```
##FILTER=<ID=TargetedLowQual,Description="Set if call has QUAL < 3.00">
```

Similarly, for cyp21a2 and gba the `TargetedLowVQL` filter is applied if the `VQL` of a target variant in low-homology region is less than `3.00`.

```
##FORMAT=<ID=VQL,Number=1,Type=Float,Description="Phred-scaled likelihood for nonvariant genotype call where overlapping deletion (*) ALT alleles are not considered to be variant.">
##FILTER=<ID=TargetedLowVQL,Description="Set if call has VQL < 3.00.">
```

The `TargetedLowGQ` filter is applied if the targeted variant has `GQ` smaller than `3`.

```
##FILTER=<ID=TargetedLowGQ,Description="Set if call has GQ < 3 and JGQ is not present.">
```

### Merging Targeted Calls In The `hard-filtered` Files

When the small variant caller is enabled, the targeted small variant VCF calls can be merged into the `<output-file-prefix>.hard-filtered.vcf.gz` and `<output-file-prefix>.hard-filtered.gvcf.gz` files, briefly `hard-filtered` files. The `--targeted-merge-vc` command line option can be used to control which targets will have their small variant VCF records merged into the `hard-filtered` files. For example, `--targeted-merge-vc rh` will enable merging of the calls from the `rh` caller into the `hard-filtered` files and `--targeted-merge-vc rh hba` will enable merging of the calls from the `rh` and `hba` targets into the `hard-filtered` files. The `true` value will merge all calls from all supported targets into the `hard-filtered` files, while the `false` value will merge no calls into the `hard-filtered` files.

The targeted calls merged into the `hard-filtered` files are marked with a `TARGETED` INFO flag.

When enabled, targeted small variants are merged into the `hard-filtered` files regardless of any regions that may be provided using the `--vc-target-bed` option.

#### Merging Strategy

The merging strategy for targeted small variant calls is to prioritize the targeted calls over small variant calls from the germline small variant caller. When a germline small variant call overlaps a targeted caller call, then the small variant call is filtered with a `TargetedConflict` filter if any of the following holds:

* The targeted caller call is `PASS`.
* The small variant call and targeted caller call have incompatible genotypes and the targeted caller call is not filtered with the `TargetedLowGQ` filter.

The strategy is summarized in the following examples.

1. The `TARGETED` call is `PASS`.

```
chr1 100 . A	C	. TargetedConflict 	.			GT 0/1
chr1 100 . A	C	. PASS 				TARGETED 	GT 1/1
```

2. The `TARGETED` call and the small variant call are not overlapping

```
chr1 110 . T	TCA	. PASS 				. 			GT 0/1
chr1 111 . G	A	. PASS 				TARGETED 	GT 0/1
```

3. The `TARGETED` call is filtered with `VARIANT_IN_HOMOLOGY_REGION` and has a discordant variant representation with the overlapping small variant call.

```
chr1 120 . ATTC A	. TargetedConflict	.			GT 0/1
chr1 121 . T	A	. TargetedLowQual	TARGETED 	GT 0/1
chr1 125 . TCAC T	. TargetedLowQual	TARGETED	GT 0/1
chr1 126 . C	G	. TargetedConflict	.			GT 0/1
```

4. The `TARGETED` call is filtered with `TargetedLowQual` and has a discordant genotype with the overlapping small variant call.

```
chr1 130 . C	G	. TargetedConflict	.			GT 0/1
chr1 130 . C	G	. TargetedLowQual	TARGETED 	GT 1/1
```

5. The `TARGETED` call is filtered with `TargetedLowGQ` and has a discordant genotype with the overlapping small variant call.

```
chr1 140 . AC	A	. PASS			.			GT:GQ 0/1:5
chr1 140 . A	T	. TargetedLowGQ	TARGETED 	GT:GQ 1/1:2
```

## Exome calling using in-run PON

Targeted calling from WES data is supported for hba and smn. It uses an in-run panel of normals (PON) for coverage normalization of the various target regions by automatically identifying copy-neutral samples from a single sequencing run. All samples in the panel are expected to be from the same sequencing run and library prep batch as the case samples being analyzed. Samples must be prepared using the Illumina CS/PGx Custom Enrichment Research Panel. If targeted calling is enabled on WES data without a PON then targeted calling is skipped and no targeted calling output files will be generated. The first step in targeted calling from WES data is to generate exome counts files for each of the samples in the PON. A minimum of 30 samples is required in the PON and the PON must be sufficiently diverse such that for a given target region, a large subset of samples is copy-neutral. For example, a PON where all samples are positive for alpha thalassemia (HBA1/2 deletion) would not be sufficiently diverse for accurately calling variants in HBA1/2. Similarly, a PON consisting of a large pedigree of related samples would not be sufficiently diverse. No more than \~6% of the samples in the PON should be related to any case sample being analyzed; a PON of 50 samples containing a quad would be acceptable since it would contain 3 samples related to a proband (Mother/Father/Sibling) or \~6% of the samples in the PON. If the samples in the sequencing run are sufficiently diverse, then it is recommended that the PON consist of as many samples from the sequencing run as possible, but can be limited to 96 samples without significantly impacting the accuracy of coverage normalization.

The table below summarizes the available options and high-level steps for running the Targeted Caller using an in-run PON. CNV and Targeted Caller require separate PON files, but the intermediate counts files can be generated in the same DRAGEN command line invocation. For additional details click on the link for each option.

| Analysis option                                                                                                                                                                                             | Steps                                                                                                                                                                                                                                                          |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| [BSSH planned sequencing run (from BCLs)](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-germline-enrichment-from-bcls-bssh-app#dragen-germline-enrichment-from-bcls-through-run-planning-tool) | <ol><li>Create run using the Run Planning tool in BSSH</li><li>Start planned run in Control Software on instrument</li></ol>                                                                                                                                   |
| [BSSH existing sequencing run (from BCLs)](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-germline-enrichment-from-bcls-bssh-app#dragen-germline-enrichment-from-bcls-from-existing-run)        | <ol><li>Run DRAGEN Germline Enrichment from BCLs App</li></ol>                                                                                                                                                                                                 |
| [ICA from FASTQs/BAMs/CRAMs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-germline-enrichment-ica-app)                                                                                        | <ol><li>Run DRAGEN Germline Enrichment App</li></ol>                                                                                                                                                                                                           |
| [BSSH from FASTQs/BAMs/CRAMs](/dragen-v4.5/product-guides/dragen-v4.5/dragen-apps/dragen-enrichment-bssh-app)                                                                                               | <ol><li>Run DRAGEN Enrichment App</li></ol>                                                                                                                                                                                                                    |
| [Local or AMI](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-wes)                                                                                                                     | <ol><li>BCL to FASTQ conversion</li><li>Generate CNV target counts and Targeted Caller exome counts for each PON sample</li><li>Generate CNV combined counts PON file</li><li>Generate Targeted Caller PON file</li><li>Perform case sample analyses</li></ol> |

### Exome counts generation

Exome counts file generation can be enabled using the command line option `--targeted-generate-exome-counts=true`. A `<output-file-prefix>.targeted.exome.counts.json.gz` file will be generated in the output directory. Note that the `--enable-targeted` option is not required, but can be used to specify a subset of targets.

### Exome PON generation

An exome PON file can be generated, using the command line option `--targeted-pon-counts-list` to pass a text file containing a list of exome counts files, one for each sample in the panel. A `<output-file-prefix>.targeted.pon.json.gz` file will be generated in the output directory. Note that this is a stand-alone independent dragen run that cannot be combined with other dragen components. A read input (bam/cram/fastq) file is not used.

### Exome case sample analysis

Exome targeted calling on a case sample is performed by passing in a PON file and a systematic noise file using the command line options `--targeted-pon` and `--targeted-systematic-noise`, respectively. Note that the PON file should be for the same batch as the case sample. A systematic noise file and corresponding pre-built pangenome reference can be downloaded from the [DRAGEN Software Support Site page](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/product_files.html). A json file, `<output-file-prefix>.targeted.json` and a vcf file, `<output-file-prefix>.targeted.vcf.gz` will be generated in the output directory with the calls for the enabled targets. For WES mode, an additional field `ponQualityFilter`, is added to the JSON output for each enabled target. It denotes the quality of the PON and the confidence of the resulting calls. If the case sample does not correlate well with the PON, the `ponQualityFilter` gets set to `LowPonCorrelation`, signaling that the calls are considered to have low confidence. Note that the `--enable-targeted` option is not required, but can be used to specify a subset of targets.

## Command-Line Examples

The Targeted Caller can be enabled in parallel with other components as part of a human WGS germline analysis workflow (see [DRAGEN Recipe - Germline WGS](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-wgs)).

The Targeted Caller can be enabled in parallel with other components as part of a human WES germline analysis workflow (see [DRAGEN Recipe - Germline WES](/dragen-v4.5/product-guides/dragen-v4.5/dragen-recipes/dna-germline-wes)).

### FASTQ Input Example

The following command-line example runs the targeted caller from FASTQ input:

```
dragen \
	-r /staging/human/reference/hg38_alt_aware/DRAGEN/${HASH_TABLE_VERSION} \
	--fastq-file1 /staging/test/data/NA12878_R1.fastq \
	--fastq-file2 /staging/test/data/NA12878_R2.fastq \
	--output-directory /staging/test/output \
	--output-file-prefix NA12878_dragen \
	--RGID DRAGEN_RGID \
	--RGSM NA12878 \
	--enable-targeted=true
```

### Prealigned BAM Input Example

The following command-line example runs cyp21a2 only using BAM input without realignment:

```
dragen \
	-r /staging/human/reference/hg38_alt_aware/DRAGEN/${HASH_TABLE_VERSION} \
	--bam-input /staging/test/data/NA12878.bam \
	--output-directory /staging/test/output \
	--output-file-prefix NA12878_dragen \
	--enable-map-align=false \
	--enable-targeted=cyp21a2
```

### Exome counts generation from prealigned BAM Input Example

```
dragen \
	-r /staging/human/reference/hg38_alt_aware/DRAGEN/${HASH_TABLE_VERSION} \
	--bam-input /staging/test/data/NA12878.bam \
	--output-directory /staging/test/output \
	--output-file-prefix NA12878_dragen \
	--enable-map-align=false \
	--targeted-generate-exome-counts=true
```

### Exome PON generation from exome counts files from a single sequencing run

```
dragen \
--output-directory /staging/test/output \
--output-file-prefix run1 \
--targeted-pon-counts-list run1_exome_counts_list.txt
```

### Exome case sample analysis from prealigned BAM Input Example

```
dragen \
-r /staging/human/reference/hg38_alt_aware/DRAGEN/${HASH_TABLE_VERSION} \
--bam-input /staging/test/data/NA12878.bam \
--output-directory /staging/test/output \
--output-file-prefix NA12878_dragen \
--targeted-pon run1.targeted.pon.json.gz \
--targeted-systematic-noise dragen4.4.targeted.systematic_noise.json.gz \
```


# CYP2B6 Caller

The CYP2B6 Caller is capable of genotyping the *CYP2B6* gene from whole-genome sequencing (WGS) data. Due to high sequence similarity with its pseudogene paralog *CYP2B7* and a wide variety of common structural variants (SVs), a specialized caller is necessary to resolve variants and identify likely star allele haplotypes.

The CYP2B6 Caller performs the following steps:

1. Determines total *CYP2B6* and *CYP2B7* copy number from read depth.
2. Determines *CYP2B6*-derived copy number at *CYP2B6*/*CYP2B7* differentiating sites.
3. Detects SV breakpoints by calculating the changes in *CYP2B6*-derived copy number along the *CYP2B6*gene.
4. Calls small variants in *CYP2B6* copies.
5. Identifies star alleles from the detected SV breakpoints and small variants.
6. Identifies the most likely genotype for the called star alleles.

## Total *CYP2B6* and *CYP2B7* Copy Number

The first step of CYP2B6 calling is to determine the combined copy number of *CYP2B6* and *CYP2B7*. Reads aligned to regions in either *CYP2B6* or *CYP2B7* are counted. The counts in each region are corrected for GC-bias, and then normalized to a diploid baseline. The GC-bias correction and normalization factors are determined from read counts in 3000 preselected 2 kb regions across the genome. These 3000 normalization regions were randomly selected from the portion of the reference genome having stable coverage across population samples. The combined *CYP2B6* and *CYP2B7* copy number is then calculated from the average sequencing depth across the *CYP2B6* and *CYP2B7* regions.

## Differentiating Sites

The *CYP2B6*-derived copy number is calculated at 99 predefined differentiating sites across the *CYP2B6* gene. The differentiating sites are selected at positions with sequence differences in *CYP2B6* and *CYP2B7* where calling the *CYP2B6*-derived copy number shows an accuracy of greater than 98% based on sequencing data from the 1000 Genomes Project.

For each differentiating site, *CYP2B6*-specific and *CYP2B7*-specific alleles are counted in reads mapping to either *CYP2B6* or the homologous region in *CYP2B7*. The *CYP2B6*-derived copy number is then calculated from the two gene-specific allele counts using the total *CYP2B6* and *CYP2B7* copy number calculated from the previous step.

## Structural Variant Calling

The *CYP2B6*-derived copy number along the *CYP2B6* gene is used to identify known population structural variants (SVs), including whole gene deletions and duplications as well as certain gene conversions and gene fusions. The following fusion variants are detected:

| Fusion Breakpoint | Hybrid Gene Structure | Star-Allele Designation |
| ----------------- | --------------------- | ----------------------- |
| intron 4-exon 5   | 2B7-2B6               | `*29`                   |
| intron 4-exon 5   | 2B6-2B7               | `*30`                   |

## Small Variant Calling

35 small variants that define various star alleles are detected from the read alignments. All of these variants are in unique (nonhomologous) regions of *CYP2B6* with high mapping quality. Only reads mapping to *CYP2B6* are used for calling variants in nonhomologous regions.

For each variant, reads containing either the variant allele or the nonvariant alleles are counted. A binomial model that incorporates the sequencing errors is then used to determine the most likely variant copy number (0 for nonvariant).

Samples with poor sequencing quality or greater than five copies of *CYP2B6* will have allele counts with higher variance. This elevated variance increases the chance that the most likely variant copy number is wrong. To handle these cases, the small variant caller also indicates alternate, less likely variant copy numbers.

## Recombinant Variant Calling

The recombinant (gene conversion) variant 18053A>G is detected by phasing the variant site with five flanking differentiating sites. When the haplotypes formed from phasing these sites supports the gene conversion in *CYP2B6*, a read depth analysis at the gene conversion breakpoints (transitions from either *CYP2B6*->*CYP2B7* or *CYP2B7*->*CYP2B6*) is performed. When the posterior probability that there is at least one gene conversion variant is above 0.7 then DRAGEN uses the variant for star allele identification.

## Star Allele Identification

The called SVs and small variant genotypes are matched against the definitions of 39 different star alleles (PharmVar version 4.17, July 3, 2020) including an unreported star allele identified as `*U1` by the caller. `*U1` is defined by the following two variant alleles: 64C>T (NC\_000019.10:g.40991369C>T) and 25505C>T (NC\_000019.10:g.41016810C>T). Multiple star allele genotypes may be reported when different sets of star alleles match the called variant genotypes, as with `*1`, `*6` and `*4`, `*49` where both sets of star alleles contain the same two small variants. Additionally, the small variant caller can emit alternate genotypes with similar likelihoods to the most likely variant genotypes. This can also result in different sets of star alleles being reported. When matching the variant genotypes to the star alleles, the number of identified star alleles must equal the number of *CYP2B6*-derived gene copies determined from previous steps. If no variant genotypes can be matched to a set of star alleles, the CYP2B6 Caller returns a no call during the genotyping step with filter value `No_call`.

## Genotyping

Given a possible set of star alleles, the genotyping step attempts to identify the two likely haplotypes that contain all star alleles in the set. The likelihood of any given genotype is determined from a table of population frequencies determined from the 1000 Genomes Project and the genotype with the highest population frequency is selected. When two or more possible genotypes are identified with similar population frequencies, then all genotypes are emitted. This results in a call with filter value `More_than_one_possible_genotype`.

## CYP2B6 Output File

The caller prints out its calls in the targeted caller output file, `<output-file-prefix>.targeted.json` that also contains calls from other targets (see [Targeted JSON File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-json-file)).

An example of the CYP2B6 caller content in the output is as follows:

```
{
    "cyp2b6": {
        "genotype": "*17/*2",
        "genotypeFilter": "PASS",
        "phenotypeDatabaseAnnotation": "Normal Metabolizer"
    }
}
```

For CYP2B6 caller, the fields are defined as follows.

| Fields in JSON              | Explanation                                                                               | Type and Possible Values                                                               |
| --------------------------- | ----------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- |
| genotype                    | star allele genotype identified for sample                                                | string                                                                                 |
| genotypeFilter              | The filter status for the genotype call                                                   | string (The value can include: PASS, No\_call, or More\_than\_one\_possible\_genotype) |
| phenotypeDatabaseAnnotation | The metabolism status corresponding to the genotype, mapped from phenotypeDatabaseSources | string or null if not present in the database                                          |


# CYP2D6 Caller

The CYP2D6 Caller is capable of genotyping the *CYP2D6* gene from whole-genome sequencing (WGS) data and is derived from the method implemented in Cyrius¹. Due to high sequence similarity with its pseudogene paralog *CYP2D7* and a wide variety of common structural variants (SVs), a specialized caller is necessary to resolve variants and identify likely star allele haplotypes.

The CYP2D6 Caller performs the following steps:

1. Determines total *CYP2D6* and *CYP2D7* copy number from read depth.
2. Determines *CYP2D6*-derived copy number at *CYP2D6*/*CYP2D7* differentiating sites.
3. Detects SV breakpoints by calculating the changes in *CYP2D6*-derived copy number along the *CYP2D6* gene.
4. Calls small variants in *CYP2D6* copies.
5. Identifies star alleles from the detected SV breakpoints and small variants.
6. Identifies the most likely genotype for the called star alleles.

## Total *CYP2D6* and *CYP2D7* Copy Number

The first step of CYP2D6 calling is to determine the combined copy number of *CYP2D6* and *CYP2D7*. Reads aligned to regions in either *CYP2D6* or *CYP2D7* are counted. The counts in each region are corrected for GC-bias, and then normalized to a diploid baseline. The GC-bias correction and normalization factors are determined from read counts in 3000 preselected 2 kb regions across the genome. These 3000 normalization regions were randomly selected from the portion of the reference genome having stable coverage across population samples. The combined *CYP2D6* and *CYP2D7* copy number is then calculated from the average sequencing depth across the *CYP2D6* and *CYP2D7* regions.

## Differentiating Sites

The *CYP2D6*-derived copy number is calculated at 117 predefined differentiating sites across the *CYP2D6* gene. The differentiating sites are selected at positions with sequence differences in *CYP2D6* and *CYP2D7* where calling the *CYP2D6*-derived copy number shows an accuracy of greater than 98% based on sequencing data from the 1000 Genomes Project.

For each differentiating site, *CYP2D6*-specific and *CYP2D7*-specific alleles are counted in reads mapping to either *CYP2D6* or the homologous region in *CYP2D7*. The *CYP2D6*-derived copy number is then calculated from the two gene-specific allele counts using the total *CYP2D6* and *CYP2D7* copy number calculated from the previous step.

## Structural Variant Calling

The *CYP2D6*-derived copy number along the *CYP2D6* gene is used to identify known population structural variants (SVs), including whole gene deletions and duplications as well as certain gene conversions and gene fusions. The following fusion variants are detected:

| Fusion Breakpoint | Hybrid Gene Structure | Star-Allele Designation                                                                       |
| ----------------- | --------------------- | --------------------------------------------------------------------------------------------- |
| exon 9            | 2D6-2D7               | <p><code>\*4.013</code>,<br><code>\*36</code>,<br><code>\*57</code>,<br><code>\*83</code></p> |
| exon 9            | 2D7-2D6               | `*13`                                                                                         |
| intron 4          | 2D7-2D6               | `*13`                                                                                         |
| intron 1          | 2D7-2D6               | `*13`                                                                                         |
| intron 1          | 2D6-2D7               | `*68`                                                                                         |

In addition to the exon 9 fusion breakpoints, exon 9 can participate in *CYP2D7* gene conversion resulting in an embedded *CYP2D7* sequence instead of a true hybrid. The structural variant caller also detects exon 9 gene conversions. Because only changes in *CYP2D6*-derived copy number yield structural variant calls, there might be rare cases where two hybrid copies result in no structural variant calls. For example, when both `*36` and `*13` with fusion breakpoint in exon 9 are present. However, the structural variant caller is capable of detecting multiple copies of the same fusion type (eg, `*36x2`) or cases where both an exon 9 gene conversion copy and an exon 9 2D6-2D7 hybrid are present.

## Small Variant Calling

118 small variants that define various star alleles are detected from the read alignments. 96 of these variants are in unique (nonhomologous) regions of *CYP2D6* with high mapping quality. Only reads mapping to *CYP2D6* are used for calling variants in nonhomologous regions. The other 22 variants occur in homologous regions of *CYP2D6* where reads mapping to either *CYP2D6* or *CYP2D7* are used for variant calling.

For each variant, reads containing either the variant allele or the nonvariant alleles are counted. A binomial model that incorporates the sequencing errors is then used to determine the most likely variant copy number (0 for nonvariant). A strand bias filter is applied to a small subset of variants that would otherwise tend to have false positive calls in the population.

Samples with poor sequencing quality or greater than five copies of *CYP2D6* will have allele counts with higher variance. This elevated variance increases the chance that the most likely variant copy number is wrong. To handle these cases, the small variant caller also indicates alternate, less likely variant copy numbers.

## Star Allele Identification

The called SVs and small variant genotypes are matched against the definitions of 128 different star alleles (PharmVar version 4.17, July 3, 2020). This might result in different sets of star alleles matching the called variant genotypes, such as with `*1`, `*46` and `*43`, `*45` where both sets of star alleles contain the same 4 small variants. When the small variant caller emits alternate, less likely variant copy numbers in addition to the most likely variant copy numbers, this might result in different sets of star alleles being identified, since these alternate sets of variant copy numbers are also matched to the star allele definitions. The number of matched star alleles must match the number of *CYP2D6*-derived gene copies determined from previous steps. When there are fewer than two *CYP2D6*-derived gene copies, then one or more `*5` deletion haplotypes are included in the output set of star alleles. If all variant genotypes cannot be matched to a set of star alleles, the CYP2D6 Caller returns a no call during the genotyping step with filter value `No_call`.

## Genotyping

Given a possible set of star alleles, the genotyping step attempts to identify the two likely haplotypes that contain all star alleles in the set. The deletion haplotype (`*5`) is considered as a possible haplotype during this process. The likelihood of any given genotype is determined from a table of population frequencies determined from the 1000 Genomes Project and the genotype with the highest population frequency is selected. When two or more possible genotypes are identified with similar population frequencies, then all genotypes are emitted. This results in a call with filter value `More_than_one_possible_genotype`.

## CYP2D6 Output File

The CYP2D6 Caller prints out its calls in the targeted caller output file, `<output-file-prefix>.targeted.json` that also contains calls from other targets (see [Targeted JSON File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-json-file)). An example of the CYP2D6 caller content in the output is as follows:

```
{
    cyp2d6": {
        "totalCopyNumber": 4,
        "genotype": "*113/*2",
        "genotypeFilter": "PASS",
        "genotypeQuality": 150,
        "phenotypeDatabaseAnnotation": "Indeterminate",
        "variants": [
        {
            "alleleId": "g.42126752C>T",
            "alleleCopyNumber": 1,
            "genotypeQuality": 28,
            "filter": "PASS"
        },
        {
            "alleleId": "g.42126938C>T",
            "alleleCopyNumber": 2,
            "genotypeQuality": 58,
            "filter": "PASS"
        },
        {
            "alleleId": "g.42126611C>G",
            "alleleCopyNumber": 1,
            "genotypeQuality": 150,
            "filter": "PASS"
        },
        {
            "alleleId": "g.42127941G>A",
            "alleleCopyNumber": 1,
            "genotypeQuality": 150,
            "filter": "PASS"
        }
        ]
    }
}
```

| Fields in JSON              | Explanation                                                                               | Type and Possible Values                                                                  |
| --------------------------- | ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- |
| totalCopyNumber             | Total combined copy number of *CYP2D6* and *CYP2D7*                                       | nonnegative integer                                                                       |
| genotype                    | called star allele genotype                                                               | string (semi-colon delimited list of possible genotypes with haplotypes separated by `/`) |
| genotypeFilter              | The filter status for the genotype call                                                   | string (The value can include: PASS, No\_call, or More\_than\_one\_possible\_genotype)    |
| genotypeQuality             | The quality value for the genotype call                                                   | integer (The value can be in the range of 0-150)                                          |
| phenotypeDatabaseAnnotation | The metabolism status corresponding to the genotype, mapped from phenotypeDatabaseSources | string or null if not present in the database                                             |
| variants                    | a json array containing info about specific *CYP2D6* variants                             | json-array                                                                                |

Each variant reported in the `variants` array will have the fields below.

| Fields in JSON   | Explanation                                      | Type and Possible Values         |
| ---------------- | ------------------------------------------------ | -------------------------------- |
| alleleId         | HGVS identifier of the variant allele            | string                           |
| alleleCopyNumber | Copy number of the allele in the called genotype | nonnegative integer              |
| genotypeQuality  | Phred-scaled quality for the called genotype     | nonnegative integer              |
| filter           | Filter for the called genotype                   | string. "PASS" when not filtered |

Each *CYP2D6* genotype contains two haplotypes separated by a slash (eg `*1/*2`). Each haplotype consists of one or more star alleles separated by a plus sign (eg `*10+*36`). When a haplotype contains more than one copy of the same star allele, that star allele only appears once and is followed by a multiplication sign, and then the number of copies (eg `*1x2` for two copies of `*1`).

CYP2D6 Caller can also print out its variant calls into VCF format which has been described in the top-level targeted-caller section (see [Targeted VCF File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-vcf-file)).

¹Chen X, Shen F, Gonzaludo N, et al. Cyrius: accurate CYP2D6 genotyping using whole-genome sequencing data. The Pharmacogenomics Journal. 2021;21(2):251-261. doi:10.1038/s41397-020-00205-5


# CYP21A2 Caller

The CYP21A2 Caller is capable of genotyping the *CYP21A2* gene from whole-genome sequencing (WGS) data. Due to high sequence similarity with its pseudogene paralog *CYP21A1P* and a wide variety of common structural variants (SVs), a specialized caller is necessary to resolve variants.

The CYP21A2 calling workflow is broken up into the following major stages:

1. Loading input configuration
2. Processing read data
3. Analyzing read data

Read data analysis is further split into the following steps:

1. Determine total *CYP21A2* and *CYP21A1P* copy number from read depth.
2. Call small variants in *CYP21A2* copies.
3. Phase reads to detect common variants and recombination events.
4. Identify most likely haplotypes.

The CYP21A2 Caller requires WGS data aligned to a human reference genome with at least 30x coverage.

## Total *CYP21A2* and *CYP21A1P* Copy Number

The first step of CYP21A2 calling is to determine the combined copy number of *CYP21A2* and *CYP21A1P*. Reads aligned to regions in either *CYP21A2* or *CYP21A1P* are counted. The counts in each region are corrected for GC-bias, and then normalized to a diploid baseline. The GC-bias correction and normalization factors are determined from read counts in 3000 preselected 2kb regions across the genome. These 3000 normalization regions were randomly selected from the portion of the reference genome having stable coverage across population samples. The combined *CYP21A2* and *CYP21A1P* copy number is then calculated from the average sequencing depth across the *CYP21A2* and *CYP21A1P* regions.

## Nonrecombinant-like Variant Calling

Of the known nonrecombinant-like variants, some are in unique (nonhomologous) regions of *CYP21A2* with high mapping quality. Only reads mapping to *CYP21A2* are used for calling variants in nonhomologous regions. The other variants occur in homologous regions of *CYP21A2*/*CYP21A1P* where reads mapping to either are used for variant calling.

For each variant, reads containing either the variant allele or the nonvariant allele is counted. A binomial model that incorporates the sequencing error rate is then used to determine the most likely variant copy number (0 for nonvariant).

For a list of the supported nonrecombinant-like variants, refer to the `targeted/cyp21a2/target_variants_*.tsv` files located in the `resources` directory of the DRAGEN install location.

## Nonallelic Homologous Recombination Variant Calling

To analyze the homologous region even further, DRAGEN phases reads covering differentiating sites and known variant sites. Whenever a detected haplotype has a *CYP21A2*->*CYP21A1P* or *CYP21A1P*-> *CYP21A2* transition that is consistent with one of the known recombinant-like variants, the transition is considered as a candidate breakpoint for calling those variants. Reads containing phasing information for the two sites flanking each candidate breakpoint are used for variant calling. When the read data supports the hypothesis that the sample contains at least one copy of a candidate breakpoint, the associated haplotype is a recombinant haplotype candidate. Recombinant haplotype candidates are sorted by likelihood and the number of variant sites. If no wild type haplotype was detected, DRAGEN reports any detected homozygous recombinant haplotype, or up to two different recombinant haplotypes (i.e. compound het) if detected. If any wild type haplotype was found, DRAGEN reports a maximum of one recombinant haplotype. When no recombinant haplotypes are detected two wild type haplotypes are reported.

For a list of recombinant variant sites, refer to the `targeted/cyp21a2/recombinant_variants_*.tsv` files located in the `resources` directory of the DRAGEN install location.

Note that NM\_000500.9:c.710\_719delinsACGAGGAGAA will be reported as the following three variants on the same haplotype: NM\_000500.9:c.710T>A NM\_000500.9:c.713T>A NM\_000500.9:c.719T>A

## CYP21A2 Output File

The CYP21A2 Caller generates its output in the targeted caller output file `<output-file-prefix>.targeted.json` that also contains calls from other targets (see [Targeted JSON File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-json-file)).

### Output File Example

An example of the CYP21A2 caller content in the `<output-file-prefix>.targeted.json` output file is shown below.

```json
{        
    "cyp21a2": {
                "totalCopyNumber": 4,
                "deletionBreakpointInGene": null,
                "recombinantHaplotypes": [
                  "NM_000500.9:c.955C>T",
                  ""
                ],
                "recombinantHaplotypesFilter": "RecombinantSiteDepthMismatch",
                "phasedHaplotypes": {
                  "targetAlleleDepths": [ 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 2, 2, 2, 1, 2],
                  "targetAlleleDepthsQual": 2.6241045542695463,
                  "rawHaplotypes": [
                    "................TT",
                    "................NT",
                    "................NN",
                    ".........TTTTT....",
                    ".........NNNN.....",
                    ".TTTTTTT..........",
                    ".NNNNNN...........",
                    "T.................",
                    "N................."
                  ],
                  "depthMatchedHaplotypes": {
                    "targetAlleleDepths": [1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 1, 2],
                    "targetAlleleDepthsQual": 0.1321771093877906,
                    "numMatchingHaplotypeSets": 52,
                    "topHaplotypeSets": [
                      {
                        "populationPriorQual": 12.124787664997037,
                        "haplotypes": [
                          {
                            "recombinantAlleleIDs": "",
                            "haplotype": "TTTTTTTT.TTTTT..TT",
                            "copyNumber": 1
                          },
                          {
                            "recombinantAlleleIDs": "NM_000500.9:c.92C>T+NM_000500.9:c.293-13C>G+NM_000500.9:c.332_339del+NM_000500.9:c.955C>T",
                            "haplotype": "NNNNNNN..TTTTT..NT",
                            "copyNumber": 1
                          },
                          {
                            "recombinantAlleleIDs": "NM_000500.9:c.92C>T+NM_000500.9:c.293-13C>G+NM_000500.9:c.332_339del+NM_000500.9:c.710T>A+NM_000500.9:c.713T>A+NM_000500.9:c.719T>A+NM_000500.9:c.955C>T+NM_000500.9:c.1069C>T",
                            "haplotype": "NNNNNNN..NNNN...NN",
                            "copyNumber": 2
                          }
                        ]
                      },
                      {
                        "populationPriorQual": 0.14664469798819596,
                        "haplotypes": [
                          {
                            "recombinantAlleleIDs": "NM_000500.9:c.955C>T",
                            "haplotype": "TTTTTTTT.TTTTT..NT",
                            "copyNumber": 1
                          },
                          {
                            "recombinantAlleleIDs": "NM_000500.9:c.92C>T+NM_000500.9:c.293-13C>G+NM_000500.9:c.332_339del",
                            "haplotype": "NNNNNNN..TTTTT..TT",
                            "copyNumber": 1
                          },
                          {
                            "recombinantAlleleIDs": "NM_000500.9:c.92C>T+NM_000500.9:c.293-13C>G+NM_000500.9:c.332_339del+NM_000500.9:c.710T>A+NM_000500.9:c.713T>A+NM_000500.9:c.719T>A+NM_000500.9:c.955C>T+NM_000500.9:c.1069C>T",
                            "haplotype": "NNNNNNN..NNNN...NN",
                            "copyNumber": 2
                          }
                        ]
                      }
                    ]
                  }
                },
            "variants": [
                {
                    "alleleId": "NM_000500.9:c.1360C>T",
                    "alleleCopyNumber": 2,
                    "genotypeQuality": 18,
                    "filter": "PASS"
                }
            ]
    }
}
```


# GBA Caller

The GBA Caller is capable of detecting both recombinant-like and nonrecombinant-like variants in the *GBA* gene from whole-genome sequencing (WGS) data. Disruption of all copies of the *GBA* gene in an individual causes the autosomal recessive disorder Gaucher disease, and carriers are at increased risk of Parkinson's disease and Lewy body dementia. Due to high sequence similarity with its pseudogene paralog *GBAP1*, calling recombinant-like variants in *GBA* requires a specialized caller.

To enable the GBA Caller, use `--enable-gba=true` as part of a germline-only WGS analysis workflow. The GBA Caller is disabled by default and requires WGS data aligned to a human reference genome with at least 30x coverage.

The GBA Caller performs the following steps:

1. Determine the total combined *GBA* and *GBAP1* copy number
2. Detect nonrecombinant-like variants from a set of 111 known variants
3. Assemble phased haplotypes in the exon 9-11 region where recombinant variants occur
4. Detect any *GBAP1* -> *GBA* breakpoints that are consistent with one of the 7 known recombinant-like variants

### Total Combined *GBA* and *GBAP1* Copy Number

A 10 kb region of unique sequence in between *GBA* and *GBAP1* is used to compute the copy number change due to reciprocal recombination events. Reads that align to this 10 kb region are counted and the count is normalized to a diploid baseline derived from 3000 preselected 2 kb regions across the genome. The 3000 normalization regions are randomly selected from the portion of the reference genome that has stable coverage across population samples. The total combined *GBA* and *GBAP1* copy number is then calculated as two more than the copy number of this 10 kb region.

### Nonrecombinant-like Variant Calling

Of the known nonrecombinant-like variants, some are in unique (nonhomologous) regions of *GBA* with high mapping quality. Only reads mapping to *GBA* are used for calling variant in nonhomologous regions. The other variants occur in homologous regions of *GBA*/*GBAP1* where reads mapping to either *GBA* or *GBAP1* are used for variant calling.

For each variant, reads containing the variant allele and the nonvariant alleles are counted. A binomial model that incorporates the sequencing error rate is then used to determine the most likely variant allele copy number (0 for nonvariant).

For a list of the supported nonrecombinant-like variants, refer to the `targeted/gba/target_variants_*.tsv` files located in the `resources` directory of the DRAGEN install location.

### Haplotype Phasing in the Exon 9-11 Region

A collection of 10 differentiating sites in the exon 9-11 region of *GBA* are used to detect the *GBA* and *GBAP1* haplotypes present in the sample. An iterative phasing algorithm is used to build up haplotypes that are supported by the read data. The phasing algorithm starts with seed sites which are then iteratively extended to neighboring sites. At each iteration, reads that can be unambiguously assigned to one of the detected partial haplotypes are used to extend the next neighboring site for each partial haplotype. Iteration continues until all sites have been extended. Some haplotypes may have sites that are unresolved (i.e. ambiguous), but these haplotypes can still participate in *GBA* -> *GBAP1* breakpoint detection.

### Nonallelic Homologous Recombination Variant Calling

If any of the 10 differentiating sites in exon 9-11 indicate that there is no wild type *GBA* allele copies, then the sample is called as homozygous variant and the recombinant-like variant that best matches the depth calls at the 10 sites is reported.

When the sample is not homozygous variant, the phased haplotypes are used to detect heterozygous variants. The detected haplotypes are compared against a set of 7 known recombinant-like variants: A495P, L483P, D448H, c.1263del, RecNciI, RecTL, c.1263del+RecTL). Whenever a detected haplotype has a *GBA*->*GBAP1* or *GBAP1*->*GBA* transition that is consistent with one of these 7 known recombinant-like variants, the transition is considered as a candidate breakpoint for calling that recombinant-like variant. Reads containing phasing information for the two sites flanking each candidate breakpoint are used for variant calling. When the read data supports the hypothesis that the sample contains at least one copy of a candidate breakpoint , the associated haplotype is a recombinant haplotype candidate. Recombinant haplotype candidates are sorted by likelihood and the number of variant sites. If no wild type haplotype was detected, DRAGEN reports any detected homozygous recombinant haplotype, or up to two different recombinant haplotypes (i.e. compound het) if detected. If any wild type haplotype was found, DRAGEN reports a maximum of one recombinant haplotype. When no recombinant haplotypes are detected two wild type haplotypes are reported.

The caller can detect the following recombinant variant haplotypes: A495P, L483P, D448H, 1263del, RecNciI, RecTL, and c.1263del+RecTL. Note: RecNciI, RecTL, and c.1263del+RecTL maye be deletion-like recombinant variants. A deletion-like recombinant variant haplotype (as opposed to a gene conversion-like recombinant variant haplotype) is defined as a haplotype with one or fewer switch sites (transitions from a *GBAP1* allele to a *GBA* allele).

The table below shows the HGVS identifiers associated with each recombinant variant haplotype.

| Recombinant variant haplotype | HGVS identifiers                                                                                     |
| ----------------------------- | ---------------------------------------------------------------------------------------------------- |
| A495P                         | NM\_000157.4:c.1483G>C                                                                               |
| L483P                         | NM\_000157.4:c.1448T>C                                                                               |
| D448H                         | NM\_000157.4:c.1342G>C                                                                               |
| c.1263del                     | NM\_000157.4:c.1265\_1319del                                                                         |
| RecNciI                       | NM\_000157.4:c.1483G>C, NM\_000157.4:c.1448T>C                                                       |
| RecTL                         | NM\_000157.4:c.1483G>C, NM\_000157.4:c.1448T>C, NM\_000157.4:c.1342G>C                               |
| c.1263del+RecTL               | NM\_000157.4:c.1483G>C, NM\_000157.4:c.1448T>C, NM\_000157.4:c.1342G>C, NM\_000157.4:c.1265\_1319del |

### GBA Caller Output File

The GBA Caller generates its output in the targeted caller output file `<output-file-prefix>.targeted.json` that also contains calls from other targets (see [Targeted JSON File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-json-file)).

### Output File Example

An example of the GBA caller content in the `<output-file-prefix>.targeted.json` output file is shown below.

```json
{
  "gba": {
    "totalCopyNumber": 4,
    "deletionBreakpointInGene": null,
    "recombinantHaplotypes": [
      "A495P",
      ""
    ],
    "recombinantHaplotypesFilter": "RecombinantSiteDepthMismatch",
    "phasedHaplotypes": {
      "targetAlleleDepths": [2, 2, 2, 2, 2, 2, 2, 2, 2],
      "targetAlleleDepthsQual": 4.122865423314576,
      "rawHaplotypes": [
        "TTTTTTTTT",
        "TTNTTTTTT",
        "NNNNNNNNN"
      ],
      "depthMatchedHaplotypes": {
        "targetAlleleDepths": [2, 2, 2, 2, 2, 2, 2, 2, 2],
        "targetAlleleDepthsQual": 4.122865423314576,
        "numMatchingHaplotypeSets": 1,
        "topHaplotypeSets": [
          {
            "haplotypes": [
              {
                "recombinantAlleleIDs": "",
                "haplotype": "TTTTTTTTT",
                "copyNumber": 2
              },
              {
                "recombinantAlleleIDs": "c.1263del+RecTL",
                "haplotype": "NNNNNNNNN",
                "copyNumber": 2
              }
            ]
          }
        ]
      }
    },
    "variants": []
  }
}
```


# HBA Caller

The HBA Caller is capable of genotyping the *HBA1* and *HBA2* genes from whole-genome sequencing (WGS) and whole-exome sequencing (WES) data. Due to high sequence similarity between the genes, a specialized caller is necessary to resolve the possible genotypes of the pair of genes. We consider regions surrounding the *HBA1* and *HBA2* sites to resolve the possible *HBA1* and *HBA2* genotypes.

The HBA Caller performs the following steps:

1. Determines total copy number from read depth of the regions surrounding the *HBA1* and *HBA2* sites.
2. Determines HBA genotypes based on the copy number of the regions surrounding the *HBA1* and *HBA2* sites.
3. Calls small variants in the *HBA1* and *HBA2* regions based on the region copy number derived from the genotype along with allele counts from read information.

For a comprehensive evaluation of the HBA caller, see [HBA targeted caller blog post](https://www.illumina.com/science/genomics-research/articles/HBA-targeted-caller.html).

For information about enabling the HBA caller see [Targeted Caller](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-caller).

## Total Copy Number of the regions surrounding the *HBA1* and *HBA2* sites

The first step of HBA calling is to determine the copy number of the regions surrounding the *HBA1* and *HBA2* sites. Reads aligned to the regions are counted. The counts in each region are corrected for GC-bias, and then normalized to a diploid baseline. The GC-bias correction and normalization factors are determined from read counts in preselected regions across the genome. Finally, a Gaussian Mixture Model (GMM) is used to obtain the integer region copy number from the region normalized counts.

## Genotyping

The genotyping step attempts to identify the two likely haplotypes described in the following table, where "a" stands for a functional copy of either *HBA1* or *HBA2*, "-" stands for a nonfunctional/missing copy of either *HBA1* or *HBA2*, while "3.7", "4.2", and "20.5" describe the recombinant event that likely caused the deletion/duplication of the functional HBA copy.

| Genotype        |
| --------------- |
| aaa3.7/aa       |
| aaa4.2/aa       |
| aaa20.5/aa      |
| aaa20.5/aaa3.7  |
| aaa20.5/aaa4.2  |
| aaa20.5/aaa20.5 |
| aaa3.7/aaa3.7   |
| aaa3.7/aaa4.2   |
| aaa4.2/aaa4.2   |
| aa/aa           |
| -a3.7/aa        |
| -a4.2/aa        |
| -a20.5/aa       |
| --/aaa3.7       |
| --/aaa4.2       |
| --/aaa20.5      |
| -a3.7/aaa4.2    |
| -a3.7/aaa20.5   |
| -a4.2/aaa20.5   |
| -a4.2/aaa3.7    |
| -a20.5/aaa3.7   |
| -a20.5/aaa4.2   |
| -a3.7/-a3.7     |
| -a4.2/-a4.2     |
| -a20.5/-a20.5   |
| -a3.7/-a4.2     |
| -a20.5/-a3.7    |
| -a20.5/-a4.2    |
| --/aa           |
| --/-a3.7        |
| --/-a4.2        |
| --/-a20.5       |
| --/--           |

The quality of fitting the region copy number calls to one of these genotypes is reported in the `FORMAT/TargetedSVModelQual` field for each SV VCF record and the `FILTER/TargetedSVModelQual` is applied to each record when the phred-scaled quality is below the threshold specified in the VCF header.

## Small Variant Calling

18 small variants are detected from the read alignments. These variants occur in homologous regions of *HBA1* and *HBA2* where reads mapping to either *HBA1* or *HBA2* are used for variant calling.

For each variant, reads containing either the variant allele or the nonvariant allele are counted and a binomial model is used to determine the likelihood for each possible variant allele copy number up to the maximum possible as determined from the *HBA1*/*HBA2* genotyping.

## HBA Output File

The HBA Caller generates its output in the targeted caller output file `<output-file-prefix>.targeted.json` that also contains calls from other targets (see [Targeted JSON File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-json-file)).

| Fields in JSON  | Explanation                                                    | Type and Possible Values                          |
| --------------- | -------------------------------------------------------------- | ------------------------------------------------- |
| totalCopyNumber | Total copy number of *HBA1* and *HBA2* genes including hybrids | nonnegative integer                               |
| genotype        | The HBA genotype.                                              | string                                            |
| genotypeFilter  | The HBA genotype filter.                                       | string, \[PASS, HBALowGQ, HBALowPValue, No\_call] |
| genotypeQual    | The HBA Phred genotype quality.                                | double                                            |
| variants        | List of detected homology region variants in *HBA1*/*HBA2*.    | Array of variants                                 |

Each variant reported in the `variants` array will have the fields below.

| Fields in JSON   | Explanation                                      | Type and Possible Values         |
| ---------------- | ------------------------------------------------ | -------------------------------- |
| alleleId         | HGVS identifier of the variant allele            | string                           |
| alleleCopyNumber | Copy number of the allele in the called genotype | nonnegative integer              |
| genotypeQuality  | Phred-scaled quality for the called genotype     | nonnegative integer              |
| filter           | Filter for the called genotype                   | string. "PASS" when not filtered |

Structural variant and homology region variants are reported in VCF format. See [Targeted VCF File](/dragen-v4.5/product-guides/dragen-v4.5/dragen-dna-pipeline/targeted-caller#targeted-vcf-file) for details about how these variants are reported in VCF.

### Output File Example

An example of the HBA caller content in the `<output-file-prefix>.targeted.json` output file is shown below.

```
  "hba": {
    "totalCopyNumber": 4,
    "genotype": "aa/aa",
    "genotypeQuality": 93,
    "genotypeFilter": "PASS",
    "variants": [
      {
        "alleleId": "NM_000517.6:c.95+1G>A",
        "alleleCopyNumber": 1,
        "genotypeQuality": 26,
        "filter": "PASS"
      }
    ]
  }
```




---

[Next Page](/llms-full.txt/1)

