> For the complete documentation index, see [llms.txt](https://help.dragen.illumina.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://help.dragen.illumina.com/dragen-v4.6/product-guides/dragen-v4.6/dragen-recipes/igg-large-cohort.md).

# iGG Large Cohort

A DRAGEN recipe, like this one, is a predefined set of analysis parameters and workflow settings tailored to a specific type of genomic analysis.

This is the default recommended recipe for most customers. The step-wise iGG workflow writes cohort and census files to disk, which enables iterative analysis (adding new sample batches without reprocessing existing ones). Use this recipe for cohorts of several thousand samples, with batching and sharding across servers.

For detailed option descriptions and extended workflow examples (hard filtering, machine learning filtering, spVCF output, and multi-batch shell workflows), see the [Iterative gVCF Genotyper user guide](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/iterative-gvcf-genotyper.md).

```
# Step 1: gVCF Aggregation (one job per batch and shard; parallelize shards across servers) 
/opt/dragen/$VERSION/bin/dragen         #DRAGEN install path 
--enable-gvcf-genotyper-iterative true  #enable iterative gVCF Genotyper 
--gvcfs-to-cohort-census true           #Step 1: aggregate gVCFs 
--variant-list $GVCF_LIST               #file listing gVCF paths for this batch 
--ht-reference $REF_FASTA               #path to reference FASTA 
--output-directory $OUTPUT/step1/batch${BATCH_NUM} 
--output-file-prefix batch${BATCH_NUM}_shard${SHARD_NUM} 
--shard ${SHARD_NUM}/100                #one genomic shard; run each shard on a separate server 
--gg-format-to-import AD,GQ,PL,DP       #FORMAT fields to import 
 
# Step 2: Census Aggregation (one job per shard; parallelize across servers) 
/opt/dragen/$VERSION/bin/dragen 
--enable-gvcf-genotyper-iterative true 
--aggregate-censuses true               #Step 2: aggregate censuses 
--input-census-list $CENSUS_LIST        #file listing per-batch census files 
--ht-reference $REF_FASTA 
--output-directory $OUTPUT/step2 
--output-file-prefix global_shard${SHARD_NUM} 
 
# Step 3: msVCF Generation (one job per shard; parallelize across servers) 
/opt/dragen/$VERSION/bin/dragen 
--enable-gvcf-genotyper-iterative true 
--generate-msvcf true                   #Step 3: generate msVCF 
--input-cohort-list $COHORT_LIST        #file listing per-batch cohort files 
--input-global-census-file $GLOBAL_CENSUS_FILE#from Step 2 
--ht-reference $REF_FASTA 
--output-directory $OUTPUT/step3 
--output-file-prefix shard${SHARD_NUM} 
--gg-msvcf-format-fields GT:GQ:LAD:FT:LPL:LAA:DP 
--gg-msvcf-info-fields "AC;AN;NS;NS_GT;IC;HWE" 
--gg-enable-indexing true 
```

## Notes and additional options

### Overview

This step-wise workflow is the default recommendation for most customers. Unlike end-to-end mode, Steps 1–3 write cohort and census files to disk, which is required for iterative analysis when new samples arrive.

For cohorts of several thousand samples, the iterative gVCF Genotyper uses a multi-step workflow that enables parallel processing and incremental updates:

1. **Step 1 (gVCF Aggregation)**: Process input gVCFs in batches of \~1000 samples, parallelized across genome shards
2. **Step 2 (Census Aggregation)**: Combine per-batch census files to create a global census
3. **Step 3 (msVCF Generation)**: Generate all-sample or per-batch msVCF files using the global census (ensures consistent variant sites across batches)

This approach allows for:

* Processing arbitrarily large cohorts
* Adding new samples (N+1 scenario) by rerunning only Steps 2–3
* Parallelization across hundreds or thousands of compute nodes

Use the [Small Cohort Recipe](/dragen-v4.6/product-guides/dragen-v4.6/dragen-recipes/igg-small-cohort.md) only for small trials that do not require iterative analysis.

### Prerequisites

Input gVCF files must be generated with the `--vc-emit-ref-confidence gVCF` option during germline variant calling. All gVCF files should use the same reference genome.

### Batching Strategy

Divide samples into batches of approximately 1000 samples each:

```bash

# Split gVCF file list into batches 

split -l 1000 -d --additional-suffix=.list gvcf_files.list batch_ 

# Creates batch_00.list, batch_01.list, batch_02.list, etc. 

```

**Recommended configuration**:

* Batch size: 1000 samples per batch
* Shards: 100 shards for whole genomes, 25-50 shards for exomes
* Total jobs in Step 1: `num_batches × num_shards`
* Total jobs in Step 2: `num_shards`
* Total jobs in Step 3: `num_shards`

### Genome Sharding and Parallel Jobs

For large cohorts, use `--shard X/N` so each DRAGEN job processes only one genomic shard (shard X of N). The recommended pattern is to run **many independent jobs in parallel**, with **each shard job on a different server or cluster node** (for example, one Slurm job per batch-and-shard pair in Step 1, or one job per shard in Steps 2 and 3).

Jobs do not need to share a host. They require the same reference, input list paths, and access to a shared output location (or equivalent staging). Submit jobs through your scheduler rather than running all shards sequentially on one machine.

| Option        | Description                                                                                                                                                          |
| ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--shard X/N` | Process only shard X out of N total shards. Run one DRAGEN invocation per job; scale out by submitting parallel jobs across servers (recommended for large cohorts). |

For illustrative shell loops and a full multi-batch example, see [Multi-batch workflow examples](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/iterative-gvcf-genotyper.md#multi-batch-workflow-examples) in the user guide.

### Step 1: gVCF Aggregation

Submit one job per batch and shard combination. Each job uses `--shard X/N` for a single genomic shard and can run on a separate server in parallel with all other Step 1 jobs.

| Option                          | Description                                                                                                          |
| ------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| `--gvcfs-to-cohort-census true` | Enable Step 1: aggregate gVCFs into cohort and census files.                                                         |
| `--variant-list $PATH`          | File listing input gVCF paths for this batch, one per line.                                                          |
| `--shard X/N`                   | Genomic shard to process. Submit one job per (batch, shard) pair; run jobs in parallel on separate servers.          |
| `--gg-format-to-import STRING`  | Comma-separated list of FORMAT fields to import (default: `AD,GQ,PL`). Include all fields needed in the final msVCF. |
| `--gg-info-to-import STRING`    | Comma-separated list of INFO fields to import (optional).                                                            |

**Output files** (per batch, per shard):

* `batch${BATCH_NUM}_shard${SHARD_NUM}.cohort` - Cohort file with aggregated gVCF data
* `batch${BATCH_NUM}_shard${SHARD_NUM}.census` - Census file with variant statistics

### Step 2: Census Aggregation

Submit one job per shard index. Each job merges census files from all batches for that shard only. Shard jobs are independent and can run in parallel on different servers.

| Option                      | Description                                                           |
| --------------------------- | --------------------------------------------------------------------- |
| `--aggregate-censuses true` | Enable Step 2: aggregate per-batch census files into a global census. |
| `--input-census-list $PATH` | File listing per-batch census files for this shard, one per line.     |

**Input**: Per-batch census files from Step 1 (one per batch for this shard)

**Output file** (per shard):

* `global_shard${SHARD_NUM}.census` - Global census containing variant statistics across all batches

### Step 3: msVCF Generation

Generate all-sample msVCF files using the global census.

Submit one job per shard index. Each job writes one msVCF for that genomic shard. Like Steps 1 and 2, shard jobs can run in parallel on separate servers.

| Option                             | Description                                                                                          |
| ---------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `--generate-msvcf true`            | Enable Step 3: generate multi-sample VCF.                                                            |
| `--input-cohort-list $PATH`        | File listing per-batch cohort files for this shard, one per line.                                    |
| `--input-global-census-file $PATH` | Global census file from Step 2 for this shard.                                                       |
| `--gg-msvcf-format-fields STRING`  | Colon-separated list of FORMAT fields to output (default: `GT:GQ:LAD:FT:LPL:LAA`).                   |
| `--gg-msvcf-info-fields STRING`    | Semicolon-separated list of INFO fields to output. Use `All` for all imported and calculated fields. |
| `--gg-enable-indexing BOOL`        | Create .tbi index (default=true).                                                                    |

**Output files** (per shard):

* `shard${SHARD_NUM}.vcf.gz` - Multi-sample VCF for this shard
* `shard${SHARD_NUM}.vcf.gz.tbi` - Tabix index

### Concatenating Shards

After Step 3 completes, concatenate shards using `bcftools`. This can potentially result in a very large file. Depending on your available disk space, proceed with caution. There is also the option to not concatenate all shards, or to generate msVCFs for a single batch or a subset of batches.

```bash

ls /results/step3/shard*.vcf.gz | sort -V > /tmp/shards.list 

bcftools concat -f /tmp/shards.list -O z -o /results/step3/merged.vcf.gz 

tabix -p vcf /results/step3/merged.vcf.gz 

```

**Output files**:

* `merged.vcf.gz` - Final merged multi-sample VCF containing all samples
* `merged.vcf.gz.tbi` - Tabix index

### Adding New Samples (N+1 Scenario)

When new samples arrive, you only need to rerun Steps 1-3 for the new batch, then Steps 2-3 for all batches:

1. **Run Step 1** for the new batch only (100 jobs if 100 shards)
2. **Rerun Step 2** with all census files including the new batch (100 jobs)
3. **Rerun Step 3** with the updated global census (100 jobs)
4. **Concatenate shards**

This approach reuses Step 1 outputs from existing batches and only processes new samples in that step.

For shell loop examples and a complete multi-batch workflow, see [Multi-batch workflow examples](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/iterative-gvcf-genotyper.md#multi-batch-workflow-examples) in the user guide.

### Licensing

The iterative gVCF Genotyper is a licensed feature metered by gigabases consumed from the input:

* Cost: metered in BioInsight Credits (BICs) per gigabase
* For OnPrem HPC, provide an Illumina BioInsight Platform API key via `--api-key-file $API_KEY_FILE`

See the [Iterative gVCF Genotyper Licensing section](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/iterative-gvcf-genotyper.md#iterative-gvcf-genotyper-licensing) for details, and [BioInsight Platform Licensing](/dragen-v4.6/reference/licensing/api_key_licensing.md#set-up-api-key-licensing) for API key setup.

Only iGG step 1 (cohort and census file generation) is metered and incurs a license fee.

### System Requirements

Before running, ensure adequate system limits:

```bash

ulimit -n 65536  # File handles 

ulimit -u 65536  # User processes 

```

### Additional Resources

For more information on the iterative gVCF Genotyper:

* [Iterative gVCF Genotyper User Guide](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/iterative-gvcf-genotyper.md) - Complete reference documentation with detailed option descriptions, output specifications, and technical details
* [Small Cohort Recipe](/dragen-v4.6/product-guides/dragen-v4.6/dragen-recipes/igg-small-cohort.md) - End-to-end mode for small trials only (no cohort or census files; no iterative analysis)
* [Pedigree Analysis](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/small-variant-calling/pedigree-analysis.md) - Information on family-based analysis and pedigree workflows
* [Hard filtering of msVCF output](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/iterative-gvcf-genotyper.md#msvcf-hard-filtering) - Options to apply hard filters on the msVCF output
* [Multi-batch workflow examples](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/iterative-gvcf-genotyper.md#multi-batch-workflow-examples) - Shell examples and distributed sharding guidance


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://help.dragen.illumina.com/dragen-v4.6/product-guides/dragen-v4.6/dragen-recipes/igg-large-cohort.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
