> For the complete documentation index, see [llms.txt](https://help.dragen.illumina.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://help.dragen.illumina.com/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/epstein-barr-virus-detection.md).

# Epstein-Barr Virus Detection

## Overview

Epstein-Barr virus (EBV) can be detected in a sample by enabling human microbe detection (described in this page) or with the [oncovirus detection feature](/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/oncovirus-detection.md):

|                                     | Human Microbe Detection | Oncovirus Detection |
| ----------------------------------- | ----------------------- | ------------------- |
| Includes EBV                        | yes                     | yes                 |
| Includes other viruses              | no                      | yes                 |
| Detects viral integration           | no                      | yes                 |
| Calls EBV Type 1 vs Type 2          | yes                     | no                  |
| Requires downloading resource files | no                      | yes                 |

The analysis takes in unmapped reads, uses the DRAGEN *k*-mer classifier to identify whether a read originates from EBV, and determines to which reference sequence(s) it best matches. If EBV is detected, the analysis attempts to determine whether the detected virus is EBV type 1 or type 2. A TSV file with the results is generated.

EBV detection can be enabled with WGS, WES, and panels, but it is expected to perform best with WGS and panels with microbial probes.

EBV detection is not currently compatible with read collapsing and UMI should not be enabled when EBV detection is enabled.

## Command-Line Arguments

Detection is enabled with `--enable-human-microbe-detection=true`. An example command is given below where sample reads are analyzed for the presence of EBV:

```shell
dragen \
  --enable-human-microbe-detection true \
  --fastq-list $fastqList \
  --ref-dir $ref \
  --output-file-prefix $prefix \
  --output-directory $out
```

| Argument                                 | Type | Description                                                    | Default |
| ---------------------------------------- | ---- | -------------------------------------------------------------- | ------- |
| `--enable-human-microbe-detection`       | bool | Enables detection of EBV                                       | `false` |
| `--microbe-detection-all-reads`          | bool | Enable to use all reads instead of just unmapped reads         | `false` |
| `--microbe-detection-softclipped-reads`  | bool | Enable to keep softclipped reads in addition to unmapped reads | `false` |
| `--microbe-detection-below-threshold`    | bool | Enable to include below-threshold microbes in detections TSV   | `false` |
| `--microbe-detection-enable-read-output` | bool | Enable to create an output file with per-read results          | `false` |
| `--microbe-detection-num-threads`        | int  | Number of threads to use for processing reads                  | 8       |

## Output File

### Description

Enabling EBV detection creates an output TSV file at `$out/$prefix.microbe_detections.ebv.tsv` with the fields described below. Empty values are denoted in the TSV with a hyphen.

| Field                            | Description                                                                                                                                              |
| -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| microbe                          | Name of microbe                                                                                                                                          |
| sample                           | Name of sample                                                                                                                                           |
| microbe\_detected                | Value is "detected" if microbe metrics are above thresholds                                                                                              |
| rpkm                             | Relative abundance of microbe                                                                                                                            |
| region\_name                     | Name of references used for detection                                                                                                                    |
| overall\_read\_count             | Number of reads that classified to any of the microbe's references                                                                                       |
| overall\_kmer\_fraction          | Weighted average of k-mer fractions for best-match references                                                                                            |
| read\_count\_threshold           | Pre-defined threshold used with overall\_read\_count                                                                                                     |
| kmer\_fraction\_threshold        | Pre-defined threshold used with overall\_kmer\_fraction                                                                                                  |
| best\_match\_ref\_accession      | The references with the highest *k*-mer fraction                                                                                                         |
| best\_match\_ref\_read\_count    | Read counts for best-match references                                                                                                                    |
| best\_match\_ref\_kmer\_fraction | *k*-mer fractions for best-match references                                                                                                              |
| best\_match\_ref\_length         | Lengths of best-match references                                                                                                                         |
| best\_match\_ref\_completeness   | Length of the best-match reference compared to the RefSeq reference for this virus; always 1.0 since only RefSeq references are used in the EBV analysis |

The file always contains rows for `Epstein-Barr virus (EBV)`, `Epstein-Barr virus (EBV) type 1`, and `Epstein-Barr virus (EBV) type 2`.

The *k*-mer fraction quantifies how much of a reference sequence is supported by the sequencing data. First, all canonical *k*-mers are enumerated from the reference sequence. The *k*-mer fraction is then calculated as the proportion of these reference *k*-mers that are observed at least once in the reads. A value close to 1 indicates broad coverage across the reference, whereas lower values indicate partial or sparse support.

### Example

The following example output file was generated with HG02790, a sample from the 1000 genomes project that was reported to have an average depth of 1.43 for EBV.

| microbe                         | sample  | microbe\_detected | rpkm      | region\_name                  | overall\_read\_count | overall\_kmer\_fraction | read\_count\_threshold | kmer\_fraction\_threshold | best\_match\_ref\_accession                        | best\_match\_ref\_read\_count | best\_match\_ref\_kmer\_fraction | best\_match\_ref\_length | best\_match\_ref\_completeness |
| ------------------------------- | ------- | ----------------- | --------- | ----------------------------- | -------------------- | ----------------------- | ---------------------- | ------------------------- | -------------------------------------------------- | ----------------------------- | -------------------------------- | ------------------------ | ------------------------------ |
| Epstein-Barr virus (EBV)        | HG02790 | detected          | 0.0356463 | genome                        | 1036                 | 0.602                   | 5                      | 0.050                     | NC\_007605.1                                       | 1008                          | 0.602                            | 85509                    | 1.000                          |
| Epstein-Barr virus (EBV) type 1 | HG02790 | detected          | 0.0531812 | EBNA2; EBNA3A; EBNA3B; EBNA3C | 193                  | 0.605                   | 5                      | 0.300                     | K03333.1; NC\_007605.1; NC\_007605.1; NC\_007605.1 | 114; 23; 30; 26               | 0.573; 0.638; 0.743; 0.483       | 3399; 2496; 2460; 2619   | 1.000; 1.000; 1.000; 1.000     |
| Epstein-Barr virus (EBV) type 2 | HG02790 | -                 | -         | -                             | -                    | -                       | 5                      | 0.300                     | -                                                  | -                             | -                                | -                        | -                              |

## Detection Logic

In order to be considered detected, a microbe must pass its read count and *k*-mer fraction thresholds. For `Epstein-Barr virus (EBV)`, the whole genome is considered. For `Epstein-Barr virus (EBV) type 1` and `Epstein-Barr virus (EBV) type 2`, only the genotyping genes *EBNA2*, *EBNA3A*, *EBNA3B*, and *EBNA3C* are considered.

|           Microbe Name          | Read Count Threshold | K-mer Fraction Threshold |                 Genes Used                |
| :-----------------------------: | :------------------: | :----------------------: | :---------------------------------------: |
|     Epstein-Barr virus (EBV)    |           5          |           0.05           |                   Genome                  |
| Epstein-Barr virus (EBV) type 1 |           5          |           0.30           | *EBNA2*, *EBNA3A*, *EBNA3B*, and *EBNA3C* |
| Epstein-Barr virus (EBV) type 2 |           5          |           0.30           | *EBNA2*, *EBNA3A*, *EBNA3B*, and *EBNA3C* |

Additional logic is applied when both types 1 and 2 are above their thresholds. Whenever type 1 is present, there is some signal for type 2 as well, and vice versa, due to the similarity in their sequence. Therefore, only the type with the highest *k*-mer fraction is reported as detected.

However, in order to be able to report clear dual positives, when both types 1 and 2 have a *k*-mer fraction above 0.95, both are reported as detected.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://help.dragen.illumina.com/dragen-v4.6/product-guides/dragen-v4.6/dragen-dna-pipeline/epstein-barr-virus-detection.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
