Structural variant (SV) preprocessing pipeline for the IMPACT workflow. This module processes SV VCF files through AnnotSV annotation and produces output files ready for visualization in IMPACT-VIS.
IMPACT-SV automates:
- Installation of AnnotSV and all dependencies
- Phenotype-aware SV annotation using candidate gene lists and HPO terms
- Quality filtering based on read support metrics
- Output formatting compatible with IMPACT-VIS
For a lightweight release smoke test that skips cleanly when AnnotSV is not installed, use tests/test_full_pipeline_optional.sh. It runs against the tiny NA12878 fixtures and validates the full release output set when AnnotSV is available.
Input: Structural variant VCF files (e.g., *.dragen.sv.vcf.gz) + SNV/indel VCF files + gene list.
Output: {sample_id}_SV_IMPACT.tsv files ready for IMPACT-VIS.
- Official host support is Linux and WSL.
- Multi-sample SV and SNV VCFs are out of scope for v1.0.0.
- SV and SNV inputs are paired by the sample ID embedded in their filename pattern.
- The basename pattern must contain exactly one
*sample placeholder. A trailing*is allowed only for extension flexibility such as.vcfversus.vcf.gz. - Sample IDs are sanitized automatically for output naming and recorded in
sample_manifest.tsv. - The pipeline fails closed if validation, AnnotSV, or post-processing fails.
- Release provenance is recorded in
run_manifest.jsonand{sample_id}_SV_IMPACT.metadata.json.
# 1. Clone the repository
git clone https://github.com/boehlernick/IMPACT-SV.git
cd IMPACT-SV
# 2. Install AnnotSV and the IMPACT-SV runtime
# On Debian/Ubuntu/WSL, add -i to let the installer provision missing host tools with apt
bash scripts/install.sh -d ~/annotsv_install
# 3. Load the recommended runtime environment
source ~/annotsv_install/env.sh
# 4. Run a quick smoke test on the bundled tiny fixture
bash scripts/run_annotSV.sh \
-w ./tests/NA12878 \
-s "*.sv.tiny.vcf.gz" \
-n "*.hard-filtered.tiny.vcf.gz" \
-a ~/annotsv_install/AnnotSV \
-o ./tests/output/quickstart
# 5. Confirm the IMPACT-VIS-ready output exists
ls ./tests/output/quickstart/*_SV_IMPACT.tsvFor production DRAGEN-style inputs, keep the default patterns and point -w at your data directory.
The defaults now accept both DRAGEN-style and plain naming conventions: *.dragen.sv.vcf.gz|*.sv.vcf.gz and *_dragen.hard-filtered.vcf*|*.hard-filtered.vcf*.
GeneList.txt is detected automatically in the working directory when you use -g GeneList.txt or omit -g entirely.
VCF index sidecars such as .tbi are ignored during input discovery.
- Operating System: Linux or WSL (Windows Subsystem for Linux)
- Permissions:
sudoaccess only if you usescripts/install.sh -ior-s - Disk Space: plan for at least 15 GB free in the install filesystem; the final AnnotSV human annotation footprint is roughly 10 GB and installation needs extra working space
- Python: Python 3.9+ with the
venvmodule available
| File Type | Naming Convention | Description |
|---|---|---|
| SV VCF | {sample_id}.dragen.sv.vcf.gz or {sample_id}.sv.vcf.gz |
Structural variant calls (bgzipped) |
| SNV VCF | {sample_id}_dragen.hard-filtered.vcf* or {sample_id}.hard-filtered.vcf* |
Small variant calls for breakpoint validation (.vcf or .vcf.gz) |
| Gene List | GeneList.txt |
One gene symbol per line (from Open Targets or similar) |
Sample ID Contract:
- The basename portion matched by
*is treated as the sample ID. - SV and SNV patterns must each expose the same sample ID for a given sample.
- Files must contain exactly one sample column in the VCF header.
- Invalid sample ID characters are sanitized for outputs and recorded in
sample_manifest.tsv.
bash scripts/install.sh [OPTIONS]
Options:
-d DIR Installation directory (default: $HOME/annotsv_install)
-r REF AnnotSV git tag or branch to install (default: v3.5.10)
-i Install missing system packages with apt (opt-in)
-k Deprecated alias for the default behavior of not installing system packages
-p PYTHON Python interpreter used to create the IMPACT-SV virtualenv (default: python3)
-s Also copy Tcllib 1.20 system-wide
-h Show helpWhat gets installed:
- System packages only when
-iis provided - Tcllib 1.20 modules required by AnnotSV
- A dedicated virtualenv at
<install_dir>/impact-sv-venv - Pinned Python requirement set from
requirements.txtinstalled into that virtualenv - AnnotSV pinned to
v3.5.10by default, with human genome annotations and the GenCC quote fix applied during installation - A generated environment file at
<install_dir>/env.sh - A generated install manifest at
<install_dir>/install_manifest.json - A persistent installer log at
<install_dir>/logs/install.log
Space Planning:
scripts/install.shnow checks for at least 15 GB of free space before downloading AnnotSV human annotations.- Expect the final installed annotations to consume about 10 GB.
- Keep additional headroom for the temporary AnnotSV build directory, the Python virtualenv, and future updates.
Install Behavior:
- The human-annotation step is large and can take a long time.
- Progress is written to
<install_dir>/logs/install.logand summarized in the terminal while downloads and extraction continue. - If installation is interrupted, the staging directory under
<install_dir>/.annotsv-buildis retained so a subsequent run can resume from partial downloads. - AnnotSV activation now uses a rollback-safe swap: if activation fails, the previous AnnotSV directory is restored.
- Installer metadata outputs (
env.shandinstall_manifest.json) are written atomically to reduce partial-file risk on interruption.
Recommended setup:
source ~/annotsv_install/env.shThe pipeline will auto-detect the install-local Tcllib module path and Python environment from that installation, but sourcing env.sh is the clearest production entrypoint.
Tip: source the generated env.sh, not the virtualenv bin/ directory.
These scripts are safe to run locally and are useful before tagging a release:
tests/test_sample_contract.shvalidates the v1.0.0 sample pairing contract using the tiny NA12878 fixtures.tests/test_read_support_filter.shchecks the IMPACT-VIS read-support post-processing without requiring AnnotSV.tests/test_full_pipeline_optional.shruns the full pipeline smoke test when AnnotSV is available and skips cleanly otherwise.
The sample-contract and full-pipeline tests use the tiny fixture patterns below because they do not follow the production DRAGEN naming convention:
*.sv.tiny.vcf.gz
*.hard-filtered.tiny.vcf.gz
The production defaults remain the DRAGEN-style patterns shown earlier in this README.
Timeout control: set IMPACT_SV_ANNOTSV_TIMEOUT_SECONDS to override the per-sample AnnotSV timeout (default: 1800 seconds).
Runtime preflight checks now fail fast if the gene list is empty, if manifest-listed VCF files are unreadable/corrupt gzip streams, or if TCLLIBPATH does not provide required Tcllib modules.
bash scripts/run_annotSV.sh [OPTIONS]
Options:
-w DIR Working directory containing VCF files (default: ./)
-g FILE Gene list file for candidate filtering (default: <work_dir>/GeneList.txt)
-p TERM HPO term for phenotype annotation (e.g., HP:0000365)
-s PATTERN SV VCF basename pattern with one '*' sample placeholder
-n PATTERN SNV VCF basename pattern with one '*' sample placeholder
-o DIR Output directory (default: <work_dir>/AnnotSV_output)
-a DIR AnnotSV installation directory (default: $HOME/annotsv_install/AnnotSV)
-j NUM Max parallel jobs (default: 4)
-h Show helpbash scripts/run_annotSV.sh \
-w ./sample_vcfs \
-g ./hearing_loss_genes.txt \
-p HP:0000365 \
-a ~/annotsv_install/AnnotSV \
-o ./resultsThe pipeline produces one TSV file per sample in the output directory:
output/
├── run_manifest.json # Run-level provenance and runtime parameters
├── sample_manifest.tsv # Raw-to-sanitized sample ID report and SV/SNV pairing
├── SAMPLE_001/ # Per-sample AnnotSV working directory
│ ├── annotsv.log # AnnotSV execution log
│ └── *.annotated.tsv # Raw AnnotSV output
├── SAMPLE_001_SV_IMPACT.tsv # ← Ready for IMPACT-VIS
├── SAMPLE_001_SV_IMPACT.metadata.json
├── SAMPLE_002_SV_IMPACT.tsv
└── ...
Each *_SV_IMPACT.tsv file contains:
- All standard AnnotSV annotation columns
- Sanitized
Samples_IDand sample-column name aligned to the output filename READ_SUPPORT_FILTERINGcolumn with values:PASSED- Variant passes QC filtersFAILED - <reason>- Variant fails with specific reason (e.g.,FAILED - GQ = 25)- Missing
SRcurrently remainsFAILED - SR Missing; the filter expects DRAGEN-style SV FORMAT fields that includePRandSRwhen available.
Each run also produces:
run_manifest.jsonwith the IMPACT-SV release version, AnnotSV revision, runtime parameters, and validated sample list{sample_id}_SV_IMPACT.metadata.jsonwith per-sample provenance linking raw identifiers, input VCFs, and final outputs
Reliability note: pipeline output files (sample_manifest.tsv, run_manifest.json, {sample_id}_SV_IMPACT.metadata.json, and *_SV_IMPACT.tsv) are written atomically to reduce the risk of partial files after interruption.
Copy output files to your IMPACT-VIS data directory:
cp output/*_SV_IMPACT.tsv /path/to/IMPACT-VIS/app/data/See IMPACT-VIS Data Preparation for details.
| Issue | Solution |
|---|---|
AnnotSV binary not found |
Set -a flag to AnnotSV install directory |
| Missing host tools during install | Install the reported packages yourself, or re-run scripts/install.sh -i on Debian/Ubuntu/WSL |
| Installer appears idle during annotation download | Check <install_dir>/logs/install.log; the annotation step can run for a long time while large archives download and unpack |
| Interrupted install | Re-run scripts/install.sh -d <install_dir> to reuse the retained staging checkout and partial downloads |
TCLLIBPATH errors |
Re-run scripts/install.sh or set TCLLIBPATH to <install_dir>/tcllib1.20 |
| Python interpreter missing pandas | Source <install_dir>/env.sh or export IMPACT_SV_PYTHON=<install_dir>/impact-sv-venv/bin/python |
| Gene list file is empty | Add at least one non-empty gene symbol line to the file passed with -g |
file is not a readable gzip stream |
Re-export or regenerate the affected .vcf.gz input and confirm it is not truncated |
| Pattern validation errors | Ensure each basename pattern contains one sample * placeholder |
| Sample ID mismatch | Check sample_manifest.tsv and confirm filename and VCF header sample IDs agree after sanitization |
| Missing SNV partner | Ensure every SV sample has exactly one SNV file exposing the same sample ID |
Check per-sample logs for AnnotSV errors:
cat output/SAMPLE_001/annotsv.logSee docs/release.md for the release checklist and provenance review steps.
GPL-3.0 - See LICENSE for details.
- AnnotSV - Structural variant annotation
- IMPACT-VIS - Visualization module
If you use IMPACT-SV in your research, please cite:
@article{impact-TBA,
title={IMPACT: An Open-Source Workflow for Unified Variant Interpretation Using Phenotype-Driven Filtering},
authors={Boehler, N. and Cheng, H. Y. M.},
journal={TBA},
year={TBA}
}