Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

IMPACT-SV

Structural variant (SV) preprocessing pipeline for the IMPACT workflow. This module processes SV VCF files through AnnotSV annotation and produces output files ready for visualization in IMPACT-VIS.

Overview

IMPACT-SV automates:

  • Installation of AnnotSV and all dependencies
  • Phenotype-aware SV annotation using candidate gene lists and HPO terms
  • Quality filtering based on read support metrics
  • Output formatting compatible with IMPACT-VIS

Validation

For a lightweight release smoke test that skips cleanly when AnnotSV is not installed, use tests/test_full_pipeline_optional.sh. It runs against the tiny NA12878 fixtures and validates the full release output set when AnnotSV is available.

Input: Structural variant VCF files (e.g., *.dragen.sv.vcf.gz) + SNV/indel VCF files + gene list. Output: {sample_id}_SV_IMPACT.tsv files ready for IMPACT-VIS.

v1.0.0 Contract

  • Official host support is Linux and WSL.
  • Multi-sample SV and SNV VCFs are out of scope for v1.0.0.
  • SV and SNV inputs are paired by the sample ID embedded in their filename pattern.
  • The basename pattern must contain exactly one * sample placeholder. A trailing * is allowed only for extension flexibility such as .vcf versus .vcf.gz.
  • Sample IDs are sanitized automatically for output naming and recorded in sample_manifest.tsv.
  • The pipeline fails closed if validation, AnnotSV, or post-processing fails.
  • Release provenance is recorded in run_manifest.json and {sample_id}_SV_IMPACT.metadata.json.

Quick Start

# 1. Clone the repository
git clone https://github.com/boehlernick/IMPACT-SV.git
cd IMPACT-SV

# 2. Install AnnotSV and the IMPACT-SV runtime
#    On Debian/Ubuntu/WSL, add -i to let the installer provision missing host tools with apt
bash scripts/install.sh -d ~/annotsv_install

# 3. Load the recommended runtime environment
source ~/annotsv_install/env.sh

# 4. Run a quick smoke test on the bundled tiny fixture
bash scripts/run_annotSV.sh \
  -w ./tests/NA12878 \
  -s "*.sv.tiny.vcf.gz" \
  -n "*.hard-filtered.tiny.vcf.gz" \
  -a ~/annotsv_install/AnnotSV \
  -o ./tests/output/quickstart

# 5. Confirm the IMPACT-VIS-ready output exists
ls ./tests/output/quickstart/*_SV_IMPACT.tsv

For production DRAGEN-style inputs, keep the default patterns and point -w at your data directory. The defaults now accept both DRAGEN-style and plain naming conventions: *.dragen.sv.vcf.gz|*.sv.vcf.gz and *_dragen.hard-filtered.vcf*|*.hard-filtered.vcf*. GeneList.txt is detected automatically in the working directory when you use -g GeneList.txt or omit -g entirely. VCF index sidecars such as .tbi are ignored during input discovery.

Requirements

  • Operating System: Linux or WSL (Windows Subsystem for Linux)
  • Permissions: sudo access only if you use scripts/install.sh -i or -s
  • Disk Space: plan for at least 15 GB free in the install filesystem; the final AnnotSV human annotation footprint is roughly 10 GB and installation needs extra working space
  • Python: Python 3.9+ with the venv module available

Input Files

File Type Naming Convention Description
SV VCF {sample_id}.dragen.sv.vcf.gz or {sample_id}.sv.vcf.gz Structural variant calls (bgzipped)
SNV VCF {sample_id}_dragen.hard-filtered.vcf* or {sample_id}.hard-filtered.vcf* Small variant calls for breakpoint validation (.vcf or .vcf.gz)
Gene List GeneList.txt One gene symbol per line (from Open Targets or similar)

Sample ID Contract:

  • The basename portion matched by * is treated as the sample ID.
  • SV and SNV patterns must each expose the same sample ID for a given sample.
  • Files must contain exactly one sample column in the VCF header.
  • Invalid sample ID characters are sanitized for outputs and recorded in sample_manifest.tsv.

Installation

bash scripts/install.sh [OPTIONS]

Options:
  -d DIR    Installation directory (default: $HOME/annotsv_install)
  -r REF    AnnotSV git tag or branch to install (default: v3.5.10)
  -i        Install missing system packages with apt (opt-in)
  -k        Deprecated alias for the default behavior of not installing system packages
  -p PYTHON Python interpreter used to create the IMPACT-SV virtualenv (default: python3)
  -s        Also copy Tcllib 1.20 system-wide
  -h        Show help

What gets installed:

  • System packages only when -i is provided
  • Tcllib 1.20 modules required by AnnotSV
  • A dedicated virtualenv at <install_dir>/impact-sv-venv
  • Pinned Python requirement set from requirements.txt installed into that virtualenv
  • AnnotSV pinned to v3.5.10 by default, with human genome annotations and the GenCC quote fix applied during installation
  • A generated environment file at <install_dir>/env.sh
  • A generated install manifest at <install_dir>/install_manifest.json
  • A persistent installer log at <install_dir>/logs/install.log

Space Planning:

  • scripts/install.sh now checks for at least 15 GB of free space before downloading AnnotSV human annotations.
  • Expect the final installed annotations to consume about 10 GB.
  • Keep additional headroom for the temporary AnnotSV build directory, the Python virtualenv, and future updates.

Install Behavior:

  • The human-annotation step is large and can take a long time.
  • Progress is written to <install_dir>/logs/install.log and summarized in the terminal while downloads and extraction continue.
  • If installation is interrupted, the staging directory under <install_dir>/.annotsv-build is retained so a subsequent run can resume from partial downloads.
  • AnnotSV activation now uses a rollback-safe swap: if activation fails, the previous AnnotSV directory is restored.
  • Installer metadata outputs (env.sh and install_manifest.json) are written atomically to reduce partial-file risk on interruption.

Recommended setup:

source ~/annotsv_install/env.sh

The pipeline will auto-detect the install-local Tcllib module path and Python environment from that installation, but sourcing env.sh is the clearest production entrypoint.

Tip: source the generated env.sh, not the virtualenv bin/ directory.

Testing the Release

These scripts are safe to run locally and are useful before tagging a release:

  • tests/test_sample_contract.sh validates the v1.0.0 sample pairing contract using the tiny NA12878 fixtures.
  • tests/test_read_support_filter.sh checks the IMPACT-VIS read-support post-processing without requiring AnnotSV.
  • tests/test_full_pipeline_optional.sh runs the full pipeline smoke test when AnnotSV is available and skips cleanly otherwise.

The sample-contract and full-pipeline tests use the tiny fixture patterns below because they do not follow the production DRAGEN naming convention:

*.sv.tiny.vcf.gz
*.hard-filtered.tiny.vcf.gz

The production defaults remain the DRAGEN-style patterns shown earlier in this README.

Running the Pipeline

Timeout control: set IMPACT_SV_ANNOTSV_TIMEOUT_SECONDS to override the per-sample AnnotSV timeout (default: 1800 seconds). Runtime preflight checks now fail fast if the gene list is empty, if manifest-listed VCF files are unreadable/corrupt gzip streams, or if TCLLIBPATH does not provide required Tcllib modules.

bash scripts/run_annotSV.sh [OPTIONS]

Options:
  -w DIR      Working directory containing VCF files (default: ./)
  -g FILE     Gene list file for candidate filtering (default: <work_dir>/GeneList.txt)
  -p TERM     HPO term for phenotype annotation (e.g., HP:0000365)
  -s PATTERN  SV VCF basename pattern with one '*' sample placeholder
  -n PATTERN  SNV VCF basename pattern with one '*' sample placeholder
  -o DIR      Output directory (default: <work_dir>/AnnotSV_output)
  -a DIR      AnnotSV installation directory (default: $HOME/annotsv_install/AnnotSV)
  -j NUM      Max parallel jobs (default: 4)
  -h          Show help

Example: Hearing Loss Analysis

bash scripts/run_annotSV.sh \
  -w ./sample_vcfs \
  -g ./hearing_loss_genes.txt \
  -p HP:0000365 \
  -a ~/annotsv_install/AnnotSV \
  -o ./results

Output

The pipeline produces one TSV file per sample in the output directory:

output/
├── run_manifest.json         # Run-level provenance and runtime parameters
├── sample_manifest.tsv        # Raw-to-sanitized sample ID report and SV/SNV pairing
├── SAMPLE_001/              # Per-sample AnnotSV working directory
│   ├── annotsv.log          # AnnotSV execution log
│   └── *.annotated.tsv      # Raw AnnotSV output
├── SAMPLE_001_SV_IMPACT.tsv # ← Ready for IMPACT-VIS
├── SAMPLE_001_SV_IMPACT.metadata.json
├── SAMPLE_002_SV_IMPACT.tsv
└── ...

Output Format

Each *_SV_IMPACT.tsv file contains:

  • All standard AnnotSV annotation columns
  • Sanitized Samples_ID and sample-column name aligned to the output filename
  • READ_SUPPORT_FILTERING column with values:
    • PASSED - Variant passes QC filters
    • FAILED - <reason> - Variant fails with specific reason (e.g., FAILED - GQ = 25)
    • Missing SR currently remains FAILED - SR Missing; the filter expects DRAGEN-style SV FORMAT fields that include PR and SR when available.

Each run also produces:

  • run_manifest.json with the IMPACT-SV release version, AnnotSV revision, runtime parameters, and validated sample list
  • {sample_id}_SV_IMPACT.metadata.json with per-sample provenance linking raw identifiers, input VCFs, and final outputs

Reliability note: pipeline output files (sample_manifest.tsv, run_manifest.json, {sample_id}_SV_IMPACT.metadata.json, and *_SV_IMPACT.tsv) are written atomically to reduce the risk of partial files after interruption.

Loading into IMPACT-VIS

Copy output files to your IMPACT-VIS data directory:

cp output/*_SV_IMPACT.tsv /path/to/IMPACT-VIS/app/data/

See IMPACT-VIS Data Preparation for details.

Troubleshooting

Issue Solution
AnnotSV binary not found Set -a flag to AnnotSV install directory
Missing host tools during install Install the reported packages yourself, or re-run scripts/install.sh -i on Debian/Ubuntu/WSL
Installer appears idle during annotation download Check <install_dir>/logs/install.log; the annotation step can run for a long time while large archives download and unpack
Interrupted install Re-run scripts/install.sh -d <install_dir> to reuse the retained staging checkout and partial downloads
TCLLIBPATH errors Re-run scripts/install.sh or set TCLLIBPATH to <install_dir>/tcllib1.20
Python interpreter missing pandas Source <install_dir>/env.sh or export IMPACT_SV_PYTHON=<install_dir>/impact-sv-venv/bin/python
Gene list file is empty Add at least one non-empty gene symbol line to the file passed with -g
file is not a readable gzip stream Re-export or regenerate the affected .vcf.gz input and confirm it is not truncated
Pattern validation errors Ensure each basename pattern contains one sample * placeholder
Sample ID mismatch Check sample_manifest.tsv and confirm filename and VCF header sample IDs agree after sanitization
Missing SNV partner Ensure every SV sample has exactly one SNV file exposing the same sample ID

Debug Mode

Check per-sample logs for AnnotSV errors:

cat output/SAMPLE_001/annotsv.log

See docs/release.md for the release checklist and provenance review steps.

License

GPL-3.0 - See LICENSE for details.

Acknowledgments

Citation

If you use IMPACT-SV in your research, please cite:

@article{impact-TBA,
  title={IMPACT: An Open-Source Workflow for Unified Variant Interpretation Using Phenotype-Driven Filtering},
  authors={Boehler, N. and Cheng, H. Y. M.},
  journal={TBA},
  year={TBA}
}

About

IMPACT-SV Preprocessing of structural variants for input into IMPACT-VIS

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages