Home » Uncategorized » Variant Annotation Tools and Co.

 
 

Variant Annotation Tools and Co.

 

Svetlana Frenkel

Svetlana Frenkel

Variant Annotation Tools and Tools for Genomic Variants processing

VCFtools is a “program package designed for working with VCF files, such as those generated by the 1000 Genomes Project. The aim of VCFtools is to provide easily accessible methods for working with complex genetic variation data in the form of VCF files.”

Used for VCF processing: cleaning, merging, calculating of allele frequencies and number of homozygotes, comparing lists of variants, filtering variants by type.

PolyPhen-2 (Polymorphism Phenotyping v2) is a “tool which predicts possible impact of an amino acid substitution on the structure and function of a human protein using straightforward physical and comparative considerations.”

I tried this tool (as web-service) only for detecting CpG SNPs. 

Real-time Genomic Analysis – iobio uses immediate visual feedback to make understanding complex genomic datasets more intuitive, and analysis more interactive

very nice graphs

ANNOVAR is an “efficient software tool to utilize update-to-date information to functionally annotate genetic variants detected from diverse genomes.

  • Gene-based annotation: identify whether SNPs or CNVs cause protein coding changes and the amino acids that are affected. Users can flexibly use RefSeq genes, UCSC genes, ENSEMBL genes, GENCODE genes, AceView genes, or many other gene definition systems.
  • Region-based annotation: identify variants in specific genomic regions, for example, conserved regions among 44 species, predicted transcription factor binding sites, segmental duplication regions, GWAS hits, database of genomic variants, DNAse I hypersensitivity sites, ENCODE H3K4Me1/H3K4Me3/H3K27Ac/CTCF sites, ChIP-Seq peaks, RNA-Seq peaks, or many other annotations on genomic intervals.
  • Filter-based annotation: identify variants that are documented in specific databases, for example, whether a variant is reported in dbSNP, what is the allele frequency in the 1000 Genome Project, NHLBI-ESP 6500 exomes or Exome Aggregation Consortium, calculate the SIFT/PolyPhen/LRT/MutationTaster/MutationAssessor/FATHMM/MetaSVM/MetaLR scores, find intergenic variants with GERP++ score < 2, or many other annotations on specific mutations.
  • Other functionalities: Retrieve the nucleotide sequence in any user-specific genomic positions in batch, identify a candidate gene list for Mendelian diseases from exome data, and other utilities.”

SnpEff – Genetic variant annotation and effect prediction toolbox.

EPACTS (Efficient and Parallelizable Association Container Toolbox) is a “versatile software pipeline to perform various statistical tests for identifying genome-wide association from sequence data through a user-friendly interface, both to scientific analysts and to method developers.”

Combined Annotation Dependent Depletion (CADD) – is a tool for scoring the deleteriousness of single nucleotide variants as well as insertion/deletions variants in the human genome.

“While many variant annotation and scoring tools are around, most annotations tend to exploit a single information type (e.g. conservation) and/or are restricted in scope (e.g. to missense changes). Thus, a broadly applicable metric that objectively weights and integrates diverse information is needed. Combined Annotation Dependent Depletion (CADD) is a framework that integrates multiple annotations into one metric by contrasting variants that survived natural selection with simulated mutations.

C-scores strongly correlate with allelic diversity, pathogenicity of both coding and non-coding variants, and experimentally measured regulatory effects, and also highly rank causal variants within individual genome sequences. Finally, C-scores of complex trait-associated variants from genome-wide association studies (GWAS) are significantly higher than matched controls and correlate with study sample size, likely reflecting the increased accuracy of larger GWAS.”

web-service, scripts and pre-calculated scores are available

GATK – “The Genome Analysis Toolkit or GATK is a software package for analysis of high-throughput sequencing data, developed by the Data Science and Data Engineering group at the Broad Institute. The toolkit offers a wide variety of tools, with a primary focus on variant discovery and genotyping as well as strong emphasis on data quality assurance. Its robust architecture, powerful processing engine and high-performance computing features make it capable of taking on projects of any size.”

GEMINI (GEnome MINIng) is a “flexible framework for exploring genetic variation in the context of the wealth of genome annotations available for the human genome. By placing genetic variants, sample phenotypes and genotypes, as well as genome annotations into an integrated database framework, GEMINI provides a simple, flexible, and powerful system for exploring genetic variation for disease and population genetics.

Using the GEMINI framework begins by loading a VCF file (and an optional PED file) into a database. Each variant is automatically annotated by comparing it to several genome annotations from source such as ENCODE tracks, UCSC tracks, OMIM, dbSNP, KEGG, and HPRD. All of this information is stored in portable SQLite database that allows one to explore and interpret both coding and non-coding variation using “off-the-shelf” tools or an enhanced SQL engine.”

Variant Annotation Integrator by UCSC Genome Bioinformatics Site.

Variant Annotation databases

dbSNP’s FTP Site – many organisms, different formats (BED, VSF…)
ClinVar – clinically significant human variants, including medical phenotypes

Lift Genome Annotations – This tool converts genome coordinates and genome annotation files between assemblies.

 

 

No comments

Be the first one to leave a comment.

Post a Comment