You can edit almost every page by Creating an account and confirming your email.

SimPlot (software)

From EverybodyWiki Bios & Wiki



SimPlot [1] is a Windows application allowing users to produce high-quality sequence similarity plots.

SimPlot++ [2] is a freely available reinterpretation of SimPlot. SimPlot++ is an open-source multi-platform application developed at the department of Computer Science of the Université du Québec à Montréal. SimPlot++ can be used to produce publication quality sequence similarity plots using 63 nucleotide and 20 amino acid distance models, to detect intergenic and intragenic recombination events using Phi [3], χ2 [4], NSS [5] or proportion tests, and to generate and analyze interactive sequence similarity networks. SimPlot++ supports multicore data processing and provides useful distance calculability diagnostics. As such, SimPlot++ improves on the original tools offered by SimPlot, such as SimPlot, BootScan and FindSites, and provides users with a new similarity network feature[2].

SimPlot++ is available on github as source code from Windows, MacOS and Linux, and as an executable for Windows.

General Use

SimPlot++ requires a multiple sequence alignment input file containing either DNA or Amino acid sequences. The input data must be in one of the following formats: FASTA, Nexus, PIR, PHYLIP, Stockholm or Clustal.

Once loaded into SimPlot++, the input sequences must be manually separated into different groups by the user, based on their evolutionary proximity in order to generate consensus sequences[6]. These consensus sequences will be used to perform any analysis available in SimPlot++. A minimum of two groups must be created in order to have full access to the application features[2].

SimPlot analysis

Typical SimPlot output provided by SimPlot++.

A SimPlot analysis uses a window of a specified size and a specified advancement step to slide this window over the multiple sequence alignment (MSA)[1]. Every sub-MSA covered by the window is extracted and a distance matrix based on a selected distance model is generated. This distance matrix is then used to produce a similarity plot for every consensus sequence against the reference sequence chosen by the user. The variations in similarity between the reference sequence and the consensus sequences can be used, for example, to detect potential recombination events[1].

The SimPlot analysis offers many features[2]:

  • 43 DNA and 20 amino acid distance models are available
  • Multiprocessing functionality is available
  • Matplotlib-based plots with a toolbar to easily customize and save the outputs in multiple formats[7]
  • A new quality control window will open an interactive HTML page to access additional information with the distance calculability diagnostic

BootScan analysis

Typical BootScan output provided by SimPlot++.

Bootscanning[8] is a pipeline consisting of 4 main steps, all done using a sliding window analysis (as in the SimPlot analysis).

Steps: [8]

  1. The subsequences extracted from the consensus groups are bootstrapped N times.
  2. For each of the N bootstrapped sub-MSAs, a distance matrix is generated.
  3. A phylogenetic tree[9][10] is inferred for each distance matrix (either with Neighbor joining[11] or UPGMA[12]).
  4. The conflicting phylogenetic signals are quantified and expressed as the percentage of trees where each sequence is the nearest neighbor of the reference sequence.

This analysis offers the following features[2]:

  1. 43 DNA distance models are available for generating the distance matrices
  2. Multiprocessing functionality is available
  3. Matplotlib-based plots with a toolbar to easily customize and save the outputs in multiple formats[7]

FindSites

The FindSites scan is used for locating possible regions of recombination by identifying Informative sites[13]. The first step of the analysis is to select a sequence assumed to be originated from a recombination event as well as two sequences of interest (one from each of the two possible parental evolutionary lines), and a fourth sequence as an outgroup. Informative sites will be identified as those where, at the same position, two of the sequences share the same nucleotide, and the other two sequences share another (different) nucleotide[13].

Similarity Network

Typical Similarity Network output provided by SimPlot++. The network nodes represent different sequences, and the edges represent both the global (in black) and local sequence similarity (in red).

The sequence similarity network[14] [15] analysis is an interactive representation of a SimPlot analysis using a window in which every group (including the reference group) is represented by a network node. These nodes are connected by an edge depending on the calculated global (over the whole sequence) or local (over sub-sequences of a selected length) similarity[2].

By adjusting the minimum similarity threshold required to show each of the edge types (global and local), it is possible to get a better insight on the relationships between every group. Furthermore, the network similarity representation can be limited to a specific range of the full MSA (in order to analyze a gene or region of interest)[2].

The graph data and visualization can be saved in an HTML file. The graph itself can be saved as either a .png or .svg directly from the toolbox in the HTML file.

Recombination analysis

Statistical tests for detecting recombination events from PhiPack[3] have been implemented in SimPlot++.

The Phi[3], Phi-profile[3], Max χ2[4] and NSS[5] tests are available for both the ungrouped (raw sequences) and grouped consensus sequences.

Moreover, a new simple Proportion test has been designed as a complement to the traditional SimPlot analysis in order to identify quickly the most likely mosaic regions[16] (i.e. possible recombination events) in the grouped sequences. This test is based on the proportion of genetic distances extracted from the SimPlot distance matrices. The Proportion score is an indicator of the signal strength but should not be always considered as a recombination signal[2].

References

  1. 1.0 1.1 1.2 Lole, Kavita S.; Bollinger, Robert C.; Paranjape, Ramesh S.; Gadkari, Deepak; Kulkarni, Smita S.; Novak, Nicole G.; Ingersoll, Roxann; Sheppard, Haynes W.; Ray, Stuart C. (January 1999). "Full-Length Human Immunodeficiency Virus Type 1 Genomes from Subtype C-Infected Seroconverters in India, with Evidence of Intersubtype Recombination". Journal of Virology. 73 (1): 152–160. doi:10.1093/bioinformatics/btac287. ISSN 0022-538X. PMID 35451456 Check |pmid= value (help).
  2. 2.0 2.1 2.2 2.3 2.4 2.5 2.6 2.7 Samson, Stéphane; Lord, Étienne; Makarenkov, Vladimir (April 2022). "SimPlot++: a Python application for representing sequence similarity and detecting recombination". Bioinformatics. arXiv:2112.09755. doi:10.1093/bioinformatics/btac287. PMID 35451456 Check |pmid= value (help).
  3. 3.0 3.1 3.2 3.3 Bruen, Trevor C; Philippe, Hervé; Bryant, David (1 April 2006). "A Simple and Robust Statistical Test for Detecting the Presence of Recombination". Genetics. 172 (4): 2665–2681. doi:10.1534/genetics.105.048975. PMC 1456386. PMID 16489234.
  4. 4.0 4.1 Smith, JohnMaynard (February 1992). "Analyzing the mosaic structure of genes". Journal of Molecular Evolution. 34 (2): 126–129. Bibcode:1992JMolE..34..126S. doi:10.1007/BF00182389. PMID 1556748. Unknown parameter |s2cid= ignored (help)
  5. 5.0 5.1 Jakobsen, Ingrid B.; Easteal, Simon (1996). "A program for calculating and displaying compatibility matrices as an aid in determining reticulate evolution in molecular sequences". Bioinformatics. 12 (4): 291–295. doi:10.1093/bioinformatics/12.4.291. PMID 8902355.
  6. Schneider, Thomas D. (2002). "Consensus Sequence Zen". Applied Bioinformatics. 1 (3): 111–119. ISSN 1175-5636. PMC 1852464. PMID 15130839.
  7. 7.0 7.1 Hunter, John D. (2007). "Matplotlib: A 2D Graphics Environment". Computing in Science & Engineering. 9 (3): 90–95. Bibcode:2007CSE.....9...90H. doi:10.1109/MCSE.2007.55. Unknown parameter |s2cid= ignored (help)
  8. 8.0 8.1 Salminen, Mika O.; Carr, Jean K.; Burke, Donald S.; McCUTCHAN, Francine E. (1 November 1995). "Identification of Breakpoints in Intergenotypic Recombinants of HIV Type 1 by Bootscanning". AIDS Research and Human Retroviruses. 11 (11): 1423–1425. doi:10.1089/aid.1995.11.1423. ISSN 0889-2229. PMID 8573403.
  9. Felsenstein, Joseph (2004). Inferring phylogenies. Sunderland, MA: Sinauer associates. Search this book on
  10. Penny, David (1 August 2004). "Inferring Phylogenies.—Joseph Felsenstein. 2003. Sinauer Associates, Sunderland, Massachusetts". Systematic Biology. 53 (4): 669–670. doi:10.1080/10635150490468530.
  11. Saitou, N.; Nei, M. (July 1987). "The neighbor-joining method: a new method for reconstructing phylogenetic trees". Molecular Biology and Evolution. 4 (4): 406–425. doi:10.1093/oxfordjournals.molbev.a040454. ISSN 0737-4038. PMID 3447015.
  12. Sokal, Michener (1958). "A statistical method for evaluating systematic relationships". University of Kansas Science Bulletin. 38: 1409–1438.
  13. 13.0 13.1 Robertson, David L.; Hahn, Beatrice H.; Sharp, Paul M. (1 March 1995). "Recombination in AIDS viruses". Journal of Molecular Evolution. 40 (3): 249–259. Bibcode:1995JMolE..40..249R. doi:10.1007/BF00163230. ISSN 1432-1432. PMID 7723052. Unknown parameter |s2cid= ignored (help)
  14. Xing, Henry; Kembel, Steven W; Makarenkov, Vladimir (1 May 2020). "Transfer index, NetUniFrac and some useful shortest path-based distances for community analysis in sequence similarity networks". Bioinformatics. 36 (9): 2740–2749. doi:10.1093/bioinformatics/btaa043. PMID 31971565.
  15. Bapteste, Eric; van Iersel, Leo; Janke, Axel; Kelchner, Scot; Kelk, Steven; McInerney, James O.; Morrison, David A.; Nakhleh, Luay; Steel, Mike; Stougie, Leen; Whitfield, James (August 2013). "Networks: expanding evolutionary thinking". Trends in Genetics. 29 (8): 439–441. doi:10.1016/j.tig.2013.05.007. PMID 23764187.
  16. Forsberg, Lars A.; Gisselsson, David; Dumanski, Jan P. (February 2017). "Mosaicism in health and disease — clones picking up speed". Nature Reviews Genetics. 18 (2): 128–142. doi:10.1038/nrg.2016.145. PMID 27941868. Unknown parameter |s2cid= ignored (help)

External links


This article "SimPlot (software)" is from Wikipedia. The list of its authors can be seen in its historical and/or the page Edithistory:SimPlot (software). Articles copied from Draft Namespace on Wikipedia could be seen on the Draft Namespace of Wikipedia and not main one.