Description
A computational study for identifying agronomically important biomarkers using nsSNP data of Korean soybeans
ABSTRACT
We present a list of biomarker candidates selected from nsSNPs of Glycine max var. Williams 82 soybean genes. The nsSNPs were obtained comparing the sequences of W 82 soybean genes with those of 10 cultured and 6 wild soybean genes in Korea. We considered the nsSNPs which have made structural and functional changes significantly as biomarker candidates accompanying phenotypic traits. To extract a short list of the nsSNPs, we used four software such as SIFT, Polyphen, PANTHER, and I-mutant 2.0 which are designed to well sense the amount of the functional and structural variations caused by the nsSNPs. Then we finally narrowed down the number of nsSNPs by taking the nsSNPs appeared in more than half of 16 soybean species. A short list of biomarker candidates extracted will be the valuable property in searching the realistic biomarkers in the future.
Key words: Biomarker; nsSNPs; Korean soybean
INTRODUCTION
Although soybean is a self-pollinating species with severe genetic bottlenecks during the course of domestication [81% of the rare alleles have been lost (Hyten et al., 2006)] and has relatively low nucleotide polymorphism rates compared with other crop species (Haun et al, 2011), it still exhibits a wide range of phenotypic variation between (or within) wild and cultivated soybeans (Fasoula and Boerma, 2005).
Soybean genetic variations tend to be caused by a change in a single nucleotide in DNA which is called single nucleotide polymorphism (SNP). Although the SNPs in soybeans do not occur frequently, most SNP with a wide range of heritable phenotypic variation seems to be neutral (Haun et al, 2011), but some non-synonymous single nucleotide polymorphisms (nsSNPs) in coding or regulatory regions alter both protein structure and protein function significantly. For instance, wild soybean (G. soja) is different from cultivated soybean (G. max) in terms of flowering time and grain weight, even though they have almost same amino acid sequences except small number of nsSNPs. These phenotype alterations may come from nsSNPs which affect gene expression and alter splice sites, eventually producing new gene products. These nsSNPs are reported to introduce or expel unexpected charged residue or inflexible amino acids with benzene ring(s) or some amino acids which significantly alter the size of the residue (Yana Bromberg and Burkhard Rost, 2007). Thus we presume that some of nsSNPs of soybeans must be biomarker candidates.
This line of reasoning led to the idea of extraction of a large nsSNP set from 16 wild and cultivated [6 wild (IT162825, IT178480, IT182869, IT182840, IT182848, and IT182932) and 10 cultivated (Sowon, Peuren, Kangyo, Ilpumgeomjeong, Seoritae, PI96983, Haman, Kumjeongol, Hwangkeum, and Williams 82k)] soybeans.
Here, we chose Cultivar William 82 as the soybean reference to compare its sequence with those of 16 cultivated soybean species in Korea. It is reported that the genetic heterogeneity in Williams 82 primarily originated from the differential segregation of polymorphic chromosomal regions following the backcross and single-seed descent generations of the breeding process (Haun et al, 2011).
However, it is unlikely to investigate their functional impact of all the nsSNPs (more than 100, 000) extracted by accomplishing wet-lab experiments. Alternatively, bioinformatics approaches have increased the prediction of the functional effect of nsSNPs. So, to predict phenotypic effects of amino acid substitution such as seed composition, seed weight, seed size, and plant height, we applied the computational methods using four representative in silico software such as SIFT (Sorting Intolerant from Tolerant)(Pauline C. Ng et al., 2003), Polyphen (Polymorphism Phenotyping) (Ramensky V et al., 2002), PANTHER (Protein Analysis Through Evolutionary Relationships) (Thomas PD et al., 2003), and I-Mutants. This process (Yuan Ming Di, 2009) has the prospective to narrow down biomarker candidates which will be needed for future QTL investigations.
These computational methods are not always guaranteed to find real biomarkers due to the severe dependence on their exclusively theoretical judgments. However, it is manifest that they will be at least the biomarker candidates in selecting and prioritizing a small number of likely and tractable biomarkers from pools of available data (Yana Bromberg and Burkhard Rost, 2007). Therefore, the main purpose of this paper is to present a short list of nsSNPs which may have significantly altered the structure and function during the course of domestication using 16 wild and cultivated soybean genes resequenced.
MATERIALS AND METHODS
The nsSNP data (from 10 cultured and 6 wild soybeans in Korea) used in this study were obtained from the Korean Bioinformation Center (KOBIC). We used four computing software to predict whether the nsSNP data were biomarker candidate or not. They were SIFT, PolyPhen, PANTHER, and I-Mutants and performed in LINUX environment at the bio-evaluation center at Korea Research Institute of Bioscience and Biotechnology.
A brief summary of the software functions is as follows.
- A SIFT software (https://sift.bii.a-star.edu.sg): It is known to distinguish between functionally tolerated and deleterious amino acid changes in mutagenesis studies. It predicts whether an amino acid substitution will affect protein function and hence, potentially alter phenotype, based on sequence homology (the degree of conservation of amino acid residues in sequence alignments derived from closely related sequences, collected through PSI-BLAST) and the physical properties of amino acids, resulting in either “Tolerated (more than 0.05) ” or “Deleterious (less than 0.05)”. To be brief, SIFT presumes that important amino acids will be conserved in the protein family, and so changes at well-conserved positions tend to be predicted as “deleterious”.
- A Polyphen-2 software (http://genetics.bwh.harvard.edu/pph2): It is a new development of the Polyphen tool for annotating coding ns SNPs. It determines the property of nsSNPs based on a combination of phylogenetic, structural, and sequence annotation information characterizing a substitution and its position in the protein. The Polyphen values range from 0 (tolerated) and 1 (most likely to be deleterious) according to the degree of structural change. Mapping of amino acid replacement to the known 3D structure reveals whether the replacement is likely to destroy the hydrophobic core of a protein, electrostatic interactions, interactions with ligands or other important features of a protein. If the spatial structure of a query protein is unknown, one can use the homologous proteins with known structure (Adzhubei IA et al., 2010).
- A PANTHER software (http://www.pantherdb.org): It predicts the functional impact of the related nsSNPs even in the absence of direct experimental evidence based on published scientific experimental evidence and evolutionary relationships (Thomas et al., 2003). The PANTHER values range from 0 (neutral) to (approx.) -10 (most likely to be deleterious) according to the degree of functional impairment. The PSCI (substitution position-specific evolutionary conservation) values less than -3 and the deleterious probability more than 0.5 are assumed to be “deleterious” [-3 of PSCI equals to 0.5 of probability deleterious].
- An I-Mutant 2.0 software (http://folding.biofold.org/i-mutant/i-mutant2.0.html): It checks the stability of the proteins based on the analysis for solvent accessibility (NetASA) and secondary structure (DSSP).
Nature generally prefers stable conditions. The values less than zero are assumed to be “unstable” or “deleterious”. So if the nsSNPs involved make any protein structure unstable, they will be classified as “deleterious” and are presumed having special aims to exist.
RESULTS
We considered multiple ways an nsSNP impacts on protein structure and function. SNPs with the highest likelihood of being functionally relevant and therefore most important to interrogate were ranked based on SIFT, PolyPhen and, PANTHER, and I Mutant 2.0 scores. We selected 91 most likely candidates that may impact their structures and functions (see APPENDIX).
| # B: a set of genes with nsSNPs exclusively appeared in wild species, A: a set of genes with nsSNPs exclusively appeared in cultivated species, and AB: a set of genes with nsSNPs simultaneously appeared in both wild and cultivated species.
# The maximum number of nsSNP per variation position of each gene: B (6), A(10), and AB (16) |
As shown in table 1, the SIFT and PolyPhen software made predictions for 39258 nsSNPs (13003 genes), which were exclusively appeared in wild-type (WT) soybeans; Of these, 13% (5175 nsSNPs) were predicted by SIFT and PolyPhen as having a potentially damaging effect on protein function. Then out of 5175, 901 nsSNPs were finally chosen with first priority by narrowing down the number of nsSNPs, only empirically considering the nsSNPs appeared in more than half of total soybean species. Similarly, the same two software made predictions for 4302 nsSNPs (2118 genes), which were exclusively appeared in cultured soybeans; Of them, 27% (1146 nsSNPs) were predicted by SIFT and PolyPhen as having a potentially damaging effect on protein function. Then out of 1146, 77 nsSNPs were finally chosen with first priority by considering the nsSNPs appeared in more than half of total soybean species. Similarly, in a result of the simulation of nsSNPs which appeared in both cultured and wild soybean genes, out of 20361, 2318 nsSNPs were included in an initial list which we are interested in.
These selected nsSNPs in these three different cases were again prioritized by the software PANTHER (more than 0.98) and I-mutant (less than 0 when pH=7 and T= 298K), eventually leaving a small number of most likely and tractable candidates. Finally, the short list was sorted again following the descent order of PANTHER. In addition, for knowing the gene flow strength of nsSNPs investigated, FST [Wright’s fixation index (1951), an estimator of the amount of gene flow between species] value of each gene was added. The FST was calculated like this:
FST in a gene = The total sum of the number of nsSNPs from two genes [one gene in 6 wild soybeans and the other in 11 cultivated soybeans (including W 82) / The total sum of the number of nsSNPs from two genes from 17 wild and cultivated soybeans.
Finally, the selected numbers of biomarker candidates are 1, 15, and 75 from A, B, and AB, respectively, totaling 81. In the consecutive study, we will investigate these 81 biomarker candidates by the approaches of more advanced bioinformatics tools including 3-D structural analysis and molecular dynamics (MD) simulation.
CONCLUSION
We suggest a short list of biomarker candidates as a result of this computational soybean study. These potentially functional nsSNPs and their effects on protein were predicted by four computational software of PolyPhen, SIFT, I-mutant 2.0, and PANTHER. They were roughly selected and most of them may not be realistic ones. However, we expect that some of the list have a higher probability to be recognized as agronomically important biomarkers in the future. Also we can search two public databases (http://www.soybase.org and http://www.ncbi.nlm.nih.gov) to find out whether some nsSNPs in the list were already reported as biomarkers or not.
We hope that this biomarker list will support the future studies: (1) to develop some methodologies to check their structural changes, (2) to extract some biologically important nsSNPs with phenotypic traits that are favorably selected by humans to enhance agronomical traits; and (3) to track their domestication process in a view of population genetics.
DISCUSSION
Because resequencing data are continuously updating, some of nsSNPs included in this list of biomarker candidates needs to be adjusted as time goes by. For reference, we are happy to provide our whole biomarker candidate data and source codes (explaining how we handled multi-data) when requested.
ACKNOWLEDGMENTS
The authors are grateful to the bio-evaluation center at Korea Research Institute of Bioscience and Biotechnology for providing the nsSNP data of 16 Korean soybeans used in this study.
I will submit this paper to a highly reputational Journal. So it needs high quality of English with concise and readable creativity.