Increasing Splicing Site Prediction by Training Gene Set Based on Species

Ahn, Beunguk;Abbas, Elbashir;Park, Jin-Ah;Choi, Ho-Jin;

doi:10.3837/tiis.2012.10.002

KSII Transactions on Internet and Information Systems (TIIS)

제6권11호
/
Pages.2784-2799
/
2012
/
1976-7277(pISSN)
/
1976-7277(eISSN)

한국인터넷정보학회 (Korean Society for Internet Information)

DOI QR Code

Increasing Splicing Site Prediction by Training Gene Set Based on Species

Ahn, Beunguk (Department of Computer Science, Korea Advanced Institute of Science and Technology (KAIST)) ;
Abbas, Elbashir (Department of Computer Science, Korea Advanced Institute of Science and Technology (KAIST)) ;
Park, Jin-Ah (Department of Computer Science, Korea Advanced Institute of Science and Technology (KAIST)) ;
Choi, Ho-Jin (Department of Computer Science, Korea Advanced Institute of Science and Technology (KAIST))

투고 : 2012.04.07
심사 : 2012.08.18
발행 : 2012.11.30

https://doi.org/10.3837/tiis.2012.10.002 인용 PDF KSCI

PDF 다운로드

⟨ 이전 논문 다음 논문 ⟩

초록

Biological data have been increased exponentially in recent years, and analyzing these data using data mining tools has become one of the major issues in the bioinformatics research community. This paper focuses on the protein construction process in higher organisms where the deoxyribonucleic acid, or DNA, sequence is filtered. In the process, "unmeaningful" DNA sub-sequences (called introns) are removed, and their meaningful counterparts (called exons) are retained. Accurate recognition of the boundaries between these two classes of sub-sequences, however, is known to be a difficult problem. Conventional approaches for recognizing these boundaries have sought for solely enhancing machine learning techniques, while inherent nature of the data themselves has been overlooked. In this paper we present an approach which makes use of the data attributes inherent to species in order to increase the accuracy of the boundary recognition. For experimentation, we have taken the data sets for four different species from the University of California Santa Cruz (UCSC) data repository, divided the data sets based on the species types, then trained a preprocessed version of the data sets on neural network(NN)-based and support vector machine(SVM)-based classifiers. As a result, we have observed that each species has its own specific features related to the splice sites, and that it implies there are related distances among species. To conclude, dividing the training data set based on species would increase the accuracy of predicting splicing junction and propose new insight to the biological research.

키워드

피인용 문헌

Combining Support Vector Machine Recursive Feature Elimination and Intensity-dependent Normalization for Gene Selection in RNAseq vol.18, pp.5, 2012, https://doi.org/10.7472/jksii.2017.18.5.47
A New Rapid Incremental Algorithm for Constructing Concept Lattices vol.10, pp.2, 2012, https://doi.org/10.3390/info10020078
Integration of geographic and hierarchical routing protocols for energy saving in wireless sensor networks with mobile sink vol.25, pp.5, 2012, https://doi.org/10.1007/s11276-019-02015-5

KSII Transactions on Internet and Information Systems (TIIS)

Increasing Splicing Site Prediction by Training Gene Set Based on Species

초록

키워드

피인용 문헌

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)