doi: 10.7763/IJIMT.2012.V3.239
Comparison of Protein Corpuses
- 1Department of Electrical and Computer Engineering, Addis Ababa Institute of Technology, Addis Ababa University, Addis Ababa, Ethiopia.
- 2Department of Computer Science and Engineering, Karunya University, Coimbatore, India.
Abstract
This paper presents a comparison of two protein corpuses. The protein corpus is a data set of four files for evaluating the performance of protein compression algorithms. Although past studies reported compression rates of protein sequences with varying degrees of success, there are wrongly stated claims and confusing results in some standard publications arising from inappropriate comparison of the data sets. To emphasize the difference and similarity of the data sets, the content of the files in the two protein corpuses are compared with respect to the size in bytes and repetitions of amino acids. In addition, comparison is made based on difficulty of compressing the files in the corpuses. The results indicate that the two protein corpuses possess different regularities. Besides, nine general purpose compression algorithms outperform the results reported by biological compressors on one of the corpus and comparable results on the other corpus.
Keywords
- Protein corpus
- compression rate
- protein compression
- biological compressors
- general purpose compressor
How to Cite
Wedajo Diribi and Kumudha Raimond, "Comparison of Protein Corpuses," International Journal of Innovation, Management and Technology, vol. 3, no. 3, pp. 281-284, 2012. https://doi.org/10.7763/IJIMT.2012.V3.239
Copyright & License
Copyright © 2012 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).