LEC-Codec: Learning-Based Genome Data Compression.

Journal: IEEE/ACM transactions on computational biology and bioinformatics
PMID:

Abstract

In this paper, we propose a Learning-based gEnome Codec (LEC), which is designed for high efficiency and enhanced flexibility. The LEC integrates several advanced technologies, including Group of Bases (GoB) compression, multi-stride coding and bidirectional prediction, all of which are aimed at optimizing the balance between coding complexity and performance in lossless compression. The model applied in our proposed codec is data-driven, based on deep neural networks to infer probabilities for each symbol, enabling fully parallel encoding and decoding with configured complexity for diverse applications. Based upon a set of configurations on compression ratios and inference speed, experimental results show that the proposed method is very efficient in terms of compression performance and provides improved flexibility in real-world applications.

Authors

  • Zhenhao Sun
  • Meng Wang
    State Key Laboratory of Urban Water Resource and Environment, School of Environment, Harbin Institute of Technology, Harbin 150001, China.
  • Shiqi Wang
    Xijing Hospital, Fourth Military Medical University, Xi'an, China.
  • Sam Kwong