Conference Article

·2017

Multi-stream word-based compression algorithm

Emır Öztürk , Altan Mesut , Banu Di̇ri̇ YTU

2017 International Conference on Computer Science and Engineering (UBMK)

Abstract

In this article, we present a novel word-based lossless compression algorithm for text files which uses a semi-static model. We named our algorithm as Multi-stream Word-based Compression Algorithm (MWCA), because it stores the compressed forms of the words in three individual streams depending on their frequencies in the text. It also stores two dictionaries and a bit vector as a side information. In our experiments MWCA obtains compression ratio over 3,23 bpc on average and 2,88 bpc on files larger than 50 MB. If a variable length encoder like Huffman Coding is used after MWCA, given ratios will reduce to 2,63 and 2,44 bpc respectively. With the advantage of its multi-stream structure MWCA could become a good solution especially for storing and searching big text data.

Keywords

Huffman coding Lossless compression Computer science Word (group theory) Algorithm Compression ratio Encoder Data compression Compression (physics) Data stream Arithmetic coding Data compression ratio Data stream mining Artificial intelligence Image compression Context-adaptive binary arithmetic coding Data mining Mathematics

Subject Areas

Algorithms and Data Compression ·Artificial Intelligence ·Physical Sciences
Advanced Data Compression Techniques ·Computer Vision and Pattern Recognition ·Physical Sciences
Parallel Computing and Optimization Techniques ·Hardware and Architecture ·Physical Sciences

Citations by Year

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Quality Education 70%