A Survey and Classification of Storage Deduplication Systems

João Tiago Paulo; José Orlando Pereira

A Survey and Classification of Storage Deduplication Systems

Files

P-009-SED.pdf (555.8 KB)

Date

2014

Authors

João Tiago Paulo

José Orlando Pereira

Abstract

The automatic elimination of duplicate data in a storage system, commonly known as deduplication, is increasingly accepted as an effective technique to reduce storage costs. Thus, it has been applied to different storage types, including archives and backups, primary storage, within solid-state drives, and even to random access memory. Although the general approach to deduplication is shared by all storage types, each poses specific challenges and leads to different trade-offs and solutions. This diversity is often misunderstood, thus underestimating the relevance of new research and development. The first contribution of this article is a classification of deduplication systems according to six criteria that correspond to key design decisions: granularity, locality, timing, indexing, technique, and scope. This classification identifies and describes the different approaches used for each of them. As a second contribution, we describe which combinations of these design decisions have been proposed and found more useful for challenges in each storage type. Finally, outstanding research challenges and unexplored design points are identified and discussed.

URI

http://repositorio.inesctec.pt/handle/123456789/4154
http://dx.doi.org/10.1145/2611778

Collections

HASLab - Indexed Articles in Journals

Full item page