UW-Madison Logo

The ADvanced Systems Laboratory (ADSL)
Publication abstract

A Five-Year Study of File-System Metadata

Nitin Agrawal, William J. Bolosky*, John R. Douceur*, Jacob R. Lorch*
Department of Computer Sciences , University of Wisconsin-Madison
* Microsoft Research

Abstract:

For five years, we collected annual snapshots of file-system metadata from over 60,000 Windows PC file systems in a large corporation. In this paper, we use these snapshots to study temporal changes in file size, file age, file-type frequency, directory size, namespace structure, file-system population, storage capacity and consumption, and degree of file modification. We present a generative model that explains the namespace structure and the distribution of directory sizes. We find significant temporal trends relating to the popularity of certain file types, the origin of file content, the way the namespace is used, and the degree of variation among file systems, as well as more pedestrian changes in sizes and capacities. We give examples of consequent lessons for designers of file systems and related software.

Full Paper: Postscript   PDF   BibTeX

Publications