教学文库网 - 权威文档分享云平台
您的当前位置:首页 > 文库大全 > 专业资料 >

Abstract Attribute-Based Prediction of File Properties(2)

来源:网络收集 时间:2026-09-04
导读: In the most extreme case,a system like AutoRAID[31] employs several different methods to store blocks with different characteristics.On a more mundane level,the performance of nearly all modern disk

In the most extreme case,a system like AutoRAID[31] employs several different methods to store blocks with different characteristics.On a more mundane level,the performance of nearly all modern disk drives is highly in?uenced by the multi-zone effect,which can cause the effective transfer rate for the outer tracks of a disk to be considerably higher than that of the inner[19].There is ample evidence that adaptive block layout can im-prove performance;we will demonstrate that we can pre-emptively determine the layout heuristics to achieve this bene?t without having to reorganize?les after their ini-tial placement.

Advances in arti?cial intelligence and machine learn-ing have resulted in ef?cient algorithms for building ac-curate predictive models that can be used in today’s?le 2

We present evidence that attributes that are known to the file system when a file is created, such as its name, permission mode, and owner, are often strongly related to future properties of the file such as its ultimate size, lifespan, and access pattern.

systems.We leverage this work and utilize a form of classi?cation tree to capture the relationships between ?le attributes and their behaviors,as further described in Section5.

The work we present here does not focus on new heuristics or policies for optimizing the?le system.In-stead it enables a?le system to choose the proper policies to apply by predicting whether or not the assumptions on which these policies rely will hold for a particular?le. 3The Traces

To demonstrate that our?ndings are not con?ned to a single workload,system,or set of users,we analyze traces taken from three servers:

DEAS03traces a Network Appliance Filer that serves the home directories for professors,graduate stu-dents,and staff of the Harvard University Divi-sion of Engineering and Applied Sciences.This trace captures a mix of research and development, administrative,and email traf?c.The DEAS03 trace begins at midnight on2/17/2003and ends on 3/2/2003.

EECS03traces a Network Appliance Filer that serves the home directories for some of the professors, graduate students,and staff of the Electrical Engi-neering and Computer Science department of the Harvard University Division of Engineering and Applied Sciences.This trace captures the canonical engineering workstation workload.The EECS03 trace begins at midnight on2/17/2003and ends on 3/2/2003.

CAMPUS traces one of14?le systems that hold home directories for the Harvard College and Harvard Graduate School of Arts and Sciences(GSAS)stu-dents and staff.The CAMPUS workload is almost entirely email.The CAMPUS trace begins at mid-night10/15/2001and ends on10/28/2003.

Ideally our analyses would include NFS traces from a variety of workloads including commercial datacenter servers,but despite our diligent efforts we have not been able to acquire any such traces.

The DEAS03and EECS03traces are taken from the same systems as the DEAS and EECS traces described in earlier work[9],but are more recent and contain infor-mation not available in the earlier traces.The CAMPUS trace is the same trace described in detail in an earlier

study[8],although we draw our samples from a longer subset of the trace.All three traces were collected with nfsdump[10].

Table1gives a summary of the average hourly oper-ation counts and mixes for the workloads captured in the traces.These show that there are differences between these workloads,at least in terms of the operation mix.

CAMPUS is dominated by reads and more than85%of the operations are either reads or writes.DEAS03has proportionally fewer reads and writes and more meta-data requests(getattr,lookup,and access)than CAMPUS,but reads are still the most common opera-tion.On EECS03,meta-data operations comprise the majority of the workload.

Earlier trace studies have shown that hourly opera-tion counts are correlated with the time of day and day of week,and much of the variance in hourly operation count is eliminated by using only the working hours[8].Table 1shows that this trend appears in our data as well.Since the“work-week”hours(9am-6pm,Monday through Fri-day)are both the busiest and most stable subset of the data,we focus on these hours for many of our analyses.

One aspect of these traces that has an impact on our research is that they have been anonymized,using the method described in earlier work[8].During the anonymization UIDs,GIDs,and host IP numbers are simply remapped to new values,so no information is lost about the relationship between these identi?ers and other variables in the data.The anonymization method also preserves some types of information about?le and direc-tory names–for example,if two names share the same suf?x,then the anonymized forms of these names will also share the same suf?x.Unfortunately,some informa-tion about?le names is lost.A survey of the?le names in our own directories leads us to believe that capital-ization,use of whitespace,and some forms of punctua-tion in?le names may be useful attributes of?le names, but none of this information survives anonymization.As we will show in the remaining sections of this paper,the anonymized names provide enough information to build good models,but we believe that it may be possible to build even more accurate models from unanonymized data.

4The Case for Attribute-Based Predic-tions

To explore the associations between the create-time attributes of a?le and its longer-term properties,we be-gin by scanning our traces to extract both the initial at-3

We present evidence that attributes that are known to the file system when a file is created, such as its name, permission mode, and owner, are often strongly related to future properties of the file such as its ultimate size, lifespan, and access pattern.

All Hours

DEAS0315.7%(55.3%)29.2%(49.3%)

EECS0312.3%(123.8%) 3.2%(263.2%)

CAMPUS21 …… 此处隐藏:6239字,全部文档内容请下载后查看。喜欢就下载吧 ……

Abstract Attribute-Based Prediction of File Properties(2).doc 将本文的Word文档下载到电脑,方便复制、编辑、收藏和打印
本文链接:https://www.jiaowen.net/wenku/266432.html(转载请注明文章来源)
Copyright © 2020-2025 教文网 版权所有
声明 :本网站尊重并保护知识产权,根据《信息网络传播权保护条例》,如果我们转载的作品侵犯了您的权利,请在一个月内通知我们,我们会及时删除。
客服QQ:78024566 邮箱:78024566@qq.com
苏ICP备19068818号-2
Top
× 游客快捷下载通道(下载后可以自由复制和排版)
VIP包月下载
特价:29 元/月 原价:99元
低至 0.3 元/份 每月下载150
全站内容免费自由复制
VIP包月下载
特价:29 元/月 原价:99元
低至 0.3 元/份 每月下载150
全站内容免费自由复制
注:下载文档有可能出现无法下载或内容有问题,请联系客服协助您处理。
× 常见问题(客服时间:周一到周五 9:30-18:00)