教学文库网 - 权威文档分享云平台
您的当前位置:首页 > 范文大全 > 行业范文 >

Andr'e Seznec y St'ephan Jourdan z Pascal Sa

来源:网络收集 时间:2026-08-14
导读: A basic rule in computer architecture is that a processor cannot execute an application faster than it fetches its instructions. This paper presents a novel costeffective mechanism called the two-block ahead branch predictor. Information f

A basic rule in computer architecture is that a processor cannot execute an application faster than it fetches its instructions. This paper presents a novel costeffective mechanism called the two-block ahead branch predictor. Information from the current i

Multiple-Block Ahead Branch Predictors to appear in Proceedings of ASPLOS VII, Boston, October 1996Andre Seznec y Stephan Jourdan z Pascal Sainrat z Pierre Michaud y IRISA y Campus de Beaulieu 35042 Rennes, France fseznec,pmichaudg@irisa.fr IRIT z Universite Paul Sabatier 31062 Toulouse, France fjourdan,sainratg@irit.fr diction mechanism allowing to increase the instruction fetch rate for both approaches.\Brainiac" processors To best exploit the available ILP,\brainiac" processors are using a large number of functional units working in parallel. Unfortunately, the instruction-fetch mechanisms implemented in current commercial microprocessors do not fully exploit the potential parallelism. For these processors, the instructions fetched in a single cycle most often belong to the same basic block, and are not usually permitted to span two cache lines. Since a processor cannot execute instructions faster than it fetches them, these constraints signi cantly impair performance, particularly on codes featuring many small basic blocks. A partial solution to the instruction fetch bottleneck, is to fetch instructions belonging to multiple consecutive basic blocks, as is done in processors such as the POWER2 18]. To solve the whole problem, multiple non-consecutive basic blocks must be fetched in a single cycle as most basic blocks are only ve instructions long. Indeed, the potential parallelism has been shown to be higher than six instructions per cycle in generalpurpose integer applications while assuming a perfect instruction-fetch mechanism 11]. A processor featuring such a mechanism would have to predict multiple targets and branch outcomes in a single cycle. In superscalar processors, blocks of consecutive instructions are fetched in parallel. The last instruction of such a block is either a branch or is determined by some implementation constraints (for instance, the boundary of a cache block or the maximum number of instructions in the block). Throughout this paper, we refer to processors that can fetch only one basic block per cycle as single I-fetch processors, to processors that can fetch two non-consecutive blocks per cycle as double Ifetch processors, and to multiple I-fetch processors as an extension to the latter case. We show in this paper that double I-fetch processors will have a major performance advantage over single Ifetch processors for a dispatch width of six or higher. Our belief is that future generations of\brainiac" processors will be multiple I-fetch processors.\Speed-demon" processors To achieve high performance,\speed demon" processors rely on moderate numbers of functional units, but a very high clock rate. In such processors, a single instruction block is dispatched in each cycle, but the branch predictor is often a critical path in the processor. In current microproces-

A basic rule in computer architecture is that a processor cannot execute an application faster than it fetches its instructions. This paper presents a

novel coste ective mechanism called the two-block ahead branch predictor. Information from the current instruction block is not used for predicting the address of the next instruction block, but rather for predicting the block following the next instruction block. This approach overcomes the instruction fetch bottleneck exhibited by wide-dispatch\brainiac" processors by enabling them to e ciently predict addresses of two instruction blocks in a single cycle. Furthermore, pipelining the branch prediction process can also be done by means of our predictor for\speed demon" processors to achieve higher clock rate or to improve the prediction accuracy by means of bigger prediction structures. Moreover, and unlike the previously-proposed multiple predictor schemes, multiple-block ahead branch predictors can use any of the branch prediction schemes to perform the very accurate predictions required to achieve high-performance on superscalar processors.

Abstract

1 Introduction

Two di erent approaches are used in current processors to achieve high performance:\brainiacs vs. speed demons" 6]. While\brainiacs" favor the parallel execution of instructions and\speed-demons" favor a high clock rate, both approaches are facing a similar di culty with fetching instructions at a su cient rate. The purpose of this paper is to propose a new branch preThis work was partially supported by PRC-GDR AMN (CNRS)

c 1996 by the Association for Computing Machinery, Inc. Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for pro t or commercial advantage and that new copies bear this notice and the full citation on the rst page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, to republish, to post on servers or to redistribute to lists, requires prior speci c permission and/or a fee. Request Permissions from Publications Dept, ACM Inc., Fax+1 (212) 869-0481, or permissions@http://doc.guandang.net.

A basic rule in computer architecture is that a processor cannot execute an application faster than it fetches its instructions. This paper presents a novel costeffective mechanism called the two-block ahead branch predictor. Information from the current i

sors, either the branch prediction and the address generation are completed in a single cycle, or pipeline bubbles are inserted on each predicted taken branch (e.g. on DEC 21164 5], PentiumPro 8] or MIPS R10000 13]), therefore potentially limiting the perfo …… 此处隐藏:33630字,全部文档内容请下载后查看。喜欢就下载吧 ……

Andr'e Seznec y St'ephan Jourdan z Pascal Sa.doc 将本文的Word文档下载到电脑,方便复制、编辑、收藏和打印
本文链接:https://www.jiaowen.net/fanwen/983893.html(转载请注明文章来源)
Copyright © 2020-2025 教文网 版权所有
声明 :本网站尊重并保护知识产权,根据《信息网络传播权保护条例》,如果我们转载的作品侵犯了您的权利,请在一个月内通知我们,我们会及时删除。
客服QQ:78024566 邮箱:78024566@qq.com
苏ICP备19068818号-2
Top
× 游客快捷下载通道(下载后可以自由复制和排版)
VIP包月下载
特价:29 元/月 原价:99元
低至 0.3 元/份 每月下载150
全站内容免费自由复制
VIP包月下载
特价:29 元/月 原价:99元
低至 0.3 元/份 每月下载150
全站内容免费自由复制
注:下载文档有可能出现无法下载或内容有问题,请联系客服协助您处理。
× 常见问题(客服时间:周一到周五 9:30-18:00)