A performance counter architecture for computing accurate CPI components

20 October 2006

journal article
Published by Association for Computing Machinery (ACM) in ACM SIGOPS Operating Systems Review

Vol. 40 (5), 175-184
https://doi.org/10.1145/1168917.1168880

Abstract

A common way of representing processor performance is to use Cycles per Instruction (CPI) `stacks' which break performance into a baseline CPI plus a number of individual miss event CPI components. CPI stacks can be very helpful in gaining insight into the behavior of an application on a given microprocessor; consequently, they are widely used by software application developers and computer architects. However, computing CPI stacks on superscalar out-of-order processors is challenging because of various overlaps among execution and miss events (cache misses, TLB misses, and branch mispredictions).This paper shows that meaningful and accurate CPI stacks can be computed for superscalar out-of-order processors. Using interval analysis, a novel method for analyzing out-of-order processor performance, we gain understanding into the performance impact of the various miss events. Based on this understanding, we propose a novel way of architecting hardware performance counters for building accurate CPI stacks. The additional hardware for implementing these counters is limited and comparable to existing hardware performance counter architectures while being significantly more accurate than previous approaches.

Keywords

This publication has 9 references indexed in Scilit:

Characterizing the branch misprediction penalty
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2006
Interaction cost and shotgun profiling
ACM Transactions on Architecture and Code Optimization, 2004
Benchmarking internet servers on superscalar machines
Computer, 2003
Pentium 4 performance-monitoring features
IEEE Micro, 2002
Performance of database workloads on shared-memory systems with out-of-order processors
Published by Association for Computing Machinery (ACM) ,1998
Continuous profiling
ACM Transactions on Computer Systems, 1997
Performance analysis using the MIPS R10000 performance counters
Published by Association for Computing Machinery (ACM) ,1996
Theoretical modeling of superscalar processor performance
Published by Association for Computing Machinery (ACM) ,1994
The Inhibition of Potential Parallelism by Conditional Jumps
International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 1972

Cited by 13 articles