A case for exploiting subarray-level parallelism (SALP) in DRAM

Top Cited Papers

1 June 2012

conference paper
conference paper
Published by Institute of Electrical and Electronics Engineers (IEEE)

p. 368-379
https://doi.org/10.1109/isca.2012.6237032

Abstract

Modern DRAMs have multiple banks to serve multiple memory requests in parallel. However, when two requests go to the same bank, they have to be served serially, exacerbating the high latency of off-chip memory. Adding more banks to the system to mitigate this problem incurs high system cost. Our goal in this work is to achieve the benefits of increasing the number of banks with a low cost approach. To this end, we propose three new mechanisms that overlap the latencies of different requests that go to the same bank. The key observation exploited by our mechanisms is that a modern DRAM bank is implemented as a collection of subarrays that operate largely independently while sharing few global peripheral structures. Our proposed mechanisms (SALP-1, SALP-2, and MASA) mitigate the negative impact of bank serialization by overlapping different components of the bank access latencies of multiple requests that go to different subarrays within the same bank. SALP-1 requires no changes to the existing DRAM structure and only needs reinterpretation of some DRAM timing parameters. SALP-2 and MASA require only modest changes (<;0.15% area overhead) to the DRAM peripheral structures, which are much less design constrained than the DRAM core. Evaluations show that all our schemes significantly improve performance for both single-core systems and multi-core systems. Our schemes also interact positively with application-aware memory request scheduling in multi-core systems.

Keywords

This publication has 37 references indexed in Scilit:

Reducing memory interference in multicore systems via application-aware memory channel partitioning
Published by Association for Computing Machinery (ACM) ,2011
Thread Cluster Memory Scheduling: Exploiting Differences in Memory Access Behavior
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2010
Complexity effective memory access scheduling for many-core accelerator architectures
Published by Association for Computing Machinery (ACM) ,2009
75nm 7Gb/s/pin 1Gb GDDR5 graphics memory device with bandwidth-improvement techniques
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2009
Parallelism-Aware Batch Scheduling: Enhancing both Performance and Fairness of Shared DRAM Systems
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2008
Stall-Time Fair Memory Access Scheduling for Chip Multiprocessors
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2007
Fair Queuing Memory Systems
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2006
Runahead execution: an alternative to very large instruction windows for out-of-order processors
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2003
The hierarchical multi-bank DRAM: a high-performance architecture for memory integrated with processors
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2002
Cached DRAM for ILP processor memory access latency reduction
IEEE Micro, 2001

Cited by 177 articles