CODA: How to Mitigate ColumnDisturb
for (Almost) Free?


Figure 1: (a) ColumnDisturb leverages shared bitlines to cause distant bitflips, both inter-subarray and intra-subarray. (b) Intra-subarray bitflips can be tolerated by SALT as it operates at subarray granularity. (c) SALT with ColumnDisturb Protection (CDP) can tolerate inter-subarray bitflips. On activation, CDP not only increments the ACTR of the demand subarray but also of the two adjacent subarrays. Such Adjacent-Counter Increment (ACI) increases the need for mitigation and incurs significant overheads (d). Our design, CODA, reduces ACI by 12x-1300x, thereby reducing the overheads of tolerating ColumnDisturb.

1 Introduction↩︎

Data-Disturbance Errors (DDEs) violate the isolation property of memory systems by allowing access to one memory location to modify the data stored in another. DDEs are not only a reliability challenge but also a significant security threat, as an attacker can use DDEs to flip bits in critical data structures, such as the Page Tables, to escalate privileges. As DRAMs scale to smaller sizes, we continue to encounter new forms of DDEs.

The most well-known DDE vulnerability is Rowhammer [1], in which frequent accesses to an aggressor row cause bitflips in nearby victim rows. The number of activations required to cause a bitflip is called the Rowhammer Threshold (TRH). Typical mitigations for Rowhammer are based on tracking aggressor rows and refreshing a small number of victim rows on either side of the aggressor row. RowPress [2] is another recent DDE, in which the aggressor row is kept open for a long time, thereby causing charge leakage in proportion to the row-open time. As RowPress reduces the number of activations required to induce a bitflip in the nearby victim rows, it reduces the effective TRH. Fortunately, RowPress can be easily tolerated with existing hardware-based Rowhammer solutions by converting the row open-time into equivalent activations [3].

A recent paper introduced ColumnDisturb [4], a new class of disturbance errors, whereby activations in aggressor row can cause bitflips in distant victim rows, which are located hundreds of rows away from the aggressor row within the same subarray (we call this intra-subarray ColumnDisturb) or in victim-rows located in adjacent subarrays (we call this inter-subarray ColumnDisturb). ColumnDisturb exploits the fact that in open-bitline architectures, adjacent subarrays share a sense amplifier and that bitlines run across the adjacent subarrays, as shown in Figure 1 (a). This sharing of bitlines breaks spatial isolation between subarrays, allowing an activation in one subarray to affect the data stored in another. The pattern of ColumnDisturb is similar to Rowhammer and RowPress, just that the time duration for ColumnDisturb is in the range of several milliseconds. Existing Rowhammer solutions that perform row-granularity tracking and mitigations cannot tolerate ColumnDisturb as they refresh only a small number of rows (typically 1-2) on either side of the aggressor row.

ColumnDisturb can be mitigated by Rowhammer mitigations that operate at a subarray granularity, such as SALT [5]. As shown in Figure 1 (b), SALT provisions each subarray with an Activation Counter (ACTR), which is incremented on each activation to the subarray. When the activation counter reaches a specific value, SALT performs a sequential refresh of several rows within the subarray. The mitigation rate of SALT is designed to ensure that SALT refreshes all rows within the subarray before the subarray receives TRH activations. As SALT provides Blast-Radius freedom (the victim row can be anywhere within the same subarray as the aggressor), SALT implicitly handles intra-subarray ColumnDisturb. However, SALT cannot handle inter-subarray ColumnDisturb.

To tolerate inter-subarray ColumnDisturb, SALT can be provisioned with ColumnDisturb Protection (CDP) [5], whereby each demand activation to a subarray also performs additional Adjacent-Counter Increments (ACI) for the two adjacent subarrays, as shown in Figure 1 (c). ACIs enable adjacent subarrays to perform mitigation even if they receive no demand activations, thereby providing protection against inter-subarray ColumnDisturb. As CDP issues two ACIs per demand activation, it significantly increases the mitigation requirements of SALT (equivalent to a bank incurring 3x activations), resulting in large performance and power overheads (e.g., CDP increases SALT’s slowdown from 0.3% to 17% at TRHD of 500). The goal of our paper is to mitigate ColumnDisturb while incurring negligible performance and power overheads.

The key insight of our paper is to reduce the overhead of ColumnDisturb protection by reducing the number of ACIs issued per demand activations. We propose CODA (ColumnDisturb Mitigation with Reduced ACIs), as shown in Figure 1 (d). We present three variants of CODA, each targeting a different inefficiency in ACI.

Our first design, CODA-E (Evade), avoids the ACI to the neighboring subarray if the neighboring subarray receives a demand activation. The key observation of CODA-E is that if a neighbor receives sufficient demand activations, it will undergo mitigations anyway due to the demand activity and does not require ACI for mitigations. We show, both theoretically and experimentally, that CODA-E can halve the rate of ACIs, reducing it from 200% to 100%.

Our second design, CODA-F (Fraction), exploits the fact that ColumnDisturb is a much slower attack compared to Rowhammer and lasts for several milliseconds. For example, the ColumnDisturb paper targets solutions for a ColumnDisturb attack duration of 8 ms. As Rowhammer solutions are designed to tolerate much faster attacks, we should lower the rate of ACI so that the neighboring subarray can fully refresh its rows within our target ColumnDisturb duration. For example, if we assume that a row-open time of 500 ns is equivalent to 1 activation, then the attacker must inflict 16K equivalent activations within 8 ms to cause ColumnDisturb. If the Rowhammer solution is designed to refresh all rows in the subarray within 1K activations, then it would be sufficient to issue one ACI every 32 demand activations, rather than 1 per demand activation. CODA-F reduces the ACI rates by 16x-2x for TRH of 500 to 4K, respectively. Thus, CODA-F is especially attractive for lower TRH. CODA-E and CODA-F are synergistic and can be combined to substantially reduce the rate of ACI from 200% for CDP to 1.4% (TRH of 500) or 17% (TRH of 4K).

Our third design, CODA-G (Gangskip), is designed for solutions that operate at multi-subarray granularity. For example, Ganged-SALT [5] is a variant of SALT that operates at the granularity of 2 (at TRH of 1K) to 8 (at TRH of 4K) subarrays. As the gang shares the same activation counter and metadata, Ganged-SALT reduces the SRAM storage overhead by 2x-8x. The key observation in CODA-G is to skip the ACI for the adjacent subarray when that subarray is within the same gang, since gang-based mitigation will also refresh all rows in the adjacent subarray. CODA-G reduces the rate of ACI by 2x (TRH of 1K) to 8x (TRH of 4K) and becomes even more attractive at higher TRH. CODA-G can be combined with CODA-E and CODA-F to reduce the rate of ACI from 200% (for CDP) to 0.23% (at TRH of 1K) or 0.16% (at TRH of 4K).

As CODA virtually eliminates ACIs, CODA makes it practical to tolerate ColumnDisturb with the same solution as Rowhammer, while incurring negligible performance and power overheads (0.01%). The storage overhead of CODA is small (4 bits per subarray).

The principles of CODA are not specific to SALT and are applicable to any Rowhammer mitigation that operates at a subarray granularity. For example, we also analyze ColumnDisturb mitigation with REGA [6]. REGA tolerates Rowhammer by provisioning the subarray with an extra row buffer and generating a refresh of one or more rows within a subarray on each demand activation. REGA can protect intra-subarray ColumnDisturb; however, to protect inter-subarray ColumnDisturb, it will also need ACI for neighboring subarrays. A design based on CDP (REGA-CDP) will issue two ACI for each demand activation. REGA-CDP incurs 3x the refresh-power overheads as REGA for mitigation. When REGA is implemented with CODA (REGA-CODA), the ACIs are reduced by 12x-140x, thereby virtually eliminating the increase in refresh-power overhead required to tolerate ColumnDisturb.

Our paper makes the following contributions:

  1. We show that the main overhead of tolerating ColumnDisturb is due to Adjacent-Counter Increment (ACI), where each activation increments the counter for three subarrays.

  2. We propose CODA, which significantly reduces the rate of ACI by leveraging demand-activations to skip ACI (CODA-E) and performing fractional increments for ACI (CODA-F).

  3. We propose CODA-G, which is tailored to mitigations that operate at multi-subarray granularity and skips ACI when the adjacent subarray is within the same gang.

As CODA reduces the rate of ACI by 12x-1300x, thereby making it practical to tolerate ColumnDisturb at nearly zero overhead.

2 Background and Motivation↩︎

2.1 Threat Model↩︎

Our threat model assumes that an attacker can issue memory requests for arbitrary addresses. The attacker knows the defense algorithm and any relevant parameters. We declare an attack successful when a row incurs sufficient charge loss to trigger a bit flip. We assume that the attacker uses the most effective pattern (including attacks that spread activations across different subarrays) based on the defenses employed. We assume that the inter-subarray bitflips from ColumnDisturb are restricted to only the adjacent subarrays. We further assume that any row within a subarray can be used to induce ColumnDisturb in adjacent subarrays, allowing the attacker to spread activations across many rows within a subarray.

2.2 DRAM Architecture and Parameters.↩︎

DRAM chips are organized into arrays comprising rows and columns. To access the data in DRAM, the row is accessed using the word-line, and the charge on the DRAM cells is sensed using a sense amplifier, and the data is stored in a Row-Buffer (RB). The DRAM chip is divided into a number of banks (32 for DDR5), with only one row-buffer per bank architecturally visible to the Memory Controller. However, internally, each bank consists of smaller independent units called the Subarray [4][10], each equipped with a row-buffer and sensing circuit. We assume an open bitline architecture where adjacent subarrays share a row-buffer, and the bitlines are spread over multiple subarrays, as shown in Figure 2. Each subarray typically contains 512 rows. In our system, each DRAM bank contains 128K rows, so each bank has 256 subarrays.

Figure 2: DRAM Architecture. A bank consists of subarrays. Open Bitline Architectures shares row-buffer between adjacent subarrays, and bitlines run across multiple subarrays.

To access data from DRAM, the memory controller must first issue an activation (ACT) to open the row. To access data from another conflicting row, the opened row must be precharged (PRE) first. To ensure data retention, all data in the DRAM must be refreshed within the retention period (tREFW of 32ms). To reduce the latency of refresh, memory is divided into 8192 groups, and a REF command, issued every tREFI (3.9 microseconds), refreshes one group. As our bank contains 128K rows, each REF must refresh 16 rows in the bank. We assume that consecutive rows in a bank are spread to consecutive subarrays, so that accesses to consecutive rows are not focused within the same subarray [7], [9].

2.3 Problem of Data-Disturbance Errors↩︎

As DRAM cells scale down, the inter-cell distance decreases, and activity in one cell starts to affect the data stored in another cell. This is called a Data-Disturbance Error (DDE). The most well-known DDE is Rowhammer [1], which was discovered more than a decade ago. Rowhammer occurs when an aggressor row is repeatedly activated, causing a small amount of charge loss in the cells of the neighboring row. When this charge loss exceeds a given threshold, it induces bit-flips in neighboring rows. The minimum number of activations to an aggressor row to cause a bit-flip in a victim row is called the Rowhammer Threshold (TRH). TRH is reported either for a single-sided pattern (TRHS) or a double-sided pattern (TRHD). TRH has dropped from 139K (TRHS) in 2014 [1] to 4.8K (TRHD) in 2020 [11]. Hardware solutions for mitigating Rowhammer typically rely on a tracking mechanism to identify the aggressor rows, and then perform mitigation by refreshing a small number of victim rows on either side of the aggressor row (as specified by the Blast-Radius).

As DRAM scales down, we encounter new DDE modalities. For example, about three years ago, RowPress [2] was released as a new form of DDE. With RowPress, the aggressor row is kept open for a long time. The open row causes a small amount of leakage on the bitlines, and the longer the row remains open, the greater the leakage. For example, keeping the row open for 500 ns results in twice as much leakage as activating and immediately precharging the row. As RowPress causes greater leakage per activation than Rowhammer, it reduces the effective TRH required to induce a bitflip in neighboring rows. RowPress can be handled by existing Rowhammer mitigation by simply converting the row open time into equivalent activations [3], so a row that is open for a long time is treated as having induced more activations. This ability to solve new DDEs using existing solutions is useful, as it avoids the complexity of maintaining separate solutions for different DDEs.

2.4 ColumnDisturb: New Modality of DDE↩︎

Both Rowhammer and RowPress affect victim rows that are within a close spatial proximity of the aggressor row. A few months ago, a new DDE vulnerability, ColumnDisturb [4], was revealed that can cause bitflips in victim rows hundreds of rows away from the aggressor row. The key property used in ColumnDisturb is the shared bitlines that span many rows and subarrays, which cause activity in one row to affect leakage in distant rows, even when the victim rows are in another subarray. The most effective pattern for ColumnDisturb is similar to RowPress (keeping a row open for a long time and repeating it), with the caveat that ColumnDisturb is much slower and must be repeated for several milliseconds [4].

We classify the bitflips caused by ColumnDisturb into two categories: First, Intra-Subarray ColumnDisturb, where the distant victim row remains within the same subarray as the aggressor row. Second, Inter-Subarray ColumnDisturb, where the victim row is in the subarray adjacent to the subarray of the aggressor row (experiments show that ColumnDisturb affects only the adjacent subarray). ColumnDisturb can be mitigated by reducing the refresh rate; however, this incurs significant performance and power overheads. For example, to tolerate a ColumnDisturb of 8ms, we would need to increase the refresh rate by a factor of 4 (from 32ms to 8ms), which would cause a 22% slowdown and quadruple the refresh power.

Existing row-granularity Rowhammer solutions cannot tolerate ColumnDisturb, as they refresh only a few nearby victim rows. To tolerate ColumnDisturb in a practical manner, we focus on solutions, such as SALT [5] and REGA [6], that operate at subarray granularity. Without loss of generality, we focus on ColumnDisturb mitigation using SALT (we discuss REGA in detail in Section 7).

2.5 SALT↩︎

SALT [5] is a recent Rowhammer mitigation that uses subarray-level tracking and mitigation to handle bitflips beyond the Blast-Radius. Figure 3 shows an overview of SALT. SALT equips each subarray with an Activation Counter (ACTR) and a Refresh Pointer (RPTR). On each activation, the ACTR of the accessed subarray is incremented. When ACTR reaches a specified value, SALT initiates an Alert-Back-Off (ABO) [12] signal to obtain the time for mitigation. The time of ABO is sufficient to refresh 7 rows, so 7 rows starting with RPTR are refreshed, and the RPTR is incremented by 7. The ACTR is decremented by APM (Activations Per Mitigation).

Figure 3: Overview of SALT with Refresh Coordination. SALT provisions ACTR and RPTR with each subarray (or gang).

The APM value dictates the threshold tolerated by SALT. For example, with an APM of 13, SALT can tolerate a TRHD of 500 (1K activations across all rows in the subarray). This would mean issuing an ABO every 13 activations, which incurs significant performance overhead (e.g., 17% at a TRHD of 500). SALT reduces ABO overheads by leveraging the time-based refresh (REF) to avoid activity-based refreshes. So, if the activity can be handled by REF, then ABO is avoided. This feature, called Refresh Coordination, is vital for keeping SALT’s performance overhead low. For example, Refresh Coordination reduces the slowdown of SALT from 17% to 0.3% at a TRHD of 500, and to 0% at a TRHD of 1K and beyond. We assume SALT is always implemented with Refresh Coordination.

As the slowdown of SALT is 0% for TRHD \(\ge\) 1K, we can reduce the storage overhead of SALT by operating at a multi-subarray granularity. Such a Ganged-SALT [5] design forms a gang of 2/4/8 subarrays for TRHD of 1K/2K/4K, respectively. As Ganged-SALT uses one ACTR and one RPTR per gang, it reduces storage overhead by 2x-8x. For example, Ganged-SALT can tolerate TRHD of 4K with only 72 bytes of SRAM per bank (which is lower than some TRR implementations already deployed in DDR4 [13], [14]).

2.6 ColumnDisturb Protection with SALT↩︎

As SALT operates at a subarray granularity, it implicitly tolerates intra-subarray ColumnDisturb, as SALT refreshes all the rows within the subarray before 2*TRHD activations are inflicted on the subarray. However, SALT cannot tolerate inter-subarray ColumnDisturb. For example, if the attacker focuses activations on a single subarray, ColumnDisturb can still cause failures in neighboring subarrays because adjacent subarrays do not get mitigated by SALT.

To mitigate inter-subarray ColumnDisturb, SALT can be equipped with a specific extension called ColumnDisturb Protection (CDP) [5]. With CDP, each demand activation not only increments the ACTR of the given subarray but also performs Adjacent-Counter Increment (ACI) for the two adjoining subarrays. The role of ACI is simply to increment the counter of the neighboring subarray without doing any activation. Thus, CDP increments 3 activation counters per demand activation. This ensures that even if activations are concentrated on a single subarray (e.g., SA-2), the neighboring subarrays (e.g., SA-1 and SA-3) are forced to undergo mitigation, and hence they are able to tolerate inter-subarray ColumnDisturb.

Figure 4: Overview of ColumnDisturb Protection (CDP). On activation, CDP issues ACI to neighboring subarrays.

2.7 Overheads of CDP↩︎

As CDP causes two ACIs per activation, it requires 3x the mitigation resources as SALT. This causes significant overhead for SALT and Ganged-SALT. Table ¿tbl:tab:cdpperf? shows the average slowdown of SALT and Ganged-SALT, both with and without CDP, as TRHD is varied from 500 to 4K. Note that SALT handles 2*TRHD activations at arbitrary locations within the subarray (e.g. focused on a single row or spread across multiple rows). CDP incurs significant overhead, increasing the slowdown of SALT/Ganged-SALT from 0.3% to 17%. We want to tolerate ColumnDisturb while avoiding these overheads.

Avg. Slowdown of CDP for SALT and Ganged-SALT
TRHD 500 1K 2K 4K
SALT 0.3% 0% 0% 0%
SALT (CDP) 17.8% 1.9% 0% 0%
Ganged-SALT N/A 0.3% 0.3% 0.3%
Ganged-SALT (CDP) N/A 17.7% 17.3% 16.8%

2.8 Goal of Our Paper↩︎

The goal of our paper is to tolerate ColumnDisturb while incurring negligible overheads. Our key observation is that the primary source of overhead for ColumnDisturb mitigation is the ACI issued at every activation (occurring at a rate of two per activation, i.e., 200%). If we reduce the rate of ACI, we can enable low-overhead ColumnDisturb mitigations. To that end, we propose CODA, a ColumnDisturb Mitigation with Reduced-ACI. We first present our experimental methodology before presenting our solution.

3 Experimental Methodology↩︎

To ensure consistency with prior work, we use the publicly available artifact [5] from the SALT paper for our evaluations. The artifact uses Memsim [5], [15], [16], a cycle-level multi-core simulator with a detailed memory model. Table 1 shows our configuration. We use the updated DDR5 timing specifications. We used a minimalist open-page mapping [17]. We use an adaptive paging policy that closes the page if there are no pending requests to the opened row. Our adaptive policy outperforms both open-page and closed-page.

For SALT, we use APM of 13/26/53/106 (and ATH of 2x of APM) for TRHD of 500/1K/2K/4K, respectively. We assume SALT always uses Refresh-Coordination.

Table 1: Baseline System Configuration
Out-of-Order Cores 8 core, 4GHz, 4-wide, 256 entry ROB
Last Level Cache (Shared) 8MB, 16-Way, 64B lines
Memory specs 32 GB, DDR5
t\({ALERT}\) 180ns (normal) + 350ns (RFM) = 530ns
Banks x Sub-channel x Rank 32\(\times\)​2\(\times\)​1
Rows 64K rows per bank, 8KB rows
Mapping and Closure Policy Minimalist Mapping, Adaptive Page Closure

We use 13 benchmarks from SPEC-2017 with at least one L3-Miss per 1K instructions (L3-MPKI), six from GAP [18], four from STREAM [19], and two data-analytics benchmarks (KMeans [20] and MassTree [21]). We run the workloads in 8-core rate-mode, until each core completes 1 billion instructions (representative slice). We measure performance using weighted speedup. Table ¿tbl:table:workloads? shows workload characteristics, including L3-MPKI and ACT-per-tREFI (per bank). Our workloads are memory-intensive.

Workload Characteristics
Suite Benchmark L3-MPKI ACT-per-tREFI
(per Bank)
SPEC2K17 bwaves 42.7 17.7
fotonik3d 28.3 26.0
lbm 26.7 27.2
parest 23.2 15.6
mcf 23.1 18.1
roms 11.6 16.6
omnetpp 9.3 23.4
xz 5.1 23.6
cam4 5.0 13.9
cactuBSSN 3.5 16.3
wrf 1.2 4.6
xalancbmk 1.1 4.7
blender 1.0 5.2
GAP ConnComp 86.2 27.6
PageRank 46.4 21.5
TriCount 52.2 11.2
BFS 37.8 19.0
BC 20.7 14.2
SSSPPath 10.3 12.5
STREAM add 15.6 14.2
triad 13.4 14.0
copy 12.5 13.0
scale 10.4 12.7
ANALYTICS kmeans 8.9 18.6
masstree 4.8 15.8
Average 21.6 16.7

4 CODA-E: Evade ACI on Demand Activation↩︎

The main source of overhead in ColumnDisturb mitigation is the Adjacent-Counter Increment (ACI) for neighboring subarrays. Thus, each demand activation increments three counters (one for demand and two ACIs for neighbors). We propose CODA to reduce the overhead of mitigating ColumnDisturb by either skipping or reducing the ACI to a much lower rate. We propose three variants of CODA: CODA-E (Evade), CODA-F (Fraction), and CODA-G (Gangskip). Each variant targets a different inefficiency in ACI. In this section, we focus on CODA-E.

4.1 Overview and Design of CODA-E↩︎

CODA-E is based on the observation that the ACI-based increment of neighboring counters is needed only in cases where the underlying rate of mitigation of the neighboring subarrays would not be enough to refresh all the rows in those subarrays. This can occur in pathological cases, where the attacker focuses all activations on a single subarray, so that neighboring subarrays receive no activations (and thus cannot undergo any mitigation). However, for benign workloads, activations are spread across subarrays, so neighboring subarrays also receive demand activations and undergo mitigations. This rate may be sufficient to refresh the neighboring subarray without requiring an ACI-based increment. The key insight for CODA-E is to skip the ACI-based increment if the subarray receives a demand increment.

To implement CODA-E, we need to delay the ACI-based increment so that the subarray has time to override it with a demand activation. Figure 5 shows an overview of our CODA-E design.

Figure 5: Overview of CODA-E. CODA-E adds a Pending-Activation Counter (PAC) to allow overriding of the ACI-based increment by a later demand ACT for that subarray.
Figure 6: Slowdown of SALT, CDP, and CODA-E at TRHD=500 compared to unprotected baseline. The average (Geometric mean) slowdown of SALT (without any ColumnDisturb mitigation) is 0.3%, SALT with CDP is 17.8%, and SALT with CODA-E is 5.9%.

To allow flexibility for ACI-based increment to be later overridden by a demand ACT, CODA-E provisions each subarray with a separate auxiliary counter, called PAC (Pending-Activation Counter). On a demand activation to a subarray, the ACTR associated with the subarray is incremented. Instead of incrementing the ACTR of the two neighboring subarrays, CODA-E increments its PAC. Further, as the subarray receiving the demand ACT already has activity that could trigger mitigation, we reduce the PAC associated with the demand subarray (if it is non-zero). Thus, in steady state, each demand ACT increments two PACs and decrements one PAC, so the total amount of PAC increments per ACT is limited to one.

If an increment to PAC would bring it to its maximum value (determined by the number of bits allocated to the PAC counter), we bypass PAC and increment ACTR directly. Having more bits for PAC increases the time between PAC increments and the ability of a later demand ACT to reduce PAC for that subarray. However, increasing the number of bits also increases the storage cost and the number of pending increments, which can affect the tolerated thresholds. Unless specified otherwise, we use a 4-bit PAC.

4.2 Impact on Rate of ACI↩︎

We note that the theoretical lower bound for CODA-E for ACI is 100%, which is much lower than the 200% for CDP. This is because, on each demand ACT, we cause two PAC increments and reduce the PAC of the demand subarray by 1, so it is effectively one PAC increment per demand ACT. Thus, CODA-E is 2x more efficient than CDP in terms of mitigations required to tolerate ColumnDisturb.

Table ¿tbl:tab:codae? shows the rate of ACI for CDP and CODA-E. For CDP, the rate remains 200% as each activation increments two additional ACTR (one for each neighbor). For CODA-E, the rate depends on the number of bits in the PAC. With a 4-bit PAC, the rate reaches the theoretical minimum for CODA-E (100%). We note that CODA-E is useful only for benign workloads and would not be effective for worst-case patterns that focus all activations on one subarray.

Rate of ACI for CDP and CODA-E
Design Rate of ACI
CDP 200%
CODA-E (1-bit PAC) 154%
CODA-E (2-bit PAC) 121%
CODA-E (3-bit PAC) 104%
CODA-E (4-bit PAC) 100%

4.3 Impact on Performance Overheads↩︎

As CODA-E reduces ACI, it has lower slowdown than CDP. Figure 6 shows the slowdown of SALT, CDP, and CODA-E for TRHD=500. On average, SALT incurs a negligible slowdown of 0.3%, CDP incurs a slowdown of 17.8%, and CODA-E incurs a slowdown of 5.9%. Thus, CODA-E removes most of the slowdown of CDP.

Note that the slowdown of CODA-E can be 2x lower than that of CDP, as the slowdown for SALT depends on whether the activity is lower than what REF can handle. If SALT had half as many activations as REF can handle, the 3x activations (due to CDP) would incur significant overhead, whereas 2x activations (due to CODA-E) would incur no slowdown, so benefits are non-linear.

4.4 Impact on Security↩︎

CODA-E has only a minor impact on the tolerated threshold. With a 4-bit PAC, at most 15 ACI-based counter increments will remain pending and uncommitted to the ACTR. Table ¿tbl:tab:codaetrhd? shows the maximum activations to subarray under SALT (Subarray-Max, which is equal to 2*TRHD), the PAC-Max (15 for 4-bit counter), the total resulting activations with CODA-E (Total), and the percentage increase in the tolerated threshold. Thus, the impact of CODA-E is 1.5% at TRHD of 500 and 0.2% at TRHD of 4K.

Impact of CODA-E on Tolerated Threshold
TRHD Subarray-Max PAC-Max Total Increase (%)
500 1000 15 1015 1.5% (508)
1K 2000 15 2015 0.8% (1.01K)
2K 4000 15 4015 0.4% (2.01K)
4K 8000 15 8015 0.2% (4.01K)

5 CODA-F: Fractional ACI-Based Increments↩︎

Our second design for CODA, CODA-F (Fraction), is based on the observation that ColumnDisturb is a much slower attack than Rowhammer. For example, at TRHD=500, one could perform a Rowhammer attack within 50 microseconds (tRC of 50 ns, multiplied by 1000 activations, 500 each on two attack rows). Comparatively, ColumnDisturb takes several milliseconds to cause failures. For example, the ColumnDisturb paper [4] targets solutions for a duration of 8ms. This also means that the ColumnDisturb solutions have a much longer time (milliseconds) to perform the mitigation.

CODA-F is based on the insight that a 1-to-1 increment between demand ACT and increment to neighboring subarrays is unnecessary. We should perform ACI at a rate such that, even if the neighboring subarray receives no demand ACTs, it can still refresh all rows within the subarray within the ColumnDisturb duration. For example, an attacker may use a RowPress pattern on the aggressor row(s) to cause intra-subarray ColumnDisturb in the adjacent subarray. We could consider each 500 ns duration of open time to be equivalent to one activation [3], [5]. Then, over 8ms [4], the attacker would need to perform 16K equivalent activations on the aggressor row(s). If SALT is designed for a TRHD of 500, then SALT would refresh all rows in the demand subarray 16 times within 8ms. The adjacent subarrays need to refresh all their rows only once over the 8ms period, so their counters need to be updated at a rate of only 1/16 per demand activation. The respective fractional rates (f) for TRHD at 1K, 2K, and 4K are 1/8, 1/4, and 1/2, respectively.

Figure 7: Slowdown of CDP, CODA-E, CODA-F, CODA-EF at TRHD=500. The average slowdown of CDP is 17.8%, of CODA-E is 5.9%, of CODA-F is 0.6%, and of CODA-EF is 0.3% (Note: bars for CODA-F and CODA-EF are vanishingly small, hence not visible).

5.1 Overview and Design of CODA-F↩︎

To implement CODA-F, we need to support a fractional increment of ACTR. However, to avoid complexity (and to be synergistic with CODA-E), we first increment the ACI in a separate Pending-Activation Counter (PAC) by a sub-integer amount and write to the ACTR only if the PAC is 1 (or, alternatively, the maximum value). Figure 8 shows the overview of CODA-F.

Figure 8: Overview of CODA-F. CODA-F performs fractional counter increments (by a value "f"=1/16 to 1/2) to PAC and increments ACTR if the PAC reaches the maximum value.

Similar to CODA-E, CODA-F also provisions an auxiliary counter (PAC) with each subarray. However, the representation of the value stored in PAC differs from that in CODA-E. On activation of a demand subarray (e.g. SA-2), CODA-F increments the associated ACTR. CODA-F also increments the PAC of the adjacent subarrays (e.g. SA-1 and SA-3) but by a smaller fractional amount (f), which is dictated by the ColumnDisturb duration and the tolerated TRHD (e.g. f=1/16 for TRHD=500 and ColumnDisturb duration of 8ms). Thus, PAC stores the value in fixed-point format instead of integer.

If an ACI increment causes PAC to reach its maximum value, we reduce PAC by 1 and increment the ACTR of the associated subarray by 1. Without loss of generality, we use a 4-bit PAC (which means it can count 15 fractional increments to reach the maximum value).

5.2 CODA-EF: Exploiting Synergy With CODA-E↩︎

CODA-F by itself can reduce the overhead of ACI increments by a factor of 1/f. However, we can combine CODA-E and CODA-F for even greater effectiveness. The combined scheme, which we call CODA-EF, can be implemented by decrementing the PAC of the demand subarray by 1 (if the PAC is less than 1, it is reset to 0).

5.3 Impact on Rate of ACI↩︎

Table ¿tbl:tab:codaf? shows the rate of ACI for CDP, CODA-E, CODA-F, and CODA-EF. While CODA-E reduces ACI by 2x, CODA-F reduces ACI by 16x-2x. When combined, the reduction is much larger, because if "f" is small (say 1/16), several activations are required for a PAC to reach 1, however, a single demand activation on the subarray resets PAC to zero. CODA-EF reduces ACI by 11.8x to 143x (the average over all of our evaluated workloads).

Rate of ACI for CDP, CODA-E, CODA-F, and CODA-EF (number with "x" denotes relative reduction versus CDP)
TRHD CDP CODA-E CODA-F CODA-EF
500 200% 100% (2x) 12.5% (16x) 1.4% (143x)
1K 200% 100% (2x) 25% (8x) 1.8% (109x)
2K 200% 100% (2x) 50% (4x) 3.6% (56x)
4K 200% 100% (2x) 100% (2x) 17% (11.8x)

5.4 Impact on Performance Overheads↩︎

Because CODA-F reduces the ACI rate more than CODA-E, it incurs a lower performance overhead than CODA-E. Figure 7 shows the slowdown of CDP, CODA-E, CODA-F, and CODA-EF at TRHD=500. The bar labeled Gmean shows the geometric mean slowdown. On average, CDP incurs a slowdown of 17.8%, CODA-E of 5.9%, CODA-F of 0.6%, and the combination CODA-EF of 0.3% (same as SALT without ColumnDisturb mitigation). Thus, CODA-F removes nearly all of the slowdown caused by ColumnDisturb mitigation.

5.5 Impact on Security↩︎

CODA-F (and CODA-EF) has a negligible impact on the tolerated threshold. With a 4-bit PAC, at most 15 ACI-based counter increments will remain pending. Table ¿tbl:tab:codaftrhd? shows the maximum activations to subarray under SALT (Subarray-Max, which is equal to 2*TRHD), the PAC-Max (depends on "f"), the total resulting activations with CODA-F (Total), and the percentage increase in the tolerated threshold. Thus, the impact of CODA-F/CODA-EF is 0.1% across all TRHD.

Impact of CODA-F/CODA-EF on Tolerated Threshold
TRHD Subarray-Max PAC-Max Total Increase (%)
500 1000 1 1001 0.1% (501)
1K 2000 2 2002 0.1% (1.001K)
2K 4000 4 4004 0.1% (2.002K)
4K 8000 8 8008 0.1% (4.004K)

6 CODA-G: Skip ACI for Intra-Gang Increments↩︎

We observe that at thresholds of 1K or higher, we are likely to implement Ganged-SALT to reduce the stored overhead (e.g., at a 4K threshold, Ganged-SALT can be implemented with just 72 bytes of SRAM per bank, which is less than even some current TRR implementations). Ganged-SALT provisions a single Activation Counter (ACTR) over multiple subarrays of the gang, and all the rows in the gang are guaranteed to get refreshed before the gang receives 2*TRHD activations. The multi-subarray nature of Ganged-SALT provides further opportunities to reduce ACI. We call the CODA variant that exploits intra-gang inefficiency as CODA-G.

6.1 Overview and Design of CODA-G↩︎

The key insight in CODA-G is to skip the ACI-based increment for the neighboring subarray when it is within the same gang. We call this optimization Gangskip. Figure 9 shows an overview of CODA-G. Ganged-SALT uses four subarrays per gang (e.g., Gang-0 contains SA-1 to SA-4, and so on). Each gang has a single ACTR. When a row is activated, the ACTR of the gang is incremented. For example, if SA-8 has an activation, the ACTR of Gang-1 is incremented.

Figure 9: Overview of CODA-G. CODA-G exploits the multi-subarray granularity to skip ACI if the adjacent subarray is within the same gang.

A straightforward way to implement ColumnDisturb Protection (CDP) for Ganged-SALT is to perform ACI on the subarrays neighboring the accessed subarray. For example, on an activation to SA-8, we not only increment ACTR for SA-8 but also for SA-7 and SA-9. However, we note that the increment for SA-7 is unnecessary as it maps to the same ACTR as SA-8, for which we have already done an ACTR increment. With CODA-G, we increment the adjacent subarray only if it is in a different gang (e.g., SA-9). CODA-G skips the ACI activation-based counter increment if the adjacent subarray is in the same gang (e.g., SA-7). The insight is that the increment from demand activation is sufficient to protect all subarrays within the gang. With CODA-G, only a single ACI occurs if the activation targets a border subarray within the gang, and ACIs are skipped entirely for non-border subarrays (as both adjacent neighbors map to the same gang as the demand subarray). Thus, CODA-G becomes even more efficient as the gang size increases.

6.2 Synergy of CODA-G with CODA-EF↩︎

The advantage of CODA-G over CODA-E and CODA-F is that CODA-G can be implemented without incurring any additional storage, whereas CODA-E and CODA-F both require an additional counter (PAC). However, both CODA-G and CODA-EF are complementary and, when combined, achieve greater reduction than either scheme alone. To implement CODA-G with CODA-EF, we implement CODA-EF at gang granularity and skip the ACI for neighboring subarrays within the same gang. We refer to the combination of CODA-G and CODA-EF as CODA-EFG.

6.3 Impact on Storage Overheads↩︎

CODA-G incurs no additional storage overhead (the decision to skip is based solely on the address of the accessed subarray). CODA-E, CODA-F, and CODA-EF require a 4-bit PAC counter with each subarray/gang. Table ¿tbl:tab:storage? shows the storage overhead (SRAM bytes per bank) for Ganged-SALT, CODA-EF, CODA-G, and CODA-EFG. The storage overhead of Ganged-SALT with CODA remains quite small, ranging from 600 bytes at TRHD=500 to 88 bytes at TRHD=4K.

Storage overhead (SRAM bytes per bank) for SALT, CODA-EF, Ganged-SALT, and CODA-EFG
TRHD
500
1K
2K
4K

6.4 Impact on Rate of ACI↩︎

With a gang of N subarrays and CODA-G standalone, the expected rate of ACI increments per demand activation is 2/N (so, 25% for 8 subarrays per gang, which is relevant for TRHD=4K). The rate of ACI for CODA-EFG can be lower than the product of either scheme, standalone, as CODA-E uses demand activation to reset the PAC if it is below 1. So, there is a non-linear benefit if the combination can frequently keep PAC below 1.

Rate of ACI for CDP, CODA-EF, CODA-G and CODA-EFG (number with "x" is relative reduction versus CDP)
TRHD CDP CODA-EF CODA-G CODA-EFG
500 200% 1.4% (143x) 200% (1x) 1.4% (143x)
1K 200% 1.8% (109x) 100% (2x) 0.23% (869x)
2K 200% 3.6% (56x) 50% (4x) 0.15% (1300x)
4K 200% 17% (11.8x) 25% (8x) 0.16% (1250x)

Table ¿tbl:tab:codag? shows the rate of ACI per demand activation for CDP, CODA-EF, CODA-G, and CODA-EFG as TRHD is varied from 500 to 4K. CODA-G reduces the rate of ACI from 200% (for CDP) to 25% at TRHD of 4K. The combination of CODA-EFG is highly effective across all thresholds, as it uses CODA-G (which becomes more effective at higher TRHD) and CODA-EF (which becomes more effective at lower TRHD). CODA-EFG provides a 143x-1300x reduction in the ACI rate compared to CDP. Thus, CODA-EFG makes it possible to tolerate ColumnDisturb while incurring negligible (virtually zero) performance and power overheads.

6.5 Impact on Performance Overheads↩︎

As CODA-G and CODA-EFG significantly reduce the ACI, the performance overhead of mitigating ColumnDisturb becomes negligible. Table ¿tbl:tab:codagperf? shows the slowdown for Ganged-SALT (without any ColumnDisturb mitigation) and Ganged-SALT implemented with CDP, CODA-G, and CODA-EFG, as the TRHD is varied from 1K to 4K (gang sizes of 2 to 8 subarrays, respectively) and the target ColumnDisturb duration is 8ms. On average, Ganged-SALT incurs a slowdown of only 0.3%, whereas CDP increases it to 17.8%. With CODA-G, the average slowdown ranges from 5.9% (at TRHD=1K) to 0.9%, with no additional storage overhead. In contrast, with CODA-EFG, the slowdown is identical to that of Ganged-SALT without any ColumnDisturb protection (average of 0.3%). Thus, CODA-EFG can protect ColumnDisturb at zero performance overhead.

Average Slowdown of Ganged-SALT, CDP, CODA-G, and CODA-EFG at TRHD 1K-4K (ColumnDisturb of 8ms)
TRHD Ganged-SALT CDP CODA-G CODA-EFG
1K 0.3% 17.7% 5.9% 0.3%
2K 0.3% 17.3% 1.9% 0.3%
4K 0.3% 16.8% 0.9% 0.3%

6.6 Impact on Security↩︎

CODA-G exploits the multi-subarray nature of Ganged-SALT implementation to skip the update for the neighboring subarray that is within the same gang. Skipping this update has no impact on security, as the update is unnecessary: the ACTR of the gang is incremented due to demand activation, thereby protecting the entire gang. When CODA-G is combined with CODA-EF, there is a minor impact of threshold from CODA-EF. Table ¿tbl:tab:codagtrhd? shows the effective TRHD with CODA-G and CODA-EFG for target TRHD of 1K to 4K. The impact of CODA-EFG is 0.1% across all TRHD.

Impact of CODA-G/CODA-EFG on Tolerated TRHD
Target-TRHD CODA-G CODA-EFG Increase
1K 2000 2002 0.1%
2K 4000 4004 0.1%
4K 8000 8008 0.1%

6.7 Impact of ColumnDisturb Duration↩︎

Similar to prior work [4], we target a ColumnDisturb duration of 8ms. The ACI of CODA-E and CODA-G do not depend on the ColumnDisturb duration, so slowdowns remain unaffected. The ACI of CODA-F and CODA-EFG depends on the ColumnDisturb duration (as the "f" value is affected), so the slowdown can vary. Table [tbl:tab:codagperfcd] shows the slowdown for Ganged-SALT (without any ColumnDisturb mitigation) and Ganged-SALT implemented with CDP, CODA-G, and CODA-EFG, as the TRHD is varied from 500 to 4K (at 500, Ganged-SALT degenerates to SALT) and the target ColumnDisturb duration is set to 4ms. The slowdowns of CDP and CODA-G remain the same. We note that CODA-G is not applicable at TRHD=500 because the gang contains only one subarray. Even at 4ms, CODA-EFG incurs 0% additional slowdown compared to Ganged-SALT.

Average Slowdown of Ganged-SALT, CDP, CODA-G, and CODA-EFG at TRHD of 500-4K (ColumnDisturb of 4ms)
TRHD Ganged-SALT CDP CODA-G (4ms) CODA-EFG (4ms)
500 0.3% 17.8% N/A 0.3%
1K 0.3% 17.7% 5.9% 0.3%
2K 0.3% 17.3% 1.9% 0.3%
4K 0.3% 16.8% 0.9% 0.3%

We also consider scaling CODA to lower ColumnDisturb duration of 1ms-2ms. However, in this regime, we must still ensure that the number of activations required for ColumnDisturb is more than what is needed for Rowhammer (e.g. if we treat row-open time of 500ns as one activation, then at 1ms duration, ColumnDisturb would need only 2K activations, so it is not meaningful to analyze TRHD of 2K or greater). So, for this analysis, we assume that row-open time of 250ns is equivalent to one activation (for our system, the adaptive page policy closes the page much earlier), and we conduct this analysis for TRHD of 1K and lower.

TRHD Ganged-SALT CDP
500 0.3% 17.8%
1K 0.3% 17.7%

7 Efficiently Protecting REGA with CODA↩︎

CODA can be implemented in any design that operates at subarray granularity, as such designs implicitly handle intra-subarray ColumnDisturb and can tolerate inter-subarray ColumnDisturb via ACI to neighboring subarrays. Our solutions and insights for efficient ColumnDisturb mitigation are not limited to SALT. In this section, we analyze another subarray-granularity mitigation, called REGA (Refresh-Generating Activations) [6], and show how CODA can reduce the overheads for mitigating ColumnDisturb.

7.1 Background on REGA↩︎

The primary difference between SALT and REGA lies in the mechanism for obtaining mitigation resources. While SALT uses ABO [12] to obtain mitigation time on demand and incurs slowdowns due to ABO, REGA modifies the DRAM circuitry to obtain mitigation time for each activation. Figure 10 shows the overview of REGA.

Figure 10: Overview of REGA. REGA uses an auxiliary row buffer to concurrently refresh one/more rows on each ACT.

REGA modifies the subarray to have a second auxiliary row buffer (AuxBuf). On an activation, the accessed row is sensed using one row buffer, while the AuxBuf is used to concurrently refresh one (or more) rows (pointed by RPTR) from the subarray. The advantage of REGA is that, as mitigation occurs in the background, there is no time penalty if the refresh can be completed within the activation and precharge period of the accessed row. However, the disadvantage of REGA is that it modifies the DRAM circuitry, making it more difficult to adopt in commodity devices. Furthermore, it also exacerbates DRAM power consumption.

If one row is refreshed per ACT, and the subarray contains 512 rows, all rows are refreshed within 512 ACTs (tolerated TRHD of 256). To tolerate larger thresholds, REGA must refresh one row per several ACTs. This can be achieved by having an Activation Counter (ACTR) per subarray. ACTR is incremented on each activation. When ACTR reaches \(N\), a refresh is performed concurrently with the ACT, and ACTR is reset. For example, for TRHD=500, we need N=2, and for TRHD=4K, we need N=16. This allows REGA to tolerate a higher threshold while incurring lower power overheads.

7.2 REGA with ColumnDisturb Protection↩︎

As REGA operates at subarray granularity, it tolerates intra-subarray ColumnDisturb. However, REGA cannot tolerate inter-subarray ColumnDisturb. We can use CDP [5] to protect REGA against ColumnDisturb, a design we call REGA-CDP. On activation, REGA-CDP not only increments the ACTR of the demand subarray but also that of the adjacent subarrays via ACI. REGA-CDP incurs 3x the refresh-power overhead for mitigation compared to REGA.

7.3 REGA-CODA: Overview and Design↩︎

The power overhead of ColumnDisturb Protection for REGA can be reduced by applying CODA principles. Figure 11 shows the overview of such a REGA-CODA design. Each subarray has two counters, ACTR and PAC. ACTR counts activations, and PAC counts delayed ACI from adjacent subarrays. We use optimizations from both CODA-E (decrementing PAC on-demand activations to a subarray) and CODA-F (incrementing PAC by a fractional value "f" instead of one) to implement REGA-CODA.

Figure 11: Overview of REGA-CODA. REGA-CODA reduces/skips the ACI to the adjacent subarrays using insights from CODA-E (override by demand) and CODA-F (fraction).

On an activation, we increment the ACTR of the demand subarray. If the ACTR reaches N (N=2, 4, 8, 16 for TRHD of 500, 1K, 2K, 4K, respectively), REGA refreshes a row and resets ACTR. Similar to CODA-F, on activation, we also increment the PAC of the adjacent subarray by a value of "f" (f=1/16, 1/8, 1/4, 1/2 for TRHD of 500, 1K, 2K, 4K, respectively). If PAC reaches its maximum value, we increment ACTR by 1 and reduce PAC by 1. Similar to CODA-E, upon activation, we reduce the PAC of the demand subarray by 1 (and reset it to 0 if it was less than 1). The fractional increment and reset-on-demand-activations significantly reduce the rate of ACI.

7.4 Impact on Rate of ACI↩︎

REGA-CDP converts a single activation into three ACTR increments. Thus, the rate of ACI is 200%. With REGA-CODA, this rate reduces significantly. Table ¿tbl:tab:rega? shows the rate of ACI (per activation) for REGA-CDP and REGA-CODA. The numbers in parentheses show the relative reduction compared to REGA-CDP. REGA-CODA reduces the ACI to 1.4% (at TRHD=500) or 17% (at TRHD=4K). Thus, REGA-CODA is much more efficient than REGA-CDP.

Rate of ACI for REGA-CDP and REGA-CODA (number with "x" shows relative reduction from REGA-CDP)
TRHD REGA-CDP REGA-CODA
500 200% 1.4% (143x lower)
1K 200% 1.8% (109x lower)
2K 200% 3.6% (56x lower)
4K 200% 17% (11.8x lower)

7.5 Impact on Power Overheads↩︎

For a mitigation rate of one (or fewer) row refreshes per ACT, REGA incurs no performance overhead, as the refresh happens concurrently with the activation. Thus, the performance of all REGA designs (REGA, REGA-CDP, REGA-CODA) is identical. The primary overhead of REGA is the additional power consumed for the mitigative refreshes. This power overhead is proportional to the mitigation rate.

Figure 12: Increase in refresh power with REGA, REGA-CDP, and REGA-CODA. REGA-CODA mitigates ColumnDisturb while incurring negligible power overheads.

Figure 12 shows the increase in refresh power due to REGA, REGA-CDP, and REGA-CODA. With REGA-CDP, the refresh power overhead triples, so the 51% increase for REGA becomes 153% with REGA-CDP. With REGA-CODA, the increase is negligible, increasing from 51% to 52%. Across all thresholds, the increase in power with REGA-CODA is within 1% of that of REGA without any ColumnDisturb mitigation. Thus, CODA can protect REGA against ColumnDisturb while incurring negligible power overheads.

8 Related Work↩︎

As ColumnDisturb is a recently disclosed vulnerability (publicly released only five months ago), there are not many studies on mitigating ColumnDisturb. In this section, we describe work closely related to ColumnDisturb, as well as other Data-Disturbance Errors.

8.1 ColumnDisturb Mitigation via Fast Refresh↩︎

The ColumnDisturb paper [4] mentions two mitigations to limit ColumnDisturb to 8ms. First, reduce the DRAM refresh time from 32ms to 8ms, which incurs significant energy and performance overhead. Second, a scheme called PRVR (Proactively Refreshing ColumnDisturb Victim Rows), which refreshes all rows in three subarrays once before the aggressor row is hammered or pressed enough times to induce a bit-flip. Unfortunately, the paper does not provide any design details for PRVR, such as the tracking mechanism, the number of activations targeted to 8ms, or the mitigation rate across the three subarrays to ensure that mitigations are completed within 8ms. The paper observes that PRVR still incurs approximately 30% of the performance overhead and 25% of the energy overhead of an 8ms refresh, indicating that the overhead of PRVR remains high.

Figure 13: Slowdown of Different ColumnDisturb Mitigations. CODA allows mitigations at 0% slowdowns.

Figure 13 shows the slowdown of handling ColumnDisturb with Ref-8ms, PRVR (estimated), SALT-CDP-500 (TRHD=500), Ganged-SALT-CDP-4K (TRHD=4K), CODA-500 (TRHD=500), and CODA-4K (TRHD=4K). For SALT-based designs, we normalized the slowdown with respect to the SALT design without any ColumnDisturb protection. For PRVR, we estimate the slowdown as 30% relative to Ref-8ms (as the paper does not provide sufficient design details). CODA-500 is built on top of SALT and uses CODA-EF. CODA-4K is built on top of Ganged-SALT and uses CODA-EFG. An 8ms REF incurs a 22.3% slowdown, and with PRVR (estimated), the slowdown remains 6.7%. CDP can incur 17% slowdown, whereas CODA has 0% slowdown for ColumnDisturb protection even at TRHD=500 (as it eliminates almost all of the ACI for neighboring subarrays).

8.2 Importance of Refresh Coordination↩︎

We assume that both SALT and Ganged-SALT use Refresh Coordination [5] to avoid ABO. CDP and CODA are also applicable to designs, such as Silver-Bullet [22], [23], that do not employ Refresh-Coordination. Such designs incur high overheads. Table ¿tbl:tab:sb? shows the slowdown of SALT(NR), SALT(NR) with CDP, and SALT(NR) with CODA-EF, where \(NR\) denotes No-Refresh-Coordination.

Slowdown of Designs with No-Refresh-Coordination (NR) for SALT, CDP, and CODA-EF.
TRHD SALT(NR) SALT(NR)+CDP SALT(NR)+CODA-EF
500 13.2% 54.0% 13.5%
1K 6.4 % 25.6% 6.6%
2K 3.0% 12.0% 3.2%
4K 1.5% 5.8% 1.9%

SALT(NR) incurs 13% at TRHD=500, and CDP increases it to 54%, so the overheads of CDP are quite high. With CODA-EF, the performance is within 0.5% of the SALT(NR) design, which does not perform any ColumnDisturb mitigation. Thus, CODA makes it practical to tolerate ColumnDisturb even for designs that do not employ Refresh-Coordination.

8.3 Inefficacy of Row-Granularity Mitigations↩︎

Typical hardware-based defenses for Rowhammer track aggressor rows and refresh a small number of victim rows adjacent to the aggressor row. Unfortunately, such row-granularity tracking cannot handle ColumnDisturb, as ColumnDisturb causes failures in rows hundreds of rows away from the aggressor row. For example, the state-of-the-art row-granularity mitigation proposal is PRAC+ABO from JEDEC [12]. Several recent works, such as Chronus [24], MOAT [16], and QPRAC [25], use PRAC to securely tolerate Rowhammer. As their mitigations are limited to a small blast radius, they cannot handle failures in distant rows.

8.4 Inefficacy of Rate-Limit Based Mitigations↩︎

Prior works have tried to mitigate Rowhammer by limiting the number of times an aggressor row gets accessed to TRHD by employing some form of Dynamic-Row-Migration (such as RRS [26], AQUA [27], SRS [28], Rubix [29], and SHADOW [30]) or explicitly enforcing rate limits (such as via Blockhammer [31]). Unfortunately, such designs are not effective at tolerating ColumnDisturb, as the attacker can spread the activations over hundreds of rows within a subarray, so each row would receive only dozens of activations and be within the permissible rate limits of these designs.

9 Conclusion↩︎

As DRAM scales down, inter-cell interference increases, leading to new modalities of Data-Disturbance Errors (DDEs). ColumnDisturb, the newest form of DDEs, is particularly challenging to mitigate because it induces failures in rows hundreds of rows away from the aggressor row, and even in adjacent subarrays. ColumnDisturb can be mitigated by subarray-granularity mitigations, such as SALT, and using Adjacent-Counter Increment (ACI) for neighboring subarrays. Unfortunately, such a solution requires 3x mitigations and incurs significant overhead. This paper proposes CODA to significantly reduce or eliminate the requirement of adjacent counter increments. We propose three variants of CODA, and together they reduce the activation overheads of ColumnDisturb mitigation by 12x-1300x. CODA makes it feasible to securely tolerate ColumnDisturb while incurring nearly zero performance overheads.

Acknowledgments↩︎

We thank the anonymous reviewers of MICRO-2026 for their feedback. This work was funded, in part, by NSF grant 33304.

References↩︎

[1]
Y. Kim et al., “Flipping bits in memory without accessing them: An experimental study of DRAM disturbance errors,” ISCA, 2014.
[2]
H. Luo et al., “RowPress: Amplifying read disturbance in modern DRAM chips,” in Proceedings of the 50th annual international symposium on computer architecture, 2023, doi: 10.1145/3579371.3589063.
[3]
A. Saxena, A. Jaleel, and M. Qureshi, “Impress: Securing dram against data-disturbance errors via implicit row-press mitigation,” in 2024 57th IEEE/ACM international symposium on microarchitecture (MICRO), 2024, pp. 935–948.
[4]
İ. E. Yüksel, A. Olgun, F. N. Bostancı, H. Luo, A. G. Yağlıkçı, and O. Mutlu, “ColumnDisturb: Understanding column-based read disturbance in real DRAM chips and implications for future systems,” in IEEE/ACM international symposium on microarchitecture (MICRO), 2025.
[5]
M. K. Qureshi, SALT: Track-and-Mitigate Subarrays, Not Rows, for Blast-Radius-Free Rowhammer Defense ,” in 2026 IEEE international symposium on high performance computer architecture (HPCA), Feb. 2026.
[6]
M. Marazzi, F. Solt, P. Jattke, K. Takashi, and K. Razavi, REGA: Scalable Rowhammer Mitigation with Refresh-Generating Activations,” in IEEE Symposium on Security and Privacy (SP), 2023.
[7]
Y. Kim, V. Seshadri, D. Lee, J. Liu, and O. Mutlu, “A case for exploiting subarray-level parallelism (SALP) in DRAM,” in 2012 39th annual international symposium on computer architecture (ISCA), 2012.
[8]
H. Hassan, A. Olgun, A. G. Yaglikci, H. Luo, O. Mutlu, and E. Zurich, Self-Managing DRAM: A Low-Cost Framework for Enabling Autonomous and Efficient DRAM Maintenance Operations ,” in 2024 57th IEEE/ACM international symposium on microarchitecture (MICRO), 2024.
[9]
H. Taneja, A. Hajiabadi, M. Marazzi, K. Razavi, and M. Qureshi, “MIRZA: Efficiently mitigating rowhammer with randomization and ALERT,” in IEEE international symposium on high performance computer architecture (HPCA), 2026.
[10]
M. Qureshi, AutoRFM: Scaling Low-Cost In-DRAM Trackers to Ultra-Low Rowhammer Thresholds ,” in HPCA, 2025.
[11]
J. S. Kim et al., “Revisiting rowhammer: An experimental analysis of modern dram devices and mitigation techniques,” in ISCA, 2020, pp. 638–651.
[12]
JEDEC, JESD79-5C: DDR5 SDRAM Specifications,” 2024.
[13]
M. Marazzi, P. Jattke, F. Solt, and K. Razavi, Protrr: Principled yet optimal in-dram target row refresh,” in IEEE Symposium on Security and Privacy (SP), 2022, pp. 735–753.
[14]
P. Jattke, V. van der Veen, P. Frigo, S. Gunter, and K. Razavi, https://comsec.ethz.ch/wp-content/files/blacksmith_sp22.pdfBLACKSMITH: Rowhammering in the Frequency Domain,” in 43rd IEEE Symposium on Security and Privacy’22 (Oakland), 2022.
[15]
M. Qureshi, S. Qazi, and A. Jaleel, “MINT: Securely mitigating rowhammer with a minimalist in-DRAM tracker,” in MICRO, 2024.
[16]
M. Qureshi and S. Qazi, MOAT: Securely mitigating rowhammer with per-row activation counters,” ASPLOS-2025, 2025.
[17]
D. Kaseridis, J. Stuecheli, and L. K. John, “Minimalist open-page: A DRAM page-mode scheduling policy for the many-core era,” in Proceedings of the 44th annual IEEE/ACM international symposium on microarchitecture, 2011, pp. 24–35.
[18]
K. A. S. Beamer and D. Patterson, “The GAP benchmark suite,” in arXiv preprint arXiv:1508.03619, 2015.
[19]
J. D. McCalpin, Memory Bandwidth and Machine Balance in Current High Performance Computers,” IEEE (TCCA) Newsletter, 1995.
[20]
S. P. Lloyd, Least squares quantization in PCM,” IEEE Trans. Inf. Theory, vol. 28, no. 2, pp. 129–136, 1982.
[21]
Y. Mao, E. Kohler, and R. T. Morris, Cache craftiness for fast multicore key-value storage,” in eurosys12, 2012, pp. 183–196.
[22]
A. G. Yağlıkçı, J. S. Kim, F. Devaux, and O. Mutlu, “Security analysis of the silver bullet technique for RowHammer prevention.” 2021, [Online]. Available: https://arxiv.org/abs/2106.07084.
[23]
F. Devaux and R. Ayrignac, Filed 2020-08-04; Priority FR2006541 (2020-06-23)“Method and circuit for protecting a DRAM memory device from the row hammer effect,” US 10,885,966 B1, Jan. 05, 2021.
[24]
O. Canpolat et al., “Chronus: Understanding and securing the cutting-edge industry solutions to DRAM read disturbance,” HPCA, 2025.
[25]
J. Woo, C. S. Lin, P. J. Nair, A. Jaleel, and G. Saileshwar, “Qprac: Towards secure and practical prac-based rowhammer mitigation using priority queues,” HPCA, 2025.
[26]
G. Saileshwar, B. Wang, M. Qureshi, and P. J. Nair, “Randomized row-swap: Mitigating row hammer by breaking spatial correlation between aggressor and victim rows,” in Proceedings of the 27th ACM international conference on architectural support for programming languages and operating systems, 2022, pp. 1056–1069.
[27]
A. Saxena, G. Saileshwar, P. J. Nair, and M. Qureshi, Aqua: Scalable rowhammer mitigation by quarantining aggressor rows at runtime,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2022, pp. 108–123.
[28]
J. Woo, G. Saileshwar, and P. J. Nair, “Scalable and secure row-swap: Efficient and safe row hammer mitigation in memory systems,” in 2023 IEEE international symposium on high-performance computer architecture (HPCA), 2023, pp. 374–389.
[29]
A. Saxena, S. Mathur, and M. Qureshi, “Rubix: Reducing the overhead of secure rowhammer mitigations via randomized line-to-row mapping,” 2024.
[30]
M. Wi et al., SHADOW: Preventing Row Hammer in DRAM with Intra-Subarray Row Shuffling,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2023, pp. 333–346.
[31]
A. G. Yağlikçi et al., “BlockHammer: Preventing RowHammer at low cost by blacklisting rapidly-accessed DRAM rows,” in 2021 IEEE international symposium on high-performance computer architecture (HPCA), 2021, pp. 345–358.