Does Testing Enhance New Learning Because It Insulates Against Proactive Interference?(1)
Jun 07, 2023
Because PI is theoretically neutral, this release-fromPI account leaves open the candidate mechanisms through which interpolated testing enhances new learning. PI can degrade subsequent learning via processes that operate either at retrieval or at encoding. In this paper, we focused on the retrieval-based explanation of PI for the forward testing effect. Szpunar et al. (2008) ascribed the forward testing effect to an enhancement in list discrimination – the idea is that testing boosts retention of the tested items, which in turn allowed participants to better discriminate list membership of the studied items and to avoid intrusions (see also Chan & McDermott, 2007; Jang & Huber, 2008; Postman et al., 1968; Postman & Keppel, 1977) – consequently, the original conceptualization of the release-from-PI account focused on retrieval processes, and we did the same here.

Effects of Cistanche-Improve Memory
Click here to view Cistanche Improve Memory products
【Ask for more】 Email:cindy.xue@wecistanche.com / Whats App: 0086 18599088692 / Wechat: 18599088692
Abstract
Taking a test on previously learned material can enhance new learning. One explanation for this forward-testing effect is that retrieval inoculates learners from proactive interference (PI). Although this release-from-PI account has received considerable empirical support, most extant evidence is correlational rather than causal. We tested this account by manipulating the level of PI that participants experience as they studied several lists while receiving interpolated tests or not. In Experiments 1 and 2, we found that testing benefited new learning similarly regardless of PI level. These results contradict those from Nunes and Weinstein (Memory, 20(2), 138–154, 2012), who found no forward testing effect when encoding conditions minimized PI. In Experiments 3 and 4, we failed to replicate their results. Together, our data indicate that reduced PI might be a byproduct, rather than a causal factor, of the forward testing effect.

Benefits of cistanche tubulosa-Improve Memory
Keywords
Retrieval practice · New learning · Forward testing effect · Proactive interference
Introduction
Many students and teachers consider tests as an assessment tool (Bonner, 2009; Karpicke et al., 2009). However, a large body of research has established that taking tests enhances learning (Chan, Meissner, et al., 2018; Karpicke & Roediger, 2008; Rowland, 2014; for exceptions, see Chan et al., 2017; Davis & Chan, 2015; Finn & Roediger, 2013). The beneficial effects of testing can be classified in two ways. On the one hand, testing promotes the learning of tested material, which is sometimes referred to as the backward testing effect. On the other hand, testing of previously studied material improves subsequent learning of new material, which is sometimes termed the forward testing effect. This latter effect is the focus of the current study.
The forward testing effect is often examined using a multi-list learning, interpolated testing paradigm. At the beginning of the study, participants are told that they might or might not take a test after studying each list. But in actuality, participants in the tested condition always take a test after each list, whereas those in the control condition either restudy the list or do a filler task such as solving math problems. All participants are tested on the final, target list (Fig. 1 shows the design of our experiments, but they also conform to the canonical design used in the literature). If the tested participants perform better on the target test than the control participants, then interpolated testing has facilitated new learning (i.e., a forward testing effect). The forward testing effect is robust and has been observed with a variety of conditions (for a review, see Chan, Manley, et al., 2018).

What does cistanche do-Improve Memory
Although there is rich literature on the forward testing effect, most studies have been empirically oriented (cf. Chan et al., 2020; Chan, Manley, et al., 2018; Kliegl & Bäuml, 2021). Recently, researchers have started to explore the mechanisms underlying the effect (for a review, see Chan, Meissner, et al., 2018; Pastötter & Bäuml, 2014; Yang et al., 2018), but the number of empirical studies far outstripped that of theoretical ones. Thus, the goal of the current study was to extend the theoretical understanding of the forward testing effect. Specifically, we provided a critical test of the release from proactive interference explanation (Szpunar et al., 2008; Yang et al., 2018), which suggests that interpolate testing facilitates subsequent learning because it insulates participants against the buildup of proactive interference (PI) from prior learning. Below, we review the relevant literature and explain the motivation of our study.

Fig. 1 Design for the four experiments. S = Study, M = Math, T = Test. Subscripts indicate list structures (ML = Mixed List, BL = Blocked List). In a mixed list, words from multiple categories were presented. In a blocked list, words from a single category were presented. The number in the subscripts indicates the list number (e.g., L1 = List 1, L2 = List 2)
Proactive interference and forward testing effect
Proactive interference plays an important role in student learning. Suppose that some students are taking Spanish and French courses on the same day. During the Spanish class, they learned that the word amor means love. A few hours later, during the French class, they learned that amour (with a “u”) means love. Later, if students are asked to write love in French, they might recall amor instead of amour. This example shows one scenario in which prior learning (of Spanish) interferes with new learning (of French), i.e., proactive interference (PI).
PI is widely considered a major source of forgetting (for reviews, see Crowder, 1976; Kliegl & Bäuml, 2021). In one of the earliest demonstrations of the forward testing effect, Tulving and Watkins (1974) considered the insulation of PI as a causal factor for the effect. In their study, participants studied two lists (the first containing A-B associations and the second containing A-C associations) and then attempted to recall words from both lists in a modified-modified free recall (MMFR) test. The key manipulation was whether participants were tested on the first, A-B list before they learned the second, A-C list (i.e., the presence or absence of interpolated testing), and their results showed that immediate testing of List 1 greatly increased recall of List 2. Tulving and Watkins interpreted this finding as showing that “the interpolation of an activity requiring explicit retrieval of stored information seems to insulate the A-B list from the A-C list in a way that removes the former as an interfering component in the learning of the latter”(p. 192).
Response competition is a mechanism through which PI impairs subsequent learning. Specifically, compared to a situation in which participants learn only a target list, prior learning increases the number of retrieval candidates for a given cue, and these items compete with each other and interfere with the recall of the target item(s). A broadly accepted way to measure the level of response competition and PI is through intrusions of items from the nontarget lists (s) during recall testing (Bäuml & Kliegl, 2013; Chan, Meissner, et al., 2018; Crowder, 1976; Darley & Murdock, 1971; Pierce et al., 2017; Postman & Hasher, 1972; Underwood, 1975; Yang et al., 2018; Zaromb & Roediger, 2010).

Cistanche deserticola experiment
As evidence that interpolated testing promotes new learning by inoculating learners against the influence of PI, numerous studies have shown that interpolated testing can reduce intrusions from prior learning. For example, Darley and Murdock (1971) found that learners are far less likely to commit intrusions of previously tested items than previously nontested ones. In their study, participants studied ten lists of words while expecting a cumulative test. After studying each list, participants took an immediate recall test for half of the lists and counted forward by threes for the other lists. When examining the frequency of prior-list intrusions during the immediate-recall tests, Darley and Murdock found that most of the intrusions (i.e., 81%) belonged to the nontested lists, which shows that tested items might be less susceptible to response competition than nontested items. This funding can be interpreted as evidence for either an early selection advantage (i.e., the tested items are less likely to come to mind when participants are attempting to retrieve the just studied, nontested items) or a late correction advantage (i.e., participants can readily withhold a retrieved item if they remember that the item had been recalled). Either way, both possibilities suggest that the benefit of testing might stem from a reduction in PI.
Release‑from‑PI account
In a study that revived interest in the phenomenon of the forward testing effect, Szpunar et al. (2008) showed that after studying five inter-related word lists, participants who had been tested on Lists 1–4 committed fewer intrusions during List 5 recall than participants who had not been tested on Lists 1–4. Moreover, the tested participants recalled more List 5 items than the control participants. Based on these results, Szpunar and colleagues proposed the release-fromPI account – such that testing inoculated participants against the buildup of PI, and this was in turn responsible for the learning benefits of List 5 amongst the tested participants. Other studies have also reported similar findings that intermittent testing reduced intrusions and boosted new learning (Chan, Manley, et al., 2018; Weinstein et al., 2011, 2014; Yang et al., 2018).
Because PI is theoretically neutral, this release-fromPI account leaves open the candidate mechanisms through which interpolated testing enhances new learning. PI can degrade subsequent learning via processes that operate either at retrieval or at encoding. In this paper, we focused on the retrieval-based explanation of PI for the forward testing effect. Szpunar et al. (2008) ascribed the forward testing effect to an enhancement in list discrimination – the idea is that testing boosts retention of the tested items, which in turn allowed participants to better discriminate list membership of the studied items and to avoid intrusions (see also Chan & McDermott, 2007; Jang & Huber, 2008; Postman et al., 1968; Postman & Keppel, 1977) – consequently, the original conceptualization of the release-from-PI account focused on retrieval processes, and we did the same here.
If interpolated testing enhances new learning because it reduces PI, then a forward testing benefit for correct recall should be accompanied by a reduction of intrusions. Consistent with this idea, in a recent meta-analysis (Chan, Meissner, et al., 2018), out of 38 studies in which a forward testing effect was observed and intrusions were reported, 36 found that interpolated testing reduced intrusions, thus giving weight to the release-from-PI account. However, this reasoning is based on a correlation (indeed, across studies, there is a robust correlation between the magnitude of the forward testing effect in correct recall and intrusions in Chan et al.’s data, r = -.60, p < .001), and few studies have provided a direct test of the release-fromPI account. Specifically, if interpolated testing potentiates new learning by reducing the influence of PI, then the magnitude of the forward testing effect should depend on the level of PI experienced by learners – that is, the benefits of testing should be more pronounced when learners are expected to experience greater PI.
Szpunar et al. (2008) provided a direct test of the release-from-PI account in Experiment 2. Specifically, they manipulated PI by varying the number of non-tested lists that participants had to study before the target list (see Fig. S1 of the Online Supplemental Material (OSM) at https://osf.io/wghbc/ for their design), and PI was designed to build up across lists. Participants in the test-every-list condition should be less susceptible to the negative effects of PI. In contrast, those in the control conditions were tested on only one list among Lists 2–5 (e.g., one group was tested on only List 2, one group on List 3, and so on), and the level of PI was expected to be proportional to the number of nontested lists before testing. Consistent with the assumption that prior testing reduces the buildup of PI when comparing the number of intrusions produced by participants in the control conditions to participants in the test-every-list condition, Szpunar et al. found that intrusions rose as the number of nontested lists increased, and this rise in intrusions was accompanied by an increase in the forward testing effect for correct recall.
Although Szpunar et al.’s (2008) results supported a release-from-PI account, manipulating the level of PI by the number of study lists raises an interpretation problem because this PI manipulation is confounded with test frequency. When comparing participants in the test-every-list condition to participants in the control condition, the benefits of interpolated testing were greater when participants in the control condition had studied more nontested lists. One interpretation of this finding is that interpolated testing reduces the influence of response competition during recall, which leads to a forward testing effect. However, other explanations are equally viable. For example, semantic clustering in the recall, which serves as an indication of memory organization, increases with test frequency (Chan et al., 2020; Zaromb & Roediger, 2010), and this funding has typically been attributed to strategy optimization. Further, increasing the number of interpolated tests can promote attentive encoding because of the frequent occurrence of testing signals to participants that they will likely be tested again shortly – such that participants experience a rise in test expectancy. Critically, neither of these accounts requires release-from-PI as an explanatory mechanism. In our review of the literature, all studies (except the Nunes & Weinstein, 2012, study described below) that have provided support for the release-from-PI account suffer from similar issues, so extant data do not provide unequivocal support for the release-from-PI account.

Cistanche powder
PI manipulation with content change method
In the present study, we aimed to isolate our PI manipulation from test frequency. To this end, we varied PI using the content change method in the traditional Brown-Peterson paradigm (Brown, 1958; Peterson & Peterson, 1959). Briefly, PI is accumulated when participants study a series of similar items (e.g., words from the same category, such as fruits), but when the content of the study item changes (e.g., to words in a different category, such as animals), learners experience a release from proactive interference because they can use the change in semantic category to constrain their memory search during retrieval (e.g., by attempting to retrieve only animals, Wickens et al., 1963, 1981).
The content change paradigm is ideally suited to investigate the release-from-PI account for two reasons. First, this method allowed us to manipulate the level of PI that participants experienced without altering the amount of information that they had to study. More importantly, the content change method provides a way to vary PI in categorized lists without changing the words that participants study across conditions. That is, all participants studied the same words, but PI was manipulated via presentation order. Second, because our theoretical objective is to investigate whether release-fromPI, as a retrieval-based phenomenon, is responsible for the occurrence of the forward testing effect, it is particularly promising that content change has been identified as a procedure that specifically affects the retrieval component of PI effects (Engle, 1975; Gardiner et al., 1972; Kliegl & Bäuml, 2021; Wixted & Rohrer, 1993).
Using the content change method, Nunes and Weinstein (2012) reported, to our knowledge, the only causal evidence for the release-from-PI account. In their experiments, participants studied five sets of Deese-Roediger-McDermott (DRM, Roediger & McDermott, 1995) words. In Experiment 1, words that belonged to each DRM theme were spread across five study lists, with each list containing three associates from six DRM themes. The tested participants took a test after every list, whereas the control participants solved math problems after Lists 1–4 and took a test only after List 5. Unsurprisingly, a forward testing effect was observed (correct = 0.77, intrusion = 1.29).
Although the results from Nunes and Weinstein’s (2012) Experiment 1 extended the forward testing effect from cat-categorized lists to associative DRM lists (Weinstein et al., 2010), the critical findings came from Experiment 2. Here, the words from each DRM theme occupied a single study list between Lists 1 and 4. The fifth list was similar to Experiment 1, such that it included three words from each of the four studied DRM themes. The key finding is that participants in the tested and control conditions recalled virtually identical proportions of the target list words – there was no forward testing effect. Furthermore, participants in both conditions produced very few intrusions. According to Nunes and Weinstein, when each list contained a single DRM theme, the switch in semantic content across lists released participants from the negative influence of PI. If interpolated testing potentiates new learning because it releases learners from PI, and if the content switch minimizes PI buildup in the control condition, then the benefits of interpolated testing become irrelevant and the forward testing effect should be eliminated. This was what happened in Nunes and Weinstein’s Experiment 2, thus providing strong support for the release-from-PI account.
To our knowledge, Nunes and Weinstein (2012) is the only study that examined the release-from-PI account using a content change method. However, several characteristics of their experiments made it difficult to draw strong conclusions. First, because Nunes and Weinstein did not manipulate presentation order in the same experiment, a cross-experimental comparison between Experiments 1 and 2 is required to determine whether a reduction in PI weakened the forward testing effect. Second, the main conclusion of Nunes and Weinstein’s study hinges on a null effect (i.e., no forward testing effect) in Experiment 2, but that experiment had only 35 participants and the critical comparison was between subjects.1 Lastly, the materials differed across their Experiments 1 and 2 (e.g., each list comprised 18 words in Experiment 1 but only 12 words in Experiment 2), and Nunes and Weinstein did not fully counterbalance their words across study lists, such as that the words that appeared in the target list never appeared in other lists. The lack of counterbalancing limited the generalizability of the results – it is possible that instead of the Low-PI manipulation eliminating the forward testing effect in Experiment 2, it was the specific words that appeared in List 5 that were, for some reason, not amenable to demonstrating a forward testing effect. Although this possibility might seem remote, it deserves attention because Nunes and Weinstein’s target list featured completely different words across their Experiments 1 and 2, which compounded the difficulty in drawing conclusions based on this cross-experimental comparison.
The current study
In our Experiments 1 and 2, we manipulated PI and intervening tasks in a factorial, between-subjects design (see Fig. 1). Specifically, PI was manipulated by presenting words from the same category across lists (High-PI) or within lists (LowPI), and the intervening task was manipulated by giving participants an interpolated test or interpolated restudy phase after each non-target list. All participants took a test on the final, target list. In the High-PI condition, participants studied mixed lists comprising words from several categories, with different words from those same categories appearing in every list, thereby providing an opportunity for PI to accumulate. In the Low-PI condition, participants studied blocked lists containing words from a unique category for Lists 1–4, and the shift in the category across lists was designed to release the buildup of PI. If the forward testing effect emerges because testing insulates the buildup of proactive interference, we should observe a smaller forward testing effect in the Low-PI than in the High-PI condition. Based on the data from Nunes and Weinstein’s Experiment 2 (2012), one might expect that the forward testing effect would be eliminated in the Low-PI condition because the content change manipulation should negate the PI-reducing advantage of interpolated testing. To preview, contrary to this prediction, we found a robust forward testing effect in both the High-PI and Low-PI conditions and the magnitude of the effect was unaffected by the PI manipulation. Thus, we aimed to directly replicate Nunes and Weinstein’s Experiment 2 in our Experiments 3 and 4.
Experiment 1
Method
Design and participants
Participants were 119 undergraduate students from Iowa State University who completed the experiment for course credits. There were 29 participants in the High-PI, Test condition, 30 in the High-PI, Restudy condition, 32 in the LowPI, Test condition, and 30 in the Low-PI, Restudy condition. Data from 14 participants were not analyzed because three (i.e., one in the High-PI, Test condition; two in the Low-PI, Restudy condition) did not follow instructions, and a programming error made the data for 11 participants unusable (i.e., six in the Low-PI, Test condition; five in the Low-PI, Restudy condition). The final sample included data from 107 participants, with 28 in the High-PI, Test condition, 30 in the High-PI, Restudy condition, 26 in the Low-PI, Test condition, and 23 in the Low-PI, Restudy condition. We conducted a sensitivity power analysis using G*Power (Faul et al., 2007) to compute the effect size that can be detected based on our collected sample size. We used the Low-PI condition (combined n = 49) for this analysis because this condition had fewer participants than the High-PI condition. The minimum effect size that can be detected with .80 power was d = 0.82, which was comparable to the meta-analytic effect size (d = 0.75, Chan, Meissner, et al., 2018). As will become evident, inadequate statistical power was never an issue in our experiments, as our observed effect sizes always exceeded this minimum detectable effect size (at .80 power) regardless of whether participants were in the High- or Low-PI condition.
Materials and procedure
For Experiments 1 and 2, we opted to employ materials that are similar to the majority of studies that have demonstrated a forward testing effect. To that end, we selected words from four categories (i.e., fruits, animals, body parts, and sports) that each had at least 20 exemplars in the Van Overschelde, Rawson, and Dunlosky norms (2004). The average typicality ratings did not difer across categories (Mfruits = .18, Manimals = .18, Mbodyparts = .18, Msports = .18), F(3, 76) = 0.01, p = .999, B01 = 14.45.
Participants studied five 16-word lists. In the High-PI condition, each list contained four exemplars from each of the four categories. The average typicality ratings did not difer across lists (range = .17 – .18), F(4, 75) = 0.01, p = .999, B01 = 19.44. All words within a list were presented in a pseudo-random order, with items from the same category never presented consecutively. In the Low-PI condition, each of the first four lists contained words from a single category presented in a random order (e.g., List 1 might contain only fruits, List 2 might contain only animals). Most importantly, List 5 was identical for participants in both the Low-PI and High-PI conditions, and it contained four new words from each of the four studied categories. Therefore, PI was manipulated by the structure of Lists 1–4. See the top of Appendix A for the full set of materials.
Participants in groups of up to eight completed the experiment at computers separated by dividers. Participants were informed that they would study several lists of 16 words, with each word presented twice. They were also told that after studying each list, they would solve some math problems, and then they would either be asked to recall the just presented list or to restudy the list. Critically, participants were informed that the computer randomly determined whether a test will be given following each list, so that they might be “asked to recall words after every list, after only one or two lists, or never at all.” In actuality, participants either received a test for every list (test condition) or they were only tested on List 5 (restudy condition). To ensure that participants would study each list intentionally, they were further told that there was a cumulative recall test for all lists regardless of how many tests they received.
At the start of each list, a 2-s prompt informed participants of the list number (e.g., “This is Word List 1”), which was followed by a 1-s fixation cross. Next, the words were presented at 4 s apiece with a 500-ms interstimulus interval. Each list was presented twice without a break and the second presentation order was different from the first. The assignment of words to lists was counterbalanced across participants. After studying each list, participants completed ten algebra problems at 6 s each to clear short-term memory and then received the recall or restudy prompt. Participants in the test condition had 60 s to recall the words from the most recent list, whereas those in the restudy condition saw the same words again but in a new random order. All participants were given 60 s to recall the words from List 5. A forward testing effect is found if participants in the test condition recall more List 5 words than participants in the restudy condition. Participants then completed a Reading Span task (Conway et al., 2005) as a filler activity for 20 min and then took the cumulative recall test on Lists 1–5. The RSPAN and cumulative recall tasks were included to ensure that the experiment lasted approximately 40 min for course credit-granting purposes and was not scored. Finally, participants completed a short survey on demographics and were debriefed.
Table 1 Nontarget, interpolated test performance across Experiments 1–4

Note. Standard deviations are in parentheses
Results
For completeness’ sake, proportions of correct recall for the nontarget lists are reported in Table 1. But the focus of our research goal was target list recall. For all experiments, we first report the frequency of intrusions during target list recall, which provides a measure of PI. We defined intrusions as the recall of words from any nontarget lists. Extralist intrusions were not considered because they do not represent proactive interference from prior list learning. If the PI manipulation was successful, there should be fewer intrusions in the Low-PI condition than in the High-PI condition.
Across all experiments, participants whose intrusions were greater than 2 SD from the mean were considered outliers. Given that we explicitly told participants to recall words from only the most recent list, those who committed a very high number of intrusions might not have followed the instructions. For example, if a participant recalled eight furnature words right after studying the animal list, we think it is likely that this participant did not follow the instructions. To be comprehensive, we conducted all of our analyses with and without outliers. We report results without outliers as the default but also highlight when the inclusion of the outliers led to different conclusions (to preview, only one analysis in Experiment 3 produced a different result). In addition, the full set of data is on the Open Science Framework at https://osf.io/wghbc/ and Figs. S2 and S3 of the OSM present results including outliers. None of the experiments were pre-registered.
After reporting intrusion data, we report the target list correct recall. We examined whether the PI level affected the magnitude of the forward testing effect. If reduced PI plays a causal role, we should observe a smaller forward testing effect in the Low-PI condition than in the High-PI condition. We also examined the correlation between intrusions and correct recall in the target list. We expect that the two variables should be negatively related such that participants with fewer intrusions should show greater correct recall — note, however, this correlation does not imply causation.
For statistical inferences, we used two-tailed tests with ⍺ = .05. We also report Bayes factors. When a result is significant, we report B10, which indicates support for the alternative hypothesis; when a result is not significant, we report B01, which indicates support for the null hypothesis. We took this approach so that a larger Bayes factor always indicates more support for the effect (null or otherwise) under consideration. All Bayesian analyses were performed using the default priors (JASP team, 2020).
Intrusions during List 5 recall
Across the entire sample, the mean frequency of intrusions was 1.03 and the standard deviation was 1.84. Accordingly, data from five participants who committed more than 4.71 intrusions (i.e., 1.03 + 2*1.84 = 4.71) were considered outliers and excluded from analyses. The left panel of Fig. 2 shows the intrusion data. We conducted a 2 (intervening task: test vs. restudy) × 2 (PI level: high vs. low) ANOVA with the frequency of intrusions during List 5 recall as the dependent variable. There was a significant main effect of intervening task, F(1,98) = 34.84, p < .001, d = 1.14, B10 = 108753.94, such that the restudied participants (M = 1.31) committed more intrusions than the tested participants (M = 0.19). It is not surprising that interpolated retrieval reduces intrusions, as the effect has been widely demonstrated (Chan, Manley, et al., 2018; Szpunar et al., 2008; Weinstein et al., 2011). More important for present purposes, the main effect of the PI level was significant, F(1,98) = 12.27, p = .002, d = 0.59, B10 = 9.36, with participants in the High-PI condition (M = 1.02) committing nearly triple the intrusions relative to participants in the Low-PI condition (M = 0.38), which shows that our manipulation of PI was successful.

Fig. 2 Frequency of intrusions during List 5 recall (left) and proportion of List 5 correct recall (right) as a function of intervening task and PI level in Experiment 1. Error bars display descriptive .95 confence intervals. Each dot represents data from an individual participant. Jitter was introduced to disperse data points horizontally to increase visibility
The interaction between the intervening task and PI level was also significant, F(1,98) = 5.61, p = .020, ηp 2 = .04, B10 = 2.73. Specifically, the tested participants produced minimal intrusions regardless of PI level (MHigh-PI = 0.29 vs. MLow-PI = 0.08), t(52) = 1.50, p = .139, d = 0.41, B01 = 1.45, but the restudied participants produced significantly fewer intrusions if they were in the Low-PI condition than in the HighPI condition (MHigh-PI = 1.81 vs. MLow-PI = 0.73), t(46) = 3.02, p = .004, d = 0.87, B01 = 9.70.
List 5 correct recall
The right panel of Fig. 2 shows the List 5 correct recall data. The same 2 × 2 ANOVA showed a main efect for intervening task, F(1,98) = 15.95, p < .001, d = 0.81, B10 = 239.75, which shows a forward testing effect overall (Mtest = .51 vs. Mrestudy = .34). PI manipulation did not affect correct recall (MHigh-PI = .43 vs. MLow-PI = .44), F(1,98) = 0.05, p = .826, d = 0.05, B01 = 4.64. Of greater interest, an interaction effect was also not significant, F(1,98) = 0.78, p = .379, ηp 2 = .01, B01 = 2.95. Contrary to the expectation that reducing PI would also reduce the magnitude of the forward testing effect, the test group vastly outperformed the restudy group in both the Low-PI condition (Mtest = .50 vs. Mrestudy = .36), t(46) = 2.78, p = .008, d = 0.81, B10 = 5.88, and the HighPI condition (Mtest = .52 vs. Mrestudy = .32), t(52) = 3.00, p < .001, d = 0.82, B10 = 9.76. The magnitude of the forward testing effect (i.e., the effect size) was nearly identical regardless of the PI condition. The robust forward testing effect in the Low-PI condition was clearly in contradiction to Nunes and Weinstein’s Experiment 2 data. Although the manipulation of PI did not affect the forward testing effect, as predicted by the release-from-PI account, we observed a substantial negative correlation between intrusions and correct recall, r(100) = -.45, p < .001, B10 = 8154.72, such that individuals who recalled more correct items also produced fewer intrusions (see Fig. 3). Closer examinations revealed that this correlation was largely driven by participants in the restudy condition, r(46) = -.45, p < .001, B10 = 25.82, as participants in the test condition had a near-floor level of intrusions, thereby restricting the range of scores on this variable and minimizing the likelihood of finding any correlations, r(52) = -.15, p = .284, B01 = 3.37.
Discussion
The purpose of Experiment 1 was to examine whether the magnitude of the forward testing effect is susceptible to differences in proactive inference. As expected, participants in the High-PI condition produced more intrusions than their Low-PI counterparts. Further, the effects of PI were alleviated by interpolated testing. In sum, the intrusion data validated our PI manipulation. Despite these results, neither proportion of correct recall nor the magnitude of the forward testing effect was influenced by the PI manipulation, which suggests that release from proactive interference might be a corollary, rather than the cause, of the forward testing effect. An additional intriguing result was the robust negative correlation between correct recall and intrusions, even though only the latter was affected by PI. At face value, this correlation is consistent with the release-from-PI account because participants who experience less PI showed greater correct recall, so testing might enhance new learning because it reduces the buildup of PI. We consider the implications of this correlation more fully in the General discussion but suffice it to say that our data indicate that the negative association between correct recall and intrusion does not imply causation. Rather, the coupling of these variables might be driven by a third variable, such as increased attention or enhanced memory organization due to testing (Chan et al., 2020).

Fig. 3 A scatter plot displaying the relationship between the frequency of intrusions during List 5 recall and the proportion of List 5 correct recall in Experiment 1. On an individual basis, frequencies of intrusions are always integers, so there was substantial overlap in the data points. To improve the visibility of data density, we introduced jitter to disperse the data points horizontally
Although the data in Experiment 1 show that reduction in PI might not play a causal role in the forward testing effect, it remained possible that we observed a forward testing effect in the Low-PI condition because our manipulation of PI was too weak. Specifically, participants in the Low-PI, Restudy condition might have still experienced considerably greater PI than their tested counterparts, because List 5 contained words from the categories that appeared in Lists 1–4. This possibility can be examined by comparing the intrusion data between the tested and the restudied participants in the LowPI conditions. Even when PI was intended to be low, the tested participants (M = 0.08) committed far fewer intrusions than the restudied participants (M = 0.73), t(46) = 3.10, p = .003, d = 0.90, B10 = 11.58. This finding shows that our manipulation did not render the PI-reducing power of testing irrelevant, which might explain why we found a forward testing effect in the Low-PI condition.
In Experiment 2, we aimed to further reduce PI for participants in the Low PI condition. To this end, every list, including the target list, included only words from a single category. The logic is to minimize the influence of PI on target list learning and recall by eliminating the possibility that words in the target list would remind participants of the previously studied lists.
Experiment 2
Method
Design and participants
PI level (high vs. low) and Intervening task (test vs. restudy) were manipulated between subjects. Participants were 120 undergraduate students from Iowa State University who participated in course credits. Data from five participants were not analyzed because three did not follow instructions and two involved a computer error. The final sample thus included data from 115 participants, with 31 in the High-PI, Restudy condition, 28 in the High-PI, Test condition, 28 in the Low-PI, Test condition, and 28 in the Low-PI, Restudy condition. For the sensitivity power analysis, we used the sample size of the Low-PI condition, which had fewer participants. The analysis showed that we could detect an effect size of d = .76 with .80 power.
Materials and procedure
The procedure of Experiment 2 was identical to Experiment 1 with the following exceptions. First, in the High-PI condition, three words from each category were spread across four lists, such that each list contained words from the same four categories. In the Low-PI condition, every list, including the target list, contained words from a single category. Note that all participants studied the same words from the four categories, with the only difference across conditions being the presentation order of the words across lists. The full set of materials is presented in Appendix A. Second, words from the categories of fruits, animals, body parts, and furniture were used, and they had the same average typicality ratings (Mfruits = .18, Manimals = .18, Mbodyparts = .18, Mfurniture = .18), F(3, 60) = 0.00, p = 1.00, B01 = 11.71. Lastly, participants studied four lists of 16 words each instead of five lists. Experiment 1 was conducted after Experiment 2 chronology cycle. But the studies were presented in the current order for exposition purposes. The logic was to first present the study with the weaker PI manipulation (i.e., Experiment 1), and then present the study with the stronger PI manipulation (i.e., Experiment 2). The number of study lists and materials differed across the two experiments because we needed more words in Experiment 1 to create List 5 that included the category words from Lists 1–4 while maintaining the list length at 16 words, and Furniture does not have enough exemplars for this purpose.
Results
Intrusions during List 4 recall
The outlier cut-off for this experiment was 4.79 intrusions (M = 0.85, SD = 1.97), and data from seven participants were excluded. The left panel of Fig. 4 shows the frequency of intrusions during List 4 recall. Because participants in the Low-PI, Test condition did not commit any intrusions at all, it was not possible to conduct a 2 (test vs. restudy) × 2 (High-PI vs. Low-PI) ANOVA. Thus, we conducted a t-test between the High- and Low-PI conditions to examine the influence of PI manipulation on overall intrusion rates. As expected, participants in the High-PI condition (M = 0.71) produced more intrusions than those in the Low-PI condition (M = 0.11), t(106) = 3.46, p < .001, d = 0.67, B10 = 36.67. Indeed, the frequency of intrusions in the Low-PI condition did not differ from zero, t(52) = 1.43, p = .159, d = 0.20, B01 = 2.57. The fact that intrusions were almost entirely removed in the Low-PI condition (M = 0.24 for participants in the restudy condition, with 23 out of 25 participants producing no intrusions, and the remaining two producing three intrusions each; see Fig. 4) showed that our PI manipulation in Experiment 2 had the intended effect. Lastly, overall, the restudied participants (M = 0.80) committed more intrusions than the tested participants (M = 0.09), t(106) = 4.24, p < .001, d = 0.82, B10 = 425.83.

Fig. 4 Frequency of intrusions during List 4 recall (left) and proportion of List 4 correct recall (right) as a function of intervening task and PI level in Experiment 2. Error bars display descriptive .95 confidence intervals. Jitter was introduced to disperse data points horizontally to increase visibility
List 4 correct recall
The right panel of Fig. 4 shows the proportion of List 4 correct recall. Given the success of our PI manipulation, one should expect that the forward testing effect would be substantially weaker (and perhaps altogether absent) in the Low-PI condition than in the High-PI condition. However, similar to Experiment 1, the interaction between PI level and interim activity was again not significant, F(1,104) = 0.82, p = .367, ηp 2 = .01, B01 = 2.82. Replicating the results of Experiment 1, reducing PI did not diminish the forward testing effect. Specifically, the tested participants outperformed the restudied participants in both the Low-PI condition (Mtest = .68, Mrestudy = .45), t(51) = 3.98, p < .001, d = 1.09, B10 = 111.69, and the High-PI condition (Mtest = .61, Mrestudy = .30), t(53) = 4.68, p < .001, d = 1.20, B10 = 894.56. The powerful forward testing effect in the Low-PI condition is especially notable given that participants in this condition produced hardly any intrusions at all. Despite our PI manipulation having little influence on the forward testing effect, there was once again a strong negative correlation between correct recall and intrusions (see Fig. 5), r(106) = -.44, p < .001, B10 = 8395.88, and this effect was again driven by participants in the restudy condition, r(48) = -.41, p = .003, B10 = 12.06, rather than those in the test condition, r(56) = -.14, p = .296, B01 = 3.58.
Discussion
In this experiment, we attempted to further reduce PI for participants in the Low-PI condition. To this end, we had the Low-PI group study words from a single category for all lists, which reduced intrusions to near-foor levels for participants in the Low-PI condition. Despite this result, the forward testing effect remained robust, and its magnitude was unaffected by the PI manipulation.
Taken together, the findings of our experiments differed from those reported by Nunes and Weinstein (2012). As stated in the Introduction, the combination of Experiments 1 and 2 is comparable to our HighPI and Low-PI conditions. Specifically, participants in their Experiment 1 studied several DRM themes in each list, which is similar to our High-PI condition, and in their Experiment 2, participants studied individual DRM themes in each list, which corresponds to our Low-PI condition. A forward testing effect was found in Experiment 1, but the effect was absent in Experiment 2.

Fig. 5 A scatter plot displaying the relationship between the frequency of intrusions during List 4 recall and the proportion of List 4 correct recall in Experiment 2. To improve the visibility of data density, we introduced jitter to disperse the data points horizontally
Why do the current study and Nunes and Weinstein (2012) show contrasting results? Although the design of our experiments was modeled after Nunes and Weinstein, several methodological differences remained. Perhaps the most obvious was the materials. Namely, we used categored lists, whereas Nunes and Weinstein used DRM lists, with the former constructed based on category norms and the latter constructed based on associative norms. Another methodological difference between the studies was the comparison task. In both studies, all participants completed 1 min of math problems after studying each word list. Following this, participants in the control condition restudied the words in our experiments, whereas they solved additional math problems in Nunes and Weinstein’s experiments.
Despite these methodological differences, we could not conceive of a compelling explanation that would account for the vastly different results. In particular, it is difficult to envision these methodological differences having different effects on the magnitude of the forward testing effect depending on whether participants are expected to experience higher or lower PI. To our knowledge, no existing theoretical explanations (Chan, Meissner, et al., 2018) would lead one to predict that manipulation of PI would affect the forward testing effect with associative DRM words but not with categorized words (or with using math problems as the control task but not with restudy as the control task). Consequently, we aimed to directly replicate Nunes and Weinstein’s Experiment 2.






