Retrieval-Enhanced Suggestibility in Forensic Contexts: A meta-analysis of practice effects on eyewitness memory contamination *
Sugestibilidad potenciada por recuperación en contextos forenses: metaanálisis de efectos prácticos en la contaminación mnésica testifical
Retrieval-Enhanced Suggestibility in Forensic Contexts: A meta-analysis of practice effects on eyewitness memory contamination *
Universitas Psychologica, vol. 25, 2026
Pontificia Universidad Javeriana
Alberto Barea Vera a alberto.barea@ucavila.es
Universidad Católica de Ávila, España
Received: 05 April 2026
Accepted: 02 july 2026
Abstract: Retrieval practice can paradoxically increase vulnerability to subsequent misinformation, a phenomenon termed retrieval-enhanced suggestibility (RES), with implications for forensic interviewing protocols involving multiple recall attempts. This meta-analysis examines RES in forensic-relevant contexts and identifies boundary conditions moderating the effect. A systematic search identified 26 studies (k = 72 comparisons, N = 4 218) in simulated forensic contexts. Random-effects meta-analyses compared misinformation susceptibility following retrieval practice versus no practice, with moderator analyses (test format, misinformation source, detail type, retention interval, warning). Retrieval practice was associated with increased misinformation acceptance, d = 0.34, 95 % CI [0.23, 0.45], p < 0.001, I² = 56.4 %. Recognition-format tests showed larger RES (d = 0.66) than free recall (d = 0.17); narrative misinformation showed stronger RES (d = 0.48) than decontextualized question-based misinformation (d = -0.19); explicit warnings were associated with no detectable RES (d = -0.07, p = 0.41). In forensic-relevant experimental contexts, initial retrieval practice was associated with greater misinformation vulnerability, whereas free-recall formats, temporal separation, and explicit warnings were associated with reduced or absent effects. These findings inform the evidence base for forensic interview protocols.
Keywords:retrieval-enhanced suggestibility, eyewitness memory, misinformation effect, forensic interviewing, testing effect, memory contamination, investigative protocols.
Resumen: La práctica de recuperación puede paradójicamente aumentar la vulnerabilidad a la desinformación posterior, fenómeno denominado sugestionabilidad potenciada por la recuperación (RES), con implicaciones para los protocolos de entrevista forense con múltiples intentos de recuerdo. Este meta-análisis examina la RES en contextos forenses e identifica las condiciones que moderan el efecto. Una búsqueda sistemática identificó 26 estudios (k = 72 comparaciones, N = 4 218) en contextos forenses simulados. Meta-análisis de efectos aleatorios compararon la susceptibilidad a la desinformación tras la práctica de recuperación frente a su ausencia, con análisis de moderadores (formato de prueba, fuente de la desinformación, tipo de detalle, intervalo de retención y advertencia). La práctica de recuperación se asoció con mayor aceptación de desinformación, d = 0.34, IC 95 % [0.23, 0.45], p < 0.001, I² = 56.4 %. Las pruebas de reconocimiento mostraron mayor RES (d = 0,66) que el recuerdo libre (d = 0.17); la desinformación narrativa, mayor RES (d = 0.48) que la basada en preguntas sin contexto (d = -0.19); las advertencias explícitas se asociaron con ausencia de RES detectable (d = -0.07, p = 0.41). La práctica de recuperación inicial se asoció con mayor vulnerabilidad a la desinformación, mientras que el recuerdo libre, la separación temporal y las advertencias se asociaron con efectos reducidos o ausentes, lo que contribuye a la base de evidencia para los protocolos de entrevista forense.
Palabras clave: sugestionabilidad potenciada por la recuperación, memoria del testigo ocular, efecto de desinformación, entrevista forense, efecto de testeo, contaminación de la memoria, protocolos investigativos.
The “testing effect” –the finding that retrieval practice enhances long-term memory retention more effectively than passive re-study–represents one of the most robust phenomena in cognitive psychology (Roediger & Karpicke, 2006; Rowland, 2014). Across hundreds of studies, initial memory testing strengthens subsequent recall, with effect sizes typically ranging from d = 0.50 to d = 1.50 (Adesope et al., 2017). This robust finding has generated enthusiasm for applying retrieval practice to forensic interviewing protocols designed to maximize eyewitness memory accuracy (Wells et al., 2020).
However, an emerging literature documents a troubling paradox: under certain conditions, initial retrieval attempts can actually increase vulnerability to subsequent misinformation –a phenomenon termed retrieval-enhanced suggestibility (RES)– (Chan et al., 2009; Gordon et al., 2015). In RES paradigms, participants who complete initial memory tests subsequently show greater acceptance of misleading post-event information compared to non-tested controls. This effect directly contradicts the beneficial testing effect and raises urgent questions about optimal forensic interviewing protocols.
These concerns are equally salient in the Spanish-speaking forensic tradition, where the psychology of testimony has produced an extensive body of work on the fragility and reconstructive nature of eyewitness memory and on the contaminating effect of post-event suggestion. Spanish American researchers have long warned that the way witnesses are questioned can distort or even create memories, and have translated these findings into best-practice protocols for investigative interviewing (Diges, 2016; Manzanero, 2010; Manzanero & González, 2013; Silva et al., 2016). This literature converges with the international evidence in emphasizing that early, uncontaminated free-recall accounts must be protected from repeated or suggestive retrieval, situating the present meta-analysis within a shared Ibero-American and Anglo-American concern for the integrity of testimonial evidence.
The retrieval-enhanced suggestibility phenomenon
Three principal theoretical mechanisms explain RES. The test-potentiated learning (TPL) account holds that initial testing reduces proactive interference, freeing attention toward subsequently encountered misinformation content (Chan et al., 2009; Gordon & Thomas, 2013), consistent with reading-time evidence showing tested participants allocate more attention to misleading narrative sentences (Gordon et al., 2015, 2020). The memory reconsolidation account proposes that retrieval destabilizes the original memory trace, rendering it labile within a limited reconsolidation window; Chan and LaPaglia (2013) demonstrated that misinformation introduced 20 minutes after retrieval disrupted original memories while a 48-hour delay did not, though RES persisting beyond this window (Chan & Langley, 2011) suggests reconsolidation is not the sole mechanism. The context processing account (LaPaglia & Chan, 2019) proposes that RES occurs specifically when misinformation reinstates contextual features of the original event: narratives reinstate event context, fostering gist-level integration of misinformation, whereas isolated questions without context prompt verbatim retrieval that discriminates original from post-event information.
Forensic context and practical importance
Standard investigative practice frequently involves multiple witness interviews separated by exposure to potentially contaminating information: leading questions, co-witness discussions, media coverage, and suggestive pre-trial preparation (Fisher & Geiselman, 1992; Gabbert et al., 2009). If initial interviews paradoxically increase vulnerability to such misinformation, standard protocols may inadvertently facilitate memory contamination, a concern amplified by documented problems with suggestive interviewing (Kassin et al., 2010), co-witness contamination in high-profile cases (Wells & Quinlivan, 2009), and mixed RES findings across studies (Chan et al., 2009; Huff et al., 2016; LaPaglia & Chan, 2019). A systematic meta-analytic integration is needed to establish overall RES magnitude, identify boundary conditions, and provide evidence-based investigative guidance.
The present meta-analysis advances this literature by (1) restricting inclusion to forensic-relevant stimuli (crimes, accidents, violent events), enhancing ecological validity; (2) systematically examining five theoretically important moderators that map directly onto manipulable features of investigative protocols; (3) synthesizing research published through 2025, substantially expanding the evidence base beyond prior narrative reviews; and (4) translating meta-analytic findings into specific, evidence-based recommendations for investigative interviewing.
Theoretical framework and hypotheses
Based on source monitoring (Johnson et al., 1993), reconsolidation (Chan & LaPaglia, 2013; Nader et al., 2000), test-potentiated learning (Chan et al., 2009), and selective attention (Anderson & Spellman, 1995) frameworks, we advance six specific hypotheses:
H1 (Overall RES): Initial retrieval practice will increase misinformation acceptance compared to no-practice controls.
H2 (Test Format): RES will follow the gradient recognition > cued recall > free recall, reflecting differential guessing demands.
H3 (Misinformation Source): Narrative-based misinformation will generate stronger RES than question-based misinformation due to context reinstatement.
H4 (Detail Type): RES will be stronger for peripheral than central details, as central details benefit more from retrieval strengthening.
H5 (Immediate vs. Delayed Misinformation): RES will be stronger when misinformation follows retrieval immediately, consistent with the reconsolidation window.
H6 (Warning Effects): Explicit warnings about potential misinformation will reduce or eliminate RES by enhancing source monitoring.
Method
Literature search strategy
A systematic search across PubMed, PsycINFO, Web of Science, Scopus, ProQuest Dissertations & Theses, and Google Scholar (accessed January 2025) used five complementary search string families: (1) “retrieval-enhanced suggestibility,” “reversed testing effect,” and “RES effect”; (2) “testing effect” or “retrieval practice” combined with “misinformation” or “false memory”; (3) “initial interview” or “repeated retrieval” combined with “eyewitness” and “misleading information”; (4) “test-induced priming” or “retrieval-induced facilitation” combined with “forensic”; and (5) targeted author searches for researchers active in this domain (Chan, LaPaglia, Gordon, Thomas, Bulevich, Wilford, Saunders, Pansky, Karanian, Wulff, Torrance). Supplementary searches included reference lists of identified articles and prior reviews (Chan & LaPaglia, 2013), forward citation searches for seminal RES articles (Chan et al., 2009; Gordon et al., 2015), conference programs from SARMAC, Psychonomic Society, EAPL, and AP-LS (2013–2024), and direct correspondence with researchers requesting unpublished data.
Inclusion and Exclusion Criteria
Studies were included if they (1) employed an experimental design with at least one retrieval-practice condition versus a no-practice control; (2) used forensic-relevant stimuli (crimes, accidents, or violent events); (3) included a misinformation phase following the retrieval manipulation; (4) assessed misinformation acceptance on a quantifiable final memory test; and (5) reported sufficient statistics to calculate effect sizes. Both peer-reviewed articles and dissertations/conference proceedings with full methodological detail were eligible. Studies were excluded if they used exclusively non-forensic stimuli (word lists, prose, neutral objects), lacked a misinformation phase or no-practice control, used clinical samples, or provided insufficient statistics despite author contact.
Study selection Process
The systematic search identified 312 potentially relevant records. After removing 89 duplicates, 223 records underwent title and abstract screening by two independent reviewers (κ = 0.91). Following abstract screening, 63 full-text articles were assessed; 37 were excluded (14 non-forensic stimuli, 9 lacked control conditions, 8 had no misinformation phase, 4 examined special populations, 2 insufficient statistics). This resulted in 26 eligible studies. Figure 1 presents the PRISMA flow diagram.

Note. Initial database searches (PubMed, PsycINFO, Web of Science, Scopus, ProQuest, Google Scholar) identified 312 records. After removing 89 duplicates, 223 records underwent title and abstract screening. Following screening, 63 full-text articles were assessed for eligibility; 37 were excluded (14 used exclusively non-forensic stimuli, 9 lacked control conditions, 8 had no misinformation phase, 4 examined special populations, 2 had insufficient statistics). The final meta-analysis included 26 studies contributing k = 72 effect size comparisons from N = 4,218 participants.
Data Extraction and Coding
Two independent coders extracted data using a standardized coding protocol. Coded variables included study characteristics (publication year, type, sample size), stimulus event characteristics (event type, modality), initial test characteristics (format, timing, number of tests), misinformation characteristics (source format, detail type, number of items), temporal intervals (test-to-misinformation delay, misinformation-to-final-test delay), warning manipulation, and final test characteristics. Inter-rater reliability was high: κ = 0.92 for categorical variables (range: 0.86–0.97), ICC (2,1) = 0.998 for continuous variables.
Effect Size Calculation
The primary effect size was Cohen’s d, calculated as the standardized mean difference in misinformation acceptance between retrieval-practice and no-practice conditions. Positive d values indicate RES effects (greater misinformation acceptance following retrieval practice); negative values indicate protective testing effects. Effect sizes were calculated from means and SDs when available, otherwise from t- or F-values, or conservatively estimated from p-values and sample sizes (Lipsey & Wilson, 2001). All effect sizes were corrected for small-sample bias using Hedges’ (1981). correction; we report Cohen’s d notation using bias-corrected values throughout. Studies reporting multiple relevant comparisons were coded as separate effect sizes with dependency accounted for in analyses.
Statistical analysis
All meta-analyses were conducted in R (version 4.3.2) using the metafor package (Viechtbauer, 2010) with random-effects models and restricted maximum likelihood (REML) estimation. For categorical moderators (test format, misinformation source, warning), we conducted subgroup analyses using mixed-effects models and tested between-group heterogeneity with Q_between. For the continuous moderator (temporal intervals), we used meta-regression: d_i = β₀ + β₁._i + ε_i. Dependent effect sizes were handled via robust variance estimation (RVE) with cluster-robust standard errors treating each study as a cluster (Hedges et al., 2010); sensitivity analyses compared RVE results to single-effect-per-study and within-study aggregation approaches, yielding consistent conclusions. Publication bias was assessed using funnel plot inspection, Egger’s regression test, trim-and-fill analysis, p-curve analysis (Simonsohn et al., 2014), and three-parameter selection models (Vevea & Woods, 2005). Sensitivity analyses examined influential studies via Cook’s distance and leave-one-out analysis, and tested robustness of results to restrictions by publication type and design type.
Results
Descriptive overview of included studies
The final meta-analysis included 26 studies published between 2002 and 2025 (median year: 2014), yielding k = 72 independent effect size estimates from N = 4,218 participants. Sample sizes ranged from 60 to 498 per study (median = 162). Twenty-three studies were peer-reviewed journal articles; three were dissertations or conference proceedings with full methodological detail. The majority (85.7 %) were from three primary research groups: Chan and colleagues (Iowa State University), Thomas and colleagues (Tufts University), and LaPaglia and colleagues. Most studies (85.7 %) used undergraduate samples; 62 % female on average. All studies used forensic-relevant stimuli, most frequently a terrorism/crime scenario from the television program 24 (57.1 % of studies), with additional event types including bank robbery scenarios (14.3 %), theft or perpetrator identification scenarios (9.5 %), and other criminal/violent events (19.1 %). Studies predominantly used between-subjects designs (71.4 %). Table 1 presents detailed characteristics of all included studies.

The three units of analysis are nested and should not be confused. The 26 studies are the independent published reports that met inclusion criteria. Because several studies reported more than one experiment, these yielded 32 independent experiment entries (i.e., separate samples). Within these experiments, the meta-analytic model was estimated over 72 individual effect-size comparisons, as a single experiment frequently contributed multiple comparisons (e.g., different retrieval formats, delays, or misinformation conditions measured on the same or on independent subsamples). Thus, 26 studies → 32 experiment entries → 72 effect-size comparisons. Dependencies among comparisons drawn from the same sample were handled through the robust variance estimation / multilevel approach described in the Method section. The full list of the 72 comparisons, together with the corresponding standardized mean differences (Cohen’s d) and their 95 % confidence intervals for each comparison, is provided as online Supplementary Material.
Overall retrieval-enhanced suggestibility effect (H1)
Random-effects meta-analysis across all 72 comparisons revealed a significant overall RES effect: d = 0.34, 95 % CI [0.23, 0.45], p < 0.001 (Figure 2). Retrieval-practice participants accepted an average of 45.4 % of misinformation items compared to 36.8 % in control conditions –an absolute increase of approximately 8.6 percentage points. The 95 % prediction interval [−0.04, 0.72] indicates that across populations the true RES effect ranges from near-zero to medium-large, justifying moderator analyses.

Note. Random-effects meta-analysis across all 72 comparisons yielded d = 0.34, 95 % CI [0.23, 0.45], p < 0.001. Each horizontal line represents one effect size comparison, with length proportional to the 95 % confidence interval and square marker area proportional to inverse variance weight. The diamond at the bottom represents the weighted mean effect size and its 95 % confidence interval. Positive values (right of zero) indicate greater misinformation acceptance in retrieval-practice conditions relative to no-practice controls (RES); negative values indicate protective testing effects. The vertical dashed line at zero represents no effect. Heterogeneity indices: Q (71) = 163.4, p < 0.001, I² = 56.4 %, τ² = 0.058.
Heterogeneity: Moderate-to-substantial heterogeneity was observed, Q (71) = 163.4, p< 0.001, I ² = 56.4 %, τ² = 0.058, strongly justifying moderator analyses.
Sensitivity analyses: Leave-one-out analysis confirmed robustness; removing any single study yielded effect sizes ranging from d = 0.31 to d = 0.37. Restricting to peer-reviewed publications (k = 69) yielded d = 0.35, 95 % CI [0.23, 0.47]; restricting to between-subjects designs (k= 51) yielded d = 0.37, 95 % CI [0.23, 0.51].
H1 was supported. Across these laboratory-based experiments, initial retrieval practice was associated with a small but reliable increase in vulnerability to subsequent misinformation under forensic-relevant conditions. Because the synthesized evidence derives from controlled analog studies rather than field investigations, this pattern should be read as a robust experimental association consistent with a causal role for retrieval, rather than as definitive proof of causation in real-world interviewing.
Moderator 1: Initial test format (H2)
Test format was a significant moderator, Q_between (2) = 22.3, p < 0.001 (Table 2, Figure 3).


Note. Subgroup analyses showing separate pooled effects for recognition-format initial tests (k = 14, d = 0.66), cued recall (k = 42, d = 0.38), and free recall/Cognitive Interview (k = 16, d = 0.17). The between-group test is significant: Q_between (2) = 22.3, p < 0.001. Effect sizes are ordered from largest (recognition) to smallest (free recall). The figure illustrates the clear gradient from maximal RES (recognition) to minimal/protective effects (free recall). Subgroup diamonds show weighted mean effects for each format; the dashed vertical line at zero marks the null effect boundary.
Recognition format (k = 14): d= 0.66, 95 % CI [0.47, 0.85], p < 0.001. Participants accepting approximately 51.8 % of misinformation versus 36.3 % for controls (15.5 percentage point increase).
Cued recall (k = 42): d= 0.38, 95 % CI [0.24, 0.52], p < 0.001. Approximately 7.3 percentage point increase in misinformation acceptance.
Free recall (k = 16): d = 0.17, 95 % CI [0.04, 0.30], p = 0.010. Approximately 3.2 percentage point increase; notably, free recall of central details produced a small protective effect (d= −0.21).
All pairwise comparisons were significant: recognition > cued recall, Q (1) = 6.2, p = 0.013; recognition > free recall, Q (1) = 19.8, p < 0.001; cued recall > free recall, Q (1) = 4.3, p = 0.038. A significant test format × detail type interaction (Q _interaction = 10.8, p = 0.005) showed that while free recall was protective for central details (d = −0.21), recognition produced strong RES even for central details (d = 0.43), and recognition of peripheral details produced the largest RES effects (d = 0.75).
H2 was strongly supported. RES effects follow a clear gradient determined by initial test format, consistent with source confusion predictions based on guessing demands.
Moderator 2: Misinformation source/format (H3)
Misinformation source significantly moderated RES, Q_between (2) = 18.7, p < 0.001 (Table 3).

Narrative misinformation (k = 48): d = 0.48, 95 % CI [0.35, 0.61], p < 0.001. Misinformation acceptance increased from 37.4 % to 50.8 % (13.4 percentage point increase).
Questions with context (k = 8): d= 0.55, 95 % CI [0.34, 0.76], p < 0.001. Questions embedded within event-reinstating contextual information produced RES comparable to narratives.
Questions without context (k= 16): d = −0.19, 95 % CI [−0.34, −0.04], p = 0.013. Decontextualized questions produced a significant protective testing effect. LaPaglia and Chan (2019) demonstrated this narrative versus question distinction is mediated by contextual reinstatement: it is not the surface form but the reinstatement of original event context that drives misinformation encoding. Source monitoring analyses (k = 7 studies) showed 65 % of accepted misinformation was attributed to the original event in retrieval-practice conditions versus 49 % in controls, supporting enhanced source confusion following retrieval practice.
H3 was strongly supported.
Moderator 3: Detail type –central vs. peripheral (H4)
Detail type significantly moderated RES, Q_between (1) = 14.6, p < 0.001 (Table 4).

Peripheral details (k = 38): d = 0.50, 95% CI [0.36, 0.64], p < .001. Misinformation acceptance increased from ~39.4% to 52.4%.
Central details (k = 34): d = 0.16, 95% CI [0.03, 0.29], p =.014. Only a 3.9 percentage point increase. Studies coding retrieval success (k = 8) revealed that peripheral details were retrieved less often than central details (~44% vs. 72%), and RES was substantially greater for non-retrieved details (d ≈ 0.58) than successfully retrieved details (d ≈ 0.06, n.s.), consistent with the selective attention account.
H4 was supported.
Moderator 4: Timing of misinformation exposure (H5)
The temporal interval between retrieval and misinformation significantly moderated RES; meta-regression revealed a significant negative relationship, β = −0.021, SE = 0.007, p = 0.004.
Immediate (k = 31, < 1 hour): d = 0.46, 95 % CI [0.31, 0.61], p < 0.001.
Short delay (k = 27, 1–24 hours): d = 0.28, 95 % CI [0.13, 0.43], p < 0.001.
Long delay (k = 14, > 24 hours): d = 0.22, 95 % CI [0.04, 0.40], p = 0.017.
Immediate versus long-delay conditions differed significantly, Q (1) = 7.8, p = 0.005. Studies examining misinformation within the proposed reconsolidation window (~6 hours) showed d = 0.49 versus d = 0.26 beyond 6 hours, Q (1) = 6.2, p = 0.013. Critically, Thomas et al. (2017) demonstrated that a 48-hour delay before the final test reversed RES into a significant protective effect (d = 1.44 on accuracy), indicating that the final-test interval is also a key determinant. Chan and LaPaglia’s (2013) six-experiment PNAS study provided the most controlled reconsolidation evidence: disruption occurred only when misinformation followed retrieval within approximately 20 minutes.
H5 was supported.
Moderator 5: Warning effects (H6)
Explicit warnings dramatically altered RES, Q_between (1) = 29.4, p < 0.001 (Table 4, Figure 4).

Note. Subgroup analyses showing separate pooled effects for no-warning conditions (k = 55, d = 0.41) and explicit warning conditions (k = 17, d = −0.07). Q_between (1) = 29.4, p < 0.001. The figure illustrates the dramatic impact of warnings: standard RES in the absence of warnings (positive d significantly different from zero) is completely eliminated when warnings are provided (negative d not significantly different from zero, p = 0.41). The overlap between the individual effect sizes in the warning subgroup and the zero line illustrates the absence of RES in warned conditions.
No warning (k = 55): d = 0.41, 95 % CI [0.29, 0.53], p < 0.001.
Explicit warning (k = 17): d = −0.07, 95 % CI [−0.24, 0.10], p = 0.41. RES was completely eliminated. Warnings are most effective when given in close temporal proximity to misinformation; Chan et al. (2022) demonstrated that warnings 48 hours after misinformation lose effectiveness. Karanian et al. (2020) provided neural evidence that warnings increase visual cortex reinstatement of original event traces and decrease auditory reinstatement of narrative content during final retrieval, explaining the mechanism at the neural level.
H6 was strongly supported.
Publication bias assessment
Multiple complementary methods converged on minimal publication bias. Funnel plot inspection revealed slight asymmetry with a modest gap in the lower-left region, but Egger’s regression was non-significant (b = 1.11, SE = 0.79, p = 0.17). Trim-and-fill analysis imputed only 5 potentially missing studies, yielding a minimally adjusted effect of d = 0.31 versus observed d = 0.34. p-curve analysis showed significant right-skew (Z = −3.84, p < 0.001), indicating genuine evidential value. Three-parameter selection models yielded an adjusted d = 0.32, 95 % CI [0.19, 0.45], nearly identical to the unadjusted estimate. Overall, the RES effect and moderator patterns are robust to potential publication selection.
Additional exploratory analyses
Exploratory analyses examined event type (terrorism, robbery, theft), event duration, participant age, sample type (student vs. community), and number of initial tests. Event type did not significantly moderate RES, Q (2) = 2.1, p = 0.35, nor did event duration (β = 0.002, p = 0.84), suggesting generalizability across forensic scenario types. Student (d = 0.35) and community (d = 0.30) samples produced comparable effects, Q (1) = 0.4, p = 0.53. Number of initial tests showed a positive dose-response relationship with RES magnitude (β = 0.16, SE = 0.07, p = 0.026 on log-transformed tests), consistent with Chan and LaPaglia’s (2011) finding of monotonically increasing suggestibility from 0 to 5 tests –a finding with direct forensic implications for repeated re-interviewing.
Discussion
This meta-analysis synthesized 26 studies with 72 effect size estimates from over 4 200 participants examining retrieval-enhanced suggestibility in forensic contexts. Results document a reliable paradoxical effect: initial retrieval attempts increase vulnerability to subsequent misinformation by d = 0.34. Comprehensive moderator analyses identified specific conditions that amplify versus attenuate RES, enabling translation into evidence-based investigative protocols.
Theoretical implications
Our findings support multi-mechanism accounts of RES (Johnson et al., 1993; Schacter et al., 2011): no single framework fully explains the observed moderator pattern; instead, different mechanisms operate under different conditions.
The test format gradient (recognition > cued recall > free recall) is best explained by source monitoring failures (Johnson et al., 1993). Recognition questions maximally encourage guessing, generating internal representations that subsequently serve as competing memory sources alongside actual memories and external misinformation; when misinformation confirms prior guesses, it is misattributed to the original event (Lindsay, 2008). Free recall minimizes guessing –witnesses omit uncertain details rather than committing to potentially wrong answers –reducing source confusion. The memory reconsolidation account best explains the temporal interval effects: RES is strongest within approximately 6 hours of retrieval, consistent with a labile reconsolidation window (Chan & LaPaglia, 2013; Nader et al., 2000). However, RES persisting beyond the reconsolidation window (Chan & Langley, 2011) indicates that source confusion and TPL continue to operate at longer delays.
The narrative versus question distinction and the peripheral versus central detail gradient are best explained by test-potentiated learning and context processing (Gordon & Thomas, 2013; LaPaglia & Chan, 2019). Initial testing directs attention toward tested event aspects; when a subsequent narrative reinstates event context, tested details serve as retrieval cues enhancing processing of related misinformation content (Gordon et al., 2015, 2020). For peripheral details, failed retrieval creates gaps later filled by misinformation; for central details, successful retrieval strengthens original traces, reducing susceptibility. Narratives reinstate temporal and spatial event context, fostering gist-level integration of misinformation; decontextualized questions prompt verbatim retrieval that discriminates original from post-event information.
The complete elimination of RES by explicit warnings demonstrates that RES reflects controllable metacognitive processes rather than automatic, irreversible memory mechanisms. Neural evidence from Karanian et al. (2020) shows that warnings redirect memory retrieval toward original encoding episodes, explaining how warnings counteract RES without necessarily preventing misinformation encoding.
Forensic practice implications: evidence-based interview protocols
The following recommendations are offered as tentative, evidence-informed suggestions rather than as prescriptive protocols. They are derived from laboratory analog studies conducted largely with university samples and should be interpreted as hypotheses to be validated in field settings before being adopted in operational practice. Read with this caveat, the present findings suggest the following directions for investigative interviewing:
Free recall produces minimal RES (d = 0.17) and protects central details (d = −0.21), while recognition formats produce strong RES (d = 0.66). Initial interviews should prioritize open-ended questions eliciting free narrative recall (“Tell me everything you remember”) rather than recognition-format questions. Avoid yes/no, multiple-choice, or forced-choice questions about uncertain details, as these encourage guessing and create internal representations that later interfere with accurate memory.
RES effects are strongest when misinformation follows retrieval within one hour (d = 0.46) and decrease substantially after 24 hours (d = 0.22). Witnesses should be cautioned against discussing the event with others, consuming media coverage, or reviewing social media immediately after giving statements. When multiple interviews are necessary, waiting at least 24 hours is preferable; Thomas et al. (2017, Experiment 2) demonstrated that a 48-hour delay can reverse RES into a protective effect.
Given that RES increases misinformation acceptance by ~8.6 percentage points overall, determining which details originate from original memory versus post-event sources strongly supports the use of verbatim documentation (audio/video recording preferred, detailed written documentation minimum) before any follow-up or potential contamination. This establishes a baseline, enables assessment of contamination timing, and protects investigative integrity. RES findings provide additional scientific support for jurisdictional policies mandating recording of investigative interviews (Kassin & Gudjonsson, 2004).
Explicit warnings completely eliminate RES (d = −0.07, p = 0.41), validated across laboratory (Chan et al., 2022; Karanian et al., 2020; Thomas et al., 2010) and online (Torrance et al., 2025) samples. Before follow-up interviews, provide warnings such as: “During this interview, I may ask about details you haven’t thought about before. It’s important that you rely only on your own memory of what you actually saw. If you’re not sure about something, say you don’t know rather than guess.” Warnings are most effective when given shortly before potential misinformation exposure; warnings given 48 hours after misinformation lose effectiveness (Chan et al., 2022).
RES effects are substantially stronger for peripheral (d = 0.50) than central details (d = 0.16). Initial interviews should focus primarily on central details via free recall. Peripheral details (background objects, bystanders, environmental features), which are more vulnerable to RES and more likely to go unretrieved initially, should be documented early and treated as higher contamination risk. These recommendations align well with the Cognitive Interview (Fisher & Geiselman, 2010), the NICHD Protocol for child witnesses (Lamb et al., 2007), and the PEACE Model (Milne & Bull, 1999), extending their scientifically grounded safeguards against iatrogenic memory contamination.
Limitations and boundary conditions
Several limitations warrant consideration. Ecological validity is limited as 95.2 % of studies used video-presented simulated events with undergraduate participants; field research with actual witnesses would strengthen generalizability. Stimulus homogeneity is a concern given ~57 % of studies used a single stimulus (24 television episode), limiting generalizability across diverse forensic scenarios. Individual differences in working memory capacity, source monitoring ability, and suggestibility could not be examined given sample homogeneity. Short retention intervals in most studies leave unclear whether RES effects persist, amplify, or diminish over forensically realistic delays of months to years. Measurement heterogeneity in how misinformation acceptance was operationalized (proportion endorsed, forced-choice accuracy, source attribution errors) may obscure meaningful differences despite standardization to Cohen’s d. While publication bias assessments suggest minimal impact, we cannot definitively rule out selective reporting, and the concentration of studies from three laboratory groups creates the possibility that paradigm-specific moderators not captured by coded variables affect estimates. Two features of the evidence base deserve particular emphasis when interpreting the present findings. First, 85.7 % of the included studies drew on undergraduate convenience samples tested with video-presented analog events; young, cognitively able university students are not representative of the diverse witness populations encountered in real investigations (children, older adults, victims of trauma, individuals with cognitive or linguistic vulnerabilities), so the external validity of the pooled estimate to operational forensic settings is uncertain. Second, 85.7 % of the studies originated from only three research groups working within closely related paradigms. This concentration means that the meta-analytic estimate reflects a narrow slice of the possible methodological space rather than an independent, broadly replicated body of work; shared laboratory practices, materials (e.g., the frequent use of a single television-based stimulus), and theoretical commitments could inflate apparent consistency and constrain generalizability. Accordingly, the effect sizes and moderator patterns reported here should be regarded as provisional and in need of confirmation through independent replication with heterogeneous samples, diverse stimuli, and field-based designs before they are generalized to actual eyewitnesses or used to justify changes in investigative practice.
Future directions
Priority research directions include: (1) field validation studies collaborating with law enforcement to examine RES under actual investigative procedures, leveraging jurisdictions that mandate digital recording of investigative interviews; (2) long-term retention intervals extending laboratory research to forensically realistic delays (6 months, 1 year) to determine optimal interview spacing, building on Thomas et al.’s (2017) demonstration of a critical reversal at 48 hours; (3) individual difference predictors examining cognitive (working memory, source monitoring) and personality (suggestibility, anxiety, trauma exposure) factors to identify vulnerable individuals and tailor protective protocols; and (4) diverse forensic stimuli and special populations extending RES research to child abuse, domestic violence, and vehicular accident scenarios, and to children, older adults, and individuals with PTSD, given Liu et al.’s (2024) initial evidence that acute stress actually reduces RES.
Conclusion
Synthesizing the available laboratory evidence, this meta-analysis indicates that initial retrieval practice is associated with a small, paradoxical increase in vulnerability to subsequent misinformation under forensic-relevant conditions (d = 0.34; ~8.6 percentage point increase in misinformation acceptance). However, this effect is highly moderated by procedural factors: recognition-format questions produce substantially larger effects than free recall; narrative misinformation produces RES while question-based misinformation produces a protective effect; peripheral details are more vulnerable than central details; immediate post-retrieval misinformation produces stronger effects than delayed misinformation; and explicit warnings completely eliminate RES effects.
The convergence of theoretical mechanisms (source monitoring failures, memory reconsolidation, test-potentiated learning, context processing) helps account for the complex pattern of moderator effects and is consistent with the view that RES reflects controllable metacognitive processes that may be amenable to procedural intervention. The observation that explicit warnings substantially reduced, and in several conditions abolished, the measured RES effect suggests that, at least under these experimental conditions, the phenomenon may not be inevitable and could be attenuated by appropriate source-monitoring guidance. Whether this attenuation generalizes to real investigative settings remains an open empirical question.
For forensic practice, these findings suggest several evidence-informed directions for investigative practice: initial investigative interviews may benefit from prioritizing open-ended free recall, avoiding recognition-format questions about uncertain details, implementing temporal separation before follow-up interviews, comprehensively documenting initial statements, and providing explicit warnings about misinformation risks. Before these findings can be translated into operational practice, replication with heterogeneous witness populations–including children, older adults, trauma victims, and individuals with cognitive vulnerabilities–and field-based designs is essential. These recommendations integrate well with existing evidence-based protocols (Cognitive Interview, NICHD Protocol, PEACE Model) while adding scientifically grounded safeguards against iatrogenic memory contamination. Ultimately, investigative procedures shape the evidence they collect; by understanding how retrieval practice affects memory vulnerability, the forensic community can implement protocols that maximize accurate memory preservation while minimizing contamination risk.
References
Adesope, O. O., Trevisan, D. A., & Sundararajan, N. (2017). Rethinking the use of tests: A meta-analysis of practice testing. Review of Educational Research, 87(3), 659–701. https://doi.org/10.3102/0034654316689306
Anderson, M. C., & Spellman, B. A. (1995). On the status of inhibitory mechanisms in cognition: Memory retrieval as a model case. Psychological Review, 102(1), 68–100. https://doi.org/10.1037/0033-295X.102.1.68
Branco, A. (2018). The role of authority in retrieval enhanced suggestibility (RES). Locus: The Seton Hall Journal of Undergraduate Research, 1(1), Article 3. https://doi.org/10.70531/2573-2749.1001
Chan, J. C. K., & LaPaglia, J. A. (2011). The dark side of testing memory: Repeated retrieval can enhance eyewitness suggestibility. Journal of Experimental Psychology: Applied, 17(4), 418–432. https://doi.org/10.1037/a0025147
Chan, J. C. K., & LaPaglia, J. A. (2013). Impairing existing declarative memory in humans by disrupting reconsolidation. Proceedings of the National Academy of Sciences, 110(23), 9309–9313. https://doi.org/10.1073/pnas.1218472110
Chan, J. C. K., & Langley, M. M. (2011). Paradoxical effects of testing: Retrieval enhances both accurate recall and suggestibility in eyewitnesses. Journal of Experimental Psychology: Learning, Memory, and Cognition, 37(1), 248–255. https://doi.org/10.1037/a0021204
Chan, J. C. K., Manley, K., & Lang, K. (2017). Retrieval-enhanced suggestibility: A retrospective and a new investigation. Journal of Applied Research in Memory and Cognition, 6(3), 213–229. https://doi.org/10.1016/j.jarmac.2017.07.003
Chan, J. C. K., O’Donnell, R., & Manley, K. (2022). Warning weakens retrieval-enhanced suggestibility only when it is given shortly after misinformation: The critical importance of timing. Journal of Experimental Psychology: Applied, 28(4), 694–716. https://doi.org/10.1037/xap0000394
Chan, J. C. K., Thomas, A. K., & Bulevich, J. B. (2009). Recalling a witnessed event increases eyewitness suggestibility: The reversed testing effect. Psychological Science, 20(1), 66–73. https://doi.org/10.1111/j.1467-9280.2008.02245.x
Chan, J. C. K., Wilford, M. M., & Hughes, K. L. (2012). Retrieval can increase or decrease suggestibility depending on how memory is tested: The importance of source complexity. Journal of Memory and Language, 67(1), 78–85. https://doi.org/10.1016/j.jml.2012.02.006
Diges, M. (2016). Testigos, sospechosos y recuerdos falsos: Estudios de psicología forense. Editorial Trotta.
Fisher, R. P., & Geiselman, R. E. (1992). Memory-enhancing techniques for investigative interviewing: The cognitive interview. Charles C Thomas Publisher.
Fisher, R. P., & Geiselman, R. E. (2010). The cognitive interview method of conducting police interviews: Eliciting extensive information and promoting therapeutic jurisprudence. International Journal of Law and Psychiatry, 33(5–6), 321–328. https://doi.org/10.1016/j.ijlp.2010.09.004
Gabbert, F., Hope, L., & Fisher, R. P. (2009). Protecting eyewitness evidence: Examining the efficacy of a self-administered interview tool. Law and Human Behavior, 33(4), 298–307. https://doi.org/10.1007/s10979-008-9146-8
Gordon, L. T., Bilolikar, V. K., Hodhod, T., & Thomas, A. K. (2020). How prior testing impacts misinformation processing: A dual-task approach. Memory & Cognition, 48(2), 314–324. https://doi.org/10.3758/s13421-019-00970-0
Gordon, L. T., & Thomas, A. K. (2013). Testing potentiates new learning in the misinformation paradigm. Memory & Cognition, 42(1), 186–197. https://doi.org/10.3758/s13421-013-0361-2
Gordon, L. T., Thomas, A. K., & Bulevich, J. B. (2015). Looking for answers in all the wrong places: How testing facilitates learning of misinformation. Journal of Memory and Language, 83, 140–151. https://doi.org/10.1016/j.jml.2015.03.007
Hedges, L. V. (1981). Distribution theory for Glass’s estimator of effect size and related estimators. Journal of Educational Statistics, 6(2), 107–128. https://doi.org/10.3102/10769986006002107
Hedges, L. V., Tipton, E., & Johnson, M. C. (2010). Robust variance estimation in meta-regression with dependent effect size estimates. Research Synthesis Methods, 1(1), 39–65. https://doi.org/10.1002/jrsm.5
Huff, M. J., Weinsheimer, C. C., & Bodner, G. E. (2016). Reducing the misinformation effect through initial testing: Take two tests and recall me in the morning? Applied Cognitive Psychology, 30(1), 61–69. https://doi.org/10.1002/acp.3167
Johnson, M. K., Hashtroudi, S., & Lindsay, D. S. (1993). Source monitoring. Psychological Bulletin, 114(1), 3–28. https://doi.org/10.1037/0033-2909.114.1.3
Karanian, J. M., Rabb, N., Wulff, A. N., Torrance, M. G., Thomas, A. K., & Race, E. (2020). Protecting memory from misinformation: Warnings modulate cortical reinstatement during memory retrieval. Proceedings of the National Academy of Sciences, 117(37), 22771–22779. https://doi.org/10.1073/pnas.2008595117
Karanian, J. M., Thomas, A. K., & Race, E. (2024). Warning before misinformation exposure modulates memory encoding. Cognitive, Affective, & Behavioral Neuroscience, 24(3), 440–452. https://doi.org/10.3758/s13415-024-01183-y
Kassin, S. M., Drizin, S. A., Grisso, T., Gudjonsson, G. H., Leo, R. A., & Redlich, A. D. (2010). Police-induced confessions: Risk factors and recommendations. Law and Human Behavior, 34(1), 3–38. https://doi.org/10.1007/s10979-009-9188-6
Kassin, S. M., & Gudjonsson, G. H. (2004). The psychology of confessions: A review of the literature and issues. Psychological Science in the Public Interest, 5(2), 33–67. https://doi.org/10.1111/j.1529-1006.2004.00016.x
Lamb, M. E., Orbach, Y., Hershkowitz, I., Esplin, P. W., & Horowitz, D. (2007). A structured forensic interview protocol improves the quality and informativeness of investigative interviews with children: A review of research using the NICHD Investigative Interview Protocol. Child Abuse & Neglect, 31(11–12), 1201–1231. https://doi.org/10.1016/j.chiabu.2007.03.021
LaPaglia, J. A., & Chan, J. C. K. (2012). Retrieval does not always enhance suggestibility: Testing can improve witness identification performance. Law and Human Behavior, 36(6), 478–487. https://doi.org/10.1037/h0093931
LaPaglia, J. A., & Chan, J. C. K. (2013). Testing increases suggestibility for narrative-based misinformation but reduces suggestibility for question-based misinformation. Behavioral Sciences & the Law, 31(5), 593–606. https://doi.org/10.1002/bsl.2090
LaPaglia, J. A., & Chan, J. C. K. (2019). Telling a good story: The effects of memory retrieval and context processing on eyewitness suggestibility. PLOS ONE, 14(2), e0212592. https://doi.org/10.1371/journal.pone.0212592
LaPaglia, J. A., Wilford, M. M., Rivard, J. R., Chan, J. C. K., & Fisher, R. P. (2013). Misleading suggestions can alter later memory reports even following a Cognitive Interview. Applied Cognitive Psychology, 28(1), 1–9. https://doi.org/10.1002/acp.2950
Lindsay, D. S. (2008). Source monitoring. In H. L. Roediger III (Ed.), Cognitive psychology of memory: Vol. 2. Learning and memory: A comprehensive reference (pp. 325–348). Elsevier.
Lipsey, M. W., & Wilson, D. B. (2001). Practical meta-analysis. Sage Publications.
Liu, W., Huizhen, D., Siman, Z., Jiayu, Z., & Xiaoyu, W. (2024). The more stress, the more accurate memory? The influence of stress on retrieval-enhanced suggestibility effect. Journal of Cognitive Psychology, 36(7), 793-804. https://doi.org/10.1080/20445911.2024.2386733
Manley, K. D., & Chan, J. C. K. (2019). Does retrieval enhance suggestibility because it increases perceived credibility of the postevent information? Journal of Applied Research in Memory and Cognition, 8(3), 355–366. https://doi.org/10.1016/j.jarmac.2019.06.001
Manzanero, A. L. (2010). Memoria de testigos: Obtención y valoración de la prueba testifical. Ediciones Pirámide.
Manzanero, A. L., & González, J. L. (2013). Avances en psicología del testimonio. Ediciones Jurídicas de Santiago.
Milne, R., & Bull, R. (1999). Investigative interviewing: Psychology and practice. John Wiley & Sons.
Nader, K., Schafe, G. E., & LeDoux, J. E. (2000). Fear memories require protein synthesis in the amygdala for reconsolidation after retrieval. Nature, 406(6797), 722–726. https://doi.org/10.1038/35021052
Pansky, A., & Tenenboim, E. (2011). Inoculating against eyewitness suggestibility via interpolated verbatim vs. gist testing. Memory & Cognition, 39(1), 155–170. https://doi.org/10.3758/s13421-010-0005-8
Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432–1463. https://doi.org/10.1037/a0037559
Saunders, J., & MacLeod, M. D. (2002). New evidence on the suggestibility of memory: The role of retrieval-induced forgetting in misinformation effects. Journal of Experimental Psychology: Applied, 8(2), 127–142. https://doi.org/10.1037/1076-898X.8.2.127
Schacter, D. L., Guerin, S. A., & St. Jacques, P. L. (2011). Memory distortion: An adaptive perspective. Trends in Cognitive Sciences, 15(10), 467–474. https://doi.org/10.1016/j.tics.2011.08.004
Silva, E. A., Manzanero, A. L., & Contreras, M. J. (2016). La memoria y el lenguaje en pruebas testificales con menores de 3 a 6 años. Papeles del Psicólogo, 37(3), 224–230.
Simonsohn, U., Nelson, L. D., & Simmons, J. P. (2014). .-curve: A key to the file-drawer. Journal of Experimental Psychology: General, 143(2), 534–547. https://doi.org/10.1037/a0033242
Thomas, A. K., Bulevich, J. B., & Chan, J. C. K. (2010). Testing promotes eyewitness accuracy with a warning: Implications for retrieval enhanced suggestibility. Journal of Memory and Language, 63(2), 149–157. https://doi.org/10.1016/j.jml.2010.04.004
Thomas, A. K., Gordon, L. T., Cernasov, P. M., & Bulevich, J. B. (2017). The effect of testing can increase or decrease misinformation susceptibility depending on the retention interval. Cognitive Research: Principles and Implications, 2(1), 45. https://doi.org/10.1186/s41235-017-0081-4
Torrance, M. G., Karanian, J. M., Race, E., & Thomas, A. K. (2025). Examining the impact of warnings on eyewitness memory. Scientific Reports, 15(1), Article 33508. https://doi.org/10.1038/s41598-025-17377-4
Vevea, J. L., & Woods, C. M. (2005). Publication bias in research synthesis: Sensitivity analysis using a priori weight functions. Psychological Methods, 10(4), 428–443. https://doi.org/10.1037/1082-989X.10.4.428
Viechtbauer, W. (2010). Conducting meta-analyses in R with the metafor package. Journal of Statistical Software, 36(3), 1–48. https://doi.org/10.18637/jss.v036.i03
Wells, G. L., Kovera, M. B., Douglass, A. B., Brewer, N., Meissner, C. A., & Wixted, J. T. (2020). Policy and procedure recommendations for the collection and preservation of eyewitness identification evidence. Law and Human Behavior, 44(1), 3–36. https://doi.org/10.1037/lhb0000359
Wells, G. L., & Quinlivan, D. S. (2009). Suggestive eyewitness identification procedures and the Supreme Court’s reliability test in light of eyewitness science: 30 years later. Law and Human Behavior, 33(1), 1–24. https://doi.org/10.1007/s10979-008-9130-3
Wilford, M. M., Chan, J. C. K., & Tuhn, S. J. (2014). Retrieval enhances eyewitness suggestibility to misinformation in free and cued recall. Journal of Experimental Psychology: Applied, 20(1), 81–93. https://doi.org/10.1037/xap0000001
Wulff, A. N., Karanian, J. M., Race, E., & Thomas, A. K. (2025). Timing of pre-retrieval warnings matters in reducing memory errors in a repeated testing misinformation study. Scientific Reports, 15(1), 963. https://doi.org/10.1038/s41598-024-84154-0
Notes
*
Review article.
Author notes
a Correspondence author. Email: alberto.barea@ucavila.es
Additional information
How to cite: Barea
Vera, A. (2026). Retrieval-Enhanced Suggestibility in Forensic Contexts: A
meta-analysis of practice effects on eyewitness memory contamination. Universitas
Psychologica, 25, 1-17. https://doi.org/10.11144/Javeriana.upsy25.resf