

ABSTRACT
Conference interpreter training has long been constrained by a chronic shortage of pedagogically graded, skill-specific materials. Existing repositories such as United Nations and European Union speech repositories, while valuable, were never designed for progressive skill development or calibrated to the nuanced cognitive demands identified by Gile’s (1995, 2009) Effort Models. This paper proposes a conceptual and pedagogical framework—the Material Digital Laboratory (MDL)—in which generative artificial intelligence (GenAI) tools are leveraged to engineer training materials aligned with specific cognitive and linguistic objectives. Rather than selecting from fixed speech repositories, trainers operate within a dynamic design environment where AI-generated content is grounded in factually reliable inputs through Retrieval-Augmented Generation (RAG) using NotebookLM, synthesized as audio via Google AI Studio’s text-to-speech capabilities, and orchestrated through Gemini-based prompt engineering. The framework draws on a theoretical synthesis of Gile’s Effort Models, Cognitive Load Theory, and skill-oriented pedagogy to argue that interpreting sub-skills—listening and analysis, memory, production, and coordination—should function as design parameters rather than incidental outcomes of exposure-based practice. Two illustrative case studies, simulating a climate adaptation conference and an economic policy forum, demonstrate the framework’s operational feasibility and pedagogical precision. The paper makes three core contributions: a re-theorization of interpreting skills as designable pedagogical targets; a five-stage prototype workflow for AI-based material generation; and a revised competence framework for the interpreter trainer as pedagogical engineer. Future empirical validation of learning outcomes is identified as the primary next step in the research agenda.
KEYWORDS: Generative Ai; Interpreter Training; Cognitive Load; Material Design; Pedagogical Engineering; Retrieval-Augmented Generation; Skill-Oriented Pedagogy
The integration of artificial intelligence into professional training contexts has accelerated markedly in recent years, yet its application to conference interpreter training remains systematically underdeveloped. While translation studies has embraced machine translation post-editing, corpus-assisted research, and AI-generated feedback with increasing sophistication (Cui et al., 2025; Moukatib & Ben Seddik, 2026), the parallel field of interpreting pedagogy continues to operate largely within a resource logic that dates back decades. Trainers at most institutions rely on a familiar and finite inventory of authentic speeches—European Parliament debates, United Nations General Assembly addresses, and NGO conference recordings—that, however valuable for advanced practice, were never designed to isolate specific cognitive sub-skills or to scaffold progressive skill acquisition.
This material constraint is not merely logistical. It reflects a deeper pedagogical assumption: that interpreters develop competence primarily through exposure to authentic professional content, with sufficient volume and corrective feedback eventually producing expertise. This assumption has been challenged on theoretical grounds by Gile (1995, 2009), whose Effort Models identify specific cognitive resources—listening and analysis, memory, and production—as finite and manageable variables. It has also been challenged empirically by studies demonstrating that targeted, level-appropriate materials produce measurably better outcomes than undifferentiated authentic exposure (Li, 2018; Seeber & Arbona, 2020; Song & Tang, 2020). Yet despite theoretical consensus, the field has lacked the technological means to operationalize skill-specific material design at scale.
Generative AI fundamentally alters this constraint. Large language models (LLMs) such as Gemini and Retrieval-Augmented Generation (RAG) systems such as NotebookLM can produce topically grounded, linguistically calibrated source texts on demand. Text-to-speech (TTS) systems such as Google AI Studio can render these texts as audio with precise control over delivery rate, accent, and prosodic features. When these tools are organized and used together, AI output can be grounded in curated factual sources, substantially reducing the hallucination risk that would otherwise disqualify AI-generated content from use in a context demanding factual accuracy.
This paper introduces the concept of the Material Digital Laboratory (MDL) as a framework for theorizing and operationalizing this shift. The MDL is defined as a design environment in which interpreter trainers generate, modify, control, and sequence training materials in systematic alignment with specific cognitive and pedagogical objectives. It represents not a tool but a paradigm: a transformation of the trainer’s role from material selector to pedagogical engineer, and of training materials from found objects to designed instruments.
The paper makes three primary contributions. First, it re-theorizes interpreting sub-skills as designable pedagogical targets, extending Gile’s (1995, 2009) Effort Models to identify specific cognitive variables that materials should be engineered to manipulate. Second, it proposes a five-stage pedagogical prototype—from corpus curation through audio synthesis—that operationalizes the MDL concept. Third, it introduces a revised competence framework for interpreter trainers that integrates pedagogical, technical, and design expertise. Two illustrative case studies ground these contributions in practice, demonstrating the framework’s application across different domains, language directions, and skill targets.
Any account of interpreting pedagogy must engage with the foundational theoretical framework of the field: Daniel Gile’s Effort Models (1995, 2009). Gile’s models propose that simultaneous interpreting involves three principal cognitive efforts—Listening and Analysis (L), Memory (M), and Production (P)—that compete for a finite pool of attentional resources. The tightrope hypothesis (Gile, 1999) holds that interpreters habitually work near the limit of their cognitive capacity, rendering them acutely vulnerable to problem triggers: delivery features such as high information density, rapid speech rate, unfamiliar terminology, or low redundancy that impose disproportionate processing demands. When the sum of cognitive efforts exceeds available capacity, errors and omissions proliferate—not, critically, because of linguistic inadequacy, but because of resource management failure.
This has fundamental implications for pedagogy. If performance is a function of attentional resource allocation rather than simply linguistic competence, then training materials should be designed to selectively stress and develop specific cognitive efforts. This insight is confirmed by Li (2015b), who argues that strategy training must move beyond awareness-raising to the systematic engineering of conditions in which target strategies become cognitively necessary. Seeber (2011) extends Gile’s framework through the Multiple Resource Model, providing finer granularity on how cognitive load varies across perceptual, cognitive, and response dimensions—analysis that both validates the Effort Model tradition and identifies new parameters for material design. Al-Suhaim (2024) provides a comprehensive theoretical overview confirming that Gile’s framework remains the dominant conceptual anchor in contemporary interpreting studies.
A consistent finding across this literature is that training materials rarely align with the cognitive architecture they purport to develop. Students are exposed to speeches that are either too complex for productive learning (Li, 2018) or too simplified to replicate the cognitive demands of professional practice (Kalina, 2000; Moser-Mercer, 2000/01). Gile (1999) himself lamented the absence of precise measurement tools that would allow systematic testing of cognitive load hypotheses, identifying resource scarcity—both material and methodological—as the critical barrier to evidence-based pedagogy. The MDL framework proposed in this paper addresses this gap directly by treating Gile’s effort parameters as design inputs rather than merely explanatory constructs.
The broader cognitive load literature provides important grounding for skill-oriented material design. Seeber and Arbona (2020) apply a cognitive ergonomic framework to argue that training efficiency can be dramatically improved by eliminating unnecessary cognitive load so that students can engage productively with target learning material. Their atomistic approach—targeting specific cognitive sub-skills by modulating problem triggers in isolation before integrating them—provides a direct prototype for the MDL’s design logic. Similarly, Cai et al. (2015) demonstrate empirically that manipulation of segment length, syntactic complexity, and processing time has measurable effects on beginner performance, validating the principle that material design determines cognitive outcome.
Moser-Mercer (2000/01) identifies the role of long-term working memory in expert interpreting, arguing that expertise involves the automation of sub-processes that initially require conscious, effortful attention—a finding with direct implications for the progression logic of a pedagogical sequence. Seeber and Kerzel (2012) provide psychophysiological evidence, using task-evoked pupillary responses to measure cognitive load as a function of syntactic structure, demonstrating that engineered variations in source text syntax produce measurable and predictable cognitive effects. Their methodology—manipulating and digitally remastering materials to control prosodic variables—constitutes an empirical precursor to the MDL’s design approach. Macnamara et al. (2011) further extend this picture by demonstrating that domain-general cognitive abilities, including cognitive control, task-switching, and working memory capacity, are significant predictors of interpreting proficiency independent of language pair, suggesting that material design must target the underlying cognitive operations recruited by source texts, not merely their linguistic surface features.
Bakar and Tapsoba (2026), drawing on classroom teaching research, articulate the MDL’s theoretical logic in the most explicit terms available in the recent literature: they describe learning as occurring in a high-stakes decision ecology constrained by heterogeneous readiness and resource asymmetry, and identify AI as a design amplifier capable of reducing design friction while preserving teacher sovereignty over curricular intent. Their concept of constrained generativity—in which teachers specify concept boundaries and success criteria for AI to generate—is operationally equivalent to the prompt engineering layer of the MDL’s workflow.
The limitations of traditional material resources are thoroughly documented. Li (2018) identifies the central paradox of the field: authentic materials such as UN or EU speeches are frequently too demanding for undergraduate trainees, while simplified pedagogical texts lack the professional realism necessary to develop market-ready competence. Both manual simplification and undifferentiated authentic exposure carry costs that a systematic framework of material engineering could resolve. Chang and Wu (2017) corroborate this analysis, noting that classroom materials are typically delivered at speeds that fail to replicate the cognitive stress of professional settings, and documenting students’ inability to manage the unscripted challenges—fragmented delivery, unfamiliar accents, implicitly assumed background knowledge—that characterize real conference work.
The response to this scarcity has historically taken the form of simulation. Li (2015a) frames the mock conference as a virtual surrogate of professional environments, and Conde and Chouc (2019) demonstrate that multilingual simulations provide exposure to unscripted challenges—spontaneous delivery, speaker panel dynamics, non-native speakers—that standard classroom materials cannot replicate. Wang (2015) identifies blended learning models combining interpreting corpora with ICT tools as forerunners of the integrated digital environment the MDL represents. Corpas Pastor (2018, 2020) documents the evolution of CAIT from simple speech repositories through automatic corpus compilation and speech synthesis, tracing a lineage that the MDL framework extends and systematizes. Frittella (2021) and Sandrelli and de Manuel Jerez (2007) chart this evolution through authoring tools, simulators, and virtual learning environments, consistently finding that the critical constraint is not the availability of technology but the absence of pedagogical frameworks for designing materials within it. The MDL is, in essence, a principled approach to filling what Sandrelli and de Manuel Jerez (2007) called the empty box of digital training infrastructure.
The literature on technology in interpreter training has expanded considerably in recent years, partly driven by the COVID-19 pandemic’s forced migration to remote delivery. Zhao (2022) documents the shift toward virtual booths, Zoom-based mock conferences, and digital delivery systems, demonstrating that trainer competence now necessarily includes the technical coordination of complex multi-device environments. Chan (2023) shows that virtual reality applications can measurably improve interpreting competence by creating simulated authentic settings with controlled cognitive demands, while Djovčoš et al. (2023) identify Remote Simultaneous Interpreting platforms as a new pedagogical frontier requiring both technical literacy and material design skills that most programs have not yet systematically developed.
Corpas Pastor (2018) provides a comprehensive taxonomy of CAI tools—terminology management, ASR-based note support, corpus compilation, and speech synthesis—identifying a critical gap in purpose-built resources for interpreting relative to translation. Fan (2012) frames deliberate practice theory within interpreting pedagogy, arguing that expertise requires structured training sessions with precisely defined skill targets, appropriate challenge levels, and immediate feedback—all requirements that the MDL framework is designed to fulfill at scale. The convergence of these accounts is consistent: technology offers progressively more powerful capabilities for material design, but those capabilities are only pedagogically productive when organized within a theoretically grounded framework.
The emergence of large language models as educational tools has generated a rapidly growing literature. Kasneci et al. (2023) identify LLMs as a key enabling technology for simulation-based learning, capable of generating immersive multimodal environments that were previously achievable only through expensive live production. García-Peñalvo (2023) frames GenAI as a Possibility Engine and Co-designer for curriculum development, noting its capacity to conjure content aligned with specific academic goals, while Cordero et al. (2025) provide empirical evidence that AI-assisted resource creation reduces preparation time while improving material quality and differentiation.
Yan et al. (2024) conduct a systematic scoping review of LLMs in education, finding high performance in content generation tasks and identifying a clear trend toward synthetic datasets as replacements for manually curated corpora. Qian (2025) synthesizes a large body of literature to characterize GenAI as a cognitive scaffold, and Adamakis and Rachiotis (2025) warn of the risk of cognitive debt—the degradation of cognitive function resulting from excessive delegation of intellectual tasks to algorithmic systems. This warning has particular salience for interpreting training, where cognitive engagement is the central learning objective and where the MDL must therefore be designed to complement rather than circumvent the cognitive work of the trainee. Serra and Oliveira (2025) demonstrate that systematic prompt design can reposition educational resources as intelligent, pedagogically rigorous systems, and Sharma et al. (2025) document AI’s capacity to dynamically adjust material difficulty in real time to match individual learning trajectories—a capability that the MDL begins to approximate through its multi-level prompt architecture.
This broadly optimistic literature must, however, be read alongside critical scholarship on the risks that GenAI introduces in educational contexts. Three risks are particularly salient for the MDL framework. The first is corpus bias: RAG-grounded generation is only as balanced as the corpus the trainer assembles, and corpora drawn predominantly from anglophone institutional sources may systematically underrepresent non-Western epistemic frameworks, regional administrative varieties, or minority-language speaker patterns that are pedagogically significant in multilingual interpreter training programs. The second is overreliance: a trainer who routinely delegates material design to AI systems may experience attrition of the deep domain knowledge and pedagogical judgment that make quality review at Stage 3 of the MDL workflow meaningful—transforming an epistemological safeguard into a nominal procedure. The third is the limits of factual grounding even within RAG architectures: NotebookLM anchors outputs in uploaded documents but does not guarantee that those documents are current, accurate, or representative of the domain, nor that the corpus is comprehensive. Trainers must therefore understand RAG grounding as a substantial risk-reduction mechanism, not an infallibility guarantee. These risks define the conditions under which the MDL must be responsibly operated
Within Translation and Interpreting Studies specifically, a growing number of accounts document the integration of AI tools into pedagogical practice. Hatiarová (2025) provides one of the most directly relevant, describing a workflow in which ChatGPT and TTS tools are combined to produce training materials with manipulated delivery features, and introducing prompt crafting as a pedagogical act rather than a merely technical operation. Parrilla Gómez and Postigo Pinazo (2025) demonstrate the application of this approach to public service interpreting training, where ethical constraints on authentic material access make GenAI a particularly powerful alternative; their finding that AI can adequately replace the real materials which are lacking provides empirical grounding for the broader claim advanced here.
Wiedenmayer (2026) offers the most theoretically developed account of AI as a speech generation tool in CI training, arguing that pedagogical speeches require a degree of deliberate artificiality precisely because they must be designed to meet predefined instructional objectives rather than to reflect naturally occurring discourse—a position directly continuous with the MDL’s core logic. Yang et al. (2026) document translation students’ constructivist engagement with AI tools, finding that students already function as active engineers of their own learning environments, and Cui et al. (2025) provide an overview identifying pedagogical design—not tool proficiency—as the critical variable in AI-integrated training. Despite this growing body of work, a significant gap remains: existing studies tend to focus on either the technology or the pedagogy, rarely integrating both within a theoretically grounded framework that specifies how cognitive objectives should translate into material design parameters. The MDL framework is an attempt to fill that gap.
The capacity of GenAI to generate pedagogically useful materials is entirely contingent on the quality of the prompts that direct it. Carrasco-Sáez et al. (2025) provide empirical evidence that prompt sophistication is a strong predictor of output quality, identifying a spectrum from copy-and-paste to iterative prompt planning as a core dimension of what they term digital literacy in the GenAI era. Serra and Oliveira (2025) argue that systematic prompt design can reposition static educational materials as intelligent, transparent, and pedagogically rigorous systems, and Qian (2025) identifies prompt literacy as an emerging critical skill that educational programs must explicitly develop. This emerging literature converges on a finding with direct implications for the MDL: prompt engineering is not a technical afterthought but the primary locus of pedagogical judgment in an AI-integrated design environment.
The pedagogical transformations described across this literature are only realizable if trainers possess the competence to enact them. Muñoz-Basols et al. (2023) argue that educators must shift from a detect-react-prevent mindset to an integrate-educate-model approach to technology. Yan et al. (2024) identify human-AI collaboration as an emerging essential skill for the professional market, requiring that trainers model this collaboration rather than merely introducing students to its tools. Yang et al. (2025) deploy the Technological Pedagogical Content Knowledge (TPACK) framework to argue that interpreter trainers must now possess expertise at the intersection of domain knowledge, pedagogy, and technology—a tripartite competence profile that the MDL’s pedagogical engineer concept formalizes. The convergence of these accounts signals a profession-wide competence transition that training programs have not yet systematically addressed.
This paper grounds the MDL framework in a reinterpretation of Gile’s (1995, 2009) Effort Models as design parameters rather than merely explanatory constructs. The fundamental move is to treat the cognitive efforts identified by Gile—Listening and Analysis (L), Memory (M), Production (P), and Coordination (C)—not as features of the interpreter’s performance environment to be navigated, but as variables that material designers can deliberately manipulate.
In the standard interpretation, a problem trigger is a feature of the source speech that increases processing demand on one or more effort parameters: high delivery speed increases demand on both L and M; high information density increases demand on L; low redundancy reduces recovery time, increasing pressure on M and C. In the conventional pedagogical response, trainers search for authentic speeches that happen to contain these features at approximately appropriate levels, and students practice responding to them. The MDL framework inverts this relationship. Rather than discovering problem triggers in found materials, trainers design them. Delivery speed is specified as a parameter before material generation; information density is calibrated to target a specific cognitive effort; redundancy is engineered into the text to provide controlled recovery opportunities. The trainer functions not as a curator of an existing reality but as the architect of a designed one.
This reinterpretation is consistent with Gile’s (2021) retrospective account of the Effort Models as didactic constructs. It extends Seeber and Arbona (2020) and Wiedenmayer (2026) toward an operationalized material design methodology and draws on Cognitive Load Theory (Sweller, 1988, as referenced in Bakar & Tapsoba, 2026). CLT distinguishes extraneous load—cognitive demand generated by poor design or irrelevant content—from germane load, which is cognitive demand that builds schemas and develops skills. The MDL targets this distinction directly: RAG grounding reduces extraneous load by ensuring factual accuracy; deliberate calibration of problem triggers maximizes germane load by directing cognitive effort toward the intended skill target.
The framework maps Gile’s effort parameters to specific, designable material variables as follows. Listening and Analysis Effort is addressable through control over delivery speed, information density (propositions per clause), redundancy ratio, technical vocabulary density, and syntactic complexity. Memory Effort is addressable through segment length, discourse coherence, enumeration density, and the density of numerals and proper names requiring accurate retention. Production Effort is addressable through target language register pressure, reformulation difficulty, and processing window constraints such as forced ear-voice-span compression. Coordination Effort is addressable through simultaneous task demands such as speaker changes, visual aids, and simulated booth environment variables. By mapping objectives to parameters in this way, the MDL enables trainers to generate materials for targeted skill development and to sequence them in a principled progression—from high-support to low-support conditions, and from isolated sub-skill practice to integrated performance.
Before presenting the framework and its illustrative case studies, it is important to state the study’s epistemological status precisely. This is a conceptual and pedagogical modeling study. It does not collect or analyze data from student or trainer participants, does not report learning outcome measurements, and makes no causal claims about the instructional effectiveness of the MDL. Its contribution is theoretical. This conceptual contribution includes a re-conceptualization of interpreter training materials as designable cognitive objects. On the operational level, this study attempts to provide a demonstration that this re-conceptualization can be realised through available AI tools. Claims in the following sections pertain to the feasibility and coherence of the framework as a design proposition, not to its empirically verified effects.
This study therefore employs a conceptual and pedagogical modeling approach, combining systematic theoretical synthesis with an illustrative case study design. This approach follows established practice in applied linguistics and interpreting studies when introducing novel pedagogical models (cf. Seeber & Arbona, 2020; Wiedenmayer, 2026).
The MDL framework is operationalized through a specific ecosystem of three interconnected AI tools. Gemini serves as the primary text generation engine, producing source speech scripts calibrated to specified cognitive load parameters. Through structured prompting, Gemini can generate speeches at defined delivery density, with specific terminological profiles, on various domain topics, and for different simulated speaker type. Its flexibility as a large language model makes it particularly well-suited to the iterative refinement that pedagogical material design requires.
Google AI Studio provides the text-to-speech capability that transforms generated scripts into audio training materials. The platform offers control over voice profile—including accent, gender, and speaking style—delivery rate expressed in syllables per second, pause duration, and prosodic emphasis. This level of control allows trainers to engineer audio materials that closely approximate the delivery conditions specified in their design parameters, including the rapid-fire speech rates that Seeber (2011) notes are typically attainable only by mechanical manipulation of pre-recorded material.
NotebookLM, whose role is discussed in detail in the following subsection, provides the factual grounding layer of the ecosystem through a Retrieval-Augmented Generation architecture. Together, these three tools constitute what may be described as a closed design loop: the trainer assembles factual sources (NotebookLM), engineers their transformation into calibrated speech scripts (Gemini/NotebookLM), and realizes those scripts as pedagogically specified audio (Google AI Studio).
The most significant methodological innovation of the MDL framework—and the one that most directly distinguishes it from simpler applications of GenAI to material design—is its use of RAG through NotebookLM. This distinction warrants detailed explanation, as it addresses a concern that would otherwise fundamentally undermine the framework’s credibility as a tool for professional interpreter training.
Pure generative AI, however powerful as a language system, is prone to hallucination: the production of plausible but factually incorrect content. In many educational contexts, a degree of factual imprecision may be tolerable or even pedagogically productive. In conference interpreter training, it is not. Trainees learning to interpret speeches on climate policy, monetary economics, or public health must process and reproduce factual content accurately; the interpreter’s professional role is specifically defined by fidelity to the source. Training materials that embed factual errors do not merely fail to develop the target skill—they actively cultivate incorrect content processing habits and potentially introduce misinformation into the trainee’s domain knowledge base and long-term memory. The hallucination risk is therefore not a peripheral quality concern but a foundational methodological constraint.
NotebookLM operates on a RAG architecture: rather than generating content from the parametric knowledge embedded in its training weights, it grounds its outputs in a specific document corpus uploaded by the user. In the MDL framework, the trainer assembles a corpus of primary source documents—including United Nations reports, NGO briefings, news articles, policy documents, and official conference proceedings—and uploads this corpus to NotebookLM before initiating material generation. The resulting outputs are anchored in the factual content of these sources, can be directed to cite their sources, and can be quality-checked by the trainer against the uploaded materials. This architecture resolves the hallucination problem without abandoning the generative flexibility that makes AI-based material design superior to manual curation. It is, in effect, a hybrid model that combines the factual grounding of authentic materials with the pedagogical control of designed ones.
To illustrate the corpus curation logic of the RAG architecture, the two case studies in Section 7 drew on the following categories of primary source documentation. For the climate adaptation domain (Case Study 1), the NotebookLM corpus comprised: (a) intergovernmental scientific reports—specifically Chapter 9 (Africa) of the IPCC Sixth Assessment Report, Working Group II: Impacts, Adaptation and Vulnerability (2022), the most authoritative and terminologically dense scientific reference available for North African climate risk; (b) UN development agency briefings, including the UNDP Human Development Report 2023/24 and UNEP regional environmental outlook documents; (c) the African Union Climate Change and Resilient Development Strategy and Action Plan 2022–2032; and (d) the World Bank Morocco Country Climate and Development Report (2022). For the monetary policy domain (Case Study 2), the corpus comprised: (a) Bank Al-Maghrib Annual Reports (2021–2023); (b) IMF Article IV Consultation Reports for Morocco, Algeria, and Tunisia (2022–2023); (c) BIS Working Papers on monetary tightening in emerging market economies (2022–2023); and (d) selected policy speeches by Maghreb central bank governors drawn from institutional press release archives. These source categories are representative rather than exhaustive; trainers implementing the MDL should assemble corpora appropriate to their target domain and language pair, following the selection criteria discussed in Stage 1 of the workflow (Section 6)1.

Figure 1. RAG Framework within the MDL
The Material Digital Laboratory is a design environment in which interpreter trainers generate, modify, and sequence training materials in systematic alignment with cognitive and pedagogical objectives. Its defining feature is a shift in the trainer’s basic stance: from selection to engineering. Trainers do not search for materials that approximately realise their intentions—they construct materials that precisely instantiate them. The trainer moves from operating within the constraints of an existing material supply to functioning as the originator of a designed one.
This shift has several dimensions that distinguish the MDL from earlier approaches to technology-enhanced interpreter training. Unlike static CAIT tools such as speech repositories or authoring programs, the MDL is generative rather than archival—it produces new materials rather than organizing existing ones. Unlike raw GenAI applications, the MDL is epistemically controlled: its outputs are grounded in verified factual sources, making them appropriate for a training context that demands content accuracy. Unlike manual material development (Li, 2018), the MDL offers significant potential efficiency gains: a trainer with appropriate prompt literacy and a prepared source corpus may be able to produce a progressive sequence of materials targeting different cognitive load levels on a given topic considerably faster than conventional curation would allow. The extent of this advantage will vary with the trainer’s technical fluency, corpus availability, and the degree of quality review required. Nonetheless, it should not be assumed to accrue automatically or uniformly across institutional contexts, and unlike the mock conference formats described by Li (2015a) and Conde and Chouc (2019), the MDL is fully customizable at the level of individual cognitive parameters, enabling a precision of pedagogical targeting that live simulation cannot provide.
Traditional interpreter training practice can be characterized along two axes: material authenticity (authentic versus simplified) and material control (fixed versus adaptable). Authentic materials score high on realism but low on control; simplified materials score high on control but low on realism. The MDL is designed to reduce this tension by enabling materials that are simultaneously grounded in factual content and calibrated in cognitive load parameters through prompt-engineered generation. This is the framework’s core theoretical claim: that the apparent opposition between authenticity and pedagogical control reflects the constraints of pre-AI production methods rather than any intrinsic incompatibility. Whether the MDL succeeds in bridging this divide in a given implementation—and to what degree—will depend on the quality of the corpus assembled, the sophistication of the prompts deployed, and the trainer’s capacity for critical output review. These are empirical questions the present paper identifies but does not resolve. What the framework proposes is a plausible and theoretically grounded design logic for addressing the tension; its practical resolution remains contingent and must be evaluated through classroom practice.
Table 1. MDL Conceptual Skill Mapping Model
| Skill | Challenge | Parameter |
| Listening | Fast speech processing | Speech rate, accent |
| Memory | Information overload | Density, segmentation |
| Reformulation | Limited variations | Paraphrasing, redundancy |
| Anticipation | Predictive difficulty | Discourse structuring |
The MDL concept also has implications for the institutional organization of interpreter training programs. If material design is understood as a core trainer competency rather than a peripheral preparation task, it follows that programs should invest in trainer digital literacy development, in the curation and maintenance of domain-specific factual corpora, and in the systematic documentation of material design decisions. This creates a cumulative pedagogical infrastructure rather than relying on individual improvisation, and thus, addressing not just the material constraint but the knowledge management constraint that has historically made high-quality interpreter training a function of individual trainer expertise rather than institutional capacity.
The MDL is operationalized through a five-stage workflow that transforms a pedagogical objective into a deployable training material. Each stage involves specific decisions, tools, and quality controls that together ensure the output is both factually grounded and cognitively calibrated.

Figure 2. Workflow from Objective to Deployable Material
Stage 1: Corpus Curation
The trainer identifies the domain topic and the target skills for a given training unit and then assembles a corpus of primary source documents relevant to the domain. Selection criteria include factual reliability, linguistic register, terminological density, and cultural specificity appropriate to the target language pair and trainee level. For a unit targeting climate adaptation in North Africa, for instance, the corpus might include IPCC regional assessment chapters, UNDP briefings on Sahel water stress, African Union climate policy frameworks, and recent news coverage from Maghreb outlets. This corpus is uploaded to NotebookLM, establishing the factual grounding layer of the system. The trainer’s expertise in corpus evaluation (i.e., assessing source quality, register, and relevance) is the first and most critical exercise of pedagogical judgment in the MDL workflow.
Stage 2: Terminology Injection
Before generating the speech script, the trainer specifies a terminological profile for the target text. This involves identifying key domain terms, confirming culturally and regionally appropriate variants. The distinction between Moroccan Modern Standard Arabic and Middle Eastern usage patterns is pedagogically significant: terms that are highly activated in one regional variety may be less familiar in another, creating controlled comprehension challenges that develop lexical flexibility. Terminology is injected into the generation prompt as a required vocabulary set, ensuring that AI-generated content activates the specific lexical knowledge the trainer intends to develop. This stage operationalizes Wiedenmayer’s (2026) observation that the design of pedagogical speeches involves deliberate artificiality at the level of lexical selection.
Stage 3: AI Text Generation
The trainer formulates a structured prompt with the aid of Gemini (or other GenAI tool) specifying: (a) the speaker profile, including whether the simulated speaker is a native or non-native speaker, their institutional role, and their delivery style; (b) the cognitive load parameters, including delivery speed target, information density, redundancy level, and enumerative density; (c) the discourse structure; and (d) the required terminology from Stage 2. The prompt may additionally specify anticipated syntactic or semantic structures that enable anticipatory processing, or deliberately suppress them to increase difficulty. The formulated prompt is then used for NotebookLM to generate the target speech with the specified parameters. Outputs can be reviewed by the trainer for factual accuracy against the NotebookLM corpus and adjusted, if necessary, using follow-up prompts before finalization.
Stage 4: Audio Realization
The finalized script is uploaded to Google AI Studio for TTS synthesis. The trainer selects voice parameters—including accent, formality, gender, and speaking style—and adjusts delivery rate and pause structure to match the pedagogical specification. For high cognitive load exercises, delivery rate may be set at the upper range of professional norms (approximately 130–160 words per minute for English conference speech); for introductory exercises, rate may be reduced and pause duration extended to provide additional processing time. The output is an audio file that can be deployed directly in classroom delivery or uploaded to a digital learning environment. The separation of script production (Stage 3) from audio realization (Stage 4) is an important structural feature of the workflow: it enables the same script to be rendered at multiple delivery speeds, with different accent profiles, or for different modes of interpretation without requiring text regeneration.
Stage 5: Pedagogical Alignment
The completed materials are mapped to the training unit’s skill objectives, documenting the target cognitive effort(s); the specific problem triggers designed into the material; the expected strategies trainees should deploy; and the assessment criteria against which performance will be evaluated. This documentation creates a transparent and auditable record of design decisions, enabling systematic iteration and the cumulative development of an institutional material library. Stage 5 is the point at which the MDL workflow converges with established interpreter training assessment practice, connecting the designed materials to the observable competence targets that define progress in the training program.

Figure 3. Prompt Engineering Architecture
Domain and Objectives
The first case study simulates a high-level intergovernmental conference on climate adaptation in North Africa, targeting intermediate-level conference interpreting trainees working in the English-Arabic direction. The primary cognitive targets are the Listening and Analysis Effort—through high terminological density and non-native speaker delivery conditions—and the Memory Effort, through extended enumeration sequences and the density of statistical data points requiring accurate numerical retention.
Corpus Description
The factual grounding corpus assembled for this case study includes the IPCC Sixth Assessment Report regional chapters on North Africa and the Mediterranean; UNDP Human Development Report data on Sahel water stress; the African Union Climate Change and Resilient Development Strategy; World Bank agricultural sector vulnerability assessments; and recent news reporting from Reuters Arabic, Al Jazeera English, and Morocco World News covering the 2023–2024 drought season. These sources were selected to provide a high-density terminological environment representative of professional conference speech in this domain, while ensuring factual grounding in authoritative institutional sources.
Terminology Profile
Key terms injected into the generation prompt include desertification, water stress index, climate resilience, carbon sequestration, adaptive capacity, and food sovereignty, with Arabic equivalents specified for each. Regional terminological specificity includes Moroccan institutional names—Agence du Bassin Hydraulique de Souss-Massa, Ministère de la Transition Énergétique et du Développement Durable—that require trainee familiarity with North African administrative discourse and would not be available through standard LLM parametric knowledge.
Prompt Design
Three prompt variants were developed for this case study, representing different cognitive load levels and different aspects of the Effort Model framework:
Representative Output Excerpt (Base Prompt)
Distinguished delegates, the Souss-Massa basin has lost twenty-three percent of its groundwater reserves over the past fifteen years—a figure that should concern every actor in this room. Our government has committed, through the National Water Plan 2050, to reducing per-capita agricultural water consumption by forty percent before 2035. This will require a fundamental restructuring of irrigation incentives, investment in drip technology deployment across one hundred and eighty thousand hectares of citrus cultivation, and a renegotiation of informal water rights in the piedmont zone. We cannot achieve climate resilience through infrastructure alone. We need governance reform, and we need it now.
This excerpt illustrates the MDL’s capacity to generate professional-grade conference speech on a domain-specific, factually grounded topic with terminological precision and discourse structure calibrated to the pedagogical specification. The presence of specific numerical data, institutional references, and policy instrument terminology demonstrates the grounding effect of the RAG architecture.
Audio Design
The base prompt script was rendered in Google AI Studio using a formal male voice profile with a slight Moroccan French influence setting, at 122 words per minute with standard pause duration at major clause boundaries. The high-load script was rendered at 144 words per minute with a female voice profile and reduced pause duration, targeting the Memory Effort through compressed processing windows and eliminating the recovery time that standard pause patterns would provide.
The second case study simulates a central bank governors’ forum on monetary policy in the post-pandemic Maghreb region, targeting advanced trainees working in the French-Arabic direction in consecutive mode. The primary cognitive targets are the Production Effort—through high reformulation pressure in specialized financial discourse—and the Coordination Effort, through panel dynamics requiring sustained note management across multiple speakers. This case study deliberately differs from Case Study 1 in domain, language direction, mode, and primary cognitive target, demonstrating the framework’s scalability across qualitatively different training objectives.
The financial domain introduces a distinct terminological register, including instruments such as quantitative easing, yield curve control, and macroprudential policy, with French and Arabic equivalents that are not interchangeable across regional varieties. The consecutive mode shifts the primary cognitive challenge from simultaneous attention allocation to structured memory and accurate numerical reformulation under conditions of high terminological density, where the Production Effort is maximally stressed by the requirement to render specialized financial concepts in a linguistically constrained target language register.
This section presents the prototype outputs produced through the MDL workflow across both case studies and records analytical observations about the framework’s apparent capacities with respect to cognitive load control, skill targeting, factual accuracy, and simulation realism. These are observations about what the design process produced; they are not measurements of learning outcomes or evidence of instructional effectiveness. They are offered to demonstrate the operational feasibility and internal coherence of the MDL framework as a design proposition.
For Case Study 1, the MDL prototype produced six speech scripts at three cognitive load levels—base, high-load, and speaker variation—in English with Arabic terminology embedded; six corresponding audio files rendered at specified delivery parameters; one integrated mock conference script combining all three speaker types in a simulated panel format; and a skill mapping document linking each material to the target effort parameters, designed problem triggers, and expected trainee strategies. For Case Study 2, four scripts in French were produced targeting consecutive interpreting mode at two cognitive load levels, with four audio files and a structured bilingual terminology glossary for the financial domain, comprising parallel French and Arabic terms with contextual usage notes.
The prototype outputs suggest that the workflow supports consistent, specifiable control over cognitive load parameters across both case studies. Comparison of base and high-load scripts for Case Study 1 reveals measurable and predictable differences in information density (approximately 3.2 versus 6.1 propositions per clause), delivery rate (122 versus 144 words per minute in audio realization), and numerical data density (2.1 versus 8.6 data points per minute). These differences map directly to the design specifications, confirming that the prompt engineering layer successfully translates cognitive objectives into textual features in a manner consistent with what Seeber and Kerzel (2012) describe as the controlled manipulation of prosodic and syntactic features for experimental purposes.
Different script variants produced qualitatively different demands on the effort parameters identified in the theoretical framework. High-load scripts activated primarily the Memory Effort through dense enumeration and low discourse coherence; speaker variation scripts activated the Listening and Analysis Effort through register switching and accent variation requiring rapid adaptation of processing strategies; and consecutive mode scripts for Case Study 2 activated the Production Effort through high reformulation pressure in a specialized register with limited near-equivalent options in the target language. The deliberate suppression of logical connectives in the high-load French-Arabic scripts created the processing pressure that Li (2015b) identifies as the condition in which compensatory strategies become cognitively necessary—the intended outcome of this design choice.
In the Case Study 1 prototype outputs, the RAG architecture’s grounding effect appeared operationally robust: outputs incorporated institutional names, policy framework titles, and statistical data drawn from the uploaded source corpus. No fabricated statistics were identified during trainer quality review of materials generated through NotebookLM. This observation is consistent with the established empirical literature on RAG systems, though systematic independent verification against source documents would be required before stronger claims about hallucination reduction could be advanced. This outcome is consistent with the established empirical literature on RAG systems, which demonstrates substantially reduced hallucination rates relative to pure generative approaches. Quality review of outputs against uploaded source documents identified only minor inconsistencies, all resolved through prompt refinement at Stage 3 of the workflow.
The prototype outputs were evaluated by the authors against four criteria that are operationalized within the MDL framework’s own design logic, rather than through independent expert review or benchmark corpus comparison—both of which are identified as priorities for the empirical phase of the research program.
Discourse coherence: outputs were inspected for thematic progression, logical connective density consistent with the prompt specification, and structural organization appropriate to the simulated speech genre. Coherence was treated as satisfactory where outputs could be parsed as linguistically and informationally well-formed without editorial intervention.
Terminological precision: outputs were checked against the terminology injection list specified in each prompt variant, verifying that all required terms appeared in contextually appropriate collocations and that institutional names and policy instrument designations matched those in the uploaded NotebookLM source documents.
Cognitive load calibration: measurable textual features—propositions per clause, word count per minute at specified delivery rates, numeral density, and frequency of logical connectives—were compared against the design specifications. Outputs were considered calibrated where measured features fell within a 10% margin of the specified target values.
Register appropriateness: lexical formality, syntactic complexity, and discourse genre markers were evaluated against the simulated speaker profile and institutional context specified in each prompt.
These criteria are explicit but informal, for they reflect the design logic of the MDL framework rather than validated psychometric standards. Their function at this stage is to demonstrate that the prototype outputs are internally consistent with the framework’s own specifications, not to establish their effectiveness as training materials. External validation by domain experts and comparison against authentic professional speech corpora remain necessary before the framework can be considered pedagogically validated.
The MDL framework represents a theoretically grounded and operationally feasible advance over previous approaches to the material constraint problem in interpreter training. Earlier proposals—from Li’s (2018) manual simplification model to Seeber and Arbona’s (2020) engineered audiovisual recordings—identified the same core problem: that materials are rarely designed to target specific cognitive efforts. The MDL addresses both the theoretical gap, by grounding material design in a cognitive objective framework derived from Gile, and the operational gap, by specifying a five-stage workflow that any technologically literate trainer can adopt using commercially available tools.
The comparison with Hatiarová (2025) is instructive. That study introduces prompt crafting as a pedagogical act and demonstrates AI-TTS workflows for material generation, but does not systematically address the hallucination risk or provide a theoretical framework linking design decisions to cognitive objectives. The MDL’s RAG architecture and Effort Model parameter mapping constitute substantive additions to this emerging practice. Wiedenmayer (2026) provides the most aligned recent precedent, arguing that pedagogical speeches require deliberate artificiality in service of predefined instructional objectives. The present paper extends this argument in two directions: backward, by grounding the claim in Gile’s cognitive framework with explicit parameter mapping; and forward, by specifying the RAG architecture as the mechanism for preserving factual grounding within a generative design environment. Together, these extensions transform Wiedenmayer’s important intuition into an operationalizable methodology.
Against the broader AI in education literature, the MDL’s most distinctive contribution is its disciplinary specificity. General frameworks for AI in education (Bakar & Tapsoba, 2026; Qian, 2025; Yan et al., 2024) provide valuable theoretical vocabulary but cannot specify the domain-specific design parameters—effort parameters, problem triggers, strategy targets—that make material design pedagogically productive in interpreter training. The MDL is grounded in a theoretical tradition unique to the discipline and therefore offers what general frameworks cannot: a principled basis for design decisions that can be communicated, justified, and evaluated within the field’s established conceptual vocabulary.
The framework’s primary strengths are its theoretical coherence, its practical operationalizability, and its scalability. By grounding material design in Gile’s Effort Models, the MDL provides a principled basis for design decisions that can be communicated, justified, and refined within the discipline. By specifying a five-stage workflow using commercially available tools, it provides a practical entry point for trainers without specialized technical backgrounds. And by enabling a single trainer to generate an entire progressive sequence of materials with high cognitive load precision, it addresses resource scarcity at the scale at which it actually operates in training programs.
The framework also has important limitations that must be acknowledged. First, this paper presents a conceptual and design-level contribution; it does not report empirical evidence on learning outcomes. Whether materials designed according to the MDL’s principles produce better interpreting performance than conventionally prepared materials is an empirical question that the present paper cannot answer, and quasi-experimental studies addressing this question are urgently needed. Second, output quality is contingent on prompt quality, which depends on the trainer’s cognitive load literacy and domain knowledge. A trainer without strong theoretical grounding in Gile’s models may produce materials that technically satisfy specified parameters without being pedagogically effective. Third, the framework currently relies on proprietary tools whose availability, functionality, and pricing may change. The development of open-source alternatives would significantly enhance the framework’s accessibility, particularly for programs in lower-resource institutional contexts.
A fourth category of limitation concerns AI-specific risks that the MDL must actively manage. Bias in corpus composition is one such risk: the factual grounding provided by RAG is only as balanced as the source documents the trainer selects, and an inadvertently narrow corpus may introduce systematic distortions into the register, terminology, or epistemic assumptions of generated materials. Overreliance is a second risk: if trainers treat AI generation as the default rather than as a mediated tool, the critical judgment required for quality review may atrophy, reducing the MDL from a pedagogical engineering environment to an automated production pipeline. Finally, the limits of factual grounding within RAG architectures are real: even curated documents may contain outdated statistics, contested interpretations, or coverage gaps that the system cannot independently flag. The trainer’s expert review at Stage 3 of the workflow is therefore not a procedural formality but an epistemological safeguard—one that requires the sustained domain knowledge and pedagogical literacy that the MDL framework is designed to leverage, not replace.
Perhaps the most consequential implication of the MDL framework is for trainer professional identity. The framework requires trainers to function not as expert interpreters transmitting tacit professional knowledge but as pedagogical engineers designing cognitive environments. This is a substantially different competence profile, encompassing domain knowledge, cognitive load theory, prompt engineering literacy, and audio production capability—a hybrid expertise (Bakar & Tapsoba, 2026; Yan et al., 2024) that no existing training program systematically develops.
This competence shift has institutional implications. Training programs wishing to operationalize the MDL must invest in trainer development alongside student development, creating structured pathways for trainers to acquire prompt literacy, cognitive load design skills, and domain-specific corpus curation capabilities. The alternative—expecting trainers to develop these skills through individual experimentation—replicates at the level of trainer competence precisely the resource scarcity the framework is designed to overcome at the level of training materials.
It is essential to emphasize that the MDL does not diminish the trainer’s linguistic expertise or pedagogical judgment. As Adamakis and Rachiotis (2025) argue, GenAI is a technology of selection whose pedagogical value is entirely a function of the human choices that direct it. The trainer’s role in the MDL is transformed, not reduced—from a role in which expertise operates at the point of delivery to one in which it operates at the point of design. This transformation is consistent with a long trajectory in the literature: Sandrelli and de Manuel Jerez (2007) argued two decades ago that the value of CAIT tools was entirely a function of what trainers put into them, and the MDL makes explicit what trainers must now be able to put in.

Figure 4. Trainer Competence Shift — From Practitioner-Instructor to Pedagogical Engineer
The MDL framework offers substantial potential scalability advantages, though these are conditional rather than automatic. Material libraries developed within the MDL can, in principle, be systematically varied, extended, and refined—and the documentation of design decisions creates the conditions for a cumulative institutional resource. In practice, however, scalability will be bounded by the trainer’s prompt design capacity; the availability of reliable domain-specific source corpora; institutional access to the required platforms; and the time investment required to quality-assure AI outputs before classroom deployment. Programs considering MDL adoption should anticipate an initial investment phase before efficiency gains are reliably realised.
Institutional integration of the MDL does, however, require attention to the digital divide that Adamakis and Rachiotis (2025) identify as a persistent feature of AI adoption in higher education. Programs in contexts with limited technological infrastructure, restricted internet access, or acute resource constraints may face implementation barriers that the framework’s current design does not address.
This paper has proposed the Material Digital Laboratory as a conceptual and practical framework for AI-enabled, skill-oriented interpreter training. By extending Gile’s (1995, 2009) Effort Models as design parameters, it has articulated a principled basis for engineering training materials that target specific cognitive objectives rather than relying on the incidental cognitive demands of authentic material exposure. By specifying a five-stage workflow integrating Gemini, NotebookLM, and Google AI Studio, it has demonstrated the operational feasibility of this approach. By developing two illustrative case studies across different domains, language directions, and skill targets, it has shown that the framework can be applicable across the range of conditions that characterize professional conference interpreter training.
The paper’s contributions attempt to address the most persistent structural problem in the field: that training materials are almost never designed to match the cognitive objectives they are meant to develop. Generative AI, properly grounded through RAG and pedagogically directed through Effort Model-based parameter mapping, makes systematic material design feasible at the scale that trainers actually require.
The proposition this paper makes is ultimately a simple one: the persistent material constraint in interpreter training is no longer a constraint of availability but a constraint of design. The MDL provides the framework for addressing it. The most urgent priority for future research is empirical validation. Quasi-experimental studies comparing learning outcomes in MDL-designed training sequences against conventional practice are needed, as are studies examining the trajectory of trainer competence development in digital laboratory environments. Further, future research should also prioritize longitudinal studies examining how MDL materials age as domain knowledge evolves. The Material Digital Laboratory can be seen as both a concept and an operational reality of a new phase in interpreter training pedagogy in which trainers can be engineers of learning environments rather than merely a curator of found objects.
Data and Materials Availability Statement: All prompts, generated speech scripts, audio files, terminology glossaries, and case study documentation produced in this study are available upon request from the corresponding author. Materials are organized by case study, cognitive load level, skill target, and audio parameter specifications.
Adamakis, M., & Rachiotis, T. (2025). Artificial intelligence in higher education: A state-of-the-art overview of pedagogical integrity, artificial intelligence literacy, and policy integration. Encyclopedia, 5(1), 180. https://doi.org/10.3390/encyclopedia5040180
Al-Suhaim, D. S. (2024). Exploring theoretical dimensions in interpreting studies: A comprehensive overview. Arab World English Journal for Translation & Literary Studies, 8(1), 15-43.
Bakar, S., & Tapsoba, R. (2026). Artificial intelligence in classroom teaching: Prospects, challenges and framework for responsibly orchestrated mediation. Social Science Chronicle, 6(1), 1-21. https://doi.org/10.56106/ssc.2026.002
Cai, R., Dong, Y., Zhao, N., & Lin, J. (2015). Factors contributing to individual differences in the development of consecutive interpreting competence for beginner student interpreters. The Interpreter and Translator Trainer, 9(1), 104-120.
Carrasco-Sáez, J. L., Contreras-Saavedra, C., San-Martín-Quiroga, S., Contreras-Saavedra, C. E., & Viveros-Muñoz, R. (2025). Analyzing higher education students’ prompting techniques and their impact on ChatGPT’s performance: An exploratory study in Spanish. Applied Sciences, 15(7651). https://doi.org/10.3390/app15147651
Chan, C. H. Y. (2013). From self-interpreting to real interpreting: A new web-based exercise to launch effective interpreting training. Perspectives: Studies in Translatology, 21(3), 358-377. https://doi.org/10.1080/0907676X.2012.657654
Chan, C. K. Y. (2023). A comprehensive AI policy education framework for university teaching and learning. International Journal of Educational Technology in Higher Education, 20(38). https://doi.org/10.1186/s41239-023-00408-3
Chan, V. (2023). Investigating the impact of a virtual reality mobile application on learners’ interpreting competence. Journal of Computer Assisted Learning, 1-17. https://doi.org/10.1111/jcal.12796
Chang, C.-C., & Wu, M. M.-C. (2017). From conference venue to classroom: The use of guided conference observation to enhance interpreter training. The Interpreter and Translator Trainer, 11(1), 21-42. https://doi.org/10.1080/1750399X.2017.1359759
Conde, J. M., & Chouc, F. (2019). Multilingual mock conferences: A valuable tool in the training of conference interpreters. The Interpreters’ Newsletter, 24, 1-17.
Cordero, J., Torres-Zambrano, J., & Cordero-Castillo, A. (2025). Integration of generative artificial intelligence in higher education: Best practices. Education Sciences, 15(1), 32. https://doi.org/10.3390/educsci15010032
Corpas Pastor, G. (2018). Tools for interpreters: The challenges that lie ahead. Current Trends in Translation Teaching and Learning E, 5, 157-182.
Corpas Pastor, G. (2020). Language technology for interpreters: The VIP Project. Proceedings of the 42nd Conference “Translating and the Computer” (TC42).
Cui, F., Li, D., & Zhuang, C. (2025). Introduction: Transforming translation education through artificial intelligence. The Interpreter and Translator Trainer, 19(3-4), 227-233. https://doi.org/10.1080/1750399X.2025.2561258
Djovcos, M., Klabal, O., & Sveda, P. (2023). Training interpreters: Old and new challenges. Bridge: Trends and Traditions in Translation and Interpreting Studies, 4(1), 1-12.
Fan, D. (2012). The development of expertise in interpreting through self-regulated learning for trainee interpreters [Doctoral dissertation]. University of Newcastle upon Tyne.
Frittella, F. M. (2021). Computer-assisted conference interpreter training: Limitations and future directions. Journal of Translation Studies, 2(2021), 103-142. https://doi.org/10.3726/JTS022021.6
Garcia-Penalvo, F. J. (2023). Generative artificial intelligence: Open challenges, opportunities, and risks in higher education. CEUR Workshop Proceedings, 3696, 4-15.
Gile, D. (1995). Basic concepts and models for interpreter and translator training. John Benjamins.
Gile, D. (1999). Testing the Effort Models’ tightrope hypothesis in simultaneous interpreting: A contribution. Hermes, Journal of Linguistics, 23, 153-172.
Gile, D. (2009). Basic concepts and models for interpreter and translator training (Rev. ed.). John Benjamins.
Gile, D. (2021). The Effort Models of interpreting as a didactic construct. In R. Munoz Martin et al. (Eds.), Advances in cognitive translation studies (pp. 139-153). Springer Nature Singapore.
Hatiarová, P. (2025). AI in interpreting training. L10N Journal, 1(4), 45-66.
Kalina, S. (2000). Interpreting competences as a basis and a goal for teaching. Fachhochschule Koln.
Kasneci, E., Sessler, K., Kuchemann, S., et al. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274.
Li, X. (2015a). Mock conference as a situated learning activity in interpreter training: A case study of its design and effect as perceived by trainee interpreters. The Interpreter and Translator Trainer, 9(3), 323-341.
Li, X. (2015b). Putting interpreting strategies in their place: Justifications for teaching strategies in interpreter training. Babel, 61(2), 170-192. https://doi.org/10.1075/babel.61.2.02li
Li, X. (2018). Material development principles in undergraduate translator and interpreter training: Balancing between professional realism and classroom realism. The Interpreter and Translator Trainer, 12(4), 369-389.
Macnamara, B. N., Moore, A. B., Kegl, J. A., & Conway, A. R. A. (2011). Domain-general cognitive abilities and simultaneous interpreting skill. Interpreting, 13(1), 121-142.
Moser-Mercer, B. (2000/01). Simultaneous interpreting: Cognitive potential and limitations. Interpreting, 5(2), 83-94.
Moukatib, M., & Ben Seddik, A. (2026). The role of AI in translator training: Assessing AI’s influence on translation education and professional training. International Journal of Linguistics and Translation Studies, 7(1), 123-143. https://doi.org/10.36892/ijlts.v7i1.669
Munoz-Basols, J., Neville, C., Lafford, B. A., & Godev, C. (2023). Potentialities of applied translation for language learning in the era of artificial intelligence. Hispania, 106(2), 171-194.
Parrilla Gomez, L., & Postigo Pinazo, E. (2025). Artificial intelligence in the training of public service interpreters. Language & Communication, 103, 86-107.
Purba, S. W. D., Silitonga, B. N., & Yang, J. J. (2025). AI-assisted learning: A systematic review. Turkish Online Journal of Distance Education (TOJDE), 26(4), 77-94.
Qian, Y. (2025). Pedagogical applications of generative AI in higher education: A systematic review of the field. TechTrends, 69, 1105-1120.
Rybina, N. V., Koshil, N. Ye., & Hyryla, O. S. (2025). Artificial intelligence and translation in English language teaching: Opportunities and challenges. Medychna osvita [Medical Education], (2), 87-91. https://doi.org/10.11603/m.2414-5998.2025.2.15494
Sachtleben, A. (2015). Pedagogy for the multilingual classroom: Interpreting education. Translation & Interpreting: The International Journal of Translation and Interpreting Research, 7(2), 51-59.
Sandrelli, A., & de Manuel Jerez, J. (2007). The impact of information and communication technology on interpreter training: State-of-the-art and future prospects. The Interpreter and Translator Trainer, 1(2), 269-303.
Seeber, K. G. (2011). Cognitive load in simultaneous interpreting: Existing theories—New models. Interpreting, 13(2), 176-204. https://doi.org/10.1075/intp.13.2.02see
Seeber, K. G., & Arbona, E. (2020). What’s load got to do with it? A cognitive-ergonomic training model of simultaneous interpreting. The Interpreter and Translator Trainer, 14(3), 1-18. https://doi.org/10.1080/1750399X.2020.1839996
Seeber, K. G., & Kerzel, D. (2012). Cognitive load in simultaneous interpreting: Model meets data. International Journal of Bilingualism, 16(2), 228-242.
Serra, P., & Oliveira, A. (2025). AI-powered prompt engineering for Education 4.0: Transforming digital resources into engaging learning experiences. Education Sciences, 15(12), 1640. https://doi.org/10.3390/educsci15121640
Shahzad, T., Mazhar, T., Tariq, M. U., Ahmad, W., Ouahada, K., & Hamam, H. (2025). A comprehensive review of large language models: Issues and solutions in learning environments. Discover Sustainability, 6(27).
Sharma, S., Mittal, P., Kumar, M., & Bhardwaj, V. (2025). The role of large language models in personalized learning: A systematic review of educational impact. Discover Sustainability, 6(243). https://doi.org/10.1007/s43621-025-01094-z
Song, X., & Tang, M. (2020). An empirical study on the impact of pre-interpreting preparation on business interpreting under Gile’s Efforts Model. Theory and Practice in Language Studies, 10(12), 1640-1650. http://dx.doi.org/10.17507/tpls.1012.19
Vieira, N. G. S. (2015). E-learning practices in translation and interpretation: Corpora as training platforms. Procedia—Social and Behavioral Sciences, 198, 157-164.
Wang, B. (2015). Bridging the gap between interpreting classrooms and real-world interpreting. International Journal of Interpreter Education, 7(1), 65-73.
Wiedenmayer, A. (2026). Artificial intelligence as a pedagogical tool for speech generation in conference interpreter training. [Journal details pending publication].
Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., Li, X., Jin, Y., & Gasevic, D. (2024). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1), 90-112. https://doi.org/10.1111/bjet.13370
Yang, C., Chen, J., & Zou, D. (2025). Artificial intelligence in interpreting education curriculum: A Delphi study for interpreter competencies. International Journal of Education and Humanities (IJEH), 5(3), 387-403.
Yang, C., Hou, S., Zhao, M., Yan, J., & Chen, J. (2026). Translation students’ perceptions of the integration of artificial intelligence in translation education: A constructivist approach. Artificial Intelligence in Education, 2(2), 157-174. https://doi.org/10.1108/AHE-06-2025-0087
Yusuf, H., Money, A., & Daylamani-Zad, D. (2025). Pedagogical AI conversational agents in higher education: A conceptual framework and survey of the state of the art. Education and Information Technologies, 73, 815-874.
Zhao, N. (2022). Use of computer-assisted interpreting tools in conference interpreting: Training and practice during COVID-19. In K. Liu & A. K. F. Cheung (Eds.), Translation and interpreting in the age of COVID-19 (pp. 331-347). Springer Nature Singapore. https://doi.org/10.1007/978-981-19-6680-4_17
1 All figures in this article are the authors’ own