Journal Description
Multimodal Technologies and Interaction
Multimodal Technologies and Interaction
is an international, peer-reviewed, open access journal on multimodal technologies and interaction published monthly online by MDPI.
- Open Access— free for readers, with article processing charges (APC) paid by authors or their institutions.
- High Visibility: indexed within Scopus, ESCI (Web of Science), Inspec, dblp Computer Science Bibliography, and other databases.
- Journal Rank: JCR - Q2 (Computer Science, Cybernetics) / CiteScore - Q1 (Neuroscience (miscellaneous))
- Rapid Publication: manuscripts are peer-reviewed and a first decision is provided to authors approximately 22.8 days after submission; acceptance to publication is undertaken in 4.6 days (median values for papers published in this journal in the first half of 2026).
- Recognition of Reviewers: reviewers who provide timely, thorough peer-review reports receive vouchers entitling them to a discount on the APC of their next publication in any MDPI journal, in appreciation of the work done.
- Journal Cluster of Artificial Intelligence: AI, AI in Medicine, Algorithms, BDCC, MAKE, MTI, Stats, Virtual Worlds, Computers and Journal of Superintelligence.
Impact Factor:
3.3 (2025);
5-Year Impact Factor:
3.4 (2025)
Latest Articles
A Mixed-Methods Feasibility Pilot Study of Medimon for High School Endocrine Education
Multimodal Technol. Interact. 2026, 10(8), 84; https://doi.org/10.3390/mti10080084 (registering DOI) - 8 Aug 2026
Abstract
Medimon is an educational role-playing game that integrates visual mnemonics, collectible creatures, disease states, treatment items, and artificial intelligence-powered non-player characters (AI-NPCs) to teach biomedical concepts. This single-arm mixed-methods feasibility pilot examined recruitment, gameplay uptake and progression, player experience, descriptive knowledge outcomes, and
[...] Read more.
Medimon is an educational role-playing game that integrates visual mnemonics, collectible creatures, disease states, treatment items, and artificial intelligence-powered non-player characters (AI-NPCs) to teach biomedical concepts. This single-arm mixed-methods feasibility pilot examined recruitment, gameplay uptake and progression, player experience, descriptive knowledge outcomes, and AI-NPC interactions during the Medimon endocrine level among high school students. Thirty students consented; 10 generated gameplay data, six completed paired 12-item pretests and posttests, and eight completed a PXI-style survey. Students played independently using a Steam-based build and encountered thyroid, pancreas, and adrenal educational content through exploration, quests, battles, and Medimon collection. Mean paired knowledge scores increased from 19.4% at pretest to 44.4% at posttest, although the change was not statistically significant in this small exploratory sample (p = 0.156). Audiovisual appeal, curiosity, and enjoyment were rated favorably, whereas Progress Feedback was the lowest-rated player-experience domain. Students who completed the posttest demonstrated greater gameplay duration, exploration, quest completion, and Medimon exposure than non-completers. Thematic analysis identified a mismatch between player expectations and AI-NPC scope, including fabricated quests, spatial directions, and game mechanics that introduced unreliable guidance into the educational environment. The study identified barriers related to gameplay uptake, progression feedback, posttest completion, and AI-NPC reliability that should be addressed before the educational efficacy of Medimon is evaluated in a larger controlled study.
Full article
(This article belongs to the Special Issue Technology-Enhanced Game-Based Approaches in Education: Learning, Emotions, and Motivation)
►
Show Figures
Open AccessArticle
From Atmospheric Tension to Embodied Regulation: A Mixed-Methods Study of Fear in VR and Non-VR Survival Horror Gameplay
by
Jianguo Fang and Yuanhao Liang
Multimodal Technol. Interact. 2026, 10(8), 83; https://doi.org/10.3390/mti10080083 - 6 Aug 2026
Abstract
►▼
Show Figures
Virtual reality (VR) survival horror is often discussed in terms of heightened fear and immersion, yet less attention has been paid to how fear is organized across different atmospheric conditions and how this organization differs from non-VR gameplay. This study approaches immersive fear
[...] Read more.
Virtual reality (VR) survival horror is often discussed in terms of heightened fear and immersion, yet less attention has been paid to how fear is organized across different atmospheric conditions and how this organization differs from non-VR gameplay. This study approaches immersive fear as a process emerging from the interaction among atmospheric configuration, embodied regulation, and post-play interpretation. A sequential mixed-methods design was employed using Resident Evil Village as the empirical context. Study 1 combined scene-based observation, synchronized gameplay recordings, and post-play interviews with eight participants to examine how fear was enacted across three contrasted atmospheric configurations: combat pressure, psychological ambiguity, and spatial disorientation. Study 2 extended this analysis through a within-subject experiment with 30 participants who completed both VR and non-VR versions of the same gameplay content under standardized conditions. The findings show that immersive fear is not a uniform increase in emotional intensity. In Study 1, different atmospheric configurations elicited distinct modes of embodied regulation, including defensive retreat, hesitant exposure, and cautious reorientation, while behavioral responses and retrospective accounts often diverged in systematic ways. In Study 2, paired-samples tests showed that VR produced lower valence, higher arousal, reduced perceived control, higher fear ratings, stronger immersion, and greater motion sickness than non-VR gameplay. Although VR increased fear ratings across all scenes, the display mode × scene interaction was not significant; descriptively, psychologically ambiguous environments produced the highest absolute fear ratings under VR. Across both studies, prior VR and genre experience appeared to shape how players interpreted and narrated threat, while short-term residual effects suggested that fear may extend beyond gameplay. These results suggest that VR modifies not only the intensity but also the organization of fear, while the scene-level and experience-related patterns should be interpreted cautiously. More broadly, the study reframes immersive fear as a temporally distributed process linking atmospheric configuration, embodied regulation, and post-play interpretation in survival horror gameplay.
Full article

Figure 1
Open AccessArticle
Prompt-Driven Fuzzing Debiasing Framework for Robust Visual Question Answering
by
Yali Fan, Gangyu Huang, Qiwen Lu and Shengbo Chen
Multimodal Technol. Interact. 2026, 10(8), 82; https://doi.org/10.3390/mti10080082 - 31 Jul 2026
Abstract
►▼
Show Figures
Visual Question Answering (VQA) systems have achieved impressive performance with the rise of large-scale vision–language models (VLMs). However, these models remain vulnerable to multiple forms of multimodal bias, severely limiting their robustness and generalization. Existing debiasing techniques mainly depend on post hoc evaluation
[...] Read more.
Visual Question Answering (VQA) systems have achieved impressive performance with the rise of large-scale vision–language models (VLMs). However, these models remain vulnerable to multiple forms of multimodal bias, severely limiting their robustness and generalization. Existing debiasing techniques mainly depend on post hoc evaluation or architectural modifications, while recent prompt-learning-based methods reveal new opportunities for aligning downstream tasks with pretrained models. In this work, we propose a unified prompt-driven debiasing framework that integrates generative prompt learning and a fuzzing-based bias correction mechanism. The generative prompt component reformulates VQA as a cloze-style masked prediction problem, leveraging pretrained language priors to improve semantic grounding. Meanwhile, the fuzzing-based module actively constructs unexpected test samples during training and employs a reflection mechanism to correct biased predictions in-loop, yielding inference-time robustness without additional test-time components. Extensive experiments on VQA-v2, VQA-CP, VQA-CE, GQA-OOD, and VQA-VS demonstrate that the proposed framework significantly improves both in-distribution (ID) accuracy and out-of-distribution (OOD) robustness, outperforming existing prompt-only or data-augmentation-only debiasing methods.
Full article

Figure 1
Open AccessReview
Co-Design Approaches in Immersive Virtual Reality for Upper Limb Stroke Rehabilitation: A Narrative Review
by
Dimosthenis Lygouras, Avgoustos Tsinakos, Ioannis Seimenis and Konstantinos Vadikolias
Multimodal Technol. Interact. 2026, 10(8), 81; https://doi.org/10.3390/mti10080081 - 29 Jul 2026
Abstract
►▼
Show Figures
Stroke-related upper-limb impairment is a leading cause of long-term disability, and immersive virtual reality (IVR) has emerged as a promising tool to enhance rehabilitation intensity and engagement. However, the effectiveness of IVR systems depends not only on technological features but also on how
[...] Read more.
Stroke-related upper-limb impairment is a leading cause of long-term disability, and immersive virtual reality (IVR) has emerged as a promising tool to enhance rehabilitation intensity and engagement. However, the effectiveness of IVR systems depends not only on technological features but also on how they are designed with end-user involvement. This narrative review examines co-design approaches used in the development of IVR systems for upper-limb stroke rehabilitation and evaluates how stakeholders contribute to system design and outcomes. A literature search was conducted in PubMed, Scopus, and IEEE Xplore, and studies involving IVR upper-limb rehabilitation with participatory, user-centered, or human-centered design were included. Thirteen studies met the inclusion criteria. Most studies demonstrated intermediate levels of participation, with seven classified as Level 4 (Collaborate), three as Level 3 (Involve), one as Level 2 (Consult), one as Level 1 (Inform), and one as Level 5 (Co-create). Common design principles included task-specific training, gamification, personalization, adaptive difficulty, meaningful activities, and clinician involvement. Despite positive findings regarding usability, motivation, and acceptability, most studies remained at feasibility or prototype stages, with limited evidence of clinical effectiveness. Recurring barriers included technological complexity, fatigue, cognitive demands, infrastructure limitations, and challenges in clinical integration. Overall, co-design is increasingly incorporated into IVR stroke rehabilitation development. However, sustained stakeholder partnership throughout the innovation lifecycle remains uncommon. Future research should prioritize longitudinal co-design processes and rigorous clinical evaluation to support the translation of IVR systems into routine rehabilitation practice.
Full article

Figure 1
Open AccessArticle
A Kinect-Based Educational Game for Embodied Activities Associated with Psychomotor Skills and Computational Thinking in Early Childhood: A Pilot Feasibility Study
by
Walter Choquehuanca-Quispe, Alexis Edmundo Gallegos-Acosta, Hector Cardona-Reyes and Claudia Acra-Despradel
Multimodal Technol. Interact. 2026, 10(8), 80; https://doi.org/10.3390/mti10080080 - 28 Jul 2026
Abstract
Early childhood education increasingly explores embodied learning, yet deploying markerless motion-controlled tools in real classrooms presents unverified operational and technical challenges regarding system stability and user interaction. This study evaluates the technical feasibility, system usability, and operational viability of a markerless, Kinect-based educational
[...] Read more.
Early childhood education increasingly explores embodied learning, yet deploying markerless motion-controlled tools in real classrooms presents unverified operational and technical challenges regarding system stability and user interaction. This study evaluates the technical feasibility, system usability, and operational viability of a markerless, Kinect-based educational game designed to support embodied interaction baseline frameworks for preschool-aged children. The platform captures full-body movement in real time across four progressively challenging spatial tasks. As an exploratory pilot study, the system was deployed in a preschool setting with nine children aged 5–6 ( ). Technical performance metrics, kinematic trajectory execution data, and qualitative feedback from evaluating educators ( ) were collected to establish operational baselines. Quantitative metrics demonstrated high technical accuracy and system stability across all stages: 93.1% in spatial navigation, 81.5% in postural coordination, 92.6% in obstacle avoidance, and 96.3% in guided movements, alongside low system-calculated error rates (6.87% to 3.70%). Evaluating teachers confirmed strong student engagement, low operational overhead, and high institutional acceptance for daily classroom routines. Due to the pilot nature and small sample size, direct impacts on psychomotor or computational thinking development cannot be empirically claimed. However, these findings successfully demonstrate the technical reliability and operational feasibility of deploying markerless gesture-based interfaces within authentic early education environments, establishing a validated framework for future large-scale longitudinal learning assessments.
Full article
(This article belongs to the Special Issue Technology-Enhanced Game-Based Approaches in Education: Learning, Emotions, and Motivation)
►▼
Show Figures

Figure 1
Open AccessArticle
HestiaInteract: A Web System for Modeling and Creating Simulations of Smart Homes
by
Artur Rodrigues Mota, Mayki dos Santos Oliveira, Eduardo Ferreira da Silva and Frederico Araújo Durão
Multimodal Technol. Interact. 2026, 10(8), 79; https://doi.org/10.3390/mti10080079 - 25 Jul 2026
Abstract
►▼
Show Figures
The rapid proliferation of Internet of Things (IoT) devices has accelerated the development of Smart Home environments. However, research in this domain remains constrained by the scarcity of high-quality datasets and the challenges associated with collecting real-world data. The HESTIA simulator addresses this
[...] Read more.
The rapid proliferation of Internet of Things (IoT) devices has accelerated the development of Smart Home environments. However, research in this domain remains constrained by the scarcity of high-quality datasets and the challenges associated with collecting real-world data. The HESTIA simulator addresses this issue by generating realistic synthetic data, but its adoption is hindered by a complex configuration process that requires manual editing and technical expertise. To address these limitations, this work introduces HestiaInteract, a web-based interaction layer that abstracts the low-level configuration required by HESTIA. Rather than merely providing a graphical interface, the proposed platform introduces a structured workflow for semantic scenario modeling, reusable simulation components, automatic input validation, and browser-based execution. By replacing manual JSON editing and script-based configuration with guided visual interactions, the platform reduces technical barriers while preserving the flexibility of the original simulator. A usability study involving 50 participants was conducted to evaluate the proposed solution. The results indicated high levels of satisfaction, ease of use, interaction quality, and clarity in task execution. Participants reported that the interface significantly simplified the simulation process and improved the overall user experience. These findings demonstrate that HestiaInteract enhances the usability and accessibility of the HESTIA simulator, facilitating its adoption and supporting research in intelligent environments.
Full article

Figure 1
Open AccessArticle
A Multi-Perspective Validation of a Gamified Virtual Reality Platform Translating CBT-Informed Techniques
by
Mashael Bin Sabbar, Gary Ushaw and Rich Davison
Multimodal Technol. Interact. 2026, 10(8), 78; https://doi.org/10.3390/mti10080078 - 24 Jul 2026
Abstract
►▼
Show Figures
Virtual reality (VR) has increasingly been explored to support mental health interventions, including those informed by Cognitive Behavioural Therapy (CBT). However, many VR-based CBT systems report limited usability evaluation and minimal practitioner involvement, raising questions about their applicability beyond clinical settings. This study
[...] Read more.
Virtual reality (VR) has increasingly been explored to support mental health interventions, including those informed by Cognitive Behavioural Therapy (CBT). However, many VR-based CBT systems report limited usability evaluation and minimal practitioner involvement, raising questions about their applicability beyond clinical settings. This study aimed to conduct a usability-focused validation of a gamified VR platform translating selected CBT-informed techniques for non-clinical use. Selected CBT techniques were implemented as short VR mini-games guided by a design rationale emphasising simplicity, symbolic interaction, and ease of use. A user study involving 58 university students evaluated usability using the System Usability Scale (SUS), ISO 9241-11-informed usability categories, and brief anxiety-related self-report measures. In addition, three psychological practitioners reviewed the platform and assessed therapeutic coherence and design alignment. Findings demonstrated good overall usability, positive immediate experiential responses, and practitioner support for the platform’s therapeutic coherence and suitability as a supportive digital wellbeing tool. The study demonstrates a structured usability-based validation approach for VR systems translating CBT-informed techniques and offers practical guidance for the design and evaluation of immersive mental health technologies.
Full article

Figure 1
Open AccessArticle
Human Factors in Teacher Readiness for Educational Virtual Reality: A CFA and SEM-Based TAM–TPB Study
by
Petru-Iulian Grigore, Corneliu Octavian Turcu and Ionela-Cristina Breahnă-Pravăţ
Multimodal Technol. Interact. 2026, 10(8), 77; https://doi.org/10.3390/mti10080077 - 23 Jul 2026
Abstract
►▼
Show Figures
Teacher adoption of virtual reality (VR) in education appears constrained despite the technology’s potential as an immersive multimodal learning environment. This cross-sectional online survey, based on convenience and snowball sampling, examined adoption perceptions among 408 Romanian teachers from primary, secondary, and tertiary levels
[...] Read more.
Teacher adoption of virtual reality (VR) in education appears constrained despite the technology’s potential as an immersive multimodal learning environment. This cross-sectional online survey, based on convenience and snowball sampling, examined adoption perceptions among 408 Romanian teachers from primary, secondary, and tertiary levels using an integrated Technology Acceptance Model and Theory of Planned Behavior framework. Seven constructs were measured on five-point Likert scales and analyzed through internal consistency indices, confirmatory factor analysis, HTMT discriminant validity assessment, structural equation modeling, Spearman correlations, nonparametric group comparisons, supplementary manifest-score regression, and descriptive thematic coding of open-ended responses. The seven-factor CFA model showed acceptable fit (CFI = 0.945, TLI = 0.936, RMSEA = 0.068), and composite reliability and AVE supported convergent validity across all constructs. However, HTMT indicated limited discriminant validity between attitude toward using VR and attitude toward the behavior of adopting VR (HTMT = 0.928). In the SEM model, perceived usefulness showed the largest standardized association with attitude toward using VR, while behavioral intention was mainly associated with attitudinal evaluations and subjective norm; perceived behavioral control showed a weaker standardized path. All scales showed acceptable internal consistency (Cronbach’s –0.95), and construct means exceeded the scale midpoint (range: 3.45–4.03), indicating generally positive but differentiated perceptions. Supplementary manifest-score regression was consistent with the SEM results: the unified attitude factor showed the strongest statistical association with behavioral intention, followed by subjective norm and perceived behavioral control. Descriptive thematic coding of open-ended responses identified training, infrastructure, equipment access, curriculum-aligned content, cost, technical support, and time as recurrent perceived conditions associated with self-reported VR adoption intentions. The findings suggest that educational VR adoption should be interpreted through self-reported human factors and perceived implementation conditions, including perceived control, access to immersive equipment, practical training, and institutional support.
Full article

Graphical abstract
Open AccessArticle
Explainable Text-Based Computational Framework for Psycho-Emotional Risk Classification: Calibration, Interpretability, and Encoder Comparison on Public Datasets
by
Orazmukhamed Bekmurat, Vassiliy Serbin, Aliya Aizhanova, Nurdaulet Niyaz, Darkhan Yerezhep, Kanibek Sansyzbay, Ayaulym Oralbekova and Malika Sagitzhanova
Multimodal Technol. Interact. 2026, 10(7), 76; https://doi.org/10.3390/mti10070076 - 21 Jul 2026
Abstract
►▼
Show Figures
The early detection of psycho-emotional risks remains challenging despite its growing importance. Most existing models rely on a single data type, mainly questionnaires or text, and operate as “black boxes”, limiting practical use. Psycho-emotional states are multidimensional, reflected in textual, structured, and temporal
[...] Read more.
The early detection of psycho-emotional risks remains challenging despite its growing importance. Most existing models rely on a single data type, mainly questionnaires or text, and operate as “black boxes”, limiting practical use. Psycho-emotional states are multidimensional, reflected in textual, structured, and temporal digital signals. Ignoring this complexity may result in information loss and lower prediction accuracy. This paper presents an explainable text-based computational framework for psycho-emotional risk classification using public datasets. Record-level structured information was used only when it was available within the same original observation and was not created by matching records across datasets. The study did not integrate observations across datasets or perform cross-dataset fusion; these aspects are considered potential directions for future research. Implemented in Python using PyTorch, the model was evaluated on two open-text datasets. The BERT-based configuration achieved accuracy of 0.593 ± 0.009 and an ROC-AUC of 0.860 ± 0.006, while RoBERTa-base improved the performance to 0.666 ± 0.005 accuracy and a 0.896 ± 0.004 ROC-AUC under five-fold cross-validation. The results demonstrate that classification quality depends on encoder selection, preprocessing, dataset characteristics, duplicate handling, calibration, and the way that available record-level information is represented. Therefore, the findings should be interpreted as record-level computational classification results rather than the clinical validation of psycho-emotional risk assessment.
Full article

Figure 1
Open AccessArticle
SMD-Net: Selective and Multiway Differential Perception-Enhanced Network for Early Multimodal Rumor Detection
by
Zhengnan Qiao, Zhekang Yang and Xianguo Zhang
Multimodal Technol. Interact. 2026, 10(7), 75; https://doi.org/10.3390/mti10070075 - 10 Jul 2026
Abstract
►▼
Show Figures
The rapid dissemination of rumors on social media at their early stages poses significant threats to public safety and social stability. While response-based methods usually depend on user comments and reposts and therefore suffer from inherent latency, existing content-based methods still struggle to
[...] Read more.
The rapid dissemination of rumors on social media at their early stages poses significant threats to public safety and social stability. While response-based methods usually depend on user comments and reposts and therefore suffer from inherent latency, existing content-based methods still struggle to extract discriminative evidence from noisy short texts and subtle visual inconsistencies under zero-response conditions. To address this issue, we propose SMD-Net, a multimodal framework for early zero-response rumor detection. In the textual branch, a selective state-space encoder is used to model fragmented and noisy posts. In the visual branch, an enhanced TransXNet backbone is designed to improve the representation of fine-grained suspicious patterns and cross-layer feature interactions. An adaptive gated fusion module is further introduced to integrate textual and visual features for final prediction. Experiments on the Weibo and PHEME datasets show that SMD-Net outperforms the compared content-based baselines, achieving 92.60% accuracy on Weibo and 90.27% accuracy on PHEME under the strict zero-response setting. These results suggest that the proposed framework provides an effective solution for early multimodal rumor detection when propagation-based evidence is unavailable.
Full article

Figure 1
Open AccessArticle
A Comprehensive Framework for Designing Metaverse-Based Learning for Engineering Education
by
Andrea Bezzina and Joseph Paul Zammit
Multimodal Technol. Interact. 2026, 10(7), 74; https://doi.org/10.3390/mti10070074 - 3 Jul 2026
Abstract
►▼
Show Figures
The growing interest in metaverse technologies for education has not been matched by structured methodologies to guide their design. This study presents a systematic review following PRISMA guidelines, analysing 33 studies to examine the use of metaverse and immersive technologies in engineering education.
[...] Read more.
The growing interest in metaverse technologies for education has not been matched by structured methodologies to guide their design. This study presents a systematic review following PRISMA guidelines, analysing 33 studies to examine the use of metaverse and immersive technologies in engineering education. The review reveals that existing frameworks remain limited in scope, lacking end-to-end guidance rooted in pedagogical foundations. To address this gap, the Metaverse Immersive Training Environment (MITE) framework is proposed. Built upon Design Thinking and integrating TPACK, Constructive Alignment and the 5E model, the framework provides a structured, replicable methodology across three sequential phases: Investigation of Requirements, Creation of Metaverse Infrastructure and Development of Learning Content. Unlike existing approaches that address pedagogical design or technical development in isolation, MITE integrates multiple pedagogical models within a unified methodology, providing end-to-end guidance for designing metaverse-based educational environments.
Full article

Figure 1
Open AccessArticle
Evaluating In-Vehicle Multimodal Interaction via Multimodal Behavioral Signals: A Theory-Driven Tool Chain and Sim-to-Real Pilot Study
by
Xinyi Li, Gang Guo, Qihang Sun, Yingzhang Wu and Wenbo Li
Multimodal Technol. Interact. 2026, 10(7), 73; https://doi.org/10.3390/mti10070073 - 29 Jun 2026
Abstract
►▼
Show Figures
Multitasking is pervasive in multimodal interaction, particularly within safety-critical domains like driving. Evaluating the impact of In-Vehicle Multimodal Interaction (IVMI) on drivers is critical, yet existing methods predominantly rely on post hoc subjective surveys or coarse unimodal monitoring. Grounded in Multiple Resource Theory
[...] Read more.
Multitasking is pervasive in multimodal interaction, particularly within safety-critical domains like driving. Evaluating the impact of In-Vehicle Multimodal Interaction (IVMI) on drivers is critical, yet existing methods predominantly rely on post hoc subjective surveys or coarse unimodal monitoring. Grounded in Multiple Resource Theory and following a Research through Design methodology, we operationalized this theory into a non-intrusive tool chain that evaluates IVMI impact from multimodal behavioral signals (visual, touch, and driving) and supports real-time, objective evaluation in both simulated and real-world domains. To mitigate the Sim-to-Real gap, the method combines real-world multimodal data acquisition with a modality-decoupled cross-domain calibration. Its feasibility was evaluated through a simulator study ( ) and a small-nscale real-world on-road pilot study ( ). The results suggest that the tool chain effectively acquires high-fidelity data to support the previously developed evaluation model (Quadratic Weighted Kappa = 0.916) and achieves a preliminary calibration of cross-domain latent feature spaces. As its reference labels are behaviorally derived and share a common basis with the model inputs, this agreement indicates internal consistency rather than independent construct validation. Crucially, while multimodal interaction behaviors (visual and touch) exhibited relatively high cross-domain consistency, real-world driving behaviors showed systematic magnitude suppression. This finding is tentatively interpreted, as a hypothesis to be tested in future work, through the lens of Risk Homeostasis Theory, and highlights the necessity of monitoring multimodal interaction behaviors rather than relying solely on vehicle telemetry. Overall, this research develops and provides preliminary feasibility evidence for a theory-driven cross-domain tool chain, indicating its potential to objectively quantify multimodal interaction impacts in real-world multitasking contexts. Given the small, homogeneous on-road sample, these pilot-stage results should be read as feasibility evidence and a methodological basis for future large-scale, demographically diverse validation.
Full article

Figure 1
Open AccessArticle
Evacuation Safety on Ships Through 360 Virtual Tour Familiarization
by
Sebastià Nicolau Morey, Julio Rodríguez, José Antonio Orosa, Juan José Cartelle Barros, Laura Castro-Santos and María Isabel Lamas
Multimodal Technol. Interact. 2026, 10(7), 72; https://doi.org/10.3390/mti10070072 - 29 Jun 2026
Abstract
In the maritime industry, it is common for crew members to be unfamiliar with their workplace until they are on board. This limits evacuation time, which is critical in emergency situations such as fire or shipwreck. The aim of this work is therefore
[...] Read more.
In the maritime industry, it is common for crew members to be unfamiliar with their workplace until they are on board. This limits evacuation time, which is critical in emergency situations such as fire or shipwreck. The aim of this work is therefore to enhance the preparedness and safety of new crew members by enabling them to become familiar with the location of escape routes before boarding the ship. For this purpose, the use of 360 technology is proposed. A 360 tour is a virtual, interactive experience that allows users to explore a location as if they were physically there. On a ship, this allows the crew to know the escape routes and make quicker and safer decisions on the most appropriate route. This paper shows how a 360 tour was developed and tested on a merchant vessel. The evacuation time from the engine room has been measured for a group of people. The results showed that the 360 tour had a very positive impact on the evacuation time for these people who had not previously been physically present in the area.
Full article
(This article belongs to the Special Issue Educational Virtual/Augmented Reality)
►▼
Show Figures

Figure 1
Open AccessArticle
Contrastive vs. Example-Based Explanations: Designing for Better User Comprehension in Smartphone Privacy Interfaces
by
Ananya Bhadauria and Andreas Riener
Multimodal Technol. Interact. 2026, 10(7), 71; https://doi.org/10.3390/mti10070071 - 26 Jun 2026
Abstract
►▼
Show Figures
Smartphone applications routinely collect and share personal data, yet users often struggle to understand these practices, particularly when third-party data sharing is involved. Existing mechanisms such as privacy policies and notices provide limited support for user understanding. To address this gap, we investigated
[...] Read more.
Smartphone applications routinely collect and share personal data, yet users often struggle to understand these practices, particularly when third-party data sharing is involved. Existing mechanisms such as privacy policies and notices provide limited support for user understanding. To address this gap, we investigated how explainability concepts can enhance contextual privacy policies for mobile apps. We designed two interface prototypes integrating explanation strategies: contrastive explanations, which clarify data-sharing boundaries, and example-based explanations, which illustrate counterfactual scenarios. In an exploratory between-subjects user study (N = 30), we evaluated their impact on comprehensibility, simplicity, cognitive load, and experience. Statistical analysis revealed no significant differences between the two types. Hence, the findings should be read as descriptive trends rather than confirmatory effects. These trends suggested example-based explanations better supported users mental models, while contrastive explanations better supported decision-making. Our findings contribute design recommendations on applying explanation strategies in privacy interfaces, offering guidance for developers and researchers seeking to improve user understanding and trust in digital systems.
Full article

Figure 1
Open AccessArticle
A Preliminary Investigation of Thai Clinical Attitudes Towards VR Adoption in Upper-Extremity Rehabilitation: Patient Usability and Clinician Perceived Usefulness
by
Sanya Utthayotha and Noppon Choosri
Multimodal Technol. Interact. 2026, 10(7), 70; https://doi.org/10.3390/mti10070070 - 26 Jun 2026
Abstract
►▼
Show Figures
Virtual reality (VR) has shown promising potential for upper-extremity rehabilitation; however, its successful integration into clinical practice depends not only on therapeutic effectiveness but also on the acceptance of the technology by patients and healthcare professionals alike. Despite growing international research in this
[...] Read more.
Virtual reality (VR) has shown promising potential for upper-extremity rehabilitation; however, its successful integration into clinical practice depends not only on therapeutic effectiveness but also on the acceptance of the technology by patients and healthcare professionals alike. Despite growing international research in this area, there is limited evidence on clinical attitudes toward VR rehabilitation in Thailand and other middle-income settings. This study investigates Thai patients’ and clinicians’ perceptions of VR for upper-extremity rehabilitation through two complementary studies focusing on perceived usability and usefulness. The first study evaluated the perceived usability of a VR rehabilitation game using the System Usability Scale (SUS) among 40 first-time VR users divided into younger and senior groups. The younger group reported a higher average SUS score (64.6) than the senior group (55.4). While both groups generally perceived VR rehabilitation positively, senior participants expressed greater concern regarding system complexity, consistency, and the need for technical assistance. Nevertheless, the findings indicate that VR remained an acceptable rehabilitation approach even among elderly first-time users in a population with relatively lower technological readiness. The second study explored clinicians’ perceptions of utilizing VR-generated movement data to support rehabilitation decision-making. Five rehabilitation professionals evaluated the potential usefulness of VR data visualizations for diagnosis and treatment monitoring. Clinicians generally perceived VR data as valuable, particularly for tracking rehabilitation progress rather than diagnostic decision-making. Feedback from interviews also highlighted practical considerations for future implementation, including the importance of normative data, simplified visualization formats, and the feasibility of clinical workflows. By combining patient usability perspectives with clinicians’ evaluations of clinical usefulness, this research provides a broader understanding of the factors influencing VR adoption for upper-extremity rehabilitation in Thailand. The findings contribute contextual evidence from an underrepresented healthcare environment and offer insights relevant to the future deployment of VR-assisted rehabilitation systems in similar socio-economic settings.
Full article

Figure 1
Open AccessSystematic Review
Effectiveness of PhET Simulations on Learning Outcomes in Science and Chemistry Education: A Systematic Review
by
Sinta Ayu Ningrum, Ijang Rohman, Gun Gun Gumilar, Ahmad Mudzakir, Muhammad Nurul Hana and Miarti Khikmatun Nais
Multimodal Technol. Interact. 2026, 10(7), 69; https://doi.org/10.3390/mti10070069 - 24 Jun 2026
Abstract
The development of digital learning technologies has introduced innovative tools to enhance science and chemistry education, including PhET simulations. This study aims to evaluate the effectiveness of PhET simulations on students’ learning outcomes through a systematic literature review following the PRISMA 2020 guidelines.
[...] Read more.
The development of digital learning technologies has introduced innovative tools to enhance science and chemistry education, including PhET simulations. This study aims to evaluate the effectiveness of PhET simulations on students’ learning outcomes through a systematic literature review following the PRISMA 2020 guidelines. A systematic search of Scopus and Crossref databases was conducted (last search: January 2026) using predefined keywords. Eligible studies were empirical research published between 2020 and 2026 that investigated PhET simulations in science-related education and reported learning outcomes, while non-empirical studies and non-Scopus-indexed articles were excluded. Risk of bias was assessed using an adapted Joanna Briggs Institute critical appraisal tool. Due to heterogeneity in study designs and outcome measures, the results were synthesized using a narrative approach. A total of 14 studies across elementary to higher education levels were included. The findings indicate that PhET simulations consistently improve learning outcomes, particularly academic achievement and conceptual understanding, with effects generally favoring simulation-based instruction over traditional methods. However, higher-order skills and affective outcomes such as motivation and attitude remain less frequently investigated. The evidence is limited by variability in study designs, incomplete reporting of non-cognitive outcomes, and the absence of quantitative synthesis. Overall, PhET simulations demonstrate strong potential as an effective interactive learning medium, although their impact depends on instructional design, teacher facilitation, and technological accessibility.
Full article
(This article belongs to the Special Issue Online Learning to Multimodal Era: Interfaces, Analytics and User Experiences)
►▼
Show Figures

Figure 1
Open AccessArticle
Automating Spatial Visualisation of Handwritten Vector Equations Using Large Vision Models in Pre-Tertiary Mathematics
by
Kenneth Y. T. Lim, Nguyen Thanh Minh Le and Sopheap Chanoudam
Multimodal Technol. Interact. 2026, 10(6), 68; https://doi.org/10.3390/mti10060068 - 14 Jun 2026
Abstract
►▼
Show Figures
Understanding advanced pre-tertiary mathematics, particularly three-dimensional vectors, demands robust spatial reasoning skills that many students find challenging to develop through traditional pedagogical methods. This study proposes and evaluates an innovative educational tool that leverages large vision models to automate the conversion of handwritten
[...] Read more.
Understanding advanced pre-tertiary mathematics, particularly three-dimensional vectors, demands robust spatial reasoning skills that many students find challenging to develop through traditional pedagogical methods. This study proposes and evaluates an innovative educational tool that leverages large vision models to automate the conversion of handwritten vector equations into accurate 3D graphical representations. By interpreting students’ handwritten input using advanced computer vision, the system provides immediate, interactive visual feedback to bridge the cognitive gap between abstract symbolic notation and tangible geometric concepts. We evaluated the system using a dataset of 1000 handwritten vector equations typical of the Singapore-Cambridge GCE ‘A’ Level H2 Mathematics syllabus. Our findings demonstrate that while GPT-4o serves as a capable baseline, achieving 84.6% accuracy with multi-shot prompting, newer variants such as GPT-4.1-mini offer superior performance, reaching 91.4% accuracy with significantly higher computational efficiency. The results confirm that AI-powered visualisation tools can effectively interpret complex spatial mathematical layouts when guided by optimal prompt engineering. Implementing such technology in educational settings presents a viable, scalable, and cost-effective method to democratise learning support, fostering independent study and enhancing students’ conceptual comprehension of spatial mathematics.
Full article

Figure 1
Open AccessArticle
Digital Tools for Innovation in Craft Design: Lessons from a Multi-Domain European Design Pilot
by
Arnaud Dubois, Zoé L’Évêque, Inés Moreno, Loïc Petitgirard, Danae Kaplanidi, Juan Carlos Bañón, Juan José Ortega, Nikolaos Partarakis and Xenophon Zabulis
Multimodal Technol. Interact. 2026, 10(6), 67; https://doi.org/10.3390/mti10060067 - 4 Jun 2026
Abstract
►▼
Show Figures
Traditional European craft practices face dual pressures: the erosion of tacit knowledge held by aging practitioners, and the risk of cultural homogenization through uninformed digital adoption. This paper presents a comparative analysis of a structured design pilot conducted across five Representative Craft Instances
[...] Read more.
Traditional European craft practices face dual pressures: the erosion of tacit knowledge held by aging practitioners, and the risk of cultural homogenization through uninformed digital adoption. This paper presents a comparative analysis of a structured design pilot conducted across five Representative Craft Instances (RCIs): glassblowing, tapestry, marble/silversmithing, porcelain, and woodcarving within the Horizon Europe CRAEFT project. Drawing on co-creative workshops, motion capture pipelines, physically based rendering (PBR), interactive simulation, and additive manufacturing, we analyze how context-specific digital tools performed as mediators rather than modernizers across heterogeneous craft domains. Cross-domain analysis reveals that digital tools achieve cultural legitimacy only when introduced through co-creative, practitioner-led cycles; that gesture and tacit knowledge are transferable via structured computational pipelines; and that methodological portability, not workflow replication, is the appropriate model for cross-context scaling. Implications are discussed for sustainable heritage policy, design education, and the development of craft-sensitive digital infrastructure in Europe. A cross-RCI comparative assessment matrix evaluates all five domains across seven analytical dimensions: practitioner adoption, perceived usefulness, cultural legitimacy, technical maturity, sustainability impact, transferability potential, and educational effectiveness. Finally, practitioner reflective accounts from participating designers and craftspeople are presented to ground the analytical findings empirically.
Full article

Figure 1
Open AccessSystematic Review
Parental Communication Strategies During Screen Time in Early Childhood: A Scoping Review of Joint Media Engagement
by
Litna A Varghese, Gagan Bajaj, Megha Mohan, Jayashree S. Bhat, Jayashree Kanthila and Aiswarya Liz Varghese
Multimodal Technol. Interact. 2026, 10(6), 66; https://doi.org/10.3390/mti10060066 - 4 Jun 2026
Abstract
►▼
Show Figures
Background: This scoping review aimed to systematically identify communication strategies used during Joint Media Engagement (JME) and examine their associations with developmental outcomes and contextual factors. Methods: A systematic search of seven databases (up to April 2025) was conducted using Rayyan,
[...] Read more.
Background: This scoping review aimed to systematically identify communication strategies used during Joint Media Engagement (JME) and examine their associations with developmental outcomes and contextual factors. Methods: A systematic search of seven databases (up to April 2025) was conducted using Rayyan, following PRISMA-ScR guidelines; 26 studies met inclusion criteria and were synthesized to categorize parent communication strategies and their theoretical underpinnings. Results: Fifteen distinct communication strategies were identified and organized into four theoretical frameworks; Social Learning, Sociopragmatic, Behaviourist, and Theory of Mind along with a fifth category for technical scaffolding. Strategies aligned with Social Learning were most frequently reported and consistently associated with improvements in children’s language, cognitive, and socio-emotional outcomes. Findings also showed that JME strategies vary based on contextual factors, including parent type, geography, device type, media content, and child characteristics. Although most studies did not explicitly focus on JME, those employing mixed methods provided deeper insights. Conclusions: JME is shaped by both interaction quality and context, with Social Learning-based strategies playing a central role in supporting child development. The findings highlight the need for more rigorous, JME-focused research across diverse digital formats to strengthen the evidence-based parent coaching approaches to optimize JME practices in early childhood.
Full article

Figure 1
Open AccessArticle
Personalizing Live Avatar Interaction for Children with ASD Through Restricted Interests: A Feasibility Study
by
Luis Fernando Guerrero-Vásquez, Martín López-Nores, Henry J. Jara-Quito, Dalila M. González-González and Jack Fernando Bravo-Torres
Multimodal Technol. Interact. 2026, 10(6), 65; https://doi.org/10.3390/mti10060065 - 2 Jun 2026
Abstract
►▼
Show Figures
Virtual avatars have shown potential as supports in Autism Spectrum Disorder (ASD) interventions, but many existing systems provide largely standardized interactions that do not account for individual variability. This study presents an exploratory evaluation of a virtual puppet system that enables real-time interaction
[...] Read more.
Virtual avatars have shown potential as supports in Autism Spectrum Disorder (ASD) interventions, but many existing systems provide largely standardized interactions that do not account for individual variability. This study presents an exploratory evaluation of a virtual puppet system that enables real-time interaction by synchronously transmitting a human model’s movements, facial gestures, and voice to a digital avatar. The system was personalized using each participant’s restricted interests (RIs), identified through a clinical triangulation process involving therapist input, caregiver reports, and observation. After an initial technical validation with 16 neurotypical children, the system was evaluated in a proof-of-concept sample of 11 children with ASD (7 in an experimental group exposed to RI-based personalization and 4 in a control group interacting with a standard interface). Data sources included eye tracking and therapist-completed observational questionnaires. Across sessions, descriptive patterns in gaze fixation and therapist reports suggested that RI-based personalization may help sustain attention to the screen and support engagement with the therapeutic environment relative to non-personalized interaction. Heatmap patterns further indicated that children under the personalized condition visually explored RI-related elements within the scene. This study provides evidence of technical and procedural feasibility and generates hypotheses for future research.
Full article

Figure 1
Highly Accessed Articles
Latest Books
E-Mail Alert
News
9 October 2025
Meet Us at the 3rd International Conference on AI Sensors and Transducers, 2–7 August 2026, Jeju, South Korea
Meet Us at the 3rd International Conference on AI Sensors and Transducers, 2–7 August 2026, Jeju, South Korea
5 August 2026
MDPI INSIGHTS: The CEO’s Letter #37 – Canada Summit, Sciforum Relaunch, 30 Years of Impactful Research & ISPRS 2026
MDPI INSIGHTS: The CEO’s Letter #37 – Canada Summit, Sciforum Relaunch, 30 Years of Impactful Research & ISPRS 2026
Topics
Topic in
AI, Inventions, MTI, Robotics, Sci, Sensors, Standards, Technologies
Toward Trustworthy Human-AI Collaboration: From Interactive Intelligence to Collaborative Autonomy
Topic Editors: George Margetis, Helmut Degen, Stavroula NtoaDeadline: 30 November 2026
Topic in
Electronics, MTI, BDCC, AI, Virtual Worlds, Applied Sciences
AI-Based Interactive and Immersive Systems
Topic Editors: Sotiris Diplaris, Nefeli Georgakopoulou, Stefanos Vrochidis, Giuseppe Amato, Maurice Benayoun, Beatrice De GelderDeadline: 31 December 2026
Topic in
AI, Arts, Computers, MTI
Artificial Intelligence and the Future of Art
Topic Editors: Ahmed Elgammal, Marian MazzoneDeadline: 31 October 2027
Special Issues
Special Issue in
MTI
Technology-Enhanced Game-Based Approaches in Education: Learning, Emotions, and Motivation
Guest Editors: Ana Manzano-León, José M. Rodríguez FerrerDeadline: 30 August 2026
Special Issue in
MTI
Educational Virtual/Augmented Reality
Guest Editor: Arun KulshreshthDeadline: 30 November 2026
Special Issue in
MTI
Intelligent and Multimodal Human–Computer Interaction: From Brain–Computer Interfaces to Assistive Systems
Guest Editors: Antonopoulos Nikos, Andreas Kanavos, Constantinos MourlasDeadline: 16 December 2026
Special Issue in
MTI
uHealth Interventions and Digital Therapeutics for Better Diseases Prevention and Patient Care
Guest Editor: Silvia GabrielliDeadline: 31 December 2026

