Adaptive virtual reality training is easy to sell and hard to demonstrate. Surgical simulation has produced randomised trials with measurable results, and they are worth reading closely. But the most important paper in this field was published in Educational Psychology Review in 2026, and it did something unusual: it sorted the existing VR studies by methodological rigour. Pooled across everything, VR showed a moderate benefit. Among the studies at low risk of bias, the advantage was no longer statistically significant. That result should govern how the rest of the evidence is read.
What the Laparoscopic Trials Actually Measured
The strongest positive evidence comes from surgery. Humm, Mohan, Fleming and colleagues published a meta-analysis in BJS Open in 2022 covering randomised clinical trials of virtual reality simulation training in laparoscopic cholecystectomy. Eighteen randomised trials were identified and 15 supplied usable data.
The results split in an instructive way. Compared with no additional training, VR improved technical skill on the OSATS scale and cut operating time. On GOALS, the Global Operative Assessment of Laparoscopic Skills, it did not. Same trials, different instrument, different answer. The authors flagged methodological heterogeneity, variation in the didactic teaching that accompanied the simulation, substantial statistical heterogeneity of 52 percent for time to completion, and stated that their conclusions are limited.
- OSATS technical-skill scores: mean difference of 6.22 points, 95 percent confidence interval 3.81 to 8.36, P below 0.001
- Time to task completion: reduced by 8.35 minutes, 95 percent confidence interval negative 13.10 to negative 3.60, P below 0.001
- GOALS scores: no significant difference between groups
Sort the Studies by Rigour and the Effect Shrinks
Layadi, Huguet, Pavic and Chevalere published a stratified meta-analysis of experimental rigour and bias in Educational Psychology Review in 2026. Pooled across all included studies, VR produced a standardised mean difference of 0.55, with modest publication bias detected. Then they stratified by risk of bias. Studies at high risk of bias returned 0.38, significant. Studies with some concerns returned 0.43, significant. Studies at low risk of bias returned 0.35, and that result was not statistically significant.
Their conclusion is quoted here in full because paraphrase softens it: “Analyses based on the best-available evidence did not yield consistent or robust support for a reliable learning advantage of VR. Effects reported in the broader literature appear to be driven largely by studies of lower methodological quality, whereas higher-quality studies provide limited and inconsistent evidence.” They add that current evidence is insufficient to support firm, general conclusions regarding when, how, and for whom VR enhances learning.
A separate 2025 systematic review in the Journal of Medical Internet Research, covering 23 randomised trials and 1,091 participants in orthopedic education, reported large pooled effects against conventional teaching, including a standardised mean difference of 1.44 for clinical operation scores and odds ratios above 4 for teaching interest and satisfaction. Its authors also noted notable heterogeneity across VR platforms and called for multicentre, double-blind, large-sample trials. Enthusiasm measures and rigour requests appearing in the same paper is the normal condition of this literature.
Scale is the other thing to check. A randomised trial published in the Journal of Surgical Education in 2020 found a VR group scored 17.5 against 7.5 on aggregate global assessment and completed 63 percent of steps correctly against 25 percent. It enrolled 20 first and second-year medical students with no procedural experience, and the comparator was a written guide, not cadaveric or box-trainer practice. Beating a printed handout is a low bar.
The Retail Figures Belong to Strivr, Not to Walmart
Almost every article about enterprise VR training cites Walmart. The figures are published on the customer page of Strivr, the vendor that supplied the training, and not by Walmart in any audited disclosure. Strivr claims 2.2 million retail associates trained across more than 4,700 store locations, training time cut by 96 percent from eight hours to 15 minutes, a 30 percent increase in employee satisfaction scores, and a 12.5 percent average improvement in post-training assessment scores, with VR learners outperforming others 70 percent of the time by 10 to 15 points.
Walmart people are quoted on that vendor page, including Andy Trainor, a vice president of US learning, who says the company previously sent three or four people to a store to train associates and now sends a pair of VR goggles. What the page does not carry is any date for the rollout or any methodology explaining how the 96 percent or the 12.5 percent were derived.
The Verizon figure has the same provenance and an additional problem. Strivr states that more than 22,000 retail associates across over 1,600 stores took a 20-minute robbery-safety module, and that 97 percent of learners reported feeling more prepared for dangerous situations. That is self-reported confidence, not measured performance. The distinction matters more here than anywhere, because the Layadi meta-analysis is precisely about the gap between what learners feel and what testing shows.
PwC’s study of virtual reality soft skills training in the enterprise, covering new managers trained in inclusive leadership at 12 US locations, reported that VR learners trained four times faster than classroom learners, were 3.75 times more emotionally connected to the material, and were 275 percent more confident, with cost break-even at 375 trainees and a claimed 52 percent cost advantage at 3,000. PwC sells immersive consulting, which makes it an interested party rather than an independent evaluator, and the account of the study available publicly does not state its sample size or publication date.
The Air Force Programme Is Not Running an Algorithm
One frequently cited example of AI-driven interactive VR turns out, on inspection, not to involve AI at all. VentureBeat reported on 30 August 2021 that the US Air Force had deployed sexual assault prevention and response training built by the New York studio Moth+Flame. The course took roughly two years to develop from about 2019, runs to two hours with the first 30-minute session complete at the time of reporting, launched at Little Rock Air Force Base in Arkansas, and could reach around 10,000 personnel. Trainees speak their side of difficult conversations aloud, and about 97 percent said they preferred it to the verbal instruction the service had used since 2005. Carmen Schott of Air Mobility Command and Moth+Flame chief executive Kevin Cornish were named, with General Jacqueline Van Ovost championing the work.
The mechanism described in that reporting is recorded actors, voice recognition and branching responses. It is a well-made piece of interactive video, not a generative system that composes anything. Secondary coverage routinely relabels projects like this as AI, and the relabelling is how a scripted product acquires claims that belong to a different technology.
Five Questions That Separate a Claim From a Result
Anyone evaluating a personalised VR training proposal can apply the same tests the research literature applies to itself.
- Who published the number, the buyer or the supplier? If the only source is the vendor’s customer page, treat it as marketing.
- Was performance measured, or was preference surveyed? A percentage of learners who felt more prepared is not an outcome.
- What was the control group doing? An effect against a written handout says little about an effect against existing hands-on practice.
- Which outcome instrument was used, and were others reported? The laparoscopic meta-analysis found a gain on one scale and nothing on another.
- What is the risk of bias in the underlying studies? On the 2026 stratified analysis, that single question moved the field’s headline result from significant to not.
Sources: BJS Open · Educational Psychology Review · Journal of Medical Internet Research · Strivr · Strivr · VentureBeat