TL;DR: The pilot is the easy part. Most enterprise XR programs stall somewhere between a successful pilot and a second site, and almost never because the software was bad. They stall on the operations layer nobody scoped: who charges, cleans and tracks the headsets, how apps and OS updates get pushed to devices nobody is holding, how the content gets changed when the procedure changes on a Tuesday, how completion data reaches the system your compliance team already trusts, and who owns the program after the champion moves teams. Here is what that layer is actually made of, what breaks first at site two, and how to scope a pilot that can survive becoming a program.
A first XR pilot is run under conditions no rollout will ever repeat. One location, the newest hardware, an engineer physically in the room, volunteers who wanted to try it, and the free tailwind of novelty. Under those conditions almost any competent build looks like a success, which is exactly why a good pilot result predicts so little about whether the program scales.
We have built XR for public safety, healthcare and shared multi-user spaces, including ARCortex's ERIS XR platform for firefighters and Nystag's clinical VR eye-tracking diagnostics on Vive Focus 3. The pattern across the ones that grew is consistent, and it has less to do with the application than buyers expect. Here is the part of the plan that decides it.
Why pilots succeed and programs stall
A pilot tests the experience. A program tests the operation around the experience. Those are different systems, and the pilot deliberately holds four variables constant that a rollout cannot.
Someone technical was in the room. During a pilot, a person who understands the device handles setup, pairing, guardian boundaries, sign-in and the reboot when something wedges. At scale that person does not exist, and every ten seconds of friction becomes a headset that stays in the cupboard.
The users volunteered. Pilot participants opted in and were curious. A rollout includes people who are indifferent, people who get motion discomfort, people who do not want to wear a device that was on someone else's face an hour ago, and people who quietly decide this is a fad.
The content was frozen. Nothing changed during the six weeks of the pilot. In production, procedures change, equipment gets replaced, policies get revised, and the day your training contradicts the real world is the day people stop trusting it.
Nothing had to be reported. A pilot produces impressions. A program has to produce records that satisfy a compliance officer, a regulator, an insurer or a training manager, in the system they already use.
If your plan does not have an answer for those four, you do not yet have a rollout plan. You have a longer pilot.
The operations layer, item by item
None of this is exotic. All of it gets discovered late, which is what makes it expensive.
Device logistics
Headsets do not behave like laptops. They have consumable parts, they arrive uncharged, they hold a charge poorly when idle, and they are shared objects worn on the face.
Plan for charging that is somebody's actual job, storage that keeps devices at hand rather than locked in an office, wipeable facial interfaces or disposable covers if devices are shared across people or shifts, lens protection because a scratched lens is a dead headset, and a spares allowance that is not zero. Devices get dropped, straps break, batteries degrade on a schedule, and a program that pauses whenever one unit fails has taught the site that the program is unreliable. Decide who owns the cart, in writing, at a named site. In practice, the single best predictor of whether site two works is whether one person there is accountable for the hardware.
Device management
This is the item most often missing from an XR proposal and the one that turns a pile of headsets into a fleet.
Standalone headsets need enrollment into a management platform built for them rather than the phone MDM your IT team already runs. Dedicated platforms exist for exactly this, ArborXR and ManageXR among them, alongside the device vendors' own business tooling, and the capability you are buying is specific: enroll devices without touching each one, push your app and its updates remotely, lock devices into kiosk or single app mode so users land in your content rather than a store, provision wifi in advance, and see remotely which units are charged, online and running the current build.
Two details cause more delays than anything else in this section. Wifi is the first: enterprise networks with certificate-based authentication or captive portals frequently refuse headsets, and that conversation with IT and security should start before hardware is ordered, not on rollout day. Vendor OS updates are the second: headset operating systems update on the manufacturer's schedule, occasionally changing behavior your app depends on, so someone has to own testing new firmware against your build and controlling when devices take it.
Who is actually in the headset
Shared devices break the assumption every software system makes about identity. If five people use one unit during a shift and the app cannot tell them apart, you have no completion records, no per-person assessment, and nothing a compliance system can consume.
Solve it deliberately and keep it fast, because the sign-in is the first thing a reluctant user meets. A badge scan, a short PIN, a QR code on a lanyard or a pairing with a phone app all work. What does not work is asking someone in a headset to type an email address and a password on a floating keyboard with two controllers.
Content updates after go-live
A procedure changes. A machine gets replaced with a different model. A regulation is revised. If updating your XR content means a change request, a quote and a six week wait, the content drifts out of date, and out-of-date training is worse than no training because it teaches the wrong thing with authority.
Decide at scoping time which categories of change your own team can make without engineering. Text, thresholds, scenario parameters, checklists, media and simple branching can usually be made editable, and that decision is one of the highest-leverage things in the contract. New geometry, new interactions and new physics are engineering work by nature, and the honest asset economics behind that are in our piece on the 3D content pipeline behind an XR app. What matters is that the split is written down, along with the release cadence and who approves a change before it reaches devices.
Getting the data out
Training that does not count competes with work, and work wins. If a session in the headset does not appear in the system your organization already trusts, supervisors will treat the headset as an extra rather than a substitute, and attendance quietly declines.
Define the destination early: your LMS, an xAPI learning record store, a compliance database, an EHR in clinical settings, or a spreadsheet a training manager already lives in. Then define the payload. Completion is the weakest possible signal, so agree on what the record contains, whether that is time on task, errors, sequence adherence, attempts, hesitation or whatever the measurable version of competence looks like for this task. That is the same discipline as designing assessment into the scenario in the first place, covered in VR training for high-risk teams and, where clinical validity is the bar, in clinical VR diagnostics.
The humans who run it
Every site that sustains an XR program has a person who does the unglamorous part: schedules sessions, gets first-timers comfortable, spots the person going pale and takes the headset off them, keeps the cart charged and reports what is broken. It is rarely a full role. It is always a named one.
Budget for onboarding that first-time users get, a comfort and accessibility path including seated modes and an opt-out that does not stigmatize anyone, and a route for reporting problems that reaches someone who can fix them. A program with no visible owner at the site becomes a cupboard full of amortizing hardware.
What actually breaks at site two
The second site is the real test, and it fails in patterns that are predictable enough to plan around.
The network is different. Different SSID, different security posture, different IT owner and a different opinion about unknown devices on the corporate network. Certificate-based wifi and captive portals block headsets routinely.
The physical space is different. Play areas, ceiling heights, reflective surfaces, glass walls and bright sunlight all affect tracking. Content designed for a five metre bay behaves differently in a corridor.
The equipment is different. Site one had one model of pump, panel, vehicle or machine. Site two has a variant. If your scenario hard-codes the first, you have just discovered a content project you did not budget.
The people are different. Different shifts, different first languages, different comfort with technology, sometimes a works council or union agreement about how new tools get introduced. Language support in particular is far cheaper to design in than to retrofit.
The champion is different, or absent. Site one had an enthusiast who wanted it to work. Site two got told. That single difference explains more failed rollouts than any technical factor.
Scoping a pilot that can actually become a program
You do not need to build the whole program up front. You need a pilot that does not build in the assumptions that will kill it.
Name site two before the pilot starts. Not to deploy it, just to specify against it. Every decision gets tested against a location you have not optimized for.
Treat device management as a day-one line item. Even for ten devices. It is far cheaper than retrofitting enrollment across a fleet, and it changes what a rollout costs per site more than anything except content.
Get IT and security into the first conversation. Network access, account model, data residency and what leaves the device are approval gates. Discovering them late costs weeks and no code fixes them.
Instrument from the first build. Usage, completion, errors, crashes and session length. Without them, the renewal conversation is a debate about anecdotes, and anecdotes lose to budget pressure.
Write down what your team can change without us. Content, thresholds, users, scenarios. This is a contract clause, not a technical detail.
Pilot with at least a few non-volunteers. The people who did not ask for this are the ones your rollout is made of.
What it costs and where the money moves
We break down build pricing in what an XR app costs, and the device decision that drives much of it in which XR device to build for. The scaling-specific point is different and worth stating plainly: the shape of the spend inverts after go-live.
Year one is dominated by the build. Year two, for a program that is genuinely running, is dominated by hardware refresh and spares, device management licensing, content updates as the real world changes, facilitator time at each site, and support. Engineering becomes the smaller line. Buyers who model only the build, and there are many, hit year two treating normal operating cost as an overrun and cut the program at exactly the point it was starting to pay back. Ask any vendor for a three year view with per-site incremental cost separated from one-time build cost. A vendor who has actually scaled a deployment will have that number close to hand.
Where AI helps, and where it does not
The genuinely useful application right now is content velocity. Drafting scenario variants, generating and localizing dialogue, producing the first pass of branching logic and keeping documentation in step with builds are exactly the mechanical work AI handles well under senior review, which is the same leverage described in our ship-in-days playbook. Conversational AI characters inside a scenario are also now practical, so a trainee can question a simulated patient, caller or bystander instead of picking from three canned options, which is where our voice and custom AI agent work meets the headset.
What AI does not touch is the operations layer above. No model charges the carts, negotiates with the network team, or makes a supervisor release people from the floor for twenty minutes. Treat claims that AI removes the rollout problem with suspicion.
Questions to ask before you commission
- How do apps and updates reach devices nobody is physically holding, and what platform manages that?
- What happens when the headset vendor ships an OS update that changes behavior?
- Which parts of the content can our team change without a change request, and which cannot?
- Where does session data land, in what format, and who owns that integration?
- What is the incremental cost of the tenth site, separated from the cost of the first?
- Who at each site is accountable for the hardware, and what does that role take weekly?
- What is your spares and replacement assumption over three years?
- Can we take the source, the content and the data elsewhere if we part ways?
Where we fit
We build XR that has to survive contact with real operations rather than a demo table. ARCortex's ERIS XR helps firefighters pre-plan and train across real buildings, Nystag runs clinical-grade VR eye-tracking diagnostics on Vive Focus 3 in clinical settings, and MR Camera puts multiple people into a shared mixed-reality space at once, the multi-user problem covered in mixed-reality collaboration. If your XR work is hands-on equipment rather than a headset scenario, enterprise AR training is the closer fit. In all of them, the questions that decided the outcome were the ones above rather than the rendering.
Sitting on a pilot that needs to become a program? Book a demo and we will walk through the device management, content update and reporting plan for your rollout, including telling you where a smaller build gets you there. See our work: ERIS XR, Planes XR, Nystag and MR Camera, shipped in days, not months.