Null StudioNullStudio

Blog · August 23, 2026 · 10 min read

From Pilot to Program: Why Enterprise XR Rollouts Stall and How to Scale One

By the Null Studio team

TL;DR: The pilot is the easy part. Most enterprise XR programs stall somewhere between a successful pilot and a second site, and almost never because the software was bad. They stall on the operations layer nobody scoped: who charges, cleans and tracks the headsets, how apps and OS updates get pushed to devices nobody is holding, how the content gets changed when the procedure changes on a Tuesday, how completion data reaches the system your compliance team already trusts, and who owns the program after the champion moves teams. Here is what that layer is actually made of, what breaks first at site two, and how to scope a pilot that can survive becoming a program.

A first XR pilot is run under conditions no rollout will ever repeat. One location, the newest hardware, an engineer physically in the room, volunteers who wanted to try it, and the free tailwind of novelty. Under those conditions almost any competent build looks like a success, which is exactly why a good pilot result predicts so little about whether the program scales.

We have built XR for public safety, healthcare and shared multi-user spaces, including ARCortex's ERIS XR platform for firefighters and Nystag's clinical VR eye-tracking diagnostics on Vive Focus 3. The pattern across the ones that grew is consistent, and it has less to do with the application than buyers expect. Here is the part of the plan that decides it.

Why pilots succeed and programs stall

A pilot tests the experience. A program tests the operation around the experience. Those are different systems, and the pilot deliberately holds four variables constant that a rollout cannot.

Someone technical was in the room. During a pilot, a person who understands the device handles setup, pairing, guardian boundaries, sign-in and the reboot when something wedges. At scale that person does not exist, and every ten seconds of friction becomes a headset that stays in the cupboard.

The users volunteered. Pilot participants opted in and were curious. A rollout includes people who are indifferent, people who get motion discomfort, people who do not want to wear a device that was on someone else's face an hour ago, and people who quietly decide this is a fad.

The content was frozen. Nothing changed during the six weeks of the pilot. In production, procedures change, equipment gets replaced, policies get revised, and the day your training contradicts the real world is the day people stop trusting it.

Nothing had to be reported. A pilot produces impressions. A program has to produce records that satisfy a compliance officer, a regulator, an insurer or a training manager, in the system they already use.

If your plan does not have an answer for those four, you do not yet have a rollout plan. You have a longer pilot.

The operations layer, item by item

None of this is exotic. All of it gets discovered late, which is what makes it expensive.

Device logistics

Headsets do not behave like laptops. They have consumable parts, they arrive uncharged, they hold a charge poorly when idle, and they are shared objects worn on the face.

Plan for charging that is somebody's actual job, storage that keeps devices at hand rather than locked in an office, wipeable facial interfaces or disposable covers if devices are shared across people or shifts, lens protection because a scratched lens is a dead headset, and a spares allowance that is not zero. Devices get dropped, straps break, batteries degrade on a schedule, and a program that pauses whenever one unit fails has taught the site that the program is unreliable. Decide who owns the cart, in writing, at a named site. In practice, the single best predictor of whether site two works is whether one person there is accountable for the hardware.

Device management

This is the item most often missing from an XR proposal and the one that turns a pile of headsets into a fleet.

Standalone headsets need enrollment into a management platform built for them rather than the phone MDM your IT team already runs. Dedicated platforms exist for exactly this, ArborXR and ManageXR among them, alongside the device vendors' own business tooling, and the capability you are buying is specific: enroll devices without touching each one, push your app and its updates remotely, lock devices into kiosk or single app mode so users land in your content rather than a store, provision wifi in advance, and see remotely which units are charged, online and running the current build.

Two details cause more delays than anything else in this section. Wifi is the first: enterprise networks with certificate-based authentication or captive portals frequently refuse headsets, and that conversation with IT and security should start before hardware is ordered, not on rollout day. Vendor OS updates are the second: headset operating systems update on the manufacturer's schedule, occasionally changing behavior your app depends on, so someone has to own testing new firmware against your build and controlling when devices take it.

Who is actually in the headset

Shared devices break the assumption every software system makes about identity. If five people use one unit during a shift and the app cannot tell them apart, you have no completion records, no per-person assessment, and nothing a compliance system can consume.

Solve it deliberately and keep it fast, because the sign-in is the first thing a reluctant user meets. A badge scan, a short PIN, a QR code on a lanyard or a pairing with a phone app all work. What does not work is asking someone in a headset to type an email address and a password on a floating keyboard with two controllers.

Content updates after go-live

A procedure changes. A machine gets replaced with a different model. A regulation is revised. If updating your XR content means a change request, a quote and a six week wait, the content drifts out of date, and out-of-date training is worse than no training because it teaches the wrong thing with authority.

Decide at scoping time which categories of change your own team can make without engineering. Text, thresholds, scenario parameters, checklists, media and simple branching can usually be made editable, and that decision is one of the highest-leverage things in the contract. New geometry, new interactions and new physics are engineering work by nature, and the honest asset economics behind that are in our piece on the 3D content pipeline behind an XR app. What matters is that the split is written down, along with the release cadence and who approves a change before it reaches devices.

Getting the data out

Training that does not count competes with work, and work wins. If a session in the headset does not appear in the system your organization already trusts, supervisors will treat the headset as an extra rather than a substitute, and attendance quietly declines.

Define the destination early: your LMS, an xAPI learning record store, a compliance database, an EHR in clinical settings, or a spreadsheet a training manager already lives in. Then define the payload. Completion is the weakest possible signal, so agree on what the record contains, whether that is time on task, errors, sequence adherence, attempts, hesitation or whatever the measurable version of competence looks like for this task. That is the same discipline as designing assessment into the scenario in the first place, covered in VR training for high-risk teams and, where clinical validity is the bar, in clinical VR diagnostics.

The humans who run it

Every site that sustains an XR program has a person who does the unglamorous part: schedules sessions, gets first-timers comfortable, spots the person going pale and takes the headset off them, keeps the cart charged and reports what is broken. It is rarely a full role. It is always a named one.

Budget for onboarding that first-time users get, a comfort and accessibility path including seated modes and an opt-out that does not stigmatize anyone, and a route for reporting problems that reaches someone who can fix them. A program with no visible owner at the site becomes a cupboard full of amortizing hardware.

What actually breaks at site two

The second site is the real test, and it fails in patterns that are predictable enough to plan around.

The network is different. Different SSID, different security posture, different IT owner and a different opinion about unknown devices on the corporate network. Certificate-based wifi and captive portals block headsets routinely.

The physical space is different. Play areas, ceiling heights, reflective surfaces, glass walls and bright sunlight all affect tracking. Content designed for a five metre bay behaves differently in a corridor.

The equipment is different. Site one had one model of pump, panel, vehicle or machine. Site two has a variant. If your scenario hard-codes the first, you have just discovered a content project you did not budget.

The people are different. Different shifts, different first languages, different comfort with technology, sometimes a works council or union agreement about how new tools get introduced. Language support in particular is far cheaper to design in than to retrofit.

The champion is different, or absent. Site one had an enthusiast who wanted it to work. Site two got told. That single difference explains more failed rollouts than any technical factor.

Scoping a pilot that can actually become a program

You do not need to build the whole program up front. You need a pilot that does not build in the assumptions that will kill it.

Name site two before the pilot starts. Not to deploy it, just to specify against it. Every decision gets tested against a location you have not optimized for.

Treat device management as a day-one line item. Even for ten devices. It is far cheaper than retrofitting enrollment across a fleet, and it changes what a rollout costs per site more than anything except content.

Get IT and security into the first conversation. Network access, account model, data residency and what leaves the device are approval gates. Discovering them late costs weeks and no code fixes them.

Instrument from the first build. Usage, completion, errors, crashes and session length. Without them, the renewal conversation is a debate about anecdotes, and anecdotes lose to budget pressure.

Write down what your team can change without us. Content, thresholds, users, scenarios. This is a contract clause, not a technical detail.

Pilot with at least a few non-volunteers. The people who did not ask for this are the ones your rollout is made of.

What it costs and where the money moves

We break down build pricing in what an XR app costs, and the device decision that drives much of it in which XR device to build for. The scaling-specific point is different and worth stating plainly: the shape of the spend inverts after go-live.

Year one is dominated by the build. Year two, for a program that is genuinely running, is dominated by hardware refresh and spares, device management licensing, content updates as the real world changes, facilitator time at each site, and support. Engineering becomes the smaller line. Buyers who model only the build, and there are many, hit year two treating normal operating cost as an overrun and cut the program at exactly the point it was starting to pay back. Ask any vendor for a three year view with per-site incremental cost separated from one-time build cost. A vendor who has actually scaled a deployment will have that number close to hand.

Where AI helps, and where it does not

The genuinely useful application right now is content velocity. Drafting scenario variants, generating and localizing dialogue, producing the first pass of branching logic and keeping documentation in step with builds are exactly the mechanical work AI handles well under senior review, which is the same leverage described in our ship-in-days playbook. Conversational AI characters inside a scenario are also now practical, so a trainee can question a simulated patient, caller or bystander instead of picking from three canned options, which is where our voice and custom AI agent work meets the headset.

What AI does not touch is the operations layer above. No model charges the carts, negotiates with the network team, or makes a supervisor release people from the floor for twenty minutes. Treat claims that AI removes the rollout problem with suspicion.

Questions to ask before you commission

Where we fit

We build XR that has to survive contact with real operations rather than a demo table. ARCortex's ERIS XR helps firefighters pre-plan and train across real buildings, Nystag runs clinical-grade VR eye-tracking diagnostics on Vive Focus 3 in clinical settings, and MR Camera puts multiple people into a shared mixed-reality space at once, the multi-user problem covered in mixed-reality collaboration. If your XR work is hands-on equipment rather than a headset scenario, enterprise AR training is the closer fit. In all of them, the questions that decided the outcome were the ones above rather than the rendering.


Sitting on a pilot that needs to become a program? Book a demo and we will walk through the device management, content update and reporting plan for your rollout, including telling you where a smaller build gets you there. See our work: ERIS XR, Planes XR, Nystag and MR Camera, shipped in days, not months.

FAQ

Why do enterprise VR pilots fail to scale?

Rarely because the software was bad, and almost always because the pilot quietly held constant four things a rollout cannot. Someone technical was in the room to handle setup, sign-in and the occasional reboot, and at scale that person does not exist. The participants volunteered, while a rollout includes people who are indifferent, people who get motion discomfort and people who did not ask for this. The content was frozen for the duration, whereas in production procedures change, equipment gets replaced and policies get revised, and training that contradicts the real world stops being trusted. And nothing had to be reported, so no one tested whether a session in the headset produces a record the compliance system accepts. A pilot tests the experience. A program tests the operation around the experience. If the plan has no answer for unattended devices, non-volunteer users, content drift and reporting, it is not a rollout plan, it is a longer pilot.

How do you manage a fleet of VR headsets across multiple sites?

With a device management platform built for headsets rather than the phone MDM your IT team already runs. Dedicated platforms exist for this, ArborXR and ManageXR among them, alongside the device vendors' own business tooling, and the specific capability you are buying is enrolling devices without touching each one, pushing your app and its updates remotely, locking devices into kiosk or single app mode so users land in your content instead of a store, provisioning wifi in advance, and seeing remotely which units are charged, online and running the current build. Budget it as a day-one line item even for ten devices, since retrofitting enrollment across a live fleet is far more expensive. Two things cause most of the delays. Enterprise wifi with certificate-based authentication or captive portals frequently refuses headsets, so that conversation with IT and security belongs before hardware is ordered. And headset operating systems update on the manufacturer's schedule, so someone has to own testing new firmware against your build and controlling when devices take it.

How do VR training results get into an LMS or compliance system?

Deliberately, and it should be specified before the build rather than bolted on, because training that does not count competes with real work and real work wins. Decide the destination first, whether that is your LMS, an xAPI learning record store, a compliance database, an EHR in clinical settings or a reporting tool a training manager already lives in. Then decide the payload, since completion alone is the weakest possible signal. Agree on what a record contains for this task, which might be time on task, errors, sequence adherence, attempts or hesitation, because that is also what forces the assessment logic to exist in the scenario in the first place. One prerequisite catches teams out: on shared headsets the app has to know who is wearing it, so plan a fast sign-in such as a badge scan, a short PIN or a QR code on a lanyard. Asking someone in a headset to type an email address and password with two controllers is how attendance data goes missing.

Want it built, not just explained?

We design, build and run these systems end-to-end — shipped in days, not months.

Book a demo →

Keep reading