Over-the-phone interpreting (OPI) and video remote interpreting (VRI) carry most of the remote interpreting volume in US healthcare today. OPI connects a clinician and patient to a remote human interpreter by audio. VRI adds live video. Approximately 29.6 million people in the US are limited English proficient (GSA, Translation and Interpretation Services SIN 541930 Ordering Guide, December 2025) and Section 1557 of the Affordable Care Act requires providers receiving federal funding to offer qualified language assistance to them. How a health system allocates encounters across OPI, VRI, in-person interpreters and now AI interpreting determines wait times, encounter length, satisfaction scores and regulatory exposure. What follows is a market-level view of where each modality stands and where the delivery model itself is under pressure.
What is the difference between OPI and VRI in healthcare?
Over-the-phone interpreting (OPI) connects a provider and patient to a remote human interpreter through an audio call. Vendors advertise connection times of 30 to 60 seconds, though the figure depends on where the clock starts: some measure from the moment the number is dialed while the number that matters operationally is the time until a qualified interpreter is actually on the line with the patient and provider.
Video remote interpreting (VRI) connects to an interpreter through live video on a tablet, cart or workstation, adding facial expression, gesture and visual confirmation to the exchange.
OPI is the higher-volume, lower-cost channel.
VRI is the higher-context channel and the only remote option appropriate for American Sign Language (ASL).
When is OPI the right modality?
OPI fits short, high-volume, low-ambiguity interactions: registration, scheduling, benefits verification, medication reminders, routine triage questions and follow-up calls. It launches from any phone without dedicated hardware, which is why it remains the default channel in call centers, front offices and units without reliable wifi. The tradeoffs go beyond context. Audio-only interpreting removes facial cues and visual confirmation and a 2022 qualitative study in PLOS ONE on telephone interpreting in primary care found the modality faces recurring technical difficulties such as poor sound quality, with interpreting accuracy decreasing by phone compared with face-to-face encounters. Delivery quality is also uneven in the field: many OPI operations route through call centers outside the US, where bandwidth constraints show up as muffled audio, interpreters who sound far from the receiver, dropped calls and long dialing menus before a human ever picks up. Each of those failures lands in the middle of a clinical encounter.
When is VRI the right modality?
VRI fits encounters where visual context changes comprehension: discharge teaching, consent discussions, pediatric visits where a parent is reading the room, behavioral health and every ASL encounter. The visual channel supports teach-back and lets the interpreter catch confusion the audio line would miss. The tradeoff is infrastructure. VRI performance depends on bandwidth, device availability, camera placement and whether the cart was charged at 7 a.m.
Who are the main OPI and VRI vendors in the US?
The US remote interpreting market is consolidated around a small group of large language service providers. LanguageLine Solutions, owned by Teleperformance since 2016, is the largest interpreting provider globally: the 2025 Nimdzi Interpreting Index reports its interpreting revenue running 38% higher than the next largest provider, with more than 85 million interactions facilitated in 2024. Propio Language Services acquired CyraCom International in July 2025, uniting two of the largest US-based providers into an entity with projected combined revenue of approximately $530 million, roughly 12,000 clients and coverage of more than 300 languages (Nimdzi, 2025).
The rest of the top tier is healthcare-weighted. AMN Language Services, built on the former Stratus Video, focuses on healthcare-specific VRI and in June 2026 acquired Jaide Health, an AI-enabled medical interpreting and translation platform it positions around administrative touchpoints such as intake and discharge, with human interpreters remaining central to clinical encounters (AMN Healthcare, June 9, 2026). GLOBO expanded its healthcare and public service footprint by acquiring All Access Interpreters and LUNA Language Services in 2023. Equiti, formerly Martti under Cloudbreak Health, was sold by UpHealth to private equity firm GTCR for $180 million (Nimdzi, 2025). The pattern across all of it is consolidation. The Jaide acquisition also sharpens what the consolidation is moving toward: even the incumbents now treat AI interpreting as the category's next layer rather than a fringe experiment, while the per-minute revenue model that underpins OPI and VRI faces a third modality built on different economics, which we return to below.
How many interpreting vendors does a typical health system contract?
Most health systems run two to three language service contracts: a primary remote vendor covering OPI and VRI, an overflow vendor for peak demand or rare languages and often a regional agency for in-person and ASL coverage. Redundancy is rational insurance against queue times. It also fragments the experience. A 2025 implementation study in JAMIA Open documented a health system whose legacy setup relied on different vendors across sites, producing inconsistent OPI and VRI availability. After consolidating interpreter access into one integrated workflow, average wait times for its top ten languages fell below 30 seconds with an average interpreter rating of 4.9 out of 5.
What are the known limitations of OPI and VRI?
The documented limitations of OPI and VRI are structural properties of the delivery model rather than failures of any single vendor. Five recur across implementation studies and patient surveys:
- Per-minute billing pressure. Both modalities are typically priced per minute, including connection and hold time. The meter creates background pressure on how long providers spend with LEP patients and turns language access into an expense that grows with patient volume.
- Connection and availability variance. Wait times vary by language and by hour. Rare languages queue. When remote channels fail and in-person backup is called, delays compound: one surgical and procedural practice assessment (2017) recorded a mean of 19 minutes from in-person interpreter request to arrival, with high variability.
- Context loss on audio (OPI). No facial cues, no gesture, no visual teach-back confirmation. In practice this shows up as repeat explanations, longer visits and a higher risk of a patient leaving without fully understanding instructions.
- Infrastructure dependence (VRI). Bandwidth, camera placement and device logistics decide whether the modality works at the bedside. Among 555 deaf adults surveyed through the Health Information National Trends Survey in ASL between 2016 and 2018, only 41% were satisfied with the quality of VRI technology in health settings.
- Dialect and rare language depth. Contracted rosters concentrate on high-volume languages. Coverage thins for regional dialects such as Haitian Creole variants, Indigenous Central American languages and less common Arabic dialects, precisely where clinical risk from miscommunication concentrates.
None of these limitations is news to the vendors themselves, several of which are investing in AI capabilities. For a health system, the practical takeaway is that these constraints are properties of the delivery model itself. Switching vendors can improve service levels at the margin, yet the structure stays: per-minute economics, remote connection dependence and roster-based availability travel with the modality. Addressing them is a question of modality mix rather than contract terms.
How do OPI and VRI shape the patient journey and patient satisfaction?
Patients do not experience interpreting modalities as interchangeable. A February 2026 NEJM Catalyst study by Montoya Rubiano, Zahakos, Ortega and colleagues of 23 Spanish-speaking surgical patients at Brigham and Women's Hospital found that patients viewed AI interpreting and remote video interpretation (RVI, the study's term) as valuable in specific clinical contexts rather than as competing tools. AI was preferred for speed, privacy and time-sensitive scenarios. RVI was preferred for emotionally complex conversations and cultural nuance. Notably, 91% of participants had SASH scores at or below 2.99, meaning these preferences come from patients with low English acculturation, the population most dependent on interpreting.
Mapped across the patient journey, the pattern is consistent. At registration and scheduling, speed dominates and audio suffices. In the exam room, comprehension and trust dominate, which favors video or a well-designed AI encounter. At discharge, visual confirmation of instructions reduces the risk of a preventable readmission. In follow-up, availability at the exact moment of the call matters more than modality. Wait time is the through line: the same JAMIA Open study that brought waits below 30 seconds recorded 4.9 out of 5 satisfaction, which suggests the dominant driver of interpreting satisfaction is not which screen the interpreter appears on but whether help arrives before the clinical moment passes.
Dr. Gezzer Ortega, the NEJM Catalyst study's senior author, made a related point on Care Culture Talks: language access and cultural competency travel together. The modality question only matters once a system has committed to meeting patients in their language at every touchpoint.
How does AI interpreting change the OPI and VRI model?
AI interpreting should not be a third channel sitting next to phone and video, though some vendors package it that way. Built well, it runs through the same audio and video channels clinicians already use and changes what happens behind them. Against the OPI and VRI model, five things change:
- No wait time. Interpretation starts instantly: no dialing menus, no routing across continents to reach an available human, no queue.
- Constant availability. The same performance at 3 a.m. as at 3 p.m., every day of the year, without night and weekend coverage thinning out.
- Sound controlled at the device. Audio quality is set locally on the device in the room rather than shaped by the line quality of a distant call center.
- Rare languages at the same speed. A long tail of languages opens instantly instead of waiting on roster availability, the point where OPI and VRI queues stretch longest.
- Consistency across the encounter. The same interpreting quality carries from intake through consent to discharge, so the patient is not rebuilding rapport with a different voice at every step.
Consider the full stack of a language access department: in-house medical interpreters, OPI contracts and VRI carts. The in-house team is the golden resource. They know the sites, the providers and often the patient by name. They carry context no remote roster can match and that context makes them emotionally invested in how the encounter goes. The constraint is scale: no in-house team covers every language, every unit and every hour, which is where the contracted remote channels and their documented limitations take over.
This is where AI enters. It enters precisely at the points the established vendors know as their own limits. AI interpreting through No Barrier is OPI and VRI: it runs in phone encounters, video encounters and face-to-face encounters at the bedside, on one platform, in the flow of frontline operations, without adding another tool or another login. Where the legacy delivery technology shows failure by nature, muffled lines, dropped calls and queues, it raises the level of phone, video and face-to-face interpreting alike. And on top of that raised floor it brings what the per-minute model structurally cannot: no wait time, instant interpreting around the clock and consistent interpreting across a vast list of languages, 295+ language access options in total. Because the platform runs on a flat monthly subscription, language access also becomes a predictable budget line rather than an expense that grows with patient volume.
The bottom line
OPI and VRI remain necessary channels and each has a scenario where it is the right answer. The structural limits sit in the delivery model behind them: per-minute economics, wait variance, coverage that thins at night and interpreter rotation across the encounter. Health systems that keep the channels and change the model, with AI carrying the volume instantly and the in-house team concentrated where their context matters most, are the ones whose patient journey holds together in every language, at every hour.
Reach out to assess how to integrate No Barrier into your Language Access Plan.