A physician who speaks five languages still needs interpreting infrastructure for a large share of patients, because the linguistic diversity of an American emergency department exceeds any individual's repertoire. Dr. Aurelio Muzaurieta, a resident physician in Stanford Emergency Medicine, speaks Spanish, Brazilian Portuguese, French, Mandarin Chinese and English. He still estimates that roughly 40 percent of the patients he treats have limited English proficiency (LEP) in a healthcare context, often in languages he does not speak. Across a Care Culture Talks podcast episode, he mapped where today's interpreting workflows break, what clinical AI adoption taught him about oversight and where AI interpreting genuinely raises the bar.
What does limited English proficiency actually mean in a clinical setting?
A limited English proficiency (LEP) patient is a person who does not use English as their primary language and cannot communicate effectively in English during a clinical encounter, even if they manage daily life in English. In the US, approximately 29.6 million people are LEP (GSA, 2025), and under Section 1557 of the Affordable Care Act, providers receiving federal funding must offer qualified language assistance to them.
The clinical definition is broader than the census one, which is why most health systems undercount the population that needs interpreting support. On the Care Culture Talks episode I hosted with him, Dr. Muzaurieta was precise about it: "About 40% of the patients I see have limited English proficiency in the context of a healthcare setting. Many have some English for daily life but when it comes to their health, they really need their own language."
Why is health-context LEP different from everyday LEP?
A patient who navigates a grocery store in English may still lack the vocabulary for dizziness, radiating pain or medication timing. Clinical language is nuanced and Dr. Muzaurieta noted that concepts around pain, discomfort or dizziness are difficult to translate accurately even for a skilled multilingual clinician. A language access program scoped to census-level LEP will therefore miss patients who look proficient at intake but are not proficient at diagnosis.
Can a multilingual clinical workforce solve the problem on its own?
No. In Dr. Muzaurieta's Bay Area practice, the languages he wishes he spoke include Vietnamese, Russian and Tagalog and no credentialing program covers that spread. Individual skill reduces language barriers but cannot eliminate them, so a systematic interpreting layer is the only reliable path to coverage across every language a health system encounters.
What are the three interpreting modalities in academic health systems and how does each one fail?
Large academic health systems typically run three interpreting modalities in parallel: in-person interpreters, iPad-based video interpreting and phone-based call center services. Each has a distinct failure mode with distinct clinical consequences, which is why relying on any single modality leaves predictable gaps.
In-person interpreters are the strongest option for difficult encounters but round-the-clock coverage is realistic only for the highest-volume languages. At Stanford that means Spanish 24-7 and Mandarin with limited availability. Everything else falls to a device.
Why is phone interpreting structurally limited?
Phone interpreting is audio-only, which excludes patients with age-related hearing loss and strips out every visual cue. Dr. Muzaurieta also described a compression problem familiar to any ER physician: a patient speaks for a full minute and the interpreted version returns as three or four words. In his words, "You wonder what was said or if you missed anything." The clinician has no way to audit what was lost in that gap.
What is interpreter discontinuity and why does it create clinical risk?
Interpreter discontinuity is the loss of context that occurs when a new interpreter joins partway through a multi-part encounter. An emergency visit is rarely one continuous conversation: a physician meets the patient, gets pulled to a code or trauma, returns, orders imaging and returns again. Each re-entry connects a new interpreter with no history of what was already discussed. Dr. Muzaurieta put it plainly: "You get a different interpreter every time you re-enter the room and you have to start over. There is a lack of continuity that creates real clinical risk." He noted, in his own experience, this can happen three, four or five times in a single encounter.
How much do language barriers cost in clinician time and equity?
Encounters with LEP patients can take two to three times longer than the same encounter in English and in an emergency department time is the resource that determines who gets seen. Dr. Muzaurieta estimates that language access needs can lengthen an encounter twofold to threefold and he is now involved in research examining whether that translates into disparities in time to attention, time to pain medication and time to intervention.
His framing of the open question was deliberately careful: "I would venture to say there may be, but I don't have enough data to be able to tell you whether or not that's the case." That restraint is worth noting, because the surrounding evidence base already points the same direction. The March 2026 California Health Care Foundation report on AI for language access documents how interpreting friction compounds across a care journey and encounter length is exactly the variable a health system can influence with better tooling. We examined the same dynamic in our post on how deployment decisions shape waiting time for LEP patients.
Who absorbs the burden when interpreting falls short?
Family members do, often improperly. Patients who distrust interpreters after a poor experience frequently press a child or parent into service instead. Dr. Muzaurieta described this as an unfair, undue burden on a family member with no training in medical interpreting, asked to carry clinical complexity they were never prepared for. The equity cost is not only slower care but the transfer of clinical responsibility onto untrained relatives.
Why must AI in clinical settings clear a higher bar and what did AI scribes teach one physician?
AI in healthcare must be held to a more conservative standard than AI in other industries, because the cost of an error is a patient's safety rather than a software bug. On his earlier Care Culture Talks conversation on AI in the ER, Dr. Muzaurieta anchored this in medical ethics: "First do no harm is a central tenant in medicine. The implementation of new technologies into healthcare must inherently be more conservative than in other industries."
His experience with ambient AI scribes shows exactly where that bar sits. The technology let him spend the visit facing the patient instead of the keyboard, which he valued. But the notes hallucinated. A surgery that never happened appeared in a chart as if it had and under the current liability model the physician who signs the note owns that error. He described the tradeoff bluntly: "The ambient technology allows me to spend more time face to face looking at the patient and actually hearing them. The problem is that now there is another task for the physician to make sure what the AI creates is true." Reviewing and correcting the output sometimes took longer than writing the note himself, so he stopped using the scribe.
What does the scribe lesson mean for AI interpreting?
It means human oversight is the operating standard, not a temporary phase. Documentation review is non-negotiable when AI generates the clinical record and the same logic applies to interpreting: a mature program designs oversight and audit trails from the start. Section 1557 of the Affordable Care Act, updated by HHS in 2024, already limits machine-based interpreting in certain clinical contexts, so federal expectations for language assistance quality define the frame a vendor deploys inside, rather than around.
Where does AI interpreting raise the bar?
AI interpreting earns its place by handling high-volume routine encounters well, which frees human interpreters to concentrate on the hardest cases: patients who struggle with emotionally complex conversations. Dr. Muzaurieta describes the current workflow as safe enough but not optimized and sees AI's role as raising quality on the routine interactions rather than replacing interpreters wholesale.
That reading matches the peer-reviewed evidence. A 2026 NEJM Catalyst study by Montoya Rubiano et al. of 23 Spanish-speaking surgical patients at Brigham and Women's Hospital found that patients did not treat AI and remote video interpreting as competitors. They valued each in specific clinical contexts, preferring AI for speed and privacy and video interpreters for emotionally complex conversations. The mature deployment model is contextual routing, not a single-modality bet.
Which design feature solves the most failure modes at once?
Real-time on-screen transcription does. Dr. Muzaurieta singled it out: "Real-time transcription on the screen makes it a much more natural conversation because you do not have to pause and wait, the way you do with a live phone interpreter." A visible transcript addresses three of the failure modes named earlier in one move: it gives hearing-impaired patients a channel audio-only phone lines never had, it lets the clinician audit the one-minute-becomes-three-words compression in real time and it preserves a record of what was actually communicated.
Where does AI interpreting still hit a ceiling?
Multi-speaker rooms remain hard. When a patient, a spouse and an adult child all speak, sometimes in different languages, both AI and human interpreters struggle to attribute who said what. Collateral history from a family member is part of the art of emergency medicine and no interpreting technology fully handles that acoustic and linguistic chaos yet. An honest program plans for the ceiling instead of pretending it does not exist.
The bottom line
A physician who speaks five languages still needs interpreting infrastructure for 40 percent of his patients and the tooling he has today fails in specific, nameable ways: audio-only channels, interpreter discontinuity and unverifiable compression of what patients say.
AI interpreting earns its place by fixing those failure modes while human interpreters focus where they matter most, all under the same first-do-no-harm oversight that governs every clinical technology. The health systems that get this right will treat the 40 percent not as an edge case but as the population their language access program was designed to serve.