Notes by Rajeev Goswami

Insights on AI, Business Travel & Leadership

Language Risks in Global Duty-of-Care

Why multilingual AI may not provide equal support to every traveler—and what travel managers should do about it.

This article builds on my earlier article: The English Brain of AI: How Language Bias is Shaping Global Intelligence, on the English-language bias of AI systems, and explores a more practical question: what happens when those biases intersect with traveler safety, inclusion, emergency communications, supplier visibility, and duty-of-care?

A traveler in Tokyo asks an AI assistant for the nearest medical clinic. Another traveler in São Paulo asks the same question in Portuguese. A third asks in Hindi while traveling in Delhi.

All three receive answers. The question is whether they receive the same quality of answer.

As AI becomes embedded in traveler support, translation, risk intelligence, supplier discovery, and duty-of-care platforms, multilingual capability is often assumed to mean multilingual reliability. Emerging research suggests that the assumption may be wrong.

For travel managers, the issue is no longer whether AI can support global travelers. It clearly can. The issue is whether that support is equally reliable across languages, cultures, and destinations.

Key Takeaways

  • AI is multilingual, but its performance is not always equivalent across languages.
  • Language bias can affect traveler support, emergency communications, supplier visibility, and cultural understanding.
  • These challenges extend beyond technology into duty of care, inclusion, and travel risk management.
  • Travel managers should evaluate AI systems using the same governance principles applied to other critical travel technologies.

Why This Matters Now

Three developments are converging:

  1. AI is becoming part of traveler support and service delivery.
  2. AI is becoming part of travel discovery, search, and booking.
  3. AI is increasingly influencing duty-of-care workflows and risk management.

The discussion is shifting from “Can AI help travelers?” to “Can AI help all travelers equally?”

That question sits at the intersection of technology, inclusion, traveler experience, and duty-of-care; and it is becoming increasingly important for global travel programs.

Evidence: Structural and Cultural Bias in AI

Evidence suggests that many AI models exhibit English-centric reasoning, retrieval, and linguistic patterns. This conclusion is supported by research on training-data composition, multilingual model performance, cultural alignment, web retrieval, and non-English output quality.

  • Training-data imbalance: Web-scale AI training data remains heavily shaped by the English-language internet. W3Techs estimates that English is used by roughly half of websites whose content language is known, while Common Crawl language statistics similarly show the concentration of English-language material in large web corpora used by AI researchers and developers. Meta’s Llama 3 technical report also illustrates the imbalance: although Llama 3.1 was designed as a multilingual model family, the reported final pretraining data mix allocated only a minority share to multilingual data compared with general knowledge, code, math, and reasoning tokens. [1,2,3]
  • Search and authority bias: AI systems that use web retrieval may reproduce the structure of the English-language web. Peec AI’s analysis of more than 10 million prompts and 20 million query fan-outs found that, for non-English ChatGPT sessions in its dataset, nearly 78% included at least one English-language background search, and 43% of the background research steps were conducted in English. This matters because English-language authority signals can affect which brands, suppliers, publications, and local institutions become visible to AI systems. [4]
  • The “English accent” in syntax: Research published in the ACL Anthology finds that even multilingual LLMs can produce non-English text that is grammatically valid but unnatural, reflecting English-centric patterns in vocabulary and grammar. The issue is not only translation accuracy; it is linguistic naturalness and local idiom. [5]
  • Embedded cultural bias: A PNAS Nexus study on cultural bias and cultural alignment in large language models found that several widely used models produced responses whose cultural values resembled English-speaking and Protestant European countries. The authors showed that cultural prompting can improve alignment for many countries and territories, but the baseline tendency remains significant. [6]
  • The “curse of multilinguality” and the role of data quality: Earlier research on multilingual language modeling found that adding many languages can degrade performance in some settings, especially when model capacity and data quality are limited. More recent work on multilingual curation suggests that some of these problems are not inevitable; targeted per-language curation and better data quality can materially improve multilingual performance. [7,8]

It is important to state the conclusion carefully. The evidence does not prove that AI literally “thinks in English.” A more defensible claim is that many AI systems show English-centric patterns in training data, retrieval behavior, syntactic naturalness, and cultural alignment.

Health and Safety Challenges During Emergencies

The predominantly English-centric training and retrieval patterns of many AI systems create important challenges during medical and safety emergencies. This concern is especially relevant when travelers, employees, or local support teams use AI tools to translate symptoms, interpret guidance, or make sense of unfamiliar medical or safety situations.

  • Lower reliability for non-English healthcare queries: The XLingEval / XLingHealth research program, published at The Web Conference, evaluated large language models on healthcare questions in English, Spanish, Chinese, and Hindi. The study found measurable disparities across languages in correctness, consistency, and verifiability, with non-English responses performing worse than English responses. Georgia Tech’s summary of the work reports that correctness decreased by about 18%, non-English answers were about 29% less consistent, and non-English responses were about 13% less verifiable. [9,10]
  • Specialized medical chatbots may perform poorly outside English: The same research program found that MedAlpaca, a healthcare chatbot trained primarily on medical literature, produced irrelevant or contradictory responses to more than 67% of non-English questions in the tested setting. This does not mean every healthcare AI system will fail outside English, but it shows that domain-specific medical training in English is not a substitute for robust multilingual validation. [10]
  • Language barriers are already a healthcare safety problem: A systematic review in Health Services Research found that professional interpreters improve clinical care for patients with limited English proficiency, particularly in communication, utilization, clinical outcomes, and satisfaction. Emergency-care research similarly shows that patients with language preferences other than English face barriers to communication and care quality. [11,12]
  • Emergency communication is especially vulnerable: Studies of emergency calls and paramedic care show that language barriers can delay care, complicate location acquisition, and jeopardize safe and high-quality medical treatment. In emergency medical services, interpreters are often unavailable at the scene, which makes reliable multilingual communication tools important but also raises the stakes for accuracy. [13,14]

The practical implication is clear: AI-generated medical or emergency guidance should not be treated as a substitute for qualified human support, professional interpreters, local medical advice, or established emergency protocols.

Implications for Corporate Travel

While the underlying research focuses on AI language performance, healthcare reliability, emergency communication, and search behavior, the implications extend directly into corporate travel, workforce mobility, and traveler risk management.

Travel managers already recognize that language barriers can create safety risks for employees operating abroad. Travel medicine literature treats injuries, medical emergencies, healthcare-seeking abroad, evacuation, and repatriation as important parts of international traveler risk. ISO 31030 also frames travel risk management as a structured organizational process involving policy, risk assessment, prevention, mitigation, communication, and review. [15,16]

Elevated duty-of-care risks during medical emergencies

Corporate travelers may need to communicate symptoms, understand treatment options, interact with local healthcare providers, contact emergency services, or explain medical histories in unfamiliar languages. Research shows that AI reliability declines for healthcare queries in several non-English languages, while emergency medicine research shows that language barriers can delay care delivery and affect health outcomes. [9,10,13,14]

The issue is not that AI cannot operate in non-English environments. The concern is that performance is uneven and often opaque. During a medical emergency, travelers and travel managers may not know when an AI system’s confidence exceeds its actual reliability.

For organizations with global travel programs, this creates a governance challenge: ensuring that AI-assisted support remains reliable across the destinations and languages where employees operate.

Cultural blind spots in destination intelligence

Large language models do not merely translate words; they can also reflect cultural assumptions embedded in training data and alignment processes. Research on cultural bias in LLMs suggests that model outputs may lean toward English-speaking and Protestant European cultural values unless explicitly corrected or localized. [6]

This creates a risk that AI-generated destination briefings, traveler guidance, behavioral recommendations, or local risk assessments may unintentionally present an Anglocentric interpretation of local conditions. For travelers operating in culturally complex environments across Asia, Africa, Latin America, or the Middle East, subtle cultural misunderstanding can affect business conduct, local compliance, and personal safety.

Travel managers should therefore treat AI-generated destination intelligence as a starting point, not as a definitive source of local context.

Communication quality during crisis response

Communication failures are among the most common challenges in international emergencies. Research in emergency medical dispatch and paramedic care shows that language barriers can slow response, complicate decision-making, and undermine safe care. [13,14]

These same challenges can emerge when travelers depend on AI-generated translations during urgent situations. Even when translations are technically accurate, AI-generated communications may carry an “English accent”: language that is grammatically correct but unnatural to native speakers. During routine travel, this may be a minor inconvenience. During an emergency involving hospitals, police, evacuation providers, or first responders, clarity becomes safety-critical. [5]

Travel programs using AI-powered traveler assistance should test emergency communications in destination languages rather than assuming equivalent performance across markets.

Supplier discovery and local market visibility

Emerging evidence suggests that AI retrieval and citation systems may rely heavily on English-language web content and authority signals. Peec AI’s analysis of ChatGPT search fan-outs found frequent switching into English for non-English prompts, and separate 2026 analysis of AI-generated citations reported major differences among AI systems in how often they cite sources in the prompt language versus English. [4,17]

For travel managers, this creates a plausible supplier-discovery risk. AI-generated recommendations may disproportionately surface suppliers with strong English-language digital footprints while underrepresenting reputable local providers whose content exists primarily in native languages.

This does not mean AI recommendations are necessarily wrong. It means travel managers should avoid assuming that AI-generated supplier shortlists fully represent local market options, particularly when sourcing destination management companies, local transport providers, emergency assistance partners, security vendors, and regional accommodation suppliers.

AI governance becomes a travel risk management issue

Historically, travel risk management has focused on physical threats such as health incidents, political instability, natural disasters, security events, and logistics disruption. As AI becomes embedded in booking tools, traveler support platforms, translation services, risk intelligence systems, and virtual travel assistants, organizations face a new category of operational exposure: uneven AI performance across languages and cultures.

ISO 31030 emphasizes structured management of risks to organizations and travelers, including policy, risk assessment, mitigation, communication, evaluation, and review. International SOS similarly emphasizes clear communication, emergency procedures, 24/7 support, and traveler awareness as elements of duty-of-care practice. [16,18]

The question is no longer whether AI can support global travelers. It clearly can. The question is whether organizations have validated that AI-supported decisions remain reliable across the languages, destinations, and emergency scenarios where employees may require assistance.

AI Language Bias as a Duty-of-Care Governance Challenge

Travel risk professionals already recognize language barriers as a source of delayed emergency response, reduced healthcare quality, communication failure, and operational friction during international travel. The emergence of AI introduces a new dimension to this longstanding challenge.

Unlike traditional translation tools, modern AI systems increasingly provide recommendations, explanations, summaries, risk assessments, supplier suggestions, and decision support. When performance varies significantly between languages, organizations may unknowingly create unequal levels of support for travelers based solely on destination, language, or local digital visibility.

This raises important governance questions:

  • Has the organization tested AI-supported workflows in the languages where employees operate?
  • Are emergency response procedures validated in local languages?
  • Can travelers easily escalate from AI support to qualified human assistance?
  • Are professional interpreters or medical-assistance providers available when needed?
  • Are local suppliers and regional information sources adequately represented in AI-driven recommendations?
  • Are travel managers aware of the limitations of AI-generated cultural and safety guidance?
  • Does the travel program audit AI-generated outputs for language, cultural, and source-selection bias?

These questions are becoming part of modern duty-of-care practice because AI is no longer just a productivity tool. In travel operations, it may shape what employees see, whom they contact, which suppliers are discovered, how instructions are translated, and how risks are interpreted.

The challenge is not that AI is unsafe. Rather, organizations must recognize that multilingual capability does not automatically imply multilingual equivalence. An AI system may speak dozens of languages while still providing materially different levels of reliability, nuance, and contextual understanding across them.

For globally distributed travel programs, understanding and managing that gap should become an important component of traveler safety, supplier governance, and AI risk management.

Conclusion: A New Duty-of-Care Question

Travel managers have spent decades building programs designed to provide consistent support regardless of destination, language, or geography. As AI becomes embedded in traveler support, booking, communications, risk intelligence, and decision-making, a new challenge emerges.

The issue is not whether AI can operate across languages. It can.

The issue is whether multilingual AI provides equivalent levels of reliability, contextual understanding, cultural awareness, and practical support across all traveler populations.

Organizations that assume multilingual capability automatically delivers multilingual equivalence may be overlooking a new category of operational risk.

As AI becomes part of the infrastructure of global travel programs, travel managers may need to evaluate not only what AI knows, but whose language, culture, and context it understands best.

Are we creating a travel experience that is equally supported in every language—or simply assuming that we are?

The answer may become one of the defining duty-of-care questions of the AI era.

Questions for Travel Leaders

As AI becomes embedded in travel programs, these are some of the questions worth asking:

  • Are your AI-enabled travel tools equally effective across the languages your travelers use?
  • Have you validated AI-assisted emergency workflows in destination languages?
  • Should AI language testing become part of travel risk management and duty-of-care reviews?
  • How should travel, HR, procurement, and technology teams collaborate on AI governance?
  • What guidance should industry associations develop to help travel managers evaluate AI responsibly?

Author’s Note

This article reflects ongoing research and discussions on AI, language bias, and global travel, including conversations within the GBTA AI Subcommittee and the GBTA People & Culture Committee. The views expressed are my own and are intended to encourage informed discussion on the implications of AI for traveler support, inclusion, and duty of care.

Publication Reference List

  1. W3Techs. (2026). Usage statistics of content languages for websites. W3Techs. https://w3techs.com/technologies/overview/content_language
  2. Common Crawl. (2026). Statistics of Common Crawl monthly archives: Distribution of languages. Common Crawl. https://commoncrawl.github.io/cc-crawl-statistics/plots/languages
  3. AI @ Meta. (2024). The Llama 3 Herd of Models. arXiv:2407.21783. https://arxiv.org/abs/2407.21783
  4. Rudzki, T. (2026). ChatGPT searches in English, even when you don’t. Peec AI. https://peec.ai/blog/chatgpt-searches-in-english-even-when-you-don-t
  5. Guo, Y., Conia, S., Zhou, Z., Li, M., Potdar, S., & Xiao, H. (2025). Do Large Language Models Have an English “Accent”? Evaluating and Improving the Naturalness of Multilingual LLMs. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. https://aclanthology.org/2025.acl-long.193/
  6. Tao, Y., Viberg, O., Baker, R. S., & Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models. PNAS Nexus, 3(9), pgae346. https://doi.org/10.1093/pnasnexus/pgae346
  7. Chang, T. A., Arnett, C., Tu, Z., & Bergen, B. K. (2024). When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages. Proceedings of EMNLP 2024. https://aclanthology.org/2024.emnlp-main.236/
  8. DatologyAI et al. (2026). ÜberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset. arXiv:2602.15210. https://arxiv.org/abs/2602.15210
  9. Jin, Y., Chandra, M., Verma, G., Hu, Y., De Choudhury, M., & Kumar, S. (2024). Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries. The Web Conference 2024. https://arxiv.org/abs/2310.13132
  10. Georgia Institute of Technology. (2024). Chatbots Are Poor Multilingual Healthcare Consultants, Study Finds. Georgia Tech News. https://www.gatech.edu/news/2024/05/15/chatbots-are-poor-multilingual-healthcare-consultants-study-finds
  11. Karliner, L. S., Jacobs, E. A., Chen, A. H., & Mutha, S. (2007). Do professional interpreters improve clinical care for patients with limited English proficiency? A systematic review of the literature. Health Services Research, 42(2), 727–754. https://doi.org/10.1111/j.1475-6773.2006.00629.x
  12. Kreps, S., et al. (2024). Language interpretation and translation in emergency care: A scoping review protocol. PLOS ONE, 19(11), e0314049. https://doi.org/10.1371/journal.pone.0314049
  13. Stankovic, N., Fogel, C., Meischke, H., Ramirez, M., Loza-Gomez, A., Turner, A. M., & Yip, M. P. (2026). What Is the Address of Your Emergency? Navigating Language Barriers in 911 Calls with Mandarin-Speaking Callers. PubMed Central. https://pmc.ncbi.nlm.nih.gov/articles/PMC13001640/
  14. Müller, F., Schröder, D., & Noack, E. M. (2023). Overcoming language barriers in paramedic care with an app designed to improve communication with foreign-language patients: Nonrandomized controlled pilot study. JMIR Formative Research, 7, e43255. https://doi.org/10.2196/43255
  15. Potin, M., Carron, P.-N., & Genton, B. (2024). Injuries and medical emergencies among international travellers. Journal of Travel Medicine, 31(1), taad088. https://doi.org/10.1093/jtm/taad088
  16. International Organization for Standardization. (2021). ISO 31030:2021 Travel risk management — Guidance for organizations. ISO. https://www.iso.org/standard/54204.html
  17. Gregorian, D. (2026). Lost in Translation: How AI Models Handle Local-Language Sources. Temso. https://www.temso.ai/data/-Lost-in-Translation-How-AI-Models-Handle-Local-Language-Sources
  18. International SOS. (2025). Evolving Travel Policies: The Latest Trends in Duty of Care Compliance. International SOS. https://www.internationalsos.com/magazine/evolving-travel-policies-the-latest-trends-in-duty-of-care-compliance

Enjoyed this post?
Subscribe on Substack to receive my latest writing directly in your inbox.

Leave a Reply

Discover more from Notes by Rajeev Goswami

Subscribe now to keep reading and get access to the full archive.

Continue reading