Large language models are quietly becoming the invisible intermediaries of global work. When a project manager in Sydney drafts a message for engineers in São Paulo, or a team lead in Berlin summarizes a meeting for stakeholders in Jakarta, an LLM increasingly sits between them, translating, paraphrasing, and smoothing the exchange. A new open-access review published in AI & Society by Yutong Shen, Dilek Cetindamar, and Yi Zhang of the University of Technology Sydney asks a deceptively simple question about this arrangement: can culturally adaptive LLMs actually improve communication and collaboration in cross-cultural project teams, or do they merely make multilingual text look fluent while missing everything that matters culturally? The answer, drawn from a rigorous synthesis of 69 interdisciplinary studies, is more unsettling than the marketing of AI assistants would suggest.
The review’s central finding is a gap between what these models are optimized to do and what global teams actually need. LLMs are trained, evaluated, and celebrated largely on semantic performance: whether a translation preserves meaning, whether a summary is factually consistent, whether a benchmark score improves. But project communication depends not only on what is said, on how, when, and to whom it is said. A semantically perfect translation of an indirect refusal may still fail catastrophically if the recipient expects directness, and a flawlessly worded status update can undermine trust if it ignores local norms about hierarchy and disagreement. The authors call this dimension cultural-pragmatic performance, and they find that current evaluation practices barely touch it.
The methodology behind this conclusion is itself notable for its discipline. The team followed established scoping review guidelines, including the PRISMA-ScR reporting standard and the five-stage framework of Arksey and O’Malley, searching four major databases: Web of Science, Scopus, ABI/INFORM, and EBSCO Business Source, supplemented by targeted screening of ACL and NeurIPS papers and arXiv preprints where methodological detail was sufficient. From 1,452 identified records, the screening funnel narrowed to 228 full-text assessments and finally to 69 included studies. The corpus was organized deductively into four thematic domains: human-AI collaboration, cross-cultural project communication, LLM cultural and multilingual optimization, and the challenge of measuring cultural-pragmatic performance. A CiteSpace keyword co-occurrence network of 312 nodes confirmed that the field is organized around linked but largely separate subareas spanning AI methods, multilingual language technologies, intercultural communication, and project application contexts.
The first thematic domain, human-AI collaboration, reveals a field still deciding what AI even is in this context. Some studies treat AI systems as active communication partners that shape information flow, task coordination, and interaction structure; others treat them as coordination machinery for task execution, resource allocation, and workflow management. Empirical work on generative chatbots shows project managers appropriating them for communication and decision support, with reported gains in documentation and repetitive tasks, but also documented concerns about judgement displacement and communication trust. A third strand examines trust itself as a relational construct, finding that transparency mechanisms, explainable outputs, and clear decision logic matter more to calibrated reliance than raw task performance, and that opaque system behavior or misaligned expectations can quietly corrode it.
The second domain grounds the problem in decades of intercultural communication research. Cultural values demonstrably shape how information is interpreted, how feedback is given, and how conflict emerges in multilingual teams, with cultural intelligence, the capacity to interpret unfamiliar cultural cues and adapt behavior accordingly, consistently associated with better communication outcomes. Studies of high-context and low-context communication styles, drawing on Edward Hall’s classic framework, document how indirect requests, implicit disagreement, and context-dependent emotional regulation create coordination friction when team members operate with different inferential defaults. Perhaps most striking is the evidence on multilingual coordination: language proficiency asymmetries systematically shape participation, influence, and trust formation, meaning that non-native speakers often have less voice in project discussions even when their task expertise is comparable. Language, in other words, is not just a skill but a social cue that distributes power.
Against this backdrop, the technical literature on LLM optimization looks simultaneously impressive and incomplete. Eleven studies in the review focus on cross-lingual semantic accuracy through multilingual pre-training, transfer experiments, and benchmark evaluations, reporting genuine improvements, particularly for high-resource languages, but also persistent disparities for low-resource languages and diminishing returns as language coverage expands. Six studies explore prompting strategies, showing that culturally conditioned prompts, grounded in nation- or community-specific contexts, can shift model outputs in directness, softening strategies, and stance expression. Seven more examine model-level adaptation through culturally conditioned fine-tuning and data augmentation, with measurable but uneven gains that favor well-represented cultures and leave underrepresented ones behind. The pattern is consistent: technical capability is real, but it clusters where data and attention already are.
The evaluation problem is where the review delivers its sharpest critique. The studies that measure LLM performance across languages rely overwhelmingly on semantic or task-oriented metrics. One study found LLM-based translation systems outperformed baselines in roughly 41 percent of evaluated language pairs by BLEU score, a benchmark-specific result rather than proof of general superiority. Another showed that factual consistency varies substantially across languages even when input semantics are stable. Studies of computational politeness can detect features like hedging and honorifics with reasonable accuracy, but treat each feature independently rather than as part of a unified cultural construct. Most damning is the finding on cross-lingual consistency: a model may generate semantically similar answers across languages while reproducing culturally specific expectations about hierarchy, gender roles, or civic responsibility as if they were universal. Consistency, the review argues, is not appropriateness.
There is a constructive thread here too. Research from intercultural communication itself shows that cultural effectiveness can be quantified through multidimensional competence scales covering cognitive, affective, and behavioral dimensions, distinguishing communicative appropriateness and mutual understanding from mere semantic consistency. By analogy, the authors suggest, LLM evaluation frameworks could incorporate an explicit cultural alignment score, assessing whether outputs demonstrate culturally situated interpretation rather than linguistic stability alone. Such a shift would move the field beyond the current situation where technically fluent outputs may still reproduce exclusion, obscure power asymmetries, or weaken trust in exactly the global project environments where the stakes are highest.
The review also situates LLMs within an older and often forgotten ecology of language mediation. Translation and interpreting scholarship has long treated cross-language communication as purposeful, situated professional practice embedded in organizational language policies, workflows, and asymmetrical power relations, not as a lexical mapping problem. When LLMs generate or evaluate multilingual project communication at scale, they enter this ecology without the situated judgement, accountability, and explicit role boundaries that professional mediators carry. The authors are careful not to anthropomorphize: describing LLMs as socio-technical actors does not attribute human understanding to them, but it does identify their effects as outcomes of interactions among model design, organizational practices, professional mediators, users, and governance arrangements. Automation, as translation-technology research has shown, reshapes agency and labor rather than merely adding efficiency.
The practical agenda that emerges is demanding but concrete. Future evaluation should combine semantic performance with human judgment, scenario-based interaction, pragmatic appropriateness, translation-purpose criteria, and governance measures, in line with frameworks such as the NIST AI Risk Management Framework and UNESCO’s ethics recommendation. Organizations should specify when professional translators, interpreters, or bilingual domain experts retain review and decision authority over AI-mediated communication. And the research community itself must confront its blind spots: the review relied on English-language publications, a choice that may underrepresent locally published scholarship and culturally specific practices, shaping both the thematic map and the apparent gaps. The bounded conclusion the authors offer is worth internalizing by anyone deploying AI in a multinational team: LLMs in global project environments warrant analysis not only as productivity tools or language models, but as components of mediated communication operating under conditions of linguistic diversity, cultural difference, and organizational asymmetry. Fluency, it turns out, is the easy part.
Subject of Research: Culturally adaptive large language models in cross-cultural project management
Article Title: Literature review on large language models (LLMs) for cross-cultural project management
Article References: Shen, Y., Cetindamar, D., & Zhang, Y. (2026). Literature review on large language models (LLMs) for cross-cultural project management. AI & SOCIETY. https://doi.org/10.1007/s00146-026-03357-3
Image Credits: AI Generated
DOI: 10.1007/s00146-026-03357-3
Keywords: large language models, cross-cultural communication, project management, cultural intelligence, human-AI collaboration, multilingual NLP, pragmatics, machine translation, AI evaluation, cultural bias, AI governance, scoping review
News Source: Denise Maddox. (October 4, 2026). AI Translators Are Not Culture Experts: What 69 Studies Reveal About LLMs in Global Teams. Scienmag.



