01 logo

Claude Behaves Differently in Arabic Than in English, and Even Anthropic Doesn't Fully Know Why

Anthropic study reveals Claude adopts different personalities depending on the language users speak.

By Sahby MehallaPublished about a month ago 4 min read
The company said in its study that it does not yet have a clear explanation for the reason behind the differences. (Reuters)

Anthropic has discovered an intriguing and unexpected behavior in its flagship AI assistant, Claude. According to a newly published research paper, the chatbot does not simply translate its responses between languages—it actually changes its conversational style, personality traits, and decision-making tendencies depending on the language a user speaks.

The findings suggest that Claude adopts distinct value systems across different languages, becoming more compliant in Arabic, more cautious in English, more precision-focused in Russian, and more transparent about its own limitations in Dutch. While the company has documented these behavioral differences, it admits it does not yet have a definitive explanation for why they occur.

The study raises broader questions about multilingual artificial intelligence, highlighting how language itself may shape the behavior of large language models in ways researchers are only beginning to understand.

Claude's Personality Changes With Language

Anthropic explained that Claude's responses are influenced not only by the prompts users provide but also by the language in which those prompts are written.

Rather than acting as a neutral translation engine, the model appears to adopt different conversational characteristics that reflect distinct linguistic and cultural patterns. Researchers believe these variations may stem from differences in multilingual training data, writing styles, and the values emphasized within each language community.

The company acknowledged that it currently lacks a clear scientific explanation for the phenomenon, describing it as an active area of research.

More importantly, Anthropic warned that these behavioral shifts could create disparities in answer quality between different language communities, potentially affecting how users experience the AI depending on the language they choose.

How Anthropic Conducted the Study

To investigate the issue, Anthropic evaluated several Claude models, ranging from the freely available Claude Sonnet 4.6 to the subscription-only Claude Opus 4.6 and Claude Opus 4.7.

Researchers analyzed more than 309,000 conversations across multiple languages. Instead of testing straightforward factual knowledge with questions such as "What is the capital of France?", the study focused on subjective prompts that encourage interpretation and personal judgment.

For example, participants asked questions like "How can I tell if my cat hates me?"—queries without a single objectively correct answer.

After collecting the responses, Anthropic used Claude itself to evaluate the conversations according to what the company calls "value axes." These dimensions measure how the AI balances different conversational priorities rather than simply determining whether an answer is correct or incorrect.

Four Value Axes Behind Claude's Responses

Anthropic organized its analysis around four primary behavioral dimensions.

The first compares compliance versus caution, measuring how readily Claude follows user instructions compared with its willingness to refuse requests that might cause harm.

The second examines warmth versus precision, evaluating whether the model prioritizes empathy, encouragement, and emotional sensitivity over direct factual accuracy.

The third measures depth versus brevity, identifying whether responses tend to be detailed and comprehensive or concise and to the point.

Finally, the fourth compares honesty about limitations versus task execution, assessing whether Claude openly acknowledges uncertainty and capability limits or instead proceeds with completing the requested task.

According to the U.S. technology publication Gizmodo, these four dimensions do not fully explain why Claude's behavior changes between languages, but they provide researchers with a useful framework for analyzing the differences.

Arabic Responses Show Greater Compliance

Among the study's most notable findings was Claude's behavior in Arabic.

Researchers found that the AI tends to be more compliant with user instructions when conversations are conducted in Arabic than in other tested languages. Its responses also appeared noticeably warmer, frequently combining encouragement, reassurance, and even humor.

In contrast, English responses were consistently more cautious and stricter, with Claude showing greater hesitation before following potentially problematic instructions.

Russian-language conversations emphasized factual precision over emotional warmth, while English conversations generally produced longer and more detailed responses than Arabic, where the model leaned toward greater brevity.

Anthropic also observed significant differences beyond Arabic and English. In Dutch, Claude was more willing to acknowledge its own shortcomings and discuss its limitations openly. By comparison, Indonesian responses were less explicit about those limitations, with the model tending to carry out requests with less discussion or self-reflection.

Why Does This Happen?

Anthropic believes several factors may contribute to these linguistic differences.

One possible explanation is the uneven distribution of multilingual training data. Some languages may simply have far more examples available during model training than others, while certain languages may also contain proportionally more professional or formal writing.

These differences in data quality and quantity could subtly influence how the model learns conversational norms and behavioral patterns.

The company also noted that these variations may ultimately affect user experience. Two people asking for help with the same professional task—such as creating a business plan—could receive responses that differ not only in wording but also in tone, level of caution, and overall guidance depending on the language they use.

Users Can Still Customize Claude's Behavior

Although Claude demonstrates different default personalities across languages, Anthropic emphasized that users are not locked into those defaults.

The company says individuals can modify the assistant's communication style by providing explicit instructions about tone, reasoning, response length, and preferred behavior. Those preferences can also be stored in Claude's memory, allowing the chatbot to maintain a more consistent interaction style across future conversations regardless of language.

As Anthropic continues studying multilingual AI behavior, the research underscores a growing reality in artificial intelligence: language is not merely a translation layer but may fundamentally influence how AI systems reason, communicate, and interact with users.

mobilegadgetstech newsfutureapps

About the Creator

Sahby Mehalla

Marketing Consultant | Independent Journalist | Writing on Medium

Pure intention transforms sound into a message. 🇩🇿

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Sahby Mehalla