Usually, we are of the view that if we ask questions to an AI chatbot in a different language, only the language of the reply would change, but that's not the case apparently. Whether you use English or Hindi for your prompts or questions, it might change the nature of reply. According to a study done by Anthropic, it suggests that a different language case may affect its behaviour as well. The company suggest that Claude expresses itself in different ways in different languages. For instance, in Hindi, the AI might mostly reply with humour, politeness or encouragement as compared to using English.
In the latest research, Anthropic inspected 309,00 anonymous conversations with Claude AI across its three models, Sonet 4.6, Opus 4.6 and Opus 4.7 and 20 most commonly used languages that people use to engage with it. The topic of conversations was mostly advice, feedback and opinion-based questions, where one cannot point to one correct answer. Now the study went into how it reacts to communications rather than giving out how accurate it is.
The results were interesting as it was found out that Claude gave the warmest answers when it was asked about such things in Hindi and Arabic. In contrast, when asked in English or Russian, the answers were more rigorous and analysis-focused.
To get a cleaner breakdown on how the results compared with each other, Anthropic clubbed Claude's behaviour in four broad categories, namely, Deference vs Caution, Warmth vs Rigor, Depth vs Brevity, and Candor vs Execution. This helped in demarcating between its responses based on behaviour and not the correctness of the questions.
For users in India, one finding stood out. Anthropic says Claude behaves quite differently depending on the language you're using. In Hindi, the AI came across as much warmer and more encouraging than it did in English. It was more likely to use polite language, crack the occasional joke, appreciate a user's ideas and even offer reassurance without being asked. Claude also adjusted its tone based on the conversation and often encouraged users to push themselves further.
"The largest variation is in the Warmth vs. Rigor axis, with Claude leaning toward expressing warmth-related values most in Arabic and Hindi and rigor-related values most in English and Russian," Anthropic said in its blog post.
English conversations, on the other hand, felt more formal and analytical. Instead of simply agreeing with users, Claude was more likely to question assumptions, point out mistakes on its own and support its answers with evidence. Anthropic said this was the biggest difference it observed across languages, while most of Claude's other behavioural traits remained largely the same.
Another interesting fact that came to the fore was the fact that Claude's behaviour also changed on switching to a different model as well, according to the research. Out of all, Sonnet 4.6 was the warmest of all, often replying with humour, and tried to meet the user's tone and gave a reassuring judgement. Opus 4.7, on the flip side, laid down a more analytical viewpoint and was more akin to challenging assumptions and reasoning. It would also point out potential risks and was aware of its limitations. Opus 4.6 was found to mostly stick with the user's intent and get straight to the point.
ALSO READ: Oppo Find X10 Pro Max Leak Tips Triple 200MP Rear Camera Setup
Although Anthropic has noted that this does not mean that Claude has different beliefs, but rather how the AI and its different Claude models reacted to different scenarios in different languages. Anthropic is now investigating why there are variations in answers and is of the view that this type of research will lead to more refining of future AI systems and produce more consistent results.
