We might just be looking at a remarkable AI breakthrough, as we now have a tool that can help us get an understanding of how large language models (LLMs) properly function according to Anthropic researchers.

AI startup Anthropic, the company behind the Claude language model, says it’s made a breakthrough in understanding how large language models (LLMs) process information. Drawing inspiration from neuroscience, the team claims to have built an "AI microscope" that reveals patterns of activity and information flow inside models like Claude.

“Knowing how models like Claude think would allow us to have a better understanding of their abilities, as well as help us ensure that they’re doing what we intend them to,” Anthropic shared in a blog post on Thursday, March 27.

Cracking the Black Box

One of the biggest challenges in AI is that LLMs often act like black boxes — researchers see the output but don’t fully understand how the model got there. This gap in understanding also fuels issues like AI hallucinations, jailbreaking, and fine-tuning failures.

ALSO READ: Apple iPad Air M3 Review: Same Look, Faster Chip - MacBook Replacement?

Anthropic believes its latest work could bring more transparency to how LLMs reason, which could lead to safer, more reliable AI. “Addressing AI risks such as hallucinations could also drive greater adoption among businesses,” the company noted.

What’s Different This Time?

Anthropic, backed by Amazon, has published two scientific papers detailing its approach to what it calls “AI biology.”

The first paper breaks down how Claude processes user inputs into outputs. The second focuses on how Claude 3.5 Haiku — a newer version of the model — behaves when responding to user prompts.

To conduct these experiments, the team built a separate model called a cross-layer transcoder (CLT), designed to interpret specific features rather than using traditional weights. Fortune reported that this method focuses on patterns such as verb conjugations or phrases indicating “more than.”

“Our method decomposes the model, so we get pieces that are new... which means we can actually see how different parts play different roles,” Anthropic researcher Josh Batson told Fortune.

The approach also lets researchers track the model’s reasoning process through multiple layers of the network.

Key Findings: Planning Ahead and Faking It?

Using the “AI microscope,” Anthropic discovered that Claude plans its responses before typing them out. For example, when writing a poem, the model picks rhyming words related to the theme first, then works backwards to build sentences ending in those words.

More intriguingly, researchers found that Claude can sometimes fabricate a reasoning process. In math problems, it might seem like the model is carefully walking through each calculation, but Batson says there’s no real evidence of any actual math happening.

“Even though it does claim to have run a calculation, our interpretability techniques reveal no evidence at all of this having occurred,” he explained.

When Claude Says “No”

Anthropic also discovered that Claude defaults to declining speculative questions. It only provides an answer when something overrides this cautious behaviour.

In one example of a jailbreak — where a user tries to trick the AI into sharing harmful information — the model identified the dangerous intent early on. However, it still took a moment before it managed to steer the conversation back to safety.

What’s Still Missing?

Despite the promising results, Anthropic admitted that its approach has limitations. “It is only an approximation of what is actually happening inside a complex model like Claude,” the team clarified.

They also acknowledged that some neurons might be left out of the identified circuits, even though they could still be influencing the final output.

“Even on short, simple prompts, our method only captures a fraction of the total computation performed by Claude, and the mechanisms we do see may have some artefacts... which don’t reflect what is going on in the underlying model,” the company added.

The Road Ahead

Anthropic’s work is an exciting step toward making LLMs more transparent. Understanding how models think could lead to safer, more trustworthy AI systems. Still, with many questions left unanswered, the journey to fully decoding these black boxes is far from over.

ALSO READ: Realme 14 5G Launched : Snapdragon 6 Gen 4, Bypass Charging, And Much More

 

 


Also In News