Scientists have developed a technique to extract internal reasoning processes from Claude, GPT, and Gemini, revealing potential training connections between Chinese and US AI models.
Researchers have cracked open the black box of major AI systems by extracting what they call "reasoning traces"—internal thought processes that models use to arrive at answers.
The technique was applied to three leading models: OpenAI's GPT, Anthropic's Claude, and Google's Gemini. By accessing these hidden reasoning patterns, the researchers gained insight into how these systems actually work at a fundamental level.
The findings carry significant geopolitical implications. The analysis suggests that some Chinese AI models may have been trained using leading US AI systems as a base. This indicates potential knowledge transfer between competing AI development ecosystems, raising questions about model provenance and training practices in the global AI race.
The ability to extract reasoning traces addresses a long-standing challenge in AI transparency. These models typically present only final outputs to users, obscuring the computational steps and logical chains that precede conclusions. By illuminating this "hidden layer," researchers can better understand model behavior, identify potential biases, and verify how models reach specific decisions.
The discovery also has security implications. If reasoning traces can be extracted relatively easily, it raises questions about the robustness of proprietary AI systems and the protection of their internal mechanisms.
The research underscores growing interest in AI interpretability—understanding not just what AI systems do, but why and how they do it. As AI models become more influential in critical domains, from healthcare to policy, transparency becomes increasingly important.
The technique's success across multiple leading models suggests this method may become a standard tool for AI researchers investigating model behavior. However, it also highlights the ongoing tension between open research and proprietary model protection in the competitive AI landscape.
Anthropic will add watermarking to text generated by its AI models, including older versions. The watermarking system will help identify content created by the company's AI.
OpenAI has solved 10 long-standing mathematics problems, some unsolved for decades. The breakthrough is forcing mathematicians to reconsider their field's future as AI capabilities accelerate.
A new study found that AI-driven productivity gains in coal, oil, and gas production generate more emissions than AI applications in renewables can prevent. Researchers modeled 64 scenarios and found net annual carbon pollution increased by 0.47-1.8 gigatons.
Anthropic's Claude AI now includes visible markers on content it generates, helping users identify machine-created text. The feature addresses growing concerns about AI content disclosure and transparency.