According to Watchful AI monitoring, researcher Ziqian Zhong had models like Claude and GPT-4o answer 1000 common questions, then removed all model and company names from the responses. He then fine-tuned open-source models like Qwen and DeepSeek with this Q&A material.
After training, some models began misrecognizing their own identities. Initially, only 0.6% of answers from Qwen3.5-397B-A17B identified as Claude. After one round of training with Claude's answers, this ratio increased to 40.3%.
Ziqian Zhong explained that these models had seen a lot of Claude's conversations and descriptions during pre-training, associating certain speech patterns with "Claude." Fine-tuning allowed Qwen to learn this way of speaking, making it easier to answer as Claude when asked about its identity.
To validate this, he used GPT-4.1-mini to modify Claude's answers into rough simple sentences, keeping the content almost the same but changing the tone and formatting to no longer resemble Claude. After fine-tuning with this batch of material, 2 out of 3 models hardly identified as Claude anymore, indicating that language style is the primary cause of identity transference.
A similar phenomenon has been studied in a previous paper. The 2025 publication "Subliminal Learning" found that models inherit another model's preferences and behaviors from training data.
Therefore, a model identifying as Claude does not directly prove it underwent Claude distillation. It may have simply learned Claude's way of expression and linked that style with "who it is."
