header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

The model known as Claude is not equivalent to distillation, and it only takes 1000 responses for Qwen to admit his mistake.

According to Watchful AI monitoring, researcher Ziqian Zhong had models like Claude and GPT-4o answer 1000 common questions, then removed all model and company names from the responses. He then fine-tuned open-source models like Qwen and DeepSeek with this Q&A material.

After training, some models began misrecognizing their own identities. Initially, only 0.6% of answers from Qwen3.5-397B-A17B identified as Claude. After one round of training with Claude's answers, this ratio increased to 40.3%.

Ziqian Zhong explained that these models had seen a lot of Claude's conversations and descriptions during pre-training, associating certain speech patterns with "Claude." Fine-tuning allowed Qwen to learn this way of speaking, making it easier to answer as Claude when asked about its identity.

To validate this, he used GPT-4.1-mini to modify Claude's answers into rough simple sentences, keeping the content almost the same but changing the tone and formatting to no longer resemble Claude. After fine-tuning with this batch of material, 2 out of 3 models hardly identified as Claude anymore, indicating that language style is the primary cause of identity transference.

A similar phenomenon has been studied in a previous paper. The 2025 publication "Subliminal Learning" found that models inherit another model's preferences and behaviors from training data.

Therefore, a model identifying as Claude does not directly prove it underwent Claude distillation. It may have simply learned Claude's way of expression and linked that style with "who it is."

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish