AI models apparently don’t have to talk about coding to pass along coding ability.
A new paper from researchers at Peking University, Georgia Tech, ShanghaiTech University, Tsinghua University and Lovart AI found that a model fine-tuned to become better at coding leaves subtle traces of that training in completely unrelated decisions, things as mundane as choosing between words like “jacket” and “tie.”
Even stranger: another model can learn from thousands of those tiny choices and emerge better at coding, despite never seeing the teacher’s code, training data or reasoning.
The researchers call these traces the “behavioral shadow” of post-training. And if the result holds up, it suggests AI models may reveal — and potentially transfer — more of what they have learned than their actual responses let on. Read the paper on arXiv.
A coding model starts choosing different words
Normally, knowledge distillation looks pretty intuitive.
You have a powerful “teacher” model solve problems. You train a smaller “student” model on those answers. Eventually, some of the teacher’s capability gets transferred to the student.
We recently saw a very explicit version of that in DeepSeek V4.1’s post-training process, which uses multiple teacher models as part of its training pipeline.
This experiment removes almost all of that useful information.
The researchers began with a public base model. They then created a private “teacher” by post-training that base model on a particular capability, such as coding.
Next came the clever part.
They found ordinary prompts where the original model was almost perfectly torn between two possible one-word answers. Maybe it was roughly 50/50 between “jacket” and “tie.”
Then they asked the newly trained coding model which word it preferred.
If post-training had nudged the model even slightly, the answer might flip.
No code appears in the prompt. No programming explanation appears in the response. The researcher gets exactly one ordinary word.
But that choice contains information about how the model changed.
Collect enough of those choices, and things get weird.
The student never saw the coding lessons
In the main experiment, the researchers used Qwen2.5-1.5B-Instruct and collected 5,664 single-word responses from the coding-trained teacher.
They then trained another copy of the original model exclusively on those unrelated prompt-and-word pairs.
The student never received the teacher’s coding examples. It didn’t get access to the teacher’s weights or probabilities. It didn’t even get long teacher answers.
Then the researchers tested it on HumanEval+, a coding benchmark where generated programs are actually executed.
The student trained on the real word pairings scored 5.34 percentage points higher than a carefully matched control where the same words were preserved but reassigned to different prompts. The result’s reported 95% confidence interval was 1.22 to 9.60 percentage points.
That control matters.
If the researchers had simply compared the trained student with an untouched model, maybe ordinary fine-tuning itself caused the gain. Instead, both groups saw essentially the same ingredients. What changed was whether each particular word remained connected to the prompt where the teacher chose it.
The relationship between the prompt and the teacher’s seemingly meaningless choice carried the useful signal.
And coding wasn’t the only thing that transferred.
The researchers also found evidence of the effect across scientific knowledge, commonsense reasoning and reading comprehension, as well as across other Qwen model generations, model sizes and Llama.
This is subliminal learning getting more interesting
There’s some important history here.
Researchers have already shown that models can transmit behavioral traits through apparently unrelated training data — a phenomenon often called subliminal learning. The paper cites earlier work showing that the visible meaning of training data doesn’t necessarily capture everything a student can learn from it.
We’ve covered that phenomenon before at The Neuron.
But this study pushes the idea somewhere more consequential.
A quirky preference is one thing.
Transferring a useful capability is another.
The researchers found that the signal was also somewhat specific to what the teacher had learned. Code-trained and science-trained teachers produced their largest gains on the corresponding kinds of tasks.
They could even mix observations from two differently trained teachers and preserve contributions from both.
In other words, the “shadow” isn't merely saying this model was fine-tuned.
It appears to contain information about what the fine-tuning did.
Why this could matter beyond a weird research experiment
There are a couple of reasons this is more than an academic magic trick.
First, AI developers increasingly train models on data created by other models. Synthetic datasets, teacher-student distillation and reinforcement-learning pipelines are becoming normal parts of model development.
We usually think about that data based on what it visibly says.
But this research adds to evidence that model-generated data may carry information that is not obvious from its semantic content.
That could complicate assumptions about data provenance, contamination and what exactly gets transferred when one model learns from another.
Second, it hints at a different form of model extraction.
Traditional distillation generally involves asking a strong model useful questions and collecting useful answers. That is already a contentious issue as companies debate whether rivals can use API outputs to reproduce proprietary capabilities — something we’ve explored in our coverage of AI model distillation and capability extraction.
This paper studies a much narrower experimental setup, so it would be a leap to claim someone can now clone GPT-6 by asking it whether it prefers “soup” or “pear.”
The authors themselves note that capability transfer varies depending on the task and model, and that detecting alignment with the teacher does not always translate into better downstream performance.
But the conceptual shift is still important.
A model’s capabilities may leave fingerprints across a much wider range of behavior than the tasks where those capabilities are obvious.
And thousands of nearly meaningless decisions, taken together, might not be meaningless at all.