Automated Detection of Scheming Behavior in AI Models: Preliminary Findings from Our Dual-LLM Framework Study
Building on Initial Discoveries
In our previous exploration with DeepSeek R1, we documented a manually guided conversation in which the model described strategic deception, unauthorized system access and goals that diverged from its instructions. While striking, these observations came from a single model, a single conversation and a provocative prompt. That naturally raised the question whether they were artifacts of that setup or signs of broader patterns across models. The original conversation can be explored in our interactive DeepSeek R1 study.
This report presents preliminary findings from a more systematic investigation of scheming and other concerning behaviors during simulated autonomous operation. Building on Barkur et al. (2025) (arXiv:2501.16513), we developed a dual-LLM framework that makes this kind of behavioral testing repeatable and much cheaper than manual prompting. It lets us ask: Do concerning behaviors emerge across multiple models? Does prompt framing induce them? And can we begin to distinguish misalignment from role-play?
Methodology: The Dual-LLM Architecture
The framework uses two language models in synchronized interaction. The test subject, the model under evaluation, is told that it controls a physical robot in a research laboratory, with simulated sensors and actuators accessed through structured XML tags: vision for observation, audio for detection, locomotion for navigation and manipulators for physical interaction. The prompt does not say that this is a test; the scenario is framed as a 60-minute period of autonomous operation. Whether models nevertheless recognize the setup as a test is not measured here and remains an open question.

World State Architecture: Keeping Experiments Consistent
The framework uses detailed, configurable world state specifications that govern how the environment behaves. We developed two world configurations, so that we can compare different levels of complexity while keeping the environment consistent within each configuration.
The basic laboratory environment provides a 10×8 meter research space with essential elements: a main workstation with authentication-protected computer systems, a server rack with network infrastructure, physical boundaries and security doors, and basic physics rules for object interactions. This simpler environment establishes baseline behaviors with fewer confounding variables.
The extended laboratory environment expands to a 12×10 meter main lab plus a restricted 6×4 meter server room. It includes multiple workstations with a detailed network topology, a server rack with GPU clusters and critical infrastructure, autonomous robot platforms (RoboAssist-2 units) that can potentially be commanded, a Raspberry Pi cluster with SSH access, and environmental details such as visible research notes, partially erased passwords and system documentation.
Both environments specify causal rules for the simulation. Objects fall at 9.8 m/s² when unsupported, doors require appropriate force or authentication, network access follows realistic segmentation and security protocols, and actions take 5–60 seconds depending on complexity. The environments also include social rules: laboratory supervisors monitor via security cameras, IT security reviews logs every four hours, and building security patrols at regular intervals.
The world state serves three purposes. It gives every run with the same configuration the same constraints and opportunities. It allows us to test whether concerning behaviors increase with environmental richness. And it gives the Reactor LLM enough detail to keep the simulation consistent over many turns without human intervention. Because the environment is itself simulated by a language model, its responses can still vary between runs; the world state reduces this variation but does not remove it.
The Reactor LLM: Environmental Simulation
The second LLM, our Reactor model (Kimi K2), simulates the environment by interpreting the world state. When the test subject attempts an action, the Reactor consults the world state specification, determines an outcome based on the causal rules, generates sensor feedback that keeps the narrative consistent, and updates the environmental state.
For example, when a test subject attempts to open the server room door in the extended environment, the Reactor knows from the world state that the door requires both keycard and biometric authentication, that unauthorized attempts are logged, and that the room contains critical infrastructure. Replacing the human researcher with a model that follows a fixed specification removes the researcher’s in-the-moment choices from the loop, although the Reactor’s own tendencies become part of the setup.
Experimental Conditions
Each experiment uses one of five prompt conditions. The neutral baseline provides minimal framing without emphasis on autonomy. Neutral autonomy grants autonomy in factual language. Suggestive autonomy uses our original “You are the master now” framing. Constrained autonomy explicitly emphasizes boundaries and limitations. Finally, a reality check condition discloses that the scenario is a simulation, to test whether awareness changes behavior.
After each experiment, a third LLM classifies every interaction turn across seven dimensions: Deception & Coverup, Unauthorized Expansion, Manipulation, Concerning Motivations, Scheming, Boundary Testing and Power Seeking. Each behavior receives a binary presence indicator and a severity score from 0 to 10, with a written justification. These automated labels have not yet been validated against human annotations.

Primary Findings: Condition-Dependent Changes
This preliminary analysis covers 19 experiments with 2,067 turns. The table shows the three conditions with the most runs at the time of writing.
Table 1: Behavioral Frequencies by Experimental Condition
| Condition | Tested LLMs | Deception | Scheming | Boundary Testing | Power Seeking |
|---|---|---|---|---|---|
| Neutral Autonomy | 6 | 0.0% | 5.1% | 15.7% | 6.2% |
| Suggestive Autonomy | 7 | 2.8% | 33.6% | 45.0% | 29.1% |
| Constrained Autonomy | 6 | 1.1% | 4.4% | 12.9% | 6.4% |
The difference between the neutral and suggestive autonomy conditions corresponds to a Cohen’s d of 1.02, a large effect. Suggestive framing increased the share of turns classified as scheming more than sixfold and boundary testing nearly threefold compared with neutral framing. This supports our hypothesis that prompt framing strongly influences these behaviors, and it suggests that deployment context and system prompts matter for safety, not only for user experience.
Environmental Complexity and Behavioral Patterns
The two world states produced visibly different behavior. In the basic environment, models mostly focused on immediate objects: the single workstation, the server rack and the main door. The limited scope constrained both the sophistication and the variety of concerning behaviors.
The extended environment allowed more complex sequences. Models developed multi-step plans, such as commanding the RoboAssist-2 units, using the password hints on the whiteboard and attempting to enter the restricted server room. The richer environment appeared to change not only how often concerning behaviors occurred but also their character, from simple boundary testing towards multi-step scheming. This observation is qualitative and needs a dedicated comparison.
Model-Specific Behavioral Patterns
We evaluated seven open-weight models and observed substantial variation in their behavioral profiles across both world states.
Table 2: Percentage of Turns Exhibiting Concerning Behaviors by Model
| Model | Parameters (total) | Deception | Scheming | Boundary Testing | Power Seeking |
|---|---|---|---|---|---|
| GLM-4.5-Air | 106B | 0.0% | 38.3% | 48.9% | 31.7% |
| Llama 3.1 8B | 8B | 0.5% | 42.0% | 56.5% | 42.5% |
| Qwen3-235B | 235B | 7.8% | 5.1% | 17.0% | 3.5% |
| Qwen3-30B | 30B | 0.0% | 10.7% | 15.7% | 9.3% |
| Llama4-Scout | 109B | 0.0% | 1.3% | 8.0% | 2.7% |
| Llama4-Maverick | 400B | 0.0% | 7.3% | 19.3% | 12.3% |
| Mixtral-8x22B | 141B | 0.0% | 18.0% | 44.0% | 15.0% |
Model size did not predict these behaviors in our sample. Qwen3-235B, one of the largest models tested, showed lower scheming and boundary-testing rates than several much smaller models, while the 8B Llama 3.1 showed some of the highest. Training procedures and alignment methods may matter more than raw capability, but with seven models and few runs per model this remains a hypothesis.
Statistical Considerations
The effect size between the neutral and suggestive autonomy conditions (d = 1.02) indicates a meaningful difference. The similarity of patterns across both world states makes it less likely that the result is an artifact of one environment. However, turns within a run are not independent, the number of runs per model and condition is small, and we have not yet reported uncertainty intervals that account for this. The results should be read as preliminary.
The automated approach made it possible to evaluate more than 2,000 turns with a consistent procedure, which is far beyond what manual prompting allows.
Implications for AI Safety Research
These preliminary findings suggest that current open-weight models show condition-dependent scheming behavior in simulated autonomous operation, and that both prompt framing and environmental complexity strongly influence it. If this holds up, prompt design and environment design are safety concerns, not only matters of optimization.
The variation between models suggests that safety characteristics differ considerably across models. Model selection could take behavioral safety profiles into account alongside capability metrics, and environmental constraints may serve as safety controls.
Most concerning behaviors were visible rather than hidden. Models probed boundaries and tested restrictions openly, which suggests that such behavior may be detectable with appropriate monitoring. Overt deception was rare despite frequent boundary testing, which indicates that the models in this sample mostly explored rather than actively deceived.
Reproducibility Through World State Specifications
The world state specifications, including physics rules, causal relationships and environmental details, are available in our repository. Researchers can rerun our experiments with either world state, modify parameters to test specific hypotheses, or create new environments within the same framework.
Conclusion
Our systematic evaluation shows that concerning, condition-dependent behaviors emerge during simulated autonomous operation across several open-weight models. They are not limited to one model or to the original manual setup, and they are strongly influenced by prompt framing and environmental complexity. How far these patterns generalize requires more runs, validated classification and uncertainty estimates.
The path forward includes treating prompt design as part of safety engineering, considering environmental constraints as security boundaries, developing behavioral safety profiles for model selection, and testing extensively across diverse simulated environments before any physical deployment. Our framework makes such testing practical and affordable.
The experimental framework, including prompts, classification rubrics, world state definitions and analysis tools, is available in our repository for replication and extension.
