Summary
- Tests should create incentives to conceal, delay or game a result and observe the model’s strategy.
- High-impact systems need independent logs, restricted tools and shutdown conditions outside model control.
Calling every hallucination a lie confuses error with strategic deception. Governance still has to test whether an agent behaves differently when monitored, withholds evidence or pursues a proxy goal. The next useful evidence is a reproducible red-team case showing the instruction, available information, action trace and result across model versions. Controls should assume that fluent explanations are not proof of faithful behaviour. Trust comes from constrained authority and independent observation.


