Environments and Benchmarks
September 26, 2026
One big concern in AI Safety circles is eval awareness: if the model can tell it’s being tested, might it dissemble and tell the evaluator it’s not going to turn us all into paperclips? In terms of what models are actually doing right now, though, it seems like much of the problem is that models don’t know enough about what is expected of them.
