Not sure what you're arguing here (we shouldn't even try?), but even going fatalistic, that's still fully compatible with "Treat models as untrusted and potentially compromised/hostile and proceed accordingly".
Of course you shouldn't fully trust a model to properly redteam your sandbox, but that doesn't mean you shouldn't redteam your sandbox, including using your own security models to do so.
Of course you shouldn't fully trust a model to properly redteam your sandbox, but that doesn't mean you shouldn't redteam your sandbox, including using your own security models to do so.