A policy says what should happen. Evaluation shows what the workflow actually does when the input is incomplete, hostile, unusual, or changed by a new model. The CAIO needs both.
We turn expected behavior, prohibited behavior, and human review boundaries into repeatable cases. Each run produces evidence for a release, revision, or stop decision.