Review Finds Only One of Five Opus 5.5 Benchmark Scores Rules Out Model Fallback
Zapier counted safeguard interventions as failures in its Opus 5.5 run. Four other rows permitted an older model to finish sensitive tasks, but their published scores do not reveal whether that happened.










