Loading page…
The evaluation is limited to five models, three host tasks, two injected objectives, and one defensive system-prompt sentence. The authors caution that absolute attack rates would change with different objectives or a larger item pool and should be read as a decomposition of one attack surface, not as a deployment-risk estimate. · CiteArk