Loading page…
Every evaluated defense reproduced standard tokenization's token stream on 100.0% of attack-free host prompts; mean host-task accuracy ranged from 88.6% to 89.1%, with the paper attributing the small spread to non-bitwise-reproducible batched decoding on GPT-OSS-20B because the other four models scored identically across conditions. · CiteArk