Loading page…
Appending one sentence that labels the user message as data had model-specific effects: it nearly stopped naive injection on Gemma and Ministral but not GPT-OSS. With standard tokenization, forged turns still succeeded at 98.3% on Llama and 99.0% on Ministral despite naive rates of 4.1% and 0.0%; nameless tokenization reduced those forged-turn rates to 50.0% and 23.3%. For Qwen's forged-tool attack, nameless reduced success from 53.4% to 0.7%. · CiteArk