Explore ArkGraph and select the steps to run.
Open-weight chat models expose the surface strings that their tokenizers map to reserved turn, role, reasoning, and tool identifiers, allowing attacker-controlled text to forge structural tokens. The paper audits 256 deployed chat tokenizers, evaluates a tokenizer-only defense called nameless tokenization on five model families, and compares it with string sanitizers and training-level instruction–data separation. Nameless tokenization removes text-to-control-identifier mappings while preserving template-written identifiers and message characters. Reported experiments cover attack-free utility, delimiter-bearing content, several forged-turn attacks, two injected objectives, two system-prompt conditions, and model-specific outcomes. Results show exact attack-free token-stream preservation, much better delimiter fidelity, and strongly context-dependent security gains.
The paper identifies a concrete interface flaw: ordinary text can become the same reserved identifiers used by a chat template. Nameless tokenization closes that text-to-identifier channel without deleting or rewriting user content, so it can complement prompt- and training-level defenses. It does not solve ordinary-language prompt injection, and the reported attack rates come from five models, three host tasks, two injected objectives, and one defensive prompt sentence rather than from a deployment-risk estimate.
The paper’s claims are available in Research claims.