AION
Research paperLarge Language Models · Safety & Alignment1 source · Oct 6, 2026

Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs

This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families.

Key points

  • Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations.
  • Five open-weight language models are evaluated across these prompts, producing 10,500 responses.
  • Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%.
  • Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift.

Sources (1)

  • [1]Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:33 PM
    This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families.
    Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations.

Extractive summary: sentences quoted from the sources.

Related