[2606.10904]
Bulat Nutfullin, Vladimir Evgrafov, Dmitry Namiot
Comparisons of inference-time defenses for multimodal large language models (MLLMs) depend on more than defense code: the image payload, model text, proxy labels, and judge protocol must refer to the same evaluated event. We audit an experimental archive covering two fixed prompt wrappers and a Gaussian image-perturbation adapter, recorded under aliases derived from RapGuard, AdaShield, and SmoothVLM, across eight InternVL and Qwen-VL models. The planned suite contained seven safety benchmarks and 9,000 inputs. Three benchmark branches fail provenance checks: the MM-SafetyBench renderer used the wrong payload field, the JailBreakV loader admitted text-only fallback, and the adversarial-patch branch did not generate the stated attack. We restrict comparative safety results to FigStep and three text-only corpora, totalling 4,820 inputs per configuration. These are descriptive outputs of a legacy keyword protocol, not validated harmlessness rates; the protocol counted empty strings as safe, and raw responses are unavailable to recover the effect. We also audit 274 adversarial responses; after excluding 28 garbage outputs, 246 comparable responses contain 13 keyword false negatives. This search is nonrandom: it detects evaluator failures but cannot estimate their prevalence or family-level rates, and cannot validate per-cell rankings. The benign audit covers 38,500 stored outputs and 11,637 unique judged responses. Pooled refusal estimate is 0.52%, largest cell 3.24%, sampling-only intervals reach 10.92%; two cells remain not estimable. The archive does not support the earlier mass-refusal interpretation. Measured cost instead appears in batch processing time and defensive-preamble contamination. This work contributes a traceable comparative audit and provenance requirements for future MLLM defense evaluation; it does not claim one defense or an adaptive router is universally superior.