Based on the framework, we design JailbreakLens, a visual analysis system that enables users to explore the jailbreak performance against the target model, conduct multi-level analysis of prompt characteristics, and refine prompt instances to verify findings. Through a case study, technical evaluations, and expert interviews, we demonstrate our system’s effectiveness in helping users evaluate model security and identify model weaknesses.
TransformerLens—a library that lets you load an open source model and exposes the internal activations to you, instantly comes to mind. I wonder if Neel’s work somehow inspired at least the name.
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models (non-peer-reviewed as of writing this)
From the abstract:
TransformerLens—a library that lets you load an open source model and exposes the internal activations to you, instantly comes to mind. I wonder if Neel’s work somehow inspired at least the name.