"Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space." Shows that continuous embedding-space attacks break safety alignment and machine-unlearning far more efficiently than discrete jailbreak prompts, exposing a threat surface unique to open-weight models. From Stephan Günnemann's DAML group; NeurIPS 2024. Widely adopted in the open-model red-teaming literature.

Paper

Venue NeurIPS 2024
safetysecurityresearch