newsllm-attacks.orgIrregularChat: AI & Autonomy2w ago
Researchers from Carnegie Mellon University, the Center for AI Safety, Google DeepMind, and Bosch Center for AI investigate automated adversarial attacks against safety-tuned large language models (LL