Keep pulling the thread on Sander Schulhoff.
AI guardrails are an ineffective defense mechanism against prompt injection and jailbreaking attacks.
According to Alex Komorosky, the absence of a massive AI-driven attack is due to the early stage of AI adoption, not the effectiveness of current security measures.
The security risks from AI vulnerabilities will increase rapidly with the growing adoption of AI agents, AI-powered browsers, and robots.
A vulnerability in ServiceNow's Assist AI allowed a second-order prompt injection attack, enabling an agent to perform CRUD actions on a database and exfiltrate data via email.
According to Alex Komorosky, there are currently no meaningful mitigations for prompt injection and jailbreaking vulnerabilities in AI systems.
Vision-language model (VLM) powered robotic systems are already being successfully jailbroken, demonstrating a vulnerability to prompt injection attacks.
All currently deployed chatbots, being based on Transformer or similar architectures, are inherently vulnerable to prompt injection, jailbreaking, and other adversarial attacks.
A recent research paper co-authored with OpenAI, Google DeepMind, and Anthropic found that human attackers can break 100% of AI model defenses within 10 to 30 attempts.
The problem of adversarial robustness in AI remains unsolved even by the top researchers at frontier labs like OpenAI, Google, and Anthropic.
The Comet AI browser was recently tricked by a malicious webpage into leaking the user's personal and account data.
There has been no meaningful progress in solving the core problems of adversarial robustness, prompt injection, and jailbreaking in large language models over the past two years.
Sander Schulhoff predicts a market correction within the next 6 to 12 months in the AI security industry, as companies realize that AI guardrail products are ineffective.