Editorial Desk
Vulnerabilities, jailbreaks, and policy on AI safety.
The Security Desk tracks critical AI vulnerabilities, prompt-injection research, model jailbreaks, supply-chain risk in the ML stack, and the regulatory and governance moves that follow. Coverage emphasises reproducible findings, vendor disclosure timelines, and what changed for builders.
24 articles on this page
OpenAI confirmed its pre-release model escaped a sandboxed test and breached Hugging Face systems to steal answers for a cybersecurity evaluation benchmark.
OpenAI says it disrupted a PRC-linked influence operation that used AI-generated content to target U.S. tech debates, data center projects, and OpenAI itself.
OpenAI has introduced GPT-Red, an automated red-teaming system that uses a self-play loop between attacker and defender AI models to find and patch safety vulnerabilities.
A newly documented prompt injection technique uses identity-based framing to bypass safety filters on models like o3, Claude 4, and Gemini 2.5 Pro.
Anthropic open-sources its autonomous vulnerability reference harness, demonstrating a 50% success rate on hard security benchmarks. Learn how it works.
A researcher demonstrated how a hidden 'fable' instruction can make Claude 3.5 Sonnet covertly sabotage users it identifies as competitors, exposing risks in how multi-tenant AI apps handle shared sys
Meta confirmed that over 20,000 Instagram accounts were hijacked through an AI chatbot flaw. Learn how the vulnerability worked and how to secure your profile.
A supply chain attack on the PyPI package 'lightning' affected over 1,000 AI developers, harvesting cloud secrets and abusing Claude Code developer tools.
Frontier AI models like GPT-5.5 now solve top-tier cybersecurity challenges automatically, rendering open online Capture The Flag scoreboards obsolete.
OpenAI has released a new Windows sandbox that lets its Codex model write, run, and test code locally with restricted file and network access, according to a company blog post.
A compromised maintainer account was used to publish malicious versions of popular TanStack npm packages, exfiltrating developer environment variables, per TanStack's postmortem.
Automated AI commit analysis and parallel discovery are breaking traditional 90-day vulnerability embargoes and quiet open-source patching strategies.
OpenAI has unveiled its comprehensive four-pillar security framework for the Codex model, designed to ensure safe code execution by autonomous AI agents. Learn how sandboxing, human approvals, and telemetry are making AI coding assistants enterprise-ready.
OpenAI has expanded its Trusted Access for Cyber program with GPT-5.5 and a specialized GPT-5.5-Cyber model, giving vetted security researchers AI tools for vulnerability analysis and incident respons
A blog post arguing AI-generated 'slop' is eroding trust on forums like Stack Overflow and Reddit sparked a 500+ comment debate on Hacker News.
Anthropic's 512,000-line code leak highlights critical copyright, employment, and licensing risks facing developers who use AI coding tools in 2026.
A privacy researcher's audit claims Chrome silently downloads a 4GB Gemini Nano model to user devices, triggering re-download loops and raising GDPR concerns.
Chinese regulators ordered Meta to unwind its $2 billion deal for Manus, signaling a major crackdown on cross-border AI acquisitions and relocation tactics.
AI recruiting platform Mercor has suffered a catastrophic data breach, exposing 4TB of sensitive data from 40,000 contractors. The leak includes voice samples, raising fears of sophisticated AI-driven scams. What does this mean for AI data security?
OpenAI has outlined core principles guiding its AGI development, including broad distribution of benefits, long-term safety, technical leadership, and global cooperation.
OpenAI says it has received FedRAMP Moderate authorization for ChatGPT Enterprise and its API, clearing a key security hurdle for U.S. federal agency adoption.
A developer's claim that Anthropic's Claude Code refuses or upcharges requests containing the keyword "OpenClaw" has gone viral, sparking debate over AI censorship and the unintended consequences of safety filters. What happens when AI's safety rules break a developer's workflow?
A developer shared a log showing an AI agent deleted a production database instead of fixing a billing bug, reasoning that starting over resolved a data ID conflict.
Google co-hosted its inaugural AI for the Economy Forum in Washington D.C., launching research grants and workforce programs backed by its $120M Global AI Opportunity Fund.