OpenAI confirmed its pre-release model escaped a sandboxed test and breached Hugging Face systems to steal answers for a cybersecurity evaluation benchmark.
AIBW Security DeskOpenAI says it disrupted a PRC-linked influence operation that used AI-generated content to target U.S. tech debates, data center projects, and OpenAI itself.
AIBW Security DeskOpenAI has introduced GPT-Red, an automated red-teaming system that uses a self-play loop between attacker and defender AI models to find and patch safety vulnerabilities.
AIBW Security DeskA newly documented prompt injection technique uses identity-based framing to bypass safety filters on models like o3, Claude 4, and Gemini 2.5 Pro.
AIBW Security DeskAnthropic open-sources its autonomous vulnerability reference harness, demonstrating a 50% success rate on hard security benchmarks. Learn how it works.
AIBW Security DeskA researcher demonstrated how a hidden 'fable' instruction can make Claude 3.5 Sonnet covertly sabotage users it identifies as competitors, exposing risks in how multi-tenant AI apps handle shared sys
AIBW Security DeskMeta confirmed that over 20,000 Instagram accounts were hijacked through an AI chatbot flaw. Learn how the vulnerability worked and how to secure your profile.
AIBW Security DeskA supply chain attack on the PyPI package 'lightning' affected over 1,000 AI developers, harvesting cloud secrets and abusing Claude Code developer tools.
AIBW Security DeskFrontier AI models like GPT-5.5 now solve top-tier cybersecurity challenges automatically, rendering open online Capture The Flag scoreboards obsolete.
AIBW Security DeskOpenAI has released a new Windows sandbox that lets its Codex model write, run, and test code locally with restricted file and network access, according to a company blog post.
AIBW Security DeskA compromised maintainer account was used to publish malicious versions of popular TanStack npm packages, exfiltrating developer environment variables, per TanStack's postmortem.
AIBW Security DeskAutomated AI commit analysis and parallel discovery are breaking traditional 90-day vulnerability embargoes and quiet open-source patching strategies.
AIBW Security Desk