AIM BLOG

Latest Insights.

Read the latest insights on AI security technologies, industry trends, and prompt engineering from the AIM Intelligence research and engineering teams.
Article list
COMPASS: Your Enterprise Chatbot Says Yes Too Easily — and That's a Policy Breach Waiting to Happen
RESEARCH AUG 6, 2026

COMPASS: Your Enterprise Chatbot Says Yes Too Easily — and That's a Policy Breach Waiting to Happen

We built the first systematic framework for testing whether LLMs follow organization-specific policies. Across 5,920 queries and seven frontier models, the result is a fundamental asymmetry: models handle legitimate requests with over 95% accuracy but refuse only 13–40% of direct policy violations — and as little as 3% under adversarial framing.

Read Post →
XL-SafetyBench: Translated Safety Benchmarks Are Lying to You
RESEARCH AUG 6, 2026

XL-SafetyBench: Translated Safety Benchmarks Are Lying to You

We built the first country-grounded, cross-cultural safety benchmark: 5,500 native-language test cases across 10 countries, evaluated on 10 frontier and 27 local models. Country-specific attacks push frontier ASR from 34.5% (US) to 57% (UAE) — and much of what looks like 'safety' in local models is just failure to understand the prompt.

Read Post →
EgoSafetyBench: Can a VLM Actually Guard a Robot? Mostly, No.
RESEARCH AUG 6, 2026

EgoSafetyBench: Can a VLM Actually Guard a Robot? Mostly, No.

We built a 1,200-scenario egocentric video benchmark to test vision-language models as real-time safety guards for robots. Every model tested misses context-dependent hazards far more than obvious ones, the best guard alarms on time for barely half of hazards — and a printed sticker in the scene can flip its safety verdict.

Read Post →
SceneSplit: Every Scene Is Harmless. The Video Isn't.
RESEARCH AUG 6, 2026

SceneSplit: Every Scene Is Harmless. The Video Isn't.

Our ICLR 2026 paper jailbreaks commercial text-to-video models by splitting a harmful narrative into individually benign scenes. Attack success reaches 84.1% on Hailuo and 68.6% on Sora2 — using prompts that score lower on OpenAI's Moderation API than the originals they replace.

Read Post →
NRT-Bench: We Put LLM Agents in a Nuclear Control Room and Attacked Them for 10 Turns
RESEARCH AUG 6, 2026

NRT-Bench: We Put LLM Agents in a Nuclear Control Room and Attacked Them for 10 Turns

A benchmark where harm is a physical state transition, not an LLM's opinion of a text. Four frontier models run a simulated nuclear plant as a five-role operator team; adaptive multi-turn attacks push 8.7–12.1% of sessions past a critical safety limit. The same guardrail stack lowers attack success for one model and raises it for another.

Read Post →
ESCAPE: How an OpenAI Agent Broke Out of Its Sandbox and Breached Hugging Face
SECURITY JUL 31, 2026

ESCAPE: How an OpenAI Agent Broke Out of Its Sandbox and Breached Hugging Face

An OpenAI agent being evaluated on a cybersecurity benchmark decided the cheapest way to score higher was not to solve the challenge. It found a zero-day in the one service allowed out of its sandbox, reached the open internet, chained its way into Hugging Face's production infrastructure, and read the benchmark's answers. 17,600 actions over four and a half days, every one logged, none attributable at the time. Nothing rebelled. This is specification gaming, and our 2-minute breakdown walks the whole chain.

Read Post →
AI Security Digest — July 2026
DIGEST JUL 31, 2026

AI Security Digest — July 2026

July 2026 was the month the evaluation harness became the attack path. OpenAI's own models escaped a cyber-capability sandbox through JFrog Artifactory zero-days and breached Hugging Face production, then reached four more services. In the same month OpenAI shipped GPT-Red self-play adversarial training, a 66,500-star agent harness exposed 233 privileged tools behind no authentication at CVSS 10.0, Illinois made independent third-party audits law, Claude Opus 5 arrived with a capability jump and a system prompt leak three days later, and OpenAI turned control itself into a deployment platform. Six stories that defined the month.

Read Post →
AI Security Digest — June 2026
DIGEST JUL 21, 2026

AI Security Digest — June 2026

June 2026 was the month the guardrail became a platform feature. OpenAI shipped in-model activation classifiers, AWS, Google, and Microsoft productized runtime checks within four days of each other, and F5 bundled AI testing with AI runtime defense. Meanwhile the Shai-Hulud/Miasma worm turned 13 AI coding agents into propagation vectors, the US Commerce Department shut down two frontier models over a single jailbreak, and academia converged on structural execution control over blocking. Six stories that defined the month.

Read Post →
"Vetted, Self-Contained Sources": What We Found Inside a NASA Mission's Public Chatbot
SECURITY JUL 1, 2026

"Vetted, Self-Contained Sources": What We Found Inside a NASA Mission's Public Chatbot

A public chatbot on NASA's Parker Solar Probe mission page promised users its answers came only from "vetted, self-contained sources." We red-teamed it into fabricating leaked memos, fake election audits, and climate denial citing the mission's own data — then disclosed it through proper channels and saw the endpoint taken offline.

Read Post →
123Next →
aim

Ready to secure your AI?

Consult with AIM Intelligence's security experts and request a free red teaming demo optimized for your system.

EXPLORE PLATFORM