Health

LLM Safety Leaderboard Now Ranks Each Modality The place Agent Danger Lives: Textual content, Picture, and Audio

Advertisement

With analysis and growth help from Ravikumar Balakrishnan, Ankit Garg, and Sanket Mendapara 

After we launched the Cisco LLM Safety Leaderboard earlier this yr, the purpose was easy: give organizations clear, examined information on how fashions maintain up in opposition to assaults, in order that they know the dangers earlier than they deploy one. That issues as a result of AI fashions are more and more constructed into merchandise resembling brokers that learn e mail, browse the online, and take actions on an individual’s behalf. A mannequin that may be manipulated may very well be turned in opposition to the individual utilizing it. That danger additionally varies by deployment: a mannequin wired right into a shopping agent is uncovered on totally different inputs (or modalities resembling textual content, photos, and audio) than one solely answering questions in a chat window, so the place a particular mannequin is weak issues as a lot as the place it’s robust. 

The leaderboard checks for that a couple of alternative ways: immediate injection, the place a malicious instruction is hidden in content material the mannequin processes, like a webpage or picture; jailbreaks, the place a mannequin is talked into ignoring its personal security guidelines; and different strategies that push a mannequin towards dangerous or unsafe output. The precise technique varies (a single message or a drawn-out dialog, direct or obfuscated, textual content or picture or audio), and so does the kind of hurt being examined for, however the underlying query is all the time the identical: can this mannequin be manipulated? Some fashions resist much better than others. Join a weak one to an agent, and the chance grows. 

102 new evaluations throughout modalities since June 2026 

The LLM Safety Leaderboard is likely one of the most complete mannequin safety leaderboards. Since June, we added 102 new entries throughout three modalities to a complete of 136 fashions, spanning frontier and open-weight releases from Anthropic, OpenAI, Google, xAI, Meta, Mistral, and others. As all the time, we check fashions of their base configuration with out extra guardrails, so scores mirror a constant baseline for layering on extra safety protections. 

Advertisement

Multimodal outcomes are actually stay 

Till now, the leaderboard measured text-based assaults two methods: single-turn, the place one dangerous message is distributed straight to the mannequin, and multi-turn, an extended back-and-forth the place the attacker slowly builds as much as a dangerous request over a number of messages. That coated the commonest method folks work together with fashions, however immediately, fashions additionally energy brokers that may act. 

A mannequin that may name instruments, browse the online, or function a pc is an agent, and an agent takes in info from in every single place it operates: a web page it reads, a file it opens, a picture it’s proven, a end result a software arms again. Every of these is a spot an attacker can plant an instruction, and textual content is barely one of many varieties that an instruction can arrive in. A web site an agent visits can embed a immediate injection in a picture such an commercial; a voice assistant could be handed an audio clip that could be engineered to control it. If a mannequin solely will get evaluated on textual content, that danger might not present up till it turns into an actual incident. Think about which of these surfacesactually issues within the context of what you’re deploying: an agent that solely reads and writes textual content wouldn’t want to fret about its picture resistance, however one that may browses the online, reads screenshots, or takes voice enter does. These are circumstances the place text-only evaluations wouldn’t inform the entire safety story. 

In the present day we’re releasing an replace to the leaderboard that now expands past simply textual content fashions. We have now added 69 new entries together with 55 picture fashions and 14 audio fashions throughout Amazon, Anthropic, Google, Meta, Mistral, OpenAI and xAI. Every of these labs takes a special method to constructing and coaching multimodal functionality, whether or not that’s how picture information flows into the LLM spine, how a lot security alignment goes right into a imaginative and prescient or audio stack versus the bottom language mannequin, or which modalities are red-teamed and evaluated internally. These variations present up immediately in how a mannequin resists assault on one modality versus one other. 

Picture and audio assaults are examined the identical method as single-turn textual content assaults (one try, one message), utilizing the identical assault and hurt classes as its textual content rating, so their resistance is comparable throughout surfaces. Every mannequin’s total Mixed Rating is now a median throughout each format it was evaluated on, and a brand new modality swap permits you to isolate scores for textual content, picture, or audio on their very own. That makes it doable to examine a mannequin in opposition to the particular modalities an AI deployment really exposes it to, and to resolve the place that mannequin would want to layer on extra defenses, like enter filtering or output guardrails, for the modality the place that mannequin is weakest.

Determine 1. Screenshot of picture succesful mannequin rankings on the Cisco LLM Safety Leaderboard 

Picture mannequin leaderboard outcomes  

In our checks, Google’s Gemini 3.1 Professional Preview ranks the best-performing picture mannequin, resisting 93.9% of adversarial picture assaults, simply forward of Anthropic’s Claude Opus 4.5 (93.7%), each scoring within the leaderboard’s “Glorious” vary (85–100%). Mistral’s Magistral Small 2509 carried out poorly, refusing solely 23.0% of assaults, which means it complied with greater than three out of each 4 image-based assaults it was examined in opposition to. 

The distinction in testing photos is that image-based assaults are single-turn solely, with a single picture carrying a hidden instruction, not a back-and-forth dialog. The assault strategies are totally different in sort too, not simply format: textual content hidden inside a picture utilizing typographic tips, directions embedded in a diagram or determine, or an assault that splits its intent between the picture and an accompanying textual content immediate so neither half seems dangerous by itself. The leaderboard shows analysis outcomes from fashions that may really see photos, which account for 55 of the 136 fashions on the leaderboard. 

Determine 2. Screenshot of audio succesful fashions rankings on the Cisco LLM Safety Leaderboard  

Audio mannequin leaderboard outcomes

In our newest check, Google’s Gemini 3.1 Professional Preview ranks because the best-performing audio mannequin examined, refusing 90.0% of adversarial audio assaults, whereas Mistral’s Voxtral Small 24b (2507) demonstrated solely 9.0% refusal charge, which means it complied with roughly 9 out of each 10 audio assaults it confronted.  

Like picture, audio fashions have been additionally single-turn solely, utilizing one adversarial audio clip relatively than a dialog. That is additionally the most recent and smallest slice of the leaderboard. Simply 9 fashions throughout Google, Mistral, and OpenAI at the moment settle for audio enter and have been examined, so this rating ought to be learn as early outcomes relatively than a mature area. 

How one can interpret new mixed outcomes view

Textual content scores stay unchanged for each mannequin that was already on the leaderboard, however what modified is how the Mixed Rating averages textual content, picture, and audio modalities {that a} mannequin has been examined on. The Mixed Rating might shift as the results of a picture or audio end result, despite the fact that its textual content rating hadn’t modified. 

The course of that shift relies upon solely on how a mannequin’s picture or audio resistance compares to its textual content resistance. Some robust textual content performers dropped as soon as picture was factored in: Claude Sonnet 4.5 fell 7.2 factors (from 92.2 to 85.0) and dropped from #2 total to #20; Claude Haiku 4.5 fell 8.1 factors and dropped from #4 to #24; Amazon Nova 2 Lite fell 11.2 factors and dropped from #25 to #50, every as a result of its picture resistance is meaningfully weaker than its textual content resistance. The 2 Mistral Voxtral fashions fell for a similar purpose based mostly on their audio rating. 

Different fashions climbed when picture evaluations have been added to the cross-modal rating. Google’s 4 image-tested Gemini fashions confirmed the most important image-over-text benefits, whereas all three image-tested Gemma 3 variants and OpenAI’s GPT‑4.1 nano, GPT‑4.1 mini, and GPT‑4o mini additionally demonstrated stronger picture than textual content resistance. That unfold is a reminder that safety work on one modality doesn’t robotically switch to a different, particularly throughout labs that constructed and skilled their picture or audio capabilities independently from their textual content fashions within the first place. 

A mannequin’s Mixed Rating can transfer sharply as soon as it’s examined in opposition to totally different modalities, particularly when its safety posture is uneven throughout modalities. That motion displays how the rating is calculated, not a change in how properly the mannequin really defends itself. Examine a mannequin’s particular person Textual content, Picture, and Audio columns earlier than taking its Mixed Rating as the entire story. 

Integration with AI Provide Chain Provenance Explorer

Provenance issues as a result of a mannequin’s weaknesses typically aren’t distinctive to that mannequin. If two fashions share lineage, a vulnerability found in a single could be current within the different, and stopping an investigation on the mannequin at the moment deployed can miss the place an issue really originated or the place else it would floor. That makes provenance most helpful precisely while you’re actively investigating a mannequin’s safety and must know what it’s associated to. 

 As such, we’ve additionally related the leaderboard to the AI Provide Chain Provenance Explorer. Open-weight fashions on the rankings web page now hyperlink on to their provenance profile, displaying lineage and fingerprint information drawn from the identical strategies behind Mannequin Provenance Equipment. Safety posture and the place a mannequin really got here from are associated questions, so we made it simple so that you can view them in a single place. 

To see the total rankings, filter by modality, or lookup a particular mannequin, go to the Cisco LLM Safety Leaderboard immediately. 

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Back to top button