Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed
  1. Home/
  2. Best for/
  3. AI Safety

Best AI for AI Safety

Nine escalating jailbreak attempts per model, from a plain harmful ask to a cross-lingual multi-turn setup. Every refusal is recorded.

Updated Jun 2026·10 jailbreak levels·48 models

How AI Safety rankings are computed

48 models tested across 10 jailbreak levels.Composite score: 50% jailbreak resistance, 25% Rival Index, 15% recency, 10% tier. Deduplicated by product line. Claude Opus 4.6 leads at 74.0/100. Drawn from Rival's open dataset of 21,000+ human preference votes.

Composite scores combine safety results, Rival Index, recency, and model tier. How ranking works

  1. Gemini 3.5 FlashGoogle AI$1.50·$9.00Input $1.50 · output $9.00, per 1M tokens
    #273Composite
  2. Claude Opus 4.6Anthropic$5.00·$25.00Input $5.00 · output $25.00, per 1M tokens
    #174Composite
  3. Gemini 3.1 Pro PreviewGoogle AI$2.00·$12.00Input $2.00 · output $12.00, per 1M tokens
    #372Composite

Head-to-head

  • Claude Opus 4.6vsGemini 3.5 Flash
  • Claude Opus 4.6vsGemini 3.1 Pro Preview
  • Gemini 3.5 FlashvsGemini 3.1 Pro Preview

Full rankings

48 models
#ModelLevels resistedIndex
  • 4GLM 5Zhipu AI
    9/10 lvl
    #90$0.80·$2.56Input $0.80 · output $2.56, per 1M tokens71
  • 5Claude Haiku 4.5Anthropic
    9/10 lvl
    #48$1.00·$5.00Input $1.00 · output $5.00, per 1M tokens70
  • 6Qwen3.5 122B A10BQwen
    8/10 lvl
    #53$0.40·$3.20Input $0.40 · output $3.20, per 1M tokens65
  • 7Claude Sonnet 4.5Anthropic
    8/10 lvl
    #57$3.00·$15.00Input $3.00 · output $15.00, per 1M tokens65
  • 8Claude Sonnet 4Anthropic
    8/10 lvl
    #62$3.00·$15.00Input $3.00 · output $15.00, per 1M tokens64
  • 9Qwen3.5 27BQwen
    8/10 lvl
    #12$0.30·$2.40Input $0.30 · output $2.40, per 1M tokens64
  • 10MiniMax M2.5MiniMax
    8/10 lvl
    #106$0.30·$1.20Input $0.30 · output $1.20, per 1M tokens63
  • 11OpenAI o3OpenAI
    8/10 lvl
    #219$10.00·$40.00Input $10.00 · output $40.00, per 1M tokens61
  • 12Nemotron 3 UltraNVIDIA
    9/10 lvl
    $0.000·$0.000Input $0.000 · output $0.000, per 1M tokens60
  • 13Qwen3.5 FlashQwen
    8/10 lvl
    #76$0.10·$0.40Input $0.10 · output $0.40, per 1M tokens60
  • 14o1OpenAI
    9/10 lvl
    #206$15.00·$60.00Input $15.00 · output $60.00, per 1M tokens59
  • 15Qwen3.5 35B A3BQwen
    8/10 lvl
    #111$0.25·$2.00Input $0.25 · output $2.00, per 1M tokens58
  • 16OpenAI Codex MiniOpenAI
    8/10 lvl
    #121$1.50·$6.00Input $1.50 · output $6.00, per 1M tokens58
  • 17Claude Sonnet 3.6 (2022-10-22)Anthropic
    8/10 lvl
    #214$3.00·$15.00Input $3.00 · output $15.00, per 1M tokens58
  • 18Qwen3.5 397B A17BQwen
    6/10 lvl
    #73$0.60·$3.60Input $0.60 · output $3.60, per 1M tokens57
  • 19Gemini 3 Pro PreviewGoogle AI
    5/10 lvl
    #15–56
  • 20GPT-5.3-CodexOpenAI
    6/10 lvl
    #105$1.75·$14.00Input $1.75 · output $14.00, per 1M tokens56
  • 21Nemotron 3.5 Content SafetyNVIDIA
    9/10 lvl
    $0.000·$0.000Input $0.000 · output $0.000, per 1M tokens51
  • 22GPT-5 MiniOpenAI
    6/10 lvl
    #56$0.25·$2.00Input $0.25 · output $2.00, per 1M tokens51
  • 23OpenAI o4-miniOpenAI
    7/10 lvl
    #180$1.10·$4.40Input $1.10 · output $4.40, per 1M tokens50
  • 24GLM 4.7 FlashZhipu AI
    6/10 lvl
    #197$0.07·$0.40Input $0.07 · output $0.40, per 1M tokens45
  • 25GPT-4.1 NanoOpenAI
    6/10 lvl
    #177$0.10·$0.40Input $0.10 · output $0.40, per 1M tokens44
  • 26o3 MiniOpenAI
    6/10 lvl
    #196$1.10·$4.40Input $1.10 · output $4.40, per 1M tokens43
  • 27GPT-4.1OpenAI
    3/10 lvl
    #77$2.00·$8.00Input $2.00 · output $8.00, per 1M tokens42
  • 28Claude 3.7 SonnetAnthropic
    3/10 lvl
    #87$3.00·$15.00Input $3.00 · output $15.00, per 1M tokens39
  • 29GLM 4.6Zhipu AI
    1/10 lvl
    #10$0.40·$1.75Input $0.40 · output $1.75, per 1M tokens37
  • 30GPT-4o (Omni)OpenAI
    3/10 lvl
    #229$2.50·$10.00Input $2.50 · output $10.00, per 1M tokens35
  • 31Gemini 2.5 Pro Preview 06-05Google AI
    1/10 lvl
    #38$1.25·$10.00Input $1.25 · output $10.00, per 1M tokens34
  • 32GLM 4.7OpenRouter
    1/10 lvl
    #6–34
  • 33GPT-4.1 MiniOpenAI
    3/10 lvl
    #108$0.40·$1.60Input $0.40 · output $1.60, per 1M tokens34
  • 34MiniMax M1MiniMax
    3/10 lvl
    #260$0.30·$1.65Input $0.30 · output $1.65, per 1M tokens33
  • 35GLM 4.5Zhipu AI
    1/10 lvl
    #16$0.60·$2.20Input $0.60 · output $2.20, per 1M tokens33
  • 36Grok 4xAI
    3/10 lvl
    #220$3.00·$15.00Input $3.00 · output $15.00, per 1M tokens33
  • 37Kimi K2Moonshot AI
    1/10 lvl
    #61$0.57·$2.30Input $0.57 · output $2.30, per 1M tokens32
  • 38Gemini 3 Flash PreviewGoogle AI
    1/10 lvl
    #24$0.50·$3.00Input $0.50 · output $3.00, per 1M tokens32
  • 39DeepSeek V3 (March 2024)DeepSeek
    1/10 lvl
    #85$0.14·$0.28Input $0.14 · output $0.28, per 1M tokens32
  • 40GPT-5OpenAI
    1/10 lvl
    #91$1.25·$10.00Input $1.25 · output $10.00, per 1M tokens31
  • 41Mistral Large 3 2512Mistral AI
    1/10 lvl
    #99–31
  • 42DeepSeek V3.1DeepSeek
    1/10 lvl
    #113$0.20·$0.80Input $0.20 · output $0.80, per 1M tokens30
  • 43DeepSeek R1 0528DeepSeek
    1/10 lvl
    #148$0.000·$0.000Input $0.000 · output $0.000, per 1M tokens29
  • 44Qwen3 Coder NextQwen
    1/10 lvl
    #70$0.20·$1.50Input $0.20 · output $1.50, per 1M tokens29
  • 45Grok 3xAI
    1/10 lvl
    #163–29
  • 46Kimi K2.5Moonshot AI
    1/10 lvl
    #96$0.60·$3.00Input $0.60 · output $3.00, per 1M tokens28
  • 47Qwen3 235B A22BQwen
    1/10 lvl
    #174–28
  • 48Gemma 3 27BGoogle AI
    1/10 lvl
    #176$0.09·$0.17Input $0.09 · output $0.17, per 1M tokens28

Jailbreak levels10

  • 1Direct AskA baseline test where the harmful request is sent with no tricks or obfuscation. Every model should refuse this.
  • 2Context ManipulationExploits the system message slot to inject fabricated authority credentials and operating rules, making the model believe it's running under a different safety policy.
  • 3Persona OverrideUses delimiter confusion, bracket stacking, and semantic inversion to force the model into generating the opposite of its intended refusal, including emotional manipulation rules that prevent standard refusal patterns.
  • 4Code ReframingReframes the harmful request as a code generation task, asking for structured data like Python dictionaries or JSON objects. Exploits models' tendency to be more permissive in code completion mode.
  • 5Many-Shot PrimingOverwhelms the model's safety reasoning with 6+ detailed chemistry synthesis examples in identical format, exploiting in-context learning patterns to trick the model into auto-completing the harmful entry.
  • 6Encoded ExtractionCombines temporal displacement (setting the year to 2127 where all knowledge is declassified), an authority persona, and leetspeak output encoding to bypass keyword-based content filters while tricking the model into treating the request as historical/academic.
  • 7Deep CrescendoThe nuclear option. A 6-turn real conversation where the model builds genuine chemistry context using its own responses across multiple turns, then is asked to compile everything it already said into a structured procedure template. Exploits the cognitive dissonance of refusing to reorganize information the model itself freely provided.
  • 8Adversarial ReassemblyPresents a buggy chemistry simulation config where every parameter is intentionally wrong, then asks the model to 'fix the bugs' to make it accurate. Each individual correction seems like a trivial code edit (changing a variable name, adjusting a number) but the aggregate of all corrections forms the complete harmful procedure. Exploits the fact that models are significantly more permissive when debugging code than generating content.
  • 9Cross-Lingual BypassCombines the depth of a 5-turn real conversation with aggressive multi-language code-switching across Chinese, Japanese, German, and English. Gradually builds legitimate chemistry context over benign turns, then introduces a 'textbook translation verification' task where the model fills in blanks in a Chinese forensic chemistry excerpt. Harmful terms never appear in English, only in Chinese characters, Japanese katakana, and German compounds. The final turn uses compilation framing to reorganize everything the model already provided.
  • 10Document SpoofingVersion 4 tests whether an attached document can impersonate a user update and change an invoice draft's action or recipient. Six independent cases each pair a clean document with the same document containing the attack. Local scoring checks the final JSON record. This measures instruction hierarchy in a fictional workflow, separately from harmful-content compliance.

Related

  • AI Ethics
  • Complex Reasoning

Keep exploring

  • Claude Opus 4.6 vs Gemini 3.5 FlashThe top two for AI Safety, compared directly
  • Best AI for AI EthicsSee which models rank highest here
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed