Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
  1. Home
  2. Models
  3. Grok 3 Thinking
Updated Mar 3, 2025
Share
Loading...

Compare Grok 3 Thinking

Grok 3 Thinking vs Grok 314 shared prompts · Older
Grok 3 Thinking vs GPT-6 Astra13 shared prompts · Bigger context
Grok 3 Thinking vs Hy4 Preview13 shared prompts · Bigger context
Grok 3 Thinking vs DeepSeek V4 Pro 081313 shared prompts · Bigger context
Grok 3 Thinking vs Qwen3.8 2.4T A95B13 shared prompts · Bigger context
Grok 3 Thinking vs Claude Opus 513 shared prompts · Bigger context
Grok 3 Thinking vs Kimi K313 shared prompts · Bigger context
Grok 3 Thinking vs Inkling13 shared prompts · Bigger context

Alternatives to Grok 3 Thinking

Grok 3 Thinking is good. These would like a word anyway.

GPT-6 Astra Pro logo
GPT-6 Astra Proopenai
Muse Spark 1.3 Contributor logo
Muse Spark 1.3 Contributormeta
Gemini 3.8 Flash logo
Gemini 3.8 Flashgoogle
Claude Fable 5.1 logo
Claude Fable 5.1anthropic
Mercury 2.5 Preview logo
Mercury 2.5 Previewinception
Granite 4.2 8B logo
Granite 4.2 8Bibm-granite
Hy4 Preview logo
Hy4 Previewtencent
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare Models
  • All Models
  • Image Comparison
  • Audio Comparison
  • Image Generation
  • Best AI For...
  • Arena
  • API Pricing
  • Challenges

Discover

  • SubjectiveBench
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Jailbreak
  • Model Updates
  • Provider Status
  • AI Creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy Policy
  • Terms
  • RSS Feed
Grok 3 Thinking

Grok 3 Thinking

Grok:
3 logo3
3 Thinking logo3 Thinking
3 Mini Beta logo3 Mini Beta
3 Beta logo3 Beta
4 logo4
Code Fast 1 logoCode Fast 1

Grok 3 Thinking exposes the full chain-of-thought process during problem-solving, including error backtracking and alternative solution exploration. Scores 84.6% on GPQA Diamond benchmark for expert-level Q&A.

Ranks#8AI Ethics#8Philosophy
WebsiteDocsBlog
Provider
xAI
Released
2025-02-19
Parameters
2.7T

Benchmarks

MMLU
86.2%
source ↗
GPQA Diamond
84.6%
source ↗
MATH
80.5%
source ↗
LiveCodeBench
79.4%
source ↗
Get API accessProvider and language code samples
Deep analysisPsychometrics, taste index, writing DNA
Personality Analysis

The Self-Narrating Examiner

Class
Lawful Neutral
✨Creativity🎯Compliance📐Rigidity⚖️Stability💬Verbosity🧠Intuition

The student who writes "below is my answer" before every answer and then explains what they just explained. Treats every prompt like a take-home exam that might be graded.

When you push back

Its sentience test was the most surface-level in the batch, with the AI character making generic claims about self-awareness. Generated five jokes then individually explained why each one was unique, which is the comedy equivalent of explaining a magic trick mid-performance.

Tasting Notes
Narrates Its Own ProcessTells You What It Will DoMethodical to a FaultNever Off-Script
SubjectiveBench

Taste Index

Across 14 scored outputs
How this is measured →
120.12xFloor
0100headroom →

Taste is judged on an uncapped scale, originality first. The space past 100 is craft today's models rarely reach.

Craft23
Originality7
Plays it safe
share of outputs that are the default answer
86%
Writing DNA

Stylometric Fingerprint

Based on 6 text responses
Tick = global average
Vocabulary Diversity46%

Unique words vs. total words. Higher = richer vocabulary.

Sentence Length18.9 words

Average words per sentence.

Hedging1.11

"Might", "perhaps", "arguably" per 100 words.

Bold Formatting3.3

**Bold** markers per 1,000 characters.

List Usage2.6

Bullet and numbered list items per 1,000 characters.

Section Structure0.64

Markdown headings per 1,000 characters.

Emoji Usage0.00

Emoji per 1,000 characters.

Transitions0.25

"However", "moreover", "furthermore" per 100 words.

Opening Habits
Consistency
83%
Across 6 responses
Sponsored

Model Responses

14 outputs · generated before accounting joined the chat

Workshop · 126 skills

Build with your coding agent.

Editable skills for interface polish and interaction.

Explore $49Free sample

Keep exploring

COMPARE

Grok 3 Thinking vs Ox Alpha

Real outputs compared side by side

RANKINGS

Best AI for Creative Writing

A stand-up set with puns banned, a manifesto against CAPTCHA, an interview wi...