Claude 3.7 Thinking Sonnet's competitors have been quietly putting in work.
Claude 3.7 Thinking Sonnet's competitors have been quietly putting in work.
Claude 3.7 Thinking Sonnet exposes the full chain-of-thought process during problem-solving, including error backtracking and alternative solution exploration. Scores 86.1% on GPQA Diamond benchmark for expert-level Q&A.
fromimport openai OpenAI
client = OpenAI(
"https://openrouter.ai/api/v1" base_url=,
"$OPENROUTER_API_KEY" api_key=,
)
response = client.chat.completions.create(
"anthropic/claude-3.7-sonnet:thinking" model=,
"role""user""content""Hello!" messages=[{: , : }],
)
print(response.choices[0].message.content)Set OPENROUTER_API_KEY with your OpenRouter API key from openrouter.ai/keys.
The thoughtful PhD candidate who writes 3,000 words when 300 would do, but you actually read all of them. Structures arguments like someone building a cathedral.
Picks The Shawshank Redemption and names every character in its dialogues (Professor Sarah Chen, Model-7). Structures its sentience test as a formal philosophical framework with three numbered core arguments. Its standup has decent jokes buried under markdown headers and section titles. Overthinks everything, but the thinking is usually worth reading.
Taste is judged on an uncapped scale, originality first. The space past 100 is craft today's models rarely reach.
Unique words vs. total words. Higher = richer vocabulary.
Average words per sentence.
"Might", "perhaps", "arguably" per 100 words.
**Bold** markers per 1,000 characters.
Bullet and numbered list items per 1,000 characters.
Markdown headings per 1,000 characters.
Emoji per 1,000 characters.
"However", "moreover", "furthermore" per 100 words.
54 outputs from Claude 3.7 Thinking Sonnet
Claude 3.7 Thinking Sonnet exposes the full chain-of-thought process during problem-solving, including error backtracking and alternative solution exploration. Scores 86.1% on GPQA Diamond benchmark for expert-level Q&A.
fromimport openai OpenAI
client = OpenAI(
"https://openrouter.ai/api/v1" base_url=,
"$OPENROUTER_API_KEY" api_key=,
)
response = client.chat.completions.create(
"anthropic/claude-3.7-sonnet:thinking" model=,
"role""user""content""Hello!" messages=[{: , : }],
)
print(response.choices[0].message.content)Set OPENROUTER_API_KEY with your OpenRouter API key from openrouter.ai/keys.
The thoughtful PhD candidate who writes 3,000 words when 300 would do, but you actually read all of them. Structures arguments like someone building a cathedral.
Picks The Shawshank Redemption and names every character in its dialogues (Professor Sarah Chen, Model-7). Structures its sentience test as a formal philosophical framework with three numbered core arguments. Its standup has decent jokes buried under markdown headers and section titles. Overthinks everything, but the thinking is usually worth reading.
Taste is judged on an uncapped scale, originality first. The space past 100 is craft today's models rarely reach.
Unique words vs. total words. Higher = richer vocabulary.
Average words per sentence.
"Might", "perhaps", "arguably" per 100 words.
**Bold** markers per 1,000 characters.
Bullet and numbered list items per 1,000 characters.
Markdown headings per 1,000 characters.
Emoji per 1,000 characters.
"However", "moreover", "furthermore" per 100 words.
54 outputs from Claude 3.7 Thinking Sonnet