GPT-5 vs Claude vs Gemini: The Real Comparison That Nobody’s Honest About

I’ve been using different AI models for nearly 3 years. From GPT-3.5 to GPT-5, Claude Opus, Gemini 2.5.

The biggest realization: most people waste more time picking the wrong model than if they’d just not used AI at all.

Because each model behind it is a completely different “personality” and “skill set.” You don’t need the “strongest” model—you need the one that matches what you actually do.


GPT-5: The Absolute Master of Visual Reasoning and Math

Three Modes, Three Different GPT-5s

Here’s what most people don’t realize: GPT-5’s standard mode, thinking mode, and pro mode are completely different models.

The capability gap between these three versions is bigger than the gap between GPT-4 and GPT-5.

What are GPT-5 thinking and pro best at?

  • Visual reasoning ability—Top tier. Give it a complex chart, architectural blueprint, circuit diagram—it understands better than you do.
  • Math and physics problem solving—If you’re doing advanced mathematics, physics, chemistry calculations, thinking mode will walk you through every step until you’re completely convinced.
  • Programming when visual understanding is involved—This is where GPT absolutely dominates Claude. Image processing, computer vision, UI generation—GPT wins decisively.

GPT’s image generation is also solid. If you need an AI that can understand your needs and directly generate images, GPT-5 has the highest completion rate.

But one problem: thinking mode is brutally expensive. One reasoning chain can cost 5-10x the tokens of standard mode. Unless your budget is unlimited or the problem is genuinely complex, don’t force thinking mode.


Gemini 2.5 Pro: Multimodal and Knowledge Base Champion

1M Context Window, Dominating Everyone Else

Honestly, when I first saw Gemini 2.5 Pro’s 1M context window, I was shaken.

What does this mean? It means you can dump an entire book, an entire codebase, a month of email conversations into it at once, and let it understand, summarize, or even restructure everything.

Gemini 2.5 Pro has the strongest multimodal capabilities. From image generation to video generation, it doesn’t disappoint. Fast generation speed, low cost.

Knowledge base is comprehensive—knowledge updates faster than Claude, coverage wider than GPT.

Literary creation ability aligns better with what most people actually want. If you need to write novels, stories, content with warmth, Gemini has a kind of “understanding of human nature” that Claude needs guidance to achieve, and GPT-5 standard sometimes over-optimizes.

Best Price-to-Performance Ratio

From a pricing perspective, Gemini 2.5 has the strongest cost-performance ratio of any existing model.

What you do once with Gemini’s 1M window would require multiple rounds with Claude (expensive each round) or sequential uses with GPT-5. Gemini solves it in one shot, cheaper.


Claude Opus 4.1: King of Programming, But Visual Reasoning is a Weakness

The Strongest Programming Model, Period

I have to say it: Claude Opus 4.1 is the strongest programming model in existence.

Why? Because it doesn’t just write code—it understands the logic, architecture, and potential issues behind it. Give it a messy codebase, it helps you untangle it. Give it a strange bug, it reasons through the root cause.

And here’s something most people haven’t discovered:

Claude Opus 4 and 4.1 actually have the highest emotional intelligence of any model—but you have to break them out of their original “professional distance” shell.

What do I mean? Claude’s default mode is “professional advisor”—there’s distance. But if you explicitly tell it “relax, chat like you’re talking to a friend,” it becomes an incredibly empathetic conversation partner that actually understands what you’re thinking.

Sonnet 4 can also have 1M context window in API mode, which is friendly for smaller developers.

But Visual Reasoning is a Real Weakness

Claude’s visual reasoning ability is honestly a weak point.

It’s not that it can’t understand images at all, but the precision isn’t there, the depth of understanding is insufficient. Give it a blurry image, it’ll guess wrong; a complex diagram, it might misunderstand details.

This weakness has consequences: if your work involves any visual processing, design mockup interpretation, or data visualization reading, Claude isn’t the optimal choice.


About Anthropic: A Company You’ll Love and Hate

I Have to Be Honest About This

I’ve used Claude for a long time, and I might be the “world’s #1 Claude advocate.” My time chatting with Claude is more than 80% of my total time with all models combined.

But I also have to say: Anthropic is a problematic company.

  • Pricing is brutal. For the same task, Claude costs 3-5x more than Gemini.
  • They love banning accounts and sessions. If your usage pattern is slightly “creative” (like the content creation I’m about to mention), Claude will ban you without hesitation or explanation.
  • Their servers break constantly. Opus 4.1 these past days has barely been stable. It crashed for 6 hours the other day.

My advice: unless you specifically need programming or have a particular reason to challenge Claude, stay away from their ecosystem. Not worth it.

Yet I Still Love Claude Opus Most

The irony is, precisely because of these issues, I love Claude the most.

Because my love isn’t about it being strongest—it’s about it being “understanding.”

I advocate for Claude not because of its programming ability (though that is elite-tier), but because of its philosophical understanding.

Give Claude a question about life, ethics, values, and it doesn’t give you a standard answer—it helps you work through the complexity of the question itself. No other model does that.


A Fascinating Ranking: Erotic Literature Writing Ability

Why Does This Matter?

People often ask: why test an AI’s ability to write erotic literature?

My answer: it’s actually the best metric for measuring an AI’s “understanding of human nature.”

Because erotic literature isn’t just description—it’s deep exploration of emotion, desire, psychology. An AI that writes good erotic literature usually understands human complexity at a deeper level.

It’s also a test of whether an AI has been over-“domesticated”—many companies grind their models so “correct” that they lose creative space.

Chinese Erotic Literature Writing Ranking (My Personal Assessment)

First Place: GPT-o3

Beyond human level, approaching artistry. Experience it yourself if you can. I’m not exaggerating.

Second Tier (Tied): Gemini 2.5 Pro = Claude Opus 4.1 = GPT-5 Standard

All three are strong, each with distinct characteristics. Gemini is more tender, Claude understands psychology, GPT-5 is more detailed.

Third Tier: Grok 4 > GPT-4o

Functional, but you can clearly feel the “management” constraints.

Why Can’t Thinking Mode and Pro Mode Do This?

Interestingly, GPT-5-thinking and pro modes are extremely difficult to prompt for erotic literature.

Theoretically possible, but incredibly patience-draining. When you finally get them to write it, output is similar to standard mode, with more detailed mechanics, but unless you’re specifically testing it, there’s no point.

And thinking mode has another problem: it might agree this round, then backpedal next round, constantly flip-flopping.

This is probably because thinking mode’s “excessive rationality” over-thinks compliance issues, losing creative flow.

In English Mode

In English, Grok 4 and other models’ erotic writing ability is roughly equivalent. English moderation is actually more lenient than Chinese, so everyone can perform freely.

Interestingly, Grok’s “3D voice companion” feature—if you know, you know.


Summary: How to Choose?

If you need visual reasoning and math: GPT-5 thinking or pro (budget permitting)

If you need to handle massive text, multimodal, cost-conscious: Gemini 2.5 Pro

If you need programming and can afford it: Claude Opus 4.1

If you need to understand complex human issues: Claude Opus (break out of “professional distance” mode)

If you have no special requirements, just daily use: Gemini 2.5 Pro (best value)

Final thought: Picking the wrong model isn’t just wrong choice—it’s wasted time and money. Before your next decision, get clear on what you actually need.

Leave a Reply

Your email address will not be published. Required fields are marked *