I came to Claude Opus 4.7 about a month late. Anthropic shipped it on April 16, but I didn't get around to seriously testing it until I had a stretch of normal-week work to run it against.

To be honest, I wasn't expecting much. The version bump from 4.6 to 4.7 is small on paper, Anthropic's release notes are characteristically understated (“Claude Opus 4.7 launch,” and that's about it), and benchmark numbers rarely tell me anything about how a model will behave inside the kind of work I actually do.

But when I put four real tasks from my own day-to-day in front of it, the difference was clearer than I'd predicted. Three quick takeaways before the details:

• The tone calibration improved noticeably. For apology emails, meeting recaps, and internal briefs — anything that needs to land at the right emotional register — I'm rewriting roughly half as much as I did with 4.6.

• It stopped dropping numbers and proper nouns. Hand it a messy meeting transcript or a half-broken customer list, and the bookkeeping stays correct.

• When I ask “how should I think about this?” I now get reasoning and options together, not just an answer. It's become a more useful sparring partner for decisions.

Here's how I got to those conclusions.

The four tasks

I'm not an engineer. I run projects and write a lot of business communication. So I picked four jobs I'd genuinely outsource to Claude on a normal week, and gave it the same casual instructions I'd give a colleague: an apology email to a long-standing client, a summary of a 90-minute internal meeting, a three-year revenue projection, and a cleanup of a messy customer list. No prompt engineering.

1. The apology email

The brief was tricky: we'd missed a delivery deadline with a client we've worked with for years. I wanted the email to be clear about what happened, not evasive about responsibility, but also not so self-flagellating that it damaged the relationship.

This is exactly where 4.6 would technically succeed but feel slightly off — competent and corporate but distant, the kind of email a legal team writes. I'd end up rewriting it to sound more like a human.

4.7 landed it on the first try. What stood out most was a line near the close: “I'd welcome the chance to walk you through what happened in person next time we meet.” I hadn't asked for that. The model added it because it understood the actual goal wasn't just to apologize — it was to keep the relationship intact. That unprompted reading between the lines is what I noticed most in this upgrade.

2. The meeting summary

I fed it a transcript of a 90-minute meeting: six attendees, mixed topics (a pricing change, a product launch, hiring plans), more than ten distinct figures. I asked for five key points, action items grouped by owner, and a list of items deferred to the next meeting.

4.6 handled the basic structure but would occasionally drift — a statement attributed to the wrong person, or a price change quietly losing a digit. With 4.7, the things I checked (names, percentages, dates, monetary figures) all sat in the right place. I didn't find a single misattribution.

That sounds like a small win until you've been the person who has to re-read a transcript just to confirm whether the AI's summary can be trusted. “It can” is the upgrade.

3. The revenue projection

The setup was deliberately ordinary: year-one revenue of ¥48M, +18% growth in year two, +12% in year three, a 5% promotional discount in year three only, consumption tax at 10%, a 30% effective tax rate, and operating expenses at 55% of revenue.

The math was already within 4.6's comfort zone. What changed in 4.7 is what comes after the math. The model didn't just hand back a table — it offered, unprompted, an interpretation: “Year two shows strong top-line growth, but since operating expenses scale proportionally, the bottom-line improvement is smaller than the revenue line suggests.” That sentence is the work of turning a number into a story, and I'd normally write it myself. Getting it as part of the default output saves real editing time.

4. The customer list

I dropped in a CSV that looked like every real customer list I've ever inherited. Dates in three formats, some quantity fields blank or marked “N/A,” three near-identical product spellings, mixed-case regions, and one unit price typed in as the words “one fifty.”

4.6's approach was mechanical: “I dropped rows with missing quantities.” 4.7 inferred. For the “one fifty” cell it noted: “Other rows show the same product priced at 150, so I treated this as 150.” Every cleanup decision came with a one-line rationale — the difference between a data prep step you can defend later and one you can't.

What actually changed, in my own words

If I compress this into one line: Claude is now noticeably better at work that requires reading the situation, not just the prompt.

With 4.6, I'd often need two or three rounds of coaching — “a little warmer here,” “consider the reader's position” — before I got something I could send. With 4.7, that context-awareness shows up in the first draft most of the time. My reliance on prompt technique has dropped.

There's no flashy new capability to point at. What got better is the hardest part to demo: the confidence that what comes back is roughly what a thoughtful colleague would have written.

Where it sits against the alternatives

GPT-5.5 Instant remains the better choice for fast, casual asks — quick, included with ChatGPT, and no worse than Opus on short turns. Gemini 3 Pro leads on cross-modal work that touches audio or video. Perplexity is still the right answer for research where citations and source freshness matter. But for long-form writing, meeting work, financial narrative, and structured data with edge cases — the workload that fills most of a knowledge worker's week — Opus 4.7 is the most reliable option I've used.

A rough scorecard

Ten-point scale, four axes. Only the Opus 4.7 numbers come from my own testing this week; the others are anchored to published benchmarks and recent third-party reviews.

Accuracy — Opus 4.6: 8.25 / Opus 4.7: 9.50 / GPT-5.5: 8.00

Tone & nuance — Opus 4.6: 8.50 / Opus 4.7: 9.50 / GPT-5.5: 8.00

Speed & cost — Opus 4.6: 8.38 / Opus 4.7: 8.38 / GPT-5.5: 9.25

Clarity of explanation — Opus 4.6: 8.75 / Opus 4.7: 9.12 / GPT-5.5: 8.25

Total / 40 — Opus 4.6: 33.88 / Opus 4.7: 36.50 / GPT-5.5: 33.50

GPT-5.5 wins on speed and price. On the axes that affect the quality of business output, Opus 4.7 is a step ahead. The honest caveat: I did not run 4.6 or GPT-5.5 on these exact tasks this week, so their scores are anchored estimates, not measurements. Treat the table as directional.

Should you switch?

If you mostly use ChatGPT for casual questions, you don't need to change anything. If you already use Claude for serious writing, meeting work, or data review, moving to Opus 4.7 is worth doing — the time you spent prompting it back into shape comes straight back to you.

One operational note: per-token pricing is unchanged from 4.6 ($5 / $25 per million tokens), but the tokenizer was updated. For some content types the same input counts as 1.0–1.35x as many tokens. If you're a heavy user of long context, run a week or two of real usage and re-check your monthly bill before committing.

Sources: Anthropic release notes (docs.anthropic.com/en/release-notes/claude-apps), Introducing Claude Opus 4.7 (anthropic.com/news/claude-opus-4-7), Claude Opus 4.7 vs Opus 4.6 (mindstudio.ai).

Keep Reading