Claude vs ChatGPT for working with your own notes

Filed under Comparison

By Gerald · 29 August 2026

Two chess knight pieces, one black and one white, facing each other on a board

Short answer: for reading and editing a connected notes and tasks workspace, Claude is the more careful writer and ChatGPT is the more casual one, and that difference matters more than either model's general benchmark score. Both support connectors well enough to use daily. Neither is "better at writing" in a way that survives past the next model update, so I am not going to pretend otherwise.

I connected both to the same Flow workspace, gave them the same set of real tasks, and watched what each one actually did rather than what it claimed it would do.

Why the general comparison is the wrong question

Every few months someone publishes a fresh "Claude vs ChatGPT" comparison. Every few months it is out of date before the next model ships. General intelligence rankings move fast. They mostly do not matter for a specific job: connecting an assistant to your own notes and tasks and letting it read, search, and occasionally write back into them.

That job depends on things that change far more slowly than benchmark scores. How well the model holds onto a long document without losing the thread. How reliably it calls a tool instead of guessing. How careful it is when it is about to overwrite something you wrote. Those are the questions worth answering, because the answer stays useful for longer than a leaderboard snapshot does.

Reading long notes without losing the thread

An open paper notebook resting beside a smartphone on a desk
The comparison that matters is what happens when the assistant can actually reach your notes, not a leaderboard score.

Context window size sets the ceiling for how much of your notes an assistant can hold at once. As of the current model lineups (checked against each vendor's own documentation in July 2026), Anthropic's flagship API models, including Claude Sonnet 5, support a 1 million token context window. OpenAI's GPT-5.5, which became the default ChatGPT model this year, also ships with a 1 million token context window in the API, up from 400,000 tokens on GPT-5.1.

On paper that is close to parity. In practice, a large context window is not the same as good long-document comprehension. A model can technically fit a hundred-page note inside its context and still lose track of an instruction buried on page forty. In my testing, both models handled a single long note (a multi-thousand-word project brief) accurately when asked direct questions about specific sections. The gap opened up with tasks that required synthesizing across several separate notes at once, where Claude was more consistent about citing which note a fact came from, and ChatGPT more often blended details from two notes into one answer without flagging that it had done so.

If you regularly ask your assistant to answer questions that span many notes rather than one, that distinction is worth testing yourself before you commit to a workflow.

Cost is a smaller factor here than it looks. Both companies price their flagship models similarly per million tokens, and a personal notes workflow rarely burns enough tokens in a day for the difference to show up on a bill. The decision worth making is about behavior, not price.

Calling tools when it should, and when it guesses

A connector is only useful if the model actually calls it instead of answering from memory or making something up. Both Claude and ChatGPT support the underlying protocol that makes this possible. Model Context Protocol, or MCP, is the open standard behind Flow's own connector, and both companies support it: Claude natively, and ChatGPT through what OpenAI now calls Apps (renamed from "connectors" in December 2025, though most people still search for the old name).

Tool-calling reliability is where I noticed the clearest difference. Asked "what's on my board due this week," Claude consistently called the list-tasks tool rather than answering generically. ChatGPT did the same most of the time, but occasionally answered from the conversation's earlier context instead of re-querying, which produced a stale answer if the board had changed since the last read. Neither failure was dramatic, but if you are relying on the assistant for a live view of your own data, a stale read is exactly the kind of small error that erodes trust over time.

Writing back into your notes without damage

This is the part I care about most, because reading your notes is low-risk and writing into them is not. A model that reads carelessly gives you a wrong answer. A model that writes carelessly can damage a note you actually need.

Both models, when asked to append a paragraph or update a task, generally did the right thing: they called the specific update tool rather than trying to rewrite the whole note from scratch. Claude was more consistent about preserving formatting, like existing checklists and headings, when appending new content. ChatGPT occasionally flattened formatting on append, turning a bulleted list into a plain paragraph in the process. That is a minor thing until it happens to a note you use daily, at which point it stops being minor.

Neither model deleted content it should not have, and both asked for confirmation before an action that looked destructive, which is the baseline you should expect from any connector-based workflow regardless of which model sits behind it.

Connector support on each side

Setup was roughly equivalent on both sides: paste a connector URL, authorize it, and the assistant gains access to a defined set of tools rather than open-ended access to your account. Claude's connector interface surfaces the available tools more explicitly before you use them, which I found reassuring the first time I connected a new workspace. ChatGPT's Apps interface is functionally similar but presents the tool list with slightly less detail up front.

Neither vendor's connector currently supports deleting data through the model, in Flow's case by design on Flow's side rather than either AI vendor's. That is a meaningful safety property regardless of which assistant you use: the worst a careless connector session can do is create clutter, not destroy something you cannot get back.

Which one I keep connected, and why

I keep Claude connected to my own Flow workspace day to day, mostly because of the formatting preservation on writes and the more consistent tool-calling on reads. Neither of those is a large gap, and either model would serve you fine if the other is the one you already pay for and prefer for everything else you do with it.

The model that writes back into your notes more carefully wins this comparison. Everything else is closer than the benchmarks suggest.

If you write far more than you read back through the assistant, the difference will matter less to you than it does to me, since I lean on both models for retrieval throughout the day.

Frequently asked questions

Which assistant is better at editing my existing notes? In my testing, Claude was more consistent about preserving existing formatting, like checklists and headings, when appending or editing content. ChatGPT occasionally flattened formatting during an append. Neither model deleted content without confirmation.

Do both Claude and ChatGPT support MCP connectors? Yes. Claude supports Model Context Protocol connectors natively. ChatGPT supports the same underlying protocol through what OpenAI now calls Apps, renamed from "connectors" in December 2025. Functionally, both let the model call a defined set of tools against an external service like Flow.

Which handles very long documents better? Both current flagship models support roughly 1 million token context windows as of mid-2026, per each vendor's own model documentation. On single long documents, both performed accurately. Claude was more consistent when a task required synthesizing facts across multiple separate notes at once.

Can I connect the same notes to both? Yes. Flow's MCP connector issues a token per account, and you can paste that same connector URL into both Claude and ChatGPT (or any other MCP-compatible client) at once. Nothing about the connector limits you to a single assistant.

Does either use my notes for training? Anthropic's consumer plans (Free, Pro, Max) use conversations for training by default unless you opt out in privacy settings, a policy that took effect in September 2025. Anthropic's business and API traffic are excluded from training. OpenAI similarly trains on individual ChatGPT conversations by default unless you opt out, while Business, Enterprise, and Edu accounts are excluded, and connector-sourced data is excluded from training by default for those account types. Check each vendor's current privacy settings directly, since these policies have shifted more than once.

Related reading

My verdict

Pick based on how the assistant writes back into your notes, not which one wins a general benchmark this month. Claude edges it on formatting care and tool-calling consistency today. Test it yourself against your own workspace before trusting either one with anything you cannot easily undo.

Own your notes and tasks

Flow puts notes, tasks, and a capture inbox in one place you pay for once, connected to your AI tools.

Create your account

Read this on flowproductivity.space · More from The Flow Journal · Try the Flow demo