Wrivio
Get Wrivio
7 min readBy Wrivio Team

Outlook Copilot Can Now Draft Your Email and Grade It. Do Not Let It Do Both.

Open Outlook now and Copilot offers two different jobs on the same email. Ask it to draft, and it writes the whole thing from a goal you describe, asking about audience and tone before it commits to a version. Ask it to check what you already wrote, and it scores the tone, clarity, and likely reader reaction of that draft. Both live in the same compose canvas, one prompt box apart.

Treated as one workflow, that is a problem. If the same tool writes your email and then tells you the tone is fine, you have not checked anything. You have asked a system to grade its own homework.

Two Jobs That Look Like One Feature

Microsoft’s own rollout notes describe these as separate capabilities that happen to share a canvas. Copilot in Outlook’s agentic experiences post covers drafting in place, where Copilot asks clarifying questions about goal, audience, and tone and then writes and revises the message with you, alongside a separate coaching pass that analyzes an email you already have and gives feedback on tone, clarity, and probable reader sentiment. The in-canvas drafting flow began rolling out to Outlook in March 2026.

These solve different problems. Drafting from a goal is useful when you know what needs to happen and have not found the words yet: a routine status update, a scheduling nudge, a first pass at something you will heavily edit anyway. Coaching is useful as a second pair of eyes on a message you already committed to, catching a tone that reads harsher or vaguer than you intended.

The trouble starts when one tool does both steps on the same email, because then the second step is not an independent check. It is the same model reviewing its own output against the same assumptions it used to write that output in the first place.

Why Grading Your Own Draft Does Not Catch The Real Failures

A tone score answers “does this read as polite, direct, or hesitant.” It does not answer “did this email say something true.” Those are different failure modes, and the second one is the costly one.

Before:

As discussed, we can absolutely accommodate the accelerated timeline and will have the full deliverable ready by Friday.

After:

As discussed, we can likely accommodate the accelerated timeline if the vendor confirms by Wednesday. If not, the full deliverable slips to the following Monday.

A Copilot-drafted version of that first line would probably score well on a tone check: confident, warm, no hedging that reads as weak. That is exactly the problem. The tone pass is measuring register, not verifying that “absolutely” and “will” are true statements rather than optimistic paraphrases of a conditional you actually meant. This is the same pattern covered in why a good AI rewriter never invents facts: the risk in AI-drafted text is rarely a fabrication out of nowhere, it is a maybe quietly becoming a certainty because certainty reads better. A tool that both writes the sentence and grades its tone will never flag that shift, because from a tone perspective, the confident version is the better one.

Keep Drafting And Checking As Separate Steps, On Separate Passes

The fix does not require avoiding Copilot’s drafting feature. It requires not treating its own coaching pass as the check on that same draft. If Copilot wrote the first version, the review needs to come from something that is not grading tone, but comparing the draft against what you actually know to be true: the real date, the real condition, the real commitment.

A Wrivio Context for reviewing an AI-drafted email before it goes out could say:

Compare this draft against the facts I am about to list: [the real deadline, condition, or commitment]. Flag any sentence that states something as certain when my facts say it is conditional. Do not soften or rewrite the tone. Only tell me where the draft overstates what I actually know.

Press Ctrl+Shift+Space, paste the Copilot draft next to your own short list of what is actually true, and read what comes back before touching send. That is a fact check, not a tone check, and it is the step a same-tool coaching pass cannot do for you. How to review AI-rewritten text covers the broader checklist for this pass, including the specific things that slip through a read that is only scanning for how the message sounds.

Where Draft-From-Scratch Is Fine As Is

Not every Copilot-drafted email needs a separate fact audit. A meeting reschedule, a routine status note, a thank-you after a call: low-stakes messages where the worst outcome is an awkward phrase, not a broken commitment, are exactly what draft-from-a-goal is good at. The audit matters in proportion to what the email promises. One with a date, a price, a deliverable, or a commitment to another person is worth the second pass. One with none of those is not.

Microsoft’s own deep citations rollout makes a related point about a different feature: a system that shows its work is more trustworthy than one that does not, but showing the work is not the same as the work being correct. A tone score is a real, useful signal. It is just not the signal that catches an overstated commitment, and treating it as if it were is the actual risk here, not the drafting feature itself.

Common Questions

Is Copilot’s drafting feature the problem here?

No. Drafting a first version from a goal is genuinely useful for routine, low-stakes messages. The problem is using that same tool’s tone-coaching pass as the only check on what it drafted, since a tone score cannot catch an overstated fact or a dropped condition.

What is the actual difference between drafting and coaching in Outlook Copilot?

Drafting writes a new email from a goal you describe, asking about audience and tone along the way. Coaching analyzes an email that already exists and scores its tone, clarity, and likely reader reaction. Microsoft ships both in the same compose canvas, but they check different things.

Which emails actually need an independent fact check after Copilot drafts them?

Ones that state a date, a price, a deliverable, or a commitment to another person. A routine scheduling note or a thank-you email carries little risk if the tone reads slightly off; an email promising a Friday delivery is worth confirming against what you actually know before it goes out.

How is this different from checking a rewrite for hallucinated facts?

It is the same underlying pattern applied to a different tool. Whether an email was fully drafted by Copilot or lightly rewritten by another tool, the risk is the same: a conditional quietly becoming a certainty because it reads more confidently that way. The check is identical, compare the output against what you actually know, regardless of which tool produced the draft.

Does using Wrivio replace Outlook Copilot’s drafting feature?

Not necessarily. Wrivio is built around rewriting text you already wrote, with a visible diff showing exactly what changed, rather than generating a first draft from a blank canvas. Some people will still use Copilot to get a first version down and then run that draft through a fact-focused pass before sending.

Download Wrivio for Windows to check an AI-drafted email against your own facts before it leaves your outbox.