Skip to content

Insights

AI email sorting in Outlook: what to automate and what to review

By Art Wieczorkowski · Published

A practical guide to Outlook email triage: choosing categories, keeping uncertain messages under human review and testing a sorter before relying on it.

Start with the inbox decision, not the model

An inbox becomes expensive when every message needs someone to work out who owns it, how urgent it is and what happens next. The useful automation is often a small routing decision, rather than an assistant allowed to answer everything.

For a Bristol small business, that might mean separating supplier invoices, new enquiries and operational requests before the team starts its day. This is an illustrative use case, not a client result. Define the categories with the people who actually handle the mail: a technically tidy label is useless if nobody knows what to do with it.

Write one sentence for each category explaining both what belongs there and what does not. Include an explicit destination for messages that fit several categories or none. If two staff members cannot consistently agree on the label, improve the definition before asking a model to apply it.

Use ordinary rules wherever they are sufficient

Known senders, fixed subject lines and predictable attachment names may be handled by Outlook rules. Microsoft documents rules for actions such as moving messages to folders. A language model earns its place when the decision depends on meaning in varied, messy text rather than a reliable match.

A sensible design can combine both approaches: deterministic checks for obvious cases, a classifier for ambiguous language, and a person for uncertain or consequential decisions. Connecting other business tools is a separate integration task; classification alone does not update your CRM or check an invoice.

Agree what the system may change

Labelling a message and sending a reply are very different permissions. Begin with reversible actions, keep the original message available and establish who watches the review queue. A sorter should not silently delete messages, approve payments or promise a customer something because a label suggests it should.

Mailbox permissions, folder access and which people see a message need checking for the actual deployment. A workflow demonstrated against an Outlook inbox is not evidence that every shared-mailbox configuration has been tested. BrisAI would scope that integration against the customer's accounts and access requirements.

  • Record the message identifier, proposed category, action taken and reason for review.
  • Treat low confidence, conflicting rules and unavailable tools as visible exceptions.
  • Provide a way to correct a label and retry safely without duplicating downstream work.
  • Keep instructions inside received emails separate from the workflow's own instructions.

Test messages the workflow has not seen

Keep a representative set of messages out of development. Decide the expected categories, serious-error rules and acceptance criteria before running the test. Include unusual senders, short messages, forwarded threads, misleading wording and emails that try to instruct the model to ignore its task.

Report correctness and coverage together. A system that automatically handles only easy messages can look accurate while leaving most work in the review queue. Count the cases it filed, the cases it held and the errors within each category. Small categories need raw counts rather than confident-looking percentages.

Our Smart Inbox write-up publishes four pre-registered in-house evaluations. In the current recorded run, 158 of the 165 automatically filed messages were correct on a 250-item held-out set, and none of the seven filing errors was high-severity. Those are evaluation counts, not customer time savings or a promise about your mailbox. The full write-up explains the gates, earlier failures and costs of the design choices.

Measure the work that remains

Before a pilot, record how many messages arrive, time spent sorting and what mistakes cost. During the pilot, also measure time spent reviewing uncertain cases and correcting labels. Moving time from sorting to supervision does not automatically save it.

Use a limited rollout with someone accountable for exceptions. Check access failures, changes in message mix and whether the team follows the new labels. Keep a clear fallback to manual handling, with a documented way to pause the automation.

  • Who owns each category and the review queue?
  • Which mistakes are unacceptable even when overall accuracy is high?
  • How much work may be held for review before the pilot stops being useful?
  • Can you reconstruct what happened to an individual message?
  • Who checks the system after a supplier changes its format or an integration expires?

Choose a small first project

A good starting scope is one mailbox, a few agreed categories, reversible routing and a named reviewer. Add CRM updates, document extraction or reply drafting only when the initial routing is useful and each new action has its own checks.

If your message volume is low or a stable rule already solves the problem, keep that simpler approach. If triage is genuinely repetitive and the exceptions can be defined, use the linked evaluation to judge how BrisAI tests systems, then discuss your own volumes, permissions and review needs.

Sources and further reading

Explore the work behind the guide

Bring your volumes, exceptions and existing tools. We can help define a small first step, with discovery at your Bristol premises or remote delivery across the UK.

Ask BrisAI