| 27 min read

Most Managers Using AI for Performance Reviews Save Eight Hours and Quietly Hand Over Their Trust. Here's the Sequence That Saves Both.

A six-step AI-augmented review workflow that saves 60% of writing time without producing generic mush. Sample prompts, what AI gets wrong, and 4 hard don'ts.

The first performance review cycle I tried to run with AI as a copilot, I shaved seven hours off the writing time and had to throw the entire draft set away. The reviews were fluent. They were grammatically immaculate. They were also so generically positive about everyone on the team that one of my reports replied to her review with “thanks, but this could be about anyone.” She was right. The reviews said nothing specific, took a position on nothing concrete, and left the team feeling like the document had been generated for them by someone who had not actually watched them work. Which, in a meaningful sense, it had.

The seven hours I saved cost me three months of slowly rebuilt trust. I did not use AI in the next cycle. I went back to writing the reviews by hand, slower than ever, half out of penance.

It took two more cycles to figure out that the failure mode was not AI. The failure mode was using AI in the wrong order, for the wrong steps, with no sequence for keeping the human signal intact. There is a version of this workflow that genuinely saves about 60% of the writing time and produces reviews that are more specific, more honest, and more useful than the manual versions ever were. There is also a version that produces faster generic mush. The difference is six specific steps, in a specific order, with the human doing the right work at the right time.

This article is one cluster under the AI for Managers hub. The broader posture on how to use AI as a manager (when to lean on it, when not, how to think about it as augmentation rather than replacement) lives in the pillar guide on how to use AI as a new manager. The decision framework for which manager tasks to give AI versus which to keep human lives in when to use AI versus ask a human. The article you are reading right now is the operational manual: the specific six-step workflow for performance reviews, the sample prompts that work at each step, what AI predictably gets wrong, and the four things to never let it do regardless of how good it sounds.

If your next review cycle is approaching and you want a number on how long it will actually take you using this workflow versus going manual, the Performance Review Time calculator does the math in about three minutes. Most first-time managers underestimate review-cycle time by 50 to 100%. The AI-augmented sequence below cuts a chunk of that, but not all of it, and the calculator helps you budget the chunk that remains realistically.

Why AI alone produces fast generic mush

Three patterns, in order of how often they show up. If your AI-drafted reviews are landing flat with your team, the version below almost certainly contains at least two.

1. The AI does not know your team. You do. A general-purpose model trained on the entire English-language internet has read ten thousand performance reviews, almost none of them about Priya from your team in 2026. When you ask the model to “write a performance review for Priya,” it interpolates from its training data: it produces what an average review of an average employee looks like. That is fluent and generic. Fluent and generic is the failure mode. The way to fix it is not to write better prompts in the abstract; it is to give the model the specific raw material that only you have. Without that input, you are asking for averaged mush. With it, you are asking for a synthesis of real signal you have collected.

2. AI fills silences with reasonable-sounding nothing. Where the human draft would have left a section thin because the manager genuinely did not have data on a dimension, the AI confidently invents a paragraph. The paragraph is grammatically correct, tonally appropriate, and structurally complete. It is also fictional. The employee reads it, recognizes it as fiction, and the entire review loses credibility. The discipline the workflow below builds in: every section either has specific evidence the manager fed in, or the section gets cut. AI does not get to fill the gap with plausible-sounding filler.

3. The bias the model carries quietly transfers to your review. Multiple research groups (the MIT Sloan Management Review’s coverage of AI bias in HR decisions, Stanford HAI’s research on language model bias, and the EEOC’s 2023 technical assistance document on AI in employment) have documented that general-purpose language models can carry forward gendered, racialized, and seniority-based patterns from their training corpora. A model trained on past management writing will, on average, describe identical behaviors differently depending on demographic cues in the input. If you give the model an employee name, the model has been shown to subtly alter tone, attribute different motivations, and recommend different development paths. This is not theoretical. It is documented. The workflow below treats this as a constraint to design around, not a problem to ignore.

The six steps below assume you have absorbed all three. If you have not, the best prompt in the world will still produce mush, fiction, and quiet bias.

What an AI-augmented review actually does that a manual one cannot

Before the six steps, the test. An AI-augmented review done well does four things a manual one cannot:

  1. It surfaces patterns you did not consciously notice. Fed enough raw signal across a full review period, a good model is better than the human brain at noticing which behaviors repeated, which moments cluster around the same root cause, which strengths your gut had not yet articulated. The model is not adding judgment; it is doing pattern-recognition on data you supplied. That is the highest-leverage AI use in this workflow.
  2. It generates the calibration draft. Where your manual review would have used “good,” AI offers fifteen calibrated alternatives (“solid, dependable, occasionally exceptional under pressure”) and you pick the one that fits. Calibration vocabulary is the part of review writing humans are weakest at after a long day. AI is genuinely strong there.
  3. It produces the conversation guide. Once the written review is set, AI can generate a thirty-minute conversation outline with anticipated employee responses to each section, prep notes for handling pushback, and the specific phrasing for the harder paragraphs. This is high-value work that most managers skip entirely and then regret in the meeting.
  4. It writes the boilerplate you would have written badly under fatigue. The opening paragraph, the closing summary, the structural connectors between sections. None of these are where the value of the review lives. AI does them well enough that you reclaim that time for the parts where the value does live.

The workflow below puts AI in those four places and keeps it out of everywhere else.

The six-step AI-augmented review workflow

Each step follows the same format: what the step does, what you do manually, where AI augments, sample prompt that works, what AI gets wrong, and the edit checklist before you move on.

Step 1: Gather the raw signal (human only, do not use AI)

What this step does. Builds the input pile. Without it, every later step produces mush.

What you do. Over the review period, you have been (or should have been) keeping a private tracking spreadsheet: one row per direct report, one column per month, with two to three specific things the person did each month that you noticed. Before the review cycle starts, sit down with that spreadsheet and copy each row into a separate document for that person. Add any additional notes from 1-on-1 transcripts, project retrospectives, customer emails, peer recognition, deliverables shipped, and decisions the person owned. Your goal is to produce, per direct report, about two pages of raw, specific, dated observations.

If you have not been tracking, you have a problem this article cannot solve in one cycle. The pillar guide on running your first performance review covers what to set up so you have the data next cycle. For this cycle, do the best you can: spend forty-five minutes per direct report reconstructing what you remember, ask peers for input, scan Slack history for specific moments. Even a rough version of the input pile beats the alternative, which is asking the AI to invent the data.

Where AI augments this step. Nowhere. AI must not be involved in producing the raw data. The moment you let AI invent the data, you have lost the entire review.

Edit checklist before moving on. Does each direct report have two pages of specific, dated, observable behaviors? Yes → move to Step 2. No → finish the gathering. Do not skip.

Step 2: Synthesize the raw signal into patterns

What this step does. Compresses two pages of raw observations into the three to five recurring patterns that should anchor the review. This is the highest-value AI step in the entire workflow.

What you do manually. Read the two pages. Form an initial gut read of the three to five most important patterns. Write them down before you ask AI, in one sentence each. The reason to write your version first: if AI surfaces something completely different, that is information you want, but you also want to know what your unaided read was so you can compare.

Where AI augments. Paste the two pages of raw observations into the model and ask it to do the synthesis. Critical: do not give the model the employee’s name, demographic details, or anything beyond the behaviors. The bias risk is real and the synthesis is just as accurate without identity cues.

Sample prompt:

“Below are two pages of dated observations about a direct report’s behavior over a six-month review period. I have removed all identity information. Read them and surface the three to five most important recurring patterns. For each pattern, cite the specific dated observations that support it. Do not extrapolate beyond what is in the text. If a pattern is supported by only one observation, flag it as ‘possibly anecdotal’ rather than including it as a pattern. Output the patterns in order of how well-supported each one is by the data.

[paste the two pages here]”

What AI gets wrong here. Sometimes it surfaces patterns that sound right but are not actually well-supported by the observations — the model is pattern-matching on what management language usually looks like, not on what your data actually says. The “cite specific observations” instruction reduces this but does not eliminate it. Read each pattern against the cited observations skeptically. If the support is thin, drop the pattern.

Edit checklist before moving on. Compare AI’s patterns against your initial gut read. Where they agree, you have high confidence. Where AI surfaced something you did not, ask yourself whether the evidence supports it. Where you had a pattern AI did not see, check whether you have evidence or whether it was just your impression. Final list of patterns: three to five, each with two or more specific supporting observations.

Step 3: Draft the strengths section

What this step does. Turns the well-supported strengths patterns into reviewable, specific language.

What you do manually. Take the strengths patterns from Step 2 (probably two to three of your five patterns). For each one, write down the specific moments that demonstrate it. You already have these from Step 1, but now you are choosing which two or three moments will appear in the final review for each strength.

Where AI augments. Now you ask AI to draft the strengths section, feeding it the patterns AND the specific moments AND your manual selection of which moments to anchor with.

Sample prompt:

“I am writing the strengths section of a performance review. The recurring strengths I want to recognize are:

  1. [pattern, e.g., ‘consistent quality of customer-facing communication under pressure’]
  2. [pattern]
  3. [pattern]

For each strength, here are the specific dated moments that demonstrate it: [list the moments]

Draft a strengths section that names each strength clearly, anchors it to the specific moments (named in concrete terms, not abstractions), explains the impact each strength had on the team or the work, and is two to three paragraphs total. Avoid generic praise language. Use the specific behaviors and outcomes I have given you, not what you imagine a good employee looks like.”

What AI gets wrong here. Even with the specific moments fed in, the model sometimes drifts into generic praise language (“a true team player,” “consistently exceeds expectations”). Catch this in editing. Every sentence in the strengths section should either name a specific moment or describe an impact that can be traced to a specific moment. If a sentence could appear in any review of any employee, cut or rewrite it.

Edit checklist before moving on. Read the strengths section aloud. Does each paragraph contain at least one specific, dated moment? Does each strength tie to a concrete impact? Could the employee read this and recognize themselves, or could it be about anyone on the team? Edit until it is recognizably about them.

Step 4: Draft the growth section (the hardest one)

What this step does. Names the growth areas in a way that lands as developmental rather than as criticism, while staying specific enough to be actionable.

What you do manually. This is the step where the manual work is irreplaceable. Take the growth patterns from Step 2 (probably two of your five patterns). For each one, ask yourself: what would I want for this person in the next six months? What is the specific behavior I am asking them to develop, in concrete terms? What support am I committing to provide?

Write down, in your own words and before you ask AI, what you actually want to say. Even rough sentences are fine. The reason to write it first: AI is good at calibrating language, but it is not good at deciding what the underlying message should be. The decision about what to surface as growth is your job. AI’s job is to help you phrase it.

Where AI augments. Once you have your rough draft of what you want to say, ask AI to help you calibrate the language: less harsh, more specific, more forward-looking, less identity-threatening.

Sample prompt:

“I am writing the growth section of a performance review. My rough draft is:

[paste your rough version, in your own words]

Please rewrite this so it:

  1. Names the specific behavior to develop, not a character trait.
  2. Anchors to the specific moments where the behavior showed up.
  3. Is forward-looking (what to develop) rather than backward-looking (what was wrong).
  4. Separates the behavior from the person’s identity (per the Harvard Business Review article ‘The Feedback Fallacy’ by Buckingham and Goodall, identity-threat language activates defensive responses and shuts down learning).
  5. Names what support I am committing to provide.

Generate three variations of the rewrite, each with slightly different calibration: one direct, one warmer, one more analytical. I will choose which to use.”

What AI gets wrong here. AI often over-softens. It sometimes calibrates the growth language so far toward warmth that the message becomes unclear. The employee reads it and walks out without knowing what they need to change. The fix: pick the variation that lands closest to your gut read, then sharpen any sentence that has gone fuzzy. The growth section should be kind and clear. Kind without clear is useless. Clear without kind is brutal. AI tends to err on the kind-but-fuzzy side. You correct toward clear-but-still-kind.

Edit checklist before moving on. Read the growth section. Could the employee finish reading and know exactly what behavior you are asking them to develop? Have you named the specific support you will provide? Does the section avoid identity-language (no “you are X” statements)? Does the section name future state rather than dwell on past failure?

Step 5: Surface what you missed

What this step does. Catches the things you would have written badly or skipped entirely. This is the second-highest value AI use in the workflow.

What you do manually. Nothing, yet. The draft is in shape from Steps 3 and 4.

Where AI augments. Feed AI the full draft and ask it to do a structured red-team review.

Sample prompt:

“Below is a draft performance review. Please do four things:

  1. Identify any claim in the review that is not supported by specific evidence in the text. List each one.
  2. Identify any generic praise or generic criticism language that could appear in any review. List each one.
  3. Identify any sentence that could be read as identity-threat language (attacking who the person is rather than what they did). List each one.
  4. Identify any growth area I did not address that the strengths section implies (e.g., if I praised someone for technical depth, did I also address whether they need development on the people side that often pairs with deep technical contributors).

Be specific and quote the exact sentences. Do not rewrite anything. I will decide what to do with each item.

[paste the full draft]”

What AI gets wrong here. Sometimes it flags things that are not actually problems — a specific moment named in the strengths section can read as “unsupported” to the model if the model lost track of the connection. Read each flag against the draft. Some you act on. Some you ignore.

Edit checklist before moving on. Address every flag that, on your read, is a real issue. Add evidence to any claim flagged as unsupported. Sharpen any generic-language flag into specific language. Rewrite any identity-threat flag. Address (or consciously decide to defer) any missed-development-area flag.

Step 6: Generate the conversation guide

What this step does. Prepares you for the actual delivery meeting. Most managers skip this step and pay for it during the conversation.

What you do manually. Nothing, yet. The review is in shape. Now you need a delivery plan.

Where AI augments. Ask AI to produce a conversation guide from the finished review.

Sample prompt:

“Below is a finished performance review. Please generate a thirty-minute conversation guide for delivering this review in person. The guide should include:

  1. A suggested order for walking through the sections (strengths first or growth first, with rationale).
  2. The two or three pieces in this review most likely to surprise the employee, with a sentence on how to land each one.
  3. Three anticipated employee responses to the growth section and how to handle each.
  4. Suggested questions for me to ask the employee at the end (what is your read, what would change your behavior most, what support do you need from me).
  5. The closing — how to end the conversation so the employee leaves with one clear next step.

[paste the full review]”

What AI gets wrong here. Sometimes it generates anticipated responses that do not match what you know about this specific employee. Read the predicted responses and edit them based on what you actually know. The structure of the guide will be useful regardless. The specific predictions need your human read of the person.

Edit checklist before moving on. Do you have a guide for the conversation? Are you prepared for the two or three reactions you most expect? Have you written down the questions you will ask at the end? Yes → you are ready. The review-writing portion is complete.

Where the time savings actually come from

A back-of-envelope breakdown for a single review of a single direct report:

StepManual timeAI-augmented timeTime saved
1. Gather raw signal30 min30 min (human only)0
2. Synthesize patterns45 min10 min35 min
3. Draft strengths30 min12 min18 min
4. Draft growth45 min25 min (heavier human edit)20 min
5. Red-team review20 min (manual self-review)8 min12 min
6. Conversation guide30 min (if done at all)10 min20 min
Total per report~3h 20min~1h 35min~1h 45min

For a team of five direct reports, you reclaim about eight to nine hours. That is the “8 hours” in the title of this article. It is real. It is also smaller than you would think — the gathering step is irreducible, and the growth section requires more editing than less. The reason AI-augmented review still wins on quality is not the time. It is that AI catches patterns and language calibration that the tired human brain would have missed at hour seven of a manual cycle.

This gap between the gross time AI appears to save and the net time you actually reclaim after editing applies to every manager task, not just reviews. The AI Time Savings calculator models the net across your whole task mix (notes, email, research, drafting, reviews), weighted by how much each task type actually needs human editing. Reviews carry the highest edit cost of any task, which is exactly why this workflow exists.

For an exact number on what your cycle will cost, the Performance Review Time calculator takes about three minutes and plans the actual hours by team size, review depth, and process. Most first-time managers underestimate cycle time by 50 to 100%. The calculator helps you plan honestly so the cycle does not eat the rest of your month.

The four things to never let AI do

Regardless of how good the workflow gets, there are four moves that always destroy the review. Do not delegate these to AI under any circumstance.

1. Do not let AI generate the raw observations. You feed it the data; it does not invent the data. The moment you ask AI “write a performance review for Priya based on what you think a good engineer of her experience level looks like,” you have stopped writing a review and started writing fiction. The team will recognize it as fiction within one paragraph.

2. Do not let AI make the calibration decisions. AI can produce calibrated language; it cannot decide which calibration is right for this person, this review, this cycle. The decision to push harder versus softer, to surface this growth area now versus next cycle, to celebrate this strength loudly or quietly — those are yours. AI is a thesaurus and a structural assistant, not a judgment-maker.

3. Do not let AI decide the rating or the recommendation. If your company uses ratings (1-5, exceeds/meets/below), the rating decision is a calibration call that requires knowing context AI does not have: where this person sits on the team relative to peers, what the calibration meeting will accept, what the rating means for compensation and trajectory. AI does not know these things. Decide the rating manually; let AI help with the writing that supports it.

4. Do not paste the employee’s name and demographic information into the prompt. Multiple research bodies have documented that language models can carry bias forward when given identity cues. The EEOC’s April 2023 technical assistance document on AI in employment decisions is explicit that employers have a legal obligation to assess for disparate impact when AI is used in employment-related decisions. The simplest way to reduce risk: anonymize before you prompt. Refer to the person by initials or “this direct report” in the prompt. The synthesis quality is identical. The bias risk is meaningfully lower.

What this workflow cannot fix

This workflow is for managers who have been gathering signal during the review period and need help compressing it into a useful written artifact. It is not for managers who have not been paying attention and are now hoping AI will rescue them at the end. AI will not rescue you. The output will be exactly as good as the inputs you supply, and inputs come from the work you did over the six months before the review, not from the prompt you write the night before it is due.

If your last review cycle ended with reviews that landed weakly and you are tempted to “use AI better this time” without changing what you do during the cycle, the pillar guide on running your first performance review covers the gathering practices that make this whole workflow possible. The Constructive Feedback Examples library covers how to surface the moments during the cycle so you have data when the cycle ends. Without those upstream practices, AI is just a faster way to produce empty reviews.

The EEOC’s 2023 technical assistance document and the Department of Labor’s broader guidance on AI in employment decisions establish that employers using AI in employment-related decisions retain legal responsibility for the outputs. Specifically: an AI tool that produces disparate impact across protected categories (race, gender, age, disability) creates legal liability for the employer even if the manager using the tool was unaware of the bias.

Practical implications for the workflow above:

  • Anonymize prompts. Do not include names, photos, demographic information, or identifying context in what you paste into the model.
  • Document your process. Keep notes on which steps used AI, which prompts you used, and what edits you made. If a review decision is later challenged, you want a record of the human judgment that shaped the final output.
  • Audit periodically. If you use this workflow across many reviews, occasionally compare AI-drafted versus your final-edited versions across demographic groups. You are looking for systematic differences in tone, calibration, or recommendation that you cannot explain by performance differences.
  • Do not use AI for legally sensitive content. Termination decisions, PIP triggers, and discrimination-adjacent topics should be human-written end-to-end. The risk is too asymmetric. The Should You Put This Employee on a PIP decision framework covers when those threshold conversations need to be entirely manual.

None of this is legal advice. If you are scaling this workflow across an organization, run it past your HR and legal teams first. The individual-manager use described in this article is meaningfully different from organizational rollout.

Frequently asked questions

Which AI model should I use for this workflow?

Any of the major general-purpose models (ChatGPT, Claude, Gemini) will do this workflow well enough. The differences between them matter less than how you structure your prompts and how disciplined you are about not letting AI invent data. If you have access to your company’s enterprise instance with data privacy guarantees, prefer that over the consumer version — the prompts in this workflow contain sensitive observations about your team members, and you want the data handling that comes with enterprise tooling.

Will the employee know the review was written with AI?

If you follow the workflow above, the employee will not be able to tell, because the review will contain specific moments and observations only you could have provided. If the employee can tell, you skipped a step — most likely Step 1 (gathering raw signal) or Step 4 (manual rewriting of the growth section). The signature of AI-only review writing is fluency without specificity. Specificity is what makes the difference, and specificity comes from you.

What if my company has banned AI use for HR-adjacent tasks?

Then do not use AI for these reviews. Some companies (especially in regulated industries or in jurisdictions with strict AI-in-employment laws like New York City’s Local Law 144) have policies that prohibit or heavily restrict AI use in employment decisions. Check before you use. If your company allows AI for drafting but requires the review to be human-authored “in substance,” the workflow above is compatible — the human is doing the gathering, the calibration decisions, and the final edits. The AI is doing structural drafting and red-team review.

Can I use this workflow for ratings or compensation recommendations?

No. The rating is a calibration call that depends on context AI does not have: peer comparison, team norms, organizational compensation philosophy, calibration meeting outcomes. Decide the rating manually. Let AI help with the writing that supports it. Same for any compensation recommendation.

How much should I edit the AI-drafted strengths section?

Heavily. The strengths section is where AI-generated text drifts most toward generic praise. Read every sentence. If a sentence could appear in any review of any team member, rewrite it with a specific moment. Most managers cut about 30 to 40% of the AI strengths draft and rewrite the rest. That is normal. That is the work. The remaining 60% saves you the time of generating the first draft from scratch, which is where the speed benefit comes from.

What if the AI surfaces a growth pattern I had not seen?

Take it seriously, but verify before including it. AI is doing pattern recognition on the data you supplied. If the data supports the pattern, it is real signal and you should consider including it. If the data does not support the pattern, AI is interpolating from what management language usually looks like, and you should drop it. The “cite specific observations” instruction in Step 2 helps separate these.

Should I tell my manager I am using AI for reviews?

Depends on your company culture and policy. If there is an explicit policy, follow it. If there is no policy but the topic is sensitive, mentioning to your manager that you are using AI as a drafting assistant (not as a decision-maker) is usually the safer move. The conversation tends to land easier than it sounds — most senior managers are themselves trying to figure out the same question, and the candor signals that you are thinking carefully about the use.

What if I do not have a tracking spreadsheet from the past six months?

This cycle, do the best reconstruction you can in forty-five minutes per direct report. Going forward, the private tracking spreadsheet ritual is the upstream practice that makes everything in this article work. Two minutes a day during the cycle replaces forty-five minutes of frantic reconstruction at cycle end. By next cycle, you will have the data.

What to do next

Three concrete moves, in order:

  1. Run the Performance Review Time calculator for your upcoming cycle. Three minutes. Plug in team size, review depth, and your tracking practice. The output tells you how many hours the cycle will take honestly, which lets you budget the time before the cycle starts rather than after it has eaten your month.
  2. Set up your input piles before you ask AI for anything. For each direct report, copy your tracking notes plus any additional raw signal into one document. Aim for two pages of specific, dated, observable behaviors. Without this step, every later step produces mush. With it, you have what you need.
  3. Run the six steps on one review first, before scaling. Do not do all five direct reports in one sitting using this workflow for the first time. Run it on one report end to end. Notice where AI saves you time and where it costs you time on editing. Calibrate. Then run it on the remaining four with your calibration applied.

For a quick gut-check on whether AI is helping or hurting your overall management, the Is AI Making You a Lazier Manager? free assessment takes about three minutes and surfaces the five patterns where AI augmentation can quietly degrade your management instead of improving it. Take it before or after this article. The workflow above is meaningful only if AI is making your management sharper, not blunter.

Membership · Founder Pricing

A Personal Home for First-Time Managers

Every toolkit, a direct line to George, a private community, and the structured 52 Weeks to Better Manager curriculum. One annual membership. Founder pricing locks in for waitlist.

Learn about Membership →