AI Meeting Notes for Performance Calibration Meetings

Published8 min read

Calibration committees adjust ratings about a quarter of the time, and the reasoning behind an adjustment usually lives nowhere except the memory of whoever argued it. Why the meeting record, not just the final number, is the part worth capturing, and where that record should stop.

Two managers rating the same work do not naturally agree

A calibration meeting exists because a single number from a single supervisor is not reliable enough to act on by itself. A 2019 meta-analysis in Frontiers in Psychology, covering 224 independent samples and more than 43,000 individuals, put the average interrater reliability of supervisory overall job performance ratings at 0.56 observed, rising to 0.61 once corrected for range restriction. That is a moderate correlation between two supervisors rating the same underlying performance, not a strong one.

The same study found the number gets worse exactly where it matters most. When ratings are used for administrative purposes, promotions, pay, and rankings rather than research or feedback, observed interrater reliability dropped to 0.45. The gap between administrative and research-purpose reliability ran as high as 0.24 across the job performance dimensions the researchers tested. In plain terms, the moment a rating is going to change someone's pay or promotion path, two managers looking at comparable work are less likely to land on the same score than when nothing is riding on it.

That is not an argument against having supervisors rate their own people, who are usually the ones with the most direct information. It is the argument for calibration in the first place: a second pass, across a group of managers, catches the disagreement that a single rater's number hides by itself.

Calibration moves a quarter of ratings, mostly downward

The most direct evidence on what calibration committees actually do comes from a 2019 study in Management Science using proprietary data from a large multinational organization. The committees in that data adjusted ratings sparingly, about 25 percent of the time, but when they did adjust, downward moves were both more frequent and larger in magnitude than upward ones. Committees tended to pull down the ratings of supervisors who had been more generous than average and pull up the ratings of supervisors who had been more critical than average.

The same research found calibration cuts both ways on bias. It improves consistency across supervisors and reduces leniency bias, the tendency for a soft rater's whole team to look better than it should. But it also exacerbates centrality bias, the pull toward the middle of the scale that flattens genuine differences in performance. And committees defer more to supervisors who sit closer to the person being rated in the org chart, adjusting less when the rater has a clear information advantage over the room.

None of that is a reason to abandon calibration. It is a reason to notice that calibration is not a rubber stamp on a spreadsheet. It is a live negotiation with a specific shape: sparing but real adjustments, a directional skew, and a deference pattern based on who in the room actually knows the work. That negotiation happens once, out loud, and then the spreadsheet is the only part of it anyone keeps.

The adjustment is a decision no one wrote down

Picture the actual moment: a supervisor's rating comes up, someone in the room pushes back, two or three people cite specific examples, the number moves half a point, and the committee moves to the next name. What gets recorded afterward is almost always the new number. What rarely gets recorded is the argument that produced it: which examples were cited, which criteria the room agreed applied, and why this case landed where it did relative to the last similar one.

That gap shows up twice. It shows up when the rating is challenged, and the manager who has to explain it is reconstructing a conversation from weeks earlier instead of pointing to what was actually said. And it shows up the next cycle, when a comparable case comes up again and the room has no record of how it reasoned through the first one, so the same debate happens from scratch with a different set of people remembering it differently.

A number without its reasoning is a decision with the working shown thrown away. The research above says the reasoning matters: it is the part that turns a moderately reliable individual rating into something closer to a consistent, defensible one. Losing it after the meeting ends undoes a meaningful share of what the meeting was for.

A meeting record, not an employee record

It is worth being precise about the boundary here, because performance conversations are an area where the wrong tool applied the wrong way causes real harm. A calibration meeting is a room of managers discussing their own ratings with each other. Everyone in that room knows it is happening and is a direct participant in the discussion, the same as a board update, a hiring debrief, or a skip-level. That is a meeting record, no different in kind from any other internal meeting a team already decides is worth capturing.

That is a different thing from monitoring an individual employee's day-to-day work, listening in on a one-on-one without both people's knowledge, or building a profile of someone from conversations they were never part of. Botless capture does not remove that line, and it does not remove any consent obligation the calibration meeting itself already carries. The record this article is describing is the committee's own reasoning about its own decision, kept the same way a company already keeps notes from any other manager meeting.

Where that line gets blurry is worth naming directly: a calibration note should capture the criteria applied and the argument made, not a transcript of every aside about the person being discussed. The next section is about drawing that line in practice, not just in principle.

What the record should hold, and what it should leave out

A useful calibration record is short and specific: which rating moved, which direction, the criteria the room agreed applied, the examples cited in support, and who raised the objection that started the discussion. That is enough to answer both of the questions that come up later, why did this land here, and how did we handle the last case like it.

It should leave out anything that is not the committee's own reasoning: offhand characterizations of the person, speculation about factors outside the review period, and side conversations that drift away from the criteria being applied. A calibration record that only captures the argument, not the aside, is both the fairer version and the one worth keeping.

In practice this looks like assigning someone in the room to note the reasoning behind every adjustment as it happens, or using a capture tool that sits in the meeting the same way it would for any other internal review, then having a named owner clean the notes down to that short, specific shape before the meeting record is stored anywhere.

Why memory alone keeps losing this

The pattern across both studies points the same direction. Two supervisors rating the same work agree only moderately, worse when something real is riding on the number. A committee that adjusts a quarter of ratings, skewed downward, is doing real work correcting for that disagreement. And then the part of the process that did the correcting, the actual argument, is the part nobody keeps.

That is not a discipline problem. It is what happens to any unrecorded reasoning in a room full of people who each have their own next meeting to get to. The fix is not a longer meeting or a stricter process. It is treating the calibration discussion the same way a well-run team already treats any other decision meeting worth remembering: write down the reasoning while it is being said, not the conclusion after everyone has moved on.

Sources

FAQ

Does capturing calibration meeting notes mean recording individual employees?

No. A calibration meeting is a discussion among managers about their own ratings, and everyone in the room is a direct participant who knows the discussion is happening. That is different from monitoring an employee's day-to-day work or a conversation they were never part of, and capturing the committee's own meeting does not change any consent obligation the meeting already carries.

How much do calibration meetings actually change ratings?

A 2019 study in Management Science using proprietary data from a large multinational organization found calibration committees adjusted ratings about 25 percent of the time, with downward adjustments more frequent and larger than upward ones.

Why do two managers rate comparable work differently in the first place?

A 2019 meta-analysis in Frontiers in Psychology covering more than 43,000 individuals found the average interrater reliability of supervisory performance ratings was moderate, 0.56 observed, and dropped further, to 0.45, specifically when ratings were used for administrative decisions like pay and promotion.

What should a calibration record leave out?

Anything that is not the committee's own reasoning: offhand characterizations of the person being discussed, speculation outside the review period, and side conversations that drift from the criteria being applied. The record should hold the argument, not the aside.

Never lose the thread of a meeting again.

Driffle keeps the decisions, owners, and context from every conversation searchable when work resumes.

Request access

Similar articles