The Decision Was Sound and the Result Was Bad
A decision and its outcome are two different things, and six months later only one of them is still visible. Here is what a meeting record has to hold if a team wants to judge how a call was made rather than how it turned out.
Six months later, the result arrives
You picked the vendor. You greenlit the migration, or hired the candidate the panel was split on, or held the launch date when half the room wanted to move it. Some months pass and you find out how it went.
If it went badly, someone will say the reasoning was obviously wrong, and they will be able to explain exactly why in a way that sounds convincing to everyone including you. If it went well, nobody will say anything at all. That second case is the one worth worrying about, and almost nobody does.
The meeting where the call was made is still in the calendar. There is even a note. It records what was decided, who owns it, and when the next check-in is. Read that note after the fact and it tells you close to nothing about whether the decision was any good, because it kept the output and dropped every input that went into it.
Knowing how it ended changes which facts look relevant
Baruch Fischhoff ran the experiment that pinned this down in 1975. Subjects read a briefing about a nineteenth century clash between British and Gurka forces and estimated how likely each possible ending had been. Being told which ending actually happened raised the estimated likelihood of that ending, which is the part everyone already expects.
The part worth an operator's attention is what outcome knowledge did to the supporting facts. One item in the briefing noted that British officers learned caution only after sharp reverses. Subjects told the British had won rated that sentence as highly relevant. Subjects told the Gurkas had won rated the same sentence as close to irrelevant. Identical words, opposite readings, and the only variable was an ending that the writer of the sentence did not know about.
Fischhoff called the effect creeping determinism. His second experiment asked people to ignore the reported outcome and judge as they would have without it, and they could not do it. They believed that even without the report they would have seen the relevance exactly as they now saw it. His third experiment found people equally wrong about what other people had understood before the outcome was known.
So the failure mode inside a decision review is not that people judge harshly. It is that a person reviewing the call is reading a different set of facts than the person who made it, while sincerely believing they are reading the same set.
Knowing about the bias buys you nothing
Jonathan Baron and John Hershey put a number on the judgement side in 1988. They gave people a medical decision, a fifty five year old man with a heart condition choosing whether to have a bypass operation that carried an eight percent chance of dying on the table, and described the decision identically to everyone. It was rated as a better decision when the patient survived and a worse one when the patient died.
A pre-registered replication published in 2023 ran the same experiment with 692 participants and found a larger effect than the original, with Cohen's d between 0.77 and 1.10 depending on whether the person being judged was the physician or the patient.
The number that should actually change how a team works is a different one from the same study. The researchers separated out the participants who had stated, before seeing any of it, that outcome information should not be taken into account when judging the quality of a decision. Those participants still showed the bias, at d = 0.64. Agreeing with the principle bought them almost nothing.
That is the practical argument in one line. Every team believes it judges process rather than results, and the evidence says that belief does not survive contact with an actual result. A review that depends on people setting the outcome aside in their heads depends on the one thing that has been measured and found not to work. What does work is an artefact: something written down before the result existed, which the result cannot edit.
What a decision-only note is actually good for
Most meeting records hold one line about the choice. Going with vendor A, owner Priya, revisit in Q3. After a bad result that line does exactly one job, which is telling you who to talk to. It contains nothing that could establish whether the reasoning was sound, so the review falls back on the only evidence anyone has, which is how it turned out.
A record like that is an accountability instrument. Making it a learning instrument means keeping the things that stop being knowable the moment the result lands.
- The options that were genuinely on the table, and the specific reason each rejected one lost. Rejected options disappear from a meeting record faster than anything else in it.
- What people thought the odds were, in their own words, before the call. "I think this is close to a coin flip and I want to take it anyway" is a completely different decision from "this is going to work", even when both end in the same choice.
- What was known to be unknown at the time. The questions nobody in the room could answer are the strongest available evidence that a decision was made carefully rather than luckily.
- The conditions attached. What would have made the answer different, and what signal would have shown that an assumption had broken.
- Who disagreed and on what grounds. The substance of the objection, not the fact that there was one.
Dissent is the first thing memory rewrites
That last item deserves its own attention, because it degrades faster than the rest. After a bad result, the person who raised a concern remembers a sharper and more certain objection than the one they made, and so does everyone who heard it. The hedges fall away. A mild question becomes a warning that was ignored.
After a good result the same objection vanishes from the story completely. Nobody is being dishonest in either direction. Both versions feel like recall.
The cost is that the team loses the ability to tell a well-argued dissent from a reflexive one. That distinction is close to the most valuable thing a group can know about its own decision-making, because it determines whose caution to weight heavily next time, and it is unrecoverable from memory once a result is in.
The reviews nobody books
Teams schedule a review when something goes wrong. Very few book one after a win. The archive that results is systematically skewed: every decision that has been examined is a decision that failed, and every decision that worked keeps its reasoning permanently unexamined.
This is how a bad call that happened to pay off turns into precedent. It gets cited in the next meeting as the thing that worked last time. Because the reasoning was never written down, nobody can check whether it worked because of the reasoning or in spite of it, and the company ends up carrying a rule it has no way to audit.
This is not an incident review, and it should not borrow that shape. Nothing broke, there is no timeline to reconstruct, and the question is not what happened. The question is whether the way this call got made is a way worth making calls.
Fixing the skew does not require more meetings. It requires the review to be cheap enough that it can be triggered by the size of a decision rather than by the badness of a result. When pulling up what was said in the original meeting takes seconds instead of an afternoon spent scrubbing recordings and searching Slack, reviewing something that went well stops feeling like a luxury.
What the missing record does to how people decide
There is a second-order effect here that is worth naming plainly. If the only durable trace of a decision is how it turned out, the safest thing any individual can do is pick the option that will read well in retrospect rather than the option with the best expected result.
Nobody announces that they are doing this. It shows up as a quiet preference for the conventional choice, the one with a recognised vendor name attached, the plan that spreads responsibility widely enough that no single person owns the downside. Every one of those preferences is rational for the person holding it and expensive for the company.
A record of the reasoning is what makes the braver option survivable. If you can point at what you believed, what you did not know, and what would have changed your mind, then a bad result stays a bad result. Without that record, a bad result is the whole case against you, and everyone watching adjusts their own behaviour accordingly.
Running the review so the record does the work
The mechanics matter more than they look, because most of the value is destroyed in the first two minutes of the conversation.
- Have everyone read the original record before the discussion opens, not during it. Whoever speaks first sets the frame, and if the first thing said out loud is what happened, the reconstruction is already contaminated.
- Ask the two questions separately and in that order. Was the reasoning sound given what was available at the time. Then, as a separate question, how did it turn out. Asked in the other order they collapse into one question, and the answer is always the result.
- Write down the answer to the first question. The decision now has two entries, the call itself and the later verdict on how the call was made. The second entry is the only thing that accumulates into a team knowing whether it decides well.
- Keep the person who made the call in the room, and have them read their own words rather than recall them. They are subject to the same effect about themselves, which is why Fischhoff's second experiment matters as much as his third.
- Decide before any of this whether the archive can be used in performance conversations. A record kept for learning and a record kept for evaluation are the same bytes with very different consequences, and people speak differently in a meeting when they know which one they are feeding.
Why this is a capture problem before it is a process problem
Everything above assumes the material exists. Usually it does not, and the reason is timing rather than discipline. To hold a usable record of a decision you have to have captured the meeting before you knew the decision mattered. By the time a call is obviously consequential the result has arrived, and the result has already edited every memory you would need to reconstruct it from.
That rules out the tidy fix, which would be deciding in advance which meetings deserve a careful record. A few announce themselves. Most do not. A large share of the decisions that get reviewed later were made in the last ten minutes of a meeting booked for something else entirely.
This is the reason Driffle treats capture as continuous rather than as something someone remembers to switch on. Meetings get captured without a bot joining the call, and the screen context around them holds the documents and dashboards that were actually open while the call was being made. All of it stays searchable afterward, so pulling up what a room believed in March is a query rather than an excavation. The reasoning survived because everything did, not because somebody correctly predicted which half hour would matter.
None of this turns a bad decision into a good one, and a record is no defence for a call that was careless at the time. What it does is make the difference visible after the fact. A team that can tell a good decision with a bad result apart from a bad decision with a bad result gets better at deciding. A team that cannot gets better at looking right.
Sources
- Hindsight is not equal to foresight: The effect of outcome knowledge on judgment under uncertainty (Journal of Experimental Psychology: Human Perception and Performance 1(3), 1975, pages 288 to 299) - Baruch Fischhoff, American Psychological Association
- Outcome bias in decision evaluation (Journal of Personality and Social Psychology 54(4), April 1988, pages 569 to 579) - Jonathan Baron and John C. Hershey, via PubMed
- Outcomes Affect Evaluations of Decision Quality: Replication and Extensions of Baron and Hershey's (1988) Outcome Bias Experiment 1 (International Review of Social Psychology 36(1), 2023, article 12) - Sriraj Aiyer, Hoi Ching Kam, Ka Yuk Ng, Nathaniel A. Young, Jiaxin Shi and Gilad Feldman
FAQ
Is a decision review just a postmortem by another name?
No. A postmortem or incident review starts from something that broke and works backwards to establish what happened. A decision review starts from a choice that was made deliberately and asks whether the way it was made holds up, given only what was available at the time. The clearest difference is that a decision review is worth running when the result was good, and a postmortem never is.
How much detail does the record actually need?
Less than people assume. Five things carry most of the value: the options considered, what people said their confidence was, the questions nobody could answer, the conditions that would have changed the answer, and the substance of any disagreement. That is a short paragraph per decision, not a document. The reason it rarely exists is not length, it is that nobody knows at the time which decisions will be worth the paragraph.
Does recording who disagreed make people less willing to disagree?
It can, and that risk is real enough to design around rather than wave off. What makes it safe is settling who can read the archive and what it can be used for before the meeting rather than after a bad quarter. A record that people know will be read back to them in a performance conversation will produce more careful speech and less useful content, which defeats the point of keeping it.
Why not just re-read the transcript?
A transcript holds the words. It rarely holds what was on the screen, what the numbers said that morning, or which of several parallel threads the room was reacting to when someone said the plan looked fine. It also still requires somebody to find the relevant five minutes inside an hour of talk, which is most of the reason transcripts get generated and then never opened again.
What if the decision genuinely was bad?
Then the record shows that as well, and shows it in a form somebody can act on. "We had the information in front of us and did not weigh it properly" is a specific finding with a fix attached. "It went badly" is not a finding at all, and a team that only ever reaches the second one keeps making the same mistake with new vocabulary.