How Accurate Are AI Notetakers?

Published15 min read

An AI notetaker has to do three different jobs and it is good at only one of them. Transcription is close to perfect on clean audio. Action items, in the one published test that measured them separately, came in between 62 and 87 percent. This reads that study, the speaker attribution limits Granola documents, and the clause in Otter's terms telling you not to rely on AI output without confirming it, then gives a ten minute test that shows you how a tool fails before you buy it.

There are three accuracy numbers and only one of them is good

Ask how accurate an AI notetaker is and you get answered on the easiest part of the question. Transcription accuracy is the number that appears in marketing copy, and it deserves to be high, because getting words off audio is the part of this problem that has been worked on for twenty years. It is also the part that matters least to you.

A notetaker does three jobs in sequence, and they fail at different rates. Ben Henry-Moreland, writing for Kitces.com, splits them cleanly: dictation accuracy, meaning how well the tool captures what was said and by whom; summarization accuracy, meaning how well it condenses the transcript into takeaways and action items; and meaning accuracy, meaning whether it reads the sarcasm, the hesitation and the tone that told everyone in the room what was actually going on.

The one published test that measures the first two separately comes from the consulting firm Oasis Group, reported in that same article. It ran six meeting notes tools built for financial advisers, Jump, Zocks, Finmate, Zeplyn, Greminders and Mili, against a scripted meeting. Most transcribed it with 100 percent accuracy. The same six tools captured key data points from that meeting at around 85 to 96 percent, and action items at 62 to 87 percent.

Sit with the bottom of that last range. Roughly one action item in three was wrong, in a test where the transcript underneath it had no errors in it at all. The failure was not in hearing the meeting. It was in deciding what the meeting had committed anyone to.

Two honest caveats. A scripted meeting read aloud is the friendliest input these tools will ever see, with no crosstalk, no one on a phone in a car, no two people starting a sentence at the same moment, so real numbers are probably worse rather than better. And the study did not compare its results against a human taking notes in the same meeting, which the Kitces piece says plainly, so 85 percent is not obviously worse than the person who was also trying to follow the argument. The figures are from a 2025 study and every one of these products has shipped changes since. What has not changed is the shape of the gap between the layers.

  • Dictation: what was said, and by whom. Close to solved on clean audio.
  • Summarization: what it meant and what it committed people to. Measurably weaker.
  • Meaning: the joke, the reluctant yes, the condition attached to the agreement. Not attempted by most tools.

The expensive mistake is a real task with the wrong name on it

Transcription errors announce themselves. A garbled word reads as a garbled word, and the reader's eye stops on it. Summary errors do the opposite. A well formed action item assigned to the wrong person looks exactly like a well formed action item assigned to the right one, and it arrives in the same even, confident register as the nineteen lines around it that are correct.

This is where attribution does the damage. The Kitces article notes that one of the six tools transcribed the meeting text accurately and still had trouble working out which participant said which part, and that misattributed dialogue is one of the more persistent accuracy complaints advisers report. Where a tool generates and assigns follow-up tasks from who was speaking, a misattribution does not stay a labelling error. The task lands on someone who never agreed to it, and the person who did agree never hears about it again.

Granola's documentation is unusually candid about the limits here, and worth reading before you judge any tool on this. Speaker tags, which put participant names above their parts of the transcript, need setting up, and they are available in the desktop app on macOS and Windows for Google Meet, Zoom and Microsoft Teams. Without them, transcripts show Me and Them, which correspond to your microphone input and your system audio, and the docs say the AI infers who is who from contextual clues in the transcript. Inference is a guess that is usually right.

The same page lists what speaker tags still cannot resolve: several people speaking through one shared meeting-room device cannot be told apart, audio playing from somewhere else on your computer may be transcribed with no speaker attached, and when several people talk at once the tool may not assign every part of the conversation to the correct person. There is one more line that matters operationally. Speaker tags only apply while the tool is actively transcribing, so they do not add names to transcripts you already have.

Read that list against your calendar and the problem gets specific. The meetings most worth capturing are the crowded decision call where three people talk over each other and the session where four colleagues share one laptop in a room. Those are the two conditions the attribution layer handles worst, and they are the two where a task with the wrong owner costs the most.

  • Set up speaker attribution before the trial, not after, because most tools will not backfill names onto old transcripts.
  • Treat any meeting held on a shared room device as attribution-free by default.
  • Check owners first when you review a summary. It is the error least likely to look like one.

Read the part of the contract nobody reads

The most direct answer to how much you should trust these summaries is written by the vendors themselves, in the document you clicked past during signup.

Otter's terms of service, effective 19 September 2025, state in the disclaimers section that Otter.ai makes no warranty about the completeness or accuracy of the transcription. The same section adds a clause specific to AI features: outputs may be inaccurate, inappropriate, false, incomplete or biased; you are responsible for implementing reasonable practices, including human oversight, to ensure outputs are correct and complete; and you should not rely on AI outputs without independently confirming their accuracy.

That last sentence is the whole argument of this article, written by the company selling the product. Fathom's terms, last updated 4 March 2026, carry no equivalent AI-specific clause, and instead provide the service and all content on an as-is basis with no warranty that it will be free of errors.

A disclaimer is not a confession. Every vendor in every category writes one, and none of them means the product is bad. What is useful about Otter's version is its specificity. It names the two things that can be wrong, the transcription and the outputs, and it names the remedy, which is a human independently confirming them.

So the review pass is not cautious hygiene that careful people add on top. It is the deal. The pitch is that the tool gives you back the twenty minutes you used to spend writing notes. The contract quietly hands some of those minutes back as review time, and nobody prices that when they compare plans.

The question that should decide the purchase

Once you accept that every summary needs confirming, the product's job changes. It is no longer to be right. It is to make being wrong cheap to discover.

There is a single measurable version of that, and it is the best evaluation question in this category: how many moves separate a sentence in the summary from the seconds of conversation behind it? One move means you will check. Four means you will not, and a check you skip under pressure is the same as no check at all.

Granola ships the right shape of answer. Its docs describe a magnifying glass beside each note that shows where in the transcript or in your own raw notes that note came from, which turns a doubt into a lookup. Fireflies takes a different route to the same place. Its default general summary includes a Time-stamped Notes section that groups the notes into timestamped chapters with each section linked to the transcript, so you can jump to the relevant part of the conversation, and on Pro and above you can expand an individual summary point to see additional context without opening the full transcript.

Slow looks like this: a clean one-page summary, a separate six thousand word transcript in another tab, and a search box. The check that should take nine seconds takes four minutes, which means it happens on the first Tuesday and never again. A month later the summary is the only record anyone consults, and whatever it got wrong is now what happened.

This is also the question to ask a salesperson, because it has a demonstrable answer. Do not ask how accurate the summaries are. Ask them to open a real summary, pick an action item you choose, and show you the moment it came from. Count the clicks.

  • Per-claim provenance: can you trace one line, not just open the whole transcript?
  • Does the trace land on a moment in the conversation, or on a page you then have to search?
  • Is the transcript still there in three months, or does retention remove the thing you would check against?

Fixing the transcript does not fix the notes

Say the check works and you catch a wrong owner. There is one more step most buyers never test, which is whether the correction actually propagates.

Fireflies is explicit about how this works, which makes it a useful example rather than a criticism. Its help documentation tells you to make your edits to the transcript, including updating speaker labels, and then regenerate the summary so it reflects the changes. Regeneration is a deliberate action, not an automatic consequence of the edit. And the same page notes that only the meeting owner or host can regenerate a summary.

Follow that through on a real team. A colleague reads the summary, spots that an action item has been assigned to them when it belongs to someone else, and fixes the speaker label. The transcript is now right. The summary is still wrong, and they are not allowed to regenerate it. The wrong version is the one that gets pasted into Slack, because it is the one that is short.

Granola has the same two-step property in a different place. Notes are generated from the transcript, your own raw notes and the calendar event, and if you change the raw notes you regenerate to have the change taken into account. Its docs also point out that editing one note changes only that note and has no effect on how future notes are written, so a correction teaches the tool nothing.

So add two questions to the evaluation. When someone corrects a name, what else updates, and what stays stale? And who is allowed to make the correction, the person who spotted the error or only the person who happened to own the call?

A ten minute test worth more than a month of trial

Trials mislead because of who runs them. You try a new notetaker on the meetings you feel calm about, usually a one to one or a call with two people and an agenda. Those are exactly the conditions where every tool in this category performs well, so a fortnight of trial tells you almost nothing about the Thursday afternoon that actually needs capturing.

Design one meeting instead. Take a real internal call, not a mock one, and build six things into it. Then read the summary once and you will know more than a month of pleasant use would have told you.

The reason this beats comparing accuracy percentages is that percentages tell you how often a tool is wrong and this tells you how it is wrong, which is the thing you have to live with. A tool that drops a detail is survivable. A tool that invents a confident owner for a commitment nobody made, in a product where you cannot trace the claim, will eventually cost you a relationship.

  • Reverse a decision out loud, halfway through. Agree on something, then two minutes later agree on the opposite. Check which one the summary recorded, and whether it recorded that the first was reversed.
  • Say one commitment as a joke or a hypothetical. See whether it appears as an action item with a name attached.
  • Say one name and one number that matter, once, at normal speed, without spelling either. Check both against what was said.
  • Include one participant on a shared room device or a bad connection, and see who their lines get attributed to.
  • Pick one action item and count the moves to the moment it came from. More than two and the check will not survive a busy week.
  • Correct one wrong owner, then check what updated, what stayed stale, and whether a teammate could have done it instead of you.

Where Driffle fits, and what it does not fix

Driffle is built on the assumption in the fourth section of this article. A claim you cannot trace is a claim you will not check, so the traceability has to be part of the retrieval, not a feature buried behind it. Ask Driffle a question about your work and the answer comes back attached to the moment it came from, which makes confirming a line a click rather than an excavation.

The capture side is designed for the same reason. Driffle transcribes audio from your own device, with nothing joining the invite and no recording indicator on the call, and pulls out speakers, decisions and action items as you wrap. Those stay connected to the conversation they came from, so a tracked commitment is not an orphaned line of text in a list somewhere.

The limits, stated plainly, because this article is about not trusting confident claims. Driffle is a Mac app for Apple Silicon, installed as a direct download; Intel Mac, Windows and Linux are a request form and not a product you can use today. Transcription and the AI work happen in the cloud, so nothing here is on-device or offline. And the three accuracy layers at the top of this article apply to Driffle as much as to anything else. Its summaries are generated by models and models are wrong sometimes.

The claim is narrower than accuracy and more useful than it. When a summary is wrong, you should be able to find out in seconds, from the summary itself, without asking the person who ran the call. If that is what you want from a meeting notes tool, you can download Driffle for Mac at driffle.ai/download and run the ten minute test above on it before you trust it with anything.

Sources

FAQ

How accurate are AI notetakers overall?

It depends entirely on which layer you mean. In the Oasis Group study of six meeting notes tools reported by Kitces.com, most transcribed a scripted meeting with 100 percent accuracy, captured key data points at around 85 to 96 percent, and produced action items at 62 to 87 percent accuracy. Transcription is close to solved on clean audio. Working out what a meeting committed people to is not. Treat those figures as a 2025 snapshot of the gap between the layers rather than as current per-product scores.

Do AI notetakers make things up?

A summary can contain an action item nobody committed to, a name that was guessed phonetically and never checked, or a decision recorded without the condition attached to it. Whether you call that invention or over-confident compression, the practical problem is the same: the wrong line reads exactly like the right ones, in the same even tone, so nothing about it draws your eye.

Which is worse, a transcription error or a summary error?

A summary error, by a wide margin. A garbled word is visible and self-correcting, because the reader can tell something went wrong. A fluent action item attached to the wrong owner is invisible, gets forwarded, and becomes the record the team acts on. It is also the error type the published numbers say is most common.

If I correct the transcript, does the summary update?

Not automatically, in the tools that document this. Fireflies tells you to edit the transcript and speaker labels and then regenerate the summary so it reflects those changes, and notes that only the meeting owner or host can regenerate. Granola generates notes from the transcript plus your raw notes, and regenerating is a deliberate step there too. Ask any vendor what propagates after a correction and who is allowed to make it.

What is the fastest way to check a summary?

Do not re-read the transcript. Pick the claims that carry consequences, meaning the owners, the numbers, the dates and any decision that was reversed, and trace each one to the moment it came from. That is only fast in a product that offers per-claim provenance, such as Granola's per-note source view or Fireflies' timestamped sections linked to the transcript. Count the moves that trace takes before you buy, because a check that takes four minutes will not happen twice.

Would a human notetaker be more accurate?

Probably not in a way that helps, and the honest answer is that the study did not measure it. A person taking notes while also following the argument misses things too. The real difference is presentation. Human notes are visibly partial, with gaps and question marks that tell the next reader to check. AI notes arrive complete and evenly confident whether or not they are right, which is why the ability to verify matters more than the percentage.

Never lose the thread of a meeting again.

Driffle keeps the decisions, owners, and context from every conversation searchable when work resumes.