Why Recording Every Meeting Makes the Important Ones Harder to Find
A team that captures every meeting assumes a bigger archive is a better archive. The best measurement anyone ever ran on that assumption used 40,000 documents of memos and meeting minutes, and found that experienced searchers stopped while confident they had found three quarters of what mattered, when they had actually found a fifth of it. Deciding which meetings stay out of the record is a retrieval decision, and almost nobody makes it.
Search quality is a property of the collection, not of the search box
Every team that starts capturing meetings starts with the same working assumption: more is better. Capture is cheap, storage is cheap, and any meeting you skip is a meeting you cannot search later. So the default becomes capture everything and let search sort it out.
The thing teams measure is capture. Did the note get made, did it land in the right folder, did the action items come out with owners. Those are easy to check the same day. The thing nobody measures is retrieval. Six months later, when somebody needed an answer that was genuinely in there, did they get it, and how would anyone find out if they did not.
That gap matters because the quality of a search is not fixed by the tool. It is set largely by what sits in the collection being searched. A search that works well over two hundred meetings can quietly stop working over five thousand, with no change to the software and no complaint from anybody using it.
The lawyers believed they had found three quarters of it
The most careful measurement of this ever run was published by David Blair and M. E. Maron in 1985. They studied an operational full text retrieval system holding just under 40,000 documents, roughly 350,000 pages, assembled for the defense of a large corporate lawsuit. The corpus was reports, correspondence, memoranda and the minutes of meetings, which is close to what a team with a meeting archive is generating every week.
Two defense attorneys generated 51 information requests. Two paralegals experienced with the system turned each request into queries and searched. The lawyers reviewed what came back, asked for refinements, and searched again. The protocol treated a request as finished only when the lawyer stated in writing that they were satisfied, meaning they judged that more than 75 percent of the relevant documents had been retrieved. The lawyers had stipulated that 75 percent was the level they needed, and they considered all of it essential to defending the case.
Only after a lawyer declared themselves satisfied did the researchers go and measure what had actually been retrieved. Average recall came out at 20 percent. Average precision was 79 percent. So four out of five relevant documents were still sitting in the database, unread, at the moment the people searching decided they were done.
The paper is direct about why that result surprises people who have used such a system. A searcher only ever sees the set they retrieved, and that set looks good, because precision is fine. They never see the relevant documents that did not come back. There is no error message for the thing you did not find.
Why a bigger archive is a worse archive to search
Blair and Maron did not stop at the number. They explained the mechanism, and the mechanism is what should worry anyone building up a meeting archive.
In their database, plenty of single search terms returned more than 10,000 documents. A result set that size is useless, so the searcher narrows it by adding another term, and another. Each added term can only shrink the set, and each one carries a real chance of excluding documents that were relevant. The paper works the arithmetic through: for a query intersecting five terms, using optimistic probabilities, expected recall lands around 0.028, under 3 percent of the relevant material. Using more realistic probabilities it falls to 0.0009, which means that if a thousand relevant documents existed, the query would likely return one of them.
The second half of the mechanism is vocabulary. One issue in the case was an accident. Queries were built around the word accident plus the obvious proper nouns. Later the researchers found the accident discussed as an event, an incident, a situation, a problem, a difficulty. People who were personally involved tended toward euphemism, so it became an unfortunate situation. Sometimes it was referenced obliquely, as the subject of your last letter, or as what happened last week, or, in the opening line of the minutes of a meeting called specifically to discuss it, as Mr. A saying that we all know why we are here.
There is a longer example in the paper that is worth sitting with. Searching for one technical topic under the phrase trap correction, the researchers found the same subject discussed as the wire warp, then as the shunt correction system, then, in documents written by the man who invented it, as the Roman circle method, then, in the records of tests run in another city, as the air truck. That chain consumed a full 40 hour week of searching and had not ended. They ran out of time.
None of this is a story about bad software. It is a story about what happens when a large body of natural conversation is put behind a search box, and it applies with full force to a folder of meeting notes written by different people describing the same thing in whatever words they reached for that day.
The failure gets quieter as the archive gets bigger
Early on, a thin archive fails loudly. Somebody searches for a decision from March, gets nothing back, shrugs, and walks over to ask a colleague. The failure is obvious in the moment and it costs one interruption.
A large archive fails politely. The same search returns fourteen results. Several of them look plausible. The person reads two, finds something that answers the question well enough, and stops. They walk away with an answer and no reason to doubt it. Whether a better answer was sitting at position nine is not knowable from where they stand, and nothing in the interface suggests they should keep going.
That is exactly the behaviour Blair and Maron measured. Their searchers were experienced, motivated, working on a case they cared about, allowed unlimited interaction and as many query revisions as they wanted, and they still stopped at 20 percent while believing they were past 75. The stopping point was set by the results feeling sufficient, and feeling sufficient is not a measurement.
So the volume problem does not announce itself. Nobody files a ticket saying search has degraded. The archive simply gets quieter and less reliable at the same time, which is the worst combination available.
The profession built around keeping records keeps almost none of them
There is a whole discipline whose entire purpose is preserving records, staffed by people who take that purpose seriously, and it is worth knowing what they actually do. The United States National Archives states it plainly: of all the documents and materials created in the course of federal government business, only 1 to 3 percent are important enough for legal or historical reasons to be kept forever.
The other 97 percent or more are not lost through neglect. They are examined and let go on purpose, through a named process called appraisal, which the agency defines as determining the value and therefore the final disposition of records, making them either temporary or permanent. Its own policy makes the distinction sharper than most teams ever do: not all records that constitute essential evidence possess archival value. Being real, and being evidence, is not the same as being worth keeping in the permanent record.
The criteria are written down. Records are appraised as permanent when they document the rights of citizens, the actions of federal officials, or the national experience. Below that sit the working questions the appraisers use, and several of them translate to a meeting archive with almost no adjustment. Is the information unique. Do the records document decisions that set precedents. What is the volume of records. Is sampling an appropriate tool. On voluminous project files specifically, the policy says a very strong justification is needed to appraise all of them as permanent.
What the archivists never do is capture everything and rely on a search interface to make the collection usable later. They decide, on stated criteria, applied consistently, accepting that any given call may turn out wrong. That is a different posture from the one most teams have adopted by default, and the discipline that has been at this the longest arrived at it deliberately.
This is a separate question from how long to keep something, and a separate question from who is allowed to read it. Retention is about expiry. Access control is about boundaries between colleagues. Appraisal is about whether a thing joins the searchable record in the first place, and it is the one of the three that almost nobody decides on purpose.
The meetings that add volume without adding a record
Some meetings reliably produce material that will dilute a search without ever answering one. They are easier to name than most teams expect.
Meetings whose durable output already lives somewhere better shaped. A standup's real record is the board. A sprint planning session's real record is the sprint. A note recapping either one is a lower fidelity copy of something people already trust more, and it competes with the original in every search that touches the project.
Meetings that were a document read aloud. If the artifact existed before the meeting and did not change during it, the artifact is the record. A transcript of somebody narrating a spreadsheet adds pages and no information.
Rehearsals. A practice run of a board update or a customer presentation is worth capturing only if the argument moved. Otherwise you have two versions of the same content in the index, one of which is the wrong one.
Recurring series where each instance closely resembles the last. This is where near duplicates come from, and near duplicates cost more to rule out than irrelevant results do, because a searcher has to read into each one to discover it is the wrong week. A weekly series running for two years is a hundred documents about the same subject, and the only thing distinguishing them is buried in the middle.
Meetings that were called to make a decision and did not make one. These are worth one line stating what was ruled out and what is still open. They are rarely worth a full record, because the useful content is a sentence and the rest is the arguing that produced it.
Tests worth running before a meeting joins the archive
The question test. Name a question somebody might ask six months from now that this note would answer. If nobody at the table can state one, the meeting is producing volume rather than record. This takes about five seconds and it is the single most useful filter available.
The duplicate test. Is the durable output of this meeting already written down somewhere the team trusts more than a meeting note. If so, keep the pointer and skip the note.
The distinctiveness test. If ten results from this recurring series came back in one search, would anybody be able to tell which one they wanted without opening all ten. If the answer is no, the series should be sampled rather than captured whole.
The precedent test, borrowed directly from the appraisers. Did this meeting set something that a person will be held to later. A commitment, a rule, a boundary, a number somebody will quote back. Precedent is the highest signal a meeting can carry and it is worth capturing even when the meeting was otherwise dull.
The uniqueness test, also borrowed. Is the information here available anywhere else. A customer saying something surprising on a call exists in exactly one place. A status update repeating the roadmap exists in five.
How Driffle fits, and what it does not solve
Driffle transcribes your computer's audio directly, so no bot joins the call and no recording announcement goes out to guests. Audio is transcribed in real time and discarded, and what stays is the text and the notes. Notes are private to their author until they are explicitly shared, workspace admins see only what has been shared into shared folders, individual notes can be removed immediately, and full account deletion clears the data within 30 days including backups. Across what you do keep, you can ask questions of past meetings and pull out action items with owners.
Three honest limits. The product does not decide which meetings matter. That call needs somebody who was in the room and knows what the meeting was for, and no vendor can make it on a team's behalf.
Capturing less does not repair a badly framed question, and it does not fix vocabulary drift. A small archive is still hard to search when four people describe the same decision four different ways. Naming things consistently in the notes themselves does more for retrieval than any amount of pruning.
Keeping a meeting out of the shared searchable record is not the same as destroying it. A private note that never enters the team index still exists for the person who wrote it. That makes this a shaping decision about the shared archive rather than a deletion policy, and deletion is a different question with its own answers.
A test you can run this week
Pick three questions you are confident your archive should be able to answer. Real ones, the kind that came up recently. Then search for each and write down two numbers: how many results came back, and how many you actually opened before you felt you had the answer.
The gap between those two numbers is the whole argument. If nine results came back and you read two, you made the same call the lawyers in the study made, on the same basis, with the same amount of evidence about whether you were right.
Then look at what came back and ask how many of those results were from meeting types you could have left out entirely. If a good share of every result set is standups, rehearsals and status readouts, the archive is charging the whole team a search tax to store material nobody has ever needed.
That is the point at which deciding what not to capture stops being a nice idea and starts being the cheapest available improvement to the archive you already have.
Sources
- An Evaluation of Retrieval Effectiveness for a Full-Text Document-Retrieval System - David C. Blair and M. E. Maron, Communications of the ACM, volume 28 number 3 (1985)
- Appraisal Policy of the National Archives and Records Administration - United States National Archives and Records Administration
- What is the National Archives and Records Administration? - United States National Archives and Records Administration
FAQ
Does this mean we should stop recording meetings?
No. Keep the meetings that carry a decision, a commitment, or information that exists nowhere else. The argument here is against capture as an unexamined default, where every meeting on the calendar joins the record because nothing stopped it. A team that captures half as many meetings and captures the right half has a better archive, not a smaller one.
How is this different from a retention policy?
A retention policy decides how long something stays once it is in the record. Appraisal decides whether it enters the record at all. Both are worth having and they answer different questions. A team can have a well designed retention schedule and still be diluting its own search every week, because everything gets captured first and expires much later.
What if we cannot tell in advance which meetings will matter?
Nobody can, item by item. That is precisely why the archival profession works from written criteria rather than prediction. Criteria applied consistently across hundreds of meetings beat individual guesses, they can be explained to the team, and they can be revised when they turn out to be wrong. Perfect foresight was never on the table. The realistic alternative is capturing everything, and that carries its own well measured cost.
Will better search technology solve this instead?
Better ranking helps with precision, which was already the healthy number in the study. The problem Blair and Maron measured is recall, and their mechanism sits mostly outside the software: people do not use the words the record used, and they stop searching when the results feel sufficient. Both of those get harder as the collection grows, whatever is doing the ranking.
Should we delete meetings we have already captured?
Not as a first move. A bulk delete is its own decision with its own risks, and some of what looks like clutter is the only copy of something. Start by narrowing what enters the shared searchable index from here on, then look at whether particular recurring series are worth thinning. Changing the inflow is reversible and cheap. Deleting the back catalogue is neither.