Meeting Notes When Half the Room Is Working in a Second Language

Published11 min read

Teams with colleagues spread across countries usually evaluate a notes tool by asking whether it will cope with everyone's accents. That question now has a reasonably good answer, and answering it hides the harder problem. In a meeting where several people are working in a second language, the transcript can be close to perfect and the record can still be missing the most useful thing anyone knew.

The accent question has a better answer than most teams expect

A team with colleagues in four countries almost always asks the same thing first when it looks at a meeting notes tool. Will it actually understand everyone. It is a fair question, and for years the honest answer was a shrug.

The answer has moved. A 2025 evaluation of five current speech recognition systems tested them against the L2-ARCTIC corpus, which covers speakers whose first languages are Arabic, Chinese, Hindi, Korean, Spanish and Vietnamese. On read speech the best systems reached a mean match error rate of 0.054, which the study describes as approaching human level accuracy, and the best result on spontaneous speech was 0.063.

That is one corpus in one study, and read sentences are an easier target than five people talking over each other on a poor connection. The same study found wide variation between systems in how they handled disfluencies, the filler words, repetitions and self-corrections that ordinary speech is full of. The direction is still clear enough to act on. For most teams, recognition is no longer the part that breaks.

Which raises the more uncomfortable question. If the transcript comes back clean and the record still feels thin, the loss happened somewhere an accuracy score cannot see.

An accurate record of a meeting that was missing its best sentence

Tsedal Neeley, Pamela Hinds and Catherine Cramton studied what happened inside several large global companies after each one made English its official working language. The behaviour they documented is not what a transcription vendor measures.

Some employees attended meetings and consciously did not contribute, withdrawing in advance of anything being said. One employee at the French company in the study put it plainly: they understood English perfectly well, having trained in it, but might say nothing or stop talking during meetings with US counterparts out of fear of looking silly and making mistakes.

The researchers describe one engineer at that company who found a technical issue during a team discussion and did not raise it, because building the argument in English was more than he wanted to take on in the moment. When the problem surfaced later it took the project team two weeks to resolve and pushed out a timeline that was already late.

Capture software would have handled that meeting flawlessly. It would have produced an accurate record of a conversation that was missing the single most valuable thing anyone in the room knew. Capture quality has a ceiling, and the ceiling is whatever made it into the room.

Fluency was not the binding constraint

The same research reports something worth sitting with. Many of the people who held back were quite fluent. What they lacked was confidence, and the anxiety that followed shaped how they behaved in meetings. Across fluency levels, almost all of the second-language speakers in the study described a feeling of diminished professional standing once English became the company language.

This changes how a meeting record should be read. If you take a person's silence as evidence they had nothing to add, you are reading a signal produced by exposure and treating it as a signal about content.

There is a structural asymmetry underneath it. Agreeing costs one word. Disagreeing costs a built argument, delivered live, at conversational speed, in front of colleagues. When those two acts cost such different amounts, a record that faithfully logs an agreement is faithfully logging the cheaper act.

What a summary quietly removes

This is the part of the story that cuts against the whole category, including this one, so it is worth stating directly.

At the German company in the study, colleagues who switched into German partway through a meeting were conscientious about it. They apologised for the switch and almost always summarised the exchange afterwards for the people who had been left out. The people receiving those summaries said the summaries were finished thoughts, and that they had been excluded from the shaping and questioning of the ideas.

A summary hands over a conclusion and drops the argument that produced it. For a colleague who was in the meeting but could not get into the discussion at full speed, a decision-only recap does the same thing to them a second time, in writing.

The practical consequence is about what you keep. In a mixed-language room the reasoning and the unresolved questions carry most of the value, and the decision line on its own is the least useful part of the record.

Move what you can out of real time

Yongle Zhang, Dennis Asamoah Owusu, Marine Carpuat and Ge Gao ran a controlled experiment on this. Twenty participants in five groups of four, each group made up of two native English speakers and two native Mandarin speakers. Each subgroup talked first in its own language, then the full group met in English.

When machine-translated logs of those subgroup conversations were shared before the meeting, meeting quality improved. The improvement showed up in how participants experienced the meeting, in task performance, and in the depth of discussion.

It is a small study on a single language pair, so treat it as a direction rather than a number. The direction is useful: the meeting got better because context arrived in advance and in writing, before anyone had to keep up with a live conversation.

You control how much has to be understood live. If prior decisions, the current state of a project, and the agenda are written down and searchable before the call, less of the meeting is spent reconstructing context at speaking speed, which is exactly the condition that pushes people into silence.

Hand the airtime over on purpose

Neeley and her co-authors are direct about what first-language speakers can do. Work to make what is being said clearer during the meeting, and make sure second-language colleagues get equal or greater airtime.

The mechanism is worth naming. The most fluent person in a meeting is usually also the fastest, and the fastest person sets the tempo at which everyone else has to decide whether to interrupt. One or two people can slow that tempo without anyone's permission and without announcing it.

Asking for input by name beats opening the floor. An open invitation goes to whoever claims it quickest, which is a competition decided by fluency before it is decided by who has the relevant information.

The errors that survive are the ones that matter

Where recognition does slip, it slips hardest on the words that have the least surrounding context to recover them. Names of people, product and company names, internal acronyms, and figures. Ordinary words sit in a sentence that makes them predictable. A colleague's name and an internal codename do not.

That should change how the record gets reviewed. Nobody proofreads a transcript, and asking a team to do it is how a good habit dies in a fortnight. Check the action items, the owners, and any number that will be repeated to someone else. A mangled filler word costs nothing. A wrong product name inside an action item travels.

Where a tool lets you correct a name once and have it stick, that is worth doing early, because the same names recur every week.

Never let the record become a language assessment

A record of a mixed-language meeting is a record of people working in a second language under time pressure, and it will look like one. Hesitation, restarts, and sentences trimmed back to what the speaker was confident they could finish.

Set the boundary explicitly, and say it out loud to the team. The record holds decisions, reasoning and commitments. It is never quoted to characterise how well somebody speaks, and it never appears in a performance conversation as evidence of communication skill.

This is not a courtesy. The research above found that fear of being judged on their English is what kept information out of the meeting in the first place. A record that gets used as proof of fluency makes that fear rational, and the team pays for it in the information it stops receiving.

Where quiet capture helps, and where it does not

Driffle transcribes the computer's audio directly, so no bot joins the call and no notification goes out to anyone. In a room where several people are already worried about how they sound, a visible recorder announcing itself is not a neutral addition. Nothing auto-joins, auto-records, or runs in the background. Notes are private to their author until they are explicitly shared, workspace admins see only what has been shared into shared folders, audio is transcribed in real time and discarded, individual notes are removed immediately on request, and full account deletion completes within 30 days, backups included.

Quiet capture removes none of the responsibility to be clear with the people in the meeting about what is being kept. None of this is a certification, and none of it is legal advice.

The honest limits are worth stating too. A record cannot recover a point that was never made. It does not change who speaks or how fast. Translating a hedge is a judgement call, and a hedge is often the most informative thing a cautious speaker says. What a record can do is reduce how much has to be absorbed live, hold on to the reasoning instead of only the verdict, and give people a written channel afterwards where the clock is not running against them.

A four week check that tells you where you actually stand

Run this for four weeks on every meeting where more than one first language is in the room. Keep two lists. First, roughly who spoke and for how long. Second, every substantive point that reached you in writing afterwards and had not come up in the meeting itself.

The second list is the measurement that matters. A short one means the meeting is receiving what the team knows. A long one means the information exists and the meeting is not the place it arrives, and no improvement in transcription accuracy will change that.

If the second list is long, the work sits in how the meeting runs and in what people can read before and write after, rather than in the quality of the capture. If it is short and the record still feels thin, then the record itself is the thing to fix, and that is a much easier problem.

Sources

FAQ

Do AI meeting notes handle non-native accents accurately enough to rely on?

Much better than they used to. A 2025 evaluation of five current speech recognition systems against the L2-ARCTIC corpus, covering speakers with six different first languages, found the best systems reached a mean match error rate of 0.054 on read speech, which the study describes as approaching human level, and 0.063 on spontaneous speech. Two caveats apply. That is one corpus in one study, and the same work found significant variation between systems in how they cope with filler words, repetitions and self-corrections, which is what real meeting speech is made of. Test on a genuinely messy meeting rather than a clean one before deciding.

Should we run meetings in one shared language or let people use their own and translate?

There is no general answer, but there is a useful finding. In a controlled experiment with twenty participants split into groups of two native English and two native Mandarin speakers, sharing machine-translated logs of each subgroup's own-language conversation before the joint English meeting improved meeting quality, task performance and depth of discussion. It is a small study on one language pair. What it suggests is that the change worth making first is often supplying context in writing beforehand rather than changing what happens during the meeting itself.

Will an AI summary help a colleague who could not follow the discussion live?

Partly, and less than you would hope if the summary only carries decisions. Research on teams that switched languages mid-meeting found that the courteous catch-up summaries afterwards were experienced as finished thoughts that excluded people from shaping and questioning the ideas. For a mixed-language team, keep the reasoning and the open questions in the record, not only the outcome, and leave a clear route for someone to reopen a point in writing after the fact.

Is it fair to look at the transcript to see who contributed?

Use it to notice a pattern in your own meetings, never to evaluate a person. Field research on companies that mandated English found that fluent employees withheld contributions because they feared being judged on their language, and that this fear was what kept useful information out of meetings. A record used as evidence of communication ability confirms that fear and costs the team the information it was trying to capture. State the boundary explicitly so people are not guessing about it.

What is worth correcting in a transcript from a multilingual meeting?

Names, internal terms, and numbers. Recognition errors land hardest where surrounding context cannot rescue the word, which means people's names, product and company names, acronyms and figures rather than ordinary vocabulary. Reviewing a whole transcript is a habit no team keeps. Reviewing the action items, their owners, and any figure that will be passed on to someone else takes a minute and catches the errors that actually propagate.

Never lose the thread of a meeting again.

Driffle keeps the decisions, owners, and context from every conversation searchable when work resumes.

Request access

Similar articles