Transcribing Interviews, Meetings and Lectures — With Speaker Labels

Multi-voice audio is where transcription earns its keep, and where it usually falls apart. A podcast, a panel, a research interview, a two-hour meeting — transcribe any of them without marking who is speaking, and you get a wall of words that's technically complete and practically useless. Was that the customer's objection or the salesperson's summary of it? The witness's account or the lawyer's rephrasing? The transcript knows; it just won't tell you.

Speaker identification fixes this at the source. Inwista separates the voices during transcription, and you attach real names once in the editor. The result reads the way the conversation actually happened: as labelled dialogue.

Here's the full workflow, the situations where it changes the game, and the practical habits that make the labels accurate.

How it works, end to end

1. Switch it on at upload

Start a + New Project and, among the upload settings, enable Speaker diarisation. You can specify how many speakers the recording contains — or, our standing recommendation, let Inwista detect the number automatically.

2. Transcription separates the voices

Inwista transcribes and splits the conversation into turns — each line attributed to Speaker 1, Speaker 2, and so on. Fast, too: expect the processing to take roughly one-fifth of the recording's length.

3. Name them in the editor

Open the file in the editor, where the speaker tools live:

  • The speaker menu — see, add, rename or delete speakers. Type "Speaker 1 → Anna Berg" once and the label updates throughout the transcript.
  • Per-line reassignment — where two voices overlap or the detection guessed wrong, assign the correct speaker to that specific line.
  • Bulk rename — change every speaker name in one operation when a sweeping adjustment is needed.

4. Export — with one checkbox that matters

Download as TXT, DOCX or PDF — and in the download window, tick "Include speakers." That checkbox is what carries the labels into the finished file; it's the single most common thing to miss in this workflow, so consider this paragraph your pre-emptive support ticket.

Where labelled transcripts change the work

Podcasts. A speaker-labelled transcript is the readable, publishable version of the episode — and the backbone of an accessible one. It slots directly into the WCAG-compliant transcript format, where "shows clearly who says what" is requirement number two.

Research interviews and focus groups. Analysis lives and dies on attribution — coding a focus group where you can't tell participants apart isn't analysis, it's archaeology. Labelled turns make quotes citable and patterns traceable to individuals. Pair with low temperature for the verbatim fidelity qualitative work demands.

Meetings. The transcript is the minutes — who raised the risk, who committed to the deadline, who disagreed — without anyone in the room burning their attention on note-taking. Name the participants once and export; the follow-up email writes itself from the record.

Lectures and teaching. Solo lectures transcribe cleanly regardless, but the moment Q&A starts, diarisation earns its place: student questions and lecturer answers stay distinct, which is exactly what a student revising from the transcript needs. And a text version of every lecture is an accessibility win before it's anything else.

Legal and formal records. Where documentation must show precisely who said what, attribution isn't a convenience — it's the point. Worth knowing alongside: recordings are processed on EU servers with AES-256 encryption on every plan, and a Data Processing Agreement is available on Enterprise for organisations that need the paperwork.

Getting accurate labels: two habits

Feed it separable audio. Diarisation distinguishes voices; heavy crosstalk — everyone talking at once — is hard for humans and machines alike. Reasonable mic placement and a chair who lets people finish do more for label accuracy than any setting.

Name speakers early, not late. Attach real names as your first editing act. Every subsequent read-through becomes easier, and misattributed lines jump out in a way "Speaker 3" errors never do.

From recording to record, in minutes

Interviews, meetings, lectures — upload, toggle diarisation, name the speakers, export with labels. Inwista's free plan lets you run the whole workflow on your own material — upload, transcribe, structure, edit and export. Current limits and plan details are on the pricing page.

Transcribe your next conversation →

Frequently asked questions

How many speakers can it handle? You can set the expected count at upload or let Inwista detect it automatically — the option we recommend. Accuracy is strongest when voices take turns; sustained crosstalk degrades any diarisation, ours included, which is why the mic-and-moderation habits above matter.

Does Inwista know the speakers' names? No — and by design. Diarisation separates voices; identity is yours to attach. Rename each detected speaker once in the editor and the label applies across the whole transcript.

A line got attributed to the wrong person — can I fix just that line? Yes. Per-line speaker assignment exists for exactly this: correct the individual line without touching the rest.

My exported transcript has no names in it. What happened? The "Include speakers" checkbox in the download window was almost certainly unticked. Re-export with it checked — the labels are in the project; the checkbox decides whether they travel into the file.

Does this work on video files too? Identically. Diarisation runs on the audio track, so a filmed panel or a video meeting recording behaves exactly like an audio file.

We record sensitive conversations — HR cases, client meetings. Where does the data go? Processing and storage on EU servers with AES-256 encryption, on every plan including Free. You control deletion, and Enterprise customers can put a Data Processing Agreement in place.