Speaker Identification: audio playback, parallel processing, and batch workflow
Speaker Identification: unreliable audio playback, parallel processing, batch workflow, and output file handling
UPDATE — September 15, 2026: Added an additional issue regarding automatic saving/exporting of transcriptions after Speaker Identification, including the ability to keep the original file or replace it automatically.
Description
I am experiencing several problems with the Speaker Identification workflow in Buzz. The issues are related to audio playback, running Speaker Identification on multiple transcriptions, the way multiple Speaker Identification sessions interact with each other, and how the resulting transcription with speaker names is saved/exported.
These problems make it difficult to use Speaker Identification efficiently when working with multiple completed transcriptions.
1. Audio playback does not work reliably in Speaker Identification
After Step 1 — Identify speakers finishes, Step 2 — Name speakers provides the option to play a sample for each detected speaker.
However, the Play sample function does not reliably produce any sound.
The important point is that audio playback itself works correctly:
- Audio plays normally in the standard Buzz transcription viewer.
- YouTube works normally.
- Other applications play audio normally.
- Windows audio works normally.
- There are no general sound-card or hardware problems that I am aware of.
Therefore, this does not appear to be a general Windows audio, sound card, or hardware problem.
The problem appears to be specific to Buzz, or possibly to an interaction between Buzz and the Windows audio system. I cannot determine the exact cause.
Bluetooth behavior
I have also noticed inconsistent behavior when using Bluetooth headphones.
Sometimes connecting Bluetooth headphones makes Speaker Identification audio playback start working.
Other times, I have to connect and disconnect the Bluetooth headphones several times before audio starts working.
Even when changing the audio output device makes the audio work, an already-open Speaker Identification session does not reliably start using the new output device. In some cases I need to restart or run Speaker Identification again before the samples can be played.
Expected behavior
The Play sample button should reliably play the corresponding speaker segment using the currently active Windows audio output device.
Changing the Windows audio output device should ideally be handled correctly by an already-running Speaker Identification session without requiring the entire process to be restarted.
2. Running Speaker Identification for multiple transcriptions
I frequently have multiple completed transcriptions that need Speaker Identification.
My desired workflow is to open several completed transcriptions and run Speaker Identification on them independently.
Currently, when I start Speaker Identification for multiple transcriptions, several problematic things can happen:
- Buzz may crash and close completely.
- Multiple Speaker Identification windows may close at the same time.
- Sometimes several identifications do complete successfully, but the sessions do not appear to behave independently.
This makes it difficult to process several transcriptions in parallel.
3. Saving one Speaker Identification session can close the others
Another problem occurs when multiple Speaker Identification sessions are open.
For example:
- I start Speaker Identification for transcription A.
- I start Speaker Identification for transcription B.
- I start Speaker Identification for transcription C.
- Step 1 finishes for all or several of them.
- I start naming speakers in transcription A.
- I save the changes.
Saving transcription A can cause the other Speaker Identification windows to close as well.
When this happens, I have to start Speaker Identification again from the beginning for transcription B, C, etc.
This is particularly problematic because Step 1 can take a significant amount of time.
Expected behavior
Each Speaker Identification session should be independent.
Saving or closing the Speaker Identification session for one transcription should not close, reset, or otherwise interfere with Speaker Identification sessions belonging to other transcriptions.
4. Speaker Identification results are not automatically saved as an output transcription
There is also a separate workflow issue after Speaker Identification has been completed.
The typical workflow is:
- Create a transcription from an audio file.
- Run Speaker Identification on the resulting transcription.
- Assign names to the detected speakers.
- Save the Speaker Identification changes.
At this point, I would expect the resulting transcription — now containing the identified speaker names — to be available automatically as the updated output file.
Instead, I still need to manually export the transcription to create/save the resulting file.
This becomes particularly cumbersome when working with many transcriptions because I have to repeatedly export the result manually, usually to the same folder where the original transcription is located.
Desired behavior
There should be an option controlling what happens to the original transcription after Speaker Identification is saved.
For example:
Option A — Keep both files
- Keep the original transcription unchanged.
- Automatically create a new output file containing the identified speaker names.
- Save the new file automatically in the same folder as the original transcription.
For example:
meeting.txt
meeting_with_speakers.txtOption B — Replace the original file
- Automatically update/replace the original transcription with the version containing the identified speaker names.
- No additional manual export should be required.
There could also be an appropriate naming/configuration option for the generated file.
The important point is that the user should not have to manually export the transcription after every Speaker Identification session.
Expected behavior
After saving Speaker Identification results, Buzz should automatically save the resulting transcription according to the user's selected preference:
- Keep original + create a new speaker-identified file, or
- Replace/update the original transcription file.
This should happen automatically without requiring the user to manually export the file and navigate back to the original transcription's folder every time.
This would be especially important for batch processing multiple transcriptions.
Desired workflow
Ideally, I would like to be able to do the following:
- Open several already-completed transcriptions.
- Start Speaker Identification for several of them.
- Let Step 1 — Identify speakers process each transcription independently and, where system resources allow, concurrently.
- Once Step 1 is complete, keep all Speaker Identification sessions open.
- Use Step 2 — Name speakers independently for each transcription.
- Play speaker samples to identify the actual people speaking.
- Assign names and save one transcription.
- Continue working on the other Speaker Identification sessions without them being closed or reset.
- Have each completed Speaker Identification result automatically saved according to the user's preferred output-file setting.
The important requirement is that each Speaker Identification session should be isolated from the others.
Ideal solution: batch Speaker Identification from the main window
A very useful addition would be the ability to select multiple completed transcriptions directly in the main Buzz window and start Speaker Identification for all of them.
For example:
- Select multiple completed transcriptions.
- Right-click them or use a toolbar action such as Identify Speakers.
- Buzz starts a separate Speaker Identification session for each selected transcription.
- Each session processes its own transcription independently.
- Processing can happen concurrently where resources allow.
- Each session remains open after Step 1 finishes.
- Saving one session does not affect the others.
- Each completed result is automatically saved according to the selected output-file behavior.
The existing window title / transcription name should make it possible to distinguish the sessions, especially when the transcription filenames are descriptive.
This would make Speaker Identification much more practical for workflows involving many recordings.
Speaker Identification with more than 3–4 speakers
I have also noticed that Speaker Identification becomes significantly less reliable when there are more than approximately 3–4 speakers.
This is less critical than the problems described above, and I understand that this may partly be a limitation of the currently selected Speaker Identification model.
I am mentioning it as an additional observation rather than necessarily considering it a Buzz bug.
Expected behavior
In summary, I would expect:
Audio
- Speaker samples should play reliably.
- Playback should use the currently active Windows audio output.
- Changing between speakers/headphones should not require restarting Speaker Identification.
- Bluetooth audio output should work consistently.
Multiple sessions
- Multiple Speaker Identification sessions should be able to exist independently.
- One session should not close or reset another session.
- Buzz should not crash simply because Speaker Identification is running for multiple completed transcriptions.
- Multiple identifications should be able to run concurrently where system resources allow.
Saving
- Saving one Speaker Identification session should only affect that transcription.
- Other active Speaker Identification sessions should remain open and retain their current progress.
- Speaker Identification results should be saved automatically without requiring a manual export.
Output files
- The user should be able to choose whether the original transcription is preserved.
- The user should be able to choose whether the original transcription is replaced/updated.
- When keeping both files, the new speaker-identified transcription should be created automatically in the appropriate folder.
- Manual export should not be required after every Speaker Identification session.
Batch workflow
- Multiple completed transcriptions should ideally be selectable in the main Buzz window.
- A single Identify Speakers action should be able to launch Speaker Identification for all selected transcriptions.
- Each transcription should then have its own independent Speaker Identification workflow.
Why this matters
Speaker Identification is particularly useful when processing many recordings involving the same or different groups of people.
At the moment, the combination of unreliable sample playback, interference between multiple Speaker Identification sessions, and the need to manually export every completed result makes this workflow unnecessarily difficult.
The biggest issue is that Step 1 can take a significant amount of time, so having other sessions close or having to restart them from the beginning after saving one transcription results in a substantial amount of duplicated processing.
Similarly, manually exporting every completed transcription becomes increasingly cumbersome when processing a larger number of files.
An independent, multi-file Speaker Identification workflow with automatic output handling would make the feature much more useful for larger transcription workloads.
Environment
- OS: Windows 10
- Buzz version: 1.4.5
- Previous Buzz version tested: 1.4.4
- Audio output: Built-in Windows audio and Bluetooth headphones
Version history / regression
The same problems were present in Buzz 1.4.4 and are still present in Buzz 1.4.5.
I originally encountered these issues while using version 1.4.4. I have since updated Buzz to 1.4.5, but the same behavior continues to occur.
Therefore, this does not appear to be a problem introduced specifically by the 1.4.5 update. The behavior is reproducible across both versions.
The problems described above are therefore present in:
- Buzz 1.4.4
- Buzz 1.4.5 (current version)
Summary
The main problems I am reporting are:
- Speaker sample playback in Step 2 is unreliable, even though audio works normally elsewhere in Buzz and in Windows.
- Bluetooth output changes can sometimes make playback work, but inconsistently, and an already-running Speaker Identification session may need to be restarted.
- Running Speaker Identification for multiple transcriptions can cause crashes or multiple sessions to close.
- Saving one Speaker Identification session can close other active sessions.
- There is no convenient batch workflow for launching Speaker Identification for multiple completed transcriptions.
- Speaker Identification results require manual export instead of being automatically saved as an output transcription.
- There is no convenient choice between keeping the original transcription and automatically replacing/updating it with the speaker-identified version.
- Speaker Identification becomes less reliable with more than approximately 3–4 speakers, although this may be related to model limitations.
The desired solution is for Speaker Identification sessions to be fully independent, with reliable audio playback, a practical multi-file/batch workflow, and automatic handling of the resulting transcription files.
Source: chidiwilliams/buzz