Jump to content

Buzz Transcription: Difference between revisions

From MediawikiCIT
Justinaquino (talk | contribs)
Add Buzz meeting transcription guide + Vulkan GPU & OOM investigation (2026-04-25)
 
Justinaquino (talk | contribs)
Mirror from wiki.gi7b.org rev 4963 (2026-08-09): Folder Watch warning, Taglish example, screenshots
Line 156: Line 156:


''Investigation date: 2026-04-25. System: Ubuntu 24.04, AMD RX 7600, Buzz 1.4.4 (Flatpak).''
''Investigation date: 2026-04-25. System: Ubuntu 24.04, AMD RX 7600, Buzz 1.4.4 (Flatpak).''
== Folder Watch (Automatic Transcription) ==
<div style="border:1px solid #a00;background:#fee;padding:0.5em 1em;">'''Do not use Folder Watch.''' It is disruptive at default settings: every audio file dropped into the watched folder is transcribed automatically without asking, silently occupying the GPU/CPU and flooding the folder with output files — and with ''Delete processed files'' ticked, the source audio is deleted afterwards. It is disabled by default and should stay that way. Use [[#File Transcription (Meeting Recording)|File → Import Media]] instead. The settings below are documented for reference only (as configured during testing on 2026-08-09).</div>
[[File:Buzz Folder Watch tab.png|thumb|400px|right|Folder Watch tab: whisper.cpp backend, Large-V3, advanced options unchecked]]
[[File:Buzz Folder Watch tab full.png|thumb|400px|right|Folder Watch tab — Enable folder watch with input/output folders set to /home/justin/Music]]
=== Settings ===
In '''Preferences → Folder Watch''':
{| class="wikitable"
|-
! Setting !! Value !! Notes
|-
| Enable folder watch || '''leave unticked''' || ticking it starts auto-transcribing everything dropped into the input folder
|-
| Input folder || <code>/home/justin/Music</code> || drop recordings here
|-
| Output folder || <code>/home/justin/Music</code> || same folder — files stay put
|-
| Delete processed files || '''unchecked''' || '''ticking this deletes the source audio after transcription''' — keep it unticked
|-
| Model || whisper.cpp — Large-V3 || engine choice matters, see below
|-
| Task || Transcribe ||
|-
| Language || Detect Language ||
|-
| Export || TXT || SRT / VTT also available
|-
| Word-level timings || '''unchecked''' ||
|-
| Extract speech || '''unchecked''' || '''do not check — crashes on long files''' (see [[#OOM Crash with Large Files|OOM Crash with Large Files]])
|}
=== Use whisper.cpp instead of the default engine ===
[[File:Buzz model dropdown.png|thumb|400px|right|Model dropdown: whisper.cpp and Faster Whisper]]
The model picker lists two engines: '''whisper.cpp''' and '''Faster Whisper'''. Faster Whisper is the default — switch to '''whisper.cpp''':
* whisper.cpp uses the <code>ggml-*.bin</code> models already cached under <code>~/.var/app/io.github.chidiwilliams.Buzz/cache/Buzz/models/</code> (Flatpak path; <code>~/snap/buzz/…</code> on snap installs) — no extra download
* it runs through the bundled Vulkan <code>whisper-cli</code>, so GPU acceleration applies (see [[#GPU Acceleration|GPU Acceleration]])
[[File:Buzz Models tab whisper.cpp.png|thumb|400px|right|Models tab, whisper.cpp group: Large-V3, Large-V3-Turbo, Medium.en, Large and Large-v2 downloaded]]
In the Models tab, the whisper.cpp group shows the downloaded sizes — Large-V3 and Large-V3-Turbo are the useful ones for meetings (Large-V3-Turbo: best quality/speed trade-off, see [[#How to Transcribe a Meeting|Recommended Settings]]).
=== Leave the advanced options unchecked ===
[[File:Buzz transcription dialog unchecked.png|thumb|400px|right|Transcription dialog: neither Word-level timings nor Extract speech checked]]
In the transcription dialog (shown when folder watch picks up a file) both '''Word-level timings''' and '''Extract speech''' are left unchecked. Checking '''Extract speech''' — the obvious option for a meeting recording — triggers the Demucs pre-processing step that OOM-crashes the app on long files (22 GB RSS kill, [[#OOM Crash with Large Files|see above]]). Keep both unchecked.
=== Result ===
Transcriptions are written next to the source as <code>&lt;original name&gt; (transcribed on &lt;date&gt;).txt</code>, e.g. <code>2026-08-08 20-56-20 3rd Apocalypse Bob's ADND1E (transcribed on 09-Aug-2026 01-50-32).txt</code>.
Config verified 2026-08-09 in <code>~/snap/buzz/570/.config/Buzz.conf</code> (snap install): <code>folder_watch\enabled</code> = <code>false</code> (off — recommended), <code>folder_watch\input_folder</code> = <code>/home/justin/Music</code>, <code>folder_watch\output_directory</code> = <code>/home/justin/Music</code>, model Whisper.cpp / large-v3, and <code>delete_processed_files</code>, <code>extract_speech</code>, <code>word_level_timings</code> all <code>false</code>.
== Example: Taglish Test (2026-08-09) ==
A Taglish (Tagalog + English code-switching) test recording, transcribed automatically by Buzz's Folder Watch during testing (now disabled — see the [[#Folder Watch (Automatic Transcription)|Folder Watch]] warning above):
* '''Audio:''' [[:File:Taglish Test 260809.mp3|Taglish Test 260809.mp3]]
* '''Evaluation report:''' [[:File:Taglish Transcription and Evaluation Report 260809.docx|Taglish Transcription and Evaluation Report 260809.docx]]
The report grades the result '''95/100 (Excellent)''' at ~96 % word accuracy. Code-switching between Filipino and English (EDSA, client, project, presentation, brain cells, lunch, work) is handled seamlessly. Noted deviations are typical ASR behavior: punctuation is collapsed into sentence breaks, "Almost two hours" became "Almost 2 hours", "na-stuck" → "nakastuck", "kailangan nating" → "kailan nating".

Revision as of 06:10, 9 August 2026

Buzz is an open-source, offline-capable audio/meeting transcription tool powered by OpenAI Whisper (via whisper.cpp). On this system it is installed as a Flatpak and runs with Vulkan GPU acceleration on the AMD RX 7600.

How to Transcribe a Meeting

Prerequisites

  • Buzz installed (Flatpak: io.github.chidiwilliams.Buzz)
  • A model downloaded — recommended: ggml-large-v3-turbo for quality, ggml-tiny for speed
  • Models are stored at ~/.var/app/io.github.chidiwilliams.Buzz/cache/Buzz/models/

File Transcription (Meeting Recording)

  1. Launch Buzz from the application menu or run: flatpak run io.github.chidiwilliams.Buzz
  2. Go to File → Import Media and select your meeting recording (MP3, WAV, M4A, etc.)
  3. In the transcription dialog:
    • Model: Select large-v3-turbo (best quality for meetings)
    • Task: Transcribe
    • Language: Select your language or leave on Auto
    • Extract Speech: Leave unchecked for files longer than ~30 minutes (see OOM Crash below)
    • Output: SRT, VTT, or TXT depending on your need
  4. Click Transcribe
  5. When complete, the transcript appears in the main window and is saved alongside the source file

Live Transcription (Real-Time)

  1. Go to File → Record and Transcribe
  2. Select your microphone input
  3. Choose model and language
  4. Click Record — transcription appears live on screen
  5. Stop recording when done; export via File → Export
Setting Value Reason
Model large-v3-turbo Best accuracy; 1.6 GB VRAM, runs on RX 7600
Extract Speech Off for files >30 min Demucs uses 8–15 GB RAM for long files (OOM risk)
Language Set explicitly Faster than auto-detect
Task Transcribe Use Translate only if you need English output from another language

Before Each Session: Clear Swap

Prior transcription sessions can leave stale pages in swap. Run this before starting a long transcription:

sudo swapoff -a && sudo swapon -a

This is safe when free RAM exceeds swap used (typical on this system: 15+ GB free, 8 GB swap).

GPU Acceleration

Buzz's Flatpak installation includes a Vulkan-enabled whisper-cli. On this system, GPU inference is automatically active — no configuration needed.

  • GPU in use: AMD RX 7600 (RDNA3, RADV NAVI33)
  • Confirmed by: whisper_backend_init_gpu: using Vulkan0 backend in runtime output
  • Buzz checks for Vulkan at startup (IS_VULKAN_SUPPORTED) and omits the --no-gpu flag when found

To force CPU inference (for debugging):

flatpak override --user io.github.chidiwilliams.Buzz --env=BUZZ_FORCE_CPU=true

To revert:

flatpak override --user --reset io.github.chidiwilliams.Buzz

OOM Crash with Large Files

Symptom

When selecting ggml-large-v3-turbo in Buzz and starting a transcription on a long file, the application crashes. VRAM usage never visibly climbs before the crash.

Root Cause

The crash is caused by the Extract Speech (Demucs) pre-processing step, not the large model.

Demucs is a PyTorch-based music source separation model that strips background noise from audio before passing it to the transcription engine. It processes raw PCM audio at full float32 precision entirely in RAM:

Audio Duration Compressed Size RAM Required (Demucs)
1 hour MP3 ~60 MB 500 MB – 1 GB
3–4 hour session ~220 MB 8–15 GB peak

The combination that caused the crash:

  1. Swap already exhausted from previous sessions
  2. Extract Speech (Demucs) triggered on a ~220 MB MP3 (~3–4 hours of audio)
  3. Python RAM usage exceeded available RAM + swap ceiling
  4. OOM killer terminated the process at 22 GB RSS
  5. whisper-cli was never launched → no VRAM activity observed

Why VRAM Never Climbed

whisper-cli runs as a subprocess of the Python GUI. The OOM kill happened during Demucs pre-processing — whisper-cli was never started, so no VRAM allocation occurred.

From the Buzz log at crash time:

~/.var/app/io.github.chidiwilliams.Buzz/.local/state/Buzz/log/logs.txt

The last entry was Will extract speech — the log never reached Starting whisper file transcription.

Fix

Uncheck "Extract Speech" in the Buzz transcription dialog for any file longer than ~30 minutes. The large-v3-turbo model has built-in noise tolerance that makes Demucs pre-processing unnecessary in most cases.

Optionally, grow the swapfile if you need Extract Speech for shorter files:

sudo swapoff /swap.img
sudo fallocate -l 16G /swap.img
sudo mkswap /swap.img
sudo swapon /swap.img

Environment Variables

Variable Default Effect
BUZZ_FORCE_CPU false Set to true to disable GPU inference
BUZZ_WHISPERCPP_N_THREADS cpu_count / 2 Override thread count for whisper-cli

Set via: flatpak override --user io.github.chidiwilliams.Buzz --env=VAR=value

Key File Paths

Path Description
/var/lib/flatpak/app/io.github.chidiwilliams.Buzz/.../buzz/whisper_cpp/whisper-cli Bundled Vulkan whisper-cli (libwhisper 1.8.3)
~/.var/app/io.github.chidiwilliams.Buzz/cache/Buzz/models/ Buzz model cache (ggml binaries)
~/.var/app/io.github.chidiwilliams.Buzz/.local/state/Buzz/log/logs.txt Buzz application log
~/.var/app/io.github.chidiwilliams.Buzz/config/Buzz.conf Buzz settings
~/whisper.cpp/build/bin/whisper-cli Custom Vulkan-enabled whisper.cpp build
~/bin/whisper-cli-vulkan Wrapper script for custom build (sets LD_LIBRARY_PATH)
~/.local/share/pipx/venvs/buzz-captions/lib/python3.12/site-packages/buzz/whisper_cpp/ pipx install whisper_cpp directory (directly modifiable)

Investigation Summary

Question Finding
Is Buzz using the RX 7600? Yes. using Vulkan0 backend confirmed inside the Flatpak sandbox.
Was --no-gpu being added? No. IS_VULKAN_SUPPORTED = True in the sandbox; flag is never appended.
Does the large model load correctly? Yes. 1.6 GB to VRAM, transcribes in ~1.25s on a short clip.
What caused the crash? Extract Speech (Demucs) OOM. Python killed at 22 GB RSS before whisper-cli launched.
Why was VRAM not climbing? whisper-cli was never started — process died during Demucs pre-processing.
Fix? Uncheck Extract Speech for long files. Clear swap before sessions.

Investigation date: 2026-04-25. System: Ubuntu 24.04, AMD RX 7600, Buzz 1.4.4 (Flatpak).

Folder Watch (Automatic Transcription)

Do not use Folder Watch. It is disruptive at default settings: every audio file dropped into the watched folder is transcribed automatically without asking, silently occupying the GPU/CPU and flooding the folder with output files — and with Delete processed files ticked, the source audio is deleted afterwards. It is disabled by default and should stay that way. Use File → Import Media instead. The settings below are documented for reference only (as configured during testing on 2026-08-09).
Folder Watch tab: whisper.cpp backend, Large-V3, advanced options unchecked
Folder Watch tab — Enable folder watch with input/output folders set to /home/justin/Music

Settings

In Preferences → Folder Watch:

Setting Value Notes
Enable folder watch leave unticked ticking it starts auto-transcribing everything dropped into the input folder
Input folder /home/justin/Music drop recordings here
Output folder /home/justin/Music same folder — files stay put
Delete processed files unchecked ticking this deletes the source audio after transcription — keep it unticked
Model whisper.cpp — Large-V3 engine choice matters, see below
Task Transcribe
Language Detect Language
Export TXT SRT / VTT also available
Word-level timings unchecked
Extract speech unchecked do not check — crashes on long files (see OOM Crash with Large Files)

Use whisper.cpp instead of the default engine

Model dropdown: whisper.cpp and Faster Whisper

The model picker lists two engines: whisper.cpp and Faster Whisper. Faster Whisper is the default — switch to whisper.cpp:

  • whisper.cpp uses the ggml-*.bin models already cached under ~/.var/app/io.github.chidiwilliams.Buzz/cache/Buzz/models/ (Flatpak path; ~/snap/buzz/… on snap installs) — no extra download
  • it runs through the bundled Vulkan whisper-cli, so GPU acceleration applies (see GPU Acceleration)
Models tab, whisper.cpp group: Large-V3, Large-V3-Turbo, Medium.en, Large and Large-v2 downloaded

In the Models tab, the whisper.cpp group shows the downloaded sizes — Large-V3 and Large-V3-Turbo are the useful ones for meetings (Large-V3-Turbo: best quality/speed trade-off, see Recommended Settings).

Leave the advanced options unchecked

Transcription dialog: neither Word-level timings nor Extract speech checked

In the transcription dialog (shown when folder watch picks up a file) both Word-level timings and Extract speech are left unchecked. Checking Extract speech — the obvious option for a meeting recording — triggers the Demucs pre-processing step that OOM-crashes the app on long files (22 GB RSS kill, see above). Keep both unchecked.

Result

Transcriptions are written next to the source as <original name> (transcribed on <date>).txt, e.g. 2026-08-08 20-56-20 3rd Apocalypse Bob's ADND1E (transcribed on 09-Aug-2026 01-50-32).txt.

Config verified 2026-08-09 in ~/snap/buzz/570/.config/Buzz.conf (snap install): folder_watch\enabled = false (off — recommended), folder_watch\input_folder = /home/justin/Music, folder_watch\output_directory = /home/justin/Music, model Whisper.cpp / large-v3, and delete_processed_files, extract_speech, word_level_timings all false.

Example: Taglish Test (2026-08-09)

A Taglish (Tagalog + English code-switching) test recording, transcribed automatically by Buzz's Folder Watch during testing (now disabled — see the Folder Watch warning above):

The report grades the result 95/100 (Excellent) at ~96 % word accuracy. Code-switching between Filipino and English (EDSA, client, project, presentation, brain cells, lunch, work) is handled seamlessly. Noted deviations are typical ASR behavior: punctuation is collapsed into sentence breaks, "Almost two hours" became "Almost 2 hours", "na-stuck" → "nakastuck", "kailangan nating" → "kailan nating".