Buzz Transcription
Buzz is an open-source, offline-capable audio/meeting transcription tool powered by OpenAI Whisper (via whisper.cpp). On this system it is installed as a Flatpak and runs with Vulkan GPU acceleration on the AMD RX 7600.
How to Transcribe a Meeting
Prerequisites
- Buzz installed (Flatpak:
io.github.chidiwilliams.Buzz) - A model downloaded — recommended:
ggml-large-v3-turbofor quality,ggml-tinyfor speed - Models are stored at
~/.var/app/io.github.chidiwilliams.Buzz/cache/Buzz/models/
File Transcription (Meeting Recording)
- Launch Buzz from the application menu or run:
flatpak run io.github.chidiwilliams.Buzz - Go to File → Import Media and select your meeting recording (MP3, WAV, M4A, etc.)
- In the transcription dialog:
- Model: Select
large-v3-turbo(best quality for meetings) - Task: Transcribe
- Language: Select your language or leave on Auto
- Extract Speech: Leave unchecked for files longer than ~30 minutes (see OOM Crash below)
- Output: SRT, VTT, or TXT depending on your need
- Model: Select
- Click Transcribe
- When complete, the transcript appears in the main window and is saved alongside the source file
Live Transcription (Real-Time)
- Go to File → Record and Transcribe
- Select your microphone input
- Choose model and language
- Click Record — transcription appears live on screen
- Stop recording when done; export via File → Export
Recommended Settings for Meetings
| Setting | Value | Reason |
|---|---|---|
| Model | large-v3-turbo | Best accuracy; 1.6 GB VRAM, runs on RX 7600 |
| Extract Speech | Off for files >30 min | Demucs uses 8–15 GB RAM for long files (OOM risk) |
| Language | Set explicitly | Faster than auto-detect |
| Task | Transcribe | Use Translate only if you need English output from another language |
Before Each Session: Clear Swap
Prior transcription sessions can leave stale pages in swap. Run this before starting a long transcription:
sudo swapoff -a && sudo swapon -a
This is safe when free RAM exceeds swap used (typical on this system: 15+ GB free, 8 GB swap).
GPU Acceleration
Buzz's Flatpak installation includes a Vulkan-enabled whisper-cli. On this system, GPU inference is automatically active — no configuration needed.
- GPU in use: AMD RX 7600 (RDNA3, RADV NAVI33)
- Confirmed by:
whisper_backend_init_gpu: using Vulkan0 backendin runtime output - Buzz checks for Vulkan at startup (
IS_VULKAN_SUPPORTED) and omits the--no-gpuflag when found
To force CPU inference (for debugging):
flatpak override --user io.github.chidiwilliams.Buzz --env=BUZZ_FORCE_CPU=true
To revert:
flatpak override --user --reset io.github.chidiwilliams.Buzz
OOM Crash with Large Files
Symptom
When selecting ggml-large-v3-turbo in Buzz and starting a transcription on a long file, the application crashes. VRAM usage never visibly climbs before the crash.
Root Cause
The crash is caused by the Extract Speech (Demucs) pre-processing step, not the large model.
Demucs is a PyTorch-based music source separation model that strips background noise from audio before passing it to the transcription engine. It processes raw PCM audio at full float32 precision entirely in RAM:
| Audio Duration | Compressed Size | RAM Required (Demucs) |
|---|---|---|
| 1 hour MP3 | ~60 MB | 500 MB – 1 GB |
| 3–4 hour session | ~220 MB | 8–15 GB peak |
The combination that caused the crash:
- Swap already exhausted from previous sessions
- Extract Speech (Demucs) triggered on a ~220 MB MP3 (~3–4 hours of audio)
- Python RAM usage exceeded available RAM + swap ceiling
- OOM killer terminated the process at 22 GB RSS
whisper-cliwas never launched → no VRAM activity observed
Why VRAM Never Climbed
whisper-cli runs as a subprocess of the Python GUI. The OOM kill happened during Demucs pre-processing — whisper-cli was never started, so no VRAM allocation occurred.
From the Buzz log at crash time:
~/.var/app/io.github.chidiwilliams.Buzz/.local/state/Buzz/log/logs.txt
The last entry was Will extract speech — the log never reached Starting whisper file transcription.
Fix
Uncheck "Extract Speech" in the Buzz transcription dialog for any file longer than ~30 minutes. The large-v3-turbo model has built-in noise tolerance that makes Demucs pre-processing unnecessary in most cases.
Optionally, grow the swapfile if you need Extract Speech for shorter files:
sudo swapoff /swap.img sudo fallocate -l 16G /swap.img sudo mkswap /swap.img sudo swapon /swap.img
Environment Variables
| Variable | Default | Effect |
|---|---|---|
BUZZ_FORCE_CPU |
false | Set to true to disable GPU inference
|
BUZZ_WHISPERCPP_N_THREADS |
cpu_count / 2 | Override thread count for whisper-cli |
Set via: flatpak override --user io.github.chidiwilliams.Buzz --env=VAR=value
Key File Paths
| Path | Description |
|---|---|
/var/lib/flatpak/app/io.github.chidiwilliams.Buzz/.../buzz/whisper_cpp/whisper-cli |
Bundled Vulkan whisper-cli (libwhisper 1.8.3) |
~/.var/app/io.github.chidiwilliams.Buzz/cache/Buzz/models/ |
Buzz model cache (ggml binaries) |
~/.var/app/io.github.chidiwilliams.Buzz/.local/state/Buzz/log/logs.txt |
Buzz application log |
~/.var/app/io.github.chidiwilliams.Buzz/config/Buzz.conf |
Buzz settings |
~/whisper.cpp/build/bin/whisper-cli |
Custom Vulkan-enabled whisper.cpp build |
~/bin/whisper-cli-vulkan |
Wrapper script for custom build (sets LD_LIBRARY_PATH)
|
~/.local/share/pipx/venvs/buzz-captions/lib/python3.12/site-packages/buzz/whisper_cpp/ |
pipx install whisper_cpp directory (directly modifiable) |
Investigation Summary
| Question | Finding |
|---|---|
| Is Buzz using the RX 7600? | Yes. using Vulkan0 backend confirmed inside the Flatpak sandbox.
|
Was --no-gpu being added? |
No. IS_VULKAN_SUPPORTED = True in the sandbox; flag is never appended.
|
| Does the large model load correctly? | Yes. 1.6 GB to VRAM, transcribes in ~1.25s on a short clip. |
| What caused the crash? | Extract Speech (Demucs) OOM. Python killed at 22 GB RSS before whisper-cli launched. |
| Why was VRAM not climbing? | whisper-cli was never started — process died during Demucs pre-processing.
|
| Fix? | Uncheck Extract Speech for long files. Clear swap before sessions. |
Investigation date: 2026-04-25. System: Ubuntu 24.04, AMD RX 7600, Buzz 1.4.4 (Flatpak).
Folder Watch (Automatic Transcription)


Settings
In Preferences → Folder Watch:
| Setting | Value | Notes |
|---|---|---|
| Enable folder watch | leave unticked | ticking it starts auto-transcribing everything dropped into the input folder |
| Input folder | /home/justin/Music |
drop recordings here |
| Output folder | /home/justin/Music |
same folder — files stay put |
| Delete processed files | unchecked | ticking this deletes the source audio after transcription — keep it unticked |
| Model | whisper.cpp — Large-V3 | engine choice matters, see below |
| Task | Transcribe | |
| Language | Detect Language | |
| Export | TXT | SRT / VTT also available |
| Word-level timings | unchecked | |
| Extract speech | unchecked | do not check — crashes on long files (see OOM Crash with Large Files) |
Use whisper.cpp instead of the default engine

The model picker lists two engines: whisper.cpp and Faster Whisper. Faster Whisper is the default — switch to whisper.cpp:
- whisper.cpp uses the
ggml-*.binmodels already cached under~/.var/app/io.github.chidiwilliams.Buzz/cache/Buzz/models/(Flatpak path;~/snap/buzz/…on snap installs) — no extra download - it runs through the bundled Vulkan
whisper-cli, so GPU acceleration applies (see GPU Acceleration)

In the Models tab, the whisper.cpp group shows the downloaded sizes — Large-V3 and Large-V3-Turbo are the useful ones for meetings (Large-V3-Turbo: best quality/speed trade-off, see Recommended Settings).
Leave the advanced options unchecked

In the transcription dialog (shown when folder watch picks up a file) both Word-level timings and Extract speech are left unchecked. Checking Extract speech — the obvious option for a meeting recording — triggers the Demucs pre-processing step that OOM-crashes the app on long files (22 GB RSS kill, see above). Keep both unchecked.
Result
Transcriptions are written next to the source as <original name> (transcribed on <date>).txt, e.g. 2026-08-08 20-56-20 3rd Apocalypse Bob's ADND1E (transcribed on 09-Aug-2026 01-50-32).txt.
Config verified 2026-08-09 in ~/snap/buzz/570/.config/Buzz.conf (snap install): folder_watch\enabled = false (off — recommended), folder_watch\input_folder = /home/justin/Music, folder_watch\output_directory = /home/justin/Music, model Whisper.cpp / large-v3, and delete_processed_files, extract_speech, word_level_timings all false.
Example: Taglish Test (2026-08-09)
A Taglish (Tagalog + English code-switching) test recording, transcribed automatically by Buzz's Folder Watch during testing (now disabled — see the Folder Watch warning above):
- Audio: Taglish Test 260809.mp3
- Evaluation report: Taglish Transcription and Evaluation Report 260809.docx
The report grades the result 95/100 (Excellent) at ~96 % word accuracy. Code-switching between Filipino and English (EDSA, client, project, presentation, brain cells, lunch, work) is handled seamlessly. Noted deviations are typical ASR behavior: punctuation is collapsed into sentence breaks, "Almost two hours" became "Almost 2 hours", "na-stuck" → "nakastuck", "kailangan nating" → "kailan nating".