Skip to content

fix(quality): disable reasoning mode and add local silence detection - #26

Merged
arrase merged 2 commits into
mainfrom
fix/transcription-quality
Sep 29, 2026
Merged

arrase merged 2 commits into
mainfrom
fix/transcription-quality

Conversation

@arrase

@arrase arrase commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Summary

Dictations were being silently discarded, and the transcription model was
hallucinating notes out of silence. Both stem from the Ollama payload and the
empty-recording guard rather than from audio capture.

Why transcription was losing notes

llm.py never set think. Gemma 4 defaults to thinking: true, and its
reasoning tokens draw from the same num_predict budget as the answer. Once
that budget is exhausted Ollama returns an empty content with
done_reason: length, which the model reports as "empty": true — so
app.py discarded a correctly transcribed dictation and told the user the
audio was empty.

Measured against Ollama's own reference clip ("Why is the sky blue?") with
noise added to 5 dB SNR:

Config Result Latency
before (no think) 4 of 5 dictations lost 3.0–7.1s
after (think: false) 5 of 5 correct 0.34s

Changes

Reasoning disabled on every request. "think": false in the payload
(llm.THINK). This is the highest-impact change: it fixes noisy-audio
transcription and is also roughly 20x faster. It also does not degrade the two
text-only phases, and produces a better note title in spot checks.

Schema field order inverted. TRANSCRIPTION_SCHEMA declared empty
before transcription. Ollama emits properties in schema order under
constrained decoding, so the model had to commit to the empty flag before
transcribing anything. Isolated with Ollama's canonical prompt: 10/10
empty: true despite a perfect transcript, versus 10/10 correct once
inverted.

Local silence detection. Asked to label silence the model answers
empty: false and invents text — pure digital silence came back as a fluent
Spanish sentence about time, which the pipeline then saved as a note. This is
a pre-existing data-integrity bug, not a consequence of the reasoning change.
audio.is_silent() now decides emptiness by RMS against a −45 dBFS threshold,
and app.py rejects silence without contacting the model. Leading and
trailing silence is trimmed before encoding, which shortens what the model must
encode; interior pauses are preserved.

Pre-existing quality issues fixed

mypy had never passed. Six errors, none introduced by this work:

  • recording_hud.py — paintEvent/mousePressEvent signatures did not match
    QWidget (PyQt6 types the event as optional), plus an unannotated
    _style_tier inferred as None
  • llm.py — untyped request payload and message lists, and a retry loop that
    could fall through without returning

Also found by a wider ruff sweep than the project configures: redundant open
modes, implicit string concatenation inside a regex, an unused os.walk
variable, and two returns inside try that belong in else. Extracted
_parse_json_response() so a non-object JSON response and missing keys fail
distinctly, naming the keys that were absent.

CI now enforces these

sonarcloud.yml ran ruff with --exit-zero and mypy with || true, so
neither could ever fail the build. Added a Verify quality gates step that
does. Flagging this explicitly as an intentional behaviour change; drop that
step if CI should stay advisory.

Verification

Everything below ran against the real shipped code paths
(AudioRecorder + llm.transcribe_audio), not replicas, n=5 per case:

Case Result
clean speech 5/5 correct, 0.38s
speech @ 5 dB SNR 5/5 correct, 0.32s
digital silence 3s / 10s 5/5 empty, model never called
mic hiss only 5/5 empty, model never called
trimming 4s+speech+4s → 1.80s sent to model

164 tests pass (13 new), 95% coverage, llm.py and audio.py at 100%.
mypy and ruff both clean.

Not changed

gemma4:12b-it-qat was evaluated and rejected as a quality improvement: same
WER on the reference clip and 46s with reasoning on. mypy --strict was left
alone — it surfaces a large amount of unrelated annotation debt.

Dictation quality regressed because the Ollama payload never set "think".
Gemma 4 defaults to thinking: true, and its reasoning tokens draw from the
same num_predict budget as the answer. Once exhausted, Ollama returns empty
content with done_reason=length and the model reports empty: true, so a
correct transcript was discarded. Against a reference clip at 5 dB SNR the
old payload dropped 4 of 5 dictations; think: false restores 5/5 and is
roughly 20x faster.

Also order TRANSCRIPTION_SCHEMA so transcription precedes empty. Ollama emits
properties in schema order under constrained decoding, so the model was
committing to the empty flag before transcribing anything.

The model's empty flag is not trustworthy: given silence it answers
empty=false and invents text, which the pipeline then saves as a note. Decide
silence locally via RMS against a -45 dBFS threshold, and trim leading and
trailing silence before encoding to cut the audio the model must encode.

Also resolve the pre-existing mypy failures: QWidget event handler signatures
in recording_hud, untyped request payload and message lists in llm, the
unreachable return in the retry loop, plus redundant open modes, implicit
string concatenation, and try/else placement found by a wider ruff sweep.
Extract _parse_json_response to name the missing keys. Make CI enforce ruff
and mypy, which previously ran with --exit-zero and could not fail the build.

Tests: 164 passing, 95% coverage (llm.py and audio.py at 100%).
@sonarqubecloud

Copy link
Copy Markdown

@arrase
arrase merged commit 383b5c4 into main Sep 29, 2026
2 checks passed
@arrase
arrase deleted the fix/transcription-quality branch September 29, 2026 20:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant