fix(quality): disable reasoning mode and add local silence detection - #26
Merged
Merged
Conversation
Dictation quality regressed because the Ollama payload never set "think". Gemma 4 defaults to thinking: true, and its reasoning tokens draw from the same num_predict budget as the answer. Once exhausted, Ollama returns empty content with done_reason=length and the model reports empty: true, so a correct transcript was discarded. Against a reference clip at 5 dB SNR the old payload dropped 4 of 5 dictations; think: false restores 5/5 and is roughly 20x faster. Also order TRANSCRIPTION_SCHEMA so transcription precedes empty. Ollama emits properties in schema order under constrained decoding, so the model was committing to the empty flag before transcribing anything. The model's empty flag is not trustworthy: given silence it answers empty=false and invents text, which the pipeline then saves as a note. Decide silence locally via RMS against a -45 dBFS threshold, and trim leading and trailing silence before encoding to cut the audio the model must encode. Also resolve the pre-existing mypy failures: QWidget event handler signatures in recording_hud, untyped request payload and message lists in llm, the unreachable return in the retry loop, plus redundant open modes, implicit string concatenation, and try/else placement found by a wider ruff sweep. Extract _parse_json_response to name the missing keys. Make CI enforce ruff and mypy, which previously ran with --exit-zero and could not fail the build. Tests: 164 passing, 95% coverage (llm.py and audio.py at 100%).
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
Dictations were being silently discarded, and the transcription model was
hallucinating notes out of silence. Both stem from the Ollama payload and the
empty-recording guard rather than from audio capture.
Why transcription was losing notes
llm.pynever setthink. Gemma 4 defaults tothinking: true, and itsreasoning tokens draw from the same
num_predictbudget as the answer. Oncethat budget is exhausted Ollama returns an empty
contentwithdone_reason: length, which the model reports as"empty": true— soapp.pydiscarded a correctly transcribed dictation and told the user theaudio was empty.
Measured against Ollama's own reference clip ("Why is the sky blue?") with
noise added to 5 dB SNR:
think)think: false)Changes
Reasoning disabled on every request.
"think": falsein the payload(
llm.THINK). This is the highest-impact change: it fixes noisy-audiotranscription and is also roughly 20x faster. It also does not degrade the two
text-only phases, and produces a better note title in spot checks.
Schema field order inverted.
TRANSCRIPTION_SCHEMAdeclaredemptybefore
transcription. Ollama emits properties in schema order underconstrained decoding, so the model had to commit to the empty flag before
transcribing anything. Isolated with Ollama's canonical prompt: 10/10
empty: truedespite a perfect transcript, versus 10/10 correct onceinverted.
Local silence detection. Asked to label silence the model answers
empty: falseand invents text — pure digital silence came back as a fluentSpanish sentence about time, which the pipeline then saved as a note. This is
a pre-existing data-integrity bug, not a consequence of the reasoning change.
audio.is_silent()now decides emptiness by RMS against a −45 dBFS threshold,and
app.pyrejects silence without contacting the model. Leading andtrailing silence is trimmed before encoding, which shortens what the model must
encode; interior pauses are preserved.
Pre-existing quality issues fixed
mypyhad never passed. Six errors, none introduced by this work:recording_hud.py—paintEvent/mousePressEventsignatures did not matchQWidget(PyQt6 types the event as optional), plus an unannotated_style_tierinferred asNonellm.py— untyped request payload and message lists, and a retry loop thatcould fall through without returning
Also found by a wider ruff sweep than the project configures: redundant
openmodes, implicit string concatenation inside a regex, an unused
os.walkvariable, and two
returns insidetrythat belong inelse. Extracted_parse_json_response()so a non-object JSON response and missing keys faildistinctly, naming the keys that were absent.
CI now enforces these
sonarcloud.ymlran ruff with--exit-zeroand mypy with|| true, soneither could ever fail the build. Added a Verify quality gates step that
does. Flagging this explicitly as an intentional behaviour change; drop that
step if CI should stay advisory.
Verification
Everything below ran against the real shipped code paths
(
AudioRecorder+llm.transcribe_audio), not replicas, n=5 per case:164 tests pass (13 new), 95% coverage,
llm.pyandaudio.pyat 100%.mypyandruffboth clean.Not changed
gemma4:12b-it-qatwas evaluated and rejected as a quality improvement: sameWER on the reference clip and 46s with reasoning on.
mypy --strictwas leftalone — it surfaces a large amount of unrelated annotation debt.