Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,18 @@
# Changelog

## 0.1.2

- Give each media playback its own player and keep native ownership separate
from UI completion. Old events and page exits cannot cancel a newer request.
- Bound each native call to eight seconds. An ambiguous timeout disables audio
for the application session, immediately attempts to stop the captured source,
and observes late results for cleanup. Learning and answer submission remain
available with a text status explaining that audio is unavailable.
- Disable unused native position polling and test the actual service against
controlled player and system-speech platform calls, including delayed failures.
- Document system-speech callback limitations and best-effort native cleanup.
These tests do not establish physical-device silence or a total exit deadline.

## 0.1.1

- Keep imported content and learning progress when the home screen refreshes.
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ not a signed store release. Platform support beyond those checks should be verif
| Bound a slow speech download | [community_gateway.dart](lib/services/community_gateway.dart) | [gateway lifecycle](test/community_gateway_test.dart) |
| Prepare a usable review from available content | [review_preparation.dart](lib/services/review_preparation.dart) | [review_preparation_test.dart](test/review_preparation_test.dart) |
| Coordinate pronunciation during review | [review_pronunciation.dart](lib/services/review_pronunciation.dart) | [review_pronunciation_test.dart](test/review_pronunciation_test.dart), [playback arbitration](test/tts_playback_arbiter_test.dart) |
| Isolate old audio events and bound native waits | [audio ownership contract](docs/audio-lifecycle.md), [tts_service.dart](lib/services/tts_service.dart) | [native lifecycle](test/tts_service_lifecycle_test.dart) |
| Keep review states explicit | [review availability tests](test/review_availability_widget_test.dart) | [loading-state tests](test/review_loading_widget_test.dart) |

```sh
Expand Down
2 changes: 2 additions & 0 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,8 @@ flutter run --dart-define=WORD_AI_API_BASE_URL=https://your-gateway.example

该地址会编译进客户端,只能放非敏感的服务地址。供应商密钥只能保存在服务端。使用自建网关时,当前需要发音的文字会发送给该网关;失败不会阻止本地学习。

发音协调器区分界面状态与原生播放归属,隔离旧播放器事件,并为每次原生调用设置期限。超时后会尝试停止音频,页面仍可继续答题;本次应用会话不再启动新的音频。具体保证、插件限制和设备验证范围见 [音频生命周期说明](docs/audio-lifecycle.md),回归入口为 [实际服务测试](test/tts_service_lifecycle_test.dart)。

## 开源与贡献

源代码采用 Apache-2.0,可修改、自托管、再分发及商用,遵守许可证和第三方许可即可。WordAI 名称与图标不构成商标授权。依赖包继续适用各自许可证。
Expand Down
80 changes: 80 additions & 0 deletions docs/audio-lifecycle.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# Audio ownership and native failure boundaries

WordAI has one audio coordinator. A review owns its pending request and native
playback by source ID and request generation. Finishing the on-screen animation
does not release native ownership: leaving the review still requests a stop.
Review answers and saved progress do not depend on audio being available.

## What the Dart service enforces

- Every media utterance gets a fresh `AudioPlayer`. The old player must
acknowledge stop and finish disposal before another backend may speak.
Media events belong to that captured player and request; a retired player's
completion or error cannot finish a new player or trigger system speech.
- Fresh media players disable the unused position updater before configuration.
Speech has no timeline UI; native position polling would otherwise introduce
plugin Futures outside the coordinator's error handling and disposal bounds.
- Device speech is created independently of media playback. Using local system
speech does not allocate or wait for an unused media player.
- Each native operation has an eight-second Dart deadline. This covers player
creation (awaited by its first configuration call), configuration, preparing
a source, resume/speak, rate changes, stop and disposal. A sequence of successful
calls can take more than eight seconds; this is a per-call bound. Downloads
have the separate [gateway deadline](speech-gateway.md).
- Timeout does not cancel a native call. A timed-out call, rejected stop or
failed disposal makes the coordinator permanently unavailable for that app
session. Queued and future playback requests finish with an error; neither
another media player nor system speech is started as a workaround.
- Late results and errors remain observed. A late result can request another
bounded stop of its captured native owner. It cannot restore availability,
update the current indicator, resume playback or create another backend.
- Entering the unavailable state immediately attempts a separate bounded stop,
even if the original call never returns. That attempt is cleanup only; a
later source/resume result still gets another stop attempt. Cleanup cannot
restore availability and a stuck cleanup stop does not retry recursively.
- Disposal rejects new work immediately and performs bounded cleanup. A media
handle with an outstanding native call is retained rather than treating
`AudioPlayer.dispose()` as a way to kill that call. Logical disposal may finish
while native cleanup remains unconfirmed. No successful return from Dart
disposal promises acoustic silence.

An ordinary media error can fall back to device speech only after the old
player's stop and disposal have both succeeded. An uncertain stop cannot fall
back. The review shows an inline unavailable notice and remains usable; restart
the app before trying audio again. Playback's boolean result means the current
request was accepted, not that a speaker produced sound or the utterance finished.

## Limits of the locked plugins

The lockfile currently resolves `audioplayers` 6.8.1 and `flutter_tts` 4.2.5.
This implementation depends on their public behavior:

| Plugin boundary | Consequence |
| --- | --- |
| AudioPlayer routes native events by `playerId`, but not by source within one player. | Separate players provide an event boundary between media utterances. |
| AudioPlayer disposal calls stop/release before native disposal. | A hung stop can also hang disposal; neither method is a forced native termination API. |
| FlutterTts instances share the `flutter_tts` channel and native engine. Constructing one replaces the Dart handler. | Creating another FlutterTts is not independent engine isolation or timeout recovery. |
| FlutterTts forwards system completion/cancel/error events without an utterance ID. | A delayed event for system utterance A can still change the indicator for system utterance B. Backend checks prevent it affecting media; retained ownership ensures B can still be stopped. |
| Android/iOS system stop returns a plugin acknowledgement without propagating the native stop result. | Acknowledgement is the strongest available handoff signal, not proof of silence. |

Strict correlation between successive system utterances needs a native bridge
that returns an utterance ID and engine generation with each event. Switching
to `awaitSpeakCompletion` does not supply this: the locked plugin uses shared
completion state. WordAI does not claim this stronger guarantee.

## Verification and remaining device checks

[`tts_service_lifecycle_test.dart`](../test/tts_service_lifecycle_test.dart)
drives the actual TTSService, AudioPlayer wrapper, per-player platform event
streams and FlutterTts method channel. Gates control pending and late native
calls. It verifies cross-backend events, retired media instances, cancellation,
failed recovery, blocked source/resume/stop/dispose, queued callers and late
cleanup. These are host tests; they do not play audio.
The fixture returns unmodified AudioPlayers and rejects native position queries,
so it also verifies that production code disables the unused updater.

On Android and iOS devices, still verify audible stop/handoff, interruptions,
background/resume, rapid A-to-B system speech, missing voices and speaker/route
changes. In particular, measure whether native audio continues after a stop
acknowledgement or plugin failure. Host tests cannot establish acoustic silence
or the OS's resource-release behavior. No device result is claimed here.
5 changes: 4 additions & 1 deletion docs/speech-gateway.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,10 @@ No server is bundled or configured. An operator may implement `POST /tts` relati
- One ten-second deadline covers connection, response headers and the complete
body. Receiving more chunks does not extend it. Timeout aborts the request,
cancels body reading and closes the client dedicated to that request.
- Failed/invalid responses fall back to system speech where available. Successful audio is cached locally.
- Failed/invalid responses can fall back to system speech where available.
An unconfirmed native stop disables further audio until the app restarts;
see [audio ownership and failure boundaries](audio-lifecycle.md). Successful
audio is cached locally.
- The client sends no provider credentials. Keep provider keys on the server, outside source control.
- Operators must add input limits, rate limits, abuse prevention and appropriate access control; this minimal client does not implement user authentication. Do not expose an unrestricted paid provider proxy.
- Configure only infrastructure you control. Use of the gateway sends the requested speech text to that operator.
26 changes: 26 additions & 0 deletions lib/pages/flash_cards/flash_cards_widget.dart
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ import '/flutter_flow/flutter_flow_util.dart';
import '/services/learning_repository.dart';
import '/services/review_preparation.dart';
import '/services/review_pronunciation.dart';
import '/services/tts_service.dart';
import '/widgets/review_loading_view.dart';

class FlashCardsWidget extends StatefulWidget {
Expand Down Expand Up @@ -764,6 +765,31 @@ class _FlashCardsWidgetState extends State<FlashCardsWidget>
: '${math.min(session.currentIndex + 1, session.targetIds.length)} / ${session.targetIds.length}',
onClose: _requestExit,
),
if (_pronunciation.playbackState case final playbackState?)
ValueListenableBuilder<TtsPlaybackSnapshot>(
valueListenable: playbackState,
builder: (context, state, child) =>
state.phase == TtsPlaybackPhase.unavailable
? Padding(
padding: const EdgeInsets.fromLTRB(24, 0, 24, 12),
child: Semantics(
liveRegion: true,
child: Text(
_t(
'Pronunciation is unavailable. You can keep reviewing. If sound continues, close the app. Reopen it to try audio again.',
'发音暂不可用,可继续复习。如仍有声音,请关闭应用;重新打开后可再试。',
'發音暫不可用,可繼續複習。如仍有聲音,請關閉應用程式;重新開啟後可再試。',
),
key: const ValueKey('audio-unavailable'),
style: theme.bodySmall.copyWith(
color: theme.secondaryText,
height: 1.4,
),
),
),
)
: const SizedBox.shrink(),
),
Expanded(
child: AnimatedSwitcher(
duration:
Expand Down
5 changes: 5 additions & 0 deletions lib/services/review_pronunciation.dart
Original file line number Diff line number Diff line change
@@ -1,5 +1,7 @@
import 'dart:async';

import 'package:flutter/foundation.dart';

import 'tts_service.dart';

/// Owns only this review's pronunciation; downloads may finish caching after
Expand All @@ -8,6 +10,7 @@ class ReviewPronunciation {
ReviewPronunciation({
required Future<void> Function(String word) play,
required Future<void> Function() stop,
this.playbackState,
}) : _play = play,
_stop = stop;

Expand All @@ -18,12 +21,14 @@ class ReviewPronunciation {
await TTSService.instance.generateAndPlay(text: word, sourceId: source);
},
stop: () => TTSService.instance.stopSource(source),
playbackState: TTSService.instance.playbackState,
);
}

static int _sequence = 0;
final Future<void> Function(String) _play;
final Future<void> Function() _stop;
final ValueListenable<TtsPlaybackSnapshot>? playbackState;
String? _lastQuestion;
bool _started = false;
bool _disposed = false;
Expand Down
Loading