Problem
load_model() is decorated with @lru_cache(maxsize=1).
Once a worker serves its first prediction, replacing models/baseline_model.joblib with a retrained model does not refresh the in-memory model. The API keeps serving predictions from the old artifact until the process restarts or the cache is cleared manually.
Why it matters
SentinelML has a retraining workflow, so stale inference after a model replacement is an operational correctness risk.
Suggested fix
Use an explicit reload strategy, for example:
- cache the model together with the file mtime/hash and reload when it changes, or
- expose a controlled reload lifecycle hook used by deployment/retraining automation.
Avoid reloading the model on every request.
Acceptance criteria
- first prediction loads model A
- replacing the artifact with model B causes a controlled reload
- subsequent predictions use model B without requiring a process restart
- tests cover the reload boundary
Problem
load_model()is decorated with@lru_cache(maxsize=1).Once a worker serves its first prediction, replacing
models/baseline_model.joblibwith a retrained model does not refresh the in-memory model. The API keeps serving predictions from the old artifact until the process restarts or the cache is cleared manually.Why it matters
SentinelML has a retraining workflow, so stale inference after a model replacement is an operational correctness risk.
Suggested fix
Use an explicit reload strategy, for example:
Avoid reloading the model on every request.
Acceptance criteria