Repository navigation
fix(adapters): raise or fall back when an aLoRA cannot activate - #1684
Merged
Merged
Conversation
An aLoRA only takes effect from the point its declared invocation token sequence appears in the assembled prompt. When that sequence is absent, PEFT silently leaves the base model running instead: no warning, no error, and list_adapters()/active_adapters() still report the adapter as loaded. Callers get a plausible-looking but fabricated result with no way to tell it came from the base model. This adds a generation-time guard (LocalHFBackend, the only backend that loads adapters itself) that checks, right before dispatch, whether the loaded aLoRA's declared alora_invocation_tokens sequence occurs in the assembled prompt tokens. On a mismatch it raises AloraActivationError before any model call: - A direct adapter-function call (core.check_certainty, etc.) lets the error propagate; there is no fallback layer below it. - Requirement.validate()'s automatic aLoRA routing catches the error and falls back to LLM-as-a-judge, mirroring the existing fallback for an adapter that was never registered at all. Also adds a comment above the pinned adapter SHAs in catalog.py naming the accuracy/false-pass evidence a future pin bump must carry before promoting an aLoRA rung to the default path. Fixes generative-computing#1678 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
…uard Fixes issues surfaced by an independent 3-reviewer panel (2 tier-1, 1 tier-3) on PR generative-computing#1684: - Correct the PR description's and the docs page's overclaim that every high-level adapter-function wrapper now raises AloraActivationError by default. resolve_adapter() still hardcodes AdapterType.LORA (Epic generative-computing#929 Phase 2), so the guard only fires for a caller-registered composed aLoRA, Requirement/ALoraRequirement routing against one, or the deprecated IntrinsicAdapter shim from its second call onward. Reword the "aLoRA activation is not guaranteed by loading" section accordingly. - Fix AloraActivationError's pickle/deepcopy round-trip: pass the four positional fields to super().__init__() and move message formatting to __str__, mirroring the sibling AdapterSchemaMismatchError's existing convention. - Fire adapter_function_invocation_complete (outcome="error") when the guard aborts a call, matching AdapterMixin.adapter_scope's own revision/binding_type derivation. Without this, a guard-aborted invocation left no telemetry signal at all for exactly the scenario the guard exists to catch. - Add missing Raises: AloraActivationError entries to the docstrings of every public wrapper that can now surface it: core.py's check_certainty/requirement_check/find_context_attributions, rag.py's five adapter functions, _util.py's call_intrinsic, and huggingface.py's _generate_from_intrinsic. - Rename alora_invocation_sequence_present to _alora_invocation_sequence_present: it had no external consumer and was not re-exported, unlike its non-underscore siblings in the same module. - Document that alora_invocation_sequence_present's contiguous match is not scoped to the last (activating) occurrence, and correct the docs' claim that the IntrinsicAdapter shim is never covered -- it is exempt only on its first call. - Add end-to-end tests: the matching-sequence arm through the real _generate_from_intrinsic call site (previously only covered via the extracted helper), and a hook-capture test proving adapter_function_invocation_complete fires on a guard abort. Verified: full test/ -m "not qualitative" suite passes (4646 passed, 0 failed), ruff format/check clean, mypy clean, docstring-quality gate clean (uv run python tooling/docs-autogen/audit_coverage.py --quality --fail-on-quality --threshold 100), markdownlint clean. Each new test was confirmed to fail against the pre-fix code before the corresponding fix, then pass after. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
A guard abort on a legacy IntrinsicAdapter read binding_type and revision
from its inert _ShimWeightsBinding ("unknown", None), while a successful
call reports through _IntrinsicPeftBinding ("local_file", catalogue SHA).
Report the same fields on both, and cover the shim path with a test.
Assisted-by: Claude Code
Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
Drop issue/PR/acceptance-criteria references from docstrings and comments this PR added where the surrounding text already explains the point, and update the shim-removal pointer from the closed generative-computing#1144 to generative-computing#1621. Regression tests keep their generative-computing#1678 reference, and test data keeps its generative-computing#1679 source. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
jakelorocco
approved these changes
Oct 1, 2026
Merged
via the queue into
generative-computing:main
with commit Oct 2, 2026
4c9a3c1
14 checks passed
gretadolcetti
pushed a commit
to gretadolcetti/mellea
that referenced
this pull request
Oct 5, 2026
…rative-computing#1684) * fix(adapters): raise or fall back when an aLoRA cannot activate An aLoRA only takes effect from the point its declared invocation token sequence appears in the assembled prompt. When that sequence is absent, PEFT silently leaves the base model running instead: no warning, no error, and list_adapters()/active_adapters() still report the adapter as loaded. Callers get a plausible-looking but fabricated result with no way to tell it came from the base model. This adds a generation-time guard (LocalHFBackend, the only backend that loads adapters itself) that checks, right before dispatch, whether the loaded aLoRA's declared alora_invocation_tokens sequence occurs in the assembled prompt tokens. On a mismatch it raises AloraActivationError before any model call: - A direct adapter-function call (core.check_certainty, etc.) lets the error propagate; there is no fallback layer below it. - Requirement.validate()'s automatic aLoRA routing catches the error and falls back to LLM-as-a-judge, mirroring the existing fallback for an adapter that was never registered at all. Also adds a comment above the pinned adapter SHAs in catalog.py naming the accuracy/false-pass evidence a future pin bump must carry before promoting an aLoRA rung to the default path. Fixes generative-computing#1678 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * fix(adapters): address code-review findings on the aLoRA activation guard Fixes issues surfaced by an independent 3-reviewer panel (2 tier-1, 1 tier-3) on PR generative-computing#1684: - Correct the PR description's and the docs page's overclaim that every high-level adapter-function wrapper now raises AloraActivationError by default. resolve_adapter() still hardcodes AdapterType.LORA (Epic generative-computing#929 Phase 2), so the guard only fires for a caller-registered composed aLoRA, Requirement/ALoraRequirement routing against one, or the deprecated IntrinsicAdapter shim from its second call onward. Reword the "aLoRA activation is not guaranteed by loading" section accordingly. - Fix AloraActivationError's pickle/deepcopy round-trip: pass the four positional fields to super().__init__() and move message formatting to __str__, mirroring the sibling AdapterSchemaMismatchError's existing convention. - Fire adapter_function_invocation_complete (outcome="error") when the guard aborts a call, matching AdapterMixin.adapter_scope's own revision/binding_type derivation. Without this, a guard-aborted invocation left no telemetry signal at all for exactly the scenario the guard exists to catch. - Add missing Raises: AloraActivationError entries to the docstrings of every public wrapper that can now surface it: core.py's check_certainty/requirement_check/find_context_attributions, rag.py's five adapter functions, _util.py's call_intrinsic, and huggingface.py's _generate_from_intrinsic. - Rename alora_invocation_sequence_present to _alora_invocation_sequence_present: it had no external consumer and was not re-exported, unlike its non-underscore siblings in the same module. - Document that alora_invocation_sequence_present's contiguous match is not scoped to the last (activating) occurrence, and correct the docs' claim that the IntrinsicAdapter shim is never covered -- it is exempt only on its first call. - Add end-to-end tests: the matching-sequence arm through the real _generate_from_intrinsic call site (previously only covered via the extracted helper), and a hook-capture test proving adapter_function_invocation_complete fires on a guard abort. Verified: full test/ -m "not qualitative" suite passes (4646 passed, 0 failed), ruff format/check clean, mypy clean, docstring-quality gate clean (uv run python tooling/docs-autogen/audit_coverage.py --quality --fail-on-quality --threshold 100), markdownlint clean. Each new test was confirmed to fail against the pre-fix code before the corresponding fix, then pass after. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * fix(adapters): report shim aLoRA guard aborts as pinned local_file A guard abort on a legacy IntrinsicAdapter read binding_type and revision from its inert _ShimWeightsBinding ("unknown", None), while a successful call reports through _IntrinsicPeftBinding ("local_file", catalogue SHA). Report the same fields on both, and cover the shim path with a test. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * docs(adapters): trim transient issue references from aLoRA guard docs Drop issue/PR/acceptance-criteria references from docstrings and comments this PR added where the surrounding text already explains the point, and update the shim-removal pointer from the closed generative-computing#1144 to generative-computing#1621. Regression tests keep their generative-computing#1678 reference, and test data keeps its generative-computing#1679 source. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> --------- Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
gretadolcetti
pushed a commit
to gretadolcetti/mellea
that referenced
this pull request
Oct 5, 2026
…rative-computing#1684) * fix(adapters): raise or fall back when an aLoRA cannot activate An aLoRA only takes effect from the point its declared invocation token sequence appears in the assembled prompt. When that sequence is absent, PEFT silently leaves the base model running instead: no warning, no error, and list_adapters()/active_adapters() still report the adapter as loaded. Callers get a plausible-looking but fabricated result with no way to tell it came from the base model. This adds a generation-time guard (LocalHFBackend, the only backend that loads adapters itself) that checks, right before dispatch, whether the loaded aLoRA's declared alora_invocation_tokens sequence occurs in the assembled prompt tokens. On a mismatch it raises AloraActivationError before any model call: - A direct adapter-function call (core.check_certainty, etc.) lets the error propagate; there is no fallback layer below it. - Requirement.validate()'s automatic aLoRA routing catches the error and falls back to LLM-as-a-judge, mirroring the existing fallback for an adapter that was never registered at all. Also adds a comment above the pinned adapter SHAs in catalog.py naming the accuracy/false-pass evidence a future pin bump must carry before promoting an aLoRA rung to the default path. Fixes generative-computing#1678 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * fix(adapters): address code-review findings on the aLoRA activation guard Fixes issues surfaced by an independent 3-reviewer panel (2 tier-1, 1 tier-3) on PR generative-computing#1684: - Correct the PR description's and the docs page's overclaim that every high-level adapter-function wrapper now raises AloraActivationError by default. resolve_adapter() still hardcodes AdapterType.LORA (Epic generative-computing#929 Phase 2), so the guard only fires for a caller-registered composed aLoRA, Requirement/ALoraRequirement routing against one, or the deprecated IntrinsicAdapter shim from its second call onward. Reword the "aLoRA activation is not guaranteed by loading" section accordingly. - Fix AloraActivationError's pickle/deepcopy round-trip: pass the four positional fields to super().__init__() and move message formatting to __str__, mirroring the sibling AdapterSchemaMismatchError's existing convention. - Fire adapter_function_invocation_complete (outcome="error") when the guard aborts a call, matching AdapterMixin.adapter_scope's own revision/binding_type derivation. Without this, a guard-aborted invocation left no telemetry signal at all for exactly the scenario the guard exists to catch. - Add missing Raises: AloraActivationError entries to the docstrings of every public wrapper that can now surface it: core.py's check_certainty/requirement_check/find_context_attributions, rag.py's five adapter functions, _util.py's call_intrinsic, and huggingface.py's _generate_from_intrinsic. - Rename alora_invocation_sequence_present to _alora_invocation_sequence_present: it had no external consumer and was not re-exported, unlike its non-underscore siblings in the same module. - Document that alora_invocation_sequence_present's contiguous match is not scoped to the last (activating) occurrence, and correct the docs' claim that the IntrinsicAdapter shim is never covered -- it is exempt only on its first call. - Add end-to-end tests: the matching-sequence arm through the real _generate_from_intrinsic call site (previously only covered via the extracted helper), and a hook-capture test proving adapter_function_invocation_complete fires on a guard abort. Verified: full test/ -m "not qualitative" suite passes (4646 passed, 0 failed), ruff format/check clean, mypy clean, docstring-quality gate clean (uv run python tooling/docs-autogen/audit_coverage.py --quality --fail-on-quality --threshold 100), markdownlint clean. Each new test was confirmed to fail against the pre-fix code before the corresponding fix, then pass after. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * fix(adapters): report shim aLoRA guard aborts as pinned local_file A guard abort on a legacy IntrinsicAdapter read binding_type and revision from its inert _ShimWeightsBinding ("unknown", None), while a successful call reports through _IntrinsicPeftBinding ("local_file", catalogue SHA). Report the same fields on both, and cover the shim path with a test. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * docs(adapters): trim transient issue references from aLoRA guard docs Drop issue/PR/acceptance-criteria references from docstrings and comments this PR added where the surrounding text already explains the point, and update the shim-removal pointer from the closed generative-computing#1144 to generative-computing#1621. Regression tests keep their generative-computing#1678 reference, and test data keeps its generative-computing#1679 source. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> --------- Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
gretadolcetti
pushed a commit
to gretadolcetti/mellea
that referenced
this pull request
Oct 5, 2026
…rative-computing#1684) * fix(adapters): raise or fall back when an aLoRA cannot activate An aLoRA only takes effect from the point its declared invocation token sequence appears in the assembled prompt. When that sequence is absent, PEFT silently leaves the base model running instead: no warning, no error, and list_adapters()/active_adapters() still report the adapter as loaded. Callers get a plausible-looking but fabricated result with no way to tell it came from the base model. This adds a generation-time guard (LocalHFBackend, the only backend that loads adapters itself) that checks, right before dispatch, whether the loaded aLoRA's declared alora_invocation_tokens sequence occurs in the assembled prompt tokens. On a mismatch it raises AloraActivationError before any model call: - A direct adapter-function call (core.check_certainty, etc.) lets the error propagate; there is no fallback layer below it. - Requirement.validate()'s automatic aLoRA routing catches the error and falls back to LLM-as-a-judge, mirroring the existing fallback for an adapter that was never registered at all. Also adds a comment above the pinned adapter SHAs in catalog.py naming the accuracy/false-pass evidence a future pin bump must carry before promoting an aLoRA rung to the default path. Fixes generative-computing#1678 Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * fix(adapters): address code-review findings on the aLoRA activation guard Fixes issues surfaced by an independent 3-reviewer panel (2 tier-1, 1 tier-3) on PR generative-computing#1684: - Correct the PR description's and the docs page's overclaim that every high-level adapter-function wrapper now raises AloraActivationError by default. resolve_adapter() still hardcodes AdapterType.LORA (Epic generative-computing#929 Phase 2), so the guard only fires for a caller-registered composed aLoRA, Requirement/ALoraRequirement routing against one, or the deprecated IntrinsicAdapter shim from its second call onward. Reword the "aLoRA activation is not guaranteed by loading" section accordingly. - Fix AloraActivationError's pickle/deepcopy round-trip: pass the four positional fields to super().__init__() and move message formatting to __str__, mirroring the sibling AdapterSchemaMismatchError's existing convention. - Fire adapter_function_invocation_complete (outcome="error") when the guard aborts a call, matching AdapterMixin.adapter_scope's own revision/binding_type derivation. Without this, a guard-aborted invocation left no telemetry signal at all for exactly the scenario the guard exists to catch. - Add missing Raises: AloraActivationError entries to the docstrings of every public wrapper that can now surface it: core.py's check_certainty/requirement_check/find_context_attributions, rag.py's five adapter functions, _util.py's call_intrinsic, and huggingface.py's _generate_from_intrinsic. - Rename alora_invocation_sequence_present to _alora_invocation_sequence_present: it had no external consumer and was not re-exported, unlike its non-underscore siblings in the same module. - Document that alora_invocation_sequence_present's contiguous match is not scoped to the last (activating) occurrence, and correct the docs' claim that the IntrinsicAdapter shim is never covered -- it is exempt only on its first call. - Add end-to-end tests: the matching-sequence arm through the real _generate_from_intrinsic call site (previously only covered via the extracted helper), and a hook-capture test proving adapter_function_invocation_complete fires on a guard abort. Verified: full test/ -m "not qualitative" suite passes (4646 passed, 0 failed), ruff format/check clean, mypy clean, docstring-quality gate clean (uv run python tooling/docs-autogen/audit_coverage.py --quality --fail-on-quality --threshold 100), markdownlint clean. Each new test was confirmed to fail against the pre-fix code before the corresponding fix, then pass after. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * fix(adapters): report shim aLoRA guard aborts as pinned local_file A guard abort on a legacy IntrinsicAdapter read binding_type and revision from its inert _ShimWeightsBinding ("unknown", None), while a successful call reports through _IntrinsicPeftBinding ("local_file", catalogue SHA). Report the same fields on both, and cover the shim path with a test. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> * docs(adapters): trim transient issue references from aLoRA guard docs Drop issue/PR/acceptance-criteria references from docstrings and comments this PR added where the surrounding text already explains the point, and update the shim-removal pointer from the closed generative-computing#1144 to generative-computing#1621. Regression tests keep their generative-computing#1678 reference, and test data keeps its generative-computing#1679 source. Assisted-by: Claude Code Signed-off-by: Nigel Jones <jonesn@uk.ibm.com> --------- Signed-off-by: Nigel Jones <jonesn@uk.ibm.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull Request
Issue
Fixes #1678
Description
Before this PR: if an aLoRA's declared activation trigger doesn't appear in the
prompt Mellea builds, PEFT quietly leaves the base model running instead and says
nothing — no warning, no error.
list_adapters()still shows the adapter loaded,generation still returns a normal-looking result, and a requirement check that should
have failed can come back as a confident pass. Nothing tells you the adapter never
switched on. This is the exact failure diagnosed in #1679: the published
requirement-checkaLoRA can never activate on Granite 4.1, and its verdicts skewtowards false passes as a result.
After this PR: Mellea checks, on every generation call, whether the adapter it
just loaded can actually activate for the prompt it just built — before that call
happens, so a doomed generation is never attempted. If it can't activate, behaviour
depends on how you reached it:
core.check_certainty(...),core.requirement_check(...), etc. — now raisesAloraActivationErrorinstead ofreturning a fabricated score.
req(...)/Requirement.validate()'s automatic aLoRA routing — the pathbug(adapters): published
requirement-checkaLoRA never activates, and scores worse than the LoRA #1679's regression actually sits on — now falls back to LLM-as-a-judgeautomatically, with a one-line warning. This is the same fallback Mellea already
gives an adapter that was never registered at all; it now also covers "registered
but proven unable to activate."
The one real behaviour change: a direct adapter-function call against a broken
aLoRA now raises where it used to silently return a value. Everything else is
additive — no behaviour change for LoRA/embedded/server-mediated adapters, or for an
aLoRA that activates correctly (the overwhelming majority of calls).
The check itself has no false positives: it's an exact match against the same
alora_invocation_tokenssequence PEFT itself looks for, run against the exact prompttokens about to be sent to
model.generate().What changed
mellea/backends/adapters/_core.py— newAloraActivationError.mellea/backends/adapters/adapter.py— newalora_invocation_sequence_present(),the pure token-subsequence check.
mellea/backends/huggingface.py— the guard itself, wired into_generate_from_intrinsicbefore any model call;_generate_from_context'sRequirement-routing branch catches the error and falls back to LLM-as-a-judge.A guard abort still fires
adapter_function_invocation_completewith outcomeerror, so adapter metrics show the failure. For the legacyIntrinsicAdaptershim it reports
local_fileand the pinned catalogue SHA, matching what asuccessful call reports (from review).
mellea/backends/adapters/catalog.py— comment above the pinned adapter SHAsnaming the evidence AC3 requires before a pin bump promotes an aLoRA rung.
docs/docs/advanced/lora-and-alora-adapters.md— documents the guard and folds itinto the existing "adapter unavailable → fallback" routing rule.
test/backends/test_adapters/test_activation_guard.py(new), plus additions totest/backends/test_huggingface_unit.py— unit coverage for the token check andevery adapter shape, plus end-to-end tests through both real call sites.
Deviations from the literal issue spec
both, split by call site to match each caller's actual recovery options, rather
than one global policy.
review): flipping the resolve-time probe order to prefer aLoRA lives in
_register_available_composed_adapter, which is part of feat(adapters)!: prompt fallback for adapter functions without a trained adapter #1654 and doesn't exist onmainyet. AC5 itself gates that reorder on "once the guard exists and the evidencequestion in bug(adapters): published
requirement-checkaLoRA never activates, and scores worse than the LoRA #1679 is settled" — bug(adapters): publishedrequirement-checkaLoRA never activates, and scores worse than the LoRA #1679 is still open, so that reorder is explicitlynot in this PR.
resolve time Mellea doesn't yet have the chat-templated prompt the check needs.
Acceptance criteria
adapters itself (currently only
LocalHFBackend).reported (raises) or skipped (falls back).
via the AC4 comment.
catalog.py.Testing
Attribution
Adding a new component, requirement, sampling strategy, or tool?
NOTE: Please ensure you have an issue that has been acknowledged by a core contributor and routed you to open a pull request against this repository. Otherwise, please open an issue before continuing with this pull request.