Skip to content

fix(ragas): resolve the judge model through the library default - #226

Open
fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/ragas-default-model
Open

fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/ragas-default-model

Conversation

@feiiiiii5

Copy link
Copy Markdown

The RAGAS scorers pick a different judge model from every other scorer in the library when nothing is configured.

_get_model is what all seven RAGAS scorers call to resolve their model. It fell back to a hardcoded constant:

# Deprecated: Use init(default_model="...") to configure the default model instead.
# This was previously "gpt-5-nano" but now defaults to the configured model.
DEFAULT_RAGAS_MODEL = "gpt-5-nano"

def _get_model(model: str | None) -> str:
    if model is not None:
        return model
    configured_default = _default_model_var.get(None)
    if configured_default is not None:
        return configured_default
    return DEFAULT_RAGAS_MODEL          # <- the deprecated path is still live

So with no explicit model= and no init(default_model=...):

init()                  # resets to the library defaults
get_default_model()     # 'gpt-5-mini'
_get_model(None)        # 'gpt-5-nano'   <- not the library default

Every RAGAS scorer is judged by a different model than every other scorer in the same library, with nothing printed. The comment above the constant already says the hardcoded fallback is deprecated. It was labelled that way by the commit that introduced init(default_model=...) (#161), whose message says the change was "replacing the hardcoded gpt-4o default" — but only half of it landed. That commit moved the scorers off model=DEFAULT_RAGAS_MODEL as a default argument while leaving _get_model returning the constant.

The TypeScript implementation never had a separate fallback — js/ragas.ts resolves through getDefaultModel() — so the same evaluation gets a different judge on the two implementations, and therefore different scores.

_get_model now delegates to get_default_model(). That also removes the manual _default_model_var lookup it was doing by hand. DEFAULT_RAGAS_MODEL had no other reader anywhere in the repo, so it goes with it, along with the now-unused import.

Test: py/autoevals/test_ragas_default_model.py fails on main at 9546b28 with AssertionError: assert 'gpt-5-nano' == 'gpt-5-mini' and passes here. It covers the unset case, an explicit model=, and a configured init(default_model=...), so the two paths that already worked stay pinned. Excluding the LLM-backed test files, the suite goes from 25 to 28 passed; the 4 failures and 8 errors are unchanged and all come from a missing OPENAI_API_KEY or a missing litellm in this environment. I did not run the LLM-backed tests, which need credentials. black, ruff and codespell are clean on both files at the versions in .pre-commit-config.yaml.

This changes which model runs by default, so if you would rather keep gpt-5-nano for RAGAS the right place to decide that is a default_model for the RAGAS group rather than a second fallback — but I did not want to make that call unilaterally, and the two implementations should not disagree either way.

_get_model is what every RAGAS scorer calls to resolve its judge model.
It fell back to DEFAULT_RAGAS_MODEL, a hardcoded "gpt-5-nano", when the
caller passed no model and had not called init(default_model=...):

    init()                    # resets to the library defaults
    get_default_model()       # 'gpt-5-mini'
    _get_model(None)          # 'gpt-5-nano'  <- not the library default

So a RAGAS scorer silently judged with a different model than every other
scorer in the library. The TypeScript implementation resolves through
getDefaultModel() and never had a separate fallback, so the same
evaluation produced different scores on the two implementations.

The constant was already labelled deprecated, by the commit that added
init(default_model=...) (braintrustdata#161): "This was previously 'gpt-5-mini' but now
defaults to the configured model." Only the first half of that change
landed -- the scorers stopped taking model=DEFAULT_RAGAS_MODEL as a
default argument, but _get_model still returned the constant.

_get_model now delegates to get_default_model(), which also removes the
manual _default_model_var lookup. The constant had no other reader, so
it and that import go with it.

Test: py/autoevals/test_ragas_default_model.py fails on main with
"assert 'gpt-5-nano' == 'gpt-5-mini'" and passes here. Excluding the
LLM-backed files, the suite goes from 25 to 28 passed; the 4 failures and
8 errors are unchanged and are all missing OPENAI_API_KEY or a missing
litellm.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant