Skip to content

Back off matchmaking when no opponent is found, instead of retrying every minute - #1254

Open
grounzero wants to merge 1 commit into
lichess-bot-devs:masterfrom
grounzero:matchmaking-no-opponent-backoff
Open

grounzero wants to merge 1 commit into
lichess-bot-devs:masterfrom
grounzero:matchmaking-no-opponent-backoff

Conversation

@grounzero

Copy link
Copy Markdown

What happens now

When Matchmaking.choose_opponent cannot produce an opponent, challenge logs
"No challenge will be created." and sets rate_limit_timer to a fixed 60
seconds. The commonest way to get there is the profile lookup failing:
get_public_data raises on a 429 from /api/user/{name}, choose_opponent
catches it as a generic Exception, and the next attempt comes a minute later.

api_get does record a 60-second delay for the rate-limited path, but
matchmaking does not consult it, so the lookup is asked again every minute for
as long as the endpoint stays limited, and each 429 seems to keep it limited.

We saw this on our bot: 37 consecutive profile lookups answered 429 over about
40 minutes, one a minute, while the bot's other endpoints answered 200. The bot
could accept challenges but could not create one. The endpoint was still
answering 429 an hour after we turned matchmaking off, and answered 200 some
hours later.

The change

The same escalation create_challenge already applies to a generic 429 on the
challenge endpoint: 60 seconds after the first attempt that finds no opponent,
doubling on each consecutive one up to 600, and back to 60 once an opponent is
found. The log line says when the next attempt will be.

Running it

We have run the equivalent as a wrapper around Matchmaking.challenge since
2026-09-24 (it doubles from 60 seconds with a cap of 960; this patch uses 600 to
match the challenge backoff). In the hours since, every lookup has succeeded, so
the backoff itself has not yet been exercised in production; it has been
exercised against a stand-in for Matchmaking in our own tests.

Tests

test_no_opponent_backs_off_and_resets_when_one_is_found in
test_bot/test_matchmaking.py drives challenge through six attempts that find
no opponent and checks the waits are 60, 120, 240, 480, 600 and 600 seconds,
then that finding an opponent resets the next wait to 60. It fails on master,
where every wait is 60. With the patch the suite gives 54 passed and 1 skipped,
the same skip as on master; ruff check --config test_bot/ruff.toml and
mypy --strict . are clean.

A failed opponent lookup, such as a 429 on the profile endpoint, was
retried every 60 seconds. It now doubles to 600, as the challenge
endpoint already does, and resets once an opponent is found.
@AttackingOrDefending

Copy link
Copy Markdown
Member

I don't think this is needed. Lichess says after a 429 you should wait a minute before the next request, and polling every minute is pretty light. And there should be many reasons why bot_username is None. Do you have the logs of the incident described as it seems quite surprising (the one where 37 consecutive profile lookups answered 429 over about
40 minutes)?

@grounzero

Copy link
Copy Markdown
Author

I don't think this is needed. Lichess says after a 429 you should wait a minute before the next request, and polling every minute is pretty light. And there should be many reasons why bot_username is None. Do you have the logs of the incident described as it seems quite surprising (the one where 37 consecutive profile lookups answered 429 over about 40 minutes)?

Thanks for looking at this. You're right on both counts. On bot_username: yes, that's a problem with the patch. choose_opponent returns None in three different cases, not just a rate limit. The patch backs off on all of them, so a bot with a narrow rating window could sit idle for ten minutes while opponents are online.

On the logs: Unfortunately the journal has rotated past that date . They said all 37 profile lookups got 429 over about 40 minutes, while other endpoints were fine.

I should have mentioned that a script of mine was hitting the same endpoint from the same address every 10 seconds at the time, so I can't tell how much of it was matchmaking.

The endpoint also stayed limited for an hour after matchmaking was off, which looks more like a long penalty than a retry keeping it alive. So the incident doesn't really show what I said it did. If any of it is worth keeping, it's narrower: only back off when the lookup itself was rate limited, which Lichess already tells us, and leave the other cases at 60 seconds. That just sends fewer requests to an endpoint that's already refusing them. Happy to push that if you'd take it, or to close this if you'd rather.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants