Repository navigation
Conversation
A failed opponent lookup, such as a 429 on the profile endpoint, was retried every 60 seconds. It now doubles to 600, as the challenge endpoint already does, and resets once an opponent is found.
|
I don't think this is needed. Lichess says after a 429 you should wait a minute before the next request, and polling every minute is pretty light. And there should be many reasons why |
Thanks for looking at this. You're right on both counts. On On the logs: Unfortunately the journal has rotated past that date . They said all 37 profile lookups got 429 over about 40 minutes, while other endpoints were fine. I should have mentioned that a script of mine was hitting the same endpoint from the same address every 10 seconds at the time, so I can't tell how much of it was matchmaking. The endpoint also stayed limited for an hour after matchmaking was off, which looks more like a long penalty than a retry keeping it alive. So the incident doesn't really show what I said it did. If any of it is worth keeping, it's narrower: only back off when the lookup itself was rate limited, which Lichess already tells us, and leave the other cases at 60 seconds. That just sends fewer requests to an endpoint that's already refusing them. Happy to push that if you'd take it, or to close this if you'd rather. |
What happens now
When
Matchmaking.choose_opponentcannot produce an opponent,challengelogs"No challenge will be created." and sets
rate_limit_timerto a fixed 60seconds. The commonest way to get there is the profile lookup failing:
get_public_dataraises on a 429 from/api/user/{name},choose_opponentcatches it as a generic
Exception, and the next attempt comes a minute later.api_getdoes record a 60-second delay for the rate-limited path, butmatchmaking does not consult it, so the lookup is asked again every minute for
as long as the endpoint stays limited, and each 429 seems to keep it limited.
We saw this on our bot: 37 consecutive profile lookups answered 429 over about
40 minutes, one a minute, while the bot's other endpoints answered 200. The bot
could accept challenges but could not create one. The endpoint was still
answering 429 an hour after we turned matchmaking off, and answered 200 some
hours later.
The change
The same escalation
create_challengealready applies to a generic 429 on thechallenge endpoint: 60 seconds after the first attempt that finds no opponent,
doubling on each consecutive one up to 600, and back to 60 once an opponent is
found. The log line says when the next attempt will be.
Running it
We have run the equivalent as a wrapper around
Matchmaking.challengesince2026-09-24 (it doubles from 60 seconds with a cap of 960; this patch uses 600 to
match the challenge backoff). In the hours since, every lookup has succeeded, so
the backoff itself has not yet been exercised in production; it has been
exercised against a stand-in for
Matchmakingin our own tests.Tests
test_no_opponent_backs_off_and_resets_when_one_is_foundintest_bot/test_matchmaking.pydriveschallengethrough six attempts that findno opponent and checks the waits are 60, 120, 240, 480, 600 and 600 seconds,
then that finding an opponent resets the next wait to 60. It fails on
master,where every wait is 60. With the patch the suite gives 54 passed and 1 skipped,
the same skip as on
master;ruff check --config test_bot/ruff.tomlandmypy --strict .are clean.