Repository navigation
Conversation
…of exiting When login or init failed, the daemon exited and container / service restart policies started it again, logging in to the server on every restart. Frequently failing deployments potentially generated thousands of such logins. The daemon now retries login and init within the same process, reusing the existing Mergin client. Consecutive failures back off exponentially; a successful sync resets the wait to sleep_time. New daemon option 'max_retries' limits consecutive failed retries of the start or of unexpected errors, after which the daemon exits as before and always sends a notification email. Sync errors after a successful start are still retried indefinitely.
This mitigates the very frequent unnecessary logins to server. After login, the auth token is stored in <working_dir>/.mergin_auth.json. Rejected tokens (401) are removed and a new login is done. Stored token is kept even when the working directory is removed by --force-init cleaning.
The server may reject a token before it expires. Failed logins are retried with longer backoff, as rejected credentials can not be fixed by retrying, to avoid locking the account.
|
|
||
| if "daemon" in config and "max_retries" in config.daemon: | ||
| max_retries = config.daemon.max_retries | ||
| if isinstance(max_retries, bool) or not isinstance(max_retries, int) or max_retries < 0: |
| return None | ||
|
|
||
| try: | ||
| mc = MerginClient( |
There was a problem hiding this comment.
Double check def validate_auth(self): method of MerginClient (see client.py in python-api-client), If it does not help here. I remember, we are using it in plugin for login magic.
| password=config.mergin.password, | ||
| plugin_version=f"DB-sync/{__version__}", | ||
| ) | ||
| AuthTokenStore.from_config().save(mc._auth_session["token"]) |
There was a problem hiding this comment.
I think be more defensive with get is better :) Just minor
| token_store.remove() | ||
| return None | ||
|
|
||
| if auth_token_expires_soon(mc): |
There was a problem hiding this comment.
Why not check expiration as soon as possible? Or directly in auth store?
| # by the server can not be fixed by retrying and too frequent failed logins may lock the account | ||
| MAX_LOGIN_RETRY_WAIT = 3600 | ||
| # default number of consecutive failed retries of startup (login / init) or unexpected errors before the daemon exits | ||
| DEFAULT_MAX_RETRIES = 10 |
There was a problem hiding this comment.
We can not use some max_failed_retries or similar to make it sure, that this is not retry of process? Just wondering.
| validate_config(config) | ||
| except ConfigError as e: | ||
| handle_error_and_exit(e) | ||
| max_retries = config.get("daemon.max_retries", DEFAULT_MAX_RETRIES) |
There was a problem hiding this comment.
I found we can use config.get("KEY", default=DEFAULT_MAX_RETRIES, cast="@int") or?
| if not args.skip_init: | ||
| try: | ||
| if not args.skip_init: | ||
| dbsync.dbsync_init(mc) |
There was a problem hiding this comment.
We do not want to exit when single run and init is not successful?
| mc = dbsync.create_mergin_client() | ||
|
|
||
| if not cleaned: | ||
| dbsync.dbsync_clean(mc) |
| dbsync.AuthTokenStore.from_config().remove() | ||
| mc = None | ||
|
|
||
| giving_up = max_retries and fatal_failures > max_retries |
There was a problem hiding this comment.
Why we are not giving up sooner?
Issue
Sometimes there are db-syncs logging in very frequently (e.g. every ~2 s). By design, db-sync logs in once per process, so the cause is processes being started over and over: a failing daemon restarted by Docker or another service manager, or
--single-runcalled frequently from cron. Each new process logged in again.Fixes
1. Retry failed start in-process instead of exiting
--force-initcleaning and init failures are retried inside the daemon, with exponential backoff fromsleep_timeup to 10 min. An existing client is reused, so a failed init doesn't log in again.DbSyncError.daemon.max_retriesoption (default 10,0= never exit). The daemon exits after this many consecutive failed retries of the start or of unexpected errors, and always sends a notification email when it gives up. Sync errors after a successful start are still retried indefinitely.2. Reuse auth token across restarts
<working_dir>/.mergin_auth.json.--force-initcleaning.3. Handle tokens rejected by the server