Repository navigation
Conversation
Rule text survives a suspend unchanged, so the watchdog reports the pf anchor intact while nothing translates, and the probe monitor is only armed by interface changes with a schedule frozen through the sleep. Detect the resume from the wall-clock gap and probe from there, with a bounded repair: at most two forced reloads across the retry schedule. Refresh the OS resolver before every probe - the probe aims at the first OS nameserver and passes when it has none, so a pre-sleep list makes it meaningless. Stand down for stabilization, which verifies on completion.
Consume resolver.split_dns and generate a policy rule per domain, plus its wildcard, so a suffix and its subdomains route to the OS resolver when no resolvers are given, or to the configured resolvers when they are. Explicit resolvers are exclusive: when all of them fail the query returns SERVFAIL rather than taking the OS catch-all or starting recovery, and it skips VPN split routing, which is auto-detected. Excludes and custom config keep precedence; a refresh regenerates the config, so removals leave no stale rule or upstream.
A failed handback restored ctrld's rule via addNRPTCatchAllRule, which also writes the GP-path rule whenever another GP rule exists - leaving CtrldCatchAll beside the administrator's catch-all. Restore only the local store while that child is still present.
… limit ctrld keeps a second log file, ctrld-journal.log, next to ctrld.log. The journal holds all warnings and errors and each event with the field journal=true. It stays on disk after a start of the service and after a self-upgrade. Each log file starts with a Log header line. Each log send upload starts with the same line, made at send time, and has a limit of about 16 MB.
The journal records the network state of the host and each change of that state. The support team reads one file to see the interfaces, the routes, the resolvers, and the recoveries around an incident. The same file shows the health of the query path. A change that touches only AirDrop or virtual interfaces runs no DNS work.
Co-authored-by: Ginder Singh <ginder@windscribe.com>
Co-authored-by: Ginder Singh <ginder@windscribe.com>
Co-authored-by: Ginder Singh <ginder@windscribe.com>
A successful macOS upgrade printed "could not get interface error=interface not found iface=en5" before reporting success, because the DNS cleanup logged every interface lookup failure at error level. An interface that is gone, an unplugged adapter or a torn down tether, has no DNS settings left to restore, so the skip is now a debug diagnostic naming the interface and the work it skipped. Only that condition is quiet. netInterface returns a sentinel error for a missing interface and no longer reports a failed enumeration as one, so any other lookup failure and any restoration failure on an interface that does exist stay at error level.
The genuine-failure cases fed a fabricated error straight to logIfaceLookupFailure, so they exercised only its own branching. A restoration failure never reaches that helper, and no test reached the enumeration classification in netInterface at all: discarding the enumeration error, or downgrading an actual restoration failure to debug, both kept the suite green. The host boundaries of the reset path are now variables, in the style the package already uses: the interface enumeration, the interface lookup, the DNS setters and the NetworkManager restore. The tests stub them, so a run changes no host DNS or NetworkManager state, and drive the real netInterface and resetDNSForRunningIface to assert that a failed enumeration keeps its cause and is not reported as a missing interface, and that a failure to restore the saved static config or to reset to DHCP on an interface that does exist stays at error level. Both surviving mutations, and three more, now fail.
NextDNS serves alternative DoH endpoints beside dns.nextdns.io: the ultralow and anycast variants of dns, dns1 and dns2. They are the same service and take the same client info headers, but isNextDNS matched the one host, so an upstream pointed at any of them was treated as a third party and sent no client info. Recognition now follows the parent domain through dns.IsSubDomain, as IsControlD already does for the ControlD domains. The label by label comparison is case insensitive and keeps lookalikes such as notnextdns.io and nextdns.io.example.com out. Based on the change proposed by Mike (Github username @mike406). See: #335.
Use Write rather than Stat to verify retained file handles are closed. Windows Stat returns ERROR_INVALID_HANDLE instead of os.ErrClosed. Retain final-cleanup content assertions and strict closed-error matching.
checkDnsLoop resolved with context.Background() and no dns.Client timeout, so the DNS client fell back to its own 2s default and the configured upstream timeout was ignored. The probes run serially on a one-minute ticker, so each unreachable local upstream stalled the loop for a fixed 2s no matter what the config asked for. This is most visible on Windows, where UDP to a closed loopback port does not surface an immediate refusal the way it does on Linux, so the probe waits out the full default. Apply the upstream timeout the way checkUpstreamOnce already does, and keep 2s for an upstream that configures none so the unconfigured case is unchanged. The probe context deliberately carries no logger. A resolver logs its own failure at error level with the endpoint in the message, which would place an Internal Domain resolver address in the retained journal; logUpstreamProbeFailure remains the reporter and bounds what the line may hold. TestInternalDomainsLoopCheckHidesResolverAddress covers this.
captureDebugMainLog swaps the process-wide mainLog, so every
mainLog.Load().Warn()/.Error() site in the package writes into the
buffer while it is installed, including goroutines left running by
other tests. retainedProbeLines returned every retained line and the
tests asserted field values over all of them, so one unrelated warn
failed the run.
Test_checkDnsLoopKeepsACustomKeyOutOfTheJournal hit this on Windows,
where the loop probe held the buffer open for two seconds:
upstream_probe_journal_test.go: field "upstream": got , want upstream.custom
The sibling tests were exposed the same way; they index retained[0]
after a strict length check, so a stray line fails them as got 2, want 1.
Let retainedProbeLines take the messages a caller owns and return only
those lines. wantNoProbeSecret keeps scanning every retained line,
which is what that leak check is for.
Decode networksetup entries once for interface lookup and DNS-target selection. Preserve first-enabled lookup and fail-closed DNS uniqueness as separate policies. Treat only the (*) index as disabled, preserving asterisks and whitespace in enabled service names. Add shared-parser, caller-policy and incomplete-read regressions.
This was referenced Sep 29, 2026
Internal Domains now support three modes. "os" (network default) is unchanged. "resolvers" becomes explicit resolver with network fallback, the default explicit selection: the configured resolvers are tried first, and a timeout, unreachable resolver, SERVFAIL, NXDOMAIN, REFUSED or NOTIMP hands the query to the network resolvers - matching VPN DNS servers, then domain-less VPN DNS servers, then the OS resolver's LAN nameservers. A valid empty answer stays final. "resolvers_only" keeps the previous strict behavior. The OS step uses a new LanOnlyQueryCtx: the OS resolver then asks only its LAN-classified nameservers (private, loopback, link-local, CGNAT) and sends nothing when it has none, so the private name never reaches a public nameserver, whether DHCP supplied it or it is ctrld's own public fallback. LanQueryCtx only dropped the latter. LAN-only queries get their own singleflight and hot-cache key, so they are never answered with an ordinary query's public answer. In fallback mode a configured resolver that the upstream monitor has marked down is skipped instead of waited on, so an endpoint off the organization network reaches the fallback without paying the resolver's 2s timeout on every query. Nothing else marks a generated resolver up again, so skipping one re-checks it in the background, at most once every 30s; an answer marks it up. Strict mode never skips. An absent mode with resolvers gets the fallback mode. An unrecognized mode with resolvers fails closed to explicit resolver only, so a wire value this build does not know never sends the domain to the Control D upstream; without resolvers it is dropped. Switching between the explicit modes is detected on refresh. Generated upstreams are now identified by an UpstreamConfig.InternalDomain marker that records their mode and that no configuration file can set, instead of by the internal_ key prefix. An upstream a local ctrld.toml or custom config defines as internal_foo, or as a copy of a generated one, is an ordinary upstream again: it keeps the OS-resolver catch-all, recovery and query-health grading. The "resolvers_only" wire value is provisional until the API agrees on it; the dashboard does not offer strict mode yet.
TestOSRecoverySkipSurvivesDebugTruncation passed time.Now() as the
argument after readLogReader, and Go evaluates arguments left to right,
so the lower bound was taken after the upload header was rendered. The
header check failed whenever a second boundary fell between the two:
header time = 11:13:44, want a time between 11:13:45.007 and now
Take the bound first, as the other splitUpload callers do.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Minor Release
This contains new features, security hardening and bug fixes.
Security
upstream.custominsteadAdded
resolver.split_dnsand routes each domain and its subdomains to the OS resolver, or to the configured resolvers when some are given. Explicit resolvers are exclusive: when all of them fail, the query returns SERVFAIL instead of falling back to the OS resolverctrld-journal.logsits next toctrld.logand keeps all warnings, errors, and the network state of the host with each change of that state (interfaces, routes, resolvers, recoveries). The journal survives a service start and a self-upgrade, so support can read one file to see what happened around an incident. Each log file and eachlog sendupload starts with a header line, and uploads are limited to about 16 MBctrld diagfor provisioning support. It collects the client version, MDM-managed preferences (macOS), the last provisioning result, the service state and API reachability in one copy-paste-safe report, with--jsonfor a machine-readable copy. It never prints the provisioning token--cd-org,--custom-hostname,--intercept-modeand conflicting flags now fail before any network call with input-stage codes (exit 20–29). Rejected provisioning codes reportTOKEN_INVALID,TOKEN_EXPIRED,TOKEN_LIMIT_REACHEDorTOKEN_DISABLED. Other terminal failures that used to crash with nothing to read now recordUNCLASSIFIED. The macOS package reports its own pre-flight failures (for example a missingProvisionToken) in the installer logdnsechotest.zscaler.comhealth-check name through the OS resolver and does not cache it, so Client Connector can validate its DNS path and enable Private Access. Every other name stays filtered by Control D, and--intercept-mode hardopts outnextdns.io, such as the ultralow and anycast variants, so these upstreams send client info too. Based on the change proposed by @mike406 (Allow including additional HTTP headers with any .nextdns.io subdomain #335)Changed
indeterminateand no longer reloads pfFixed
networksetupservice name. An unknown or unreadable native state still leaves DNS unchanged~.catch-all is no longer treated as a split-DNS routescutil, and stopped binding scoped resolvers to the default interface's source addresslog senduploads to the API's direct IP with the full body, and moved the log to the API-configuredlog_pathwithout losing earlier linesmongoexecutable is not available on the router