Skip to content

Release branch v1.5.8 - #344

Open
cuonglm wants to merge 29 commits into
v1.0from
release-branch-v1.5.8
Open

cuonglm wants to merge 29 commits into
v1.0from
release-branch-v1.5.8

Conversation

@cuonglm

@cuonglm cuonglm commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Minor Release

This contains new features, security hardening and bug fixes.

Security

  • Stopped DNS64-synthesized AAAA answers from claiming DNSSEC authentication. The AD bit is set only for DO clients, and only when both the AAAA denial and the A answer were authenticated upstream. DO and non-DO results are cached separately, and requests with CD set are never synthesized
  • Kept Internal Domain resolver addresses and operator-defined upstream keys out of the retained log journal. The upstream loop-check probe and recovery-skip events now report a redacted upstream.custom instead

Added

  • DNS64 synthesis for IPv6-only networks without CLAT. When the network has no IPv4 and no CLAT interface, ctrld discovers the NAT64 prefix (RFC 7050) and maps A records into it for AAAA queries that the policy already allowed. IPv4-only destinations stay reachable on these networks. Blocked answers are never synthesized
  • Organization Internal Domains from managed config. ctrld reads resolver.split_dns and routes each domain and its subdomains to the OS resolver, or to the configured resolvers when some are given. Explicit resolvers are exclusive: when all of them fail, the query returns SERVFAIL instead of falling back to the OS resolver
  • A log journal on disk. ctrld-journal.log sits next to ctrld.log and keeps all warnings, errors, and the network state of the host with each change of that state (interfaces, routes, resolvers, recoveries). The journal survives a service start and a self-upgrade, so support can read one file to see what happened around an incident. Each log file and each log send upload starts with a header line, and uploads are limited to about 16 MB
  • ctrld diag for provisioning support. It collects the client version, MDM-managed preferences (macOS), the last provisioning result, the service state and API reachability in one copy-paste-safe report, with --json for a machine-readable copy. It never prints the provisioning token
  • More provisioning failure codes. Invalid --cd-org, --custom-hostname, --intercept-mode and conflicting flags now fail before any network call with input-stage codes (exit 20–29). Rejected provisioning codes report TOKEN_INVALID, TOKEN_EXPIRED, TOKEN_LIMIT_REACHED or TOKEN_DISABLED. Other terminal failures that used to crash with nothing to read now record UNCLASSIFIED. The macOS package reports its own pre-flight failures (for example a missing ProvisionToken) in the installer log
  • Support for Zscaler Private Access on macOS DNS intercept. ctrld resolves Zscaler's dnsechotest.zscaler.com health-check name through the OS resolver and does not cache it, so Client Connector can validate its DNS path and enable Private Access. Every other name stays filtered by Control D, and --intercept-mode hard opts out
  • Recognition of every NextDNS endpoint under nextdns.io, such as the ultralow and anycast variants, so these upstreams send client info too. Based on the change proposed by @mike406 (Allow including additional HTTP headers with any .nextdns.io subdomain #335)

Changed

  • Limited OS-triggered recovery to real outages. An OS-only resolver failure no longer starts global recovery while a configured upstream is still healthy. ctrld refreshes its OS resolver list in memory instead. When every configured upstream is down, recovery probes those upstreams rather than waiting only for OS DNS
  • Made the macOS interception probe tell a real interception failure apart from a probe that could not run. Only a query that was sent but never reached ctrld triggers a pf reload. A missing target or a helper failure is now reported as indeterminate and no longer reloads pf
  • Improved recovery and DNS-target diagnostics. Network transitions, recoveries, probe results and DNS-target decision failures now carry correlation IDs and outcomes. Repeated identical failures are suppressed

Fixed

  • Fixed DNS on IPv6-only macOS networks with native CLAT, such as an iPhone hotspot or USB tethering. When DHCPv4 is unavailable and macOS confirms CLAT on the current primary service, ctrld now installs its DNS target. USB CLAT targets use the networksetup service name. An unknown or unreadable native state still leaves DNS unchanged
  • Made DNS target ownership on macOS safer. Ownership is saved before a target is installed, and a target that someone else cleared or edited is no longer reinstalled. A failed cleanup keeps its record, so a later sweep does not restore an outdated static DNS backup
  • Chose the most specific VPN split-DNS zone, so a parent VPN domain can no longer take queries for a child zone that another VPN serves. The Linux ~. catch-all is no longer treated as a split-DNS route
  • Kept link-local IPv6 resolver zones from scutil, and stopped binding scoped resolvers to the default interface's source address
  • Closed losing DoT and DoQ/HTTP/3 dials, and fully retired replaced DoT pools and HTTP/3 transports, so reloads and network changes no longer leak connections or sockets
  • Released runtime resources on startup failure and shutdown. Network-change callbacks and recoveries are drained before host DNS is restored, and host DNS is restored while the listeners are still serving
  • Serialized mobile controller teardown, so a restart can no longer overlap a session that is still stopping
  • Stopped a failed Windows NRPT handback from writing ctrld's rule next to the administrator's Group Policy catch-all rule
  • Skipped DNS restoration quietly for an interface that is gone, such as an unplugged adapter or a disconnected tether, instead of logging an error after a successful upgrade. Other restore failures are still logged as errors
  • Bounded the DNS loop-check probe by the configured upstream timeout, instead of a fixed 2s wait for each unreachable local upstream (most visible on Windows)
  • Retried log send uploads to the API's direct IP with the full body, and moved the log to the API-configured log_path without losing earlier lines
  • Skipped UniFi client discovery quietly when the mongo executable is not available on the router

cuonglm and others added 25 commits September 28, 2026 16:40
Rule text survives a suspend unchanged, so the watchdog reports the pf
anchor intact while nothing translates, and the probe monitor is only
armed by interface changes with a schedule frozen through the sleep.
Detect the resume from the wall-clock gap and probe from there, with a
bounded repair: at most two forced reloads across the retry schedule.

Refresh the OS resolver before every probe - the probe aims at the first
OS nameserver and passes when it has none, so a pre-sleep list makes it
meaningless. Stand down for stabilization, which verifies on completion.
Consume resolver.split_dns and generate a policy rule per domain, plus its
wildcard, so a suffix and its subdomains route to the OS resolver when no
resolvers are given, or to the configured resolvers when they are.

Explicit resolvers are exclusive: when all of them fail the query returns
SERVFAIL rather than taking the OS catch-all or starting recovery, and it
skips VPN split routing, which is auto-detected. Excludes and custom config
keep precedence; a refresh regenerates the config, so removals leave no
stale rule or upstream.
A failed handback restored ctrld's rule via addNRPTCatchAllRule, which
also writes the GP-path rule whenever another GP rule exists - leaving
CtrldCatchAll beside the administrator's catch-all. Restore only the
local store while that child is still present.
… limit

ctrld keeps a second log file, ctrld-journal.log, next to ctrld.log. The
journal holds all warnings and errors and each event with the field
journal=true. It stays on disk after a start of the service and after a
self-upgrade. Each log file starts with a Log header line. Each log send
upload starts with the same line, made at send time, and has a limit of
about 16 MB.
The journal records the network state of the host and each change of
that state. The support team reads one file to see the interfaces, the
routes, the resolvers, and the recoveries around an incident. The same
file shows the health of the query path. A change that touches only
AirDrop or virtual interfaces runs no DNS work.
Co-authored-by: Ginder Singh <ginder@windscribe.com>
Co-authored-by: Ginder Singh <ginder@windscribe.com>
Co-authored-by: Ginder Singh <ginder@windscribe.com>
A successful macOS upgrade printed "could not get interface
error=interface not found iface=en5" before reporting success, because
the DNS cleanup logged every interface lookup failure at error level.
An interface that is gone, an unplugged adapter or a torn down tether,
has no DNS settings left to restore, so the skip is now a debug
diagnostic naming the interface and the work it skipped.

Only that condition is quiet. netInterface returns a sentinel error for
a missing interface and no longer reports a failed enumeration as one,
so any other lookup failure and any restoration failure on an interface
that does exist stay at error level.
The genuine-failure cases fed a fabricated error straight to
logIfaceLookupFailure, so they exercised only its own branching. A
restoration failure never reaches that helper, and no test reached the
enumeration classification in netInterface at all: discarding the
enumeration error, or downgrading an actual restoration failure to
debug, both kept the suite green.

The host boundaries of the reset path are now variables, in the style
the package already uses: the interface enumeration, the interface
lookup, the DNS setters and the NetworkManager restore. The tests stub
them, so a run changes no host DNS or NetworkManager state, and drive
the real netInterface and resetDNSForRunningIface to assert that a
failed enumeration keeps its cause and is not reported as a missing
interface, and that a failure to restore the saved static config or to
reset to DHCP on an interface that does exist stays at error level.
Both surviving mutations, and three more, now fail.
NextDNS serves alternative DoH endpoints beside dns.nextdns.io: the
ultralow and anycast variants of dns, dns1 and dns2. They are the same
service and take the same client info headers, but isNextDNS matched
the one host, so an upstream pointed at any of them was treated as a
third party and sent no client info.

Recognition now follows the parent domain through dns.IsSubDomain, as
IsControlD already does for the ControlD domains. The label by label
comparison is case insensitive and keeps lookalikes such as
notnextdns.io and nextdns.io.example.com out.

Based on the change proposed by Mike (Github username @mike406).
See: #335.
Use Write rather than Stat to verify retained file handles are closed.
Windows Stat returns ERROR_INVALID_HANDLE instead of os.ErrClosed.
Retain final-cleanup content assertions and strict closed-error matching.
checkDnsLoop resolved with context.Background() and no dns.Client
timeout, so the DNS client fell back to its own 2s default and the
configured upstream timeout was ignored. The probes run serially on a
one-minute ticker, so each unreachable local upstream stalled the loop
for a fixed 2s no matter what the config asked for.

This is most visible on Windows, where UDP to a closed loopback port
does not surface an immediate refusal the way it does on Linux, so the
probe waits out the full default.

Apply the upstream timeout the way checkUpstreamOnce already does, and
keep 2s for an upstream that configures none so the unconfigured case
is unchanged.

The probe context deliberately carries no logger. A resolver logs its
own failure at error level with the endpoint in the message, which
would place an Internal Domain resolver address in the retained
journal; logUpstreamProbeFailure remains the reporter and bounds what
the line may hold. TestInternalDomainsLoopCheckHidesResolverAddress
covers this.
captureDebugMainLog swaps the process-wide mainLog, so every
mainLog.Load().Warn()/.Error() site in the package writes into the
buffer while it is installed, including goroutines left running by
other tests. retainedProbeLines returned every retained line and the
tests asserted field values over all of them, so one unrelated warn
failed the run.

Test_checkDnsLoopKeepsACustomKeyOutOfTheJournal hit this on Windows,
where the loop probe held the buffer open for two seconds:

    upstream_probe_journal_test.go: field "upstream": got , want upstream.custom

The sibling tests were exposed the same way; they index retained[0]
after a strict length check, so a stray line fails them as got 2, want 1.

Let retainedProbeLines take the messages a caller owns and return only
those lines. wantNoProbeSecret keeps scanning every retained line,
which is what that leak check is for.
Decode networksetup entries once for interface lookup and DNS-target
selection. Preserve first-enabled lookup and fail-closed DNS uniqueness
as separate policies. Treat only the (*) index as disabled, preserving
asterisks and whitespace in enabled service names.

Add shared-parser, caller-policy and incomplete-read regressions.
Internal Domains now support three modes. "os" (network default) is
unchanged. "resolvers" becomes explicit resolver with network fallback,
the default explicit selection: the configured resolvers are tried
first, and a timeout, unreachable resolver, SERVFAIL, NXDOMAIN, REFUSED
or NOTIMP hands the query to the network resolvers - matching VPN DNS
servers, then domain-less VPN DNS servers, then the OS resolver's LAN
nameservers. A valid empty answer stays final. "resolvers_only" keeps
the previous strict behavior.

The OS step uses a new LanOnlyQueryCtx: the OS resolver then asks only
its LAN-classified nameservers (private, loopback, link-local, CGNAT)
and sends nothing when it has none, so the private name never reaches a
public nameserver, whether DHCP supplied it or it is ctrld's own public
fallback. LanQueryCtx only dropped the latter. LAN-only queries get
their own singleflight and hot-cache key, so they are never answered
with an ordinary query's public answer.

In fallback mode a configured resolver that the upstream monitor has
marked down is skipped instead of waited on, so an endpoint off the
organization network reaches the fallback without paying the
resolver's 2s timeout on every query. Nothing else marks a generated
resolver up again, so skipping one re-checks it in the background, at
most once every 30s; an answer marks it up. Strict mode never skips.

An absent mode with resolvers gets the fallback mode. An unrecognized
mode with resolvers fails closed to explicit resolver only, so a wire
value this build does not know never sends the domain to the Control D
upstream; without resolvers it is dropped. Switching between the
explicit modes is detected on refresh.

Generated upstreams are now identified by an UpstreamConfig.InternalDomain
marker that records their mode and that no configuration file can set,
instead of by the internal_ key prefix. An upstream a local ctrld.toml or
custom config defines as internal_foo, or as a copy of a generated one,
is an ordinary upstream again: it keeps the OS-resolver catch-all,
recovery and query-health grading.

The "resolvers_only" wire value is provisional until the API agrees on
it; the dashboard does not offer strict mode yet.
TestOSRecoverySkipSurvivesDebugTruncation passed time.Now() as the
argument after readLogReader, and Go evaluates arguments left to right,
so the lower bound was taken after the upload header was rendered. The
header check failed whenever a second boundary fell between the two:

    header time = 11:13:44, want a time between 11:13:45.007 and now

Take the bound first, as the other splitUpload callers do.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants