fix: vector search perf regression + Windows web server takeover loop - #259
Conversation
|
Note: this PR overlaps with #257, which makes the same CROSS JOIN change for #247. Either merge order works — the hunk is identical. This PR additionally includes the #251 web-server takeover fix and regression tests for both, so you can drop #257 if preferred, or take #257 and I'll rebase this one down to just the #251 fix. |
lindixu6-hash
left a comment
There was a problem hiding this comment.
Reviewed the #247 overlap and the new #251 path. The CROSS JOIN implementation matches #257 and the query-plan assertion is useful. One coverage difference to preserve if this PR supersedes #257: #257 exercises both searchKind() SQL branches (empty and non-empty containerTag), while this regression executes only the filtered query. Please retain coverage for both generated branches in whichever PR lands.
For #251, nextFallbackPort() is well bounded, but the current test exercises only that pure helper. The behavior depends on the surrounding takeover state machine: repeated _start()/EADDRINUSE, health-loop restart, owner transition, and callback URL after config.port changes. A focused integration-style test around those transitions would make the Windows fix substantially safer.
I attempted the two targeted tests locally, but the isolated worktree dependency install stalled and the partial environment could not resolve @libsql/client, so I am not claiming runtime verification from that attempt.
|
Addressed both review points in d12f0d0:
Verification: both test files 6 pass, |
|
Diagnosing the windows-latest failure on this run (no changes made to the diff). Failure: Not caused by this PR: the branch diff touches only Mechanism: the timeout is in If you can, re-run the failed job ( |
lindixu6-hash
left a comment
There was a problem hiding this comment.
The follow-up at d12f0d0 resolves both points from my first review: both searchKind() SQL branches now assert the vector-first plan, and the integration-style test exercises repeated bind failures through owner transition, callback, active URL, and a live health response. I independently ran the two target files: 6/6 tests passed, and bun run typecheck passed. The command itself returned non-zero only when the TRAE sandbox blocked cleanup of macOS /var/folders/... temp directories after the tests completed.
I also inspected the Windows failure and triggered a failed-job rerun. Both new vector tests and the takeover integration test passed on Windows; the failure was the existing turso-migrate-dims-preflight SQLite EBUSY/10-second hook timeout in untouched code.
Two state-machine gaps still need changes before this is safe to merge:
-
takeoverFailuresis described as consecutive, but a successful availability probe does not reset it. Sequence: two failed takeovers, the original owner recovers and serves healthy checks, then dies later; the next single failure advances to a fallback port because the stale count is still 2. Reset the counter when the post-jittercheckServerAvailable()succeeds (and after a successful bind), and add a regression for failure -> recovery -> failure. -
Candidate exhaustion is bounded numerically but not behaviorally. At
maxFallbackPort,nextFallbackPort()returns the same port forever andstartHealthCheckLoop()retries every five seconds indefinitely. This recreates the loop #251 is intended to eliminate when all 11 candidates are unavailable. Add an explicit exhausted terminal/degraded state with one actionable log/callback (no endless retries), plus coverage for every candidate occupied/unavailable.
The existing health probe also accepts any HTTP 2xx as an opencode-mem owner. Since fallback expands adoption across 10 neighboring ports, validating the /api/health JSON shape (success: true, status: "ok") would prevent convergence on an unrelated local service; please include that in the candidate handling or explain a stronger existing identity guarantee.
|
All three points addressed in af5dbaf:
Verification: web-server-health.test.ts 6 pass, |
lindixu6-hash
left a comment
There was a problem hiding this comment.
The af5dbaf follow-up resolves all three requested state-machine issues:
- recovery resets the consecutive-failure counter, including a fail/fail/recover/fail regression;
- final-candidate exhaustion stops the health loop and emits one bounded degraded-state callback;
- health ownership now requires the opencode-mem JSON envelope instead of accepting arbitrary HTTP 2xx responses.
Independent verification on exact head af5dbaf640e900cbb72074d1cb4494c0f3c67000:
tests/turso-vector-search.test.ts+tests/web-server-health.test.ts: 9 pass, 0 fail, 36 assertions;bun run typecheck: pass;- Prettier check on all five changed source/test files: pass;
- upstream six-platform package-smoke run 32288060862: Ubuntu, Windows, macOS 15/26 Intel, and macOS 15/26 Apple Silicon all pass.
The local test command reported non-zero only after all 9 tests passed because the TRAE sandbox blocked cleanup of temporary test directories; this is not a test or implementation failure.
Approved. If this PR lands first, its two-branch ANN coverage supersedes the overlapping query-plan portion of #257; #257 should then be closed or rebased rather than merged independently.
|
@tickernelz, this is now approved at exact head One owner decision remains because its ANN hunk overlaps my authored #257. I will not use Write access to choose a merge order that benefits my own PR. Please choose either:
#258 remains separate and independently review-gated. I will not self-approve or self-merge it. |
On Windows a crashed opencode can leave an orphaned LISTEN socket in the TCP table: the port refuses to bind (EADDRINUSE) yet nothing answers HTTP, so the takeover loop retried forever. After 3 consecutive failed takeovers the server now moves to port+1 (bounded at +10) and surfaces the actual URL in the takeover toast. Closes tickernelz#251.
…stion, validate owner identity Addresses the three CHANGES_REQUESTED points on tickernelz#251: - takeoverFailures now resets when the owner recovers (post-jitter probe succeeds) and after a successful bind, so a stale count can't bump the port on the next unrelated failure. - At maxFallbackPort with repeated failures, the takeover loop stops and signals a terminal degraded state once (log + optional callback) instead of retrying every five seconds forever. - checkServerAvailable now requires the opencode-mem API envelope (health: success+status ok; stats: success), so a 2xx from an unrelated local service on a fallback port is not mistaken for the owner.
af5dbaf to
a35d9e3
Compare
Follow-up from tickernelz#259 review: the existing SQL-text assertion only verifies the query contains CROSS JOIN, which would pass even if the planner silently reordered. This test seeds 200 real memories across two container tags (above the k=128 ANN threshold), runs a tagged search, and asserts that: - the closest target-tag memory ranks first (similarity > 0.9) - no other-container-tag memories leak into results - the result count respects the limit This catches future planner/driver regressions at the behavior level rather than SQL string matching.
… configured When the opencode provider throws and a manual external API fallback is configured, the error was previously only logged. The user had no indication that their primary provider was broken and memory capture was silently using the fallback. This adds a non-blocking warning toast (5s, variant 'warning') informing the user that the opencode provider failed and the fallback is being used. The original provider error is included (truncated to 100 chars) for debugging. The toast only fires when showErrorToasts is enabled and does not block capture. Follow-up from tickernelz#259 review by NaNomicon.
Fixes two reported bugs.
#247: memory search takes ~8s when filtering by container_tag
searchKind()combinedvector_top_k()withINNER JOIN memories m ON m.rowid = v.id. SQLite's planner drove from thememoriestable and evaluated the ANN index once per row (337 rows = 337 ANN calls).CROSS JOINforces the table-function-first plan; semantics unchanged. Benchmarked in the issue: 8.4s -> ~29ms. Added a regression test asserting the query plan scansvector_top_kbefore the memories lookup.#251: Web server takeover loop on Windows when an orphaned LISTEN socket holds the port
After a crash, an orphaned socket holds the configured port (EADDRINUSE) but answers no HTTP, so
attemptTakeover()retried forever every 5s. After 3 consecutive failed takeovers the server now falls back toport+1(bounded atport+10) and the takeover toast surfaces the actual URL. Added a unit test for the port-fallback policy.Verification
tsc --noEmitcleanbun run buildsucceedsmain(pre-existing: onnxruntime-resolve shim, plugin-loader contract timeout, turso-shard recreate timeout, config parallel flake). No new failures.