Skip to content

Sync and advanced dashboard health audit: repeated repairs, ownership conflicts, private Brain update, false drift and source readiness #237

Description

@stuinfla

Summary

After running ak sync (twice, back to back), the dashboard (ak dashboard → Overview) still shows 6 warnings. Each warning card tells the user to run ak sync, but sync cannot clear any of them. I traced each one to its cause in the installed source. I found four separate defects. In two of them, sync prints ✓ for a step that did not fix what ak status flags. There are also three smaller dashboard and status problems.

# Subsystem Dashboard says What ak sync does Root cause (file) Severity
1 mcp legacy ruflo-keyed MCP registration → "sync migrates it" prints ✓ mcp: claude-flow registered, never removes the legacy entry lib/mcp.mjs register() / registrationStatus() disagree on what is migratable High: the same ~290-tool MCP server loads twice in every Claude session
2 blocks ruflo-providers-reference→upserted → "sync reconciles blocks" prints ✓ blocks(CLAUDE.md): in sync status/nudge build detector context without guidanceContextFromConfig() Medium: drift that can never clear
3 agentdb "agentdb CLI not installed" → "setup/sync installs it" fails every run: npm error EEXIST … ~/.npm-global/bin/agentdb bin collision with agentic-flow's agentdb proxy; detector checks only node_modules/agentdb Medium: sync fails on every run
4 ruvnet-brain "v4.3.22 → v4.3.28 available → sync refreshes the KB" fails every run: fresh-install activation refused because this brain contains a private overlay ak always runs the fresh-install path, even when the installer requires --update Medium: sync fails on every run, and the KB stays stale

Environment

agentic-kit 4.0.0-alpha.55 (next tag; latest tag is 4.0.0-alpha.0)
ruflo 3.45.0
agentic-qe 3.14.3
Claude Code 2.1.282
ruvnet-brain plugin 4.3.28, KB bundle 4.3.22
Node / npm v24.18.0 / 11.16.0 (Homebrew node@24)
OS macOS 27.0 (Apple Silicon)
kit.json hosts {"claude": true, "codex": false, "opencode": false}. The codex and opencode CLIs are installed but not enabled

Reproduction

ak sync                 # run 1
ak sync                 # run 2: identical result
ak status --deep        # still lists the same warnings
ak dashboard            # Overview → "6 warnings", each one saying "run ak sync"

Output of ak sync (identical on both runs):

sync plan (5 action(s)):
  • [ruvnet-brain] sync refreshes the KB — because: ruvnet-brain release v4.3.22, release v4.3.28 available
  • [agentdb] setup/sync installs it (pinned to ruflo's bundled agentdb) — because: agentdb CLI not installed
  • [mcp] sync migrates it to claude-flow at user scope — because: legacy 'ruflo'-keyed MCP registration present (user)
  • [blocks] sync reconciles blocks — because: 1 CLAUDE.md block(s) drifted: ruflo-providers-reference→upserted
  • [aqe-embedding] verify selected backend and repair missing opted-in local model

✗ ruvnet-brain: ✗ install stopped: fresh-install activation refused because this brain contains a private overlay
✗ agentdb: npm error with --force to overwrite files recklessly. ...
✓ mcp: claude-flow registered (user scope), 0 tool(s) denied per kit.json
✓ blocks(CLAUDE.md): in sync
✗ aqe-embedding: Local embedding setup incomplete. Recommended: local Ollama + MiniLM ...

Note the two ✓ lines. ak status, run straight afterwards, still reports mcp and blocks as drifted.


1. Legacy ruflo MCP entry is never migrated, but sync reports ✓ (High)

State on disk (~/.claude.json → mcpServers):

ruflo        → /Users/<me>/.npm-global/bin/ruflo  args: ["mcp"]
claude-flow  → ruflo                               args: ["mcp","start"]

claude mcp list shows both as ✔ Connected. Impact: every Claude Code session loads the full Ruflo tool surface twice (mcp__ruflo__* and mcp__claude-flow__*, about 290 tools each). This doubles deferred-tool noise and context cost, and tools can be routed to either copy.

Cause:

  • registrationStatus() in src/lib/mcp.mjs marks every user-scope legacy entry as auto-migratable:
    const autoMigratableLegacyScopes = topology.legacyRufloScopes.filter((scope) => scope === 'user');
    As a result, status/sections/mcp.mjs promises "sync migrates it to claude-flow at user scope".
  • register() only removes a legacy entry that passes replaceableRufloRegistration(). That check requires command === 'ruflo' and args deep-equal to ['mcp','start']:
    const removableLegacy = legacy && replaceableRufloRegistration(legacy) ? legacy : null;
    This legacy entry uses an absolute path and args ["mcp"], so it is skipped silently. Because claude-flow is already canonical, alreadyDesired is true, register() returns true, and sync prints ✓.

Suggested fix (either option):

  • (a) Treat any user-scope ruflo entry whose command resolves to the ruflo binary (basename ruflo, or realpath equal to which ruflo) and whose args start with mcp as replaceable, or
  • (b) If it is intentionally preserved, have registrationStatus() apply the same replaceableRufloRegistration() predicate, so that status says "preserved — remove with claude mcp remove ruflo -s user" instead of promising a migration. Sync should also print ⚠, not ✓, while a legacy entry remains.

A regression test should cover an absolute-path legacy entry (/…/bin/ruflo, args ["mcp"]).


2. blocks drift can never clear: status and sync evaluate the detector in different contexts (Medium)

ruflo-providers-reference is gated on codexEnabled, with a fallback to {type:'command', target:'codex'} (PATH).

  • Sync calls reconcileGuidance(), which wraps the context with guidanceContextFromConfig(cfg). That gives codexEnabled = false (from kit.json), so the block is correctly omitted and sync reports "in sync".
  • Status (src/commands/status/sections/blocks.mjs:23) and the nudge (src/lib/nudge.mjs:43) build a bare context:
    const ctx = { flags: { dualMode: bothHostsEnabled(cfg), opencodeEnabled: !!cfg.integrations?.hosts?.opencode } };
    codexEnabled is absent, so the detector falls back to "is codex on PATH?". It is (installed but disabled), so status reports →upserted on every run.

~/.claude/CLAUDE.md does not contain the block, which is correct for kit.json. So sync is right and status/nudge are wrong. The comment in nudge.mjs states its contract as "never disagrees with ak status". Both call sites should instead share sync's context.

Suggested fix: in status/sections/blocks.mjs and lib/nudge.mjs, call reconcileGuidance({ …, dryRun: true }), or at least guidanceContextFromConfig(cfg, ctx), so that all three paths share one context builder. Status also passes retiredForTarget(rowsReg, t.name) without knownTargets, while reconcileGuidance passes it. That is a second, latent divergence. A good regression test: kit.json with codex:false plus a codex binary on PATH should mean status reports no drift.


3. agentdb install fails on every sync with EEXIST (bin owned by agentic-flow); the detector misreports "not installed" (Medium)

$ ls -l ~/.npm-global/bin/agentdb
agentdb -> ../lib/node_modules/agentic-flow/dist/cli/agentdb-proxy.js
$ agentdb --version
agentdb v3.0.0-alpha.20          # ← the exact version ak wants to pin
$ ls ~/.npm-global/lib/node_modules/agentdb
No such file or directory

npm log (npm install --global --allow-scripts … agentdb@3.0.0-alpha.20):

error code EEXIST
error path ~/.npm-global/bin/agentdb
error File exists: ~/.npm-global/bin/agentdb

Cause:

  • globalVersion() in src/lib/agentdb.mjs checks only <globalRoot>/agentdb/package.json. It never considers that a working agentdb bin can be supplied by another global package (agentic-flow ships an agentdb proxy bin).
  • The install in setup.mjs (around line 341) then collides with that bin. npm refuses, and the same failure repeats on every sync forever.

Suggested fix:

  • Detect bin ownership first (readlink of $(npm prefix -g)/bin/agentdb). If agentic-flow owns it, either (a) accept it when agentdb --version matches the bundled target and report "provided by agentic-flow proxy", or (b) report a clear, specific conflict with a guarded remediation (for example, a Maintenance action that removes the proxy symlink and retries), rather than re-running a doomed install every sync.
  • Do not blindly add --force. That would clobber agentic-flow's bin.
  • ak x verify (verify.mjs:259) uses have('agentdb'), which is satisfied by the proxy. So verify and status already disagree about whether agentdb is present.

4. ruvnet-brain refresh fails on every sync: ak uses the fresh-install path, but an existing KB with private stores requires --update (Medium)

ak runs npx ruvnet-brain@latest with INSTALL_ARGS = ['--yes','--no-stack','--no-enhance','--no-nightly-prompt','--no-telemetry'] (src/lib/ruvnet-brain.mjs:33-34). The ruvnet-brain 4.3.28 installer (bin/install.mjs, around line 673) deliberately refuses that path when the live KB has private overlays:

if (stores.some((store) => store?.updateManaged === false)) {
  die('fresh-install activation refused because this brain contains a private overlay',
      `Run npx ruvnet-brain --update so the bundle updater preserves those private stores.`);
}

This machine's ~/.cache/ruvnet-brain/kb/SOURCE.json has 8 stores with updateManaged: false. grep -rn "'--update'\|updateManaged" src/ in agentic-kit returns nothing, so ak has no path that can succeed here. The KB is stuck at 4.3.22 while the plugin is at 4.3.28, and the brain's own SessionStart hook raises an "INSTALL ALARM: plugin and knowledge bundle are out of sync" on every session.

Suggested fix: when ~/.cache/ruvnet-brain/kb/SOURCE.json exists, and particularly when any store has updateManaged === false, invoke the updater (--update, together with the relevant --no-* opt-outs) instead of the fresh installer. Alternatively, surface the installer's own remediation text in the status/dashboard card instead of "sync refreshes the KB".


Smaller issues found while reviewing the dashboard

  1. aqe-embedding: sync and status disagree about severity. Sync prints ✗ aqe-embedding: Local embedding setup incomplete… and lists it under still failing. The dashboard shows the same subsystem as a grey informational · ("backend configured-unverified"), not counted among the 6 warnings. On this machine the cause is environmental: Ollama 0.20.4 is installed at /usr/local/bin/ollama but not running (ollama serve). The message could detect "installed but not running" and say that, instead of "Install Ollama". Status and sync should also agree on whether this is a failure.

  2. Permanent warnings with no remediation. Two warning cards count toward "attention advised" but offer no action, and sync cannot change them:

    • ruvnet-brain-plugin: "Automatic hook declarations differ from the exact reviewed 4.3.17/4.3.26 contracts for plugin version 4.3.28". src/lib/brain-hook-contract.mjs deep-equals the hooks against two hard-coded contracts, so every new brain plugin release trips a warning until ak ships a new contract. Consider an info level for "unreviewed newer version", or a documented way to review and accept it.
    • agent-browser: "external agent-browser 0.38.1 is outside Ruflo >=0.27.0 <0.28.0; preserved". This is intentional preservation (Track agent-browser + vibium as managed components — visibility, not a fix request #189), but it is permanently amber with no suggested action.

    Together with issues 1–4, all 6 dashboard warnings on this machine are warnings that ak sync cannot clear. For a user, "attention advised" can then never go green.

  3. GET /api/system returns a 22.9 MB JSON payload (22,923,686 bytes; served in about 0.1 s locally). Almost all of it is catalog: catalog.items is 13.8 MB for 1,258 items, catalog.artifacts is 4.2 MB for 4,986 entries, and catalog.consumerBindings is 2.1 MB for 5,841 entries. The server is fast, but the browser has to parse about 23 MB on the System tab. This machine has 73 tracked projects, so the payload will grow with project count. Consider pagination, or lazy-loading items and artifacts on demand.


What I did NOT verify

  • I did not apply any fix locally. The suggested fixes have not been tested.
  • I did not test on Linux, or with a clean npm prefix where agentic-flow is not installed.
  • I did not check whether the duplicate MCP registration causes functional conflicts (for example, double hook firing). I only confirmed that both servers connect and both tool sets load.
  • aqe-embedding was not retested with Ollama running.
  • The dashboard session token is omitted from this report on purpose.

Follow-up audit — 2026-09-26: exact ownership, unresolved repairs, and live-source health

This supplements the original report with a fresh read-only audit of installed 4.0.0-alpha.55. It separates product defects from local configuration and tightens the proposed fixes. Related dashboard evidence is tracked in #238. The original measurements above are historical; they were not all repeated in this audit.

Current verification and scope

  • macOS 27.0, build 26A428, Apple Silicon; Node v24.18.0; npm 11.16.0.
  • Installed Agentic Kit: 4.0.0-alpha.55. Live npm tags: next=4.0.0-alpha.55, latest=4.0.0-alpha.0; no newer published next was found in this check. Do not recommend switching to latest as an upgrade.
  • Live npm Ruflo latest=3.45.0; current kit status reports installed Ruflo 3.45.0.
  • Brain plugin 4.3.28 / installed KB SOURCE.json brainVersion 4.3.22. Eight entries carry updateManaged:false (private identities omitted).
  • The installed KB does contain coverage-integrity.mjs. A missing-validator incident in older sessions must not be assumed to explain this failure.
  • Persisted host intent is {claude:true,codex:false,opencode:false}. Codex is nevertheless used independently of Kit.
  • Persisted AQE intent is {mode:"endpoint",endpoint:"http://127.0.0.1:11434",provisioning:"ollama"}. GET /api/tags fails with connection refused.
  • Both configured advanced-dashboard sources are absent: .claude-flow/live-events.jsonl and .agentic-qe/live-events.jsonl.
  • Current launch directory is not a Git repository and has no project memory DB or AQE initialization. ruflo memory search --query ... --path <project>/.swarm/memory.db returned Database not found. This is absence of initialization, not evidence of a broken existing memory store.
  • ak sync --dry-run still lists five actions: Brain, AgentDB, MCP, blocks, and AQE embeddings. No mutating sync/setup was run during this follow-up.
  • No dashboard listener was found on port 7431 during the follow-up. Browser refresh was unavailable, so the earlier rendered dashboard observations are not presented as fresh screenshots.

A. Confirmed false block drift: minimal reproduction and acceptance test

src/commands/status/sections/blocks.mjs constructs only dualMode and opencodeEnabled. It does not call guidanceContextFromConfig. src/lib/blocks.mjs:reconcileGuidance does call that function and supplies codexEnabled:false from persisted intent. The ruflo-providers-reference detector therefore falls back to executable presence in status, but uses explicit intent during repair.

Read-only comparison:

import { loadKitConfig } from './src/lib/config.mjs';
import { reconcileGuidance } from './src/lib/blocks.mjs';
await reconcileGuidance({
  cwd: process.cwd(), cfg: loadKitConfig(), pkgRoot: process.cwd(), dryRun: true,
});

Run this from the Kit package/source root with equivalent user config. In the affected environment, all returned targets had changed:"", while ak status --json still reported ruflo-providers-reference→upserted.

Fix: share the reconciliation context/target selection among status, nudge, setup, and sync; also align the retiredForTarget known-target arguments. Do not add a providers block simply to silence this false warning.

Acceptance: with Codex installed but explicitly disabled in Kit, status, nudge, and dry-run reconciliation all agree on no drift. Repeat for enabled hosts, explicitly disabled Brain/AQE, and re-scoped custom blocks. Actual changes still produce actionable drift.

B. MCP preservation is correct; promising an automatic repair is not

Current user-scope entries are:

ruflo        command=<npm-global>/bin/ruflo  args=["mcp"]             env={}
claude-flow  command=ruflo                  args=["mcp","start"]    env keys=[AGENT_BROWSER_CONFIG]

src/lib/mcp.mjs:replaceableRufloRegistration accepts only the exact canonical command/argument shape. register can return success because the canonical entry already exists, while leaving the legacy entry. Status's auto-migratable classification is broader than that mutation predicate.

Fix: use one ownership/migratability predicate in planning, status, and execution. Recognize verified equivalent global binary paths and specifically supported legacy arguments only when safe. Otherwise say “custom registration preserved; manual migration required” and disclose both entries. Confirm replacements, retain rollback, and verify final topology.

Safety correction to the original suggestion: basename ruflo or arbitrary arguments beginning with mcp alone are not sufficient authority to remove a registration. Preserve unknown environment variables, wrappers, custom arguments, and foreign registrations.

Acceptance: tests for absolute global paths, mcp versus mcp start, custom env, wrappers, project scope, failed registration with restoration, and canonical-already-present. A preserved entry must never be described as migrated. Duplicate connection/tool counts and cost impact require separate runtime evidence; this follow-up does not remeasure them.

C. Brain repair must distinguish update from replacement

src/commands/sync.mjs invokes heal.installRuvnetBrain({force:true}). src/lib/heal.mjs builds an installer invocation with the selected release and --force, but no --update. An existing private-overlay installation rejects fresh activation. This guard is protecting data and must remain.

Fix: select the supported existing-install update flow; preserve telemetry/nightly opt-outs; verify the requested release is actually applied, rather than stamping it merely because a command exits zero. If the installer cannot honor the selected target or safely preserve overlays, surface the precise blocked condition.

Acceptance: fixture with public and updateManaged:false private stores; update to a different public release; private bytes/provenance unchanged; public bundle/plugin alignment verified; interrupted/failed update retains a usable prior generation; no fallback that overwrites private data. Cover fresh install separately. Existence of the trusted validator must be verified, but should not be misdiagnosed as missing here.

D. AgentDB: command ownership, compatibility, and explicit policy

The current command symlink belongs to agentic-flow/dist/cli/agentdb-proxy.js; no standalone global agentdb package is present. src/lib/agentdb.mjs:present checks only the standalone package. healAgentdb then attempts a global version-pinned install and npm encounters the owned executable.

Fix: inventory executable ownership first. Distinguish standalone package, proxy, compatible required capabilities, and unresolved conflict. A version string is not sufficient to establish that a proxy supports harvest or uses the intended store/schema. Verify required operations against isolated fixtures before accepting it. Provide an explicit opt-out/ownership boundary for installations that mandate Ruflo-only project-memory access and disallow standalone installs/version pinning. That mandate is a local policy conflict, not a universal claim that standalone AgentDB is invalid.

Acceptance: clean prefix, compatible proxy, incompatible proxy, absent target package, foreign executable, and explicit unmanaged/opt-out cases. No blind --force, bin deletion, disconnected store, or repeated doomed install. status, verify, and sync must agree on which capabilities are available and which remain blocked.

E. Additional product defect: missing structured sources can look healthy

The earlier advanced launch used:

ak dashboard --no-open --port 7431 \
  --live-source 'ruflo=.claude-flow/live-events.jsonl' \
  --live-source 'aqe=.agentic-qe/live-events.jsonl'

Both files are absent. Earlier Sources displayed one file / zero events / zero errors / “ok” for each. Source inspection explains why:

  • src/lib/live/live-sessions-service.mjs:#add increments files when it registers a tailer, not when an actual file is successfully read.
  • #reconcile marks the adapter status:'ok' after tailer.reconcile() returns.
  • src/lib/live/jsonl-tailer.mjs:reconcile catches ENOENT and returns without reporting failure or presence state.

Fix: model configured-source count, present/readable-file count, ingestion state, accepted/rejected-record counts, and last accepted event independently. Late creation is legitimate; report “awaiting file” rather than healthy ingestion. Explain that --live-source observes an existing producer and does not enable Ruflo/AQE event production. structured-adapter.mjs requires a session ID, actor ID, and action; expose schema rejection diagnostics without exposing record contents.

Acceptance: absent file, empty file, valid record, schema-invalid record, malformed JSON, permission failure, removal/recreation, truncation, rotation, and multiple sources with mixed health. A missing or non-ingesting source must not be represented as verified operational. Test the adapter with fixtures, then prove a real producer event reaches the UI; synthetic records alone are not an end-to-end acceptance test.

F. Additional product defect: post-sync convergence excludes unresolved repairs

At the end of src/commands/sync.mjs, remaining initially includes rows whose level is fail, plus a special companion case. Recorded apply failures are added later, but a repairable warning that persists despite a successful step result can be omitted. This can produce “converged — no failing subsystems” even though status still proposes repairs.

Fix: separate applied steps, preserved/unmanaged advisories, blocked actions, and unresolved planned repairs. Recompare the post-state to the original plan using shared ownership predicates. Do not simply fail on every amber row: intentionally preserved external versions and unreviewed contracts are different from a promised repair that did not happen.

Acceptance: successful helper result plus persistent actionable warning yields partial/not-converged status and names the unresolved action. An intentionally preserved advisory does not trigger an endless repair loop. A real failure remains nonzero. Two successive runs after repair are idempotent. Pure verification steps, such as embedding verification, may run again, but must not repeatedly attempt a known impossible repair.

Local setup conditions: improve diagnostics; do not automatically mutate

  1. Non-Git folder: ak setup --help explicitly says project setup normally requires .git, with --project available. Missing memory/AQE here is therefore expected until opted in, not an initialization bug by itself. Explain machine versus project readiness and point to the exact project initialization action. A dashboard-only launch folder does not need every project feature initialized.
  2. Codex used outside Kit: separate installed/authenticated/in-use state from Kit routing intent (details in Dashboard: Codex shown as 'Disabled' while it's the main host in use; Live omits running sessions in non-git folders; 6 more display/diagnostic defects (with screenshots) #238). Do not automatically enable routing because activity is observed.
  3. Ollama unavailable: selected loopback endpoint is refused. Explain selected backend, reachability, and synthetic embedding verification separately. Supported in-process mode is an alternative, but dependencies/model cache and existing corpus compatibility must be validated before switching. Do not create login services, migrate vectors, or silently enable a new backend.
  4. Dashboard lifecycle: the command intentionally runs foreground and stops with its owning process. The absent listener is not itself a dashboard defect. Document this clearly rather than implying a persistent service.
  5. Browser/plugin warnings: compatibility-range or unreviewed-hook warnings do not prove actual runtime failure. Provide an actionable review path and clearly state the evidence limit; do not downgrade or accept contracts just to make the dashboard green.

Recommended maintainer sequence

  1. Fix shared status/repair predicates and add regression fixtures for block context, MCP ownership, missing sources, and post-sync convergence.
  2. Implement safe Brain update selection and AgentDB conflict handling, preserving foreign/private state.
  3. Improve readiness diagnostics for host routing, project initialization, and the chosen embedding backend.
  4. Verify on a clean disposable macOS environment and an upgraded environment with existing bins, private overlays, custom MCP entries, and a non-Git launch folder. Keep fixtures isolated from real user data.
  5. Publish a forward update with an explicit remediation guide and prove repeat-run idempotence. Measure any claimed performance improvement with a representative workload.

Limits of this follow-up

No suggested patch was implemented or tested. No Keychain, login item, permission, corpus, MCP, or user-guidance changes were made. No new actual ak sync was run; the mutating failures above are from the original report, while source/state/dry-run checks are fresh. No Linux/Windows regression, live embedding result, real structured producer event, fresh browser render, or performance benchmark was established. Paths, private store identities, transcript contents, credentials, and dashboard tokens are omitted.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions