The Zero-Overhead, Agent-Ready AI Memory Engine.
"A database returns what was written. A memory engine reconstructs what is reachable from a cue — under decay, association, and tier physics."
— Memory Fundamentals MF-001
Legacy AI stacks bolt memory onto stateless vector databases — storage without cognition. Spector is a cognitive memory engine for AI agents: it remembers, forgets, consolidates, and forms associations across working, episodic, semantic, and procedural tiers, then retrieves with fused semantic and hybrid scoring at sub-millisecond latency. Traces are linked by co-activation, temporal, and entity edges, so recall can surface what is related, not only what matches a vector. Connect any agent through the built-in MCP server, call it over REST/gRPC, drive it from the Python or TypeScript SDKs, or embed it directly in the JVM. Every user, agent, or tenant is physically isolated in its own on-disk namespace — true data separation, not a shared-store filter. Under the hood, Java Project Panama and the Vector API deliver SIMD scoring with measured near-zero garbage-collection pressure, and no external database.
Connect an agent, install an SDK, or launch a local node in seconds:
Run instantly via NPX — connects to a running local Synapse daemon on :7070 if healthy, or runs an embedded memory kernel (requires OpenJDK 25+):
# Agent runner: connects to local daemon or launches in-process MCP kernel
npx -y @spectrayan/spector mcp
# Or build and launch from source:
mvn clean package -pl synapse/spector-cli -am -DskipTests
java --enable-preview --add-modules=jdk.incubator.vector -jar synapse/spector-cli/target/spector.jar mcpInteract with Spector over HTTP / SSE from your language of choice:
Python:
pip install spector-clientfrom spector_client import SpectorClient, MemoryTier
client = SpectorClient.builder().with_rest("http://localhost:7070").build()
client.memory.remember(
text="User prefers concise answers and dark mode",
tier=MemoryTier.SEMANTIC,
tags=["preferences", "ui"],
)
memories = client.memory.recall("user preferences", top_k=3)TypeScript / Node.js:
npm install @spectrayan/spector-clientimport { SpectorClient, MemoryTier } from '@spectrayan/spector-client';
const client = SpectorClient.createDefault('http://localhost:7070');
await client.memory.remember({
text: 'User prefers concise answers and dark mode',
tier: MemoryTier.SEMANTIC,
tags: ['preferences', 'ui'],
});
const memories = await client.memory.recall('user preferences', { topK: 3 });docker compose up -d # Core engine (:7070) + Cortex Neural Dashboard (:7700)
docker compose --profile embeddings up -d # Adds local Ollama container for embeddings# Linux / macOS (POSIX)
curl -fsSL https://raw.githubusercontent.com/spectrayan/spector/main/scripts/install.sh | sh
# Windows (PowerShell)
irm https://raw.githubusercontent.com/spectrayan/spector/main/scripts/install.ps1 | iex(For Homebrew, Scoop, Helm, and Java embed instructions, see the Installation Guide).
Connect Spector to your favorite AI coding assistant or desktop agent in seconds:
Add to claude_desktop_config.json:
{
"mcpServers": {
"spector": {
"command": "npx",
"args": ["-y", "@spectrayan/spector", "mcp"]
}
}
}Add to .cursor/mcp.json:
{
"mcpServers": {
"spector": {
"command": "npx",
"args": ["-y", "@spectrayan/spector", "mcp"]
}
}
}Add to ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"spector": {
"command": "npx",
"args": ["-y", "@spectrayan/spector", "mcp"]
}
}
}claude mcp add spector -- npx -y @spectrayan/spector mcpEvery stored trace is a fixed record:
embedding, tag Bloom, valence, base importance, write time, recall count.
Decay (power-law, bucketed lookup in the hot loop):
with
Rank (after live-bit, tag-containment, and valence-range gates):
Association (after top-$K$, not inside the SIMD product): neighbours on the co-activation graph enter with
Similarity is dense retrieval. Decay and importance change what remains reachable. The graph walk changes what else is admitted. That is the whole distinction from a vector store. Details and parameters: scoring-pipeline.md and MF-001.
Spector is structured as a modular four-tier engine: a sealed off-heap kernel, graph-backed association, and in-process MCP between the SIMD path and the agent runtime:
- Nucleus (Foundation): Core configurations, off-heap storage layouts (Panama MemorySegment), and standard utilities.
- Memory (Cognitive Engine): The flagship hybrid retrieval and cognitive memory system combining dense vector, sparse (SPLADE/Li-LSR), keyword (BM25), association graphs, and consolidation pipelines.
- Synapse (Gateway & APIs): Spring Boot entry points, Armeria-based REST/gRPC gateways, and stdio/HTTP Model Context Protocol (MCP) servers.
- Cortex (UI): Three.js and Angular-powered observability dashboard for real-time visualization of memory graphs, decay, and search metrics.
For a comprehensive analysis of the system architecture, data flows, thread scheduling model, and detailed Mermaid diagrams, see the Architecture Overview Docs.
Spector is an MCP-native cognitive memory engine — not an afterthought adapter. The MCP server runs in-process with the memory system (zero network, zero serialization), giving agents direct SIMD-accelerated access to 16 tools across memory storage, recall, and introspection.
| Spector (MCP-native) | Typical MCP adapter | |
|---|---|---|
| Architecture | Memory + MCP in one JVM | Python wrapper → HTTP → DB |
| Memory recall | Ultra-low latency (fused scoring) | 50–200ms (Mem0/Letta/Zep) |
| Tools | 16 (memory, recall, introspection) | 3–5 basic CRUD |
| Cognitive features | Decay, Hebbian, consolidation, valence | Key-value store |
| GC pressure | Zero (Panama off-heap) | Full GC overhead |
Spector Memory gives AI agents a cognitive memory layer that can remember, forget, consolidate, and associate — with sub-millisecond in-process recall and measured near-zero garbage-collection pressure.
| Capability | What it does |
|---|---|
| Four memory tiers | Working → Episodic → Semantic → Procedural |
| ⚡ Ultra-Fast Recall | Sub-millisecond in-process execution (vs. 50–200ms for Mem0/Letta/Zep) |
| 🔗 Fused SIMD Scoring | Similarity × importance × decay in a single pass — no truncation trap |
| 🛏️ Offline Consolidation | Demote or drop low-utility traces, rebuild partitions |
| 😱 Valence Weighting | Signed valence weight in the recall score |
| 🚫 Zero GC | 100% off-heap Panama storage (≤0.01% overhead measured) |
| Capability | What makes it different |
|---|---|
| 🧠 Cognitive memory tiers | Working → Episodic → Semantic → Procedural, with decay, consolidation, and valence — memory that behaves like memory, not a key-value store |
| 🔗 Associative memory graphs | Co-activation, temporal, and entity edges — recall surfaces what is related, not only what matches |
| 🤖 In-process MCP server | Cognitive tools over stdio + Streamable HTTP — agents call memory directly, zero network hops |
| ⚡ Fused SIMD scoring | Similarity × importance × decay in one pass — ultra-fast in-process fused recall |
| 🔍 Hybrid retrieval | Dense + sparse + late-interaction reranking, fused with RRF, with graceful degradation |
| 🔒 Physical namespace isolation | Every user, agent, or tenant's memory lives in its own on-disk directory tree — true data separation, not a logical filter — hash-sharded to millions of namespaces, encrypted at rest (AES-256-GCM) |
| 🧊 Zero-GC off-heap storage | 100% off-heap via Panama — ~0.01% GC overhead measured |
| 🗜️ Quantization | SVASQ-8/4 + IVF-PQ — 4–32× compression at ~99.5% recall |
| 🖥️ GPU acceleration | Optional CUDA via Panama FFM, zero-copy transfer |
| 📦 Flexible deployment | Embedded JAR, standalone, or distributed |
🎥 Watch the Neural Graph in action →
📊 Dashboard — 12+ live cognitive panels
Real-time scoring pipeline, SIMD lanes, decay curves, vector space, Hebbian graph, cognitive profiles, live metrics — all rendered with Three.js, Canvas 2D, and Angular Signals.🌌 Graph Explorer — 3D neural galaxy
Interactive 3D graph with glowing star nodes, Hebbian/temporal/entity edges, fly-to navigation, and real-time topology stats.🧠 Memory Table — browse & manage memories
Full CRUD with tier filtering, importance bars, valence indicators, synaptic tags, recall counts, and bulk actions.🔬 Memory Detail — deep cognitive inspection
Identity, cognitive state (importance/valence/arousal), synaptic tags, and full relationship graph (Hebbian associations, temporal chains, entity links).Prerequisites: OpenJDK 25+, Maven 3.9+
| Component | Minimum JDK | Notes |
|---|---|---|
Client SDK (spector-client) |
25 | Thin HTTP client; Java 25 reactor aligned |
Server image (deploy/docker) |
25 | Ships with --enable-preview --add-modules=jdk.incubator.vector |
Embedded engine (spector-memory) |
25 + Vector API | Requires --add-modules jdk.incubator.vector --enable-native-access=ALL-UNNAMED --enable-preview |
Note: There is no startup preflight for the Vector API module.
SimdCapabilityreports SIMD availability but does not probe or fail fast.spector-kernel'smodule-info.javahard-requiresjdk.incubator.vector, andSimdCapabilitytouchesFloatVectorin a static initializer — so a missing module causes aNoClassDefFoundErrorat class load, not a graceful scalar fallback. The scalar fallback atAcceleratorRegistry/SmartSimilarityKernelcovers a missing accelerator, not a missing module.
git clone https://github.com/spectrayan/spector.git
cd spector
mvn clean test # Build reactor & run tests
mvn package -pl synapse/spector-cli -am -DskipTests # Package standalone spector.jarLaunch the standalone engine:
java --add-modules jdk.incubator.vector \
--enable-native-access=ALL-UNNAMED --enable-preview \
-jar synapse/spector-cli/target/spector.jar doctorAll numbers measured on Intel Core Ultra 9 285K, Java 25, AVX2 256-bit.
| Benchmark | Result | Notes |
|---|---|---|
| Vector search p50 | 88–143µs | 10K–100K docs, HNSW M=16 |
| Cognitive recall | Ultra-low latency | Hardware-accelerated in-process SIMD |
| Peak QPS (16 threads) | 61,011 | Concurrent vectorSearch |
| GC overhead | 0.01% | 1 pause / 100K searches |
| vs. Python MCP servers | 23–113× faster | In-process SIMD, zero network |
| I want to... | Start here |
|---|---|
| Get started in 30 seconds | Quick Start · Installation Guide |
| Connect an AI agent | MCP Server Setup · Claude & Cursor Guide |
| Use client SDKs | TypeScript SDK · Python SDK · Java SDK · Spring AI |
| Explore cognitive memory | Memory Overview · Cognitive Profiles · Scoring Pipeline |
| Deploy to production | Docker & Compose · Kubernetes Helm · Terraform Cloud |
| Review Architecture & ADRs | Architecture Decision Records (ADRs) · System Architecture |
| Contribute & Governance | Developer Guide · Contributing Guide · Project Governance |
We welcome contributions of all kinds — code, docs, tests, benchmarks, and ideas!
- 🏛️ Open Governance → See GOVERNANCE.md for our 4-tier contributor ladder and decision mechanics
- 📋 Architecture Decisions (ADRs) → See docs/adr/README.md for the living ADR framework and proposal process
- 🔧 Want to contribute code? → See CONTRIBUTING.md
- 🐛 Found a bug? → Open an Issue
- 💡 Have an idea? → Start a Discussion
- 🤖 AI-assisted PRs welcome!
Spector is free and open-source software licensed under the Apache License 2.0. All modules, client SDKs, tooling, and connectors are distributed under Apache 2.0.
For branding and trademark guidelines, see the NOTICE file.
See SECURITY.md for our security policy and vulnerability reporting.
See ACKNOWLEDGMENTS.md for our Open Source Contributors hall of fame, as well as credits to the cognitive science researchers, open-source frameworks, and AI coding tools that made Spector possible.
Built with ⚡ by Spectrayan


