You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Feedback Loop: Native OpenWebUI RLHF (π/π) with JSON exports to create training datasets
Traceability: n8n journal of read/write operations, Langfuse for tracing prompts/latencies/errors from Quick Wins phase, unlimited retention via periodic summaries and manual purging
2. Architecture & Deployment
2.1 Target Infrastructure
Proxmox: Two distinct VMs (Dev/Prod) for AI stack, extensible resources
Nextcloud: Remains hosted on OpenMediaVault
GPU: NVIDIA Container Toolkit operational; GTX 1660 Ti dedicated to Ollama (prod), RTX 4070 Super mobilized occasionally (manual routing)
Backups: Proxmox Backup Server covers ./AI_Data/openwebui, ./AI_Data/n8n, ./AI_Data/qdrant, ./AI_Data/pgdata
2.2 Services & Interconnections
Orchestration: Stack maintained on Docker Compose (Kubernetes/k3s not retained at this stage)
OpenWebUI: Main UI, native pipelines and RAG enabled (VECTOR_DB=qdrant, RAG_VECTOR_DB=qdrant, CHAT_HISTORY_LIMIT=20). OpenWebUI doesn't write chat history to Qdrant and only uses document collections (docs_public, docs_prive). Long-term conversation memory is managed by n8n in the convos_long collection. Custom Python pipelines for OCR, ingestion, specialized agents
Ollama: Local models (Qwen2.5-7B-Instruct default, Llama-3-3B fallback); simple routing, option to force 7B and external GPU delegation. Tests of quantizations (Q4_K_M, Q8) planned for Qwen2.5-7B with P95/P99 latency measurement via Langfuse; Triton/FasterTransformers not retained at this stage (GPU 1660 Ti)
n8n: Main orchestrator, internal webhooks (/webhook/owui-router), runners enabled, agent state storage in Postgres for long-running workflows
Qdrant: Vector memory (public/safe/private spaces); not exposed, accessible from OpenWebUI and n8n via internal network. Possibility to add Redis or Memcached if embedding latency becomes a friction point
PostgreSQL: n8n database + agent/pipe state persistence, LAN-only
SearxNG: Search engine; exposable via Cloudflare if needed
Prometheus/Grafana: Existing VM, OpenWebUI metrics export (HTTP exporter), n8n and Ollama; alerts routed by n8n to Discord/Telegram
Network: Minimal exposure via pfSense/Cloudflare Zero Trust (if remote access); IP allowlist between VMs for OWUI β n8n β Qdrant/Postgres. Keycloak or Authelia can be introduced later if MFA/SSO need appears, without immediate priority
2.3 Volumes & Environments
Dynamic volumes co-located with docker-compose.yml (./AI_Data/...)
Distinct .env files for Dev/Prod (Ollama URL, secrets, ports, webhook keys) with semiannual rotation
settings.yml managed for SearxNG (limiter disabled locally, OCR and DOI configured)
OpenWebUI Pipeline activation to connect automations to n8n workflows
2.4 Multi-Agents & Pipelines
OpenWebUI handles short ingestion (high top_k) and delegates to n8n for multi-agent workflows
n8n Webhook Trigger + Merge nodes to aggregate responses (PDF, infra, FPV) and return unified result to UI
Agent context storage in Postgres (Database node) for conversation resumption and deferred tasks
2.5 Compliance Controls
OpenWebUI: Native RAG + pipeline enabled, history limited to 20 messages, RBAC enabled to limit pipeline actions
n8n generates monthly report (Markdown/PDF) summarizing new documents, created embeddings and ingestion errors β Nextcloud export
4. AI Intelligence & Workflows
4.1 Memory & RAG
OpenWebUI: Short memory (20 messages) + native RAG (adaptable top_k) with pipeline heuristics; integrated RLHF to refine responses
Qdrant: Long memory, segmentation by project and sensitivity; tags to retrieve conversations and sources. Weaviate multimodal can be evaluated if vision+text need appears
n8n: Triggers long RAG (reduced top_k, tag filters) when pipeline heuristic indicates extended context need
4.2 Multi-Agent Orchestration
n8n pipeline orchestration:
Main agent (router) receives OpenWebUI request via webhook
Specialized agents (PDF/log summary, infra monitoring, FPV tuning, SearxNG search) execute in parallel
Merge node assembles responses, adds used sources, returns to OpenWebUI
Agent state persistence in Postgres for long-running workflows (ex: Freqtrade backtest, heavy log analysis)
Possibility to embed MCP if future need, but pipeline priority for isolated context flexibility
4.3 Feedback & Improvement
OpenWebUI RLHF feeds exportable dataset; n8n records feedback in Qdrant (tag "feedback") without automatic score modification
Automatic test pipeline (n8n) to replay critical prompts and validate agents after update
Sequential Startup: Optimized service startup order with health checks
DEV Mode: Optional complete reset (.env, AI_Data, logs) for development
This technical specification reflects the current architecture, operational priorities, and roadmap of the FlowTech personal AI system. Any major evolution (new services, external exposure, critical automations) must be validated then documented in the infra wiki.