You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Sep 8, 2026. It is now read-only.
Mirror the sandbox backend split (BYO daemon vs Vercel Sandbox) for inference: operators and tenants should be able to run Invincible against a TokenPony / TPX OpenAI-compatible provider or keep today’s Vercel AI Gateway path — without hard-coding a single vendor into the product architecture.
Spec (source of truth for the wire): TPX — Token Pony Express
Reference provider origin (example): https://api.tokenpony.dev
This is a product issue, not an implementation plan.
Why
Today
Gap
Chat/agent inference is server-side via Vercel AI Gateway (AI_GATEWAY_API_KEY)
Forks / BYO operators who do not want (or cannot use) Vercel Gateway credits still need a first-class path
Tenancy BYOK still routes through Gateway (providerOptions.gateway.byok + paid team credits)
Provider billing can be the tenant’s, but routing fabric stays Vercel-centric
Sandbox already has per-row / env backend (byo | vercel)
Inference has no analogous backend seam
Reusable-product north star (AGENTS.md / docs/bring-your-own.md): connect your Vercel (or host) project and choose your inference fabric — same spirit as choosing your sandbox.
POST /v1/chat/completions — OpenAI-compatible, SSE stream + non-stream
Models
GET /models (ids + OpenRouter-style pricing strings; "0" = free)
Credits
GET /credits — grant spend (total_purchased / total_used USD)
Errors
e.g. 401 invalid token; 402budget_exhausted / balance_exhausted (OpenAI-shaped error body)
Trust model
No embedded provider API keys in the app; user/operator names issuer; budget is a hard USD damage cap
Important distinction from current tenant BYOK:
Today’s BYOK stores provider API keys and still sends them through Vercel AI Gateway. TPX is a different inference backend: OAuth grant to a TPX-compatible issuer, then OpenAI-compatible calls with DPoP/bearer — not “paste OpenAI key into Gateway.”
Goals
Config seam for inference backend: at least vercel_gateway | tokenpony (names bikesheddable), analogous to sandbox backend.
Host/operator path can run chat + agent using TokenPony when configured (tools still work when sandbox is configured).
Per-tenant path is designed so a tenant (or host) can prefer TPX without forcing every deploy to keep paying Gateway BYOK credit tax — exact tenancy UX can be phased, but the seam must not paint us into Gateway-only forever.
Model catalog for harness protocol v3 picker comes from the active backend (GET /api/models stays session-honest).
Streaming agent/chat paths keep working (SSE to host; thinking/tools unchanged at product level).
Errors map cleanly (401/402 budget vs Gateway 401/402 credits) — never leak tokens; never look like “sandbox not configured” 503 chat fallback.
Docs: docs/bring-your-own.md, AGENTS.md where-to-change, .env.example / operator checklist — BYO inference next to BYO sandbox.
Non-goals (v1)
Replacing Vercel hosting (app still deploys on Vercel or equivalent Node host)
Putting TPX tokens or DPoP keys in Wasm / browser
Supporting every OpenAI-compatible random base URL without a clear trust profile (prefer TPX-shaped issuers; generic OpenAI base URL may be a later stretch)
Dropping Gateway support (both backends remain valid)
Changing sandbox BYO design
Dual DOM chat or moving inference into the canvas
Constraints
Do
Do not
Inference only on server (/api/chat, /api/agent)
NEXT_PUBLIC_ Gateway or TPX secrets
Reuse OpenAI-compatible tool/streaming patterns where the AI SDK allows a custom base URL / fetch
Hard-code only api.tokenpony.dev as the sole legal issuer if the spec is multi-provider
Treat OAuth/DPoP token storage like other server secrets (env and/or DEK ciphertext under tenancy)
Log access tokens, refresh tokens, or DPoP proofs
Fail closed when backend is tokenpony but auth/budget missing
Silent fallback to host Gateway spend
Cloud-native ops if Production needs migrate/seed for new tables
Laptop-only cutover as the primary path
Open design questions (for plan later)
Env / operator shape: e.g. INFERENCE_BACKEND=tokenpony + issuer base URL + how operators obtain/store grant tokens (device flow vs long-lived operator grant).
Per-tenant: per-tenant issuer + grant vs host-level TPX only for dogfood.
AI SDK integration: custom OpenAI-compatible provider vs Gateway-only providerOptions — must support tools + streaming used by runAgent / runAgentStream.
Model ids: TPX GET /models vs today’s grant catalog tables — one catalog abstraction.
402 budget_exhausted UX in harness (EMBER error, no chat fallback).
Prefer a follow-up create-plan issue (phases: wire seam → OAuth/token storage → chat/agent provider → catalog → tenancy → docs/ops) before large implementation. This ticket captures product intent only.
Review note
Tenancy is now always-on / multi-tenant only (the tenancy-off code path was removed and merged). References to a "tenancy-off" origin/env shape in this ticket were the only stale mentions — cleaned up to host/operator and per-tenant framing. No functional change to this feature's intent.
Summary
Mirror the sandbox backend split (BYO daemon vs Vercel Sandbox) for inference: operators and tenants should be able to run Invincible against a TokenPony / TPX OpenAI-compatible provider or keep today’s Vercel AI Gateway path — without hard-coding a single vendor into the product architecture.
Spec (source of truth for the wire): TPX — Token Pony Express
Reference provider origin (example):
https://api.tokenpony.devThis is a product issue, not an implementation plan.
Why
AI_GATEWAY_API_KEY)providerOptions.gateway.byok+ paid team credits)byo|vercel)Reusable-product north star (AGENTS.md / docs/bring-your-own.md): connect your Vercel (or host) project and choose your inference fabric — same spirit as choosing your sandbox.
Analogy (locked product shape)
SANDBOX_URL+ token,/v1/*)Same rules as sandbox BYO:
NEXT_PUBLIC_*/ git)/api/chat,/api/agent)TokenPony / TPX (what we’re targeting)
From https://tokenpony.dev/spec/ (TPX v0.3 summary for implementers later):
GET /.well-known/oauth-protected-resource,GET /.well-known/oauth-authorization-serverPOST /v1/chat/completions— OpenAI-compatible, SSE stream + non-streamGET /models(ids + OpenRouter-style pricing strings;"0"= free)GET /credits— grant spend (total_purchased/total_usedUSD)budget_exhausted/balance_exhausted(OpenAI-shaped error body)Important distinction from current tenant BYOK:
Today’s BYOK stores provider API keys and still sends them through Vercel AI Gateway. TPX is a different inference backend: OAuth grant to a TPX-compatible issuer, then OpenAI-compatible calls with DPoP/bearer — not “paste OpenAI key into Gateway.”
Goals
vercel_gateway|tokenpony(names bikesheddable), analogous to sandboxbackend.GET /api/modelsstays session-honest).docs/bring-your-own.md,AGENTS.mdwhere-to-change,.env.example/ operator checklist — BYO inference next to BYO sandbox.Non-goals (v1)
Constraints
/api/chat,/api/agent)NEXT_PUBLIC_Gateway or TPX secretsapi.tokenpony.devas the sole legal issuer if the spec is multi-providertokenponybut auth/budget missingOpen design questions (for plan later)
INFERENCE_BACKEND=tokenpony+ issuer base URL + how operators obtain/store grant tokens (device flow vs long-lived operator grant).providerOptions— must support tools + streaming used byrunAgent/runAgentStream.GET /modelsvs today’s grant catalog tables — one catalog abstraction.Acceptance (issue-level)
/api/chatand/api/agentcomplete a real turn (including tool loop when sandbox available)..env.exampleupdated for BYO inference operators.References
/v1/chat/completions,/models,/credits)AI_GATEWAY_API_KEYlib/chatServer.ts,lib/agent/runAgent.ts,app/api/chat,app/api/agent,lib/tenancy/resolveInference*,lib/gateway/*,app/api/modelsNote
Prefer a follow-up create-plan issue (phases: wire seam → OAuth/token storage → chat/agent provider → catalog → tenancy → docs/ops) before large implementation. This ticket captures product intent only.
Review note
Tenancy is now always-on / multi-tenant only (the tenancy-off code path was removed and merged). References to a "tenancy-off" origin/env shape in this ticket were the only stale mentions — cleaned up to host/operator and per-tenant framing. No functional change to this feature's intent.