Skip to content
This repository was archived by the owner on Sep 8, 2026. It is now read-only.
This repository was archived by the owner on Sep 8, 2026. It is now read-only.

feat: BYO inference via TokenPony (TPX) alongside Vercel AI Gateway #326

Description

@btipling

Summary

Mirror the sandbox backend split (BYO daemon vs Vercel Sandbox) for inference: operators and tenants should be able to run Invincible against a TokenPony / TPX OpenAI-compatible provider or keep today’s Vercel AI Gateway path — without hard-coding a single vendor into the product architecture.

Spec (source of truth for the wire): TPX — Token Pony Express
Reference provider origin (example): https://api.tokenpony.dev

This is a product issue, not an implementation plan.

Why

Today Gap
Chat/agent inference is server-side via Vercel AI Gateway (AI_GATEWAY_API_KEY) Forks / BYO operators who do not want (or cannot use) Vercel Gateway credits still need a first-class path
Tenancy BYOK still routes through Gateway (providerOptions.gateway.byok + paid team credits) Provider billing can be the tenant’s, but routing fabric stays Vercel-centric
Sandbox already has per-row / env backend (byo | vercel) Inference has no analogous backend seam

Reusable-product north star (AGENTS.md / docs/bring-your-own.md): connect your Vercel (or host) project and choose your inference fabric — same spirit as choosing your sandbox.

Analogy (locked product shape)

Concern Option A (origin default) Option B (BYO)
Sandbox Vercel Sandbox SDK BYO HTTP daemon (SANDBOX_URL + token, /v1/*)
Inference (this issue) Vercel AI Gateway TokenPony TPX (OAuth-metered, OpenAI-compatible chat)

Same rules as sandbox BYO:

  • Secrets server-only (never client / Wasm / NEXT_PUBLIC_* / git)
  • Fail closed when misconfigured
  • Living docs timeless (no phase theater)
  • Feature divide unchanged: Wasm harness is product UI; inference stays on Next host APIs (/api/chat, /api/agent)

TokenPony / TPX (what we’re targeting)

From https://tokenpony.dev/spec/ (TPX v0.3 summary for implementers later):

Area Behavior
Discovery GET /.well-known/oauth-protected-resource, GET /.well-known/oauth-authorization-server
Auth OAuth 2.0 + PKCE (S256); DPoP recommended; short-lived tokens; refresh rotation; audience-restricted grants
Chat POST /v1/chat/completions — OpenAI-compatible, SSE stream + non-stream
Models GET /models (ids + OpenRouter-style pricing strings; "0" = free)
Credits GET /credits — grant spend (total_purchased / total_used USD)
Errors e.g. 401 invalid token; 402 budget_exhausted / balance_exhausted (OpenAI-shaped error body)
Trust model No embedded provider API keys in the app; user/operator names issuer; budget is a hard USD damage cap

Important distinction from current tenant BYOK:
Today’s BYOK stores provider API keys and still sends them through Vercel AI Gateway. TPX is a different inference backend: OAuth grant to a TPX-compatible issuer, then OpenAI-compatible calls with DPoP/bearer — not “paste OpenAI key into Gateway.”

Goals

  1. Config seam for inference backend: at least vercel_gateway | tokenpony (names bikesheddable), analogous to sandbox backend.
  2. Host/operator path can run chat + agent using TokenPony when configured (tools still work when sandbox is configured).
  3. Per-tenant path is designed so a tenant (or host) can prefer TPX without forcing every deploy to keep paying Gateway BYOK credit tax — exact tenancy UX can be phased, but the seam must not paint us into Gateway-only forever.
  4. Model catalog for harness protocol v3 picker comes from the active backend (GET /api/models stays session-honest).
  5. Streaming agent/chat paths keep working (SSE to host; thinking/tools unchanged at product level).
  6. Errors map cleanly (401/402 budget vs Gateway 401/402 credits) — never leak tokens; never look like “sandbox not configured” 503 chat fallback.
  7. Docs: docs/bring-your-own.md, AGENTS.md where-to-change, .env.example / operator checklist — BYO inference next to BYO sandbox.

Non-goals (v1)

  • Replacing Vercel hosting (app still deploys on Vercel or equivalent Node host)
  • Putting TPX tokens or DPoP keys in Wasm / browser
  • Supporting every OpenAI-compatible random base URL without a clear trust profile (prefer TPX-shaped issuers; generic OpenAI base URL may be a later stretch)
  • Dropping Gateway support (both backends remain valid)
  • Changing sandbox BYO design
  • Dual DOM chat or moving inference into the canvas

Constraints

Do Do not
Inference only on server (/api/chat, /api/agent) NEXT_PUBLIC_ Gateway or TPX secrets
Reuse OpenAI-compatible tool/streaming patterns where the AI SDK allows a custom base URL / fetch Hard-code only api.tokenpony.dev as the sole legal issuer if the spec is multi-provider
Treat OAuth/DPoP token storage like other server secrets (env and/or DEK ciphertext under tenancy) Log access tokens, refresh tokens, or DPoP proofs
Fail closed when backend is tokenpony but auth/budget missing Silent fallback to host Gateway spend
Cloud-native ops if Production needs migrate/seed for new tables Laptop-only cutover as the primary path

Open design questions (for plan later)

  • Env / operator shape: e.g. INFERENCE_BACKEND=tokenpony + issuer base URL + how operators obtain/store grant tokens (device flow vs long-lived operator grant).
  • Per-tenant: per-tenant issuer + grant vs host-level TPX only for dogfood.
  • AI SDK integration: custom OpenAI-compatible provider vs Gateway-only providerOptions — must support tools + streaming used by runAgent / runAgentStream.
  • Model ids: TPX GET /models vs today’s grant catalog tables — one catalog abstraction.
  • 402 budget_exhausted UX in harness (EMBER error, no chat fallback).
  • Parity matrix: chat JSON, agent SSE, reasoning/thinking models, tool calls, abort/Stop.

Acceptance (issue-level)

  • Documented inference backend seam (Gateway | TokenPony/TPX), parallel to sandbox backend story.
  • With TokenPony configured and Gateway path unused, /api/chat and /api/agent complete a real turn (including tool loop when sandbox available).
  • Model list for the harness reflects the active backend; unauthenticated clients never see secrets.
  • Misconfig / expired grant / budget exhausted → clear 4xx (not 503 sandbox-not-configured).
  • No TPX or Gateway secrets in client bundle, Wasm, or git.
  • Living docs + .env.example updated for BYO inference operators.
  • Gateway-only deploys keep working unchanged.

References

Note

Prefer a follow-up create-plan issue (phases: wire seam → OAuth/token storage → chat/agent provider → catalog → tenancy → docs/ops) before large implementation. This ticket captures product intent only.

Review note

Tenancy is now always-on / multi-tenant only (the tenancy-off code path was removed and merged). References to a "tenancy-off" origin/env shape in this ticket were the only stale mentions — cleaned up to host/operator and per-tenant framing. No functional change to this feature's intent.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions