Skip to content

Add experimental Jev features: natural-language workflow search and event history review - #3948

Draft
rossnelson wants to merge 4 commits into
mainfrom
nl-search-jev
Draft

rossnelson wants to merge 4 commits into
mainfrom
nl-search-jev

Conversation

@rossnelson

@rossnelson rossnelson commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Two experimental features, both off by default, that use the TypeSafe System One model "Jev". Jev returns typed judgments (choice / yes-no / score) with probabilities. It never generates text. Code finds the candidates, Jev selects, and code builds the result.

nl-search-jev.mp4

1. Natural-language workflow search. An input above the workflow filter bar turns a sentence such as "failed order workflows from yesterday" into the existing filter pills. The frontend builds the visibility query with the existing pipeline, so the model cannot invent an attribute or a value. The pills are the verification step: the user sees what was understood and can edit it.

2. "Review with Jev" on the event history and the timeline. A button scores each row for importance to the workflow outcome. Runs of routine rows collapse into one toggle row. Each run has its own toggle, "Show all events" opens every run, and "Clear review" removes the review. One review is shared by the history and timeline views.

A short screen recording follows in a comment.

How it works

  • POST /api/v1/nl-search and POST /api/v1/history-review in ui-server. The TypeSafe client and the two translators are plain Go packages with no Echo or config dependency.
  • Both endpoints share one gate, in this order: 404 when off → rate limit → auth header → body limit → validation → DescribeNamespace with the caller's credentials → the paid call. A caller that the Temporal frontend rejects causes zero TypeSafe calls (covered by tests with a bufconn fake frontend).
  • Rate limit: a deployment bucket, a per-IP bucket, and a per-caller bucket, with an overflow bucket so token rotation cannot grow the map or evict callers. On by default even when the config has no rateLimit block; rateLimit.disabled: true is the explicit opt-out.
  • History review cost control: rows with the same signature share one question (300 identical workflow-task rows cost one question), at most 200 distinct signatures and 4 model requests per call.
  • Fail visible: a row with no usable answer always stays visible. A failed batch returns 502 with no partial scores.

Privacy

  • Nothing is sent before the user acts (submit / click).
  • The history review request holds identifiers only: event type, category, classification, attempt, and names such as activity type, signal name, timer id. No payloads, failure messages, stack traces, user metadata, memo or search-attribute values. A unit test builds a history that contains all of those (plain and base64) and asserts that none reaches the request body, and that the body has a closed key set. markerName is the one name a workflow author can set freely; this is noted in the code.
  • The API key stays in ui-server config. Settings expose two booleans only.
  • Operators should know that search text and workflow type names go to a third party when they enable the features.

Configuration

typesafe:            # one key per ui-server deployment
  apiKey: ""
  model: jev-latest
  baseUrl: https://api.typesafe.ai
  timeout: 5s
  rateLimit: { disabled: false, requestsPerMinute: 30, burst: 10, deploymentRequestsPerMinute: 300 }
nlSearch:
  enabled: false
historyReview:
  enabled: false

Local: TEMPORAL_TYPESAFE_API_KEY=... pnpm dev:ui-server turns both features on in development.yaml. Docker: TEMPORAL_TYPESAFE_*, TEMPORAL_NL_SEARCH_ENABLED, TEMPORAL_HISTORY_REVIEW_ENABLED.

Testing

  • Go: gofmt, go vet, go test -count=1 -race ./... (14 packages pass). httptest fake for TypeSafe, bufconn fake for the Temporal frontend, no network.
  • Frontend: pnpm check (0 errors), pnpm lint, pnpm test (270 files, 3487 tests pass), pnpm vite build. New render tests for the search input and the hidden-run row (pnpm check does not detect a missing .svelte import; a render test does).
  • Playwright integration: history-review.spec.ts (history + timeline, endpoint mocked).
  • Manual, against a local dev server with a real key: search and history review both return 200 in roughly 0.2–0.4 s.

Also in this PR

  • Fix: the config refresh goroutine kept looping after Close() (break inside select).
  • Retry-After is added to the CORS exposed headers.

Open items before this leaves draft

  • The default collapse threshold (0.35) hides completed activities (score ≈ 0.33) as well as bookkeeping rows. 0.20 would hide bookkeeping only. It is a frontend constant; no new model call is needed to change it.
  • GROUP BY WorkflowType is not available on standard visibility, so the known-workflow-types hint is empty there; the search then selects a type only from the words in the text.
  • The config loader uses html/template, so a key that contains + & < > ' " must be set as a quoted string in a config file. Startup rejects a key that was HTML-escaped.
  • Docs for the new config block and environment variables.
  • Whether the translator packages should later move to the visibility service so the CLI and SDKs get the same search.

Add two experimental, off-by-default endpoints that use the TypeSafe
System One model "Jev". Jev returns typed judgments with probabilities
and never generates text: code finds the candidates, Jev selects, and
code builds the result.

- POST /api/v1/nl-search turns a sentence into typed workflow filters.
  The frontend builds the visibility query with the existing pipeline,
  so the model cannot invent an attribute or a value.
- POST /api/v1/history-review scores event history rows for importance
  to the workflow outcome. Rows with the same signature share one
  question, failures and pending rows are pinned, and a row without an
  answer always stays visible.

Both endpoints share one gate: 404 when off, rate limit, auth header,
body limit, validation, a DescribeNamespace call with the caller's
credentials, and only then the paid call. The rate limiter has a
deployment bucket, a per-IP bucket and a per-caller bucket, with an
overflow bucket so key rotation cannot grow the map or evict callers.

Config lives in a shared `typesafe` block (one key per deployment) with
separate `nlSearch.enabled` and `historyReview.enabled` gates. The key
never reaches the browser; settings expose two booleans only.

Also fix the config refresh goroutine, which kept looping after Close.
Natural-language search: an experimental input above the workflow
filter bar sends a sentence to ui-server and turns the typed filters
into the existing filter pills, so the pills, the query, the counts and
saved views all keep working. Low confidence still applies the filters
and asks the user to check the pills.

History review: a "Review with Jev" button on the event history and the
timeline collapses runs of routine rows into one toggle row. Each run
has its own toggle, "Show all events" opens every run, and "Clear
review" removes the review. One review is shared by both views. Nothing
is sent before the click, and the request holds identifiers only: no
payloads, failure messages, user metadata, memo or search attribute
values. A test builds a history with all of those and asserts that none
reaches the request body.

Both features show only when ui-server reports them as enabled.
@vercel

vercel Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
holocene Ready Ready Preview Sep 26, 2026 1:41pm UTC

Request Review

The translator now sends Jev every search attribute with the comparisons
its type allows, and asks which attribute and comparison each value in
the search uses. Each answer is recorded as a trace step with its score,
threshold, probabilities, and outcome, and the response returns the trace.

On the Workflows page, Why? opens a panel beside the table that shows
each stage of the trace. A carousel keeps the last searches; selecting
one applies its query again. The search text is kept in the nlSearch URL
parameter so a reload does not clear it.
Starts a fixed set of workflows with CustomerTier and Attempts search
attributes, so each demo search has one right answer. record.ts records
the searches, the Why? panel, and a history review with Playwright.

This branch was successfully deployed

1 active deployment
Preview — d872ca1b Deployed Sep 26, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant