Skip to content

Group rankings by model line, fetch source URLs, show output cost - #3

Open
jammyfu wants to merge 5 commits into
mainfrom
cursor/family-grouped-rankings-0e4b
Open

jammyfu wants to merge 5 commits into
mainfrom
cursor/family-grouped-rankings-0e4b

Conversation

@jammyfu

@jammyfu jammyfu commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

Summary

Follow-up polish on the merged AA ranking demo (PR #2), including later same-PR asks.

  • Ranking UI is grouped by mainstream lines/families instead of one flat list: OpenAI (GPT), Anthropic (Claude), Google (Gemini), SpaceXAI (Grok), DeepSeek, Alibaba (Qwen), Z AI (GLM), Kimi, MiniMax, plus speed specialists / other.
  • Each snapshot row has a family field. Global AA rank numbers and median TPS stay. Doubao and ERNIE are omitted because the 2026-09-13 AA public table has no median Tokens/s for those rows (no invented numbers).
  • Switching A or B to a model in a different line resets that pane’s stream. Same-line SKU / TPS tweaks keep existing text.
  • Source text is editable on every breakpoint. Paste a public http(s) URL and Fetch readable title + body into the source box (Vite /api/fetch-page + client fallbacks). Fetch updates input only.
  • Models with an AA-published output $/1M show that price, a $/s · $/min · $/h burn rate, and cumulative spend while streaming. Missing prices render as —. Formula: spend ≈ tokens × output $/token; rate ≈ TPS × output $/token.
  • One primary Play action. zh/en labels, 100dvh layout, no paid AA API.

How grouping + clear work

  • src/data/modelFamilies.ts orders household Western + CN lines, then specialists/other.
  • Desktop: sticky family headers. Mobile: family pills filter the compact chip row.
  • isSameFamily(prev, next) decides whether streamA.reset() / streamB.reset() runs.

How source / URL fetch works

  • Controls always expose a source <details>.
  • /api/fetch-page fetches the URL, strips HTML, rejects localhost/private addresses.

How cost works

  • Field: outputPricePerMillionUsd (USD per 1M output tokens), set only when an AA model page quotes it (Celeris, Mercury, DeepSeek V4.1 Flash, Claude Opus 5, Grok 4.6, Qwen3.8 Max, GLM-5.3 max, MiniMax-M3).
  • Rate = active TPS × $/token; spend = tokens emitted × $/token.
  • Stats show $/s · $/min · $/h plus running spend; compare shows A/B per-second rates. Footer has the formula.

Verification

  • python3 tools/verify.py passed.
  • Browser-checked family grouping, same-line keep, cross-line clear (incl. compare B), URL fetch of example.com, ~390px pills + Play, and cost ticking on Celeris-1.
Open in Web Open in Cursor 

cursoragent and others added 2 commits September 13, 2026 10:29
Present AA median TPS rows under household lines (OpenAI, Anthropic,
Google, DeepSeek, specialists/other). Switching A or B to a different
line resets that pane; same-line SKU changes keep streamed text.

Co-authored-by: PaintingCoder <jammyfu@users.noreply.github.com>
Restore the source editor on every breakpoint and add a secondary
Fetch action that loads a public page into the source box via a
same-origin Vite proxy, with client-side fallbacks. Fetching updates
input only; output still clears only when the model line changes.

Co-authored-by: PaintingCoder <jammyfu@users.noreply.github.com>
@cursor cursor Bot changed the title Group rankings by model line and clear output on family change Group rankings by model line, clear on family change, fetch source URL Sep 13, 2026
cursoragent and others added 2 commits September 13, 2026 11:03
Stops the token loop from gluing the last word of a page onto the
next repeat. Notes browser verification of grouping and line-change clear.

Co-authored-by: PaintingCoder <jammyfu@users.noreply.github.com>
Group Qwen, GLM, Kimi, MiniMax, and Grok from the AA public table
(Doubao/ERNIE omitted — no published median TPS). Show AA output
$/1M, burn rate, and cumulative spend when a list price exists.

Co-authored-by: PaintingCoder <jammyfu@users.noreply.github.com>
@cursor cursor Bot changed the title Group rankings by model line, clear on family change, fetch source URL Group rankings by model line, fetch source URLs, show output cost Sep 13, 2026
@jammyfu
jammyfu marked this pull request as ready for review September 13, 2026 11:13
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-13T11:17:36.009246Z c50d402 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Single-model stats now read $/s · $/min · $/h from active TPS × AA
output list price. Compare mode still shows A/B per-second rates.

Co-authored-by: PaintingCoder <jammyfu@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c50d402bb0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vite.config.ts
Comment on lines +35 to +36
const upstream = await fetch(target.href, {
redirect: 'follow',

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Revalidate redirect targets before following them

When the dev or preview server is reachable by an untrusted user, redirect: 'follow' lets an initially public URL redirect to a loopback, metadata, or other private address without passing the new destination through parsePublicHttpUrl. Because the endpoint returns the fetched body, this creates an SSRF path; follow redirects manually and validate every destination before requesting it.

Useful? React with 👍 / 👎.

Comment thread src/lib/sourceUrl.ts
Comment on lines +28 to +31
throw new SourceFetchError('blocked');
}
if (/^\d+\.\d+\.\d+\.\d+$/.test(host) && isPrivateIPv4(host)) {
throw new SourceFetchError('blocked');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reject bracketed private IPv6 destinations

When fetching a literal IPv6 URL such as http://[::1]:8080/, URL.hostname retains the brackets, so the value does not match the ::1 entry in BLOCKED_HOSTS, and the IPv4-only check below also ignores it. The proxy therefore accepts loopback and other private IPv6 destinations and can expose internal HTTP services; normalize IPv6 literals and reject loopback, link-local, and private IPv6 ranges.

Useful? React with 👍 / 👎.

Comment thread src/App.tsx
Comment on lines +139 to +140
const spentA = cumulativeCostUsd(streamA.tokensCount, selected?.outputPricePerMillionUsd);
const spentB = cumulativeCostUsd(streamB.tokensCount, compare?.outputPricePerMillionUsd);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Accumulate cost across same-family price changes

When tokens have already streamed and the user switches to another model in the same family, the intentional no-reset behavior preserves tokensCount, but this calculation reprices every previous token using the newly selected model. For example, switching from Celeris-1 to Mercury 2 retroactively changes the displayed total, while switching to an unpriced specialist erases it entirely; track spend incrementally per price segment or reset cost whenever the price changes.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants