This is close to the work you would do here: a claim-aware editor where a real model reviews the content as you write, annotations that have to stay put through every edit, and a document that more than one person works at once. We have given you a working starter repo to extend, so you are not starting from a blank page.
Do the core well. The collaboration piece is the headline, and we will mostly dig into it together in a follow-up; do not feel you have to finish everything. We would rather see the core done carefully, with tests, than everything done loosely.
Please keep this to about four hours of focused work. The core done well in that time is a complete submission — we are not expecting collaboration or the stretch items to be finished. Where you run out of time, a line in DECISIONS.md about what you would reach for next counts for as much as the code. We mean this: do not spend a weekend on it. (Deploying is still required and takes only a few minutes on top of the build — see Deploy.)
Use it. We do, all day. Claude, Cursor, Copilot, whatever you reach for. We are not checking whether you can hand-write a React component, because a model will give you the happy path in a minute.
What we read is the judgment around it: a document model that keeps annotations correct under real editing, a streaming chat that is genuinely pleasant to use, and code the next person can extend. Put a short note in DECISIONS.md about where you leaned on AI and how you checked it.
An editor for a short one-pager, with an AI reviewer beside it. Three things share one core idea:
- Claims, compliance flags, highlights, and comments are all annotations anchored to a range of text. When the text changes, by your edit, a block move, or someone else's edit, every anchor has to remap so it still points at the right words. This is the heart of the exercise.
- The reviewer is a streaming chat. It runs on a real model, streams its thinking and suggestions, drops accept or reject cards inline, and asks you to settle a claim it cannot decide on its own.
- The document is collaborative. Open it in two browser windows and edits, presence, and annotations should stay in sync.
You will not be writing copy or inventing a product. The one-pager, its three sources, and the initial flags all ship in the seed. The example is a made-up consumer product, the Halo sleep band, backed by three short sources:
- Spec sheet: 30 hours of battery, water-resistant to 50m, tracks heart rate and sleep stages.
- Study summary: in a 200-person internal study, its sleep-stage detection agreed with a clinical sleep lab 87% of the time.
- Compliance note: not a medical device, not meant to diagnose or treat anything.
One flagged claim, "clinically proven sleep tracking," overstates an internal study. Another, "more than a day on a single charge," is supported by the 30-hour spec. A third, "the most accurate sleep tracker you can buy," is a superlative no source settles. Telling those apart is the reviewer's job; keeping the flags on the right words as the draft changes is yours.
Read this part closely, it shapes everything else.
- The document model is plain and testable:
lib/model/has the blocks, the anchors, the operations, and the collaboration transform. No React. The anchoring logic lives inanchors.tsand is the core task. - The reviewer runs on a real model through
app/api/suggestions/[runId]/route.ts. Running it needs an API key, which we provide (see Setup). The chat consumes the stream through oneSuggestionSourceinterface. The live implementation is the default; a deterministicScriptedSuggestionstest double drives the same interface in the tests, so you can build and test the chat with no key. - The reviewer's prompt ships and works, but shaping it is part of the task: tune what it flags, how it phrases a rewrite, and when it should ask you rather than guess. Be ready to talk through it.
- Collaboration runs over a real channel (
lib/collab/) usingBroadcastChannel, so two windows of the same browser sync with no server. The transform that keeps concurrent edits converging is yours to build.
npm install
npm run dev # http://localhost:3000
npm test # the unit suite (no key needed)
npm run test:e2e # the Playwright smokeTo run the reviewer against the real model, put a key in the environment as ANTHROPIC_API_KEY. We provide one for this assignment, so do not spend your own; just do not commit it. Everything except a live review works without it.
content-studio/
app/
page.tsx the three-pane editor
api/suggestions/[runId]/ the reviewer SSE route (real model)
components/
editor/ the editor, blocks, selection toolbar
review/ the streaming review chat
ui/ a small set of primitives
lib/
model/ document, anchors, ops, transform (start here)
suggestions/ the SuggestionSource: live + scripted test double
collab/ the cross-window channel and presence
sse-client.ts a working SSE client
store/document-store.ts Zustand state for the document
fixtures/ the seeded doc, sources, and the test-double stream
tests/ some passing, some skipped that point at the work
Start with lib/model/anchors.ts and tests/unit/.
1. Anchoring. Make annotations survive editing. remapAnchors in lib/model/anchors.ts handles only the easy case today. Make it correct across a delete, an insert inside an annotation's span, the boundary cases, and a block move, so flags, highlights, and comments stay on their text. Make the skipped tests in tests/unit/anchors.test.ts pass.
2. Editing. The block body is editable: typing flows through editBlockBody into insert/delete ops, and a flag should stay on its words as you type around it. The baseline handles plain typing and keeps the caret in place, but leaves the hard cases (IME/composition, paste, restoring a full selection, the caret when a remote edit lands under it, and re-rendering every segment per keystroke on a long doc) for you. Wire highlight and comment so a text selection becomes a correctly anchored annotation (the selection mapping in the editor is a first pass and is wrong once a block holds an annotation). Make a section reorder keep its annotations.
3. The review chat. Build the chat against the streaming reviewer: render tokens as they arrive with a thinking indicator, autoscroll that yields when the reader scrolls up, an optimistic send with a Stop button that aborts the stream, honest loading and error states, accept and reject cards that apply to the document, and the human-in-the-loop question. Shape the reviewer prompt while you are here.
4. Collaboration. Open the document in two windows and make edits, annotations, presence, and cursors stay in sync and converge. lib/model/transform.ts is a no-op today, so concurrent edits diverge. Build the transform so applying a local and a remote op in either order reaches the same document, and so a remote edit never corrupts a local annotation. Showing two windows converging is a strong signal; you do not have to finish it before we talk.
Undo and redo that restore the selection and annotations; performance on a long document; an accessibility pass (the selection toolbar and reorder are pointer-only today); overlapping annotations.
Put it somewhere we can click through. It is a Next.js app, so Vercel is the easy path, though use whatever you prefer. The live reviewer needs an API key set in your host's environment; without one the editor and the rest of the app still demo the work, so deploy it either way and tell us which. Send the live URL.
- The code. The previously-skipped tests should pass, and add your own for what you build, especially the anchoring and the chat.
- A short DECISIONS.md: the two or three calls you thought hardest about and the trade-offs, where you used AI and how you checked it, what you would reach for in production for collaboration and why, and any risk you would flag before this saw real traffic.
- Hand it over the way you would to a teammate: a branch, readable commits, and whatever we need to run it.
- A live URL for the deployed app (Vercel or equivalent).
We will not be tallying features, and we will go through your submission with you. In rough order of what counts:
- Anchoring. Do annotations stay correct under editing, reorder, and (the headline) a remote edit? This is where strong separates from solid.
- The streaming chat. Does it stream well, handle the question, and recover from a dropped or aborted stream? Is it actually pleasant to use?
- The live run. Does the real reviewer work, can you defend the prompt, and can you show or reason about two windows converging?
- Craft and taste. Readable diffs, tests that pin the hard cases, accessibility, and a DECISIONS.md that shows your thinking. Where the design is underspecified — the reviewer's question card, the streaming feel, an empty or error state — we look at the call you made and why. Polish one interaction end-to-end rather than many loosely; we are not tallying features.
Where something is underspecified, make a call, write down why, and keep going. We read that as a good sign.