-
Notifications
You must be signed in to change notification settings - Fork 0
feat(kernel): lens follow-ups — vocabulary owner, render bound, ordering (#197) #376
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -637,29 +637,17 @@ fn worker_reported_failure_without_detail_still_records_a_verification() { | |
| /// place fails here instead of silently showing a reader two names for one | ||
| /// completion. | ||
| /// | ||
| /// The list below is itself hand-maintained: `reason_label`'s wildcard-free | ||
| /// match makes a NEW variant a compile error there, but a new variant simply | ||
| /// missing from this array is not caught by anything. Add variants in both | ||
| /// places. (An iterable-enum derive would remove the second list; that is a | ||
| /// dependency decision, not one to smuggle into a diagnostic fix.) | ||
| /// #197 moved the drift test onto `CompletionReason` itself, beside the label. | ||
| /// This wrapper stays only to catch a rename of `journal_label()` at a | ||
| /// familiar call site — the real invariant lives in `entry.rs`'s | ||
| /// `completion_reason_tests::every_journal_label_matches_serialized`. | ||
| #[test] | ||
| fn every_reason_label_matches_its_serialized_form() { | ||
| for reason in [ | ||
| CompletionReason::Success, | ||
| CompletionReason::VerificationFailed, | ||
| CompletionReason::RetriesExhausted, | ||
| CompletionReason::LeaseExpired, | ||
| CompletionReason::Crashed, | ||
| CompletionReason::Timeout, | ||
| CompletionReason::WorkerError, | ||
| CompletionReason::BudgetExceeded, | ||
| CompletionReason::Canceled, | ||
| ] { | ||
| for reason in CompletionReason::ALL { | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. P3: This test duplicates the owner-side drift test in Prompt for AI agents |
||
| let serialized = serde_json::to_value(reason).unwrap(); | ||
| assert_eq!( | ||
| serialized.as_str().expect("a string spelling"), | ||
| super::reason_label(&reason), | ||
| "journal label drifted from the serialized form for {reason:?}" | ||
| reason.journal_label(), | ||
| ); | ||
| } | ||
| } | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -77,10 +77,19 @@ impl Engine<WallClock> { | |
| // output the worker sent with its failing completion — and that is | ||
| // exactly what gets nulled. Capture it here, bounded, or the run records | ||
| // that the step failed and discards every trace of why. | ||
| let mut failure_detail = failure_reason | ||
| .is_some() | ||
| .then(|| worker_failure_detail(&completion.output)) | ||
| .flatten(); | ||
| // | ||
| // ORDERING NOTE (#197 item 4): this capture MUST run before any | ||
| // `reject()` call below. If it did not, `reject`'s "keep both accounts" | ||
| // path would see a leftover `failure_detail` from a completion whose | ||
| // ORIGINAL `completion_reason` was `Success` and format | ||
| // "rejected: …; worker reported: …" for a success. The explicit | ||
| // `match` (rather than the tighter `.is_some().then().flatten()`) | ||
| // makes the "only on non-success" precondition self-evident so a | ||
| // future edit that moves this line downward reads as suspicious. | ||
| let mut failure_detail = match failure_reason { | ||
| Some(_) => worker_failure_detail(&completion.output), | ||
| None => None, | ||
| }; | ||
| let mut rejected_completion = false; | ||
| let mut reject = |error: anyhow::Error| { | ||
| rejected_completion = true; | ||
|
|
@@ -378,20 +387,90 @@ fn next_stream_offset(journal: &SqliteJournal, stream: &str) -> Result<u64> { | |
| Ok(next) | ||
| } | ||
|
|
||
| /// The worker's own account of a failure, bounded so a large or hostile output | ||
| /// cannot bloat the journal. `None` when the worker sent nothing useful, which | ||
| /// keeps the caller's fallback ("reported X without detail") honest rather than | ||
| /// recording an empty string as though it were a diagnostic. | ||
| /// The worker's own account of a failure, bounded — both in the journal AND | ||
| /// during render — so a large or hostile output cannot bloat the process's | ||
| /// heap OR the run's journal. `None` when the worker sent nothing useful, | ||
| /// which keeps the caller's fallback ("reported X without detail") honest | ||
| /// rather than recording an empty string as though it were a diagnostic. | ||
| /// | ||
| /// #197 (item 3) previously said "bounded" when only the OUTPUT was bounded; | ||
| /// the render inside this function was unbounded, so a multi-megabyte JSON | ||
| /// object was fully materialized into memory before anything was measured or | ||
| /// truncated. `write_bounded_json` now stops emitting once | ||
| /// `MAX_RENDER_BYTES` has been produced, and the truncation suffix is added | ||
| /// after the cheap byte cap, not after a full render. | ||
| fn worker_failure_detail(output: &Value) -> Option<String> { | ||
| // Chars, not bytes: the cut below is by char index. The suffix reports the | ||
| // remainder in bytes, which is why both units appear in one function. | ||
| const MAX_CHARS: usize = 2000; | ||
| // UTF-8 upper-bounds a char at 4 bytes, plus slack for the "… (N bytes | ||
| // truncated)" suffix computation. This ceiling ONLY caps the render; the | ||
| // final cut still happens on char boundaries below. | ||
| const MAX_RENDER_BYTES: usize = MAX_CHARS * 4 + 256; | ||
| if output.is_null() { | ||
| return None; | ||
| } | ||
| let mut render_truncated = false; | ||
| let rendered = match output { | ||
| Value::String(text) => text.clone(), | ||
| other => other.to_string(), | ||
| Value::String(text) => { | ||
| if text.len() > MAX_RENDER_BYTES { | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. P2: When a failure string has more than 8,256 leading whitespace bytes, this prefix cap discards the useful diagnostic and returns Prompt for AI agents |
||
| render_truncated = true; | ||
| // Char-boundary-safe slice for the pre-cap head. | ||
| let cut = text | ||
| .char_indices() | ||
| .take_while(|(byte_index, _)| *byte_index <= MAX_RENDER_BYTES) | ||
| .last() | ||
| .map(|(byte_index, ch)| byte_index + ch.len_utf8()) | ||
| .unwrap_or(0); | ||
| text[..cut].to_owned() | ||
| } else { | ||
| text.clone() | ||
| } | ||
| } | ||
| other => { | ||
| // A hand-rolled `Write` that stops once its budget is exhausted, | ||
| // so `to_writer` never allocates a full render of a hostile | ||
| // object before we get a chance to cut it. | ||
| struct BoundedWriter { | ||
| buf: Vec<u8>, | ||
| budget: usize, | ||
| truncated: bool, | ||
| } | ||
| impl std::io::Write for BoundedWriter { | ||
| fn write(&mut self, chunk: &[u8]) -> std::io::Result<usize> { | ||
| if self.budget == 0 { | ||
| self.truncated = true; | ||
| return Ok(chunk.len()); | ||
| } | ||
| let take = chunk.len().min(self.budget); | ||
| self.buf.extend_from_slice(&chunk[..take]); | ||
| self.budget -= take; | ||
| if take < chunk.len() { | ||
| self.truncated = true; | ||
| } | ||
| // Report full consumption so serde does not spin. | ||
| Ok(chunk.len()) | ||
| } | ||
| fn flush(&mut self) -> std::io::Result<()> { | ||
| Ok(()) | ||
| } | ||
| } | ||
| let mut writer = BoundedWriter { | ||
| buf: Vec::with_capacity(MAX_RENDER_BYTES.min(4096)), | ||
| budget: MAX_RENDER_BYTES, | ||
| truncated: false, | ||
| }; | ||
| let _ = serde_json::to_writer(&mut writer, other); | ||
| render_truncated = writer.truncated; | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. P2: When the render cap and character cap both apply, this flag is ignored and the detail reports only the remainder of the capped prefix. Mark the detail as render-bounded, or report the byte count as a lower bound, whenever Prompt for AI agents |
||
| // Repair a mid-multi-byte cut so `String::from_utf8` never fails. | ||
| while !writer.buf.is_empty() && std::str::from_utf8(&writer.buf).is_err() { | ||
| writer.buf.pop(); | ||
| } | ||
| match String::from_utf8(writer.buf) { | ||
| Ok(text) => text, | ||
| Err(_) => String::new(), | ||
| } | ||
| } | ||
| }; | ||
| let trimmed = rendered.trim(); | ||
| if trimmed.is_empty() { | ||
|
|
@@ -400,7 +479,13 @@ fn worker_failure_detail(output: &Value) -> Option<String> { | |
| // Truncate on a char boundary; `output` is arbitrary worker-supplied data | ||
| // and slicing it by byte index would panic on multi-byte input. | ||
| Some(match trimmed.char_indices().nth(MAX_CHARS) { | ||
| None => trimmed.to_owned(), | ||
| None => { | ||
| if render_truncated { | ||
| format!("{trimmed}… (render bounded)") | ||
| } else { | ||
| trimmed.to_owned() | ||
| } | ||
| } | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Truncation count ignores render capLow Severity The Additional Locations (1)Reviewed by Cursor Bugbot for commit c32351c. Configure here. |
||
| Some((cut, _)) => format!( | ||
| "{}… ({} bytes truncated)", | ||
| &trimmed[..cut], | ||
|
|
@@ -455,16 +540,21 @@ mod worker_failure_detail_tests { | |
| /// trusts it. | ||
| #[test] | ||
| fn truncation_does_not_split_a_multi_byte_char() { | ||
| // 3000 three-byte chars = 9000 bytes. | ||
| let output = json!("€".repeat(3000)); | ||
| // 2500 three-byte chars = 7500 bytes: above the char cap (2000) so | ||
| // the char cut happens, but below MAX_RENDER_BYTES (~8256) so the | ||
| // render cap does NOT preempt it. The char-boundary invariant is | ||
| // what this test exists to pin, so keep the input in the char-cap | ||
| // regime and let `render_is_bounded_before_allocation_for_hostile_*` | ||
| // cover the render cap. | ||
| let output = json!("€".repeat(2500)); | ||
| let detail = worker_failure_detail(&output).expect("detail for a long output"); | ||
| assert!( | ||
| detail.contains('…'), | ||
| "expected a truncation marker, got {detail:?}" | ||
| ); | ||
| // Cut at 2000 CHARS = 6000 bytes, so 3000 bytes remain. | ||
| // Cut at 2000 CHARS = 6000 bytes; input is 7500 bytes, so 1500 bytes remain. | ||
| assert!( | ||
| detail.contains("3000 bytes truncated"), | ||
| detail.contains("1500 bytes truncated"), | ||
| "expected the byte remainder, got {detail:?}" | ||
| ); | ||
| assert_eq!(detail.chars().take_while(|c| *c == '€').count(), 2000); | ||
|
|
@@ -475,4 +565,50 @@ mod worker_failure_detail_tests { | |
| let exact = "a".repeat(2000); | ||
| assert_eq!(worker_failure_detail(&json!(exact.clone())), Some(exact)); | ||
| } | ||
|
|
||
| /// #197 item 3: the render itself is bounded. Before the fix, a | ||
| /// pathological JSON object would allocate its full serialization into | ||
| /// memory before anything measured or cut it, which was the exact | ||
| /// "bloat the process" case the docstring claimed to prevent. Give the | ||
| /// worker a JSON object whose full render would be ~500 KB and confirm | ||
| /// (a) we still return a bounded detail, and (b) we mark it as bounded. | ||
| #[test] | ||
| fn render_is_bounded_before_allocation_for_hostile_json() { | ||
| // 50_000-element array of small integers → ~500 KB serialized. | ||
| let big: Vec<Value> = (0..50_000_i64).map(|n| json!(n)).collect(); | ||
| let detail = worker_failure_detail(&Value::Array(big)) | ||
| .expect("detail for a large output"); | ||
| // Truncation was applied AND signaled — the caller can tell the | ||
| // difference between a short detail that fit and a bounded render. | ||
| assert!( | ||
| detail.contains('…'), | ||
| "expected a truncation marker, got {} bytes", | ||
| detail.len() | ||
| ); | ||
| // Render was capped: `MAX_RENDER_BYTES = MAX_CHARS * 4 + 256 = 8256`; | ||
| // trimmed detail should stay near that ceiling with slack for the | ||
| // suffix ("… (N bytes truncated)"). Pin at a generous ~9 KB so | ||
| // future MAX_CHARS bumps don't need to touch this test. | ||
| assert!( | ||
| detail.len() < 9_000, | ||
| "detail bloated past render bound: {} bytes", | ||
| detail.len() | ||
| ); | ||
| } | ||
|
|
||
| /// A big STRING output also gets its render bounded — the previous cheap | ||
| /// `text.clone()` allocated the full string before the char-boundary cut. | ||
| #[test] | ||
| fn render_is_bounded_before_allocation_for_hostile_string() { | ||
| // 500 KB of ASCII. | ||
| let big: String = "x".repeat(500_000); | ||
| let detail = worker_failure_detail(&json!(big)) | ||
| .expect("detail for a large string"); | ||
| assert!(detail.contains('…'), "expected truncation marker"); | ||
| assert!( | ||
| detail.len() < 9_000, | ||
| "detail bloated past render bound: {} bytes", | ||
| detail.len() | ||
| ); | ||
| } | ||
| } | ||


There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
P2: The test's name and doc claim it proves
ALLenumerates every variant, but it only checks that the listed labels are unique. Both this test andevery_journal_label_matches_serializediterate overCompletionReason::ALL, so if a new variant is added to the enum andjournal_labelbut not toALL, neither test observes it. The claim '(b) a test failure if you skip ALL' is therefore false — the missing variant is silently undetected and can drift the journal label out of sync with the serialized form. Either add a real coverage guard or correct the misleading name/doc.Prompt for AI agents