Revert "perf(world-vercel): avoid repeated frame buffer copies" (#3586) - #3650
Revert "perf(world-vercel): avoid repeated frame buffer copies" (#3586)#3650pranaygp wants to merge 2 commits into
Conversation
This reverts commit 52526e1.
🦋 Changeset detectedLatest commit: 5e7a90a The changes in this PR will be included in the next version bump. This PR includes changesets to release 17 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results✅ All tests passed 🛠 Infra Events (absorbed by the harness)Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.
E2E Test SummarySummary
Details by Category✅ ▲ Vercel Production
✅ 💻 Local Development
✅ 📦 Local Production
✅ 🐘 Local Postgres
✅ 🪟 Windows
✅ 🌐 Cross-language Conformance
✅ vercel-multi-region
|
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 240925ms → this run 225149ms (Δ -15776ms, -7%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): 📜 Previous results (1)91578ddWed, 19 Aug 2026 01:05:22 GMT · run logs
Streams
ℹ️ Metric definitions & methodologyStreams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
There was a problem hiding this comment.
Pull request overview
This PR reverts the recent @workflow/world-vercel frame-decoding optimization that deferred concatenation across awaits, restoring copy-on-receipt behavior to prevent corruption when upstream chunk sources reuse backing buffers (which was causing fleet-wide Vercel-world e2e/WS lane hangs).
Changes:
- Revert
decodeFrames’srefillimplementation to immediately copy each received chunk into an owned buffer. - Add a new changeset describing the revert and remove the prior performance-oriented changeset.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
| packages/world-vercel/src/frames.ts | Reverts deferred chunk-view accumulation; returns to copy-on-receipt buffering to avoid payload corruption under buffer reuse. |
| .changeset/revert-fast-frame-refill.md | Adds a patch changeset documenting the revert and its motivation. |
| .changeset/fast-frame-refill.md | Removes the prior changeset that described the reverted optimization. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| '@workflow/world-vercel': patch | ||
| --- | ||
|
|
||
| Revert the deferred frame-buffer concatenation in `decodeFrames`: stashing yielded chunk views across `await`s let the stream reuse their backing buffers before the copy, corrupting encrypted frame payloads (vercel-world e2e and WS transport lanes failed fleet-wide). Chunks are copied on receipt again. |
| if (!value || value.byteLength === 0) continue; | ||
| const next = new Uint8Array(buffer.byteLength + value.byteLength); | ||
| next.set(buffer, 0); | ||
| next.set(value, buffer.byteLength); | ||
| buffer = next; |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 41 total
Full trace: |
E2E Vercel Prod Tests (nextjs-turbopack - quickjs) has outgrown the 30 minute job budget: it was killed mid-final-test-group with every test passing on three consecutive attempts (and was the sole non-green job on unrelated runs earlier this week). GitHub reports the timeout as conclusion "cancelled", which reads like infra noise and repeatedly failed the E2E Required Check aggregate for otherwise-green runs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Green across the board with the 40m lane budget — E2E Required Check passed (the three earlier failures were all the turbopack-quickjs lane hitting the old 30m ceiling mid-final-group). Ready to land; main's vercel lanes stay red until this merges. 🤖 Generated with Claude Code |
Summary & Motivation
Every
mainrun since #3586 merged (22:27Z tonight) fails fleet-wide on the vercel-world lanes: 17+E2E Vercel Prod Testsjobs per run, and all fourE2E Vercel WS Transport Testjobs hitting the 30-minute job ceiling with hung runs. The last green vercel-lane run (PR #3333's768401c9, 22:16Z) was based one commit before #3586; every run containing it is red. Same signature on PRs: test failures dumping raw encrypted frame payloads (Error: {"0":101,"1":110,"2":99,"3":114,...}— bytes spellingencr…), e.g. this nextjs run on #3333's branch.Root cause
The new
refilldefers copying: it stashes yielded chunk views inpartsand keepsawait chunks.next()-ing until enough bytes arrive before concatenating. If the stream reuses a previously-yielded chunk's backing buffer while producing the next chunk (pooled/BYOB-style readers do), the stashed views are overwritten before the copy — corrupting frame payloads. Corrupted encrypted frames then surface as theencr…blob errors, and the framed replay/WS readers wedge (the 30m WS timeouts). The old code copied each chunk on receipt, which is what this revert restores. (The rewrite also dropped the!valueguard on yielded chunks.)A safe re-land needs the deferred-concat to copy each chunk before the next
await(which is most of the copies back) or prove the upstream iterator never reuses buffers.Validation
@workflow/world-vercelunit suite: 528/528. The e2e verdict is this PR's own vercel lanes going green again.cc @NathanColosimo
🤖 Generated with Claude Code