check-events v2 — streaming test results as a ForgeGraph standardShipped

2026-08-24 · ForgeGraph monorepo @ main + bob monorepo @ master · author: Claude for Graham · follows the Bob Cockpit plan & bob-check (shipped 2026-08-24)
1
NDJSON envelope, two producers (Bob worktrees, ForgeGraph CI)
9
runtime adapters (Go, Rust, Py, Ruby, JS, JVM, PHP, .NET, shell)
4
fallback rungs: native JSON → TeamCity → TAP → scrape
CTRF
stored report format on every build row
4
phases; 1–2 shippable independently

Overview

Research (2026-08-24) confirmed no widely-adopted cross-framework protocol exists for streaming test results. The streaming formats are ecosystem-bound (TAP, subunit, go test -json, Bazel BEP); the modern cross-framework standard, CTRF, is batch-only. Every live-test product invents its own NDJSON ingestion — and Bob just did too: bob-check (worktree shim → .bob/check-events.ndjson → runner tail → check WS events → cockpit phase bars) shipped in the cockpit V2 cut.

This plan promotes that proven pipeline into a shared standard — fg-check — that covers Rust, Python, Ruby, Go, JS/TS and more by adapting each framework's native live output, and bakes it into ForgeGraph best practices: conventional target names, sample package.json scripts / justfiles / Makefiles, a forge-ci.toml [check] section, and CI pages that render live per-phase bars and per-test results. Bob's cockpit and ForgeGraph's CI UI become two consumers of one contract.

Goals

Non-goals

Decisions what the exploration settled

QuestionDecisionWhy
Wire formatOwn NDJSON envelope, v2 of bob-check's eventsFile-append + tail already proven in prod; JSON end-to-end (hub WS, event bus, React UIs); binary subunit would need bespoke encoders anyway.
"subunit with TAP fallback"?Rejected as wire format; adopted as designsubunit's timestamps/routing/attachments shape the envelope; TAP demoted to one adapter rung, not the primary.
Test object shapeCTRF's test schema inside our envelopeEnd-of-run fold becomes a reduce into a valid CTRF report; off-the-shelf reporters/tooling; mergeable, comparable, storable.
Fallback laddernative JSON → TeamCity service messages → TAP → regex scrapeTeamCity messages are the de-facto cross-framework streaming stdout format (PHPUnit built-in, pytest/mocha/RSpec libs, every IntelliJ runner); ~50-line parser.
TransportAppend to .fg/check-events.ndjson; consumer tailsIdentical to bob-check today; works for any agent or CI runner; watcher failure degrades to "no events", never a failed run.
ParallelismRequired stream discriminator per eventsubunit's route-code lesson: nextest per-binary, go per-package, sharded vitest interleave; counts stay coherent per-source.
Repo contractConventional targets typecheck · lint · test · e2e · build in npm scripts / justfile / Makefile; no new manifest section — fg-check wraps the existing forge-ci.toml test.command / test.integration / build.commandbob-check already detects package.json scripts; the manifest schema already models unit vs integration — reuse it instead of inventing [check].
ConfigurationsUnit tests, e2e/integration, and each pre-check (lint, typecheck) stream as separate configurationsDistinct phase bars and verdicts per configuration; Bob identifies pre-checks as streaming checks explicitly instead of inferring from output.
CTRF storageFull CTRF report uploaded as a build artifact; ci-report carries an inline summary (counts, duration, top failures)Keeps the report row small and the API payload bounded; full per-test detail fetched on demand by the CI page.
Vitest / JestBuild first-class reporter injection (in-process NDJSON reporters), not TAP/scrapeThese are the fleet's dominant frameworks; exact per-test events there pay for themselves immediately.
JVM adapterDeferred — no JVM apps in the fleet; envelope reserves nothing JVM-specificOpen Test Reporting adapter can be added later without schema changes.

The envelope fg-check events v2

{"v":2,"phase":"test","event":"run_started","framework":"pytest","stream":"api-tests","command":"pytest -q","at":"…"}
{"v":2,"phase":"test","event":"test_finished","stream":"api-tests",
 "test":{"name":"test_streak_gaps","suite":"tests/lib/test_streak.py","status":"failed",
         "duration":812,"message":"expected 3, got 4","trace":"…"},
 "counts":{"passed":41,"failed":1,"total":58},"at":"…"}
{"v":2,"phase":"test","event":"run_finished","stream":"api-tests","status":"failed",
 "counts":{"passed":57,"failed":1,"skipped":0,"total":58},"durationMs":8120,"at":"…"}

Adapter matrix what each runtime streams natively

RuntimeFirst choiceNotesFallback
Gogo test -jsonStable; Package field → stream—
Rustcargo nextest --message-format libtest-json-plus (NEXTEST_EXPERIMENTAL_LIBTEST_JSON=1)Immediate per-test results; nextest subobject disambiguates binaries. Experimental — pin nextest in toolchain evidence.cargo test scrape; cargo build --message-format=json (stable) for build phase
Pythonpytest --report-log=FILE (pytest-dev's reportlog)JSON-lines, flushed per line by design; tail the fileteamcity-messages (unittest); pytest-tap
RubyRSpec formatter: --require .fg/fg_formatter.rb --format FgFormatter --out .fg/rspec-events.ndjsonOne .rb file we ship into the worktree — no gem install; stable formatter APIminitest-reporters; rspec_tap_formatter
JS/TSnode:test --test-reporter=tap; mocha json-stream; vitest/jest reportersNDJSON or TAP, both streamcurrent bob-check scrape (vitest/jest human output)
JVM deferredJUnit Open Test Reporting event stream (file or …open.xml.socket)Not built now (no JVM apps in fleet); adapter slot reserved, no envelope changes needed laterJUnit XML at end (Gradle/Maven)
PHPPHPUnit --teamcity (built-in)Streams service messages—
.NETTeamCity VSTest adapterService messagesTRX at end
Shellbats (native TAP)——
rung 1 native machine stream — exact per-test events, trust fully
rung 2 TeamCity service messages — per-test events, parse ##teamcity[…] lines
rung 3 TAP v12/13/14 — ok/not-ok + YAML diagnostics, no timing
rung 4 regex scrape (today's bob-check) — counts + failure lines only, mark "confidence":"scraped"

Repo conventions best practices + samples

The contract: a repo exposes some subset of typecheck · lint · test · build as runnable targets. fg-check detects them in priority order — package.json scripts → justfile → Makefile → forge-ci.toml [check] explicit override — and picks the adapter from lockfiles/manifests (Cargo.toml, pyproject.toml, Gemfile, go.mod…). Samples ship in docs and in the fg-onboard-repo skill.

package.json

{
  "scripts": {
    "typecheck": "tsc --noEmit",
    "lint": "oxlint .",
    "test": "vitest run",
    "build": "next build"
  }
}

justfile

typecheck:
    cargo check --all-targets

lint:
    cargo clippy -- -D warnings

test:
    cargo nextest run

build:
    cargo build --release

Makefile

.PHONY: typecheck lint test build
typecheck:
	mypy src
lint:
	ruff check .
test:
	pytest -q
build:
	python -m build

forge-ci.toml (existing schema — no new section)

version = 1

[build]
command = "pnpm build"

[test]                    # unit tests → configuration "test"
command = "pnpm vitest run --project unit"

[test.integration]        # e2e → separate configuration "e2e"
command = "pnpm playwright test"
stage   = "beta"

# lint/typecheck configurations come from conventional
# script targets — detected, streamed, never invented
Detection detail
fg-check never invents commands: a missing target is skipped, not failed (bob-check semantics). For just/make, targets are discovered via just --summary / parsing .PHONY + top-level rules; anything unparseable falls back to attempting the conventional names. Where a forge-ci.toml exists, its test.command / test.integration / build.command win over detection — one source of truth, already deployed fleet-wide.

Architecture

fg-check wrapper (per run)

repo conventions

go test -json

nextest libtest-json

pytest --report-log tail

rspec formatter --out tail

mocha json-stream / node:test tap

teamcity msgs

TAP

scrape fallback

tail

WS events

fold at end

ci-report tests field

npm scripts / justfile / Makefile
typecheck · lint · test · build

forge-ci.toml [check] overrides

detect runtime + targets

run command
tee stdout to human log

adapter

normalize

.fg/check-events.ndjson
NDJSON, CTRF-shaped tests

node agent / bob runner

hub ws.forgegraf.com
bob ws-gateway

CI run page · cockpit tiles
live phase bars + counts

CTRF report

builds row
evidence

Diagram source (mermaid)
flowchart LR
  subgraph Repo["repo conventions"]
    C1["npm scripts / justfile / Makefile<br/>typecheck · lint · test · build"]
    C2["forge-ci.toml [check] overrides"]
  end
  subgraph Wrapper["fg-check wrapper (per run)"]
    D[detect runtime + targets] --> R[run command<br/>tee stdout to human log]
    R --> A{adapter}
    A -->|go test -json| N[normalize]
    A -->|nextest libtest-json| N
    A -->|pytest --report-log tail| N
    A -->|rspec formatter --out tail| N
    A -->|mocha json-stream / node:test tap| N
    A -->|teamcity msgs| N
    A -->|TAP| N
    A -->|scrape fallback| N
  end
  N --> E[".fg/check-events.ndjson<br/>NDJSON, CTRF-shaped tests"]
  E -->|tail| AG[node agent / bob runner]
  AG -->|WS events| HUB[(hub ws.forgegraf.com<br/>bob ws-gateway)]
  HUB --> UI["CI run page · cockpit tiles<br/>live phase bars + counts"]
  E -->|fold at end| CT["CTRF report"]
  CT -->|ci-report tests field| DB[(builds row<br/>evidence)]
  C1 --> D
  C2 --> D
  classDef new stroke-dasharray: 5 5,stroke:#e6a23c,color:#e6a23c;
  class D,R,A,N,E,CT new;

Two injection modes cover every adapter: parse-stdout (go, nextest, TAP, TeamCity, mocha json-stream — tee and parse, human log stays pristine for logHead/logTail evidence) and side-channel file/socket (pytest --report-log, RSpec --out, JUnit socket — tail it, exactly like bob-check's events file today).

Mockups animated · what the consumers render as events arrive

Mockups A, C, D, and E are live simulations, not stills: each plays a looped run driven by pure CSS keyframes (postplan executes no scripts), showing exactly what a viewer sees as v2 events stream in — counts ticking, bars creeping, failures popping the moment they happen. Under prefers-reduced-motion every mockup degrades to its final static frame, matching the cockpit's own motion-level philosophy.

A — ForgeGraph CI run page, mid-run

forgegraf.com/ci/4821
CI RUN · FORGEGRAPH/FORGEGRAPH

Build #4821 — fix(agent): stream check events

RUNNINGhetzner-worker · slot 1/2c9e4f21
Status
running
Duration
1m 48s2m 14s2m 39s
Toolchain
mise
Tests
38/5845/5852/1✗/5857/1✗/58 new
Checks new section
typecheck✓ passed12.4s
lint✓ passed8.1s
test38/5845/5852/58 · 1 failed57/58 ✗ failed0m 41s1m 02s1m 08s
e2e · buildqueuedrunning✓ passed—12s24s
▸ api-tests · running verify.test.ts  ·  latest: ✓ webhook signature verify (121ms)▸ api-tests · running streak.test.ts  ·  latest: ✗ computeStreak › gaps — expected 3, got 4▸ test finished · 57/58 · 1 failed  ·  folding CTRF report → e2e/build started
Animated (CSS-only): one 18-second loop plays a run — counts tick 38 → 45 → 52 as counts events arrive, the test bar creeps then flips red on the failure, the live line follows the current file, then the e2e/build configuration takes over. Everything above CHECKS exists today; the strip, Tests card, and live region are new, fed by the hub WS. Reduced-motion viewers see the final state.

B — same page, after the run: per-test results

forgegraf.com/ci/4821
FAILED3m 41stypecheck ✓ · lint ✓ · test 57/58 ✗ · build ✓
Test results · 58 tests · 2 streams new section
TestSuiteStreamTime
FAILcomputeStreak › gapssrc/lib/streak.test.tsapi-tests812ms
AssertionError: expected 3, got 4  ·  streak.test.ts:42  ·  passed locally in agent worktree ⚠ drift
PASScomputeStreak › empty historysrc/lib/streak.test.tsapi-tests3ms
PASSwebhook signature verifysrc/hooks/verify.test.tsweb-tests121ms
SKIPhyperdrive round-tripsrc/db/live.test.tsapi-tests—
▤ CTRF report attached to build evidence⇅ sortable: status · duration · streamfilter: failed ▾
Stored view rendered from the folded CTRF report on the builds row — sortable, filterable, and diffable against the previous run. The drift marker appears when the same test id passed in the agent-worktree run for this changeset (see the cockpit section).

C — Bob cockpit wall tile: today vs v2

bob.blder.bot/cockpit
today · v1
session #1284 · fix-streak-gapsclaude
▸ Running tests…
$ pnpm test
⠸ (spinner — UI knows nothing until scrape)
typecheck ✓lint ✓test 57/58 ✗build —
counts scraped from vitest human output at phase end
with check-events v2
session #1284 · fix-streak-gapsclaude
typecheck ✓ 12slint ✓ 8stest 41/58 ◐test 48/58 ◐test 53/58 ◐build —
✗ computeStreak › gaps · streak.test.ts:42 · expected 3, got 4
▸ running: streak.test.ts · api-testsverify.test.ts · api-testslive.test.ts · api-tests
live per-test events · counts per stream, summed
Same tile, same WS subscription, same check event type — and the contrast is the animation itself: the v1 tile sits frozen until the phase ends, while the v2 tile ticks per-stream counts, cycles the running file, and pops the failure in the moment it happens, pulsing amber while streaming. Wall mode contract: glanceable at 3m. Ops mode expands the phase bar into the full per-test table (same component as mockup B).

D — cockpit PR pipeline strip with local-vs-CI drift

bob.blder.bot/cockpit · ops mode
PR #162 · fix(streak): handle gap days  — session #1284 → forgejo CI run #4821
code ✓ 58/58──▶ CI ✗ 57/58──▶ review──▶ repair──▶ merge──▶ deploy
⚠ DRIFT · computeStreak › gaps — ✓ passed in worktree (mise · node 22.6) → ✗ failed in CI (mise · node 22.4) · same commit c9e4f21 · likely env, not code → suggested action: pin node in mise config
The convergence payoff, animated the way ops mode would show it: the CI stage pulses red while the failure is live, then the drift diagnosis materializes once both runs' evidence is in. Both runs emit the same envelope for the same commit, so the strip names the exact diverging test and juxtaposes each side's toolchain evidence — the flake/env-drift signal neither system can produce today. (Stretch task, phase 4.)

E — what the agent sees in the terminal

agent worktree · $ bob-check
$ bob-check
▶ fg-check typecheck: pnpm run typecheck  · adapter: reporter
✓ typecheck passed in 12.4s
▶ fg-check test: pnpm vitest run  · adapter: vitest-reporter · streams: api-tests, web-tests
…vitest output passes through untouched…
✗ test failed in 68.2s — 57 passed · 1 failed · 1 skipped (58)
  ✗ computeStreak › gaps src/lib/streak.test.ts:42 — expected 3, got 4

fg-check: FAILURES ABOVE  · 60 events → .fg/check-events.ndjson · ctrf: .fg/ctrf-report.json
Plays as a typewriter: lines land in the order the agent would see them, cursor blinking. The human summary contract from bob-check v1 is unchanged — the agent reads the same output it does today, now annotated with which adapter and streams were used. Structured events and the CTRF fold are side effects, never the agent's problem.

Bob in the cockpit first consumer, already live

Bob is not a future consumer — it's the migration source. bob-check v1 shipped with cockpit V2: shim in every agent worktree, worktree-watch.ts tails .bob/check-events.ndjson, the ws-gateway fans check session events, tiles render per-phase bars. fg-check replaces the shim in place; everything downstream keeps working, then gets richer:

Convergence point
A PR flowing through Bob's SDLC gets check events twice: locally in the agent worktree (cockpit "code" stage) and again in ForgeGraph/Forgejo CI (CI run page). Same envelope both times — so the cockpit's pipeline strip can diff "passed locally, failed in CI" per test, which is the flake/environment-drift signal neither system can see today.
2026-08-25 — deployed to the fleet; onboarding became a product. Agent 0.1.56 → 0.1.57 released and converged (7/7 nodes) — every CI worktree carries fg-check on PATH, with the routes-strip guard protecting custom domains. The dogfood repo (check-events-demo) exposed that "zero-config onboarding" required three hand-set pieces of state, each now a real feature: #447 Forgejo webhook provisioning (repo create Step 2.6 + backfill route — new repos pushed into the void), #441 host_node_id adoption from the app's stages, #439 ci_provider auto-enroll when forge-ci.toml exists at head. Also landed: #435 pipeline matrix per-test dots + collapsed PREVIEWS panel (live on forgegraf.com), #443 CTRF artifact fetch (forge ci checks --ctrf + node-disk route), #444 release-agent fixes (HOME under set -u; tag-token fallback with loud degradation — secret since re-minted properly). Bob PR #43 (drift markers) awaits Bob's own review loop. CLI 0.3.5 dispatch chained to #443's merge. The end-to-end demo (push → webhook → adopt → enroll → build → stream → ci checks) is armed as an autonomous chain behind #447.
2026-08-24 — remainder lands via PR #421 (delta on #423). Two sessions built in parallel; #423 landed the core (library 0.1.3, wiring, Checks section), and #421 was reduced to the strict delta: fg-check on every CI worktree PATH (agent/internal/fgcheck embeds the built CLI, .fg/bin prepended via sandbox.DefaultCIPath — command = "fg-check test" needs zero installs), 0.1.4 reporters (per-suite stream so multi-suite counts sum; published), forge ci checks + token route, CTRF artifact beside the build log (Go FoldToCtrf), monorepo CI dogfooding (reporter in apps/web + fold-to-job-summary step), docs/check-events.md + onboarding step, the ChecksSummary story, and Bob's cockpit drift markers (cockpit-v2). Remaining after both merge: agent release (branch release/agent-0.1.56 already cut) + first fleet repo on fg-check test.
2026-08-24 — Phase 3 landed (ForgeGraph) + Phase 4 consumer live (Bob). Agent: manifest/mise runner tails .fg/check-events.ndjson, posts ci-progress every 2 s, ships CIReport.tests. Server: ci-report accepts tests → builds.metadata.tests; ci.progress bus event on /api/fg/events; ci/gate + runs/:id/failures expose tests. Web: CI run page Checks section (per-configuration bars, exact/scraped counts, failing tests) polling /api/ci/:id/checks. Bob: bob-check → fg-check launcher, runner persists end-of-run rollups, cockpit reads FG summaries + bridges ci.* SSE. Remaining: agent release so fleet runners emit; Forgejo-runner PATH prep (fg-check on runner PATH); best-practices docs/skill; "local vs CI" diff marker (stretch).

Phases

Phase 1 — envelope spec + shared wrapper Done

TaskFilesVerificationStatus
Write the v2 envelope spec (schema + JSON Schema for the event union; CTRF test object embedded)new packages/check-events/ in ForgeGraph monorepo (spec + TS types + Go types)Schema validates the example corpus; v1 bob-check lines parse as legacyDone
Extract bob-check into fg-check: zero-dep CLI, detection (package.json → justfile → Makefile → manifest), tee-stdout runner, events file appendpackages/check-events/cli/ (published binary), vendored shim buildSmoke: run in a JS repo, events match v1 behavior + v2 fieldsDone
Rung 2–4 adapters: TeamCity service-message parser, TAP parser, scrape (port existing regexes, tag confidence)same packageFixture corpus per format (recorded real outputs) → golden event streamsDone

Phase 2 — native runtime adapters Done

TaskFilesVerificationStatus
Go: go test -json adapter (Package → stream)packages/check-events/cli/adapters/go.*Golden stream from ForgeGraph agent's own test suiteDone
Rust: nextest libtest-json-plus adapter + env injection; pin nextest version into toolchain evidenceadapters/rustFixture from a sample workspace; parallel binaries → distinct streamsDone
Python: inject --report-log, tail the jsonl; map reportlog $report_type eventsadapters/pythonpytest suite fixture incl. skips/xfailsDone
Ruby: ship fg_formatter.rb (RSpec formatter → NDJSON via --out); minitest reporter fallbackadapters/rubyRSpec fixture; formatter works without Gemfile changesDone
JS: first-class vitest + jest reporters — small in-process NDJSON reporter packages injected via CLI flags (--reporter); mocha json-stream / node:test TAP for the restadapters/js, packages/check-events/reporters/vitest + /jestPer-test golden streams from the monorepo's own suites; scrape rung becomes unreachable for vitest/jest reposDone

Phase 3 — ForgeGraph integration Done

TaskFilesVerificationStatus
Manifest CI runner invokes fg-check (when targets detected) or wraps test.command; tails events into hub session eventsagent/internal/manifestci/runner.go, new tail (mirror bob's worktree-watch drain)Runner test: sentinel repo emits events end-to-endDone
ci-report gains optional tests field: inline summary (per-configuration counts, duration, top failures) on the builds row; full CTRF report uploaded as a build artifact, fetched on demand by the CI pageapps/web/src/app/api/agent/ci-report/route.ts, artifact upload path, schema migrationRoute test; mock-module (not row-insert) per deploy-route mock convention; artifact round-trips the CTRF JSON SchemaDone
CI run page: live phase bars (WS) + per-test results table (stored CTRF) alongside toolchain evidenceapps/web/src/app/ci/[runId]/Page renders live during a real run; failed tests listed with messagesDone
Forgejo-runner jobs: fg-check available on runner PATH via toolchain prep (mise/nix strategies)agent/internal/miseenv/ hookWorkflow job in monorepo emits eventsDone

Phase 4 — Bob migration + best-practices rollout Bob PR open

TaskFilesVerificationStatus
Swap bob-check shim to fg-check build; keep shim name/path; cockpit tiles read v2 test payloads (failing names live, ops-mode per-test expansion)bob: apps/ooda-runner/src/bob-check-cli.mjs → vendored fg-check; components/cockpit/*Cockpit renders richer tiles on a live session; v1 sessions unaffectedIn review
"Local vs CI" diff on the PR pipeline strip (same test failed/passed in both runs)bob services/cockpit/pipeline.tsSynthetic divergence fixture renders the drift markerStretch
Best-practices docs + samples (npm scripts / justfile / Makefile / forge-ci.toml [check]) and fg-onboard-repo skill updatedocs/, skillsOnboarding a fresh Rust + Python repo yields live bars with zero extra configDone

What shipped, and what production taught us 2026-08-26

Every phase above is built and running. The parts a plan cannot predict are below — each one cost a real outage or a real near-miss, so they are recorded as findings rather than as tidy checkmarks.

The documented config was a fork bomb. The plan told repos to write [test] command = "fg-check test" in forge-ci.toml, and separately said manifest commands win over detection. Following both is a loop: fg-check read the manifest, found itself as the test target, and spawned itself without end. It swap-killed a CI node twice — the first time read as a capacity problem, which is why it took two incidents to see.

Fixed in check-events 0.1.7 / agent 0.1.59: a manifest command naming fg-check is a delegation, so the phase resolves through package.json / justfile / Makefile instead; and FG_CHECK_ACTIVE_PHASES makes a nested fg-check refuse a phase its parent already owns. The docs now say this explicitly.

Containment, not just a fix. The sandbox had no resource ceiling, so any runaway build could take its host down. Agent 0.1.60 wraps every CI command in a transient systemd-run scope: MemoryMax at half of RAM, MemorySwapMax=0 so the kernel OOM-kills inside the cgroup instead of the host thrashing into swap-death, and TasksMax=512 so a fork bomb hits a wall. Verified in the runner's journal — Started run-<id>.scope wrapping both the build and the test command — not merely assumed from the version number.

A thrashing node is not a silent node. The self-healing reaper only reset nodes silent past 15 minutes. A fork-bombed box still squeezes out a heartbeat every few minutes, so it never looked dead and had to be reset by hand — twice. The reaper now also judges cadence: beating slower than every 5 minutes opens a node_cadence_collapsed alert, and a streak past 15 minutes makes the node a reset candidate. The firing alert row is itself the "degraded since" timestamp, so this needed no schema change.

Proven end to end, in production

Consumer sweep

1,246 package.json files scanned across the dev tree: the only consumers of @forgegraph/check-events are ForgeGraph itself (via workspace:*, so always the fixed source) and Bob (root bob-check plus ooda-runner, both moved to 0.1.7). Bob's lockfile had pinned 0.1.4 and 0.1.6 exactly — the vulnerable versions — and now carries a single deduplicated 0.1.7. Nothing vulnerable was ever installed on a node, so this was caught before it shipped.

Risks & mitigations

RiskMitigation
nextest libtest-json is experimental and may change shapePin nextest version in toolchain evidence; adapter tolerant of unknown fields; fixture corpus catches drift in CI
Interleaved stdout from parallel runners corrupts rung-2/3 parsingstream discriminator where the format provides one; TAP explicitly documented as unsafe under interleaving → prefer file side-channels
Event volume on huge suites (10k+ tests) floods hub WSWrapper coalesces test_finished into batched count updates above a threshold; full per-test detail only in the folded CTRF report
Scrape rung silently wrong counts"confidence":"scraped" flag; UI renders scraped counts dimmed/approximate
Two checkouts of the contract drift (bob vendored copy vs ForgeGraph package)Single source: fg-check built in ForgeGraph, vendored into bob by version; envelope spec has a JSON Schema both CIs validate fixtures against

Verification

Open questions all resolved 2026-08-24