Skip to content

feat(core): the agent indicator in the diagnostics log (#152) - #153

Open
Maxaubert wants to merge 2 commits into
mainfrom
feat/152-diag-agent
Open

Maxaubert wants to merge 2 commits into
mainfrom
feat/152-diag-agent

Conversation

@Maxaubert

Copy link
Copy Markdown
Owner

Closes #152.

The diagnostics log (#140) now records the agent indicator, so a wrong mark like #151 (restored tabs showed Finished after a restart) can be read from diag.jsonl rather than guessed. Quiet level, always on.

What changed

  • core/renderer/lib/agentDiag.ts (new): the record, pure helpers plus the page's one instance.
    • agent-hook: every OSC 777 prism-agent signal (id, state, kind). A run of the same state on a tab is folded: the first is written at once, the rest counted and written as one line with repeats when the state changes, the tab closes, or after 5 s with no signal from it (t is the last repeat's). A turn of 40 tool calls is two lines.
    • agent-title: the title's meaning (idle / working / starting / question / none), only when it changes, never the text and never each spinner frame.
    • agent-mark: each move of a tab's mark (from, to, held, why, agent, restored ms when within a minute of a restore). why is the rule that fired: hook <state> (, box on screen, , tab in front), spinner after <phase>, idle title after work (an Esc), title <state>, poll found/lost agent, screen read: question box, answer key, box gone, output scored working / output went quiet, stopped working while away, tab looked at, working again, agent gone, closed.
    • agent-restore: a tab restored with an agent resume (id, cwd, resume: true).
  • useAgentIndicator.ts: noteWhy at each decision point and one read-only effect that writes the moves. No rule changed (the gate e2e below all pass unchanged).
  • diag.ts: record(k, fields) for a page's own record kinds; diagIpc.ts: main accepts the four agent-* kinds from the page (and no other agent- kind).
  • termActivity.markResume(id, session, cwd?): logs agent-restore; App passes the folder.
  • tools/diag.mjs: --agent (the timeline: agent lines, tab and shell crumbs, sessions, marks), --tab <id>; --all and --kinds show them; the default problem view shows agent-mark / agent-restore among the context before a problem. --help prints the whole header.
  • docs/diagnostics.md (kinds table, a section on reading the marks), PRIVACY.md (one clause: the indicator's states, never screen, title or conversation text).
  • Versions: core 0.28.1 -> 0.29.0, app 0.35.1 -> 0.35.2.

Prism gets this too

This is a core/ change, so Prism gets the record when it takes this core. Nothing to wire beyond what #140 already needs (startDiag in the page): the hook and title lines and the marks come from useAgentIndicator. Optional: pass the tab's folder as the third argument of markResume so agent-restore carries cwd (without it, cwd is null).

Tests

  • Unit (test-first): agentDiag.test.ts (folding, the quiet sweep, kinds, the meaning-change filter, mark order and held marks, why freshness, restore, close) and a pageLine case for the new kinds.
  • e2e agentDiag (new): a shell stands in for Claude; reads the scenario profile's logs\diag.jsonl and asserts the hook sequence ["working","workingx2","done","question","working","done"], the mark moves with their why (none>working: hook working, working>done: hook done, done>none: tab looked at, none>question: hook question, question>working: hook working, tab in front, working>none: hook done, tab in front), titles ["idle","working","none"], and that a marker string on the screen and in the title never reaches the file.
  • Gates: typecheck, lint (0 errors), npm test 110 files / 1571 tests, e2e agentHooks 37/0, attention 25/0, indicator 12/0, indicatorStyles 74/0, diagLog 26/0, agentDiag 12/0.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FHHaWKR4M5QtW7Wecyuk4t

Maxaubert and others added 2 commits October 10, 2026 05:45
agent-hook (folded runs with a count), agent-title (meaning changes only),
agent-mark (from, to, held and the rule that fired) and agent-restore, at
the quiet level. npm run diag -- --agent shows the timeline. No screen,
title or conversation text is written.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FHHaWKR4M5QtW7Wecyuk4t
… whys (#152)

- agent-hook: a run of one state keeps folding across the model's thinking
  gaps (repeats at most every 30 s, a new line after a minute of silence);
  a 5 s quiet gap made nearly every signal a line of its own.
- A failure kind outside Claude Code's own list is written as `other`
  (it is pty bytes).
- agent-title `none` only once it has held 2 s: Codex's child titles
  between its spinners no longer write a none/working pair each.
- At most 40 hook, title and mark lines per tab per 10 s; the next line
  says `dropped`, and diag.mjs shows it.
- `why` lists only the rules noted in the last second: a busy tab used to
  keep every older rule alive in the list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FHHaWKR4M5QtW7Wecyuk4t
@Maxaubert

Copy link
Copy Markdown
Owner Author

Adversarial review, fixed in 1f685f7 (tests first, each failed on the old code):

  • Folding did not hold for real turns: a 5 s quiet gap ended a run, and the model's thinking between tool calls is usually longer, so nearly every hook was its own line. A run now folds across gaps, writes its count at most every 30 s (and at once on a state change or close), and starts a new line after a minute of silence. Ten minutes of tool calls every 8 s: 150 signals, at most 23 lines.
  • kind was pty bytes (any [a-z0-9_]{1,40} passed): only Claude Code's own StopFailure kinds are written, anything else is other.
  • Codex title flapping: its child processes set titles between its spinners, which wrote a none/working pair each time. none is written only once it has held 2 s.
  • Flood cap: any program can print the OSC or a title; at most 40 hook/title/mark lines per tab per 10 s, the next line carries dropped (diag.mjs shows it).
  • Stale why: each new note refreshed the whole list's time, so a busy tab kept older rules in the reason forever. Only rules noted in the last second are listed now.

Behaviour: no indicator rule changed (diff of useAgentIndicator.ts is notes and one read-only effect). Volume: 8 working Claude tabs are some tens of lines a minute, a few hundred KB an hour of the 10 MB.

Gates: typecheck, lint (0 errors), npm test 1577 pass; e2e one at a time: agentDiag 12, agentHooks 37, attention 25, indicator 12 + indicatorStyles 74, all pass. CI green.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FHHaWKR4M5QtW7Wecyuk4t

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Diagnostics: record the agent indicator (hooks, titles, marks and why, restores)

1 participant