Skip to content

feat: automatic issue remediation via stack trace analysis (replace agent-directed filing) #31

Description

@markballew

Problem

Today, mcp_remediation_wrapper and install_cli_exception_handler in agent_remediation.py emit markdown that tells the agent to search GitHub issues, thumbs-up duplicates, comment with new info, or open new issues. This has problems:

  1. It's noisy — every error response includes a multi-paragraph remediation block that pollutes tool output
  2. It's unreliable — the agent may or may not follow the instructions, may file low-quality issues, or may get distracted from its primary task
  3. It's not machine-oriented — a block of markdown instructions is the wrong interface for automated issue correlation
  4. It consumes tokens — the remediation text is included in every error response, costing context window on every failure

Proposal

Replace the agent-directed remediation approach with an automatic system that handles issue correlation without involving the agent at all:

1. Stack trace → fingerprint → issue correlation pipeline

When a failure occurs:

  1. log_trace_event() computes an error_fingerprint (already implemented in Design unified logging infrastructure for MCP/CLI requests #17/logging.py)
  2. A background process or post-hoc analysis tool groups traces by fingerprint
  3. Matching GitHub issues are found automatically (fingerprint in issue body or label)
  4. New issues are opened automatically when no match exists, with structured context

2. Remove agent-facing remediation text

  • Stop including "search GitHub issues and file one" instructions in ToolError responses
  • Error responses should contain only: exception type, message, and a fingerprint ID
  • The agent's job is to handle the failure in the primary task, not to do ops triage

3. Implementation options

Approach Pros Cons
Post-hoc log analysis script Simple, runs on collected system logs from #30 Not real-time
GitHub Action on log upload Automated, runs in CI Requires log shipping
In-process async reporter Real-time, immediate issue creation Adds complexity to MCP runtime
Separate MCP tool (report_failure) Agent can optionally invoke it Still agent-directed
Webhook/sidecar Decoupled, listens on syslog Infrastructure overhead

Recommendation: Start with a post-hoc analysis script that reads structured trace logs (from system log via #30), groups by fingerprint, and creates/updates GitHub issues. Graduate to real-time later if needed.

4. Transition plan

  1. feat: route MCP/CLI logs to system log, structured stack traces, and timing telemetry #30 lands: stack traces flow through log_trace_event() to system log with fingerprints
  2. Build the analysis script (this issue)
  3. Strip remediation markdown from mcp_remediation_wrapper / install_cli_exception_handler — they still catch and log, but stop instructing the agent
  4. Downstream MCPs update to new mcp-common version

Scope

  • Design the fingerprint → issue correlation pipeline
  • Build post-hoc analysis tool that reads trace logs and correlates with GitHub issues
  • Strip agent-facing remediation instructions from error wrappers
  • Slim down error responses to: exception type, message, fingerprint
  • Document the new error handling flow
  • Tests for issue correlation logic

Not in scope

  • Real-time issue creation (future enhancement)
  • Integration with external incident management (PagerDuty, etc.)

Depends on

Replaces

  • The current format_agent_exception_remediation() behavior in agent_remediation.py
  • The serverUseInstructions snippet that tells agents to search/file issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions