Skip to content

feat(gateway): explain compute driver discovery decisions #3211

Description

@elezar

User Story

As an OpenShell operator, I want the gateway to explain how it selected or rejected local compute drivers, so that I can diagnose startup failures without reproducing them under a debugger.

Problem Statement

Gateway auto-detection probes local driver candidates, such as Docker Unix sockets, but silently treats probe failures as unavailable. When a candidate socket exists but Snap confinement denies access, the gateway eventually reports only that no suitable driver was found and systemd may restart it. Operators cannot tell which candidates were examined, whether a driver was selected, or whether a candidate failed because of a missing socket, an access denial, a timeout, or an unexpected API response.

Impact / Why This Matters

Package and confinement failures are difficult to distinguish from an absent runtime. The current workaround is to manually inspect socket paths, Snap interface connections, and journal output, then infer the cause. That is insufficient for CI and user installations because the decisive probe result is not recorded.

Proposed Design

When gateway driver auto-detection runs, expose diagnostics at debug level that identify each candidate driver and socket probe outcome without logging credentials or other sensitive data. Startup errors should provide a concise summary of the attempted drivers and why no usable driver was selected. A successful selection should state the selected driver and endpoint category in diagnostics.

Acceptance Criteria

  • With debug logging enabled, each local compute-driver candidate records whether it was selected, skipped, or rejected.
  • Docker socket diagnostics distinguish a missing/non-socket path, connection denial, timeout, and an incompatible or unsuccessful API response.
  • Normal startup errors summarize the failed discovery decision without exposing secrets or request contents.
  • A successful auto-detection records the selected driver and does not expose sensitive endpoint data.
  • Coverage verifies representative successful, missing, and permission-denied probe outcomes.

Alternatives Considered

  • Relying only on systemd or Snap logs: these do not report the gateway discovery decision or every candidate it attempted.
  • Treating socket existence as availability: this would hide confinement failures and select unusable drivers.
  • Requiring an explicitly configured driver: useful as a workaround, but does not make default installations diagnosable.

Agent Investigation

Docker discovery currently probes each candidate with a Unix-socket HTTP ping and returns only a boolean result. Failed metadata, connect, write, read, timeout, and response-validation paths are intentionally collapsed into unavailable. This behavior is shared by local API socket discovery and leaves no per-candidate diagnostic trail.

Related: #2869 (comment)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions