Skip to content

feat(uptimekuma): add Uptime Kuma monitor provider - #678

Open
davidgrldo wants to merge 3 commits into
stakater:masterfrom
davidgrldo:enh/627-uptime-kuma-support
Open

davidgrldo wants to merge 3 commits into
stakater:masterfrom
davidgrldo:enh/627-uptime-kuma-support

Conversation

@davidgrldo

Copy link
Copy Markdown
Contributor

Fixes #627

What

Adds Uptime Kuma (2.x) as a monitor provider, following the Adding support for a new Monitor conventions in CONTRIBUTING.md. Uptime Kuma has no REST API for monitor management — its officially sanctioned management API is socket.io, so this uses github.com/breml/go-uptime-kuma-client (MIT, actively maintained, tested against Kuma 2.5.x, also used by the Terraform providers).

Changes

  • New provider pkg/monitors/uptimekuma: UpTimeKumaMonitorService implementing the MonitorService interface. One socket.io client held for the service lifetime, lazily (re)connected after failures under a mutex. GetByName matches on the full monitor list (Kuma has no lookup-by-name); Update fetches the current monitor and overrides controller-managed fields, matching Kuma's editMonitor full-object semantics.
  • CRD: new uptimeKumaConfig (interval min 20, monitorType http|keyword, keywordValue, keywordExists, notifications). Generated deepcopy + CRD manifests updated in all three copies.
  • Wiring: monitor-proxy.go (OfType/ExtractConfig cases, TypeUptimeKuma) and controller dispatch.
  • Config: reuses the existing provider apiURL/username/password fields — no new global config.
  • Tests: offline unit tests for config defaults, http/keyword mapping, keyword-invert mapping, notification-ID parsing, and Equal comparisons (325 lines, stdlib testing).
  • Docs: new docs/uptimekuma-configuration.md (config example, CRD example, Kuma 2.x-only note, auth notes: API keys are not accepted for monitor management, 2FA must be off for the service account, recommend pointing at the internal HTTP service), README provider list, CONTRIBUTING test-config sample.

Verification

  • go build ./... ✅
  • go vet ./pkg/monitors/uptimekuma/... ./api/... ./pkg/monitors/... ✅
  • go test ./... — all 15 packages ✅ (offline)
  • gofmt clean ✅

Notes for reviewers

  • go.mod: the new dependency requires go >= 1.25.2, so the go directive was bumped 1.24.0 → 1.25.2 (CI workflows use go-version-file: go.mod, so they follow automatically). The rest of the go.mod churn is MVS-forced by the new module's requirements — verified minimal by re-resolving from a clean state (identical result).
  • Uptime Kuma 1.23.x: not supported by the Go client (login flow differs); 2.0 has been stable since Oct 2025. Documented in the docs page.
  • Live testing against a real Kuma instance is recommended during review (unit tests are offline; live integration tests follow the existing self-skip pattern).

Add support for Uptime Kuma (2.x) as an uptime monitor provider, managed
through its socket.io API via github.com/breml/go-uptime-kuma-client:

- New UpTimeKumaMonitorService implementing the MonitorService interface
  with a single lazily (re)connected client, guarded by a mutex
- New uptimeKumaConfig in the EndpointMonitor CRD: interval (min 20),
  monitorType (http|keyword), keywordValue, keywordExists, notifications
- Auth via username/password reusing the existing provider config fields
  (Uptime Kuma API keys are not accepted for monitor management)
- Update fetches the full monitor object and overrides controller-managed
  fields, matching Uptime Kuma's editMonitor semantics
- Docs: docs/uptimekuma-configuration.md, README provider list, sample
  test config in CONTRIBUTING

Dependency note: github.com/breml/go-uptime-kuma-client requires go >= 1.25.2
(go directive bumped; CI reads the Go version from go.mod). The remaining
go.mod churn is MVS-forced by the new module's requirements, verified minimal.

Fixes stakater#627

Signed-off-by: davidgrldo <davidgrldo123@gmail.com>

@SyedaFatimaKazmi SyedaFatimaKazmi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @davidgrldo , your contribution is much appreciated! The provider is well structured.

I reviewed the code and tested the PR in a kind cluster against Uptime Kuma 2.5.5.

Summary

  • With the code as is, the controller cannot create any monitor. Three things block it: the connect context is cancelled right after login, RetryInterval is 0 (Kuma rejects it), and the Docker image does not build because the Dockerfile still uses Go 1.24. CI lint also fails for the same Go 1.25 reason.
  • After two small local fixes (background context and a retry interval), the rest works well: create, update, delete, switching between http and keyword, notifications, and reconnect after a Kuma restart. With 15 monitors and 5 concurrent reconciles there were no duplicates or errors.
  • A few behaviours need a look: the keywordExists mapping is reversed compared to the CRD text and the UptimeRobot provider, redirects show as DOWN (max redirects 0), there is no request timeout, and repeated reconnects can trigger Kuma's login rate limit.

I left inline comments on each point and suggested fixes.

Happy to re-test once the changes are in. Thanks again!

Comment thread go.mod
module github.com/stakater/IngressMonitorController/v2

go 1.24.0
go 1.25.2

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We now need Go 1.25.2, but the image in Dockerfile #L2 is still golang:1.24 - making the build stop at go mod download with go.mod requires go >= 1.25.2. Please change it to golang:1.25. The CI "Build image" step will hit this too once lint passes

similarly in the pull_request.yaml workflow
L#37 lint fails here because golangci-lint v1.64 is built with Go 1.24. It cannot lint a Go 1.25.2 module. Please bump the version to a release built with Go 1.25 (v2.x).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3832d8b.

  • Dockerfile: golang:1.24 -> golang:1.25.
  • .github/workflows/pull_request.yml and push.yml: golangci-lint-action@v6 -> @v8 with version: v2.14.0. v1.64 refuses to run at all on this module, which I could reproduce with the released binary:
    Error: can't load config: the Go language version (go1.24) used to build golangci-lint is lower than the targeted Go version (1.25.2)
    
  • Makefile: GOLANGCI_LINT_VERSION -> v2.14.0 and the install path -> github.com/golangci/golangci-lint/v2/cmd/golangci-lint so make verify works too.

Because v2 enables checks that v1 did not (stylecheck + quickfix are merged into staticcheck), the v2 run reported 12 pre-existing issues in other providers. I added a small .golangci.yml that keeps the v1 defaults (exclusions.presets is the old exclude-use-default, and the staticcheck.checks list reproduces v1's default set), then fixed the 6 remaining real findings (ST1005/ST1017) so lint is at 0 issues. Locally verified with golangci-lint v2.14.0 and the CI Lint step passes.

ctx, cancel := context.WithTimeout(context.Background(), connectTimeout)
defer cancel()

client, err := kuma.New(ctx, service.url, service.username, service.password,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This context is cancelled when ensureClient returns (defer cancel()). The client keeps that context for the whole socket connection. So the connection closes right after login. Every call then fails with context canceled. I could not create one monitor with this code. Please pass context.Background() here. WithConnectTimeout already limits the connect time. The TestUpTimeKumaMonitorLifecycle live test fails on this too.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, this was the showstopper. Fixed in 3832d8b — ensureClient now passes context.Background() to kuma.New. WithConnectTimeout(connectTimeout) already bounds the connect attempt, so nothing is lost.

Name: m.Name,
Interval: int64(normalized.Interval),
NotificationIDs: parseNotificationIDs(normalized.Notifications),
IsActive: true,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RetryInterval is not set, so it is 0. Kuma 2.5.5 rejects it: Retry interval cannot be less than 1 seconds. No monitor can be created. Please set it, for example to the same value as Interval. That is the Kuma UI default.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3832d8b. RetryInterval is now set to the interval on create (buildKumaMonitor) and on update, using the Kuma UI default.

details := monitor.HTTPDetails{
URL: m.URL,
Method: "GET",
AcceptedStatusCodes: []string{"200-299"},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MaxRedirects and Timeout are not set, so both are 0. With 0 redirects, any URL that redirects (for example http to https) is DOWN with status code 302. With timeout 0, Kuma waits interval × 800 seconds (about 13 hours at interval 60). A site that hangs never goes DOWN. Please use the Kuma UI defaults: MaxRedirects: 10 and Timeout: interval × 0.8.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3832d8b, on create and on update, with the Kuma UI defaults: MaxRedirects: 10 and Timeout: interval * 0.8 (requestTimeoutFor). Covered by tests for 20/60/120/300 second intervals.

// keywordInvert maps KeywordExists to Uptime Kuma's invertKeyword flag:
// "no" means alert when the keyword does NOT exist
func keywordInvert(keywordExists string) bool {
return strings.EqualFold(keywordExists, "no")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This mapping is reversed. In Kuma, invertKeyword=false means DOWN when the keyword is missing. So keywordExists: "yes" alerts when the keyword is missing. The CRD says yes = alert if the value exists. The UptimeRobot provider does that. Test: keywordExists: "no" went DOWN with keyword is present. Please flip this function and invertToKeywordExists, or change the CRD text.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I checked this against the Kuma 2.5.5 source, and I believe the mapping is the other way round. In server/model/monitor.js the check is:

let keywordFound = data.includes(this.keyword);
if (keywordFound === !this.isInvertKeyword()) {
    bean.status = UP;
}

So invertKeyword=false -> UP when the keyword is found, and invertKeyword=true -> UP only when it is absent. That is what this provider does today (keywordExists: "no" -> invertKeyword=true), it is the same as UptimeRobot's keyword_type 1 (contains) and 2 (does not contain), and it is what the docs example relies on: keywordValue: "404" with keywordExists: "no" stays UP while the page is healthy and alerts as soon as a 404 page shows up.

Your own test result confirms it: with keywordExists: "no" the monitor went DOWN while the keyword was present, which is the documented behaviour rather than a bug.

The confusing part is the CRD wording "Alert if value exist (yes) or doesn't exist (no)", read literally it implies the opposite. I took your second option and changed the text instead, in the CRD description and in the docs:

  • yes (default) — up while the keyword is in the response, down when it disappears
  • no — up only while the keyword is absent, so it alerts as soon as the keyword shows up

If you did mean the literal reading of the old text (i.e. yes = alert when present), say so and I will flip keywordInvert/invertToKeywordExists instead, but that would make the provider inconsistent with UptimeRobot and invert the example above.

}

// invalidateClient drops the connection so the next call reconnects
func (service *UpTimeKumaMonitorService) invalidateClient() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every failed call drops the client. The next reconcile logs in again. Kuma allows only 20 logins per minute for all clients together. With a wrong password, the controller used all 20 in a few seconds. Kuma then blocked my own UI login too. Please add a backoff before a reconnect, or do not reconnect on auth errors

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3832d8b. ensureClient now tracks a reconnect deadline: a dropped connection is re-established after 5s, then 10s, 20s, ... capped at 5m, and ensureClient returns an error instead of logging in again while that backoff is active. An authentication failure (the client reports Incorrect username or password / authIncorrectCreds) backs off for 5 minutes and logs which account/URL to check, so a wrong password can no longer burn the 20 logins per minute.


if normalized.MonitorType == "keyword" {
if len(normalized.KeywordValue) == 0 {
log.Error(nil, "Monitor is of type Keyword but the `keyword-value` is missing")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This logs an error but still creates a keyword monitor with an empty keyword. Please return here instead.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3832d8b. buildKumaMonitor now returns an error and Add bails out instead of creating a keyword monitor with an empty keyword; Update checks the same before it sends anything. The conversion failures of the current monitor in Update now return as well instead of continuing with a zero-value monitor.

Comment thread docs/uptimekuma-configuration.md Outdated
interval: 120
monitorType: keyword
keywordExists: no
keywordValue: 404

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The API server rejects this example. YAML reads no as a boolean and 404 as a number, but both fields are strings. Please quote them: keywordExists: "no" and keywordValue: "404"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3832d8b, both values are quoted now (keywordExists: "no", keywordValue: "404"). I also added a note about why they have to be quoted, since YAML reads no as a boolean and 404 as a number.

return nil, nil
}

func (service *UpTimeKumaMonitorService) Add(m models.Monitor) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With a notification ID that does not exist, Kuma returns a foreign-key error but still creates the monitor without notifications. The controller then retries the update on every reconcile and fails each time. A check against GetNotifications, or a clearer error, would help.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3832d8b. Before creating or updating, the configured IDs are checked against client.GetNotifications() for the account the controller logs in with, and an unknown ID aborts the operation with an error naming the missing IDs and pointing at Settings > Notifications. That stops the create-without-notifications plus retry-every-reconcile loop.

davidgrldo and others added 2 commits October 1, 2026 15:45
Blockers:

- pass context.Background() to kuma.New: the client keeps that context for
  the whole socket connection, so the deferred cancel closed it right after
  login and every call afterwards failed with "context canceled"
- set retryInterval (Kuma rejects a monitor without one), maxredirects (10)
  and timeout (interval * 0.8) to the Kuma UI defaults, on create and on
  update; with maxredirects 0 every redirect was reported DOWN and with
  timeout 0 a hanging endpoint never went down
- Dockerfile: golang:1.24 -> golang:1.25, the go.mod directive bump made
  `go mod download` fail
- golangci-lint v1 is built with go 1.24 and refuses to lint a go 1.25
  module, so bump the workflows and the Makefile to v2 and add a
  .golangci.yml that keeps the v1 default checks

Behaviour:

- do not create a keyword monitor without a keyword value, and return on a
  failed conversion of the current monitor in Update instead of continuing
  with an empty monitor
- reconnect backoff: Uptime Kuma allows only 20 logins per minute for all
  clients, so a dropped connection is re-established after 5s, 10s, ... up to
  5m, and an authentication error backs off for 5m with an actionable log
- validate notification IDs against the notifications of the account: Kuma
  answers with a foreign key error but still creates the monitor without
  them, which made the controller retry a failing update on every reconcile

Docs:

- quote keywordExists/keywordValue in the example, the API server rejects
  the unquoted `no` (boolean) and `404` (number)
- describe what keywordExists actually does and why it matches the
  UptimeRobot provider, and update the CRD description accordingly

Tests: cover the retry interval, timeout and redirect defaults, the missing
keyword value, the auth error detection, the reconnect backoff and the
"not reconnecting yet" guard.
@davidgrldo

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review — the three blockers were all real, and they are fixed in 3832d8b.

Blockers

Fix
Dockerfile still on golang:1.24 bumped to golang:1.25
golangci-lint v1.64 cannot lint a go 1.25 module workflows (pull_request.yml, push.yml) and Makefile on v2.14.0, plus a .golangci.yml that keeps the v1 default checks so the migration does not turn on new checks for existing code
connection closed right after login kuma.New now gets context.Background() (WithConnectTimeout already bounds the connect)
no RetryInterval set to the interval on create and update
MaxRedirects/Timeout were 0 Kuma UI defaults: 10 redirects, interval * 0.8 timeout, on create and update

Behaviour

  • A keyword monitor without keywordValue is no longer created, and Update returns on a failed conversion of the current monitor instead of continuing with an empty monitor.
  • Reconnect backoff: a dropped connection is re-established after 5s, 10s, 20s … capped at 5m instead of logging in again on the next reconcile, and an authentication error backs off for 5m with an actionable log — that was the 20-logins-per-minute problem.
  • Notification IDs are validated against the notifications of the account before create/update, so an unknown ID fails with a clear error instead of the create-without-notifications + retry loop.
  • Docs: keywordExists/keywordValue are quoted in the example, plus a section on what keywordExists actually does (see the inline comment on keywordInvert — I went with changing the CRD text rather than flipping the mapping, and explained why).

Verification: go build ./..., go vet ./..., go test ./... (all packages), gofmt, golangci-lint v2.14.0 run → 0 issues, helm lint. New unit tests cover the retry interval/timeout/redirect defaults, the missing keyword value, auth-error detection, the backoff growth and the "not reconnecting yet" guard.

CI: the Lint and Helm Lint steps pass, the multi-arch image build is running. Live re-testing against a Kuma 2.5.5 instance would still be appreciated for the notification-ID check and the keyword polarity.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[ENHANCE] Support for Uptime Kuma

2 participants