Repository navigation
feat(uptimekuma): add Uptime Kuma monitor provider - #678
davidgrldo wants to merge 3 commits into
Conversation
Add support for Uptime Kuma (2.x) as an uptime monitor provider, managed through its socket.io API via github.com/breml/go-uptime-kuma-client: - New UpTimeKumaMonitorService implementing the MonitorService interface with a single lazily (re)connected client, guarded by a mutex - New uptimeKumaConfig in the EndpointMonitor CRD: interval (min 20), monitorType (http|keyword), keywordValue, keywordExists, notifications - Auth via username/password reusing the existing provider config fields (Uptime Kuma API keys are not accepted for monitor management) - Update fetches the full monitor object and overrides controller-managed fields, matching Uptime Kuma's editMonitor semantics - Docs: docs/uptimekuma-configuration.md, README provider list, sample test config in CONTRIBUTING Dependency note: github.com/breml/go-uptime-kuma-client requires go >= 1.25.2 (go directive bumped; CI reads the Go version from go.mod). The remaining go.mod churn is MVS-forced by the new module's requirements, verified minimal. Fixes stakater#627 Signed-off-by: davidgrldo <davidgrldo123@gmail.com>
SyedaFatimaKazmi
left a comment
There was a problem hiding this comment.
Hi @davidgrldo , your contribution is much appreciated! The provider is well structured.
I reviewed the code and tested the PR in a kind cluster against Uptime Kuma 2.5.5.
Summary
- With the code as is, the controller cannot create any monitor. Three things block it: the connect context is cancelled right after login,
RetryIntervalis 0 (Kuma rejects it), and the Docker image does not build because the Dockerfile still uses Go 1.24. CI lint also fails for the same Go 1.25 reason. - After two small local fixes (background context and a retry interval), the rest works well: create, update, delete, switching between http and keyword, notifications, and reconnect after a Kuma restart. With 15 monitors and 5 concurrent reconciles there were no duplicates or errors.
- A few behaviours need a look: the
keywordExistsmapping is reversed compared to the CRD text and the UptimeRobot provider, redirects show as DOWN (max redirects 0), there is no request timeout, and repeated reconnects can trigger Kuma's login rate limit.
I left inline comments on each point and suggested fixes.
Happy to re-test once the changes are in. Thanks again!
| module github.com/stakater/IngressMonitorController/v2 | ||
|
|
||
| go 1.24.0 | ||
| go 1.25.2 |
There was a problem hiding this comment.
We now need Go 1.25.2, but the image in Dockerfile #L2 is still golang:1.24 - making the build stop at go mod download with go.mod requires go >= 1.25.2. Please change it to golang:1.25. The CI "Build image" step will hit this too once lint passes
similarly in the pull_request.yaml workflow
L#37 lint fails here because golangci-lint v1.64 is built with Go 1.24. It cannot lint a Go 1.25.2 module. Please bump the version to a release built with Go 1.25 (v2.x).
There was a problem hiding this comment.
Fixed in 3832d8b.
Dockerfile:golang:1.24->golang:1.25..github/workflows/pull_request.ymlandpush.yml:golangci-lint-action@v6->@v8withversion: v2.14.0. v1.64 refuses to run at all on this module, which I could reproduce with the released binary:Error: can't load config: the Go language version (go1.24) used to build golangci-lint is lower than the targeted Go version (1.25.2)Makefile:GOLANGCI_LINT_VERSION->v2.14.0and the install path ->github.com/golangci/golangci-lint/v2/cmd/golangci-lintsomake verifyworks too.
Because v2 enables checks that v1 did not (stylecheck + quickfix are merged into staticcheck), the v2 run reported 12 pre-existing issues in other providers. I added a small .golangci.yml that keeps the v1 defaults (exclusions.presets is the old exclude-use-default, and the staticcheck.checks list reproduces v1's default set), then fixed the 6 remaining real findings (ST1005/ST1017) so lint is at 0 issues. Locally verified with golangci-lint v2.14.0 and the CI Lint step passes.
| ctx, cancel := context.WithTimeout(context.Background(), connectTimeout) | ||
| defer cancel() | ||
|
|
||
| client, err := kuma.New(ctx, service.url, service.username, service.password, |
There was a problem hiding this comment.
This context is cancelled when ensureClient returns (defer cancel()). The client keeps that context for the whole socket connection. So the connection closes right after login. Every call then fails with context canceled. I could not create one monitor with this code. Please pass context.Background() here. WithConnectTimeout already limits the connect time. The TestUpTimeKumaMonitorLifecycle live test fails on this too.
There was a problem hiding this comment.
Good catch, this was the showstopper. Fixed in 3832d8b — ensureClient now passes context.Background() to kuma.New. WithConnectTimeout(connectTimeout) already bounds the connect attempt, so nothing is lost.
| Name: m.Name, | ||
| Interval: int64(normalized.Interval), | ||
| NotificationIDs: parseNotificationIDs(normalized.Notifications), | ||
| IsActive: true, |
There was a problem hiding this comment.
RetryInterval is not set, so it is 0. Kuma 2.5.5 rejects it: Retry interval cannot be less than 1 seconds. No monitor can be created. Please set it, for example to the same value as Interval. That is the Kuma UI default.
There was a problem hiding this comment.
Fixed in 3832d8b. RetryInterval is now set to the interval on create (buildKumaMonitor) and on update, using the Kuma UI default.
| details := monitor.HTTPDetails{ | ||
| URL: m.URL, | ||
| Method: "GET", | ||
| AcceptedStatusCodes: []string{"200-299"}, |
There was a problem hiding this comment.
MaxRedirects and Timeout are not set, so both are 0. With 0 redirects, any URL that redirects (for example http to https) is DOWN with status code 302. With timeout 0, Kuma waits interval × 800 seconds (about 13 hours at interval 60). A site that hangs never goes DOWN. Please use the Kuma UI defaults: MaxRedirects: 10 and Timeout: interval × 0.8.
There was a problem hiding this comment.
Fixed in 3832d8b, on create and on update, with the Kuma UI defaults: MaxRedirects: 10 and Timeout: interval * 0.8 (requestTimeoutFor). Covered by tests for 20/60/120/300 second intervals.
| // keywordInvert maps KeywordExists to Uptime Kuma's invertKeyword flag: | ||
| // "no" means alert when the keyword does NOT exist | ||
| func keywordInvert(keywordExists string) bool { | ||
| return strings.EqualFold(keywordExists, "no") |
There was a problem hiding this comment.
This mapping is reversed. In Kuma, invertKeyword=false means DOWN when the keyword is missing. So keywordExists: "yes" alerts when the keyword is missing. The CRD says yes = alert if the value exists. The UptimeRobot provider does that. Test: keywordExists: "no" went DOWN with keyword is present. Please flip this function and invertToKeywordExists, or change the CRD text.
There was a problem hiding this comment.
I checked this against the Kuma 2.5.5 source, and I believe the mapping is the other way round. In server/model/monitor.js the check is:
let keywordFound = data.includes(this.keyword);
if (keywordFound === !this.isInvertKeyword()) {
bean.status = UP;
}So invertKeyword=false -> UP when the keyword is found, and invertKeyword=true -> UP only when it is absent. That is what this provider does today (keywordExists: "no" -> invertKeyword=true), it is the same as UptimeRobot's keyword_type 1 (contains) and 2 (does not contain), and it is what the docs example relies on: keywordValue: "404" with keywordExists: "no" stays UP while the page is healthy and alerts as soon as a 404 page shows up.
Your own test result confirms it: with keywordExists: "no" the monitor went DOWN while the keyword was present, which is the documented behaviour rather than a bug.
The confusing part is the CRD wording "Alert if value exist (yes) or doesn't exist (no)", read literally it implies the opposite. I took your second option and changed the text instead, in the CRD description and in the docs:
yes(default) — up while the keyword is in the response, down when it disappearsno— up only while the keyword is absent, so it alerts as soon as the keyword shows up
If you did mean the literal reading of the old text (i.e. yes = alert when present), say so and I will flip keywordInvert/invertToKeywordExists instead, but that would make the provider inconsistent with UptimeRobot and invert the example above.
| } | ||
|
|
||
| // invalidateClient drops the connection so the next call reconnects | ||
| func (service *UpTimeKumaMonitorService) invalidateClient() { |
There was a problem hiding this comment.
Every failed call drops the client. The next reconcile logs in again. Kuma allows only 20 logins per minute for all clients together. With a wrong password, the controller used all 20 in a few seconds. Kuma then blocked my own UI login too. Please add a backoff before a reconnect, or do not reconnect on auth errors
There was a problem hiding this comment.
Fixed in 3832d8b. ensureClient now tracks a reconnect deadline: a dropped connection is re-established after 5s, then 10s, 20s, ... capped at 5m, and ensureClient returns an error instead of logging in again while that backoff is active. An authentication failure (the client reports Incorrect username or password / authIncorrectCreds) backs off for 5 minutes and logs which account/URL to check, so a wrong password can no longer burn the 20 logins per minute.
|
|
||
| if normalized.MonitorType == "keyword" { | ||
| if len(normalized.KeywordValue) == 0 { | ||
| log.Error(nil, "Monitor is of type Keyword but the `keyword-value` is missing") |
There was a problem hiding this comment.
This logs an error but still creates a keyword monitor with an empty keyword. Please return here instead.
There was a problem hiding this comment.
Fixed in 3832d8b. buildKumaMonitor now returns an error and Add bails out instead of creating a keyword monitor with an empty keyword; Update checks the same before it sends anything. The conversion failures of the current monitor in Update now return as well instead of continuing with a zero-value monitor.
| interval: 120 | ||
| monitorType: keyword | ||
| keywordExists: no | ||
| keywordValue: 404 |
There was a problem hiding this comment.
The API server rejects this example. YAML reads no as a boolean and 404 as a number, but both fields are strings. Please quote them: keywordExists: "no" and keywordValue: "404"
There was a problem hiding this comment.
Fixed in 3832d8b, both values are quoted now (keywordExists: "no", keywordValue: "404"). I also added a note about why they have to be quoted, since YAML reads no as a boolean and 404 as a number.
| return nil, nil | ||
| } | ||
|
|
||
| func (service *UpTimeKumaMonitorService) Add(m models.Monitor) { |
There was a problem hiding this comment.
With a notification ID that does not exist, Kuma returns a foreign-key error but still creates the monitor without notifications. The controller then retries the update on every reconcile and fails each time. A check against GetNotifications, or a clearer error, would help.
There was a problem hiding this comment.
Fixed in 3832d8b. Before creating or updating, the configured IDs are checked against client.GetNotifications() for the account the controller logs in with, and an unknown ID aborts the operation with an error naming the missing IDs and pointing at Settings > Notifications. That stops the create-without-notifications plus retry-every-reconcile loop.
Blockers: - pass context.Background() to kuma.New: the client keeps that context for the whole socket connection, so the deferred cancel closed it right after login and every call afterwards failed with "context canceled" - set retryInterval (Kuma rejects a monitor without one), maxredirects (10) and timeout (interval * 0.8) to the Kuma UI defaults, on create and on update; with maxredirects 0 every redirect was reported DOWN and with timeout 0 a hanging endpoint never went down - Dockerfile: golang:1.24 -> golang:1.25, the go.mod directive bump made `go mod download` fail - golangci-lint v1 is built with go 1.24 and refuses to lint a go 1.25 module, so bump the workflows and the Makefile to v2 and add a .golangci.yml that keeps the v1 default checks Behaviour: - do not create a keyword monitor without a keyword value, and return on a failed conversion of the current monitor in Update instead of continuing with an empty monitor - reconnect backoff: Uptime Kuma allows only 20 logins per minute for all clients, so a dropped connection is re-established after 5s, 10s, ... up to 5m, and an authentication error backs off for 5m with an actionable log - validate notification IDs against the notifications of the account: Kuma answers with a foreign key error but still creates the monitor without them, which made the controller retry a failing update on every reconcile Docs: - quote keywordExists/keywordValue in the example, the API server rejects the unquoted `no` (boolean) and `404` (number) - describe what keywordExists actually does and why it matches the UptimeRobot provider, and update the CRD description accordingly Tests: cover the retry interval, timeout and redirect defaults, the missing keyword value, the auth error detection, the reconnect backoff and the "not reconnecting yet" guard.
|
Thanks for the thorough review — the three blockers were all real, and they are fixed in Blockers
Behaviour
Verification: CI: the Lint and Helm Lint steps pass, the multi-arch image build is running. Live re-testing against a Kuma 2.5.5 instance would still be appreciated for the notification-ID check and the keyword polarity. |
Fixes #627
What
Adds Uptime Kuma (2.x) as a monitor provider, following the
Adding support for a new Monitorconventions in CONTRIBUTING.md. Uptime Kuma has no REST API for monitor management — its officially sanctioned management API is socket.io, so this usesgithub.com/breml/go-uptime-kuma-client(MIT, actively maintained, tested against Kuma 2.5.x, also used by the Terraform providers).Changes
pkg/monitors/uptimekuma:UpTimeKumaMonitorServiceimplementing theMonitorServiceinterface. One socket.io client held for the service lifetime, lazily (re)connected after failures under a mutex.GetByNamematches on the full monitor list (Kuma has no lookup-by-name);Updatefetches the current monitor and overrides controller-managed fields, matching Kuma'seditMonitorfull-object semantics.uptimeKumaConfig(intervalmin 20,monitorTypehttp|keyword,keywordValue,keywordExists,notifications). Generated deepcopy + CRD manifests updated in all three copies.monitor-proxy.go(OfType/ExtractConfigcases,TypeUptimeKuma) and controller dispatch.apiURL/username/passwordfields — no new global config.Equalcomparisons (325 lines, stdlib testing).docs/uptimekuma-configuration.md(config example, CRD example, Kuma 2.x-only note, auth notes: API keys are not accepted for monitor management, 2FA must be off for the service account, recommend pointing at the internal HTTP service), README provider list, CONTRIBUTING test-config sample.Verification
go build ./...✅go vet ./pkg/monitors/uptimekuma/... ./api/... ./pkg/monitors/...✅go test ./...— all 15 packages ✅ (offline)gofmtclean ✅Notes for reviewers
go >= 1.25.2, so thegodirective was bumped 1.24.0 → 1.25.2 (CI workflows usego-version-file: go.mod, so they follow automatically). The rest of the go.mod churn is MVS-forced by the new module's requirements — verified minimal by re-resolving from a clean state (identical result).