fix: drop the Alloy healthcheck that could only ever fail - #96
Merged
Merged
Conversation
It reported unhealthy for twenty hours while working perfectly: Prometheus scraped it the whole time and Loki holds the logs of all eleven datamap_ containers. The image carries a shell but no http client, so the check was running a command that does not exist. Loki had the same problem and was fixed when its image turned out to be distroless; Alloy was not checked the same way, because the local test only ever saw it as `health: starting`. The scrape is the signal, and TargetDown is the alert. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EQda9NZvkbStEeNU54Tqgh
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Found while validating the deploy:
datamap_alloyhad beenunhealthyfor twenty hours.The image carries a shell but no HTTP client, so the check was running a command that does not exist. It never had a chance to pass.
It was working the whole time
And Loki holds the logs of all eleven
datamap_*containers, and none of the three neighbouring projects' — so the discovery, the filter and the shipping are all fine.Why I missed it
Loki had exactly this problem and I fixed it when its image turned out to be distroless. I did not check Alloy the same way: the local run only ever showed
health: startingbefore I moved on to the next thing, so the check never had time to fail in front of me.The lesson is the narrow one —
health: startingis nothealthy, and waiting for the difference is the whole point of looking.What replaces it
Nothing new. The Prometheus scrape was already the real signal, and
TargetDownwas already the alert — which is why a container reporting unhealthy for a day changed nothing about whether the logs arrived.🤖 Generated with Claude Code
https://claude.ai/code/session_01EQda9NZvkbStEeNU54Tqgh