Kubernetes Monitoring with Uptimy Agent
Install the agent with Helm, label what you want watched, and get your first alert. About ten minutes.
Overview
Uptimy Agent is an open-source monitor that runs inside your cluster. Instead of creating monitors one by one, you label resources and the agent works out how to check them:
- Service: an HTTP request to
name.namespace.svcon its pods' readinessProbe path, or a TCP connect for ports that aren't HTTP. - Ingress or HTTPRoute: each hostname, over HTTPS when the Ingress has TLS for it.
- Deployment, StatefulSet or DaemonSet: all replicas ready.
- CronJob: a heartbeat on its schedule, with each run read from its Jobs.
Monitors appear within 30 seconds of labeling and go away when the label or the resource does. The agent only reads the cluster; it never changes it.
Before you start
helm3.8 or later andkubectlwith access to the cluster.- A default StorageClass for the agent's 1 Gi volume (or an existing PVC).
- Somewhere to send alerts: Slack, email, PagerDuty, Teams, ntfy or a webhook.
1. Install
helm install uptimy-agent oci://ghcr.io/uptimy/charts/uptimy-agent \
-n monitoring --create-namespaceThe chart creates a read-only ClusterRole for checks and discovery. To keep the agent to its own namespace, add --set rbac.clusterWide=false. The chart is on Artifact Hub, and its image and chart are signed with cosign.
2. Sign in and add an alert channel
kubectl -n monitoring logs deploy/uptimy-agent | grep -i password
kubectl -n monitoring port-forward svc/uptimy-agent 8080:80Open http://localhost:8080, sign in as admin with the password from the log, and choose your own. Then add a channel under Notifications. New channels alert for every monitor, including the ones discovered later, so this is the only alerting setup you need.
đź’ˇ Set the password up front
--set existingSecret=<secret> (a Secret with key ADMIN_PASSWORD) to skip the generated one.3. Label what you want monitored
Straight in the cluster, to try it:
kubectl -n shop label service checkout upti.my/monitor=true
kubectl -n shop label deployment checkout upti.my/monitor=true
kubectl -n shop label cronjob backup upti.my/monitor=trueFor good, put the label in your manifests or Helm chart, so every app and every environment is monitored when it's deployed:
apiVersion: v1
kind: Service
metadata:
name: checkout
namespace: shop
labels:
upti.my/monitor: "true"
annotations:
upti.my/name: Checkout API # optional
upti.my/interval: 30s # optional, default 1mWithin 30 seconds the monitors show up under Healthchecks and Heartbeats, marked “Discovered in Kubernetes”. Under Settings → Kubernetes you see every resource in the cluster, which ones are monitored, and the kubectl label command for any of them.
4. Your first alert
Scale a labeled Deployment to fewer ready replicas, or break its readiness:
kubectl -n shop scale deployment checkout --replicas=0The Service check fails on its next run and alerts after two failures in a row (about two minutes at the default interval, 20 seconds with upti.my/interval: 10s). Scale it back and a recovery alert follows.
CronJobs without pings
A labeled CronJob gets a heartbeat on its own schedule and time zone. The agent reads its Jobs every 10 seconds and records each run at the times Kubernetes saw: when it started, how long it took, and whether it completed or failed. A failure carries the reason and the container's exit code, for example BackoffLimitExceeded; container backup exited with code 137 (OOMKilled). A run that never happens, or doesn't finish in time, alerts as missed.
A run has to finish within its grace period after it's due. By default that's the Job's activeDeadlineSeconds plus a minute, or else the time between runs, at most an hour. For longer jobs, set it:
metadata:
labels:
upti.my/monitor: "true"
annotations:
upti.my/grace: 2hSuspending the CronJob pauses its heartbeat, and resuming it resumes it.
The agent's check or Kubernetes' own
By default the agent checks a Service with its own request, end to end through DNS, the Service and the app. That catches a pod that's “ready” but returning 500s. With upti.my/type: kubernetes it reuses the kubelet's readinessProbes instead: the Service is up while it has a ready endpoint, and the app gets no extra traffic. That works for any port and protocol, databases and gRPC included, but only knows what the readinessProbe tests.
Use the agent's check for what users reach, and Kubernetes' for internal plumbing. Labeling both a Service and its Deployment gives you both: one says users can't reach it, the other why.
Annotations
| Annotation | For | Default |
|---|---|---|
upti.my/name | all | namespace/name, or the hostname |
upti.my/path | Service, Ingress, HTTPRoute | the readinessProbe path, else / |
upti.my/port | Service | the first port |
upti.my/type | Service | http, tcp or kubernetes, picked from the port |
upti.my/scheme | Service, Ingress, HTTPRoute | http or https, picked from the port or TLS |
upti.my/interval | all but CronJobs | 1m |
upti.my/grace | CronJob | see above |
upti.my/expected-status, upti.my/keyword | HTTP checks | 200-399, none |
upti.my/monitor also accepts yes, 1 and on; false, no, 0 and off opt a resource out.
When a resource doesn't show up
Settings → Kubernetes lists what the last scan couldn't use: a mistyped label value, an invalid annotation, or missing RBAC. A resource with a problem keeps the monitor it had, so a typo never deletes history. If the scope there says one namespace, the agent was installed with rbac.clusterWide=false.
When the whole cluster goes down
Nothing inside the cluster can tell you the cluster is gone. In the agent, Settings → Connect to Uptimy makes it check in with an Uptimy heartbeat every minute, and Uptimy alerts you from outside when it goes quiet. The agent's alerts can also open incidents in Uptimy through the Uptimy Agent integration.
Next steps
- Chart values and monitors in values.yaml
- Uptimy Agent vs Uptime Kuma, including the importer
- Status pages for what customers should see