Server Monitoring and Log Management Tools for Sysadmins

Server monitoring and log management tools for sysadmins: Prometheus, Grafana Loki and Uptime Kuma

This guide is for sysadmins, MSP technicians and SREs who run mixed Linux and Windows servers and need three answers fast: is it up, is it about to break, and what happened last night. The short answer: a free, self-hosted stack of Prometheus for metrics, Alertmanager for alert routing, Grafana Loki for logs and Grafana for dashboards covers servers, VMs, containers and cloud instances. Add Uptime Kuma or Healthchecks for outside-in checks and cron jobs. If you prefer a web GUI over YAML files, Zabbix or Checkmk do most of this in one product. For a small Windows office, LANState plus a desktop log viewer such as LogFusion may be all you need.

The short list

ToolBest forLicencePlatformsStatus
PrometheusMetrics and alert rules for servers, containers, KubernetesFree, open source (Apache-2.0)Linux, Windows, macOS, FreeBSD, DockerActive, 3.15.0 (September 2026)
Grafana LokiCentral log aggregation queried with LogQLFree, open source (AGPL-3.0)Linux, Docker, KubernetesActive, 3.7.8 (September 2026)
LogFusionLive tailing of text, IIS and Event Logs on WindowsShareware (conditionally free)Windows 10/11, Server 2016+Active, 7.2 (July 2026)
LANStateNetwork map with agentless up/down checksShareware (conditionally free), free mode up to 25 devicesWindowsActive, 11.0 (September 2026)
NetCrunch ToolsPing, port, SNMP and syslog testing during troubleshootingFree (freeware)Windows 10/11 x64Active, version not published
MaltrailDetecting malicious traffic on a SPAN port or gatewayFree, open source (MIT engine)Linux, BSD, macOS, Windows 10+Active, 3.4 (September 2026)
GlassWirePer-app traffic on a single Windows PCShareware (conditionally free)Windows 10/11 (x64, ARM64)Active, 3.10 (September 2026)
GrafanaDashboards over Prometheus, Loki and SQL dataFree, open source (AGPL-3.0)Linux, Windows, macOS, DockerActive
AlertmanagerGrouping, silencing and routing Prometheus alertsFree, open source (Apache-2.0)Same as PrometheusActive
ZabbixAgent-based monitoring with a web GUI and templatesFree, open source (AGPL-3.0 since 7.0)Server on Linux; agents for Linux, Windows and moreActive
Checkmk CommunityAuto-discovered server and network checksFree, open source (GPL-2.0); Pro and Ultimate editions are paidServer on Linux or Docker; agents for Linux, Windows and moreActive, 2.5 series
Uptime KumaHTTP, TCP, DNS and push checks plus status pagesFree, open source (MIT)Docker, Linux, WindowsActive, 2.5.5 (September 2026)
HealthchecksDead-man’s-switch monitoring for cron jobsSelf-hosted: free, open source (BSD-3-Clause); hosted healthchecks.io: shareware (conditionally free)Docker or Python/Django; hosted serviceActive

Build the core: metrics, alerts and dashboards

For a server fleet an “observability stack” means metrics (numbers over time), logs (events) and, if you write your own services, traces. Start with metrics: they drive almost every useful alert.

Prometheus scrapes an HTTP /metrics endpoint on each target, stores samples in its own time-series database (15 days by default) and evaluates alert rules. Install node_exporter on Linux (TCP 9100) and windows_exporter on Windows (TCP 9182; older guides call it the WMI exporter):

scrape_configs:
  - job_name: node
    static_configs:
      - targets: ['web01:9100', 'db01:9100']
  - job_name: windows
    static_configs:
      - targets: ['fs01:9182']

Run promtool check config /etc/prometheus/prometheus.yml before reloading, then confirm every target is UP under Status > Targets on port 9090. For fifty hosts, deploy exporters with an Ansible playbook (see our DevOps automation tools guide).

Prometheus decides that something is wrong; Alertmanager decides who hears about it. It deduplicates and groups alerts, routes them by label to email, Slack, PagerDuty or a webhook, and supports silences and inhibition, so a dead core switch does not page you for every host behind it. Grafana adds dashboards over Prometheus, Loki and SQL databases.

If you would rather click than edit YAML, Zabbix and Checkmk are the mature alternatives: agents with ready-made templates for Linux, Windows, databases and network gear, plus full web GUIs. Checkmk’s free edition is now called Checkmk Community (the Raw Edition before the 2.5 renaming). Prometheus wins in dynamic, containerised environments; Zabbix and Checkmk are often faster to roll out across a traditional server room.

Collect and search logs from many servers

Log aggregation puts every server’s logs in one place with retention and a query language, so you stop SSH-ing into boxes to grep. Grafana Loki is the lightest self-hosted option if you already run Grafana and Prometheus: it indexes only labels such as host and job and stores compressed chunks on disk or object storage. Promtail reached end of life on 2 March 2026, so ship logs with Grafana Alloy, which tails files under /var/log and reads Windows event logs.

Three LogQL queries cover most daily work:

{job="varlogs", host="web01"} |= "error"
sum by (host) (count_over_time({job="varlogs"} |= "Failed password" [5m]))
{job="nginx"} | json | status >= 500

The second query makes a good alert: Loki’s ruler evaluates it and sends it to the same Alertmanager as your metrics. Keep labels low-cardinality (host, job, environment); request IDs as labels create millions of streams. Loki has no built-in authentication, so keep port 3100 on a private network or behind an authenticating proxy.

Network devices usually speak syslog. Before blaming the collector, send a test message with the Syslog Tester that ships with NetCrunch Tools. If you need full-text search across every word, the ELK or OpenSearch stack does that at a much higher storage cost.

On a single Windows server, LogFusion tails text and IIS logs, reads local and remote Event Logs and captures OutputDebugString output; filtering, tabs and highlighting are Pro features. With nothing installed, Get-Content C:applog.txt -Wait -Tail 50 or journalctl -fu nginx does the job.

Check that services and websites are up

Inside checks (CPU, disk, a Windows service) and outside checks (does HTTPS answer, is the certificate valid) are different, and you need both. For outside checks Prometheus has blackbox_exporter: an http_2xx module turns each URL into a target, and probe_success and probe_ssl_earliest_cert_expiry give you uptime and certificate-expiry alerts.

Uptime Kuma is easier when you just want monitors in a web UI: HTTP(S), keyword and JSON-query checks, TCP, ping, DNS records, Docker containers and push monitors, with notifications to Slack, Telegram, email and many other channels. It runs in Docker (port 3001) or on Node.js.

For a small office, LANState draws a network map from switch port tables over SNMP and runs agentless ping, TCP, HTTP, SNMP and WMI checks from one Windows PC. Free mode covers 25 devices, so map infrastructure (switches, servers, firewall, NAS) rather than desktops. When a check goes red, NetCrunch Tools or Angry IP Scanner show quickly whether it is one host or the whole subnet.

Alert when a cron job silently stops

Backups and certificate renewals fail in the worst way: nothing happens and nobody notices. The fix is a dead man’s switch: the job pings a URL when it finishes, and a missing ping raises an alert. Healthchecks is built for this. Add a check with the expected period and grace time, then append the ping:

# /etc/cron.d/nightly-backup
30 2 * * * root /usr/local/bin/backup.sh; curl -fsS -m 10 --retry 5 -o /dev/null https://hc-ping.com/<uuid>/$?

Appending /$? reports the script’s exit status, so a non-zero exit alerts immediately instead of waiting for the grace period. Ping /start at the beginning of long jobs to catch jobs that hang. Healthchecks can be self-hosted with Docker or used as the hosted healthchecks.io.

Uptime Kuma’s Push monitor works the same way. In a Prometheus shop, have the job write a timestamp to a .prom file in node_exporter’s textfile collector directory (--collector.textfile.directory) and alert on time() - backup_last_success_timestamp_seconds > 26*3600. The backups themselves are covered in our backup and server hardening tools guide.

Monitor cloud servers and hybrid estates

Cloud monitoring does not need a tool per provider. Prometheus has service discovery for EC2, Azure, Kubernetes, Consul, DNS and file-based lists, so new instances are picked up automatically. Scrape exporters inside the VMs over a private network or VPN, never the public internet: exporters have no authentication unless you configure TLS and basic auth. For managed databases and load balancers, the provider’s own monitoring service is the data source; Grafana can show it next to your metrics. For Kubernetes clusters and hypervisors, see our virtualization and container tools guide.

Track SLOs and error budgets

SLO tracking and platform reliability work start with one number per service: the share of good requests. For an HTTP service instrumented with a request counter:

sum(rate(http_requests_total{job="api",code!~"5.."}[30d]))
/
sum(rate(http_requests_total{job="api"}[30d]))

With a 99.9% objective, the error budget is 0.1% of requests over the window. Alerting on raw error rate pages people for blips; alerting on burn rate (how fast the budget is being spent) pages only when the SLO is at risk. Rather than writing burn-rate rules by hand, use Sloth or Pyrra (both free, Apache-2.0): they turn a short YAML SLO definition into Prometheus recording rules and multi-burn-rate alerts, and Pyrra adds a web UI showing remaining budget. Budget left means you can ship; budget gone means reliability work comes first.

Publish a status page

A status page saves your helpdesk from answering the same question fifty times during an outage. Uptime Kuma publishes several status pages from its monitors, each on its own domain. Host it outside the infrastructure it reports on (a small VPS elsewhere is enough) and show customer-facing services (website, email, VPN), not every internal check.

Plan capacity before disks and CPUs run out

Capacity planning with Prometheus is mostly trend queries. Predict which filesystems will fill up within four days based on the last 24 hours:

predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[24h], 4*24*3600) < 0

That beats a static 90% threshold, which fires too late on small volumes and too early on large ones. For CPU and memory, use quantile_over_time(0.95, ...) over weeks, because averages hide the peaks users feel. The default 15-day retention is too short for quarterly trends, so raise it or remote-write to long-term storage.

Write incident runbooks people actually use

A runbook is only useful if the on-call engineer finds it from the alert. Put the link in the alert itself:

- alert: DiskAlmostFull
  expr: node_filesystem_avail_bytes / node_filesystem_size_bytes < 0.10
  for: 15m
  labels:
    severity: warning
  annotations:
    summary: "Less than 10% free on {{ $labels.instance }} {{ $labels.mountpoint }}"
    runbook_url: "https://wiki.example.local/runbooks/disk-full"

Keep each runbook short: what the alert means, how to confirm it, safe first actions, when to escalate and to whom. Name the tools the responder needs, for example mRemoteNG for RDP and SSH sessions to the affected servers and WinSCP for pulling log files. During planned work, silence alerts instead of ignoring them: amtool silence add alertname=DiskAlmostFull instance=db01:9100, then amtool silence expire <id> afterwards. Update the runbook after every real incident.

Watch the network for suspicious traffic

Maltrail watches a SPAN port or gateway, matches domains, IPs and URLs against public blacklists and heuristics, and shows hits in a web dashboard on port 8338. Change the default password immediately and forward events to syslog so they land in Loki. Since 3.x the sensor is written in Rust; guides describing sensor.py are out of date. On a single Windows PC, GlassWire shows which application caused a traffic spike and alerts on the first connection from new software; its free tier keeps one day of history.

How to choose

  1. Count what you monitor. Under 25 devices in one Windows office: LANState in free mode plus Uptime Kuma. Dozens to hundreds of servers: Prometheus with Alertmanager and Grafana, or Zabbix or Checkmk Community if you want a GUI.
  2. Match your team’s habits. Config in Git and Ansible or Kubernetes already in use: Prometheus and Loki. A team that lives in web consoles will maintain Zabbix or Checkmk better.
  3. Cover the silent failures first. Put dead man’s switch checks on backups and certificate renewals before building dashboards.
  4. Add central logs once you have more than a handful of servers. Loki if you use Grafana; a desktop viewer like LogFusion is enough for one or two Windows boxes.
  5. Route alerts before adding more. Every alert needs an owner, a severity and a runbook link. An alert nobody acts on should be deleted.
  6. Monitor the monitor. Run an external check (Uptime Kuma or Healthchecks on another host) against Prometheus or Zabbix itself.

FAQ

What is the best free server monitoring tool?

For Linux and container-heavy estates, Prometheus with Alertmanager and Grafana. For a GUI with templates, Zabbix or Checkmk Community. For simple uptime checks, Uptime Kuma.

Is Grafana Loki a replacement for ELK?

For most sysadmin use, yes, at lower storage cost. It does not index full text, so broad searches over long periods are slower than in Elasticsearch.

How do I monitor a cron job?

Make the job ping a monitoring URL when it succeeds (Healthchecks or an Uptime Kuma push monitor) and alert when the ping is late.

What is the difference between Prometheus and Grafana?

Prometheus collects and stores metrics and evaluates alert rules. Grafana visualises data from Prometheus, Loki and other sources.

Can I monitor Windows servers with Prometheus?

Yes. Install windows_exporter on each server (TCP 9182) and add it as a scrape target. Ship Windows event logs to Loki with Grafana Alloy.

Do I need a separate status page tool?

Not usually. Uptime Kuma builds status pages from its own monitors. Host it outside the infrastructure it reports on, so the page stays up during the outage it describes.

Last updated: 30 September 2026 · ITForgePro editorial team. Licence, version and platform details are checked against each developer's official documentation.

Other articles

Submit your application