Key takeaways
- The stack is four components on four ports: Prometheus (9090), which fetches metrics itself over HTTP from each target's
/metricsendpoint (pull model), node_exporter (9100), which exposes system metrics, Grafana (3000) for visualization and Alertmanager (9093), which deduplicates, groups and routes alerts. - Always validate before restarting:
promtool check config /etc/prometheus/prometheus.yml,promtool check rulesfor alerting rules andamtool check-configon the Alertmanager side. With the--web.enable-lifecycleflag, a simplecurl -X POST http://localhost:9090/-/reloadhot-reloads the configuration without interrupting scraping. - Three PromQL queries cover most needs: CPU
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100), memory(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100and latencyhistogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))—rate()only applies to counters. In Grafana, importing the Node Exporter Full dashboard (ID 1860) gives you those views without writing a single query. - An alerting rule is only as good as its
for:clause, which filters out transient spikes:up == 0for 2 minutes for a dead instance, CPU > 85 % for 5 minutes, disk > 85 % for 10 minutes. In Alertmanager,group_wait: 30s,group_interval: 5mandrepeat_interval: 4h(1 h for critical) keep the noise down, andinhibit_rulesmute warnings as soon as a critical alert exists for the same alertname/instance pair. - Never expose ports 9090, 9100 and 9093 to the Internet. Prometheus has supported basic authentication natively since version 2.24 through
--web.config.fileand a bcrypt hash generated withhtpasswd -nBC 10; otherwise put it behind an Nginx reverse proxy with TLS. On the Grafana side, change the defaultadmin/adminand setallow_sign_up = false.
This tutorial covers the complete deployment of a monitoring stack with Prometheus for metrics collection, Grafana for visualization, Node Exporter for system metrics and Alertmanager for notifications. By the end of this guide, you will have a working monitoring infrastructure capable of monitoring your servers and alerting you when something goes wrong.
Prerequisites
- Operating system: Debian 12+ or Ubuntu 22.04+ (the commands can be adapted to CentOS/RHEL)
- Privileges: Root or sudo access on the monitoring server
- Resources: Minimum 2 GB of RAM and 20 GB of disk space (adjust according to the number of targets)
- Network: Ports 9090 (Prometheus), 9100 (Node Exporter), 3000 (Grafana), 9093 (Alertmanager) accessible
- Knowledge: Basics of Linux administration and systemd service management
Stack architecture
Prometheus works on a pull model: it is the Prometheus server that periodically queries the targets to retrieve their metrics. This model has several advantages over push (StatsD/Graphite style):
- Prometheus: Central server that scrapes metrics over HTTP, stores them locally as time series and evaluates alerting rules
- Exporters: Lightweight agents that expose metrics in the Prometheus format on a
/metricsendpoint. The most common isnode_exporterfor system metrics - Grafana: Visualization interface that connects to Prometheus as a data source to create interactive dashboards
- Alertmanager: Component that receives the alerts fired by Prometheus, deduplicates them, groups them and routes them to the right recipients (email, Slack, PagerDuty)
+------------------+ +-------------------+ +------------------+
| Node Exporter | <---- | Prometheus | ----> | Alertmanager |
| (port 9100) | | (port 9090) | | (port 9093) |
+------------------+ +-------------------+ +------------------+
|
v
+-------------------+
| Grafana |
| (port 3000) |
+-------------------+
Premium Content
This advanced tutorial is reserved for premium members.
- All advanced tutorials
- New content every week
- Progress tracking
- Cancel anytime