Home

Prometheus + Grafana: Complete Monitoring Stack

Monitoring
Difficulty: Advanced
14 min read

Step-by-step guide to install and configure a monitoring stack with Prometheus, Grafana and Alertmanager on a Linux infrastructure.

Back to tutorials

Key takeaways

  • The stack is four components on four ports: Prometheus (9090), which fetches metrics itself over HTTP from each target's /metrics endpoint (pull model), node_exporter (9100), which exposes system metrics, Grafana (3000) for visualization and Alertmanager (9093), which deduplicates, groups and routes alerts.
  • Always validate before restarting: promtool check config /etc/prometheus/prometheus.yml, promtool check rules for alerting rules and amtool check-config on the Alertmanager side. With the --web.enable-lifecycle flag, a simple curl -X POST http://localhost:9090/-/reload hot-reloads the configuration without interrupting scraping.
  • Three PromQL queries cover most needs: CPU 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100), memory (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 and latency histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) — rate() only applies to counters. In Grafana, importing the Node Exporter Full dashboard (ID 1860) gives you those views without writing a single query.
  • An alerting rule is only as good as its for: clause, which filters out transient spikes: up == 0 for 2 minutes for a dead instance, CPU > 85 % for 5 minutes, disk > 85 % for 10 minutes. In Alertmanager, group_wait: 30s, group_interval: 5m and repeat_interval: 4h (1 h for critical) keep the noise down, and inhibit_rules mute warnings as soon as a critical alert exists for the same alertname/instance pair.
  • Never expose ports 9090, 9100 and 9093 to the Internet. Prometheus has supported basic authentication natively since version 2.24 through --web.config.file and a bcrypt hash generated with htpasswd -nBC 10; otherwise put it behind an Nginx reverse proxy with TLS. On the Grafana side, change the default admin/admin and set allow_sign_up = false.
Prometheus + Grafana monitoring stack
This tutorial covers the complete deployment of a monitoring stack with Prometheus for metrics collection, Grafana for visualization, Node Exporter for system metrics and Alertmanager for notifications. By the end of this guide, you will have a working monitoring infrastructure capable of monitoring your servers and alerting you when something goes wrong.

Prerequisites

  • Operating system: Debian 12+ or Ubuntu 22.04+ (the commands can be adapted to CentOS/RHEL)
  • Privileges: Root or sudo access on the monitoring server
  • Resources: Minimum 2 GB of RAM and 20 GB of disk space (adjust according to the number of targets)
  • Network: Ports 9090 (Prometheus), 9100 (Node Exporter), 3000 (Grafana), 9093 (Alertmanager) accessible
  • Knowledge: Basics of Linux administration and systemd service management

Stack architecture

Prometheus works on a pull model: it is the Prometheus server that periodically queries the targets to retrieve their metrics. This model has several advantages over push (StatsD/Graphite style):

  • Prometheus: Central server that scrapes metrics over HTTP, stores them locally as time series and evaluates alerting rules
  • Exporters: Lightweight agents that expose metrics in the Prometheus format on a /metrics endpoint. The most common is node_exporter for system metrics
  • Grafana: Visualization interface that connects to Prometheus as a data source to create interactive dashboards
  • Alertmanager: Component that receives the alerts fired by Prometheus, deduplicates them, groups them and routes them to the right recipients (email, Slack, PagerDuty)

+------------------+       +-------------------+       +------------------+
|   Node Exporter  | <---- |    Prometheus      | ----> |   Alertmanager   |
|   (port 9100)    |       |    (port 9090)     |       |   (port 9093)    |
+------------------+       +-------------------+       +------------------+
                                    |
                                    v
                           +-------------------+
                           |     Grafana        |
                           |    (port 3000)     |
                           +-------------------+

Premium Content

This advanced tutorial is reserved for premium members.

9,90€ / month
  • All advanced tutorials
  • New content every week
  • Progress tracking
  • Cancel anytime
MR

Written by

Morgann Riu

Cybersecurity and Linux administration expert. I share my knowledge through free tutorials and training to help system administrators and developers secure their infrastructures.

Frequently asked questions

What is the difference between Prometheus and Grafana?
Prometheus is a metrics collection and storage system based on a pull model: it periodically queries targets to retrieve their metrics. Grafana is a visualization tool that connects to Prometheus (and other sources) to create interactive dashboards. The two are complementary: Prometheus collects, Grafana displays.
How does Prometheus collect metrics?
Prometheus uses a pull model: it sends HTTP GET requests to a /metrics endpoint exposed by each target. Exporters (node_exporter, blackbox_exporter, etc.) convert system or application metrics into the text format expected by Prometheus. The collection frequency is defined in the scrape_interval of the configuration.
Can Prometheus and Grafana be used in production?
Yes, Prometheus and Grafana are used in production by many companies (SoundCloud, DigitalOcean, CERN). Prometheus is designed for reliability: each instance is autonomous and keeps working even if the network or remote storage is unavailable. For high availability, deploy two identical Prometheus instances and use Thanos or Cortex for long-term storage.
What is PromQL and how do I learn it?
PromQL (Prometheus Query Language) is the query language of Prometheus. It lets you select, filter, aggregate and transform time series. The essential functions are rate() for counters, avg/sum/max for aggregation, and histogram_quantile() for percentiles. The best way to learn it is to experiment in the Prometheus interface or in Grafana Explore.
How do I configure alerts with Prometheus?
Alerts are configured in two steps. First, define alerting rules in Prometheus (YAML files with PromQL expressions and thresholds). Then, configure Alertmanager to route those alerts to the right recipients (email, Slack, PagerDuty). Alertmanager handles grouping, silencing and inhibition of alerts to avoid noise.
Which Prometheus exporters should I use to monitor a Linux server?
The node_exporter is essential: it exposes CPU, memory, disk, network and filesystem metrics. Add the blackbox_exporter to monitor HTTP/TCP/ICMP availability, the mysqld_exporter for MySQL/MariaDB, and the nginx-prometheus-exporter for Nginx. Each exporter is deployed as an independent systemd service.

Share this tutorial

Did you enjoy this article?

Was this article helpful?

Thanks for your feedback!

Comments

Recommended for you

In-depth article on the topic

Checklist Sécurité Linux

30 points essentiels pour sécuriser un serveur Linux. Recevez aussi les nouveaux tutoriels par email.

Pas de spam. Désabonnement en 1 clic.

↑