Monitoring and Metrics
EMQX monitoring and Metrics: Dashboard statistics, node status, system resources, the concepts behind the Prometheus integration, and message rate metrics.
Why this matters
A broker that works is not the same as a broker that is healthy under load. To keep the message backbone of a house running, you need to know how many connections and subscriptions there are right now, how much memory each node has left, and whether the rate of messages in and out has blown up. The EMQX Monitoring module is where all of those numbers are gathered in one place.
This chapter walks you through the cluster statistics, the node status and the system resource columns in the Dashboard, and through the two kinds of figures EMQX gives you: Stats and Metrics. It closes with Prometheus, as an introduction to moving EMQX monitoring into a monitoring system of your own, so you can build your own Grafana dashboard at home.
Core concepts
EMQX splits monitoring data into two kinds. Statistics are integer gauges: a single number for one moment in time. Metrics are integer counters that accumulate, such as the running total of bytes and messages. Both are the most direct thing to watch in the Dashboard.
Expand Monitoring in the Dashboard sidebar and you get a handful of sub-items: Cluster Overview, Clients, Subscriptions and Retained Messages, plus Delayed Publish and Alarms — the last two are EMQX Enterprise features.
If you would rather not open the Dashboard, you have three other routes: pull the same data over the REST API, subscribe to the system topics that start with $SYS/, or push the metrics to third-party monitoring such as Prometheus. Prometheus comes up at the end of this chapter.
Terms at a glance
| Term | Plain English | What it means |
|---|---|---|
| Monitoring | cluster monitoring | The Dashboard module for watching the state of the cluster |
| Cluster Overview | cluster-wide summary | The page you land on after login, with cluster-level numbers |
| Node | one EMQX instance | A single EMQX instance in the cluster |
| Statistics | point-in-time counts | An integer gauge for one moment in time |
| Metrics | running totals | Counters that accumulate (bytes, packets, messages, events) |
| Prometheus | metrics collector | A third-party service that monitors systems and collects metrics |
Hands-on
Open the Monitoring module
In the sidebar, choose Monitoring → Cluster Overview. The first thing you see is the cluster-level connection, subscription and topic counts.
Switch to the Nodes tab
In the top half, switch the group to Nodes to expand every node in the cluster (a home add-on install is usually a single node).
Read the system resources and the version
The node row shows the connection count, the version, uptime, the number of Erlang processes and memory/CPU usage. Click the node name to go into the node details.
Use the Metrics tab for the cumulative counters
Switch to the Metrics tab and work through the four groups of counters one by one: bytes, packets, messages and events. Read the numbers together with the time range you have selected.
Cluster Overview: connections and nodes at a glance
Cluster Overview is the first item under Monitoring. Switch to the Nodes tab and you see each node's name, status, uptime, connection count, version, Erlang process count, and memory and CPU usage. A node drawn in gray has stopped; click its name to open the node details and read the system paths and the log path.
A home setup is usually a single node, so the first entry on the Node tab is your add-on itself. The Version column is the most direct place to confirm that the EMQX you are running is 5.8.9.
The time series chart below can be set to the last 1 hour, 6 hours … up to 7 days, so you can watch how connection counts and message volume trend.
To make the Dashboard start counting again, use Reset Monitoring Data to clear what has accumulated so far and start over.
Statistics and Metrics
On the Metrics tab of Cluster Overview, EMQX's Metrics come in four dimensions: bytes (bytes in and out), packets (MQTT packets of each kind, sent and received), messages (message counts, including QoS 0/1/2, received, sent, forwarded and dropped) and events (event counts, such as connections, sessions, authentication and access).
Another part of it is Statistics, for example:
connections.count/connections.max: the current connection count and the all-time maximum.sessions.count/sessions.max: the current session count.subscriptions.count: the total number of subscriptions right now (shared subscriptions included).topics.count: the number of unique topics right now.retained.count: the number of retained messages right now.delayed.count: the number of delayed-publish messages right now.
Most of these are shown as a current value plus an all-time maximum, which suits long-term trends.
Integrating with Prometheus
EMQX can expose its monitoring data to a third-party monitoring system, and the most common one is Prometheus. That puts EMQX on the same chart as everything else you measure on the host, or lets you draw a visual dashboard with Grafana and have Alertmanager tell you when something is off.
On the Dashboard's Management → Monitoring page, open the "Integration" tab and choose the Prometheus settings. EMQX supports two modes:
- Pull mode: Prometheus polls EMQX's REST API on a schedule, for example
/api/v5/prometheus/stats(basic metrics and counters),/api/v5/prometheus/auth(access control) and/api/v5/prometheus/data_integration(rules, Connectors, Actions and Sinks). - Push mode: EMQX pushes the metrics to a Pushgateway (off by default) and Prometheus collects them from there; right now this mode covers only the basic metrics.
Pull mode is what most people use: point the url in your Prometheus configuration at EMQX and the metrics start arriving.
The pull-mode API needs no authentication by default. To turn Basic Auth on, create an API key in EMQX and put the key into your Prometheus configuration.
Troubleshooting
- The numbers on the home page look stuck: the Overview charts have a time range and an aggregation interval you can change; widen the range and look again, or press Reset Monitoring Data to start accumulating from scratch.
- A node is drawn in gray: that node has stopped. Do not read the data as a missing entry; check the node status column first.
- You need the EMQX version you are running: go to Monitoring → Nodes; the Version column is the EMQX version currently running (this guide targets 5.8.9).
- Delayed Publish/Alarms have vanished from the sidebar: both are EMQX Enterprise features. Open Source 5.8.9 does not offer them, and nothing is broken.
FAQ
What is the actual difference between Statistics and Metrics?
Statistics are integer gauges read in one shot (how many subscriptions there are right now, what the all-time maximum was). Metrics are counters that accumulate (how many bytes have been received in total, how many packets have gone out). The Dashboard keeps both under Monitoring.
Why can't I see Alarms/Delayed Publish in the Dashboard?
Because they are EMQX Enterprise features. Open Source 5.8.9 does not show them in the Monitoring sidebar. That is a licensing boundary, not a setting you got wrong.
Should I pick Pull or Push for Prometheus?
Most people pick Pull, and the official EMQX documentation recommends Pull too, because the Pushgateway push mode currently covers only the basic metrics — it is not as complete as auth or data_integration.
How do I check one node's system resources?
Go to Monitoring → Nodes and click the node name to open its details page, which holds memory/CPU usage, the Erlang process count, the connection count and so on. On a single node at home, read that page as the health of your EMQX instance.