Introduction

Apache Ignite 3 brings a modernized, redesigned architecture to distributed database systems. However, monitoring distributed clusters effectively requires clear visibility into cluster topology, JVM performance, and storage metrics.

While earlier versions of Apache Ignite relied heavily on JMX exporters for metric ingestion, Apache Ignite 3 introduces native OpenTelemetry (OTLP) support. This allows Ignite to directly push metrics over OTLP directly to receivers like Prometheus and OpenTelemetry Collectors.

In this guide, we will walk through setting up a complete 3-node Apache Ignite 3 cluster using Docker Compose, configuring OTLP metric streaming to Prometheus, and visualizing key cluster health metrics — such as total active cluster nodes — in Grafana.

Finally, we'll validate our dashboard by simulating a node drop (docker compose stop node3) and recovery (docker compose start node3).

Architecture Overview

Our setup consists of five lightweight Docker containers connected on a shared bridge network:

3 Ignite Nodes (node1, node2, node3): Running Apache Ignite 3.1.0 with dynamic metric sources pre-enabled in bootstrap configuration.

Prometheus: Configured with --web.enable-otlp-receiver to accept OTLP metric payloads pushed directly from Ignite 3.

Grafana: Visualizing metrics ingested by Prometheus.

None

Step 1: Define the docker-compose.yml

Create a new directory for your project and create a docker-compose.yml file. Notice how we enable metric sources (jvm, os, topology.cluster, topology.local) directly inside the HOCON bootstrap configuration block (node_config).

name: ignite3

x-ignite-def: &ignite-def
  image: apacheignite/ignite:3.1.0
  environment:
    JVM_MAX_MEM: "4g"
    JVM_MIN_MEM: "4g"
    BOOTSTRAP_NODE_CONFIG: /opt/ignite/etc/ignite-config.conf
  configs:
    - source: node_config
      target: /opt/ignite/etc/ignite-config.conf
      mode: 0644

services:
  node1:
    <<: *ignite-def
    command: --node-name node1
    ports:
      - "10300:10300"
      - "10800:10800"
    volumes:
      - ./data/node1:/opt/ignite/work

  node2:
    <<: *ignite-def
    command: --node-name node2
    ports:
      - "10301:10300"
      - "10801:10800"
    volumes:
      - ./data/node2:/opt/ignite/work

  node3:
    <<: *ignite-def
    command: --node-name node3
    ports:
      - "10302:10300"
      - "10802:10800"
    volumes:
      - ./data/node3:/opt/ignite/work

  prometheus:
    image: prom/prometheus:latest
    container_name: prometheus
    command:
      - '--config.file=/etc/prometheus/prometheus.yml'
      - '--web.enable-otlp-receiver'
    ports:
      - "9090:9090"
    configs:
      - source: prometheus_config
        target: /etc/prometheus/prometheus.yml

  grafana:
    image: grafana/grafana:latest
    container_name: grafana
    ports:
      - "3000:3000"
    environment:
      - GF_SECURITY_ADMIN_PASSWORD=admin
    depends_on:
      - prometheus

configs:
  node_config:
    content: |
      ignite {
        network {
          port: 3344
          nodeFinder.netClusterNodes = ["node1:3344", "node2:3344", "node3:3344"]
        }
        storage {
          profiles: [
            {
              name: "rocksDbProfile"
              engine: "rocksdb"
            }
          ]
        }
        metrics {
          sources {
            jvm { enabled: true }
            os { enabled: true }
            topology.cluster { enabled: true }
            topology.local { enabled: true }
          }
        }
      }

  prometheus_config:
    content: |
      global:
        scrape_interval: 15s

      scrape_configs:
        - job_name: 'prometheus'
          static_configs:
            - targets: ['localhost:9090']

Start the service stack:

docker compose up -d

Step 2: Define a CLI Helper Function

Since our Ignite nodes are running inside Docker, using localhost with a local CLI won't route correctly to the container network. Define a local bash alias function that executes ignite-cli inside the project's Docker network:

ignite-cli() {
  docker run --rm -it --network=ignite3_default \
    -e JAVA_TOOL_OPTIONS="--add-opens=java.base/java.lang=ALL-UNNAMED --add-opens=java.base/java.util=ALL-UNNAMED" \
    apacheignite/ignite:3.1.0 cli "$@"
}

Step 3: Initialize the Ignite 3 Cluster

Next, initialize the metastorage group across all three active nodes using node1 as the gateway URL:

ignite-cli cluster init --url http://node1:10300 --name my-cluster \
    --metastorage-group node1,node2,node3

Upon success, you will see a confirmation message: Cluster was initialized successfully.

Step 4: Configure the OTLP Exporter Endpoint

Now, configure the cluster's global exporter to push metric payloads over HTTP/Protobuf to the Prometheus OTLP endpoint:

ignite-cli cluster config update --url http://node1:10300 \
  "ignite.metrics.exporters.otlpExporter={exporterName:otlp, \
   endpoint:\"http://prometheus:9090/api/v1/otlp/v1/metrics\", \
   protocol:http/protobuf}"

To confirm the configuration was applied:

ignite-cli cluster config show --url http://node1:10300 ignite.metrics.exporters

Step 5: Configure Prometheus Data Source in Grafana

  1. Open Grafana in your web browser: http://localhost:3000
  2. Log in using the default credentials (admin / admin).
  3. In the left navigation menu, go to Connections ➡ Data sources ️️➡ Add data source.
  4. Select Prometheus.
  5. Set the Prometheus server URL to:
http://prometheus:9090

Click Save & Test. You should see the message "Data source is working".

None

Step 6: Create Your First "Total Cluster Nodes" Dashboard Panel

In Grafana, navigate to Dashboards ➡ New ➡ New Dashboard.

Click + Add visualization and select your newly added Prometheus data source.

Set Title to Total Active Cluster Nodes.

Switch to the Code query view and enter either of the following PromQL expressions TotalNodes and click on "Run queries".

None

First result:

None

I'm going to change the style to Gauge and refine it.

None

Step 7: Test Node Disconnection & Rebalancing

Now, let me demonstrate what happens when a node abruptly drops out of the cluster.

Open a separate terminal on your host machine and stop node3:

docker compose stop node3
None

node1 and node2 report 2 total nodes because they are active members of the updated cluster view. node3 reports 3 because, after being dropped, it is isolated and continues pushing its last recorded (stale) topology snapshot to Prometheus.

Step 8: Test Node Recovery & Scale-Up

Now start node3 back up:

docker compose start node3
None

After restarting node3, the cluster topology automatically heals. node1 and node2 detect node3 rejoining, causing all three metrics to converge back to 3 active nodes.

Step 9: Complete Clean Teardown

To stop the cluster, clean up Docker networks, and wipe host-mounted data directories for a fresh start:

# Stop containers and remove internal network/volumes
docker-compose down -v

# Clear host bind-mount data
rm -rf ./data

Exploring Available Metrics

To discover the complete list of exported metrics and metric sources available in Apache Ignite 3 — such as transactions (transactions), SQL execution (sql.queries), client connectivity (client.handler), and underlying page memory/storage engines—refer to the official Apache Ignite 3 Metrics List Documentation.

Conclusion

Integrating OpenTelemetry directly into Apache Ignite 3 drastically simplifies monitoring distributed databases. Rather than maintaining heavy agent wrappers or JMX sidecars, Ignite 3 pushes metrics directly over standard OTLP protocols to Prometheus and Grafana.

By observing how topology metrics respond dynamically to node state changes (docker compose stop/start), you can easily build robust alerting rules—such as triggering alerts when active cluster size drops below expected thresholds.

Futher Readings

📣 Call to Action

If you are interested in following along with my journey, I invite you to dive into all the details provided below:

Thanks for reading