Technical Insight How-to Guide

How to Troubleshoot OpenStack–Ceph Integration Issues

Diagnose OpenStack–Ceph integration failures across Cinder, Nova and libvirt by isolating RBD authentication, pool access, connectivity and configuration issues.

Technical Context

Where this fits.

Knowledge area Openstack Section Operations Purpose How-to Guide

OpenStack–Ceph problems rarely stay within one component.

A failed Cinder volume may originate in Cinder configuration, Ceph authentication, network connectivity or the RBD backend. A Nova boot failure may involve Nova, libvirt, Ceph credentials or an existing volume.

The fastest way to troubleshoot these issues is to start with the failing OpenStack operation, identify which component actually returned the error, and then test the Ceph integration from the affected host using the same client identity and configuration.

Before you start

Service names, log locations, configuration paths and commands vary between OpenStack distributions and deployment methods. Adapt them to your environment and verify the applicable OpenStack and Ceph versions before making changes.

Where OpenStack–Ceph integration usually breaks

Most integration problems occur somewhere along this path:

OpenStack service → Ceph client configuration → authentication → network → Ceph cluster

Before changing configuration, establish:

  1. Which OpenStack operation failed?
  2. Which service actually reported the failure?
  3. Which Ceph client identity is involved?
  4. Which Ceph pool is being accessed?
  5. Can that identity access the pool directly from the affected host?
  6. Is the Ceph cluster itself healthy?

Start with Ceph cluster health

ceph status
ceph health detail

Inspect relevant pools and client capabilities:

ceph osd pool ls detail | grep -E "volumes|vms|images"
ceph auth get client.cinder
ceph auth get client.nova
ceph auth get client.glance

If the cluster is unhealthy or the expected pool is unavailable, resolve or understand that condition before changing OpenStack configuration.

Cinder: volume creation fails

Start with the cinder-volume logs around the failed request:

journalctl -u openstack-cinder-volume -n 100 --no-pager

If your distribution uses a different service name or containers, inspect the corresponding Cinder volume service logs.

Verify the Cinder Ceph client

ls -la /etc/ceph/ceph.client.cinder.keyring
ceph auth get client.cinder

Test RBD access from the Cinder host:

rbd --id cinder --conf /etc/ceph/ceph.conf -p volumes ls

--id cinder selects the Ceph identity client.cinder; it does not switch the operating-system user.

If direct RBD access succeeds, basic Ceph authentication, connectivity and pool access are working. Investigate the Cinder backend configuration and service environment. If it fails, continue on the Ceph side.

Nova: instance boot or volume attachment fails

Inspect compute logs on the affected host:

journalctl -u nova-compute -n 200 --no-pager
journalctl -u nova-compute -n 200 --no-pager | grep -Ei "error|libvirt|rbd|ceph"

Check the libvirt secret

virsh secret-list

Compare the UUID expected by the Nova/libvirt RBD configuration with the secret available on the compute host. Do not recreate or replace a secret before determining whether it is already used by running instances or other storage integrations.

Test RBD access from the compute host

rbd --id nova --conf /etc/ceph/ceph.conf -p vms ls

If this fails, investigate Ceph authentication, capabilities, configuration and network connectivity. If it succeeds but Nova still fails, focus on Nova and libvirt.

Test Ceph connectivity from the affected host

Inspect the monitor configuration:

grep -E "mon_host|mon host" /etc/ceph/ceph.conf

Modern Ceph deployments commonly use Messenger v2 on TCP port 3300; Messenger v1 uses TCP port 6789. Test the endpoints actually configured for your cluster:

nc -zv <monitor-address> 3300
nc -zv <monitor-address> 6789

A successful TCP connection proves network reachability, not authentication or pool access. Test the complete client path with RBD:

rbd --id nova -p vms ls
rbd --id cinder -p volumes ls

This distinguishes network reachability, authentication and authorization.

Check keyring permissions

ls -la /etc/ceph/ceph.client.nova.keyring
ls -la /etc/ceph/ceph.conf

Do not assume a particular owner or group across deployments. If Nova runs under a nova OS account, permissions might look similar to:

chown nova:nova /etc/ceph/ceph.client.nova.keyring
chmod 640 /etc/ceph/ceph.client.nova.keyring

Treat this as an example. Verify the effective service user, group and deployment model first.

Verify Ceph capabilities

ceph auth get client.cinder
ceph auth get client.nova
ceph auth get client.glance

Compare those capabilities with the pools each service must access. A valid key can authenticate successfully while lacking authorization for the target pool.

Diagnosing live migration failures

Live migration can involve:

Nova orchestration → libvirt → migration network → Ceph/RBD access → destination compute host

Inspect Nova and libvirt logs on source and destination:

journalctl -u nova-compute -n 200 --no-pager
journalctl -u libvirtd -n 100 --no-pager
journalctl -u libvirtd -n 100 --no-pager | grep -Ei "migration|error|rbd|ceph"

Do not assume a particular libvirt transport or migration port range. Check the configured transport, authentication, firewall rules and migration network. Verify that both compute nodes can access the relevant Ceph-backed resources.

Inspect the migration state

openstack server migration list <instance-id>

If a diagnosed migration must be aborted:

openstack server migration abort <instance-id> <migration-id>

Aborting a migration is a recovery action, not a substitute for identifying the underlying failure.

A practical troubleshooting matrix

SymptomStart withThen verify
Cinder volume creation failscinder-volume logsCeph client, pool access and backend configuration
Nova cannot boot or attach a volumenova-compute and libvirt logsSecret configuration and RBD access
RBD command fails from a compute nodeCeph client configurationAuthentication, capabilities, monitors and network
Live migration failsNova and libvirt logsDestination access, migration transport and Ceph access
Multiple storage operations failCeph healthMON, OSD, pool and cluster state

Useful Ceph diagnostic commands

ceph status
ceph health detail
ceph osd pool ls detail
ceph auth get client.cinder
ceph auth get client.nova
ceph auth get client.glance
rbd --id cinder -p volumes ls
rbd --id nova -p vms ls

Run client-specific tests from the host where the failing OpenStack service is running whenever possible.

Troubleshooting pattern

Start with the failed OpenStack operation. Identify the component returning the error. Then reproduce the relevant Ceph access from the affected host using the intended client identity.

This separates OpenStack service configuration, Ceph client configuration, authentication, pool authorization, network connectivity, Ceph cluster health and libvirt integration.

If the corresponding RBD operation fails directly from the affected host, investigate the Ceph side first. If it succeeds, move back up the stack toward the OpenStack service and its configuration.

Avoid changing keyrings, libvirt secrets or service configuration until you have identified which layer is actually failing.

Related Articles