OpenStack–Ceph problems rarely stay within one component.
A failed Cinder volume may originate in Cinder configuration, Ceph authentication, network connectivity or the RBD backend. A Nova boot failure may involve Nova, libvirt, Ceph credentials or an existing volume.
The fastest way to troubleshoot these issues is to start with the failing OpenStack operation, identify which component actually returned the error, and then test the Ceph integration from the affected host using the same client identity and configuration.
Before you start
Service names, log locations, configuration paths and commands vary between OpenStack distributions and deployment methods. Adapt them to your environment and verify the applicable OpenStack and Ceph versions before making changes.
Where OpenStack–Ceph integration usually breaks
Most integration problems occur somewhere along this path:
OpenStack service → Ceph client configuration → authentication → network → Ceph cluster
Before changing configuration, establish:
- Which OpenStack operation failed?
- Which service actually reported the failure?
- Which Ceph client identity is involved?
- Which Ceph pool is being accessed?
- Can that identity access the pool directly from the affected host?
- Is the Ceph cluster itself healthy?
Start with Ceph cluster health
ceph status
ceph health detailInspect relevant pools and client capabilities:
ceph osd pool ls detail | grep -E "volumes|vms|images"
ceph auth get client.cinder
ceph auth get client.nova
ceph auth get client.glanceIf the cluster is unhealthy or the expected pool is unavailable, resolve or understand that condition before changing OpenStack configuration.
Cinder: volume creation fails
Start with the cinder-volume logs around the failed request:
journalctl -u openstack-cinder-volume -n 100 --no-pagerIf your distribution uses a different service name or containers, inspect the corresponding Cinder volume service logs.
Verify the Cinder Ceph client
ls -la /etc/ceph/ceph.client.cinder.keyring
ceph auth get client.cinderTest RBD access from the Cinder host:
rbd --id cinder --conf /etc/ceph/ceph.conf -p volumes ls--id cinder selects the Ceph identity client.cinder; it does not switch the operating-system user.
If direct RBD access succeeds, basic Ceph authentication, connectivity and pool access are working. Investigate the Cinder backend configuration and service environment. If it fails, continue on the Ceph side.
Nova: instance boot or volume attachment fails
Inspect compute logs on the affected host:
journalctl -u nova-compute -n 200 --no-pager
journalctl -u nova-compute -n 200 --no-pager | grep -Ei "error|libvirt|rbd|ceph"Check the libvirt secret
virsh secret-listCompare the UUID expected by the Nova/libvirt RBD configuration with the secret available on the compute host. Do not recreate or replace a secret before determining whether it is already used by running instances or other storage integrations.
Test RBD access from the compute host
rbd --id nova --conf /etc/ceph/ceph.conf -p vms lsIf this fails, investigate Ceph authentication, capabilities, configuration and network connectivity. If it succeeds but Nova still fails, focus on Nova and libvirt.
Test Ceph connectivity from the affected host
Inspect the monitor configuration:
grep -E "mon_host|mon host" /etc/ceph/ceph.confModern Ceph deployments commonly use Messenger v2 on TCP port 3300; Messenger v1 uses TCP port 6789. Test the endpoints actually configured for your cluster:
nc -zv <monitor-address> 3300
nc -zv <monitor-address> 6789A successful TCP connection proves network reachability, not authentication or pool access. Test the complete client path with RBD:
rbd --id nova -p vms ls
rbd --id cinder -p volumes lsThis distinguishes network reachability, authentication and authorization.
Check keyring permissions
ls -la /etc/ceph/ceph.client.nova.keyring
ls -la /etc/ceph/ceph.confDo not assume a particular owner or group across deployments. If Nova runs under a nova OS account, permissions might look similar to:
chown nova:nova /etc/ceph/ceph.client.nova.keyring
chmod 640 /etc/ceph/ceph.client.nova.keyringTreat this as an example. Verify the effective service user, group and deployment model first.
Verify Ceph capabilities
ceph auth get client.cinder
ceph auth get client.nova
ceph auth get client.glanceCompare those capabilities with the pools each service must access. A valid key can authenticate successfully while lacking authorization for the target pool.
Diagnosing live migration failures
Live migration can involve:
Nova orchestration → libvirt → migration network → Ceph/RBD access → destination compute host
Inspect Nova and libvirt logs on source and destination:
journalctl -u nova-compute -n 200 --no-pager
journalctl -u libvirtd -n 100 --no-pager
journalctl -u libvirtd -n 100 --no-pager | grep -Ei "migration|error|rbd|ceph"Do not assume a particular libvirt transport or migration port range. Check the configured transport, authentication, firewall rules and migration network. Verify that both compute nodes can access the relevant Ceph-backed resources.
Inspect the migration state
openstack server migration list <instance-id>If a diagnosed migration must be aborted:
openstack server migration abort <instance-id> <migration-id>Aborting a migration is a recovery action, not a substitute for identifying the underlying failure.
A practical troubleshooting matrix
| Symptom | Start with | Then verify |
|---|---|---|
| Cinder volume creation fails | cinder-volume logs | Ceph client, pool access and backend configuration |
| Nova cannot boot or attach a volume | nova-compute and libvirt logs | Secret configuration and RBD access |
| RBD command fails from a compute node | Ceph client configuration | Authentication, capabilities, monitors and network |
| Live migration fails | Nova and libvirt logs | Destination access, migration transport and Ceph access |
| Multiple storage operations fail | Ceph health | MON, OSD, pool and cluster state |
Useful Ceph diagnostic commands
ceph status
ceph health detail
ceph osd pool ls detail
ceph auth get client.cinder
ceph auth get client.nova
ceph auth get client.glance
rbd --id cinder -p volumes ls
rbd --id nova -p vms lsRun client-specific tests from the host where the failing OpenStack service is running whenever possible.
Troubleshooting pattern
Start with the failed OpenStack operation. Identify the component returning the error. Then reproduce the relevant Ceph access from the affected host using the intended client identity.
This separates OpenStack service configuration, Ceph client configuration, authentication, pool authorization, network connectivity, Ceph cluster health and libvirt integration.
If the corresponding RBD operation fails directly from the affected host, investigate the Ceph side first. If it succeeds, move back up the stack toward the OpenStack service and its configuration.
Avoid changing keyrings, libvirt secrets or service configuration until you have identified which layer is actually failing.