Troubleshooting
Most problems announce themselves in one of three places: the error a command prints, the instance’s state and console, or the daemon’s log. This guide starts from what you see and works back to why.
Where to look
| To learn | Run |
|---|---|
| Why an instance is not running, and how it last ended | dicer inspect NAME |
| What the guest said while booting and running | dicer logs NAME |
| Why a guest never booted at all | dicer logs --source hypervisor NAME |
| What happened to it, and when | dicer events --name NAME |
| What the daemon did | journalctl -u dicerd |
| What the command line sent and got back | dicer --debug … |
For more detail from the daemon, set log_level: debug in its
configuration and restart it; running
guests are not affected.
The command line cannot reach the daemon
Error: there is no socket at /run/dicer/dicer.sock; is dicerd running?The daemon is not running. Start it, and see why it stopped:
$ sudo systemctl start dicerd
$ journalctl -u dicerd -n 50
Error: connection error: desc = "transport: Error while dialing: dial unix /run/dicer/dicer.sock: connect: permission denied"Only root can use the socket: run the command with sudo.
For a remote daemon, check which one a command reaches with dicer info,
and see Remote access. A TLS error names what did not
match: usually the daemon’s certificate lacks the name or address you
connect to, which --tls-server-name or a new certificate fixes.
A start is refused
The error says why; nothing was started.
Error: 1 vCPU, 4 TiB is more than this host can give instances in total (16 vCPU, 30 GiB)The host has too little room left, counting what running instances hold. Stop something, give the instance less, or change the overcommit; see Capacity.
Error: 999 vCPUs is more than the host's 4 CPUsAn instance cannot have more vCPUs than the host has CPUs, whatever the overcommit.
Error: port 18080:80/tcp is already published by instance "web", which is runningAnother instance, or a process on the host, has the port. Publish another, or stop the other instance.
Error: volume "pgdata" is attached to instance "db", which is running, …A volume is used read-write by one running instance at a time; see Files and volumes.
Error: image "no-such-image:1" not found on docker.io (or it is private)The name or tag is wrong, or the registry wants a login; see Managing images.
Error: cannot pull image "…": no child with platform linux/amd64 in index …The image has no build for the host’s architecture.
An instance ends right after starting
dicer inspect says how it ended:
$ dicer inspect web
● web — docker.io/library/nginx:1.27
Active: failed, exited (1) 5 seconds ago
…
An exit code is the workload’s own: read what it printed with
dicer logs web. It is the same failure the image would have in a
container, such as a missing setting or a wrong command.
An instance that ended in a kernel panic, a reset, or with its hypervisor gone says so instead. The console’s last lines usually tell which:
$ dicer logs -n 30 web
A guest that powers itself off, or a workload that exits 0, ends Stopped, not Failed: that is the workload finishing, not failing. Give it a command that keeps running, or a restart policy.
An instance runs, but its workload does not
An instance whose command cannot be started, because the path is wrong or the image does not have it, stays Running, with nothing working. Its console says why:
$ dicer logs web
… level=ERROR msg="entrypoint start failed" err="fork/exec /usr/bin/app: no such file or directory"
The guest is still up, so look around in it, then fix the command:
$ dicer exec web ls -l /usr/bin
$ dicer stop web
$ dicer update web -- /usr/local/bin/app
$ dicer start web
An instance never boots
When dicer logs is empty or stops during the kernel’s boot messages, the
guest did not get far. The hypervisor’s own log says why:
$ dicer logs --source hypervisor web
Look for:
- A kernel that does not suit the hypervisor or the host. The kernel
must be built for the host’s architecture and for a virtual machine;
Dicer’s kernel works under both
hypervisors. A kernel without EROFS boots, but the console then says
mount /dev/vda: no such device. See Kernels. - Kernel arguments that were replaced.
--kernel-argsreplaces the defaults whole, console and panic handling included. /dev/kvmmissing or not usable, in which case no instance boots: check that the host supports virtualisation and that it is enabled.
The hypervisor’s log is kept only while the instance runs, and is lost when it stops.
The network does not work
The guest cannot reach the outside world.
Starting an instance fails with
IPv4 forwarding is not enabledwhen the host does not forward.install.shturns it on; to do it yourself:$ sudo sysctl -w net.ipv4.ip_forward=1 $ echo net.ipv4.ip_forward=1 | sudo tee /etc/sysctl.d/99-dicer.confTraffic leaves through the interface of the host’s default route. On a host with several, set
network.uplink_interfacein the configuration.Names are resolved with
8.8.8.8unless the network says otherwise. If your network blocks it, create the network with--nameservers.
A published port cannot be reached.
- From the host itself,
localhostdoes not reach published ports; use the host’s own address, or the guest’s. - Check that the workload listens on all of the guest’s addresses, not only
its loopback:
dicer exec web ss -ltn. - Check the host’s firewall lets the port in.
Two instances cannot reach each other. Instances on different networks
cannot, by design, and neither can instances on the same network created
with --isolated. See Networking.
dicer exec, cp or health checks fail
These go through the guest’s agent. instance "web" is paused, not running
means just that: resume it first. An UNAVAILABLE error means the agent
did not answer:
- The guest may still be booting: try again in a moment.
- In a guest that boots systemd, the agent is a unit,
dicer-agent.service, and starts once systemd has; a guest stuck early in its boot has none.dicer logsshows how far it got.
The daemon misbehaves
Restarting it is safe: running guests keep running, and the new daemon takes them over.
$ sudo systemctl restart dicerd
$ journalctl -u dicerd -n 100
If the daemon will not start, its log says why. The usual cause is an error
in /etc/dicerd/config.yaml, which the log names with its line.
Reporting a problem
When asking for help, include the output of dicer version and
dicer info, the commands you ran and what they printed, and the relevant
part of dicer logs, dicer logs --source hypervisor and
journalctl -u dicerd.