Frequent 'database is locked error' 500 on lease update
Summary
Seeing as below whereby a small number of lease updates in k8s results in a large number of duplicate update queries in k8s-dqlite.
This issue is observed in v1.32 and v1.33 microk8s, not seen in a previous version used v1.29 so suspect some change has introduced the behaviour.
Example below 1 lease update for kube-controller-manager
kubectl get leases --all-namespaces --watch | grep kube-controller-manager
2025-07-21 15:04:50.475 pods kube-system kube-controller-manager dailybuild4-host1_53ce2bb9-14e4-46af-83c0-bde7007e05c3 9hDebug logging turned up in '/var/snap/microk8s/current/args/k8s-dqlite' this results in 327 of these update queries in one second. For some reason that same update is getting fired many times internally.
labuser@dailybuild4-host1:~$ sudo journalctl -u snap.microk8s.daemon-k8s-dqlite | grep '15:04:50' | grep kube-controller-manager | grep INSERT | wc -l 327
Jul 21 15:04:50 dailybuild4-host1 microk8s.daemon-k8s-dqlite[971627]: time="2025-07-21T15:04:50+01:00" level=debug msg="EXEC (try: 499) [/registry/leases/kube-system/kube-controller-manager 0 [107 56 115 0 10 31 10 22 99 111 111 114 100 105 110 97 116 105 111 110 46 107 56 115 46 105 111 47 118 49 18 5 76 101 97 115 101 18 253 2 10 160 2 10 23 107 117 98 101 45 99 111 110 116 114 111 108 108 101 114 45 109 97 110 97 103 101 114 18 0 26 11 107 117 98 101 45 115 121 115 116 101 109 34 0 42 36 100 99 52 101 100 101 101 98 45 51 97 99 51 45 52 98 55 101 45 98 102 49 54 45 101 102 55 51 97 97 51 98 55 52 51 50 50 0 56 0 66 8 8 245 138 247 195 6 16 0 138 1 190 1 10 8 107 117 98 101 108 105 116 101 18 6 85 112 100 97 116 101 26 22 99 111 111 114 100 105 110 97 116 105 111 110 46 107 56 115 46 105 111 47 118 49 34 8 8 129 146 249 195 6 16 0 50 8 70 105 101 108 100 115 86 49 58 124 10 122 123 34 102 58 115 112 101 99 34 58 123 34 102 58 97 99 113 117 105 114 101 84 105 109 101 34 58 123 125 44 34 102 58 104 111 108 100 101 114 73 100 101 110 116 105 116 121 34 58 123 125 44 34 102 58 108 101 97 115 101 68 117 114 97 116 105 111 110 83 101 99 111 110 100 115 34 58 123 125 44 34 102 58 108 101 97 115 101 84 114 97 110 115 105 116 105 111 110 115 34 58 123 125 44 34 102 58 114 101 110 101 119 84 105 109 101 34 58 123 125 125 125 66 0 18 88 10 54 100 97 105 108 121 98 117 105 108 100 52 45 104 111 115 116 49 95 53 51 99 101 50 98 98 57 45 49 52 101 52 45 52 54 97 102 45 56 51 99 48 45 98 100 101 55 48 48 55 101 48 53 99 51 16 60 26 12 8 188 198 248 195 6 16 248 236 250 173 3 34 12 8 129 146 249 195 6 16 128 173 136 153 3 40 6 26 0 34 0] /registry/leases/kube-system/kube-controller-manager 202186] : INSERT INTO kine(name, created, deleted, create_revision, prev_revision, lease, value, old_value) SELECT ? AS name, 0 AS created, 0 AS deleted, CASE WHEN kine.created THEN id ELSE create_revision END AS create_revision, id AS prev_revision, ? AS lease, ? AS value, value AS old_value FROM kine WHERE id = (SELECT MAX(id) FROM kine WHERE name = ?) AND deleted = 0 AND id = ?"The result is the following error in syslog, and the action fails.
Jul 21 15:04:50 dailybuild4-host1 microk8s.daemon-kubelite[283598]: E0721 15:04:50.129980 283598 status.go:71] "Unhandled Error" err="apiserver received an error that is not an metav1.Status: &status.Error{s:(*status.Status)(0xc01fcd4378)}: rpc error: code = Unknown desc = exec (try: 500): database is locked"
Jul 21 15:04:50 dailybuild4-host1 microk8s.daemon-kubelite[283598]: E0721 15:04:50.132155 283598 leaderelection.go:429] Failed to update lock optimistically: rpc error: code = Unknown desc = exec (try: 500): database is locked, falling back to slow path
Jul 21 15:04:50 dailybuild4-host1 microk8s.daemon-k8s-dqlite[971627]: time="2025-07-21T15:04:50+01:00" level=error msg="failed to update key" error="exec (try: 500): database is locked"
Jul 21 15:04:50 dailybuild4-host1 microk8s.daemon-k8s-dqlite[971627]: time="2025-07-21T15:04:50+01:00" level=debug msg="UPDATE /registry/leases/kube-system/kube-controller-manager, value=425, rev=202186, lease=0 => rev=0, updated=false, err=exec (try: 500): database is locked"
Jul 21 15:04:50 dailybuild4-host1 microk8s.daemon-k8s-dqlite[971627]: time="2025-07-21T15:04:50+01:00" level=error msg="error in txn: exec (try: 500): database is locked"Seen on a single node microk8s and also 3 node cluster. Introducing disk i/o provokes the issue further, assuming due to queries taking longer to run and therefore more likely to hit a lock condition. Not particular to this one lease update event, happens for multiple different pods in the deployment. Function being called in k8s-dqlite looks to be 'updateSQL' in pkg/backend/sqlite/driver.go. Is there some situation in microk8s that this gets retried in quick succession or there is a very short timeout value for retry?
Happy to provide further logs if needed.
What Should Happen Instead?
No database is locked errors. This causes other k8s actions to fail.
Reproduction Steps
I can reproduce by introducing some disk load on the microk8s server eg by running
while true; do dd if=/dev/zero of=test1.img bs=40G count=1 oflag=dsync; done
Source: canonical/microk8s