Is there an existing issue already for this bug?
I have read the troubleshooting guide
I am running a supported version of CloudNativePG
Contact Details
No response
Version
1.29 (latest patch)
What version of Kubernetes are you using?
1.35
What is your Kubernetes environment?
Cloud: Google GKE
How did you install the operator?
Helm
What happened?
Occasionally when we CNPG decided to shutdown primary we get:
Failed to execute CHECKPOINT command
error reported in logs.
Looking into it seems that #8867 added optimization aiming to speedup final checkpoint by issuing online CHECKPOINT command. As I understand idea was to leave for a final checkpoint less stuff to do, and thus speeding it up, by issuing early online checkpoint while DB is still serving requests in smart shutdown mode.
Implementation disables checkpointing for immediate shutdown, but decided to leave it on for the fast shutdown mode (that is smartShutdownDelay=0). I think there is no much to win for the fast shutdown with online checkpoint. Fast mode closes all in-flight queries and proceeds to checkpoint and then exits. There is no speedup to be had with this optimization if server doesn't stay up meaningful amount of time.
Cluster resource
Relevant log output
{
"level": "error",
"logger": "instance-manager",
"msg": "Failed to execute CHECKPOINT command",
"logging_pod": "XYZ",
"error": "failed to connect to `user=postgres database=postgres`: /controller/run/.s.PGSQL.5432 (/controller/run): server error: FATAL: the database system is shutting down (SQLSTATE 57P03)",
"stacktrace": "github.com/cloudnative-pg/machinery/pkg/log.(*logger).Error\n\tpkg/mod/github.com/cloudnative-pg/machinery@v0.5.0/pkg/log/log.go:128\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).tryCheckpointBeforeShutdown\n\tpkg/management/postgres/instance.go:1591\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).Shutdown\n\tpkg/management/postgres/instance.go:574\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).TryShuttingDownSmartFast\n\tpkg/management/postgres/instance.go:647\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).HandleInstanceCommandRequests\n\tpkg/management/postgres/instance.go:1547\ngithub.com/cloudnative-pg/cloudnative-pg/internal/cmd/manager/instance/run/lifecycle.(*PostgresLifecycle).Start\n\tinternal/cmd/manager/instance/run/lifecycle/lifecycle.go:157\nsigs.k8s.io/controller-runtime/pkg/manager.(*runnableGroup).reconcile.func1\n\tpkg/mod/sigs.k8s.io/controller-runtime@v0.24.1/pkg/manager/runnable_group.go:260"
}
Code of Conduct
Is there an existing issue already for this bug?
I have read the troubleshooting guide
I am running a supported version of CloudNativePG
Contact Details
No response
Version
1.29 (latest patch)
What version of Kubernetes are you using?
1.35
What is your Kubernetes environment?
Cloud: Google GKE
How did you install the operator?
Helm
What happened?
Occasionally when we CNPG decided to shutdown primary we get:
error reported in logs.
Looking into it seems that #8867 added optimization aiming to speedup final checkpoint by issuing online
CHECKPOINTcommand. As I understand idea was to leave for a final checkpoint less stuff to do, and thus speeding it up, by issuing early online checkpoint while DB is still serving requests in smart shutdown mode.Implementation disables checkpointing for immediate shutdown, but decided to leave it on for the fast shutdown mode (that is
smartShutdownDelay=0). I think there is no much to win for the fast shutdown with online checkpoint. Fast mode closes all in-flight queries and proceeds to checkpoint and then exits. There is no speedup to be had with this optimization if server doesn't stay up meaningful amount of time.Cluster resource
Relevant log output
{ "level": "error", "logger": "instance-manager", "msg": "Failed to execute CHECKPOINT command", "logging_pod": "XYZ", "error": "failed to connect to `user=postgres database=postgres`: /controller/run/.s.PGSQL.5432 (/controller/run): server error: FATAL: the database system is shutting down (SQLSTATE 57P03)", "stacktrace": "github.com/cloudnative-pg/machinery/pkg/log.(*logger).Error\n\tpkg/mod/github.com/cloudnative-pg/machinery@v0.5.0/pkg/log/log.go:128\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).tryCheckpointBeforeShutdown\n\tpkg/management/postgres/instance.go:1591\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).Shutdown\n\tpkg/management/postgres/instance.go:574\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).TryShuttingDownSmartFast\n\tpkg/management/postgres/instance.go:647\ngithub.com/cloudnative-pg/cloudnative-pg/pkg/management/postgres.(*Instance).HandleInstanceCommandRequests\n\tpkg/management/postgres/instance.go:1547\ngithub.com/cloudnative-pg/cloudnative-pg/internal/cmd/manager/instance/run/lifecycle.(*PostgresLifecycle).Start\n\tinternal/cmd/manager/instance/run/lifecycle/lifecycle.go:157\nsigs.k8s.io/controller-runtime/pkg/manager.(*runnableGroup).reconcile.func1\n\tpkg/mod/sigs.k8s.io/controller-runtime@v0.24.1/pkg/manager/runnable_group.go:260" }Code of Conduct