Skip to content

Improve robustness of rake work_jobs #411

Description

@zlogic

First of all, many thanks on creating this awesome RSS reader! Just what I was looking for and was about to create myself, but Stringer is so much better than anything I wanted :)

I installed it into Heroku and noticed a problem with reliability of the Heroku-specific rake lazy_fetch approach:

While the bundle exec rake work_jobs is launched from config/unicorn.rb, it's never checked to be running and its delayed_job_pid is never used. It seems that Heroku kills long-running database connections either on purpose, or just randomly, which also causes an unhandled exception to kill the bundle exec rake work_jobs process (example log). While rake lazy_fetch continues to submit jobs, the worker process is dead and never restarted.

It seems the best way would be to run rake work_jobs from a Procfile. I've tried two other workarounds and can confirm they work:

  1. Restart the rake process in bash:
diff --git config/unicorn.rb config/unicorn.rb
--- config/unicorn.rb
+++ config/unicorn.rb
@@ -9,7 +9,7 @@ before_fork do |_server, _worker|
   # as there's no need for the master process to hold a connection
   ActiveRecord::Base.connection.disconnect! if defined?(ActiveRecord::Base)

-  @delayed_job_pid ||= spawn("bundle exec rake work_jobs")
+  @delayed_job_pid ||= spawn("while true; do bundle exec rake work_jobs; sleep 60; done")

   sleep 1
 end
  1. Catch exceptions in the work_jobs job and try to reconnect:
diff --git Rakefile Rakefile
--- Rakefile
+++ Rakefile
@@ -46,11 +46,25 @@ desc "Work the delayed_job queue."
 task :work_jobs do
   Delayed::Job.delete_all

-  3.times do
-    Delayed::Worker.new(
-      min_priority: ENV["MIN_PRIORITY"],
-      max_priority: ENV["MAX_PRIORITY"]
-    ).start
+  logger = Logger.new(STDOUT)
+  logger.level = Logger::DEBUG
+
+  loop do
+    begin
+      logger.info "(Re-)starting worker"
+      if !ActiveRecord::Base.connection.active?
+        logger.info "Restarting DB connection"
+        ActiveRecord::Base.connection.reconnect!
+      end
+      Delayed::Worker.new(
+        min_priority: ENV["MIN_PRIORITY"],
+        max_priority: ENV["MAX_PRIORITY"]
+      ).start
+    rescue => e
+      logger.error e.message
+      e.backtrace.each { |line| logger.error line }
+      sleep 60
+    end
   end
 end

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions