Skip to content

Sync loop guard doesn't mark the attachment as failed, causing endless requeue and server overload #1301

Description

@nickstewart95

Bug Description

The Cloudinary WordPress plugin can get stuck repeatedly processing the same media item in the background sync queue.

process_assets() in php/sync/class-push-sync.php has a guard against attachments that keep landing on the same sync step. If that happens twice for the same attachment, it gives up on that attachment for the rest of the request:

if ( isset( $stat[ $attachment_id ][ $type ] ) ) {
    // Loop prevention.
    break;
}

The problem is that the attachment is not marked as failed when this happens. After the break, Utils::clean_up_sync_meta( $attachment_id ) removes the queued/thread metadata, but no _cld_error is stored. If the attachment’s _cloudinary metadata/signature was not saved, or if a stale cached value is returned, the next queue rebuild sees the attachment as unsynced, unqueued, and not errored. It then adds the same attachment back to the queue.

My site is hosted on Cloudways and uses Breeze and Object Cache Pro. Object Cache Pro or server/object caching may be relevant because this issue appears to depend on Cloudinary processing an attachment and then reading sync metadata/signature state that has not advanced. I'm not totally sure that's the cause, but it seemed worth noting.

This all results in a good number of internal REST requests to the Cloudinary queue endpoint and severe CPU/load impact, taking my site down.

Expected Behaviour

When Cloudinary detects that the same attachment is reaching the same sync step more than once in a single processing request, it should treat that attachment as failed or temporarily blocked instead of allowing it to be requeued immediately.

Expected behavior:

  • The affected attachment should be marked with a sync error, such as _cld_error, when the loop guard is triggered.
  • Queue cleanup should not leave the attachment in a state where it is unsynced, unqueued, and not errored.
  • The next queue rebuild should exclude the affected attachment until the sync error is manually cleared or the asset is otherwise repaired.
  • Background queue processing should stop once no valid syncable work remains.
  • The plugin should not repeatedly send internal requests to /wp-json/cloudinary/v1/queue for the same attachment when its sync state is not advancing.

Steps to reproduce

I'm not totally sure how to cause Step 3 in a live environment, outside of its happening on my site. I did provide a test filter to recreate it however.

  1. Install and connect the Cloudinary WordPress plugin.
  2. Enable Cloudinary media sync or auto sync.
  3. Add the test filter shown below. This prevents Cloudinary’s _cloudinary post meta from being saved for one specific attachment and simulates the condition where Cloudinary processes an attachment, but the attachment’s sync metadata/signature does not advance.
  4. Upload a normal JPG to the WordPress Media Library.
  5. Set the media item title to cld-loop-test.
  6. Start a Cloudinary bulk sync, or wait for autosync to process the attachment.
  7. The Cloudinary sync step runs, but the attachment’s _cloudinary metadata/signature is not saved because of the test hook.
  8. Cloudinary\Sync\Push_Sync::process_assets() asks for the next sync step and receives the same sync step again for the same attachment.
  9. The existing loop guard exits the current processing pass, but no _cld_error is stored.
  10. Utils::clean_up_sync_meta() removes the queued/thread metadata.
  11. The next queue rebuild sees the attachment as unsynced, unqueued, and not errored, then adds it back to the queue.
  12. The plugin sends another internal request to /wp-json/cloudinary/v1/queue, and the same attachment is processed again.

Test filter:

add_filter(
    'update_post_metadata',
    function ( $check, $object_id, $meta_key ) {
        if ( '_cloudinary' !== $meta_key ) {
            return $check;
        }

        $post = get_post( $object_id );

        if ( $post && 'cld-loop-test' === $post->post_title ) {
            return true; // Pretend the update succeeded, but prevent the meta write.
        }

        return $check;
    },
    10,
    3
);

Additional context

Cloudinary system report details:

  • WordPress version: 7.1.2
  • PHP version: 8.2.34
  • Cloudinary plugin version: 3.3.7
  • Cloudinary status: ok
  • Auto sync: on
  • Cloudinary folder: wordpress_assets/
  • Offload setting: dual_full
  • Image delivery: on
  • Video delivery: on
  • Image optimization: on
  • Video optimization: on
  • Responsive breakpoints: on
  • Cloudinary plugin cron system: off
    • update_asset_paths: off
    • rest_api: off
    • check_status: off
  • Cloudinary debug log: empty

Caching / hosting context:

  • Site is hosted on Cloudways.
  • Breeze 2.5.17 is active.
  • Object Cache Pro is active.
  • Redis PHP extension is loaded.

I also have an AI generated insight from the hosting provider, Cloudways, of when my site went down

Your WordPress application has been experiencing severe CPU overload caused by a runaway background process in the Cloudinary image management plugin. The investigation examined server load metrics, PHP processing engine status, database performance, web traffic patterns, and error logs over the past several hours.
Key findings: The server reached extreme load levels (39+ on a 4-core server, nearly 10 times normal capacity) with 96% of web requests failing with timeout errors. The root cause is the Cloudinary plugin stuck in an infinite processing loop, continuously attempting to sync your media library to Cloudinary's CDN. Your server is making hundreds of internal requests per hour to its own REST API endpoint trying to process the Cloudinary upload queue, but these requests never complete successfully.
Each Cloudinary queue processing attempt triggers slow database queries that take 25-50+ seconds to complete, primarily related to post metadata lookups. While these requests wait for the database, they occupy PHP processing slots. With hundreds of these stuck requests accumulating, your PHP processing engine reached its 40-worker limit, causing new visitor requests to time out. This created a cascading failure: timeouts cause retries, retries add more load, more load causes more timeouts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions