Skip to content

[Bug] BatchProducer benchmark counts a failed batch twice and mixes failed-send latency into average success RT #11280

Description

@chennaji9

Description

The batch benchmark producer example/src/main/java/org/apache/rocketmq/example/benchmark/BatchProducer.java reports incorrect statistics in two ways.

1. One interrupted send is counted twice

In the catch (InterruptedException e) block of the send thread, the failure counters are incremented twice — once before the Thread.sleep(3000) and once after:

} catch (InterruptedException e) {
    statsBenchmark.getSendRequestFailedCount().increment();
    statsBenchmark.getSendMessageFailedCount().add(msgs.size());
    try {
        Thread.sleep(3000);
    } catch (InterruptedException e1) {
    }
    statsBenchmark.getSendRequestFailedCount().increment();
    statsBenchmark.getSendMessageFailedCount().add(msgs.size());
    logger.error("[BENCHMARK_PRODUCER] Send Exception", e);
}

A single failed batch of size N is therefore reported as 2 failed requests and 2N failed messages, which doubles the failure rate shown by the benchmark report.

2. Failed requests inflate the average success RT

After the send result is obtained, the elapsed time is always added to the success-time accumulator, even when the status is not SEND_OK:

if (sendResult.getSendStatus() == SendStatus.SEND_OK) {
    statsBenchmark.getSendRequestSuccessCount().increment();
    statsBenchmark.getSendMessageSuccessCount().add(msgs.size());
} else {
    statsBenchmark.getSendRequestFailedCount().increment();
    statsBenchmark.getSendMessageFailedCount().add(msgs.size());
}
long currentRT = System.currentTimeMillis() - beginTimestamp;
statsBenchmark.getSendMessageSuccessTimeTotal().add(currentRT);

printStats divides sendMessageSuccessTimeTotal by the success counters (averageRT = (end[5] - begin[5]) / (end[1] - begin[1])), so failed sends' latency pollutes the reported average success RT. For example one success of 10 ms plus one failure of 90 ms reports an average success RT of 100 ms.

Expected behavior

  • An interrupted failed batch is counted exactly once.
  • sendMessageSuccessTimeTotal only accumulates the latency of SEND_OK sends so the reported average success RT reflects successful requests only.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions