Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The reporter runs PDF2zh 1.9.11 on Windows 11/Python 3.12.11 and sees repeated OpenAI RateLimitError warnings while the GUI appears stuck; Google translation works. The supplied logs name translator.py and show exponential retry delays with a 100-attempt budget. Current OpenAITranslator.do_translate still has that budget, but TranslateConverter.receive_layout wraps its paragraph worker in an unlimited retry decorator, restarting exhausted translation attempts forever. The only reply attributes the rejection to the API account's limits; the response payload is absent, so neither insufficient quota nor a transient request limit is established.
Testing
OpenAI succeeds immediately in streaming and non-streaming modes: text and normal request options remain unchanged.
A genuine mocked openai.RateLimitError followed by success retries and returns the translation; construct errors with a synthetic HTTP 429 response and no credentials or external requests.
Persistent RateLimitError exhausts exactly the configured 100 application-level attempts with sleeping disabled and raises the original final exception, not tenacity.RetryError. Mock the SDK call directly so SDK-internal retries do not distort the expected count.
Feed a minimal nonempty text layout into real TranslateConverter.receive_layout with one worker and a real OpenAITranslator whose mocked client always fails: the combined path exits after the same 100 calls, never starts call 101, and does not emit translated output. Use a fail-fast sentinel on call 101 so the pre-fix regression cannot hang even under the unlimited outer decorator; do not mock receive_layout itself.
A non-rate-limit ordinary exception followed by success retains the converter's existing retry behavior; blank text/formula-only paragraphs remain skipped.
Run the focused translator and converter suites with existing installed dependencies (pytest test/test_translator.py test/test_converter.py), then the project's unit suite and Black check for changed files in the implementation lane. No live API quota, PDF downloads, or model downloads are required for the new tests.
Summary
Set reraise=True on the existing OpenAITranslator.do_translate Tenacity decorator so exhausting its existing 100-attempt budget propagates the original openai.RateLimitError instead of a RetryError wrapper. In pdf2zh/converter.py, import the exception and configure the existing paragraph worker retry predicate to exclude RateLimitError while retaining retries for other ordinary Exceptions; this caller change is essential to prevent resetting the translator's budget. Preserve the existing exponential backoff, streaming behavior, subclass inheritance, and attempt count, and introduce no new retry abstraction or test-only seam.
Fixes #1147