Skip to content

ScriptEngine.eval(Reader)/compile(Reader) fails on scripts over ~97,800 characters -- is the fixed 100,000-char retry budget in parseExpressionOrModule intentional? #459

Description

@blwfish

Environment

  • jython-standalone-2.7.4.jar
  • OpenJDK Corretto 21.0.9 (also observed inside JMRI's bundled Java 21 runtime)
  • macOS 15.7.9 (aarch64)

Question

ParserFacade.parseExpressionOrModule() (used by the JSR223 python engine's
eval(Reader)/compile(Reader)) always attempts to parse the input as a bare
expression first, and only falls back to parsing it as a module if that fails
— which it always will for any real script. The retry relies on
reader.mark(MARK_LIMIT) (100,000, a fixed constant) taken before the first
(doomed) expression attempt. If that failed attempt reads past 100,000
characters before giving up, the mark is invalidated and the module-parse
retry's reset() throws IOException: Mark invalid — the whole
compile/eval fails, with no indication that size was the issue.

Was 100,000 chosen as "plenty for any script anyone would run through
eval()", or is there a reason the retry can't just re-read from the
original source (or size the mark to the input) instead of relying on a
fixed budget? I ran into this via JMRI, which loads Jython startup scripts
through this exact API, and wanted to check whether this is expected/known
before treating it as a bug on my end.

What I found (verified against source, not just bytecode)

org.python.core.ParserFacade:

private static int MARK_LIMIT = 100000;
public static mod parseExpressionOrModule(Reader reader, String filename, CompilerFlags cflags) {
    ExpectedEncodingBufferedReader bufReader = null;
    try {
        bufReader = prepBufReader(reader, cflags, filename);
        // first, try parsing as an expression
        return parse(bufReader, CompileMode.eval, filename, cflags);
    } catch (Throwable t) {
        if (bufReader == null) {
            throw Py.JavaError(t);
        }
        try {
            // then, try parsing as a module
            bufReader.reset();   // <-- throws here once the failed expression parse read >100,000 chars
            return parse(bufReader, CompileMode.exec, filename, cflags);
        } catch (Throwable tt) {
            throw fixParseError(bufReader, tt, filename);
        }
    }
}

private static mod parse(ExpectedEncodingBufferedReader reader, CompileMode kind, String filename, CompilerFlags cflags) throws Throwable {
    reader.mark(MARK_LIMIT); // We need the ability to move back on the
                             // reader, for the benefit of fixParseError and
                             // validPartialSentence
    ...
}

I initially assumed this was related to the PEP-263 encoding-declaration
sniff (prepBufReader/findEncoding also use MARK_LIMIT), but checked and
ruled that out: findEncoding() only reads the first two lines before giving
up, matching PEP 263 exactly, and isn't the culprit. The actual failure is in
the expression-then-module retry above — confirmed directly by the exception
stack trace, which names ParserFacade.parseExpressionOrModule(ParserFacade.java:140) (the bufReader.reset() line) as the frame that throws:

java.io.IOException: Mark invalid
	at java.base/java.io.BufferedReader.reset(BufferedReader.java:518)
	at org.python.core.ParserFacade.parseExpressionOrModule(ParserFacade.java:140)
	at org.python.util.PythonInterpreter.compile(PythonInterpreter.java:321)
	at org.python.util.PythonInterpreter.compile(PythonInterpreter.java:313)
	at org.python.jsr223.PyScriptEngine.compileScript(PyScriptEngine.java:101)
	at org.python.jsr223.PyScriptEngine.compile(PyScriptEngine.java:80)

Since every real script (anything with a statement, not just a bare
expression) is parsed as an expression first and is guaranteed to fail that
attempt, this isn't really about script size in the abstract — it's that
the doomed first parse attempt can read arbitrarily far into a large file
before the grammar rules it out as a valid expression, and once that read
exceeds 100,000 characters, the mark meant to support the retry is gone.

Minimal reproduction

import javax.script.*;
import java.io.*;

public class TestJythonEval {
    public static void main(String[] args) throws Exception {
        ScriptEngineManager mgr = new ScriptEngineManager();
        ScriptEngine engine = mgr.getEngineByName("python");
        try (Reader r = new BufferedReader(new FileReader(args[0]))) {
            ((Compilable) engine).compile(r);
            System.out.println("COMPILE OK");
        } catch (ScriptException e) {
            System.out.println("SCRIPT EXCEPTION: " + e);
        }
    }
}
javac -cp jython-standalone-2.7.4.jar TestJythonEval.java

# Any real script over ~100KB reproduces this; a plain comment-only file works too:
python3 -c "open('big.py','w').write(('# ' + 'x'*76 + chr(10)) * 1300)"

java -cp .:jython-standalone-2.7.4.jar TestJythonEval big.py
# -> SCRIPT EXCEPTION: javax.script.ScriptException: java.io.IOException: java.io.IOException: Mark invalid

I bisected the exact threshold on one specific real-world script and found
the boundary a few thousand characters short of the literal 100000
(compiles at 97,801 characters of that file, fails at 97,802) — consistent
with the mark being consumed partway through the doomed expression-parse
attempt rather than at some clean, predictable byte count, which tracks with
the mechanism above (how far the failed expression-parse gets before giving
up depends on the input's actual content, not just its length).

One more data point

The same content loads without any error if it's imported as a module
from a small top-level script instead of being passed to
ScriptEngine.eval/compile directly — that path goes through
org.python.core.imp's own bytecode-compiling loader, never attempts the
expression-first parse at all, and isn't affected regardless of the
imported module's size.

Ask

If the fixed 100,000-character retry budget is intentional: is there a
config knob to raise it, and would a clearer failure (naming the actual
cause) be reasonable instead of a bare IOException: Mark invalid? If not:
would it make sense to size the mark relative to the input's own length
(when that's knowable), or to skip the expression-parse attempt entirely for
input that's obviously multi-statement (e.g. contains a newline followed by
another statement) rather than always trying it first?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions