Environment
jython-standalone-2.7.4.jar
- OpenJDK Corretto 21.0.9 (also observed inside JMRI's bundled Java 21 runtime)
- macOS 15.7.9 (aarch64)
Question
ParserFacade.parseExpressionOrModule() (used by the JSR223 python engine's
eval(Reader)/compile(Reader)) always attempts to parse the input as a bare
expression first, and only falls back to parsing it as a module if that fails
— which it always will for any real script. The retry relies on
reader.mark(MARK_LIMIT) (100,000, a fixed constant) taken before the first
(doomed) expression attempt. If that failed attempt reads past 100,000
characters before giving up, the mark is invalidated and the module-parse
retry's reset() throws IOException: Mark invalid — the whole
compile/eval fails, with no indication that size was the issue.
Was 100,000 chosen as "plenty for any script anyone would run through
eval()", or is there a reason the retry can't just re-read from the
original source (or size the mark to the input) instead of relying on a
fixed budget? I ran into this via JMRI, which loads Jython startup scripts
through this exact API, and wanted to check whether this is expected/known
before treating it as a bug on my end.
What I found (verified against source, not just bytecode)
org.python.core.ParserFacade:
private static int MARK_LIMIT = 100000;
public static mod parseExpressionOrModule(Reader reader, String filename, CompilerFlags cflags) {
ExpectedEncodingBufferedReader bufReader = null;
try {
bufReader = prepBufReader(reader, cflags, filename);
// first, try parsing as an expression
return parse(bufReader, CompileMode.eval, filename, cflags);
} catch (Throwable t) {
if (bufReader == null) {
throw Py.JavaError(t);
}
try {
// then, try parsing as a module
bufReader.reset(); // <-- throws here once the failed expression parse read >100,000 chars
return parse(bufReader, CompileMode.exec, filename, cflags);
} catch (Throwable tt) {
throw fixParseError(bufReader, tt, filename);
}
}
}
private static mod parse(ExpectedEncodingBufferedReader reader, CompileMode kind, String filename, CompilerFlags cflags) throws Throwable {
reader.mark(MARK_LIMIT); // We need the ability to move back on the
// reader, for the benefit of fixParseError and
// validPartialSentence
...
}
I initially assumed this was related to the PEP-263 encoding-declaration
sniff (prepBufReader/findEncoding also use MARK_LIMIT), but checked and
ruled that out: findEncoding() only reads the first two lines before giving
up, matching PEP 263 exactly, and isn't the culprit. The actual failure is in
the expression-then-module retry above — confirmed directly by the exception
stack trace, which names ParserFacade.parseExpressionOrModule(ParserFacade.java:140) (the bufReader.reset() line) as the frame that throws:
java.io.IOException: Mark invalid
at java.base/java.io.BufferedReader.reset(BufferedReader.java:518)
at org.python.core.ParserFacade.parseExpressionOrModule(ParserFacade.java:140)
at org.python.util.PythonInterpreter.compile(PythonInterpreter.java:321)
at org.python.util.PythonInterpreter.compile(PythonInterpreter.java:313)
at org.python.jsr223.PyScriptEngine.compileScript(PyScriptEngine.java:101)
at org.python.jsr223.PyScriptEngine.compile(PyScriptEngine.java:80)
Since every real script (anything with a statement, not just a bare
expression) is parsed as an expression first and is guaranteed to fail that
attempt, this isn't really about script size in the abstract — it's that
the doomed first parse attempt can read arbitrarily far into a large file
before the grammar rules it out as a valid expression, and once that read
exceeds 100,000 characters, the mark meant to support the retry is gone.
Minimal reproduction
import javax.script.*;
import java.io.*;
public class TestJythonEval {
public static void main(String[] args) throws Exception {
ScriptEngineManager mgr = new ScriptEngineManager();
ScriptEngine engine = mgr.getEngineByName("python");
try (Reader r = new BufferedReader(new FileReader(args[0]))) {
((Compilable) engine).compile(r);
System.out.println("COMPILE OK");
} catch (ScriptException e) {
System.out.println("SCRIPT EXCEPTION: " + e);
}
}
}
javac -cp jython-standalone-2.7.4.jar TestJythonEval.java
# Any real script over ~100KB reproduces this; a plain comment-only file works too:
python3 -c "open('big.py','w').write(('# ' + 'x'*76 + chr(10)) * 1300)"
java -cp .:jython-standalone-2.7.4.jar TestJythonEval big.py
# -> SCRIPT EXCEPTION: javax.script.ScriptException: java.io.IOException: java.io.IOException: Mark invalid
I bisected the exact threshold on one specific real-world script and found
the boundary a few thousand characters short of the literal 100000
(compiles at 97,801 characters of that file, fails at 97,802) — consistent
with the mark being consumed partway through the doomed expression-parse
attempt rather than at some clean, predictable byte count, which tracks with
the mechanism above (how far the failed expression-parse gets before giving
up depends on the input's actual content, not just its length).
One more data point
The same content loads without any error if it's imported as a module
from a small top-level script instead of being passed to
ScriptEngine.eval/compile directly — that path goes through
org.python.core.imp's own bytecode-compiling loader, never attempts the
expression-first parse at all, and isn't affected regardless of the
imported module's size.
Ask
If the fixed 100,000-character retry budget is intentional: is there a
config knob to raise it, and would a clearer failure (naming the actual
cause) be reasonable instead of a bare IOException: Mark invalid? If not:
would it make sense to size the mark relative to the input's own length
(when that's knowable), or to skip the expression-parse attempt entirely for
input that's obviously multi-statement (e.g. contains a newline followed by
another statement) rather than always trying it first?
Environment
jython-standalone-2.7.4.jarQuestion
ParserFacade.parseExpressionOrModule()(used by the JSR223pythonengine'seval(Reader)/compile(Reader)) always attempts to parse the input as a bareexpression first, and only falls back to parsing it as a module if that fails
— which it always will for any real script. The retry relies on
reader.mark(MARK_LIMIT)(100,000, a fixed constant) taken before the first(doomed) expression attempt. If that failed attempt reads past 100,000
characters before giving up, the mark is invalidated and the module-parse
retry's
reset()throwsIOException: Mark invalid— the wholecompile/eval fails, with no indication that size was the issue.
Was 100,000 chosen as "plenty for any script anyone would run through
eval()", or is there a reason the retry can't just re-read from theoriginal source (or size the mark to the input) instead of relying on a
fixed budget? I ran into this via JMRI, which loads Jython startup scripts
through this exact API, and wanted to check whether this is expected/known
before treating it as a bug on my end.
What I found (verified against source, not just bytecode)
org.python.core.ParserFacade:I initially assumed this was related to the PEP-263 encoding-declaration
sniff (
prepBufReader/findEncodingalso useMARK_LIMIT), but checked andruled that out:
findEncoding()only reads the first two lines before givingup, matching PEP 263 exactly, and isn't the culprit. The actual failure is in
the expression-then-module retry above — confirmed directly by the exception
stack trace, which names
ParserFacade.parseExpressionOrModule(ParserFacade.java:140)(thebufReader.reset()line) as the frame that throws:Since every real script (anything with a statement, not just a bare
expression) is parsed as an expression first and is guaranteed to fail that
attempt, this isn't really about script size in the abstract — it's that
the doomed first parse attempt can read arbitrarily far into a large file
before the grammar rules it out as a valid expression, and once that read
exceeds 100,000 characters, the mark meant to support the retry is gone.
Minimal reproduction
I bisected the exact threshold on one specific real-world script and found
the boundary a few thousand characters short of the literal
100000(compiles at 97,801 characters of that file, fails at 97,802) — consistent
with the mark being consumed partway through the doomed expression-parse
attempt rather than at some clean, predictable byte count, which tracks with
the mechanism above (how far the failed expression-parse gets before giving
up depends on the input's actual content, not just its length).
One more data point
The same content loads without any error if it's
imported as a modulefrom a small top-level script instead of being passed to
ScriptEngine.eval/compiledirectly — that path goes throughorg.python.core.imp's own bytecode-compiling loader, never attempts theexpression-first parse at all, and isn't affected regardless of the
imported module's size.
Ask
If the fixed 100,000-character retry budget is intentional: is there a
config knob to raise it, and would a clearer failure (naming the actual
cause) be reasonable instead of a bare
IOException: Mark invalid? If not:would it make sense to size the mark relative to the input's own length
(when that's knowable), or to skip the expression-parse attempt entirely for
input that's obviously multi-statement (e.g. contains a newline followed by
another statement) rather than always trying it first?