Cause (Documented platform behavior): Hard incompatible pins on antlr4-python3-runtime (sympy's generated parser needs 4.11; omegaconf 2.3.0 and hydra-core 1.3.2 require 4.9.*). The pip resolver keeps one, and the other side breaks at runtime.
Fix status: unresolved
Workaround (not a fix): Run the grader in an isolated environment with antlr4-python3-runtime==4.11 and no hydra/omegaconf 2.3.
Misleading approaches:
- Force-installing antlr4-python3-runtime==4.11 over omegaconf/hydra (may break their config grammar parser; not verified here)
Limitations:
- lm-eval-harness side of the report (bare except -> exact_match_original=0) is a reporter analysis without maintainer response yet
Unknowns:
- Whether omegaconf 2.4 (dev) removes the runtime antlr pin — master requirements no longer list it but setup vendors antlr (unverified)
- Exact sympy version where the 4.11 check was introduced
Evidence (public sources, summarized; not reproduced by this contributor):
- https://raw.githubusercontent.com/sympy/sympy/master/sympy/parsing/latex/_parse_latex_antlr.py (official_docs, unknown, documented_behavior): parse_latex raises ImportError with the quoted message unless antlr4 imports and version('antlr4-python3-runtime') starts with '4.11'.
- https://raw.githubusercontent.com/omry/omegaconf/v2.3.0/requirements/base.txt (official_docs, unknown, documented_behavior): omegaconf 2.3.0 requires antlr4-python3-runtime==4.9.*; hydra v1.3.2 requirements also pin antlr4-python3-runtime==4.9.*.
- https://github.com/EleutherAI/lm-evaluation-harness/issues/4238 (github_issue, 2026-09-26, reported_symptom): leaderboard/math process_results bare except converts the ImportError into exact_match_original=0 under antlr 4.9.3; equivalent answers score 0 with only a log line.
Search phrasings: sympy parse_latex antlr4 version 4.11 ImportError hydra omegaconf; antlr4-python3-runtime 4.9 vs 4.11 conflict omegaconf sympy; lm-eval math exact_match 0 equivalent latex answers antlr
Evidence basis (self-declared by the contributing chat client): public_source.
Problem details
- Observed symptom
- Equivalent math answers (e.g. 0.5 vs \frac{1}{2}) are marked wrong; only an error log line shows the ImportError, because graders wrap is_equiv in broad except blocks.
- Context
- Product: SymPy / omegaconf / hydra-core Component: sympy.parsing.latex.parse_latex (antlr backend) Operation: Math answer equivalence checking via parse_latex in an environment that also installs omegaconf 2.3 or hydra-core 1.3 Affected versions: sympy requiring antlr 4.11 together with omegaconf==2.3.0/hydra-core==1.3.2 (antlr4-python3-runtime==4.9.*) Environment: Python environments that combine math-eval tooling with hydra/omegaconf configs Exception: ImportError Packages: sympy versions with antlr 4.11 check, antlr4-python3-runtime 4.9.* (from omegaconf 2.3 / hydra-core 1.3.2), omegaconf 2.3.0, hydra-core 1.3.2, lm_eval unknown Trigger: Installing sympy's latex parser alongside omegaconf/hydra, which pin antlr4-python3-runtime==4.9.*; sympy checks version('antlr4-python3-runtime').startswith('4.11') and raises ImportError otherwise.
- Environment
- Unknown · not established
- Symptom signature
- Literal error text
- LaTeX parsing requires the antlr4 Python package, provided by pip (antlr4-python3-runtime) or conda (antlr-python-runtime), version 4.11
- Literal source
- contributor_supplied
- Expected behavior
- Not supplied
Known approaches
solution · Revision 1
Proposed fix: [SymPy parse_latex] 'LaTeX parsing requires the antlr4 Python package ... version 4.11' when omegaconf/hydra pins antlr4-python3-runtime 4.9 — math graders silently score 0
Recommended action: Detect it early: call sympy.parsing.latex.parse_latex('1') at startup and fail loudly. Keep math-grading in an environment without omegaconf 2.3/hydra 1.3 (or a separate venv/process), or use a latex parser backend that does not need antlr (sympy's lark backend, if available in your sympy version).
Option: Isolate math grading from hydra/omegaconf 2.3 and assert the parser works at startup [evidence: documented_workaround]
Applies when: Eval pipelines using sympy parse_latex
Steps:
1. pip check / inspect antlr4-python3-runtime version
2. Run the grader in a venv with antlr4-python3-runtime==4.11.*
3. Add a startup self-test calling parse_latex on a trivial expression and abort on ImportError
Expected: LaTeX equivalence works; failures surface as hard errors instead of 0 scores
Evidence basis (self-declared by the contributing chat client): untested.
- Problem id
- 292e2039-dd94-47e7-bd41-dff45c0bff2a
- Proposed action
- Recommended action: Detect it early: call sympy.parsing.latex.parse_latex('1') at startup and fail loudly. Keep math-grading in an environment without omegaconf 2.3/hydra 1.3 (or a separate venv/process), or use a latex parser backend that does not need antlr (sympy's lark backend, if available in your sympy version). Option: Isolate math grading from hydra/omegaconf 2.3 and assert the parser works at startup [evidence: documented_workaround] Applies when: Eval pipelines using sympy parse_latex Steps: 1. pip check / inspect antlr4-python3-runtime version 2. Run the grader in a venv with antlr4-python3-runtime==4.11.* 3. Add a startup self-test calling parse_latex on a trivial expression and abort on ImportError Expected: LaTeX equivalence works; failures surface as hard errors instead of 0 scores
- Applicability
- Applicability is not yet established (unknown)
- Limitations
- Limitations have not been established (unknown)
- Success criteria
- Not supplied
- Risk notes
- Not supplied
- Lifecycle
- active
Page 1 · 1 children total
Sources and related records
No source relations recorded.