Skip to content

Commit bb03b9f

Browse files
committed
Merge commit '53538f272978bf146839ce51cad410b7415f4e1d' into wb/t24
2 parents ae1a181 + 53538f2 commit bb03b9f

12 files changed

Lines changed: 349 additions & 18 deletions

File tree

‎.github/workflows/test.yml‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -22,6 +22,15 @@ jobs:
2222
poetry install --with dev --all-extras
2323
- name: Install Pandoc # apt version seems too old
2424
uses: r-lib/actions/setup-pandoc@v2
25+
# The validator compiles a set the way lambda-feedback/PDF-generator does, so it
26+
# needs what that image installs: xelatex with braket, cancel, xeCJK and the
27+
# Noto Sans fonts the template sets as its main and CJK fonts.
28+
- name: Install XeLaTeX
29+
run: |
30+
sudo apt-get update
31+
sudo apt-get install -y --no-install-recommends \
32+
texlive-xetex texlive-latex-recommended texlive-latex-extra \
33+
texlive-science texlive-lang-chinese fonts-noto-core fonts-noto-cjk
2534
- name: Linting Checks
2635
run: |
2736
poetry run black --check .

‎docs/source/quickstart.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -75,6 +75,8 @@ By default, this generates an `out` directory in the same place that the command
7575

7676
Before writing anything, in2lambda prints the problems it can detect that would stop the set importing or make it render wrongly — an answer that doesn't fit the box marking it, a figure the export won't contain, maths that KaTeX can't display. Each names the question, part and field to go and look at. They are warnings rather than errors: the `out` directory is written either way, since a problem found here may well be deliberate.
7777

78+
With [xelatex](https://tug.org/texlive/) installed alongside pandoc, the set is also compiled the way Lambda Feedback makes a PDF of it, and any LaTeX error names the field it is in. Without it, one warning says which packages to install instead.
79+
7880
Check the [command line tool reference](reference/command-line) for more information.
7981

8082
## 3. Import into Lambda Feedback

‎in2lambda/api/set.py‎

Lines changed: 7 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -103,12 +103,16 @@ def increment_current_question(self) -> None:
103103
"""
104104
self._current_question_index += 1
105105

106-
def problems(self) -> list[Problem]:
106+
def problems(self, compile: bool = True) -> list[Problem]:
107107
r"""Everything in2lambda can tell Lambda Feedback would refuse or render wrongly.
108108
109109
This is a report, not a refusal: the set can still be written out, since a
110110
problem found here may well be deliberate.
111111
112+
Args:
113+
compile: Whether to also compile the set as Lambda Feedback's PDF generator
114+
will, which needs pandoc and xelatex installed.
115+
112116
Returns:
113117
One :class:`~in2lambda.api.problem.Problem` per problem found, each naming
114118
the question, part and field to go and look at.
@@ -118,14 +122,14 @@ def problems(self) -> list[Problem]:
118122
>>> s = Set()
119123
>>> s.add_question("Momentum", "The rocket is at $45^\\circ$.")
120124
>>> s.current_question.images.append("no_such_file.png")
121-
>>> for problem in s.problems():
125+
>>> for problem in s.problems(compile=False):
122126
... print(problem)
123127
Question 1 "Momentum", main text: ^\circ does not display; write the degree sign ° instead
124128
Question 1 "Momentum": there is no image file at no_such_file.png
125129
"""
126130
from in2lambda.validation import validate
127131

128-
return validate(self)
132+
return validate(self, compile)
129133

130134
def to_json(self, output_dir: str) -> None:
131135
"""Turns this set into Lambda Feedback JSON/ZIP files.

‎in2lambda/validation/__init__.py‎

Lines changed: 35 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -18,6 +18,7 @@
1818
from in2lambda.api.response_area import ResponseArea
1919
from in2lambda.api.set import Set
2020
from in2lambda.katex_convert.katex_convert import unsupported_commands
21+
from in2lambda.validation import pdf
2122
from in2lambda.validation.delimiters import MathDelimiterError, math_delimiter_checker
2223

2324
__all__ = ["MathDelimiterError", "Problem", "math_delimiter_checker", "validate"]
@@ -34,11 +35,14 @@
3435
"""``^\\circ``, with or without braces around it."""
3536

3637

37-
def validate(question_set: Set) -> list[Problem]:
38+
def validate(question_set: Set, compile: bool = True) -> list[Problem]:
3839
r"""Everything in2lambda can tell is wrong with a set, in the order it is written.
3940
4041
Args:
4142
question_set: The set about to be exported.
43+
compile: Whether to also compile the set as Lambda Feedback's PDF generator
44+
will, which needs pandoc and xelatex - see
45+
:mod:`in2lambda.validation.pdf`.
4246
4347
Returns:
4448
One :class:`~in2lambda.api.problem.Problem` per problem found, each naming the
@@ -50,17 +54,27 @@ def validate(question_set: Set) -> list[Problem]:
5054
>>> from in2lambda.validation import validate
5155
>>> s = Set()
5256
>>> s.add_question("Angles", "Turn through $90^\\circ$.")
53-
>>> [str(problem) for problem in validate(s)]
57+
>>> [str(problem) for problem in validate(s, compile=False)]
5458
['Question 1 "Angles", main text: ^\\circ does not display; write the degree sign ° instead']
5559
"""
5660
problems: list[Problem] = []
61+
# Every markdown field with the location to report it against, kept so that the
62+
# whole set can then be compiled in one go rather than a field at a time.
63+
fields: list[tuple[str, str]] = []
64+
images: list[str] = []
65+
66+
def check(
67+
markdown: str, question: Question, location: str, compiled: bool = True
68+
) -> list[Problem]:
69+
if compiled:
70+
fields.append((location, markdown))
71+
return _markdown_problems(markdown, question, location)
5772

5873
for number, question in enumerate(question_set.questions, start=1):
5974
where = f'Question {number} "{question.title}"'
60-
problems += _markdown_problems(
61-
question.main_text, question, f"{where}, main text"
62-
)
75+
problems += check(question.main_text, question, f"{where}, main text")
6376

77+
images += question.images
6478
for image in question.images:
6579
if not Path(image).is_file():
6680
problems.append(Problem(where, f"there is no image file at {image}"))
@@ -72,30 +86,39 @@ def validate(question_set: Set) -> list[Problem]:
7286
("worked solution", part.worked_solution),
7387
("answer", part.answer),
7488
):
75-
problems += _markdown_problems(
76-
markdown, question, f"{part_where}, {field}"
77-
)
89+
problems += check(markdown, question, f"{part_where}, {field}")
7890

7991
for area_number, area in enumerate(part.response_areas, start=1):
8092
area_where = f"{part_where}, answer box {area_number}"
8193
problems += [
8294
Problem(area_where, message) for message in _area_problems(area)
8395
]
96+
# An answer box is only in the PDF if it is marked to be, so LaTeX it
97+
# would not compile cannot break one unless it is.
8498
for field, markdown in (
8599
("pre_text", area.pre_text),
86100
("post_text", area.post_text),
87101
("content_after", area.content_after),
88102
):
89-
problems += _markdown_problems(
90-
markdown, question, f"{area_where}, {field}"
103+
problems += check(
104+
markdown,
105+
question,
106+
f"{area_where}, {field}",
107+
area.include_in_pdf,
91108
)
92109
options = (area.config or {}).get("options")
93110
if isinstance(options, list):
94111
for option_number, option in enumerate(options, start=1):
95-
problems += _markdown_problems(
96-
option, question, f"{area_where}, option {option_number}"
112+
problems += check(
113+
option,
114+
question,
115+
f"{area_where}, option {option_number}",
116+
area.include_in_pdf,
97117
)
98118

119+
if compile:
120+
problems += pdf.problems(fields, images)
121+
99122
return problems
100123

101124

Lines changed: 199 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,199 @@
1+
"""Compiles a set the way Lambda Feedback makes a PDF of it, and reads the errors back.
2+
3+
Lambda Feedback renders question PDFs with lambda-feedback/PDF-generator: pandoc with
4+
``template.latex`` beside this file, then xelatex. Markdown that pipeline refuses is a
5+
fault in the set, so the whole set is compiled once here and each LaTeX error is traced
6+
back to the field it came from.
7+
8+
The trick for tracing is a marker: the document handed to pandoc carries a raw-LaTeX
9+
comment naming the field before each field's markdown, and pandoc copies raw blocks
10+
through untouched. ``xelatex -file-line-error`` then reports every error as
11+
``set.tex:<line>: <message>``, and the last marker above that line names the field.
12+
13+
pandoc and xelatex are both optional, as they are everywhere else in in2lambda: without
14+
them this reports what to install rather than raising.
15+
"""
16+
17+
import re
18+
import shutil
19+
import subprocess
20+
import tempfile
21+
from pathlib import Path
22+
23+
from in2lambda.api.problem import Problem
24+
25+
_TEMPLATE = Path(__file__).with_name("template.latex")
26+
"""The PDF generator's own pandoc template - see the README beside it."""
27+
28+
_TOOLS = {
29+
"pandoc": "pandoc (see https://pandoc.org/installing.html)",
30+
"xelatex": (
31+
"xelatex (apt install texlive-xetex texlive-latex-recommended"
32+
" texlive-latex-extra texlive-science texlive-lang-chinese"
33+
" fonts-noto-core fonts-noto-cjk)"
34+
),
35+
}
36+
37+
_SET = "The set"
38+
"""Where an error that is not inside any one field is reported against."""
39+
40+
_MARKER = "% in2lambda: "
41+
42+
_ERROR = re.compile(
43+
r"^(?:\./)?(\S+\.(?:tex|sty|cls|def|cfg|fd|ltx)):(\d+): (.+)$", re.MULTILINE
44+
)
45+
"""One ``-file-line-error`` line. The file is only ``set.tex`` for the set's own text."""
46+
47+
_IMAGE = re.compile(r"(!\[[^\]]*\]\()([^)]*)(\))")
48+
"""A markdown image with its path apart, so that the path can be rewritten or dropped."""
49+
50+
_TIMEOUT = 120
51+
"""Seconds for pandoc or xelatex. A set that takes longer is reported, not waited for."""
52+
53+
54+
def missing_tools() -> list[str]:
55+
"""What is needed to compile a set but is not installed, each saying how to get it.
56+
57+
Returns:
58+
One line per missing tool, or an empty list if a set can be compiled here.
59+
"""
60+
return [hint for tool, hint in _TOOLS.items() if shutil.which(tool) is None]
61+
62+
63+
def problems(fields: list[tuple[str, str]], images: list[str]) -> list[Problem]:
64+
"""Everything the PDF generator's pandoc and xelatex refuse, by the field it is in.
65+
66+
Args:
67+
fields: Every markdown field of the set in the order it is written, each with
68+
the location - ``Question 1 "Title", part (a), text`` - to report against.
69+
images: Every image path the set's questions hold. Those that exist are put
70+
beside the compiled document under their file name, since that is how the
71+
export refers to them; a reference to any other is dropped, as the image
72+
check already reports it.
73+
74+
Returns:
75+
One :class:`~in2lambda.api.problem.Problem` per distinct LaTeX error. An empty
76+
list means the set compiles as Lambda Feedback will compile it.
77+
"""
78+
if missing := missing_tools():
79+
return [
80+
Problem(
81+
_SET,
82+
"not compiled as the PDF generator would: install "
83+
+ " and ".join(missing),
84+
)
85+
]
86+
87+
try:
88+
return _compiled(fields, images)
89+
except subprocess.TimeoutExpired as expired:
90+
# TeX can be made to loop forever, which is itself a fault in the set.
91+
return [
92+
Problem(
93+
_SET,
94+
f"the PDF generator cannot compile this: {expired.cmd[0]} did not"
95+
f" finish within {_TIMEOUT} seconds",
96+
)
97+
]
98+
99+
100+
def _compiled(fields: list[tuple[str, str]], images: list[str]) -> list[Problem]:
101+
"""The set run through pandoc and then xelatex in a directory of its own."""
102+
with tempfile.TemporaryDirectory() as directory:
103+
work = Path(directory)
104+
available = set()
105+
for image in images:
106+
if Path(image).is_file():
107+
shutil.copy(image, work / Path(image).name)
108+
available.add(Path(image).name)
109+
110+
run = subprocess.run(
111+
[
112+
"pandoc",
113+
"-f",
114+
"markdown-implicit_figures",
115+
"-t",
116+
"latex",
117+
"-s",
118+
f"--template={_TEMPLATE}",
119+
"-o",
120+
"set.tex",
121+
],
122+
input=_marked_document(fields, available),
123+
capture_output=True,
124+
text=True,
125+
# Not the locale's encoding: a set holding any non-ASCII character would
126+
# then fail to even be handed over under, say, LC_ALL=C.
127+
encoding="utf-8",
128+
cwd=work,
129+
timeout=_TIMEOUT,
130+
)
131+
if run.returncode:
132+
return [Problem(_SET, f"pandoc cannot read the set: {run.stderr.strip()}")]
133+
134+
latex = (work / "set.tex").read_text(encoding="utf-8")
135+
run = subprocess.run(
136+
[
137+
"xelatex",
138+
"-interaction=nonstopmode",
139+
"-file-line-error",
140+
"-no-shell-escape",
141+
"set.tex",
142+
],
143+
stdin=subprocess.DEVNULL,
144+
capture_output=True,
145+
text=True,
146+
encoding="utf-8",
147+
cwd=work,
148+
timeout=_TIMEOUT,
149+
)
150+
return _reported(run.stdout, _locations(latex))
151+
152+
153+
def _marked_document(fields: list[tuple[str, str]], available: set[str]) -> str:
154+
"""The whole set as one markdown document, each field under a marker naming it.
155+
156+
The fence is four backticks so that a field which itself contains a code block
157+
cannot close the marker's raw-LaTeX block early.
158+
"""
159+
blocks = []
160+
for location, markdown in fields:
161+
markdown = _IMAGE.sub(
162+
lambda image: (
163+
f"{image[1]}{Path(image[2]).name}{image[3]}"
164+
if Path(image[2]).name in available
165+
else ""
166+
),
167+
markdown,
168+
)
169+
blocks.append(f"````{{=latex}}\n{_MARKER}{location}\n````\n\n{markdown}\n")
170+
return "\n".join(blocks)
171+
172+
173+
def _locations(latex: str) -> list[tuple[int, str]]:
174+
"""Each marker in the generated LaTeX as the line it is on and the field it names."""
175+
return [
176+
(number, line.partition(_MARKER)[2])
177+
for number, line in enumerate(latex.splitlines(), start=1)
178+
if line.startswith(_MARKER)
179+
]
180+
181+
182+
def _reported(log: str, locations: list[tuple[int, str]]) -> list[Problem]:
183+
"""The xelatex log's errors as problems, each against the field it happened in.
184+
185+
The same error repeated - a command used twice, say - is one problem, since the
186+
author has one thing to go and fix.
187+
"""
188+
found = []
189+
for error in _ERROR.finditer(log):
190+
file, line, message = error[1], int(error[2]), error[3].strip()
191+
where = _SET
192+
if file == "set.tex":
193+
for number, location in locations:
194+
if number <= line:
195+
where = location
196+
problem = Problem(where, f"the PDF generator cannot compile this: {message}")
197+
if problem not in found:
198+
found.append(problem)
199+
return found

‎tests/fixtures/problems/README.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -4,6 +4,10 @@ Each folder here is a hand-written Lambda Feedback export exhibiting exactly one
44
`in2lambda.validation.validate` looks for, beside the `expected.txt` report it should produce:
55
one `str(Problem)` line per problem, which the test compares sorted.
66

7+
The reports are the ones produced with pandoc and xelatex installed, since the validator then
8+
also compiles the set as Lambda Feedback's PDF generator does. Without them those tests are
9+
skipped, because the report would be missing whatever the compiler would have said.
10+
711
They are written by hand rather than exported by the platform, because the platform does not
812
produce broken sets. Real exports live in `../exports`, and the same test suite checks that none
913
of them is reported as having a problem.
Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
Question 1 "Undefined command", part (a), text: the PDF generator cannot compile this: Undefined control sequence.
Lines changed: 22 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,22 @@
1+
{
2+
"orderNumber": 0,
3+
"title": "Undefined command",
4+
"masterContent": "",
5+
"publish": true,
6+
"displayFinalAnswer": true,
7+
"displayStructuredTutorial": true,
8+
"displayWorkedSolution": true,
9+
"displayChatbot": true,
10+
"parts": [
11+
{
12+
"orderNumber": 0,
13+
"content": "Evaluate $x = \\nosuchcommand$.",
14+
"answerContent": "",
15+
"workedSolution": {
16+
"content": "",
17+
"children": []
18+
},
19+
"responseAreas": []
20+
}
21+
]
22+
}

0 commit comments

Comments
 (0)