You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
* Add S3 file download support to evaluation pipeline
Implements `params["files"]` feature to allow downloading S3 objects into a per-request working directory. Updates `evaluation.py` and security logic to enable read-only file access. Introduces `s3_files.py` for managing S3 interactions and accompanying unit tests in `s3_files_test.py`. Expands documentation in `CLAUDE.md` and adds integration tests to verify functionality.
* Replace S3-based file downloads with HTTPS support
Replaces S3 bucket-based file download logic with direct downloads via HTTPS URLs, removing AWS dependency. Updates `evaluation.py`, rewrites `s3_files.py` to handle streaming downloads securely, and adjusts tests and documentation (`CLAUDE.md`, `evaluation_test.py`, `s3_files_test.py`) to reflect the changes.
* Update Dockerfile to use the latest Python base image
Switches from a version-specific base image to the `latest` tag for the evaluation function, ensuring up-to-date dependencies and compatibility.
* Refactor file handling to replace `filename` with `name` key and improve validation
Standardizes file specification by replacing the `filename` key with `name`. Enhances validation to handle malformed or incomplete file specifications gracefully. Adds error handling to catch unexpected exceptions and ensures temporary directories are cleaned up. Updates tests and documentation to reflect changes.
* Support file uploads via response payload and prioritize over `params["files"]`
Adds `_resolve_submission` to handle responses containing code and file specifications in `"code"` and `"files"` keys. Updates `evaluation.py` to normalize and prioritize file sources from the response payload over `params["files"]`. Enhances tests to validate this logic. Updates documentation to reflect the new behavior.
* Add support for unwrapping `{"code", "files"}` payloads in answers and submissions
Refactors `_resolve_submission` into `_unwrap_payload` and updates answer handling to support the `{"code", "files"}` structure. Adds unit tests for evaluation behavior with structured answers.
* Refactor file-handling logic to unify `answer_files` and `response_files` handling
Replaces `_resolve_submission` and `_unwrap_payload` with `_collect_file_specs` to streamline file normalization and prioritization. Updates test suite to validate the new logic and modifies documentation to reflect the updated file specification and submission flow.
* Revert "Refactor file-handling logic to unify `answer_files` and `response_files` handling"
This reverts commit ee80ab9.
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|`s3_files.py`| Downloads `params["files"]` objects from S3 into the per-request working directory |
14
15
|`dev.py`| CLI wrapper for local manual testing |
15
16
16
17
### Evaluation pipeline (`evaluation.py`)
17
18
18
19
1. Run AST security check on student code
19
-
2. Dispatch by `params["mode"]` (required):
20
+
2. Resolve the submission into `(code, file_specs)` via `_resolve_submission`: the response may be a bare code string, or a `{"code", "files"}` object (or JSON string of one) as sent by the LF web client's upload widget. If `file_specs` (from `response["files"]`, else `params["files"]`) is non-empty, download the listed objects once into a per-request working directory (see `s3_files.py`), used as the subprocess `cwd` for every run in this request
21
+
3. Dispatch by `params["mode"]` (required):
20
22
-**`demo`**: execute code with no stdin; return stdout/plots as `output` feedback (no pass/fail)
21
23
-**`io_test`**: for each test in `params["tests"]`, execute with `test["input"]` as stdin and compare stdout against `test["expected_output"]`; upload matplotlib plots on pass or fail
3. Upload any captured matplotlib figures via `lf_toolkit``upload_image` (`_UPLOAD_FOLDER = "evaluatePython"`); backend is GCS or S3 per `IMAGE_UPLOAD_BACKEND`
24
-
4. Return a `Result` with feedback tags: `pass`, `fail`, `hidden_fail`, `error`, `output`, `summary`
25
+
4. Upload any captured matplotlib figures via `lf_toolkit``upload_image` (`_UPLOAD_FOLDER = "evaluatePython"`); backend is GCS or S3 per `IMAGE_UPLOAD_BACKEND`
26
+
5. Return a `Result` with feedback tags: `pass`, `fail`, `hidden_fail`, `error`, `output`, `summary`
25
27
26
28
### Request shape
27
29
@@ -90,16 +92,45 @@ All source lives in `evaluation_function/`:
90
92
"pep8_feedback": ["E225", "E231"], # custom rule list
91
93
"tests": [...]
92
94
}
95
+
96
+
# files — optional, works with all modes
97
+
# Downloads files into a per-request working directory (the subprocess's
98
+
# cwd) before student code runs, given a pre-signed or public HTTPS URL per
99
+
# file (fetched directly with a GET — no AWS credentials needed here). Data
100
+
# files can be read with open()/pandas.read_csv()/etc.; .py files are
101
+
# importable by student code since they're co-located with the generated
102
+
# script. The same files are also available to the answer code when
103
+
# use_answer_as_expected_output/use_answer_as_test_code is set.
-**Dunder attribute access**: any `__attr__` style attribute
102
131
132
+
`open`/`pathlib` are intentionally **not** blocked here — they're needed to read files loaded via `params["files"]` (see above). **Important caveat**: `preview_function` (this check) and `evaluation_function` (actual grading) are registered as two independent RPC methods in `main.py`; `evaluation.py` never calls `preview.py`. This check only powers editor-time linting feedback — it does not gate what code can do at grading time. The real, load-bearing control for file access is a runtime-injected restricted `open`/`io.open` in `evaluation.py`'s subprocess preamble (`_safe_open`), which blocks *write* access to anything inside the per-run files directory. It is not a hard sandbox boundary — since `os`/`subprocess` remain fully importable and runnable at grading time regardless of this feature, a student can bypass file restrictions entirely via `os`. Treat this as scoping the intended file-access path, not as isolation.
133
+
103
134
## Key commands
104
135
105
136
```bash
@@ -153,7 +184,7 @@ CI runs on Python 3.12 and uploads JUnit XML results (`.github/workflows/test-li
153
184
|`LOG_LEVEL`|`debug`| Logging verbosity |
154
185
|`IMAGE_UPLOAD_BACKEND`|`gcs`| Plot upload backend in lf_toolkit (`gcs` set in Dockerfile; override to `s3` on the service to use AWS) |
155
186
|`GCS_BUCKET`| Runtime env | Target bucket for matplotlib plot uploads; set per-environment on the Cloud Run service. Auth is via the runtime service account (ADC) — no keys |
156
-
|`AWS_*` / `S3_BUCKET_URI`| Runtime env | Only for the legacy S3 plot-upload backend (`IMAGE_UPLOAD_BACKEND=s3`) |
187
+
|`AWS_*` / `S3_BUCKET_URI`| Runtime env | Only for the legacy S3 plot-upload backend (`IMAGE_UPLOAD_BACKEND=s3`). Not needed for `params["files"]` / response-payload file downloads — those are plain HTTPS GETs from a pre-signed/public URL|
157
188
|`SANDBOX_ENABLED`|`true`| Wrap the worker in shimmy's nsjail sandbox (needs `--privileged` at run time) |
0 commit comments