Skip to content

fix(grep): skip binary and non-UTF-8 input instead of aborting - #124

Open
so0k wants to merge 1 commit into
strands-agents:mainfrom
so0k:fix/grep-binary-files
Open

so0k wants to merge 1 commit into
strands-agents:mainfrom
so0k:fix/grep-binary-files

Conversation

@so0k

@so0k so0k commented Sep 30, 2026

Copy link
Copy Markdown

Description

grep read every line into a String, so the first file that wasn't valid UTF-8 made read_line fail with InvalidData. cmd_grep then propagated that with ?, which aborted the whole file loop: the command exited 1 with grep: stream did not contain valid UTF-8, and every file after the offending one (in directory order) was silently never scanned.

This change makes grep byte-oriented and handles binary input the way GNU grep (3.11, UTF-8 locale) does:

  • Lines are read with read_until(b'\n'), so file content can no longer make grep fail.
  • A NUL byte in the first buffer marks the whole file binary. Its lines are never printed; a match produces grep: FILE: binary file matches on stderr, with exit status 0, and reading stops after the first match.
  • An individual line with invalid UTF-8 (or a NUL) is matched lossily but not printed. The file's other lines print normally, followed by the same stderr note.
  • -c, -l, -L and -q count binary matches as usual and print no note.
  • Hidden binary lines still take up their slot in the -A/-B/-C window, so context output never shows lines from outside the requested range.
  • New options: -a/--text, -I, and --binary-files=binary|text|without-match (a small BinaryFiles enum).
  • A read error on one file is now reported as grep: PATH: ERR and grep moves on to the next file, the same way open errors are already handled. A closed stdout (BrokenPipe, e.g. grep -r … | head -n 1) still stops grep straight away.
  • grep takes stderr once up front and shares it across files. Previously it called io::stderr() for each open error, and the second call failed because the writer had already been taken.

Before (0.3.3) and after, using the repro from #123 (Node binding, direct read-only bind):

$ grep -rn needle /mnt
- status: 1  stdout: "/mnt/c.txt:1:needle in c.txt too\n"  stderr: "strands-shell: grep: stream did not contain valid UTF-8\n"
+ status: 0  stdout: "/mnt/c.txt:1:needle in c.txt too\n/mnt/a.txt:1:needle in a.txt\n"  stderr: ""
$ grep -rnI needle /mnt
- status: 1  stderr: "strands-shell: grep: invalid option '-I'\n"
+ status: 0  (same two matches)

Known differences from GNU grep (not addressed here)

  • Whole-file detection looks only at the first BufReader fill (8 KiB); GNU uses a larger buffer. A NUL that appears later hides only the line it is on.
  • -a decodes invalid bytes lossily (U+FFFD) rather than writing the raw bytes.
  • -o skips matches on binary lines, where GNU prints them.
  • In a few -B/-C layouts where a hidden binary line splits two groups, the -- separators differ slightly from GNU.

Out of scope, possible follow-ups

  • GNU exits with 2 when any file errored. This PR keeps the existing exit-status behaviour for open errors and read errors.
  • -r visits files in unsorted directory order, as before.
  • cargo clippy --workspace --all-targets -- -D warnings already fails on main (find.rs, lua.rs, ls.rs, exec.rs). This PR adds no new warnings.

Related Issues

Fixes #123

Documentation PR

None yet. The command reference page could list the new -a, -I and --binary-files options.

Type of Change

Bug fix

Testing

Checklist

  • I have read the CONTRIBUTING document
  • I have reviewed and understand every line of code in this PR, including any generated by AI tools, and I can explain why it works
  • My change is focused and reasonably small; I have split unrelated work into separate PRs
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

🤖 Generated with Claude Code

https://claude.ai/code/session_01AXaT7zwomYh4CtqSGrCY1Y

grep read lines into a String, so the first non-UTF-8 file failed
read_line with InvalidData, and the `?` in cmd_grep aborted the file
loop: the command exited 1 and every later file was never scanned.

Read lines as bytes and treat binary input like GNU grep: a NUL in the
first buffer marks the file binary, and invalid-UTF-8 lines are matched
lossily but never printed. Either way a match reports
"grep: FILE: binary file matches" on stderr. Add -a/--text, -I and
--binary-files=TYPE. Per-file read errors are reported and grep moves
on; BrokenPipe still stops it.

Fixes strands-agents#123

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXaT7zwomYh4CtqSGrCY1Y
@so0k
so0k requested a review from a team as a code owner September 30, 2026 10:01
@so0k
so0k requested a review from Unshure September 30, 2026 10:01

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] grep -r aborts the whole command (and silently drops matches from other files) when a bound directory contains a non-UTF-8 file

1 participant