Skip to content

feat: discover files with git ls-files (respect .gitignore, skip nested worktrees) - #4

Open
Kevinrob wants to merge 3 commits into
ThinkyMiner:mainfrom
qoqa:feat/git-file-discovery
Open

Kevinrob wants to merge 3 commits into
ThinkyMiner:mainfrom
qoqa:feat/git-file-discovery

Conversation

@Kevinrob

@Kevinrob Kevinrob commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #3. Only the last commit is new; review feat: discover files with git ls-files….

Summary

File discovery used rglob("*"). It walked into every directory, including .git and node_modules, before filtering, and it ignored .gitignore. Nested git worktrees, such as the ones Claude Code creates under .claude/worktrees/, were indexed as part of the repository. Every symbol then showed up once per worktree in find_references, resolve_symbol, get_blast_radius, etc., and the index was several times larger than needed. On a repository with eight worktrees, codetree indexed ~86 000 files instead of ~9 500.

Changes

  • Git work tree: files come from git ls-files.
    • Tracked files are always indexed: .gitignore decides, so a tracked build/ or env/ is included.
    • Untracked, non-ignored files (--others --exclude-standard) are also checked against SKIP_DIRS, so an un-ignored .venv or node_modules is never crawled. python -m venv on Python < 3.13 writes no .gitignore.
    • Git does not descend into nested worktrees, repositories or submodules.
  • No git (not a repository, git missing or refusing the repository, e.g. safe.directory, or the root ignored by a parent repository): os.walk prunes SKIP_DIRS and directories whose .git is a file (worktrees, submodules) before descending. A root that groups several full repositories still indexes them.
  • Paths are decoded with os.fsdecode (non-UTF-8 names), de-duplicated (git lists an unmerged path once per stage) and sorted. The index keeps that order, so results no longer depend on which files came from the cache.
  • Only discovered files are re-injected from the skeleton cache, and the cache is rebuilt from the index. Stale entries for deleted or newly ignored files can no longer come back.
  • README: new "What Gets Indexed" section.

Behaviour changes

  • In git repositories, tracked files under build/, dist/, env/, etc. are now indexed. Ignored files are no longer indexed even if they sit outside SKIP_DIRS.
  • Nested worktrees, repositories and submodules inside a git repository are no longer indexed.

Test plan

  • Ran pytest: all tests pass (1182)
  • Added tests/test_file_discovery.py (20 tests): .gitignore, untracked files, a real nested git worktree, a nested repository, a submodule, root inside a worktree, deleted-but-tracked files, tracked SKIP_DIRS names, un-ignored .venv/node_modules, non-UTF-8 names, merge conflicts, cache-independent order, walk fallback (no repository, ignored root, git missing, nested worktrees vs. full repositories), stale cache entries.

Related issues

Second of three stacked PRs: #3 → #4 (this one) → #5.

🤖 Generated with Claude Code

Kevinrob and others added 3 commits October 1, 2026 15:02
delete_edges_for_file used `source_qn LIKE 'file::%' OR target_qn LIKE
'file::%'`. LIKE folds ASCII case and treats `_` as a wildcard, so
deleting the edges of `my_file.py` also deleted those of `myXfile.py`
and `My_File.py`. The OR also prevented SQLite from using the edge
indexes, so every call scanned the whole edges table, which made graph
builds quadratic in the number of files.

Use two range scans on the indexed columns instead: every qualified name
starting with "file::" sorts in ["file::", "file:;").

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Compile each tree-sitter Query once per (language, pattern) with a
  bounded, thread-safe cache (`_query()` in languages/base.py). Compiling
  a query costs several milliseconds, more than parsing a whole file, and
  every plugin method compiled its queries on each call: ~98% of
  skeleton extraction time.
- `CachedParser` wraps each plugin parser and reuses the tree for
  consecutive calls on the same source (per-thread, 2 trees, keyed by
  the source bytes, so a changed file never gets a stale tree).
- GraphBuilder: lookup tables replace the O(files^2) scans in import
  resolution; the content hash uses the source already in memory.
- GraphBuilder writes symbols, edges and file rows with multi-row
  INSERT statements and resolves callees from an in-memory name map
  loaded once per build. sqlite3 releases the GIL on every row, even with
  executemany, so row-by-row writes stall behind any CPU-bound thread.

The plugin diffs are a mechanical swap (`Query(` -> `_query(`,
`Parser(...)` -> `CachedParser(Parser(...))`) with no logic change.

On a ~1 100-file repository a cold start drops from 193 s to under 7 s,
with an identical graph (symbols, edges, files). The test suite runs
about 4x faster.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Discovery used rglob("*"), which walked into every directory (including
.git and node_modules before filtering) and ignored .gitignore. Nested
git worktrees, such as the ones Claude Code creates under
.claude/worktrees/, were indexed as part of the repository, so every
symbol showed up once per worktree in find_references, resolve_symbol
and friends, and the index was several times larger than needed.

- In a git work tree, files come from `git ls-files`: tracked files are
  always indexed (.gitignore decides, so a tracked build/ or env/ is
  included), and untracked, non-ignored files are also checked against
  SKIP_DIRS so an un-ignored .venv or node_modules is never crawled.
  Git does not descend into nested worktrees, repositories or
  submodules.
- Without git (no repository, git missing or refusing the repository,
  root ignored by a parent repository), os.walk prunes SKIP_DIRS and
  directories whose .git is a file (worktrees, submodules) before
  descending. A root that groups several full repositories still
  indexes them.
- Paths are decoded with os.fsdecode (non-UTF-8 names), de-duplicated
  (unmerged paths are listed once per stage) and sorted. The index keeps
  that order, so results no longer depend on which files came from the
  cache.
- Only discovered files are re-injected from the skeleton cache, and the
  cache is rewritten from the index. Stale entries for deleted or newly
  ignored files never come back.

On a repository with eight Claude Code worktrees, this indexes ~9 500
files instead of ~86 000.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant