Conversation
delete_edges_for_file used `source_qn LIKE 'file::%' OR target_qn LIKE 'file::%'`. LIKE folds ASCII case and treats `_` as a wildcard, so deleting the edges of `my_file.py` also deleted those of `myXfile.py` and `My_File.py`. The OR also prevented SQLite from using the edge indexes, so every call scanned the whole edges table, which made graph builds quadratic in the number of files. Use two range scans on the indexed columns instead: every qualified name starting with "file::" sorts in ["file::", "file:;"). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Compile each tree-sitter Query once per (language, pattern) with a bounded, thread-safe cache (`_query()` in languages/base.py). Compiling a query costs several milliseconds, more than parsing a whole file, and every plugin method compiled its queries on each call: ~98% of skeleton extraction time. - `CachedParser` wraps each plugin parser and reuses the tree for consecutive calls on the same source (per-thread, 2 trees, keyed by the source bytes, so a changed file never gets a stale tree). - GraphBuilder: lookup tables replace the O(files^2) scans in import resolution; the content hash uses the source already in memory. - GraphBuilder writes symbols, edges and file rows with multi-row INSERT statements and resolves callees from an in-memory name map loaded once per build. sqlite3 releases the GIL on every row, even with executemany, so row-by-row writes stall behind any CPU-bound thread. The plugin diffs are a mechanical swap (`Query(` -> `_query(`, `Parser(...)` -> `CachedParser(Parser(...))`) with no logic change. On a ~1 100-file repository a cold start drops from 193 s to under 7 s, with an identical graph (symbols, edges, files). The test suite runs about 4x faster. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Discovery used rglob("*"), which walked into every directory (including
.git and node_modules before filtering) and ignored .gitignore. Nested
git worktrees, such as the ones Claude Code creates under
.claude/worktrees/, were indexed as part of the repository, so every
symbol showed up once per worktree in find_references, resolve_symbol
and friends, and the index was several times larger than needed.
- In a git work tree, files come from `git ls-files`: tracked files are
always indexed (.gitignore decides, so a tracked build/ or env/ is
included), and untracked, non-ignored files are also checked against
SKIP_DIRS so an un-ignored .venv or node_modules is never crawled.
Git does not descend into nested worktrees, repositories or
submodules.
- Without git (no repository, git missing or refusing the repository,
root ignored by a parent repository), os.walk prunes SKIP_DIRS and
directories whose .git is a file (worktrees, submodules) before
descending. A root that groups several full repositories still
indexes them.
- Paths are decoded with os.fsdecode (non-UTF-8 names), de-duplicated
(unmerged paths are listed once per stage) and sorted. The index keeps
that order, so results no longer depend on which files came from the
cache.
- Only discovered files are re-injected from the skeleton cache, and the
cache is rewritten from the index. Stale entries for deleted or newly
ignored files never come back.
On a repository with eight Claude Code worktrees, this indexes ~9 500
files instead of ~86 000.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All indexing ran inside create_server(), before mcp.run() answered the MCP handshake. On a large repository a cold build took longer than the client's startup timeout, so the server never came up. Indexing and the graph build now run in a background thread owned by the new IndexState (src/codetree/index_state.py): - Single-file tools (get_file_skeleton, get_symbol, get_imports, get_skeletons, get_symbols, get_complexity, analyze_dataflow flow/taint) answer right away, parsing the requested files on demand (only files discovery would index). - Repo-wide and graph tools wait up to 20 s (CODETREE_WAIT_TIMEOUT), then return a "still building" message with progress. - index_status never blocks. It adds status, files_discovered, files_indexed, index_ready, graph_ready, startup_seconds and error; during a rebuild it reports the last committed graph. - Failures release every waiter. A graph failure rolls back and leaves structural tools working. create_server(root) stays synchronous by default and re-raises build errors (tests, embedding); `codetree` uses create_server(root, background=True). No tool signature changes. Also: - Require fastmcp>=3.0.0, the first release that runs sync tools in a thread pool. On 2.x a waiting tool would block the event loop. - Cache writes are atomic, with a unique temp file per save, and keep the file mode. A failed save no longer fails indexing. The cache now stores has_errors, so syntax warnings survive warm starts. - Indexer: index_file(), build(files=, progress=), and a locked, build-then-publish lazy call graph for concurrent tool calls. - GraphStore.rollback(). On a ~9 500-file repository the handshake completes in about 1 s (previously more than 25 s). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
All indexing ran inside
create_server(), beforemcp.run()could answer the MCP handshake. On a large repository a cold build still took longer than the client's startup timeout, so the server never came up. Indexing and the graph build now run in a background thread. On a ~9 500-file repository the handshake completes in about 1 s (previously more than 25 s), and single-file tools answer within ~1.5 s while indexing continues.Changes
src/codetree/index_state.py.IndexStateowns the lifecycle (discovery → indexing with the skeleton cache → graph build) and exposesfiles_ready,index_readyandgraph_readyevents.get_file_skeleton,get_symbol,get_imports,get_skeletons,get_symbols,get_complexity,analyze_dataflowflow/taint) answer immediately by parsing the requested files on demand. Only files that discovery would index are served.CODETREE_WAIT_TIMEOUTto change it), then return a "still building … (indexing N/M files)" message.index_statusnever blocks.status,files_discovered,files_indexed,index_ready,graph_ready,startup_seconds,error. Existing keys are unchanged.GraphStore.rollback()) and leaves structural tools working.create_server(root, background=False): synchronous by default, re-raising build errors (tests, embedding). ThecodetreeCLI usesbackground=True.has_errorsis now cached, so syntax warnings survive warm starts. Older cache entries are re-parsed once.index_file(),build(files=, progress=), and a locked build-then-publish lazy call graph for concurrent tool calls.fastmcp>=3.0.0: 3.0 is the first release that runs sync tools in a thread pool. On 2.x, a tool waiting for the index would block the event loop. The tests already rely on the 3.xlocal_providerAPI.Behaviour changes
index_status.index_statusreturns additional keys.fastmcpversion is now 3.0.0.Test plan
pytest: all tests pass (1206)tests/test_async_startup.py:.gitignore;atexitregistration, and committed stats during a rebuild;has_errorsacross warm starts, and the env timeout.index_status), cold and warm.Related issues
Third of three stacked PRs: #3 → #4 → #5 (this one).
🤖 Generated with Claude Code