-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathspec.txt
More file actions
330 lines (197 loc) · 6.32 KB
/
Copy pathspec.txt
File metadata and controls
330 lines (197 loc) · 6.32 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
Product spec — Research Agent v0 (Carter)
1) Goal
Validate an auditable, bounded, citation-grounded research agent that produces a 1–2 page research brief from a user question, with traceability from every claim to evidence.
2) Target users
Engineers / PMs needing quick technical landscape scans
Internal platform/infra teams producing “decision briefs”
Researchers creating fast lit-review scaffolds
3) User experience
Input
Research question (required)
Optional: constraints
time budget (minutes)
source constraints (domain allowlist/denylist)
recency window (e.g., last 2 years)
number of sources (e.g., 5–10)
Output
Research brief (markdown)
Executive summary
Key findings (bullet points)
Contradictions / disagreements (if found)
Gaps / open questions
References
Claims table (machine-readable + human readable)
claim → evidence snippets → source URLs
Run report
cost/time/tool-call counts
stopping reason
audit trail pointer (event log)
4) Success criteria (v0)
Grounding: ≥ 90% of non-trivial claims have at least one evidence snippet + source.
No hallucinated references: every cited source must be retrievable and appear in the final reference list.
Bounded: run completes within configured budgets (time + tool calls).
Auditable: can reconstruct “why” a claim exists via event log + evidence IDs.
5) Non-goals (v0)
Paywalled / proprietary retrieval solving
Fully automated credibility scoring
Domain-specific deep reading (PDF parsing/OCR heavy)
Multi-agent debate
Fine-tuning or learning across runs (save this for v1)
6) Key product decisions
Evidence-first writing: the agent must collect evidence snippets before drafting conclusions.
Strict citation policy: claims without evidence are either removed or placed under “Speculation / Hypotheses” with explicit labeling.
Stop rules are explicit: completion is a function of coverage score + diminishing returns + budget.
7) Core metrics
Time to draft (p50/p90)
Sources used (count + diversity)
Claim grounding rate
Reviewer acceptance rate (manual rubric)
Cost per brief (tokens + tool calls)
8) MVP rollout plan
Week 1: CLI-only internal prototype (markdown output + JSON event log)
Week 2: Simple web UI (optional), add replay + regression suite
Engineering spec — Research Agent v0 (Carter)
1) Architecture (kernel + plugins + policy)
Kernel (stable)
Orchestrator (control loop)
Event Log (append-only)
Artifact Store (sources, snippets, drafts)
Policy Engine (gates + redaction rules)
Evaluator (coverage/grounding/cost + stop)
Plugins (replaceable)
Retriever: web search + fetch (later: internal docs)
Extractor: snippet extraction + metadata
Synthesizer: outline → brief
Verifier: claim→evidence mapping checks
2) Data model (minimal)
TaskSpec (input)
{
"question": "string",
"constraints": {
"max_tool_calls": 25,
"max_minutes": 10,
"max_sources": 10,
"recency_days": 3650,
"allow_domains": [],
"deny_domains": []
}
}
Event Log (append-only)
Each entry is immutable and timestamped.
{
"event_id": "ulid",
"ts": "iso8601",
"type": "TASK_CREATED | SEARCHED | SOURCE_SELECTED | SNIPPET_EXTRACTED | OUTLINE_DRAFTED | CLAIMS_BUILT | BRIEF_RENDERED | STOPPED | ERROR",
"payload": {}
}
Source record
{
"source_id": "ulid",
"title": "string",
"url": "string",
"published_date": "string|nullable",
"retrieved_at": "iso8601",
"content_hash": "sha256",
"notes": "string|nullable"
}
Evidence snippet
{
"evidence_id": "ulid",
"source_id": "ulid",
"quote": "string",
"char_start": 123,
"char_end": 234,
"relevance": 0.0
}
Claim
{
"claim_id": "ulid",
"text": "string",
"evidence_ids": ["ulid"],
"confidence": 0.0,
"tags": ["finding|contradiction|gap|definition"]
}
3) Control loop (v0)
Normalize TaskSpec (defaults + policy)
Retrieve candidates (search)
Select sources (diversity + relevance)
Extract snippets (evidence-first)
Build claims table (claim drafting grounded in evidence)
Draft brief from claims + outline
Verify:
every claim has ≥1 evidence snippet (unless explicitly labeled as speculation)
every evidence_id resolves
every source_id resolves + URL present
Evaluate & stop:
stop if: coverage score ≥ threshold AND diminishing returns hit
or budgets exhausted → “partial brief” with gaps + explicit stop reason
4) Stopping rules (concrete)
Hard stops
tool calls >= max_tool_calls
time >= max_minutes
sources >= max_sources
Soft stop: diminishing returns
run one more retrieval round only if it improves:
coverage score by ≥ Δ (e.g., 0.05), or
adds new contradiction/gap category
5) Evaluator (v0 scoring)
Coverage: fraction of brief sections populated with grounded claims
Grounding: % claims with evidence_ids length ≥ 1
Diversity: # unique domains / source types
Cost: tool calls + tokens
Return:
{
"coverage": 0.82,
"grounding": 0.94,
"diversity": 0.6,
"stop_reason": "DIMINISHING_RETURNS"
}
6) Security & privacy (v0)
No secrets ingestion: redact tokens/keys if detected in fetched text (basic regex)
Domain allow/deny enforcement
Store only what’s needed:
source URL + hash + snippets used (avoid saving full page where possible)
7) Repo structure
carter/
README.md
docs/
product-spec.md
engineering-spec.md
schemas.md
src/
main.py (or main.ts)
kernel/
orchestrator.*
event_log.*
policy.*
evaluator.*
artifacts.*
plugins/
retriever_web.*
extractor_basic.*
synthesizer_llm.*
verifier_basic.*
tests/
fixtures/
regression_cases.json
test_replay.*
8) Minimal deliverables checklist (v0)
CLI: carter run --question "..." --max-sources 8 --max-minutes 8
Outputs:
brief.md
claims.json
run_report.json
event_log.jsonl
Replay: carter replay --run-id ... reproduces outputs with cached artifacts
Regression: 5 canned questions with expected minimum grounding score
9) Tech stack suggestion
Python (fast prototype, good text tooling)
pydantic for schemas
sqlite or plain jsonl for event log (v0: jsonl is fine)
pluggable retriever interface
(If you prefer Node/TS, same architecture maps cleanly.)
10) Milestones (tight)
M0 (day 1–2): schemas + event log + CLI skeleton
M1 (day 3–4): web retriever + snippet extractor + claims table
M2 (day 5–6): brief synthesis + verifier + evaluator + stop rules
M3 (day 7): replay + regression harness