生物医学 × 计算交叉领域的文献自动追踪 + AI 深度精读模板
用 GitHub Actions 按周期自动追踪 PubMed · arXiv · bioRxiv · medRxiv · chemRxiv 上的最新文献,存档落盘、生成 Issue 周报,再用 LLM 为每篇产出「评分 + 一句话 + 摘要翻译 + 16 节 Paper Card + 审稿人评审」,并同步部署一个可全文检索的 GitHub Pages 站点,配合 Zotero 完成筛选与入库。
三步上手 · 设计思路 · 抓取与产出 · AI 深度精读 · 站点与搜索 · 平台检索要点 · 密钥与仓库配置 · 本地运行与调试 · 真实案例:5 年全量回填 · 每周阅读工作流
模板即实例。 仓库内的
config.yaml与prompts.yaml已内置一套完整可跑的示例检索式与提示词 —— 拿到后你只需改两个文件:config.yaml(换成你的研究领域)与prompts.yaml(换成你领域的措辞)。工具与平台默认面向 生物医学 × 计算交叉课题。核心驱动为我们自研的文献检索获取工具 pyPaperFlow。
| 步骤 | 你要做的 | 具体内容 |
|---|---|---|
| 1 · 换领域 | 编辑 config.yaml 与 prompts.yaml |
改检索式与命名、换 4 个 prompt 的领域措辞,具体见下方「换课题(最终操作)」 |
| 2 · 设定时与密钥 | 配置 workflow、Secrets 与 Pages |
确认 monitor.yml 的 cron;在 Settings → Secrets and variables → Actions 添加 ENTREZ_EMAIL(必填)、NCBI_API_KEY(可选)与 LLM_BASE_URL / LLM_API_KEY / LLM_MODEL(AI 必填);并在 Settings → Pages 启用 Pages(来源选 GitHub Actions) |
| 3 · 每周收报 | 等待或手动触发 actions | 有新文献时自动开一条 Issue 周报,同时更新 GitHub Pages 站点:按相关度评分排序,逐篇可读摘要翻译、16 节 Paper Card 与审稿人评审,并支持全文检索 |
- GitHub 上 Use this template 建新仓库
- 改
config.yaml:topic、title、site_base_url、platforms.*.query - 改
prompts.yaml:4 个 prompt 里与领域相关的措辞 - 配 Secrets:
ENTREZ_EMAIL、LLM_BASE_URL/LLM_API_KEY/LLM_MODEL - 先启用 Pages:Settings → Pages 把来源设为 GitHub Actions(不启用则站点不会部署)
- 手动
workflow_dispatch跑一次验证
推送到
main后也可在仓库 Actions 页手动触发试跑;想先本地验证见「本地运行与调试」。
一句话:
每周自动「替你搜一遍研究领域」→ 新文献永久存档、另出一张可筛选清单 → LLM 为每篇写深度阅读材料 → 开一条 Issue 提醒、部署一个站点沉淀下来。 最后是否精读、入库、去重,都留给你自己判断。
设计取舍:
- 元数据分发,不碰版权全文 —— 只抓题录 / 作者 / 摘要 / DOI,轻量合规;要全文时按需走你自己的文献工具,或继续使用我们自研的 pyPaperFlow。
- 双份数据落盘,职责分离 —— 只增的
Archive/当文献库底账;按周的Discovery/快照当筛选清单;site/再把这些数据渲染成可检索的静态站点。 - AI 负责「读」,你负责「判」 —— 每篇自动产出评分、翻译、Paper Card 与评审报告,把「要不要读原文」的决策成本压到最低,但最终判断仍在你。
- Issue 即提醒,站点即沉淀 —— 不额外接邮件推送,有命中才开、零命中不打扰;静态站点把每周产出永久沉淀、可全文检索。
- 密钥不入库 —— PubMed 凭据与 LLM 密钥一律走环境变量 / Actions Secret。
- 一个课题一个仓 —— 同一骨架可复制成多个独立仓库,并行追踪多个方向。
抓什么: 每周期一次,逐平台用 pyPaperFlow 检索「窗口内新文献」;窗口 = 运行日往前 window_days 个自然日(不含当天)。
三类产物:
| 产物 | 内容 | 归档路径 | 口径 |
|---|---|---|---|
Archive/ |
逐篇完整元数据 JSON + LLM 深度阅读材料 | Archive/<source>/<year>/<month>/<id>/ |
只增不删,按月归档 |
Discovery/ |
合并 CSV + _ids.txt |
Discovery/<year>/<month>/<topic>_<date>.csv |
按抓取日归档 |
site/ |
静态站点(解析页 + 搜索索引) | site/,由 Pages 部署 |
每次运行重建 |
Archive/ 下每篇文献是一个目录:
Archive/{source}/{year}/{month}/{id}/
{id}.json # 完整元数据(pyPaperFlow 原始 JSON)
fulltext.md # 全文原文(含获取来源;拿不到则回落摘要)
analysis.json # 元数据 + 评分 / 一句话 / 摘要翻译 + 各产物路径
paper-card.md # 16 节深度阅读卡片(score ≥ min_score 才生成)
review.md # 审稿人评审报告(score ≥ min_score 才生成)
CSV 固定 9 列(utf-8-sig,Excel 友好):
source, id, doi, title, authors, journal, published_date, url, abstract
id 为平台主键(PubMed = PMID,预印本 = DOI)。_ids.txt 每行一个标识符,供 Zotero「按标识符添加」批量导入:pmid:xxx / arXiv:xxx(自动去版本号)/ 其余预印本裸 DOI。
日期口径(entrez date): PubMed 的 DP 常残缺且存在标引时滞,故搜索、归档、Issue 三处统一用 [edat](被 PubMed 收录的日期);真实 DP 仍保存在各 JSON 的 data.source.pub_date。预印本用 posting 日期。不做跨平台去重、不判重,当周某平台 0 命中时 CSV 仅含表头。
🗂️ 仓库结构
.
├── scripts/
│ ├── monitor.py # 主脚本:读 config → 逐平台检索 → 落盘 → 生成 Issue → 触发 LLM 流水线
│ ├── backfill.py # 历史回填:按周驱动 monitor.py 补齐过去各周(前置探活 + 断点续跑)
│ ├── agent.py # LLM 层:OpenAI 兼容客户端 + 评分/翻译/Paper Card/评审 四个生成函数
│ ├── fulltext.py # 全文获取封装:优先原生全文,失败回落摘要
│ └── build_site.py # 静态站点生成:遍历 Archive/ 渲染 Jinja2 模板
├── templates/ # 站点模板(本周 / 归档 / 搜索 / 单篇解析页)+ 样式与搜索脚本
├── config.yaml # 检索配置:topic / title / window_days / llm / 每平台一条 query
├── prompts.yaml # 四个 system prompt(score / translate / paper_card / review)集中于此
├── assets/banner.svg # README 头图
├── requirements.txt # 依赖:pyPaperFlow + openai + jinja2 等
├── tests/ # pytest 自检用例(不参与部署,可删除)
└── .github/workflows/
├── monitor.yml # 定时任务 + 手动触发;跑完 commit+push 并开 Issue
└── deploy_pages.yml # monitor 成功后自动把 site/ 部署到 Pages
每篇命中文献跑一次 LLM 流水线(OpenAI 兼容网关,system prompt 全部来自 prompts.yaml),产出深度阅读材料:
| 阶段 | 函数 | 产出 | 输出上限 |
|---|---|---|---|
| 评分 + 一句话 | score_paper |
analysis.json 的 score / one_liner_zh:0–10 相关度打分 + 一句中文概括,进入 Issue 表格与站点排序 |
500 token(JSON 模式) |
| 摘要翻译 | translate_abstract |
analysis.json 的 abstract_zh:原文摘要的忠实中文译文 |
2000 token |
| Paper Card | build_paper_card |
paper-card.md:固定 16 节的深度阅读卡片 |
16000 token |
| 审稿人评审 | build_review |
review.md:Nature 风格单人评审报告 |
12000 token |
config.yaml 的 llm.min_score(默认 5)控制后两类深度产物的生成:仅当 score >= min_score 时才调用 Paper Card + 评审报告;低于阈值则只保留「评分 + 一句话 + 摘要 + 摘要翻译」。对应评分分档 0–4 / 5–7 / 8–10,5 为「部分命中」入口。
站点单篇解析页的顺序为 head 信息 → 摘要 → 摘要翻译 →(高分时)Paper Card → (高分时)评审报告;低分篇只有前三者。
LLM 调用以网关 IO 等待为主,用线程池并发(llm.concurrency,默认 8)把墙钟时间从「每篇耗时之和」压到约「每篇耗时 × (篇数 / 并发)」,避免顺序跑 80+ 篇超出 GitHub Actions 单 job 6 小时上限。单篇失败只跳过该篇、不中断整批。
并发上限(SJTU 网关限流): 若你的 LLM 网关对单个 API key 有请求速率限制(例如 SJTU 网关约 10 次/窗口),
concurrency要明显低于该限制——因为concurrency是「同时跑的论文数」,而每篇论文会串行调用 LLM 24 次(评分 + 翻译 + 高分时的卡片 + 评审),失败还会重试,所以并发 10 篇的实际请求量是 2040 次,照样打爆「10 次/窗口」的硬顶。收到大量429时,两类要分清:
Rate limit exceeded for api_key ... Current limit: 10, Remaining: 0—— key 级请求限速;No deployments available for selected model ... cooldown_list=[...]—— 网关把模型标进冷却,暂时无可用实例。这两种都会让该篇
[agent] ... FAILED、analysis.json不落盘(只剩元数据 + 全文),即静默缺 score/翻译/卡片/评审。backfill.py会读取 monitor 打出的Warning: N paper(s) agent-failed: [...]汇总行,把该周标为需重跑。默认concurrency: 8即为此而设(低于 10、留出重试与峰值余量)——不要盲目调大,出现 429 时反而应继续调低。
build_site.py 遍历 Archive/ 渲染出无后端静态站点,由 deploy_pages.yml 在 monitor 成功后自动部署到 GitHub Pages。页面包括:本周推荐(按评分排序)、按年/周归档、全文搜索、单篇解析页。
站内搜索(两段式索引):站点无后端,搜索完全在浏览器端完成,索引在导出阶段预生成:
| 索引 | 文件 | 字段 | 加载时机 |
|---|---|---|---|
| 头索引(head) | site/data/search.json |
标题、一句话、作者/期刊/日期、摘要(原文 + 中文) | 页面加载即取 |
| 深索引(deep) | site/data/search-deep.json |
仅 {id, deep},deep 为 Paper Card + 评审的纯文本 |
勾选「深度搜索」时惰性拉取 |
检索为朴素加权子串匹配:查询做 NFKD 折叠 + 小写 + 去除非字母数字,按 AND 语义打分;字段权重 title(18) > summary(6) > meta(4) > abstract(1)。深索引是头索引的超集而非替换——勾选深度搜索后摘要等字段仍在,只是追加了 card/review 正文。
全文获取来源(
fulltext_source):PubMed → PMC 全文,arXiv → ar5iv,bioRxiv/medRxiv → Europe PMC,chemRxiv 无全文路由恒回落摘要。
各平台检索语法差异很大,详细的召回与噪声踩坑结论都写在各 query 正上方的注释里,改写时请务必保持:
- PubMed ——
(对象) AND (方法)括号两段式;Mesh 受标引时滞影响,周窗召回主要靠[tiab]精确词;[edat]时间窗由代码自动拼接。 - arXiv —— 布尔项须写成
all:"phrase"/all:word,裸词会被强 AND、OR 失效;建议设max_results上限,否则会翻整周全部命中导致限速挂起。 - bioRxiv / medRxiv —— 同一检索器(Europe PMC + Crossref 超集)。必须保留括号两段式,不可拍平成无括号 DNF——Europe PMC(Lucene)会把无括号 AND/OR 混排错乱。Europe PMC 不可达时自动降级为纯 Crossref 检索(丢失仅出现在全文里的词命中),monitor 在 stderr 打
DEGRADED告警,backfill 据此把该周标为需重跑。 - chemRxiv —— 仅 Crossref 收录(Europe PMC 不覆盖),是唯一无严格索引的平台,接受一定噪声,交给 Zotero 兜底。
PubMed 与 LLM 凭据一律走环境变量,不写入仓库:
| 用途 | 环境变量 | 必填 |
|---|---|---|
| PubMed 邮箱 | ENTREZ_EMAIL |
必填 |
| NCBI API key | NCBI_API_KEY |
可选 |
| LLM 网关地址 | LLM_BASE_URL |
AI 必填 |
| LLM API key | LLM_API_KEY |
AI 必填 |
| LLM 模型 | LLM_MODEL |
可选(默认 deepseek-chat) |
- 本地:
export ENTREZ_EMAIL=you@example.com等。 - GitHub Actions:Settings → Secrets and variables → Actions → Repository secrets / Variables 新建对应项,workflow 以
${{ secrets.* }}/${{ vars.* }}注入。
推送周期在 monitor.yml 的 schedule.cron(默认周一 09:23 UTC,避开整点);workflow 权限为 contents: write + issues: write:跑完 commit+push 产物,再按是否命中决定开 Issue。站点部署由 deploy_pages.yml 监听 monitor 成功后自动触发,需先在仓库 Settings → Pages 启用 Pages(来源设为 GitHub Actions),否则 deploy_pages.yml 会因 Pages 未开启而失败。
模型选型提示: 评分阶段走
response_format={"type":"json_object"},所选模型必须支持 JSON mode(如 DeepSeek 的deepseek-reasoner通常不支持,不能直接替换);评分/翻译要忠实、不宜用思考/推理模式;卡片/评审单次输出达 16k / 12k token,需留意网关超时与 token 成本。
pip install -r requirements.txt # 或 pip install pyPaperFlow
export ENTREZ_EMAIL=you@example.com
export LLM_BASE_URL=... LLM_API_KEY=... # 可选 LLM_MODEL=...
python scripts/monitor.py \
--config config.yaml \ # 指定检索配置
--out-dir . \ # Archive/ 与 Discovery/ 落在仓库根
--issue-body /tmp/issue.md \ # (可选)生成的 Issue 正文
--issue-title /tmp/issue.title # (可选)生成的 Issue 标题
python scripts/build_site.py \
--out-dir . \ # 从 Archive/ 重建 site/
--config config.yaml| 常用参数 | 作用 |
|---|---|
--window-days 1 |
把窗口收窄到 1 天快速试跑,无需改 config |
--run-date 2026-09-03 |
固定运行日,便于回测某周 |
--platforms arxiv,pubmed |
只跑逗号分隔列出的平台,跳过其余平台(默认跑 config 里全部已配置平台) |
窗口默认不含运行当天;单平台失败仅告警,全部失败才非零退出。
新建仓库后想补齐过去各周的文献(而不是只从今天往前 window_days),用 scripts/backfill.py 按周驱动 monitor.py 逐周回填:
export ENTREZ_EMAIL=... LLM_BASE_URL=... LLM_API_KEY=... # 与 monitor 相同的环境变量
# 1) 先 dry-run 看周计划:--since/--until 之间每个周一各跑一次,窗口 = 该周一往前 7 天
python scripts/backfill.py --since 2025-09-15 --until 2026-09-14 --dry-run
# 2) 正式回填
python scripts/backfill.py --since 2025-09-15 --until 2026-09-14某平台因故漏抓时(例如 arXiv 曾因 API 拒绝请求而整段缺失),不必整段重跑所有平台,用 --platforms 只回补指定平台:
# 只重跑 arxiv,其余平台跳过;ENTREZ_EMAIL 不再必需
python scripts/backfill.py --since 2025-09-15 --until 2026-09-14 --platforms arxiv--platforms 为逗号分隔,透传给 monitor.py 的 --platforms。前置探活与 env 校验随平台联动:
- 只跑
arxiv:不再要求ENTREZ_EMAIL(只有 LLM 网关的LLM_API_KEY仍必需);pre-flight 只探测 arXiv 与 LLM 网关。 - 只跑
pubmed:仍需ENTREZ_EMAIL,pre-flight 探测 NCBI eutils + LLM 网关。
要点:
- 为什么按周:
--since/--until之间每个周一拆成独立一次monitor.py运行(--window-days 7),避免多年窗口撞上 PubMedretmax=500/ arXivmax_results截断,也保证老论文按自身所在周归档。 - 前置探活:正式跑之前先并行探测 NCBI / Europe PMC / arXiv / LLM 网关,任一不通就拒绝开始(fresh tmux 里常漏代理,脚本会提示检查
http_proxy/https_proxy)。 - 断点续跑:某周崩溃、平台缺失/降级、或某篇 LLM 调用失败(
agent-failed)时,脚本非零退出并列出该周,用打印出的重试命令只补那一周,不必整段重跑(重复某周会连同其 LLM 工作一并重跑)。 - 跑完手动部署 Pages:回填只提交
Archive/,不提交site/(已 gitignore)。跑完后重建站点并手动触发一次部署:
python scripts/build_site.py --out-dir . --config config.yaml
git add Archive/ Discovery/ && git commit -m "backfill: capture <范围>" && git push
# 再到仓库 Actions → deploy_pages.yml → Run workflow 手动触发测试: tests/ 是 pytest 自检用例(不参与部署,可安全删除)。pip install pytest && pytest 即可。
以本模板的 idp-interaction-ai 实例为例,从 2021-01-01 起按周回填(backfill.py),此后每周自动续跑,截至 2026-09 的真实规模:
| 指标 | 实测值 |
|---|---|
| 归档论文 | 3087 篇(analysis.json 计数) |
| 覆盖周 | 299 周(2021-01-03 .. 2026-09-20),单周最多约 30 篇 |
Archive/ |
265 MB / 14422 文件(只增) |
site/ |
254 MB(papers 139 MB + data 111 MB + weeks 4.5 MB),3389 个 HTML 页面 |
| 最大单文件 | site/data/search-deep.json 89 MB |
build_site.py 全量重建 |
约 48 s |
.git/ |
约 85 MB |
五年、3000 篇的量级,单仓库 + 单 Pages 站点完全扛得住——距 GitHub Pages 站点 1 GB、单文件 100 MB 两条上限都还有余量(当前最大文件即 89 MB 的 search-deep.json,是全仓最需留意的一条)。
耗时与并发。 回填逐周串行驱动 monitor.py:每周一次多平台检索 + 该周 LLM 流水线,周与周之间不并行,总时长 ≈ 周数 × 单周耗时。真正吃时间的是 LLM,检索本身占比很小;llm.concurrency 只在同一周内并行论文,不跨周。所以五年回填是一次以小时计的长跑,务必放进 tmux / nohup,并让 stderr 可见——backfill.py 逐周回显 monitor 的 stderr,中断后可用它打印的重试命令断点续跑。
回填后的 push。 回填产物只有 Archive/ 与 Discovery/(site/ 已 gitignore),git add Archive/ Discovery/ && git commit && git push 即可。注意两点:单次提交 265 MB / 上万文件,GitHub 接受(无单文件超 100 MB),但体积不小;git push 到 GitHub 常需代理(http_proxy / https_proxy),这与「回填命中的 PubMed / arXiv / Europe PMC 直连即可、无需代理」并不矛盾——需要代理的是 GitHub 这一跳。
回填后的部署。 deploy_pages.yml 平时由 monitor 的 workflow_run 触发,而手动回填的 push 不会触发它(回填提交里没有 site/ 变化,也就没有触发源),跑完必须手动触发一次:
gh workflow run deploy_pages.yml --repo <owner>/<repo> --ref main部署 job 是从 Archive/ 重新构建站点(并非复用回填时构建的产物),只做重新渲染、不重新抓取、不重跑 LLM,全量 3000 篇约 1~3 分钟,远低于 Actions 单 job 6 小时上限。
gh用 snap 安装时:sandbox 里看不到仓库目录,在仓库内直接gh workflow run会报fatal: not a git repository (or any parent up to mount point /var/lib)。加--repo owner/repo --ref main显式指定即可,与当前目录无关。
站点仍按历史周分类。 回填论文按各自 analysis.json 的 window_end 落入其自身所在周(site/weeks/<window_end>/),与周更 cron 的周 key 完全一致;首页「本周推荐」只显示最新一周,历史各周在归档页按周翻阅,因此不会出现「首页一次推荐 3000 篇」(单周 ≤ 约 30 篇)。
- 打开仓库 Issues 看当周推送报告,按平台与「评分 / 一句话」粗筛,点标题直达原文、点链接直达pages站点;
- 打开 Pages 站点按相关度排序细读——需要深入了解时点进单篇解析页,读摘要翻译、16 节 Paper Card 与审稿人评审;
- 感兴趣的先入库:用 Zotero 按本次
Discovery/下_ids.txt「按标识符添加」批量导入,去重与精筛在这一层完成; - 需要原文或深挖时,用文献工具按 DOI / PMID 二次获取,再走你自己的精读、检索与分析流程——下游工具(Zotero、agent、skill 等)自由接入。
Issue提醒:assign给你
网页部署:依次为周推送 - > 站点沉淀 - > 可全文搜索
文献WIKI积淀:仓库即WIKI理念,你可以畅所使用LLM WIKI/RAG技术来消化你的文献仓库!
历史文献回填,沉淀效果:
高通量简略搜索效果:
Warning
具体效果取决于RAG实现技术与你私有语料的质量,以及接入的LLM模型能力
尽量整合使用: 成熟的RAG技术栈工具+优化你私有文献语料的质量(可只保留全英文原文)+接入能力强的LLM模型,才能获得更好的效果
文献 RAG 不替你读、不替你写,它把你私有语料里"归档的文本内容"变成"可问、可溯源到原文段落"的记忆——省掉的是检索与引用核对的摩擦,留下的是你自己的判断
此处我们使用mcp-local-rag来尝试对文献仓库进行RAG式QA,
我们此处以预印本Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments为例展示效果
./node_modules/.bin/mcp-local-rag ingest /tmp/rag-t1/src/
Found 1 file(s) to ingest.
VectorStore initialized: /tmp/rag-t1/db
Parsed MD: /tmp/rag-t1/src/fulltext.md (58418 characters)
Embedder: First use detected. Initializing model (downloading ~90MB, may take 1-2 minutes)...
Embedder: Setting cache directory to "./models/"
Embedder: Loading model "Xenova/all-MiniLM-L6-v2" on device "cpu"...
Embedder: Model loaded successfully (device=cpu)
VectorStore: Skipping deletion as table does not exist
VectorStore: Created table "chunks"
VectorStore: FTS index "fts_index_v2" created successfully
VectorStore: Inserted 135 chunks
[1/1] /tmp/rag-t1/src/fulltext.md ... OK (135 chunks)
VectorStore connection closed
--- Ingest Summary ---
Succeeded: 1
Failed: 0
Total chunks: 135我们的问题是: "What is the role of the disordered N-terminus of 4E-BP2? "(4E-BP2 自身无序 N 端区域的功能)
RAG返回的文本引用即为该问题的回答,见下文每个条目的"text"
原始输出如下:
./node_modules/.bin/mcp-local-rag query "What is the role of the disordered N-terminus of 4E-BP2? "
VectorStore initialized: /tmp/rag-t1/db
Embedder: First use detected. Initializing model (downloading ~90MB, may take 1-2 minutes)...
Embedder: Setting cache directory to "./models/"
Embedder: Loading model "Xenova/all-MiniLM-L6-v2" on device "cpu"...
Embedder: Model loaded successfully (device=cpu)
[
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 16,
"text": "It demonstrates contacts between the C‐terminus of 4E‐BP2 and both binding sites on eIF4E from the crystal structure, corroborating previous observations of binding‐induced changes of this region of 4E‐BP2 (Smyth et al., 2022). The ensemble also showed significant intramolecular eIF4E contacts between a region of the N‐IDR that was reported to act as a 4E‐BP2 binding inhibitor and the canonical binding site (Abiko et al., 2007).",
"score": 0.23694531197987462,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 80,
"text": "This indicates that the secondary binding site of the 4E‐BP2 to eIF4E acts primarily in concert with the canonical binding site and that instances where only the secondary binding site is occupied are rare. Critically, it demonstrates that the crystal structure of the analogous 4E‐BP1:eIF4E complex presentation of Both at 100% is not representative of the ensemble present in solution, providing a clear basis for the accessibility of regulatory kinases within the context of the 3 nM complex. It also aligns with deletion studies that report that the canonical but not the secondary binding site can bind to eIF4E in isolation (Paku et al., 2012). The resulting ensemble also shows increased contacts between residues in the N‐IDR of eIF4E known to affect the binding of 4E‐BP2 and the canonical binding surface of 4E‐BP2.",
"score": 0.2580291783246409,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 49,
"text": "Figure 5a shows the assembly process of the 4E‐BP2:eIF4E conformational ensemble; starting from the crystal structure all residues that are not fixed in each specific instance are removed, then ensembles of the disordered loops or tails are screened with LDRS to select those consistent with the required geometry.",
"score": 0.2610657507110806,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 61,
"text": "Binding induces changes to the NMR intensity ratios, which decrease ~50%–60% at residues ~95–120, indicating that there are possible transient interactions with eIF4E (Lukhele et al., 2013). Electrostatics could explain the dynamic interaction mode of the C‐terminus of 4E‐BP2.",
"score": 0.2653705713278768,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 42,
"text": "This broadening is a signature of 4E‐BP2 conformational dynamics on the ~100 μs time scale that, remarkably, is not inhibited when it forms a complex with eIF4E, but consistent with a previous PET‐FCS study (Smyth et al., 2022). It is also interesting to note that segments A and C , which do not overlap at all with known binding site residues, compact rather than expand in the complex with eIF4E, suggesting that binding favors a larger number of intramolecular 4E‐BP2 contacts compared to the free state.",
"score": 0.27343735052867923,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 0,
"text": "# Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments ## Abstract\n Eukaryotic cap‐dependent translation initiation is regulated by binding of the predominantly folded eukaryotic initiation factor 4E (eIF4E) to the intrinsically disordered eIF4E binding proteins (4E‐BPs). Here, we report full‐length atomistic conformational ensembles generated by IDPConformerGenerator and optimized by X‐EISDv2 workflow for both apo 4E‐BP2, the neuronal 4E‐BP, and 4E‐BP2 in complex with eIF4E, using data from single‐molecule fluorescence and nuclear magnetic resonance (NMR), together with select coordinates from a 4E‐BP1:eIF4E crystal structure. Structural sampling within dynamic complexes is often underappreciated, with NMR and crystal structure data for 4E‐BP:eIF4E suggesting different degrees of structural heterogeneity. Our ensemble models validated by solution spectroscopy data enable comparison of free 4E‐BP2 and its complex with eIF4E. This shows a delocalization of contacts around canonical regions, which supports previous findings of unidirectional conditional occupancy of the binding sites. Two new contact regions emerged: one between the disordered N‐termini of eIF4E and 4E‐BP2, which may play an allosteric role in tuning the binding affinity, and the other between the C‐terminus of 4E‐BP2 and an extended region of eIF4E, which is consistent with the extended, dynamic binding interface that we reported previously. These results support a model of translation regulation in which the dynamic 4E‐BP2:eIF4E complex facilitates accessibility of regulatory sites of 4E‐BP2 when bound.",
"score": 0.2787180244922638,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 60,
"text": "Two new regions of intermolecular contact emerge: (RN1) contacts of the N‐IDR of eIF4E (residues 1–40) with the N‐terminus of 4E‐BP2 (residues 1–20), and (RN1) and sparse contacts between the C‐terminus of 4E‐BP2 (residues 110–120) and eIF4E, in particular residues 75–85 (Figure 6c,i). These enriched contacts are consistent with the binding‐induced changes at the C‐terminus of 4E‐BP2 measured by time‐resolved fluorescence spectroscopy, such as slower segmental dynamics and increased quenching dynamics (Smyth et al., 2022).",
"score": 0.3000167072069846,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 57,
"text": "As such, given the dynamic nature of the secondary binding site, the fraction of conformers where the secondary binding site is fixed to the crystal structure geometry is expected to be low, consistent with our optimized ensemble. ### Binding‐induced changes in the 4E‐BP2 ensemble\n 2D inter‐residue contact difference and normalized distance maps were constructed to compare the optimized ensembles of 4E‐BP2 in the bound and free states (Figure 6a). 4E‐BP2 when bound to eIF4E is expanded overall compared to the apo state, especially between residues 1–60 with 60–120 which reverses the contraction of these separations seen in Figure 3a. Topologically, more extended conformations of 4E‐BP2 that wrap around eIF4E enable it to interact with the extended binding interface on the surface of eIF4E (Lukhele et al., 2013). At the same time, there is compaction compared to the apo state observed for N‐ and C‐terminal stretches of ~20–30 residues. An increase in the number of intramolecular contacts within segments A and C in the bound state has been confirmed by contact map analysis (Table S4). Thus, the phospho‐regulatory 15RAIP18 motif is brought closer to the first two phosphorylation sites T37 and T46, potentially pre‐forming conformations that facilitate the initial phosphorylation steps.",
"score": 0.30860637301963434,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 54,
"text": "Interestingly, when all 4E‐BP2 residues in the crystal structure were fixed in the initial pool (Both), no subset of conformations that agreed with all restraints could be found (Smyth, 2024). In contrast, starting from the mixed pool containing Canonical, Secondary, and Both ensembles, subsets of conformations that satisfied all experimental restraints were found. This points to additional structural heterogeneity of the 4E‐BP2:eIF4E complex which is present in solution (Lukhele et al., 2013), but is not accurately represented in the crystal structure.",
"score": 0.3095265984739414,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
},
{
"filePath": "/tmp/rag-t1/src/fulltext.md",
"chunkIndex": 45,
"text": "To address this need, TIRF‐measured smFRET data were obtained on surface‐immobilized 4E‐BP2:eIF4E particles having a different fluorophore attached to each protein (Figure 4c). ### Generation and optimization of the eIF4E:4E‐BP2 ensemble\n In recent years, an increasing number of IDP structural ensembles have been deposited in the Protein Ensemble Database (Ghafouri et al., 2024), although most of these entries are for proteins in their free, unbound states. Calculating ensembles of complexes of disordered proteins bound to their (folded) molecular targets remains a challenging prospect, inhibiting understanding of sequence‐ensemble‐function relation for IDPs (Hadži et al., 2021). To model the eIF4E‐bound state of 4E‐BP2, we made use of existing structural information from a high resolution crystal structure of the 4E‐BP1:eIF4E complex (PDB ID: 4UED) (Peter et al., 2015). This structure was used because it resolves both canonical and secondary binding sites and because 4E‐BP1 and 4E‐BP2 have the same eIF4E binding mechanism with high sequence identity for binding residues (Fletcher et al., 1998). The overall sequence identity of 4E‐BP1 and 4E‐BP2 is 57%; for the fragment of 4E‐BP1 observed in the crystal structure, the sequence identity is 85%. The crystal structure has coordinates for 4E‐BP1 residues 50–83 and most of the eIF4E residues, only lacking the N‐terminal disordered tail (residues 1–32). The initial conformation pool of the 4E‐BP2:eIF4E complex was built starting from the x‐ray crystal structure (PDB ID: 4UED) (Peter et al., 2015), see Figure 1b,c. Coordinates for residues of 4E‐BP2 and eIF4E that were not resolved in the crystal structure were generated with IDPConformerGenerator (Teixeira et al., 2022), using the local disordered region sampling (LDRS) tool (Liu et al., 2023).",
"score": 0.3098001675474414,
"fileTitle": "Conformational ensembles of the disordered 4E‐BP2:eIF4E complex restrained by smFRET experiments",
"images": []
}
]VectorStore connection closed
逐条判断
- chunk16(0.237)❌ 讲4E-BP2 C端 + eIF4E N-IDR抑制作用,主体不是4E-BP2 N端,干扰。
- chunk80(0.258)❌ eIF4E的N-IDR和4E-BP2经典结合面的接触,是eIF4E的无序区,容易读错。
- chunk49(0.261)❌ 构象集合搭建流程,纯方法,无关生物学功能。
- chunk61(0.265)❌ 4E-BP2 C端动态相互作用,完全偏离N端。
- chunk42(0.273)
⚠️ 背景辅助 讲4E-BP2结合后整体构象动力学,A/C区段压缩;不专门讲N端功能,可以留作背景,不能用来直接回答问题。 - chunk0(0.279)✅ 核心摘要
one between the disordered N-termini of eIF4E and 4E-BP2, which may play an allosteric role in tuning the binding affinity 关键点:4E-BP2 N端 ↔ eIF4E N端 形成分子间接触,变构调节复合物亲和力。
- chunk60(0.300)✅ 最强直接证据
(RN1) contacts of the N‐IDR of eIF4E (residues 1–40) with the N‐terminus of 4E‐BP2 (residues 1–20) 精准定位:4E-BP2 N端 1–20aa 与 eIF4E N-IDR(1–40aa)形成RN1互作区域。这是原文专门定义的新接触区。
- chunk57(0.309)✅ 核心构象&调控功能 4E-BP2残基1–60(包含N端无序区)结合eIF4E后发生重排;N端20–30残基发生压缩,把磷酸调控基序 ¹⁵RAIP¹⁸拉近T37/T46磷酸位点,预组织构象,方便激酶磷酸化修饰。
- chunk54(0.310)❌ 晶体结构与溶液构象异质性,不涉及N端功能。
- chunk45(0.310)❌ 建模方法、PDB、IDPConformerGenerator,纯方法描述。
筛选总结
- ✅ 核心可用于回答的chunk:chunk0、chunk60、chunk57(这三段联合就能完整回答query)
⚠️ 可选背景:chunk42- ❌ 其余全部是干扰,直接丢弃,避免混淆eIF4E N-IDR和4E-BP2 N端
具体细节见 RAG-test-demo
❓ 常见问题
- 为什么按周而不是每日? 默认
window_days: 7+ cron 每周一运行;想改频次,改 config 与 cron 两处即可。 - 抓不全 / 有噪声怎么办? 平台 query 的召回与噪声实测都记录在各 query 上方的注释里,按注释微调;残余噪声由 Zotero 层也就是人工筛掉。
- 能不能不要 AI 精读? 把
config.yaml的llm.enable_card/enable_reviewer设为false,或直接不配置 LLM 密钥——检索、落盘、Issue 与站点仍照常运行,只是少了评分 / 翻译 / Card / 评审。 - 评分模型换了会影响什么? 评分能力变化会改变分档分布,进而影响
min_score门控通过率——换模型后应观察高分篇数量与 Card/Review 生成量是否失控。 - 能抓全文吗? 本模板只做元数据分发,全文仅用于喂给 LLM 精读、不落版权文件;原文按需走你已有的文献工具获取。
这是一个刻意保持通用的模板仓库:属于你的只有
config.yaml与prompts.yaml,脚本、模板、工作流、测试基本不必动。把它复制成「一个课题 topic = 一个独立仓库」,即可同时追踪多个研究方向。











