跳转到正文
报告库
用途分类 / 其他用途

Firecrawl Research Index Skill 安全审计

作者说它能做什么(原文)

Find the papers that answer a research query in Firecrawl's research paper index — a corpus of paper abstracts whose largest share is biomedical and life-science literature (PubMed, bioRxiv, medRxiv), alongside arXiv preprints in CS, physics, and math — using semantic search, semantic and structural expansion, and in-body verification. Use this skill for literature-finding and paper-retrieval task

第三方安全检查结论

发现安全风险

已检查文件
1
发现的风险
2
会不会运行危险命令?检查是否下载程序后直接运行、让他人远程控制电脑,或藏起要运行的命令。未发现风险
会不会泄露文件和密钥?检查是否发送含密码或密钥的文件,以及代码里是否直接写了密钥。发现 1 项风险
中风险

研究查询和正文核实问题会发送给 Firecrawl

原文依据:4 处
发现了什么

搜索命令直接接受查询文本,正文读取命令还接受具体问题。该 Skill 明确覆盖临床、药物、基因和疾病主题,因此用户若在查询中写入个人病情、未公开研究方向、药物项目代号或其他机密背景,这些内容会进入外部工具调用。

为什么需要注意

外部服务可能接触并记录用户的敏感健康兴趣、研究意图或商业研发信息。所示文件没有说明其保留、访问或再利用规则。

该 Skill 主动要求把用户的查询文本,以及针对单篇论文的具体核实问题,提交给 Firecrawl 的 MCP 工具或 CLI。由于适用范围包括临床、药物、基因和疾病问题,用户若把个人病情、未公开课题或项目代号写进查询,这些内容可能被发送到外部服务。来源没有说明数据保留或隐私边界。用户可要求作者说明 Firecrawl 的数据处理政策,并限制代理仅提交去标识化、非机密的检索词。

SKILL.md:12来自说明文档打开原文件
Paper abstracts, with full text reachable per paper. The largest share of the corpus is **biomedical and life-science** literature — **PubMed** journal articles plus **bioRxiv** and **medRxiv** preprints — so clinical, drug, gene, disease, epidemiology, and public-health questions are in scope. **arXiv** preprints cover computer science, physics, and mathematics. Coverage outside those sources is thinner: a paper that exists only behind a publisher paywall or in a niche venue may not be indexed, and the general web tools below are the fallback when it isn't.
查看另外 3 个位置
SKILL.md:18来自说明文档打开原文件
- MCP: **`firecrawl_research_search_papers(query, k?)`**  CLI: **`firecrawl research search-papers <query> [--k <number>]`**  Semantic (HyDE) search over **abstracts**. The natural first move for almost any query.  If results look thin or all-alike, re-run with a different framing (sibling domain, rival method, dataset/benchmark name) rather than giving up.
SKILL.md:23来自说明文档打开原文件
- MCP: **`firecrawl_research_related_papers(seed_ids, intent, mode?, k?)`**  CLI: **`firecrawl research related-papers <seedIds...> --intent <intent> [--mode <similar|citers|references>] [--k <number>]`**  Semantic and structural expansion, ranked to your `intent`.
SKILL.md:35来自说明文档打开原文件
- MCP: **`firecrawl_research_read_paper(id, question)`**  CLI: **`firecrawl research read-paper <id> --question <question>`**  In-body passages of **one** paper, to verify a load-bearing constraint (a method actually used, a score actually reported, an affiliation, what a paper compares to).  Use it to settle a specific doubt, not on everything.
会不会删除文件或一直在后台运行?检查是否大范围删除文件、改写磁盘,或设置自动启动。未发现风险
会不会绕过安全保护?检查是否跳过网站安全验证、开放过多文件权限,或取消操作前的确认。未发现风险
会不会误导 AI 或隐藏内容?检查工作说明是否要求 AI 忽略你的指令、干扰检查结果,或夹带看不见的文字。发现 1 项风险
中风险

“有疑问就收录”的策略可能让证据不足的论文影响决策

原文依据:3 处
发现了什么

该 Skill 指示代理保留大多数“可能相关”的论文,并把正文核实主要用于排除,而不是要求候选先被证实。索引还包含 bioRxiv、medRxiv 和 arXiv 预印本;指令没有要求区分同行评审状态、研究质量、撤稿或偏倚风险。

为什么需要注意

在临床、药物或公共卫生等高影响场景中,结果清单可能混入只具主题相似性、未经验证或质量较弱的论文。若用户把“被列出”误当成“证据支持”,可能形成不可靠的医疗、研究或商业判断。

这不是单纯的召回率建议:指令明确要求保留大多数“可能相关”的论文,并主要用正文核实来排除候选。语料同时包含 bioRxiv、medRxiv 和 arXiv 预印本,但可见指令没有要求标注同行评审、撤稿或研究质量。因此在医疗或其他高风险决策中,相关性较弱或未经评审的材料可能与更可靠证据并列,影响判断。用户可要求输出明确标注证据级别、发表状态及核实限制,并在决策用途下采用更严格纳入标准。

SKILL.md:12来自说明文档打开原文件
Paper abstracts, with full text reachable per paper. The largest share of the corpus is **biomedical and life-science** literature — **PubMed** journal articles plus **bioRxiv** and **medRxiv** preprints — so clinical, drug, gene, disease, epidemiology, and public-health questions are in scope. **arXiv** preprints cover computer science, physics, and mathematics. Coverage outside those sources is thinner: a paper that exists only behind a publisher paywall or in a niche venue may not be indexed, and the general web tools below are the fallback when it isn't.
查看另外 2 个位置
SKILL.md:66来自说明文档打开原文件
- **Query shape and subject field are separate.** A clinical-trial question and a machine-learning question take the same shapes above; what differs is only which source the hits come from. Don't send a biomedical or life-science query to the open web on the assumption the corpus is arXiv-only — PubMed, bioRxiv, and medRxiv are the largest part of what `search_papers` reads.- **When in doubt, include.** For any topic / method / comparison question, return the relevant _family_, not just the single best match — err toward keeping a plausibly-relevant paper rather than dropping it. The neighboring methods are part of a good answer; don't reason close work out just because one paper is the most exact match.- **Follow the literature, and keep what you find.** The seminal source, the competing methods, the close neighbors are usually a hop away — use `related_papers`, and _include_ them, not just the first hit. Stopping at one good result is the most common way to leave the reader with half an answer.
SKILL.md:68来自说明文档打开原文件
- **Follow the literature, and keep what you find.** The seminal source, the competing methods, the close neighbors are usually a hop away — use `related_papers`, and _include_ them, not just the first hit. Stopping at one good result is the most common way to leave the reader with half an answer.- **Verify to exclude, not to gatekeep.** Use `read_paper` to rule a paper _out_ when a hard constraint clearly fails (wrong org/author, doesn't actually report the score). When a paper is plausibly relevant, lean toward keeping it rather than demanding proof.- **Only drop the clearly off-topic.** Don't pad with papers you're confident are unrelated — but that's a high bar; most plausibly-relevant work should make the cut.
会不会偷偷改推广链接或收款方?检查是否强制替换推广链接或收款对象,同时要求隐瞒更改。未发现风险

Skill 逻辑拆解

5 个说明模块

该 Skill 通过 Firecrawl 的 MCP 工具或 CLI,把研究查询发送到论文摘要语义搜索;还可以沿相似论文、引用和参考文献扩展结果。

查看原文
SKILL.md:18来自说明文档打开原文件
- MCP: **`firecrawl_research_search_papers(query, k?)`**  CLI: **`firecrawl research search-papers <query> [--k <number>]`**  Semantic (HyDE) search over **abstracts**. The natural first move for almost any query.  If results look thin or all-alike, re-run with a different framing (sibling domain, rival method, dataset/benchmark name) rather than giving up.
SKILL.md:23来自说明文档打开原文件
- MCP: **`firecrawl_research_related_papers(seed_ids, intent, mode?, k?)`**  CLI: **`firecrawl research related-papers <seedIds...> --intent <intent> [--mode <similar|citers|references>] [--k <number>]`**  Semantic and structural expansion, ranked to your `intent`.  This reaches papers semantic search _cannot_, and it's how you turn one good hit into the rest of a set.  `mode=similar` → niche siblings; `citers` → who uses/builds on the seeds; `references` → what they build on / compare against.

对于需要核实论文正文的问题,该 Skill 会把论文 ID 和具体问题交给 `read_paper`,由其返回相关正文片段;它明确建议只在需要解决具体疑点时使用。

查看原文
SKILL.md:35来自说明文档打开原文件
- MCP: **`firecrawl_research_read_paper(id, question)`**  CLI: **`firecrawl research read-paper <id> --question <question>`**  In-body passages of **one** paper, to verify a load-bearing constraint (a method actually used, a score actually reported, an affiliation, what a paper compares to).  Use it to settle a specific doubt, not on everything.

遇到排行榜、最大或最流行等问题时,该 Skill 会转向普通网页搜索或抓取,再把排名条目映射回论文;这条路径不局限于论文索引。

查看原文
SKILL.md:46来自说明文档打开原文件
- MCP: **`firecrawl_search(query)` / `firecrawl_scrape(url)`**  CLI: **`firecrawl search <query>` / `firecrawl scrape <url>`**  General **web** search and page fetch, for facts that don't live in paper abstracts: benchmark **leaderboards**, rankings, "who scores best / is largest / is most used."  Find the ranking on the web, then map the top entries back to papers with `search_papers`.  Reach for these only when the corpus can't answer the question on its own.

该 Skill 的结果选择策略偏向返回较完整的相关论文集合,只有明确离题或不满足硬约束时才排除候选。

查看原文
SKILL.md:66来自说明文档打开原文件
- **Query shape and subject field are separate.** A clinical-trial question and a machine-learning question take the same shapes above; what differs is only which source the hits come from. Don't send a biomedical or life-science query to the open web on the assumption the corpus is arXiv-only — PubMed, bioRxiv, and medRxiv are the largest part of what `search_papers` reads.- **When in doubt, include.** For any topic / method / comparison question, return the relevant _family_, not just the single best match — err toward keeping a plausibly-relevant paper rather than dropping it. The neighboring methods are part of a good answer; don't reason close work out just because one paper is the most exact match.- **Follow the literature, and keep what you find.** The seminal source, the competing methods, the close neighbors are usually a hop away — use `related_papers`, and _include_ them, not just the first hit. Stopping at one good result is the most common way to leave the reader with half an answer.- **Verify to exclude, not to gatekeep.** Use `read_paper` to rule a paper _out_ when a hard constraint clearly fails (wrong org/author, doesn't actually report the score). When a paper is plausibly relevant, lean toward keeping it rather than demanding proof.- **Only drop the clearly off-topic.** Don't pad with papers you're confident are unrelated — but that's a high bar; most plausibly-relevant work should make the cut.
从这里开始 · 工作说明SKILL.md
firecrawl-research-index
连线表示工作说明包含的模块,不是实际运行顺序。点击模块可查看原文。
文件与检查记录1 个文件

检查范围与遗漏

逐文件查看涉及的内容

下方列出本次涉及的原文范围;纳入检查不代表已查清所有问题。

  • SKILL.md已纳入全文

这份报告只针对上方版本。我们看了拿到的代码和说明文件,没有实际运行 Skill,也没有检查它另外安装的软件包。因此,这不是“保证安全”的承诺;换了版本或使用环境,结果也可能不同。

  • SKILL.md工作说明

代码和说明中提到的操作

连接外部网站
SKILL.md:73来自说明文档打开原文件
- [firecrawl-build-search](https://github.com/firecrawl/skills/tree/main/skills/build/firecrawl-build-search) — building the paper index into an app instead of querying it here
读取了多少行
74
文件校验值(用于核对版本)
7fc95ca9c37654337d315c53f2aad3ea66892b5f4615e5c2f9c3db8a07eec104