doc-crawler
用 Node js 实现一个 CLI 爬虫,递归爬取页面内容。
参数
- 接收一个 URL,必传
- 可选参数 selector,-s,如果传入了,则仅保存 selector指定的内容
- 范围限定参数 -p:仅爬取和此参数前缀匹配的内容,如果不传,则将主页面的 parent 作为此参数。可以指定多个
- 转换为 Markdown 格式。
- -c 限定数量,超过此数量则停止。
明确:1、应该先要做路径规范化,转为绝对路径,再判断是否是前缀包含关系。 2、整个页面的链接都要用来做递归,而非仅仅selector内的 3、广度优先 4、-v 输出详细日志,包括某个链接是递归还是被过滤了
other
- https://chat.deepseek.com/a/chat/s/42bf37b5-ccee-4857-9c69-4085c276a334
- 不要搜索网页,用附件文档查询,结果给出对应的出处
change:--exclude 改为匹配原始的 url,添加 --exclude-last 匹配最终的 url(代替现在的 --exclude)
- 很多网站有大量的多语言 url,需要前置过滤,否则下载量大,很慢
demo cmds
node doc-crawler https://learn.chatgpt.com/docs/configuration?site_locale=en -c 4 -s "#mainContent" -v -p https://learn.chatgpt.com/docs -p https://learn.chatgpt.com/codex
node doc-crawler https://learn.chatgpt.com/docs/configuration -s "#mainContent" -p https://learn.chatgpt.com/docs -p https://learn.chatgpt.com/codex --exclude =
node doc-crawler https://developer.chrome.com/docs/extensions/reference/manifest/web-accessible-resources -s "#main-content" -p https://developer.chrome.com/docs/extensions
node doc-crawler https://developer.mozilla.org/en-US/docs/Web/API/Document/readystatechange_event -s "#content" -p https://developer.mozilla.org/en-US/docs
node doc-crawler https://www.electronjs.org/docs/latest -s "main.docMainContainer_TBSr" -p https://www.electronjs.org/docs/latest
node doc-crawler https://react.dev/learn/ -s "main" -p https://react.dev/reference/react -p https://react.dev/learn
node doc-crawler https://babeljs.io/docs/usage -s "main"
node doc-crawler https://vuejs.org/guide/introduction.html -s "main"
node doc-crawler https://pnpm.io/motivation -s "main" --exclude '\//pnpm.io/(\w{2}|\w{2}-\w{2})/'
node doc-crawler https://nodejs.org/docs/latest-v26.x/api/ -s "main"
node doc-crawler https://docs.npmjs.com/getting-started -s "div.layout-module--Box_2--c505b"
node doc-crawler https://code.visualstudio.com/docs/debugtest/debugging -p https://code.visualstudio.com/docs -s "main"
design
#!/usr/bin/env node
// ============================================================
// doc-crawler.js — 零依赖 Node.js CLI 文档爬虫
// 递归爬取页面,按前缀限定范围,可选 CSS 选择器截取内容,
// 可选输出 Markdown,可限定爬取数量。
//
// 用法:
// node doc-crawler.js <url> [选项]
//
// 参数:
// <url> 入口 URL(必传)
// -s, --selector <css> 仅保存匹配 CSS 选择器的内容
// (支持 标签 / #id / .class / [属性] /
// 空格后代 / > 子代 / 逗号分组)
// -p, --prefix <前缀> 仅爬取 URL 以该前缀开头的页面,可重复指定;
// 不传时取入口 URL 的父目录作为前缀
// -x, --exclude <正则> 正则匹配的文件名被过滤(不保存),可重复指定;
// 仅作用于最终生成文件的文件名(basename,含扩展名),
// 不影响页面抓取与链接入队
// -c, --count <数量> 最多爬取的页面数量,达到后立即停止
// -l, --limit <n> 限速:每秒最多抓取 n 个页面,可为小数(默认不限)
// --cache <目录> 指定缓存目录(默认 .cache),命中缓存时跳过网络请求
// --no-cache 禁用缓存读取,但抓取结果仍会刷新缓存
// -v, --verbose 输出详细日志(每条链接的递归/过滤状态)
// -h, --help 显示帮助
//
// 产物:默认写入当前工作目录下 doc-crawler-output/<域名>/<URL 路径>/,
// 目录结构镜像 URL 路径,目录型 URL 落为 index.html / index.md。
// 抓取缓存写入 <缓存目录>(默认 .cache/<域名>/,JSON 格式)。
//
// 运行环境:Node.js 18+(使用内置 fetch),无第三方依赖。
队列 exclude(-last)/prefix filter
- 比如a 最终变成 b,则 前缀匹配在两处都要做,b还要用 --exclude-last 匹配过滤一次,匹配了则丢弃,不获取页面
- 不需要的资源不要进队列
- exclude-last 解决重定向 url 变更问题
因为shell有转意处理,一开始有打印,确认正则是否传递正确
$ node doc-crawler https://pnpm.io/motivation -s "main" --exclude '\//pnpm.io/(\w{2}|\w{2}-\w{2})/'
编译 -x/--exclude 正则:\//pnpm.io/(\w{2}|\w{2}-\w{2})/
- windows cygwin:'/' 不用 '/' 这样写,但开头的需要,否则有诡异问题。
不再支持下载为 html,固定 md 格式
pic 图片
- 默认连带图片一起下载,但仅下载 selector 范围内的图片。可以通过 --no-pic 禁用图片下载。
- 如果图片已经存在,默认不重复下载。但如果开开启了 --no-cache 就重新下载。
react
E:\projects\pjkit\tools\doc-crawler\doc-crawler-output\react.dev\learn\writing-markup-with-jsx.md
问题:react.dev 使用 Next.js 图片优化器 /_next/image?url=%2Fimages%2F...&w=750&q=75。旧逻辑把整个查询串拼进文件名,导致:
文件名形如 image_url=%2F...&w=750&q=75
path.extname 识别出 .png&w=750&q=75 这种非法扩展名
Markdown 中图片路径不可读,图片无法正常显示
修复(doc-crawler.js):在 imgOutputPath 中检测 /_next/image URL,提取 url 查询参数作为实际图片路径命名。
others
非法字符处理:https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/Selectors/::placeholder
检索打包
https://chat.deepseek.com/a/chat/s/a95d657e-9346-4ae4-8330-e0d34277b108
- prompt:我需要动态加载一个 js 文件,这个 js 文件放我自己的服务器,怎样实现?用这个问题给我生成一条检索命令,我要在本地 *.md 文档中查询
或者更稳的写法(处理带空格的文件名):
# E:\codex\test
rg -l -0 -g "*.md" "dynamic(ally)? load|load.*script" | xargs -0 cat
摘要抽取
比如将这种内容:
## Overview
### What are extensions?
Chrome extensions enhance the browsing experience by customizing the user interface, observing browser events, and modifying the web. Visit the [Chrome Web Store](https://chromewebstore.google.com/) for more examples of what extensions can do.
...
抽取为这种摘要:
## Overview
### What are extensions?
提取核心内容,忽略demo,引用链接等
chrome ext test
从文档中回答,并给出引用的位置,来源给出链接,每部分开头是来源,如:<!-- source: https://developer.chrome.com/docs/extensions/develop/concepts/activeTab --> 我需要动态加载一个 js 文件,这个 js 文件放我自己的服务器,怎样实现?
总结 chrome 插件开发
扩展更新了用户需要手动更新吗?
如何方便的重新加载扩展,不想每次都去扩展管理页面手动点击
优势:
- 可以添加自己的注释,文档更新后还可以合并。
- 翻译
- 特定版本查询(快照)
- 或者排查老版本干扰
Accept-Language header
developer.chrome.com 是 Google 站点,当请求不带 Accept-Language 头、也没有语言 Cookie 时,它会按出口 IP 的地理位置猜语言,并把 ?hl=xx 追加到重定向后的 URL 上。
- 有的站点需要把其他语言的链接过滤掉,eg:队列 exclude(-last)/prefix filter
