Uninote
Uninote
用户根目录
brdr
common
programming
docs
后端试题
问题讨论

doc-crawler

用 Node js 实现一个 CLI 爬虫,递归爬取页面内容。

参数

  • 接收一个 URL,必传
  • 可选参数 selector,-s,如果传入了,则仅保存 selector指定的内容
  • 范围限定参数 -p:仅爬取和此参数前缀匹配的内容,如果不传,则将主页面的 parent 作为此参数。可以指定多个
  • 转换为 Markdown 格式。
  • -c 限定数量,超过此数量则停止。

明确:1、应该先要做路径规范化,转为绝对路径,再判断是否是前缀包含关系。 2、整个页面的链接都要用来做递归,而非仅仅selector内的 3、广度优先 4、-v 输出详细日志,包括某个链接是递归还是被过滤了

other

change:--exclude 改为匹配原始的 url,添加 --exclude-last 匹配最终的 url(代替现在的 --exclude)

  • 很多网站有大量的多语言 url,需要前置过滤,否则下载量大,很慢

demo cmds

node doc-crawler https://learn.chatgpt.com/docs/configuration?site_locale=en -c 4 -s "#mainContent" -v -p https://learn.chatgpt.com/docs -p https://learn.chatgpt.com/codex
node doc-crawler https://learn.chatgpt.com/docs/configuration -s "#mainContent" -p https://learn.chatgpt.com/docs -p https://learn.chatgpt.com/codex --exclude =
node doc-crawler https://developer.chrome.com/docs/extensions/reference/manifest/web-accessible-resources -s "#main-content" -p https://developer.chrome.com/docs/extensions
node doc-crawler https://developer.mozilla.org/en-US/docs/Web/API/Document/readystatechange_event -s "#content" -p https://developer.mozilla.org/en-US/docs
node doc-crawler https://www.electronjs.org/docs/latest -s "main.docMainContainer_TBSr" -p https://www.electronjs.org/docs/latest
node doc-crawler https://react.dev/learn/ -s "main" -p https://react.dev/reference/react -p https://react.dev/learn
node doc-crawler https://babeljs.io/docs/usage -s "main"
node doc-crawler https://vuejs.org/guide/introduction.html -s "main"
node doc-crawler https://pnpm.io/motivation -s "main" --exclude '\//pnpm.io/(\w{2}|\w{2}-\w{2})/'
node doc-crawler https://nodejs.org/docs/latest-v26.x/api/ -s "main"
node doc-crawler https://docs.npmjs.com/getting-started -s "div.layout-module--Box_2--c505b"
node doc-crawler https://code.visualstudio.com/docs/debugtest/debugging -p https://code.visualstudio.com/docs -s "main"

design

#!/usr/bin/env node
// ============================================================
// doc-crawler.js — 零依赖 Node.js CLI 文档爬虫
//   递归爬取页面,按前缀限定范围,可选 CSS 选择器截取内容,
//   可选输出 Markdown,可限定爬取数量。
//
// 用法:
//   node doc-crawler.js <url> [选项]
//
// 参数:
//   <url>                   入口 URL(必传)
//   -s, --selector <css>    仅保存匹配 CSS 选择器的内容
//                           (支持 标签 / #id / .class / [属性] /
//                            空格后代 / > 子代 / 逗号分组)
//   -p, --prefix <前缀>     仅爬取 URL 以该前缀开头的页面,可重复指定;
//                           不传时取入口 URL 的父目录作为前缀
//   -x, --exclude <正则>    正则匹配的文件名被过滤(不保存),可重复指定;
//                           仅作用于最终生成文件的文件名(basename,含扩展名),
//                           不影响页面抓取与链接入队
//   -c, --count <数量>      最多爬取的页面数量,达到后立即停止
//   -l, --limit <n>         限速:每秒最多抓取 n 个页面,可为小数(默认不限)
//   --cache <目录>          指定缓存目录(默认 .cache),命中缓存时跳过网络请求
//   --no-cache              禁用缓存读取,但抓取结果仍会刷新缓存
//   -v, --verbose           输出详细日志(每条链接的递归/过滤状态)
//   -h, --help              显示帮助
//
// 产物:默认写入当前工作目录下 doc-crawler-output/<域名>/<URL 路径>/,
//       目录结构镜像 URL 路径,目录型 URL 落为 index.html / index.md。
//       抓取缓存写入 <缓存目录>(默认 .cache/<域名>/,JSON 格式)。
//
// 运行环境:Node.js 18+(使用内置 fetch),无第三方依赖。

队列 exclude(-last)/prefix filter

  • 比如a 最终变成 b,则 前缀匹配在两处都要做,b还要用 --exclude-last 匹配过滤一次,匹配了则丢弃,不获取页面
  • 不需要的资源不要进队列
  • exclude-last 解决重定向 url 变更问题

因为shell有转意处理,一开始有打印,确认正则是否传递正确

$ node doc-crawler https://pnpm.io/motivation -s "main" --exclude '\//pnpm.io/(\w{2}|\w{2}-\w{2})/'
    编译 -x/--exclude 正则:\//pnpm.io/(\w{2}|\w{2}-\w{2})/
  • windows cygwin:'/' 不用 '/' 这样写,但开头的需要,否则有诡异问题。

不再支持下载为 html,固定 md 格式

pic 图片

  • 默认连带图片一起下载,但仅下载 selector 范围内的图片。可以通过 --no-pic 禁用图片下载。
  • 如果图片已经存在,默认不重复下载。但如果开开启了 --no-cache 就重新下载。

react

E:\projects\pjkit\tools\doc-crawler\doc-crawler-output\react.dev\learn\writing-markup-with-jsx.md

问题:react.dev 使用 Next.js 图片优化器 /_next/image?url=%2Fimages%2F...&w=750&q=75。旧逻辑把整个查询串拼进文件名,导致:

文件名形如 image_url=%2F...&w=750&q=75
path.extname 识别出 .png&w=750&q=75 这种非法扩展名
Markdown 中图片路径不可读,图片无法正常显示
修复(doc-crawler.js):在 imgOutputPath 中检测 /_next/image URL,提取 url 查询参数作为实际图片路径命名。

others

非法字符处理:https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/Selectors/::placeholder

检索打包

https://chat.deepseek.com/a/chat/s/a95d657e-9346-4ae4-8330-e0d34277b108

  • prompt:我需要动态加载一个 js 文件,这个 js 文件放我自己的服务器,怎样实现?用这个问题给我生成一条检索命令,我要在本地 *.md 文档中查询

或者更稳的写法(处理带空格的文件名):

# E:\codex\test
rg -l -0 -g "*.md" "dynamic(ally)? load|load.*script" | xargs -0 cat

摘要抽取

比如将这种内容:

## Overview

### What are extensions?

Chrome extensions enhance the browsing experience by customizing the user interface, observing browser events, and modifying the web. Visit the [Chrome Web Store](https://chromewebstore.google.com/) for more examples of what extensions can do.

...

抽取为这种摘要:

## Overview

### What are extensions?

提取核心内容,忽略demo,引用链接等

chrome ext test

从文档中回答,并给出引用的位置,来源给出链接,每部分开头是来源,如:<!-- source: https://developer.chrome.com/docs/extensions/develop/concepts/activeTab --> 我需要动态加载一个 js 文件,这个 js 文件放我自己的服务器,怎样实现?

总结 chrome 插件开发

扩展更新了用户需要手动更新吗?

如何方便的重新加载扩展,不想每次都去扩展管理页面手动点击

优势:

  • 可以添加自己的注释,文档更新后还可以合并。
  • 翻译
  • 特定版本查询(快照)
  • 或者排查老版本干扰

Accept-Language header

developer.chrome.com 是 Google 站点,当请求不带 Accept-Language 头、也没有语言 Cookie 时,它会按出口 IP 的地理位置猜语言,并把 ?hl=xx 追加到重定向后的 URL 上。

codex

ds-mynote-总结

点赞(0) 阅读(7) 举报
目录
标题