网站内容爬虫:面向 AI 与 RAG 管线的纯净 Markdown

从任意网站提取纯净 Markdown 与纯文本,针对 AI 摄取、RAG 管线与 LLM 上下文窗口优化。Readability 风格主体内容提取去除导航、页脚、侧栏与广告,让您的 AI 只获得有价值的内容。Flat fetch(深度=0)适用于 URL 列表,或最大深度 5 整站爬取。最多 20 个并行 worker。

在 Apify 打开 → 立即试用
定价
$1/千页 + $0.25 启动
内存
128MB
输出
Markdown + 文本
并发
最多 20 worker
爬取深度
0 至 5 层
引擎
HTTP-only Go

每页可获得

主要用例

API 示例

# 直接抓取一组文档页(不爬取)
curl -X POST "https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/runs?token=YOUR_TOKEN" \
 -H "Content-Type: application/json" \
 -d '{
 "startUrls": [
 "https://docs.example.com/api/overview",
 "https://docs.example.com/api/authentication"
 ],
 "maxDepth": 0,
 "extractMainContent": true
 }'

# 或通过 MCP 与 AI 智能体一起使用:
# https://mcp.apify.com?tools=santamaria-automations/website-content-crawler

集成

输出字段

字段类型描述
urlstring已爬取页面 URL
titlestring页面标题
descriptionstringMeta description
markdownstring纯净 Markdown,最多 50,000 字符
textstring纯文本,最多 10,000 字符
word_countinteger纯文本字数
content_typestringarticle、blog、documentation、generic
depthinteger爬取深度(0 = 起始 URL)
status_codeintegerHTTP 状态码
scraped_atstringISO 8601 UTC 时间戳

相关 Actor

直接在 Apify 上运行此 actor(无需代码)

点击上方 在 Apify 上打开,在浏览器中运行 website-content-crawler — 无需代码,无需安装。在 Apify 控制台,您可以获得:

获取您的 Apify API 令牌

要从自己的代码运行此 actor,您需要一个 Apify API 令牌。大约一分钟完成:

  1. 免费注册 Apify 账户(Google、GitHub 或电子邮件)。
  2. 前往 设置 → 集成 → API 令牌
  3. 点击 创建新 API 令牌。复制并妥善保存。

免费额度:Apify 每月为您的账户提供 5 美元平台使用额度,无需信用卡。足以有意义地测试任何 actor — 按典型 scraper 每条结果 0.001 美元计算,每月大约 5,000 条免费结果。

从代码调用

通过 Apify REST API 从任何语言运行此 actor。将 YOUR_TOKEN 替换为您的 API 令牌,并根据需要调整输入 JSON。

curl -X POST "https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{}'
// npm install apify-client
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('santamaria-automations/website-content-crawler').call({});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
# pip install apify-client
from apify_client import ApifyClient

client = ApifyClient('YOUR_TOKEN')
run = client.actor('santamaria-automations/website-content-crawler').call(run_input={})
items = list(client.dataset(run['defaultDatasetId']).iterate_items())
print(items)
// dotnet add package Apify.Client
using Apify.Client;

var client = new ApifyClient("YOUR_TOKEN");
var run = await client.Actor("santamaria-automations/website-content-crawler").CallAsync(new { });
var items = await client.Dataset(run.DefaultDatasetId).ListItemsAsync();
// Maven: com.apify:apify-client
import com.apify.client.ApifyClient;

ApifyClient client = new ApifyClient("YOUR_TOKEN");
ActorRun run = client.actor("santamaria-automations/website-content-crawler").call(Map.of());
List<Map<String,Object>> items = client.dataset(run.getDefaultDatasetId()).listItems();

与 AI 代理配合使用 (MCP)

此 actor 可在 Apify MCP 服务器上使用,因此您可以从任何兼容 MCP 的 AI 客户端驱动它 — Claude Desktop、Claude.ai、Cursor、VS Code、LangChain、LlamaIndex 或自定义 agent — 无需编写代码。

https://mcp.apify.com?tools=santamaria-automations/website-content-crawler

连接后的示例提示:

"使用 website-content-crawler 运行一次 scrape,并将结果以表格形式返回给我。"

支持动态工具发现的客户端(Claude.ai、VS Code)会通过 add-actor 自动接收完整的输入 schema。

与无代码平台配合使用

从您喜欢的自动化工具触发此 actor。下面的每个平台都可以点几下就调用 Apify API — 无需代码。

  1. 添加 n8n Apify 节点(n8n Cloud 默认已安装;self-hosted 请安装 @apify/n8n-nodes-apify)。
  2. Actor 设置为 santamaria-automations/website-content-crawler
  3. 选择操作 Run Actor and get dataset,粘贴您的输入 JSON,然后连接下游节点(Google Sheets、Postgres、Webhook 等)。
  1. 使用任意触发器创建新 Zap。
  2. 添加一个 Webhooks by Zapier POST 操作到 https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN
  3. Body:您的输入 JSON。Content-Type:application/json
  1. 创建新场景。
  2. 添加 HTTP 模块:URL https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN,方法 POST,body raw JSON。
  3. 使用 JSON 模块解析响应以在下游迭代 items。
  1. 创建新工作流。
  2. 添加 HTTP Request 步骤:POST 到 https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN
  3. 在后续步骤中访问 steps.http.$return_value
  1. 安装 API Connector 插件。
  2. 添加新 API:POST https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN,body 类型 JSON。
  3. 用您的输入初始化 call,然后在 workflow 操作中引用响应数组。
在 Apify 打开 → 立即试用(提供免费额度)