Extraire un Markdown propre et du texte brut de tout site web, optimisé pour l'ingestion IA, les pipelines RAG et les fenêtres de contexte LLM. L'extraction du contenu principal façon Readability supprime navigation, footers, sidebars et publicités, pour que votre IA ne reçoive que le contenu utile. Mode flat fetch (profondeur=0) pour listes d'URL ou crawl complet jusqu'à profondeur 5. Jusqu'à 20 workers parallèles.
# Récupérer une liste de pages de documentation (sans crawl) curl -X POST "https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/runs?token=YOUR_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "startUrls": [ "https://docs.example.com/api/overview", "https://docs.example.com/api/authentication" ], "maxDepth": 0, "extractMainContent": true }' # Ou utiliser avec agents IA via MCP : # https://mcp.apify.com?tools=santamaria-automations/website-content-crawler
| Champ | Type | Description |
|---|---|---|
| url | string | URL de la page crawlée |
| title | string | Titre (og:title ou title HTML) |
| description | string | Meta description |
| markdown | string | Markdown propre, jusqu'à 50 000 caractères |
| text | string | Texte brut, jusqu'à 10 000 caractères |
| word_count | integer | Nombre de mots du texte brut |
| content_type | string | article, blog, documentation, generic |
| depth | integer | Profondeur de crawl (0 = URL de départ) |
| status_code | integer | Code HTTP |
| scraped_at | string | Horodatage ISO 8601 UTC |
Cliquez sur Ouvrir sur Apify ci-dessus pour lancer website-content-crawler dans votre navigateur – sans code, sans installation. Dans la console Apify, vous obtenez :
Pour lancer cet actor depuis votre propre code, il vous faut un jeton API Apify. Environ une minute :
Quota gratuit : Apify crédite votre compte de 5$ d'usage plateforme par mois, sans carte bancaire. Assez pour tester n'importe quel actor — à 0,001$ par résultat, c'est environ 5 000 résultats gratuits chaque mois.
Lancez cet actor depuis n'importe quel langage via l'API REST Apify. Remplacez YOUR_TOKEN par votre jeton API et adaptez le JSON d'entrée.
curl -X POST "https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{}'// npm install apify-client
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('santamaria-automations/website-content-crawler').call({});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);# pip install apify-client
from apify_client import ApifyClient
client = ApifyClient('YOUR_TOKEN')
run = client.actor('santamaria-automations/website-content-crawler').call(run_input={})
items = list(client.dataset(run['defaultDatasetId']).iterate_items())
print(items)// dotnet add package Apify.Client
using Apify.Client;
var client = new ApifyClient("YOUR_TOKEN");
var run = await client.Actor("santamaria-automations/website-content-crawler").CallAsync(new { });
var items = await client.Dataset(run.DefaultDatasetId).ListItemsAsync();// Maven: com.apify:apify-client
import com.apify.client.ApifyClient;
ApifyClient client = new ApifyClient("YOUR_TOKEN");
ActorRun run = client.actor("santamaria-automations/website-content-crawler").call(Map.of());
List<Map<String,Object>> items = client.dataset(run.getDefaultDatasetId()).listItems();Cet actor est disponible sur le serveur MCP Apify, vous pouvez donc le piloter depuis n'importe quel client IA compatible MCP – Claude Desktop, Claude.ai, Cursor, VS Code, LangChain, LlamaIndex ou un agent maison – sans écrire de code.
https://mcp.apify.com?tools=santamaria-automations/website-content-crawlerExemple de prompt une fois connecté :
"Utilise website-content-crawler pour lancer un scrape et donne-moi les résultats sous forme de tableau."
Les clients avec découverte dynamique des outils (Claude.ai, VS Code) reçoivent le schéma d'entrée complet automatiquement via add-actor.
Déclenchez cet actor depuis votre outil d'automatisation. Toutes les plateformes ci-dessous peuvent appeler l'API Apify en quelques clics — sans code.
@apify/n8n-nodes-apify).santamaria-automations/website-content-crawler.https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN.application/json.https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN, méthode POST, body raw JSON.https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN.steps.http.$return_value dans les étapes suivantes.https://api.apify.com/v2/acts/santamaria-automations~website-content-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN, body type JSON.