PDF Text Extractor - Bulk PDF to Text & Metadata

The NanoScrape Pdf Extractor Scraper is an Apify actor that extracts structured web data from Pdf Extractor. It returns 12 fields per result including url, file_size_bytes, success, error, page_count, priced at $5/1k pdf, with no API key or monthly subscription required.

Extract text and metadata from PDF URLs. Returns page content, page count, author, title, and scanned/encrypted flags. Bulk processing supported. Pay-per-result.

Open on Apify →
Pricing
$5.00 / 1,000 pdf
Runtime
Cloud (Apify)
Proxy
Datacenter

What you can scrape — 12 fields per result

FieldTypeExampleGroupFill %
urlstringhttps://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdfCore100%
scraped_atstring2026-09-15T06:57:03ZMetadata100%
file_size_bytesinteger13264Other100%
successbooleantrueOther100%
errorstringOther
page_countinteger1Other100%
textstring Dummy PDF file Other100%
text_lengthinteger16Other100%
metadataobject{"title": "", "author": "Evangelos Vlachogiannis", "subject": "", "keywords": []Other100%
is_encryptedbooleanfalseOther100%
is_scannedbooleantrueOther100%
needs_ocrbooleantrueOther100%

Input example — showcase (full detail)

{
  "pdfUrls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://www.irs.gov/pub/irs-pdf/f1040.pdf"
  ],
  "maxFileSizeMB": 50,
  "extractMetadata": true,
  "extractText": true,
  "timeoutSeconds": 60
}

Output example — real dataset item (PII scrubbed)

{
  "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
  "file_size_bytes": 13264,
  "success": true,
  "error": "",
  "page_count": 1,
  "text": "\nDummy PDF file\n",
  "text_length": 16,
  "metadata": {
    "title": "",
    "author": "Evangelos Vlachogiannis",
    "subject": "",
    "keywords": [],
    "creator": "Writer",
    "producer": "OpenOffice.org 2.1",
    "creation_date": "D:20070223175637+02'00'",
    "modification_date": ""
  },
  "is_encrypted": false,
  "is_scanned": true,
  "needs_ocr": true,
  "scraped_at": "2026-09-15T06:57:03Z"
}

Cost math

Priced at $0.0050 per pdf. 10 pdfs ≈ $0.05, 100 ≈ $0.50, 1,000 ≈ $5.00. The first runs land inside Apify's $5/mo free-tier credit — pay only for what you extract, no monthly subscription.

Integrations

Run this actor directly on Apify (no code)

Click Open on Apify above to run pdf-extractor in your browser - no code, no install. In the Apify console you get:

Get your Apify API token

To run this actor from your own code you need an Apify API token. Get one in about a minute:

  1. Sign up for a free Apify account (Google, GitHub, or email).
  2. Go to Settings → Integrations → API tokens.
  3. Click Create a new API token. Copy it and keep it secret.

Free tier: Apify credits your account with $5 of platform usage every month, no credit card required. Enough to test any actor meaningfully - at $0.001 per result on typical scrapers, that is roughly 5,000 results for free every month.

Call from code

Run this actor from any language via the Apify REST API. Replace YOUR_TOKEN with your API token and adapt the input JSON to your needs.

curl -X POST "https://api.apify.com/v2/acts/santamaria-automations~pdf-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{}'
// npm install apify-client
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('santamaria-automations/pdf-extractor').call({});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
# pip install apify-client
from apify_client import ApifyClient

client = ApifyClient('YOUR_TOKEN')
run = client.actor('santamaria-automations/pdf-extractor').call(run_input={})
items = list(client.dataset(run['defaultDatasetId']).iterate_items())
print(items)
// dotnet add package Apify.Client
using Apify.Client;

var client = new ApifyClient("YOUR_TOKEN");
var run = await client.Actor("santamaria-automations/pdf-extractor").CallAsync(new { });
var items = await client.Dataset(run.DefaultDatasetId).ListItemsAsync();
// Maven: com.apify:apify-client
import com.apify.client.ApifyClient;

ApifyClient client = new ApifyClient("YOUR_TOKEN");
ActorRun run = client.actor("santamaria-automations/pdf-extractor").call(Map.of());
List<Map<String,Object>> items = client.dataset(run.getDefaultDatasetId()).listItems();

Use with AI agents (MCP)

This actor is available on the Apify MCP server, so you can drive it from any MCP-compatible AI client - Claude Desktop, Claude.ai, Cursor, VS Code, LangChain, LlamaIndex, or a custom agent - without writing any code.

https://mcp.apify.com?tools=santamaria-automations/pdf-extractor

Example prompt once connected:

"Use <code>pdf-extractor</code> to run a scrape on my target list and give me the results as a table."

Clients that support dynamic tool discovery (Claude.ai, VS Code) receive the full input schema automatically via add-actor.

Use with no-code platforms

Trigger this actor from your favorite automation tool. Every platform below can call the Apify API in a few clicks - no code required.

  1. Add the n8n Apify node (installed by default on n8n Cloud; on self-hosted install @apify/n8n-nodes-apify).
  2. Set Actor to santamaria-automations/pdf-extractor.
  3. Choose operation Run Actor and get dataset, paste your input JSON, and connect a downstream node (Google Sheets, Postgres, webhook, ...).
  1. Create a new Zap with any trigger.
  2. Add a Webhooks by Zapier POST action to https://api.apify.com/v2/acts/santamaria-automations~pdf-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN.
  3. Body: your input JSON. Content-Type: application/json.
  1. Create a new scenario.
  2. Add an HTTP module: URL https://api.apify.com/v2/acts/santamaria-automations~pdf-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN, method POST, body raw JSON.
  3. Parse the response with a JSON module to iterate items downstream.
  1. Create a new workflow.
  2. Add an HTTP Request step: POST to https://api.apify.com/v2/acts/santamaria-automations~pdf-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN.
  3. Access steps.http.$return_value in subsequent steps.
  1. Install the API Connector plugin.
  2. Add a new API: POST https://api.apify.com/v2/acts/santamaria-automations~pdf-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN, body type JSON.
  3. Initialize the call with your input, then reference the response array in workflow actions.

FAQ

How much does the pdf-extractor cost?

$5.00 / 1,000 pdf, billed per result on Apify. Apify's $5/month free tier covers the first runs. No monthly subscription — you only pay for what you extract.

What fields does it return?

12 fields per result. Full field catalog is listed in the 'What you can scrape' table above — includes core identifiers, descriptive text, dates, and any category-specific data.

Is scraping this site legal?

Scraping publicly available data is generally permitted in most jurisdictions, but GDPR applies to any personal data you collect. Review your local rules and consult a lawyer for commercial use.

How fresh is the data?

Every run pulls live data at execution time. Set up a schedule (daily/hourly/weekly cron) in the Apify console to keep your dataset current automatically.

Related actors in this category

Open on Apify →