Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions api-reference/errors.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@ All non-2xx responses return JSON with a top-level `success: false` and a string
| 402 | `Payment Required: Insufficient credits` | Plan credits are exhausted or billing is not configured. | Turn on pay-as-you-go, or upgrade your plan. | No |
| 403 | `Forbidden` | Key lacks permission for this endpoint or feature. | Use a key with the required scope, or upgrade the plan that gates this feature. | No |
| 403 | `SCRAPE_PROMPT_INJECTION_DETECTED` | JSON mode with `checkPromptInjection: true` detected a prompt injection attempt in the scraped page content, so extraction was aborted. | Inspect the page content manually. If it is a false positive, retry without `checkPromptInjection`. See [Prompt injection detection](/features/llm-extract#prompt-injection-detection). | No |
| 403 | `THIRD_PARTY_DATA_TERMS_REQUIRED` | An [Alexandria](/features/alexandria) provider needs your organization to accept its terms before the request can run. The provider did not run. The response has this value in `code`, and includes `requiresAction.url`. | Send `requiresAction.url` to an org admin, who reviews and accepts the terms there. Agents must not accept terms on their own. After acceptance, send the same request again. | After acceptance |
| 404 | `Not Found` | The job ID, resource, or endpoint path does not exist. | Verify the resource ID and endpoint URL. | No |
| 408 | `Request Timeout` | The page took longer than the request `timeout` to load. | Increase `timeout`, simplify actions, or use `fastMode`. | Yes, with backoff |
| 409 | `Conflict` | Resource is in a state that prevents the operation (e.g. already deleted). | Re-fetch state and reconcile before retrying. | No |
Expand Down
53 changes: 53 additions & 0 deletions features/scrape.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -572,6 +572,59 @@ curl -s -X POST "https://api.firecrawl.dev/v2/scrape" \

See [Verifying Freshness and Liveness](/developer-guides/usage-guides/verifying-freshness-and-liveness) for a checklist and examples.

## Alexandria data tools

Scrape can also run a data tool from [Alexandria](/features/alexandria), the Firecrawl catalogue of third-party data providers. Send an `alexandria` object with `provider`, `capability`, and `options` instead of a `url`. Get these values from [Search](/features/search#alexandria-tools) or [Find Tools](/features/alexandria#find-tools).

<CodeGroup>

```python Python
result = firecrawl.scrape(
alexandria={
"provider": "particle",
"capability": "podcasts/episodes/search",
"options": {"semantic_search": "AI agents", "limit": 2},
},
)

print(result.alexandria[0].data)
```

```js Node
const result = await firecrawl.scrape({
alexandria: {
provider: "particle",
capability: "podcasts/episodes/search",
options: { semantic_search: "AI agents", limit: 2 },
},
});

console.log(result.alexandria[0].data);
```

```bash cURL
curl -X POST https://api.firecrawl.dev/v2/scrape \
-H "Content-Type: application/json" \
-H "Authorization: Bearer fc-YOUR_API_KEY" \
-d '{
"alexandria": {
"provider": "particle",
"capability": "podcasts/episodes/search",
"options": { "semantic_search": "AI agents", "limit": 2 }
}
}'
```

</CodeGroup>

- The response has one item for each call in `data.alexandria`. Each item has `data` or an `error`. Examine each item, because one call can fail when the request succeeds.
- Send an array of up to 10 calls to run them in one request. You cannot use `alexandria` together with `url` or other scrape options.
- Each call costs the price listed on its tool. `data.creditsCost` shows the total.
- To retry, send the same `x-request-id` header with the same body. A completed request replays its result and does not run again. Without this header, a retry is a new charge.
- Some providers need an org admin to accept their terms first. Then the request fails with [`THIRD_PARTY_DATA_TERMS_REQUIRED`](/api-reference/errors).

To find tools for the page that you scrape, add `domainTools: true` to an ordinary URL scrape. Matching tools come back in `data.tools`. See [Alexandria](/features/alexandria#pricing-and-access) for pricing and access.

## Batch scraping multiple URLs

You can now batch scrape multiple URLs at the same time. It takes the starting URLs and optional parameters as arguments. The params argument allows you to specify additional options for the batch scrape job, such as the output formats.
Expand Down
22 changes: 22 additions & 0 deletions features/search.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -94,9 +94,31 @@ In addition to regular web results, Search supports specialized result types via
- `web`: standard web results (default)
- `news`: news-focused results
- `images`: image search results
- `alexandria`: data tools from [Alexandria](/features/alexandria), the Firecrawl catalogue of third-party data providers

You can request multiple sources in a single call (e.g., `sources: ["web", "news"]`). When you do, the `limit` parameter applies **per source type** — so `limit: 5` with `sources: ["web", "news"]` returns up to 5 web results and up to 5 news results (10 total). If you need different parameters per source (for example, different `limit` values or different `scrapeOptions`), make separate calls instead.

### Alexandria tools

When `sources` includes `alexandria`, Search returns matching data tools in `data.tools` (`result.tools` in the SDKs). A tool is a provider and a capability that you can run with [Scrape](/features/scrape#alexandria-data-tools). Tool discovery is free. Use `sources: ["alexandria"]` to get tools only, or `["web", "alexandria"]` to also get web results, which are billed normally.

```bash cURL
curl -X POST https://api.firecrawl.dev/v2/search \
-H "Content-Type: application/json" \
-H "Authorization: Bearer fc-YOUR_API_KEY" \
-d '{
"query": "podcast conversations about AI agents",
"sources": ["web", "alexandria"],
"limit": 2
}'
```

- By default, each tool has only its `provider`, `capability`, and `description`. Set `toolDetail` to `"summary"` for more metadata, or to `"full"` for the input and output contracts.
- Search also adds tools that match the websites in your results. Set `domainTools: false` to turn off this domain matching.
- Tool discovery needs an API key. It is not available with zero data retention.

See [Alexandria](/features/alexandria) for the full workflow and [Find Tools](/features/alexandria#find-tools) to browse the catalogue.

## Search Categories

Filter search results by specific categories using the `categories` parameter:
Expand Down
14 changes: 14 additions & 0 deletions mcp-server/tools.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ Start with [Get Started](/mcp-server) and pick [For Agents](/mcp-server/keyless)
| Extract structured data | `firecrawl_scrape` with JSON format | You have the URL and want data matching a prompt or JSON schema. |
| Discover site URLs | `firecrawl_map` | You need to find pages before deciding what to extract. |
| Search the web | `firecrawl_search` | You have a query rather than a known URL. |
| Get data from a third-party provider | `firecrawl_find_tools`, then `firecrawl_scrape` with `alexandria` | You need typed records from a provider in the [Alexandria](/features/alexandria) catalogue. Authenticated `firecrawl_search` also returns matching tools in `data.tools`. |
| Parse a file | `firecrawl_parse` | You need content from a PDF, document, spreadsheet, or HTML file. |
| Extract many pages | `firecrawl_crawl` and `firecrawl_check_crawl_status` | You need to traverse a site or section. The crawl tool polls the job to a terminal state before returning. |
| Run autonomous research | `firecrawl_agent` and `firecrawl_agent_status` | The task spans multiple sources and the exact pages are not known. |
Expand Down Expand Up @@ -70,6 +71,16 @@ Use the schema shown by your MCP client for the current arguments. The feature g
The `firecrawl_monitor_*` family creates, lists, updates, runs, and inspects recurring monitors. `firecrawl_monitor_delete` permanently removes a monitor and should be called only when the user explicitly intends to delete it.
</Accordion>

<Accordion title="Use Alexandria data tools">
[Alexandria](/features/alexandria) tools need an authenticated session. Keyless sessions get no Alexandria tools.

1. **Discover.** Authenticated `firecrawl_search` uses `sources: ["web", "alexandria"]` by default and returns matching tools in `data.tools`. Use `sources: ["alexandria"]` for tools only, or `sources: ["web"]` to omit semantic tool discovery.
2. **Inspect.** Call `firecrawl_find_tools` with no arguments to browse categories, then providers, then tools. Use `query` or `urls` to find tools for a task or a website. Use `capabilities` to read the full contract of a tool. Follow the returned `nextTool` for the next step.
3. **Run.** Call `firecrawl_scrape` with `alexandria: {provider, capability, options}` instead of `url`. Send an array of up to 10 calls to run them together. Each result in `data.alexandria` has `data` or an `error`.

Discovery is free. Each run costs the price listed on its tool. If a provider needs accepted terms, the tool returns `THIRD_PARTY_DATA_TERMS_REQUIRED` with `requiresAction.url`. Give that URL to an org admin, who accepts the terms in the Firecrawl dashboard. The agent must not accept terms. After the admin confirms, send the same call again.
</Accordion>

<Accordion title="Enable optional feedback tools">
Set `FIRECRAWL_NO_SEARCH_FEEDBACK=1` to prevent `firecrawl_search_feedback` from being registered. Set `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` to prevent `firecrawl_feedback` from being registered.
</Accordion>
Expand Down Expand Up @@ -105,6 +116,9 @@ Use the schema shown by your MCP client for the current arguments. The feature g
<Card title="Monitoring" href="/features/monitoring" icon="bell">
Track page changes and receive notifications.
</Card>
<Card title="Alexandria" href="/features/alexandria" icon="books">
Find and run data tools from third-party providers.
</Card>
</CardGroup>

## Troubleshooting
Expand Down
51 changes: 50 additions & 1 deletion sdks/cli.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,10 @@ Scrape a single URL and extract its content in various formats.
| `--json` | | Force JSON output even with single format |
| `--pretty` | | Pretty print JSON output |
| `--timing` | | Show request timing and other useful information |
| `--alexandria <tool>` | | Run an [Alexandria](/features/alexandria) tool, given as `provider/capability`, instead of scraping a URL. Repeat to run more than one tool |
| `--options <json>` | | JSON input for each Alexandria tool, in the same order |
| `--request-id <id>` | | Request ID for an Alexandria run. Use the same ID to recover the same run |
| `--domain-tools` | | Also return Alexandria tools that match the scraped URL. The tools do not run |

---

Expand All @@ -170,7 +174,10 @@ Search the web and optionally scrape the results.
| Option | Description |
| ---------------------------- | ------------------------------------------------------------------------------------------- |
| `--limit <number>` | Maximum results (default: 5, max: 100) |
| `--sources <sources>` | Sources to search: `web`, `images`, `news` (comma-separated) |
| `--sources <sources>` | Sources to search: `web`, `images`, `news`, `alexandria` (comma-separated, default: `web,alexandria`). Use `web` to omit Alexandria tools |
| `--domain-tools` | Include Alexandria tools for the domains in web results (on by default with `alexandria`) |
| `--no-domain-tools` | Turn off domain matching for Alexandria tools |
| `--tool-detail <detail>` | Alexandria tool detail: `compact` (default), `summary`, or `full` (includes contracts) |
| `--categories <categories>` | Filter by category: `research`, `pdf`, `developer` (comma-separated) |
| `--tbs <value>` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) |
| `--location <location>` | Geo-targeting (e.g., "Berlin,Germany") |
Expand All @@ -186,6 +193,48 @@ Search the web and optionally scrape the results.

---

### Alexandria

Find and run data tools from [Alexandria](/features/alexandria), the Firecrawl catalogue of third-party data providers. Discovery is free. A tool run costs the price listed on the tool.

```bash CLI
# Find tools for a task (tools only, no web results)
firecrawl search alexandria "podcast conversations about AI agents"

# Browse categories, then providers and their tools
firecrawl alexandria list
firecrawl list particle

# Read the inputs and response of one tool
firecrawl find-tools --options '{"providers":["particle"],"capabilities":["podcasts/episodes/search"],"level":"tools","expand":["options","response"]}' --json

# Run the tool
firecrawl scrape particle/podcasts/episodes/search --options '{"semantic_search":"AI agents","limit":2}'
```

`firecrawl scrape provider/capability` is a short form of `firecrawl scrape --alexandria provider/capability`. Tool addresses go directly to Alexandria. They never fall back to a URL scrape.

#### Provider terms

When a provider returns `THIRD_PARTY_DATA_TERMS_REQUIRED`, read its terms first:

```bash CLI
firecrawl alexandria terms show particle --pretty
```

To accept, give the exact version and digest that you reviewed. Acceptance applies to the organization of your API key:

```bash CLI
firecrawl alexandria terms accept particle \
--terms-version '<reviewed-version>' --digest '<reviewed-sha256>' --confirm
```

An agent must not accept terms on its own. It must show the terms to the user, get explicit approval, and wait. An org admin can also accept terms at [Data sources settings](https://www.firecrawl.dev/app/settings?tab=data-sources). After acceptance, run the original command again.

Use `firecrawl alexandria feedback` to report issues or request capabilities. Run `firecrawl alexandria feedback --help` for options.

---

### Developer

Search the [Developer Index](/features/developer) — issues, merged pull requests, and READMEs from public code repositories, alongside curated documentation sites.
Expand Down
34 changes: 34 additions & 0 deletions sdks/go.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -222,6 +222,40 @@ for _, result := range results.Web {
}
```

### Alexandria Data Tools

[Alexandria](/features/alexandria) is the Firecrawl catalogue of third-party data providers. Use `FindTools` to browse providers and read tool contracts for free. Then run a tool with `ScrapeAlexandria`.

```go
catalogue, err := client.FindTools(ctx, &firecrawl.FindToolsOptions{
Providers: []string{"particle"},
Limit: firecrawl.Int(2),
})
if err != nil {
log.Fatal(err)
}
fmt.Println(catalogue.Items)

result, err := client.ScrapeAlexandria(ctx, []firecrawl.AlexandriaCall{{
Provider: "particle",
Capability: "podcasts/episodes/search",
Options: map[string]interface{}{"semantic_search": "AI agents", "limit": 2},
}}, nil)
if err != nil {
log.Fatal(err)
}

for _, item := range result.Alexandria {
if item.Error != nil {
fmt.Println(item.Error.Code, item.Error.Message)
continue
}
fmt.Println(item.Data)
}
```

Examine `Error` on each result, because one call can fail when the request succeeds. To retry after an uncertain result, pass the same `AlexandriaOptions.RequestID` with the same calls. `Search` with `Sources: []interface{}{"alexandria"}` also returns matching tools in `Tools`.

### Batch Scraping

Scrape multiple URLs in parallel using `BatchScrape`.
Expand Down
24 changes: 24 additions & 0 deletions sdks/node.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -123,6 +123,30 @@ Agent runs are asynchronous. Use `startAgent` to get a job ID back immediately,

Every run also records an execution trace and output snapshots, which you can read with `getAgentTrace` and `getAgentSnapshot`. See [Agent](/features/agent) for the event schema and the full parameter list.

### Alexandria Data Tools

[Alexandria](/features/alexandria) is the Firecrawl catalogue of third-party data providers. Use `findTools` to browse providers and read tool contracts for free. Then run a tool with `scrape` and an `alexandria` request instead of a URL.

```js Node
const catalogue = await firecrawl.findTools({ providers: ["particle"], limit: 2 });
console.log(catalogue.items);

const result = await firecrawl.scrape({
alexandria: {
provider: "particle",
capability: "podcasts/episodes/search",
options: { semantic_search: "AI agents", limit: 2 },
},
});

for (const item of result.alexandria) {
if (item.error) console.error(item.error.code, item.error.message);
else console.log(item.data);
}
```

`alexandria` also accepts an array of up to 10 calls. Each item in `result.alexandria` has `data` or an `error`, and `result.creditsCost` shows the total cost. To retry after an uncertain result, pass the same `requestId` with the same request. `search` with `sources: ["alexandria"]` also returns matching tools in `result.tools`.

### Crawling a Website with WebSockets

Stream crawl results in real time with `watcher(jobId, options)`. You receive each page as it is crawled instead of waiting for the entire job to finish.
Expand Down
25 changes: 25 additions & 0 deletions sdks/python.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -123,6 +123,31 @@ Agent runs are asynchronous. Use `start_agent` to get a job ID back immediately,

Every run also records an execution trace and output snapshots, which you can read with `get_agent_trace` and `get_agent_snapshot`. See [Agent](/features/agent) for the event schema and the full parameter list.

### Alexandria Data Tools

[Alexandria](/features/alexandria) is the Firecrawl catalogue of third-party data providers. Use `find_tools` to browse providers and read tool contracts for free. Then run a tool with `scrape(alexandria=...)` instead of a URL.

```python Python
catalogue = firecrawl.find_tools(providers=["particle"], limit=2)
print(catalogue.items)

result = firecrawl.scrape(
alexandria={
"provider": "particle",
"capability": "podcasts/episodes/search",
"options": {"semantic_search": "AI agents", "limit": 2},
},
)

for item in result.alexandria:
if item.error:
print(item.error.code, item.error.message)
else:
print(item.data)
```

`alexandria` also accepts a list of up to 10 calls. `scrape_alexandria(calls)` does the same as `scrape(alexandria=calls)`. Each item in `result.alexandria` has `data` or an `error`, and `result.credits_cost` shows the total cost. To retry after an uncertain result, pass the same `request_id` with the same request. `search` with `sources=["alexandria"]` also returns matching tools in `result.tools`.

### Crawling a Website with WebSockets

To crawl a website with WebSockets, start the job with `start_crawl` and subscribe using the `watcher` helper. Create a watcher with the job ID and attach handlers (e.g., for page, completed, failed) before calling `start()`.
Expand Down
Loading
Loading