From 78828ca7c889500678f68d79cde05eea7248afaf Mon Sep 17 00:00:00 2001 From: Claude Date: Fri, 2 Oct 2026 14:11:31 +0000 Subject: [PATCH] docs: link Alexandria from the Search, Scrape, MCP, CLI, SDK and error pages Add short Alexandria sections that link to /features/alexandria: - Search: alexandria source, data.tools, toolDetail and domainTools - Scrape: alexandria body with a code example, results and retries - MCP tools: firecrawl_find_tools row and an Alexandria workflow - CLI: scrape and search options, find-tools, list, terms show/accept - Node, Python, Go and Rust SDKs: find tools and run a tool - Errors: THIRD_PARTY_DATA_TERMS_REQUIRED (403, requiresAction.url) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01LKCJJgeohVR5dnd1MDBzAR --- api-reference/errors.mdx | 1 + features/scrape.mdx | 53 ++++++++++++++++++++++++++++++++++++++++ features/search.mdx | 22 +++++++++++++++++ mcp-server/tools.mdx | 14 +++++++++++ sdks/cli.mdx | 51 +++++++++++++++++++++++++++++++++++++- sdks/go.mdx | 34 ++++++++++++++++++++++++++ sdks/node.mdx | 24 ++++++++++++++++++ sdks/python.mdx | 25 +++++++++++++++++++ sdks/rust.mdx | 31 +++++++++++++++++++++++ 9 files changed, 254 insertions(+), 1 deletion(-) diff --git a/api-reference/errors.mdx b/api-reference/errors.mdx index 01d1e9d52..fbaed37f6 100644 --- a/api-reference/errors.mdx +++ b/api-reference/errors.mdx @@ -35,6 +35,7 @@ All non-2xx responses return JSON with a top-level `success: false` and a string | 402 | `Payment Required: Insufficient credits` | Plan credits are exhausted or billing is not configured. | Turn on pay-as-you-go, or upgrade your plan. | No | | 403 | `Forbidden` | Key lacks permission for this endpoint or feature. | Use a key with the required scope, or upgrade the plan that gates this feature. | No | | 403 | `SCRAPE_PROMPT_INJECTION_DETECTED` | JSON mode with `checkPromptInjection: true` detected a prompt injection attempt in the scraped page content, so extraction was aborted. | Inspect the page content manually. If it is a false positive, retry without `checkPromptInjection`. See [Prompt injection detection](/features/llm-extract#prompt-injection-detection). | No | +| 403 | `THIRD_PARTY_DATA_TERMS_REQUIRED` | An [Alexandria](/features/alexandria) provider needs your organization to accept its terms before the request can run. The provider did not run. The response has this value in `code`, and includes `requiresAction.url`. | Send `requiresAction.url` to an org admin, who reviews and accepts the terms there. Agents must not accept terms on their own. After acceptance, send the same request again. | After acceptance | | 404 | `Not Found` | The job ID, resource, or endpoint path does not exist. | Verify the resource ID and endpoint URL. | No | | 408 | `Request Timeout` | The page took longer than the request `timeout` to load. | Increase `timeout`, simplify actions, or use `fastMode`. | Yes, with backoff | | 409 | `Conflict` | Resource is in a state that prevents the operation (e.g. already deleted). | Re-fetch state and reconcile before retrying. | No | diff --git a/features/scrape.mdx b/features/scrape.mdx index 655f93c3a..e8f41d1eb 100644 --- a/features/scrape.mdx +++ b/features/scrape.mdx @@ -572,6 +572,59 @@ curl -s -X POST "https://api.firecrawl.dev/v2/scrape" \ See [Verifying Freshness and Liveness](/developer-guides/usage-guides/verifying-freshness-and-liveness) for a checklist and examples. +## Alexandria data tools + +Scrape can also run a data tool from [Alexandria](/features/alexandria), the Firecrawl catalogue of third-party data providers. Send an `alexandria` object with `provider`, `capability`, and `options` instead of a `url`. Get these values from [Search](/features/search#alexandria-tools) or [Find Tools](/features/alexandria#find-tools). + + + +```python Python +result = firecrawl.scrape( + alexandria={ + "provider": "particle", + "capability": "podcasts/episodes/search", + "options": {"semantic_search": "AI agents", "limit": 2}, + }, +) + +print(result.alexandria[0].data) +``` + +```js Node +const result = await firecrawl.scrape({ + alexandria: { + provider: "particle", + capability: "podcasts/episodes/search", + options: { semantic_search: "AI agents", limit: 2 }, + }, +}); + +console.log(result.alexandria[0].data); +``` + +```bash cURL +curl -X POST https://api.firecrawl.dev/v2/scrape \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer fc-YOUR_API_KEY" \ + -d '{ + "alexandria": { + "provider": "particle", + "capability": "podcasts/episodes/search", + "options": { "semantic_search": "AI agents", "limit": 2 } + } + }' +``` + + + +- The response has one item for each call in `data.alexandria`. Each item has `data` or an `error`. Examine each item, because one call can fail when the request succeeds. +- Send an array of up to 10 calls to run them in one request. You cannot use `alexandria` together with `url` or other scrape options. +- Each call costs the price listed on its tool. `data.creditsCost` shows the total. +- To retry, send the same `x-request-id` header with the same body. A completed request replays its result and does not run again. Without this header, a retry is a new charge. +- Some providers need an org admin to accept their terms first. Then the request fails with [`THIRD_PARTY_DATA_TERMS_REQUIRED`](/api-reference/errors). + +To find tools for the page that you scrape, add `domainTools: true` to an ordinary URL scrape. Matching tools come back in `data.tools`. See [Alexandria](/features/alexandria#pricing-and-access) for pricing and access. + ## Batch scraping multiple URLs You can now batch scrape multiple URLs at the same time. It takes the starting URLs and optional parameters as arguments. The params argument allows you to specify additional options for the batch scrape job, such as the output formats. diff --git a/features/search.mdx b/features/search.mdx index 9f787f4c4..2f564f833 100644 --- a/features/search.mdx +++ b/features/search.mdx @@ -94,9 +94,31 @@ In addition to regular web results, Search supports specialized result types via - `web`: standard web results (default) - `news`: news-focused results - `images`: image search results +- `alexandria`: data tools from [Alexandria](/features/alexandria), the Firecrawl catalogue of third-party data providers You can request multiple sources in a single call (e.g., `sources: ["web", "news"]`). When you do, the `limit` parameter applies **per source type** — so `limit: 5` with `sources: ["web", "news"]` returns up to 5 web results and up to 5 news results (10 total). If you need different parameters per source (for example, different `limit` values or different `scrapeOptions`), make separate calls instead. +### Alexandria tools + +When `sources` includes `alexandria`, Search returns matching data tools in `data.tools` (`result.tools` in the SDKs). A tool is a provider and a capability that you can run with [Scrape](/features/scrape#alexandria-data-tools). Tool discovery is free. Use `sources: ["alexandria"]` to get tools only, or `["web", "alexandria"]` to also get web results, which are billed normally. + +```bash cURL +curl -X POST https://api.firecrawl.dev/v2/search \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer fc-YOUR_API_KEY" \ + -d '{ + "query": "podcast conversations about AI agents", + "sources": ["web", "alexandria"], + "limit": 2 + }' +``` + +- By default, each tool has only its `provider`, `capability`, and `description`. Set `toolDetail` to `"summary"` for more metadata, or to `"full"` for the input and output contracts. +- Search also adds tools that match the websites in your results. Set `domainTools: false` to turn off this domain matching. +- Tool discovery needs an API key. It is not available with zero data retention. + +See [Alexandria](/features/alexandria) for the full workflow and [Find Tools](/features/alexandria#find-tools) to browse the catalogue. + ## Search Categories Filter search results by specific categories using the `categories` parameter: diff --git a/mcp-server/tools.mdx b/mcp-server/tools.mdx index 2e35cc9b8..9b2790e28 100644 --- a/mcp-server/tools.mdx +++ b/mcp-server/tools.mdx @@ -26,6 +26,7 @@ Start with [Get Started](/mcp-server) and pick [For Agents](/mcp-server/keyless) | Extract structured data | `firecrawl_scrape` with JSON format | You have the URL and want data matching a prompt or JSON schema. | | Discover site URLs | `firecrawl_map` | You need to find pages before deciding what to extract. | | Search the web | `firecrawl_search` | You have a query rather than a known URL. | +| Get data from a third-party provider | `firecrawl_find_tools`, then `firecrawl_scrape` with `alexandria` | You need typed records from a provider in the [Alexandria](/features/alexandria) catalogue. Authenticated `firecrawl_search` also returns matching tools in `data.tools`. | | Parse a file | `firecrawl_parse` | You need content from a PDF, document, spreadsheet, or HTML file. | | Extract many pages | `firecrawl_crawl` and `firecrawl_check_crawl_status` | You need to traverse a site or section. The crawl tool polls the job to a terminal state before returning. | | Run autonomous research | `firecrawl_agent` and `firecrawl_agent_status` | The task spans multiple sources and the exact pages are not known. | @@ -70,6 +71,16 @@ Use the schema shown by your MCP client for the current arguments. The feature g The `firecrawl_monitor_*` family creates, lists, updates, runs, and inspects recurring monitors. `firecrawl_monitor_delete` permanently removes a monitor and should be called only when the user explicitly intends to delete it. + + [Alexandria](/features/alexandria) tools need an authenticated session. Keyless sessions get no Alexandria tools. + + 1. **Discover.** Authenticated `firecrawl_search` uses `sources: ["web", "alexandria"]` by default and returns matching tools in `data.tools`. Use `sources: ["alexandria"]` for tools only, or `sources: ["web"]` to omit semantic tool discovery. + 2. **Inspect.** Call `firecrawl_find_tools` with no arguments to browse categories, then providers, then tools. Use `query` or `urls` to find tools for a task or a website. Use `capabilities` to read the full contract of a tool. Follow the returned `nextTool` for the next step. + 3. **Run.** Call `firecrawl_scrape` with `alexandria: {provider, capability, options}` instead of `url`. Send an array of up to 10 calls to run them together. Each result in `data.alexandria` has `data` or an `error`. + + Discovery is free. Each run costs the price listed on its tool. If a provider needs accepted terms, the tool returns `THIRD_PARTY_DATA_TERMS_REQUIRED` with `requiresAction.url`. Give that URL to an org admin, who accepts the terms in the Firecrawl dashboard. The agent must not accept terms. After the admin confirms, send the same call again. + + Set `FIRECRAWL_NO_SEARCH_FEEDBACK=1` to prevent `firecrawl_search_feedback` from being registered. Set `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` to prevent `firecrawl_feedback` from being registered. @@ -105,6 +116,9 @@ Use the schema shown by your MCP client for the current arguments. The feature g Track page changes and receive notifications. + + Find and run data tools from third-party providers. + ## Troubleshooting diff --git a/sdks/cli.mdx b/sdks/cli.mdx index 19d4101b5..85316b70f 100644 --- a/sdks/cli.mdx +++ b/sdks/cli.mdx @@ -152,6 +152,10 @@ Scrape a single URL and extract its content in various formats. | `--json` | | Force JSON output even with single format | | `--pretty` | | Pretty print JSON output | | `--timing` | | Show request timing and other useful information | +| `--alexandria ` | | Run an [Alexandria](/features/alexandria) tool, given as `provider/capability`, instead of scraping a URL. Repeat to run more than one tool | +| `--options ` | | JSON input for each Alexandria tool, in the same order | +| `--request-id ` | | Request ID for an Alexandria run. Use the same ID to recover the same run | +| `--domain-tools` | | Also return Alexandria tools that match the scraped URL. The tools do not run | --- @@ -170,7 +174,10 @@ Search the web and optionally scrape the results. | Option | Description | | ---------------------------- | ------------------------------------------------------------------------------------------- | | `--limit ` | Maximum results (default: 5, max: 100) | -| `--sources ` | Sources to search: `web`, `images`, `news` (comma-separated) | +| `--sources ` | Sources to search: `web`, `images`, `news`, `alexandria` (comma-separated, default: `web,alexandria`). Use `web` to omit Alexandria tools | +| `--domain-tools` | Include Alexandria tools for the domains in web results (on by default with `alexandria`) | +| `--no-domain-tools` | Turn off domain matching for Alexandria tools | +| `--tool-detail ` | Alexandria tool detail: `compact` (default), `summary`, or `full` (includes contracts) | | `--categories ` | Filter by category: `research`, `pdf`, `developer` (comma-separated) | | `--tbs ` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) | | `--location ` | Geo-targeting (e.g., "Berlin,Germany") | @@ -186,6 +193,48 @@ Search the web and optionally scrape the results. --- +### Alexandria + +Find and run data tools from [Alexandria](/features/alexandria), the Firecrawl catalogue of third-party data providers. Discovery is free. A tool run costs the price listed on the tool. + +```bash CLI +# Find tools for a task (tools only, no web results) +firecrawl search alexandria "podcast conversations about AI agents" + +# Browse categories, then providers and their tools +firecrawl alexandria list +firecrawl list particle + +# Read the inputs and response of one tool +firecrawl find-tools --options '{"providers":["particle"],"capabilities":["podcasts/episodes/search"],"level":"tools","expand":["options","response"]}' --json + +# Run the tool +firecrawl scrape particle/podcasts/episodes/search --options '{"semantic_search":"AI agents","limit":2}' +``` + +`firecrawl scrape provider/capability` is a short form of `firecrawl scrape --alexandria provider/capability`. Tool addresses go directly to Alexandria. They never fall back to a URL scrape. + +#### Provider terms + +When a provider returns `THIRD_PARTY_DATA_TERMS_REQUIRED`, read its terms first: + +```bash CLI +firecrawl alexandria terms show particle --pretty +``` + +To accept, give the exact version and digest that you reviewed. Acceptance applies to the organization of your API key: + +```bash CLI +firecrawl alexandria terms accept particle \ + --terms-version '' --digest '' --confirm +``` + +An agent must not accept terms on its own. It must show the terms to the user, get explicit approval, and wait. An org admin can also accept terms at [Data sources settings](https://www.firecrawl.dev/app/settings?tab=data-sources). After acceptance, run the original command again. + +Use `firecrawl alexandria feedback` to report issues or request capabilities. Run `firecrawl alexandria feedback --help` for options. + +--- + ### Developer Search the [Developer Index](/features/developer) — issues, merged pull requests, and READMEs from public code repositories, alongside curated documentation sites. diff --git a/sdks/go.mdx b/sdks/go.mdx index 179a49159..c28ffab2a 100644 --- a/sdks/go.mdx +++ b/sdks/go.mdx @@ -222,6 +222,40 @@ for _, result := range results.Web { } ``` +### Alexandria Data Tools + +[Alexandria](/features/alexandria) is the Firecrawl catalogue of third-party data providers. Use `FindTools` to browse providers and read tool contracts for free. Then run a tool with `ScrapeAlexandria`. + +```go +catalogue, err := client.FindTools(ctx, &firecrawl.FindToolsOptions{ + Providers: []string{"particle"}, + Limit: firecrawl.Int(2), +}) +if err != nil { + log.Fatal(err) +} +fmt.Println(catalogue.Items) + +result, err := client.ScrapeAlexandria(ctx, []firecrawl.AlexandriaCall{{ + Provider: "particle", + Capability: "podcasts/episodes/search", + Options: map[string]interface{}{"semantic_search": "AI agents", "limit": 2}, +}}, nil) +if err != nil { + log.Fatal(err) +} + +for _, item := range result.Alexandria { + if item.Error != nil { + fmt.Println(item.Error.Code, item.Error.Message) + continue + } + fmt.Println(item.Data) +} +``` + +Examine `Error` on each result, because one call can fail when the request succeeds. To retry after an uncertain result, pass the same `AlexandriaOptions.RequestID` with the same calls. `Search` with `Sources: []interface{}{"alexandria"}` also returns matching tools in `Tools`. + ### Batch Scraping Scrape multiple URLs in parallel using `BatchScrape`. diff --git a/sdks/node.mdx b/sdks/node.mdx index 17c9395c4..0b53152c7 100644 --- a/sdks/node.mdx +++ b/sdks/node.mdx @@ -123,6 +123,30 @@ Agent runs are asynchronous. Use `startAgent` to get a job ID back immediately, Every run also records an execution trace and output snapshots, which you can read with `getAgentTrace` and `getAgentSnapshot`. See [Agent](/features/agent) for the event schema and the full parameter list. +### Alexandria Data Tools + +[Alexandria](/features/alexandria) is the Firecrawl catalogue of third-party data providers. Use `findTools` to browse providers and read tool contracts for free. Then run a tool with `scrape` and an `alexandria` request instead of a URL. + +```js Node +const catalogue = await firecrawl.findTools({ providers: ["particle"], limit: 2 }); +console.log(catalogue.items); + +const result = await firecrawl.scrape({ + alexandria: { + provider: "particle", + capability: "podcasts/episodes/search", + options: { semantic_search: "AI agents", limit: 2 }, + }, +}); + +for (const item of result.alexandria) { + if (item.error) console.error(item.error.code, item.error.message); + else console.log(item.data); +} +``` + +`alexandria` also accepts an array of up to 10 calls. Each item in `result.alexandria` has `data` or an `error`, and `result.creditsCost` shows the total cost. To retry after an uncertain result, pass the same `requestId` with the same request. `search` with `sources: ["alexandria"]` also returns matching tools in `result.tools`. + ### Crawling a Website with WebSockets Stream crawl results in real time with `watcher(jobId, options)`. You receive each page as it is crawled instead of waiting for the entire job to finish. diff --git a/sdks/python.mdx b/sdks/python.mdx index 2cbf0cdd2..68d0bcca9 100644 --- a/sdks/python.mdx +++ b/sdks/python.mdx @@ -123,6 +123,31 @@ Agent runs are asynchronous. Use `start_agent` to get a job ID back immediately, Every run also records an execution trace and output snapshots, which you can read with `get_agent_trace` and `get_agent_snapshot`. See [Agent](/features/agent) for the event schema and the full parameter list. +### Alexandria Data Tools + +[Alexandria](/features/alexandria) is the Firecrawl catalogue of third-party data providers. Use `find_tools` to browse providers and read tool contracts for free. Then run a tool with `scrape(alexandria=...)` instead of a URL. + +```python Python +catalogue = firecrawl.find_tools(providers=["particle"], limit=2) +print(catalogue.items) + +result = firecrawl.scrape( + alexandria={ + "provider": "particle", + "capability": "podcasts/episodes/search", + "options": {"semantic_search": "AI agents", "limit": 2}, + }, +) + +for item in result.alexandria: + if item.error: + print(item.error.code, item.error.message) + else: + print(item.data) +``` + +`alexandria` also accepts a list of up to 10 calls. `scrape_alexandria(calls)` does the same as `scrape(alexandria=calls)`. Each item in `result.alexandria` has `data` or an `error`, and `result.credits_cost` shows the total cost. To retry after an uncertain result, pass the same `request_id` with the same request. `search` with `sources=["alexandria"]` also returns matching tools in `result.tools`. + ### Crawling a Website with WebSockets To crawl a website with WebSockets, start the job with `start_crawl` and subscribe using the `watcher` helper. Create a watcher with the job ID and attach handlers (e.g., for page, completed, failed) before calling `start()`. diff --git a/sdks/rust.mdx b/sdks/rust.mdx index b83d14cde..6b3c83392 100644 --- a/sdks/rust.mdx +++ b/sdks/rust.mdx @@ -351,6 +351,37 @@ for doc in &docs { } ``` +### Alexandria Data Tools + +[Alexandria](/features/alexandria) is the Firecrawl catalogue of third-party data providers. Use `find_tools` to browse providers and read tool contracts for free. Then run a tool with `scrape_alexandria`. + +```rust +use firecrawl::{AlexandriaCall, FindToolsOptions}; +use serde_json::json; + +let catalogue = client.find_tools(FindToolsOptions { + providers: Some(vec!["particle".into()]), + limit: Some(2), + ..Default::default() +}).await?; +println!("{:?}", catalogue.items); + +let result = client.scrape_alexandria(vec![AlexandriaCall { + provider: "particle".into(), + capability: "podcasts/episodes/search".into(), + options: json!({ "semantic_search": "AI agents", "limit": 2 }).as_object().cloned(), +}], None).await?; + +for item in result.alexandria { + match item.error { + Some(error) => println!("{} {}", error.code, error.message), + None => println!("{:?}", item.data), + } +} +``` + +Examine `error` on each result, because one call can fail when the request succeeds. To retry after an uncertain result, pass the same `AlexandriaOptions.request_id` with the same calls. `search` with `SearchSource::Alexandria` in `sources` also returns matching tools in `data.tools`. + ### Batch Scraping Scrape multiple URLs in parallel using `batch_scrape`.