Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 22 additions & 4 deletions mcp_server/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,9 +3,10 @@
Anonymous, read-only Model Context Protocol access to the public
[Papers With Code](https://paperswithcode.co) catalog.

The server uses MCP `2026-07-28` over Streamable HTTP and serves legacy
2025-era clients on the same `/mcp` endpoint. It returns versioned structured
output with compact text fallbacks.
The server uses stock-client MCP `2025-11-25` over Streamable HTTP and also
serves the experimental `2026-07-28` discovery protocol on the same `/mcp`
endpoint. It returns versioned structured output with compact Markdown
fallbacks.

## Run locally

Expand All @@ -27,17 +28,27 @@ curl http://127.0.0.1:7860/health
- `get_paper_info`
- `read_paper`
- `get_related_papers`
- `get_trending_papers`
- `get_paper_evaluations`
- `get_paper_lineage`
- `get_task`
- `list_tasks`
- `get_method`
- `list_methods`
- `list_benchmarks`
- `get_benchmark`

All tools are annotated read-only and idempotent. Search is deterministic;
All tools are annotated read-only and idempotent. Expected failures use typed
messages (`not_found`, `ambiguous`, `no_markdown`, and `upstream_timeout`), with
candidate IDs and slugs for ambiguous references. Search is deterministic;
the caller controls keyword or semantic mode. `read_paper` fetches at most one
64 KiB catalog chunk per call and returns a signed, one-hour continuation cursor
when more Markdown remains. Continuations stay pinned to the resolved paper and
content version, so a changed paper fails with an explicit restart response.
Benchmark and paper evaluations are paginated. Equivalent duplicate rows are
merged while preserving every task-specific rank scope; rows also expose the
reported protocol, split, shot count when stated, source, update timestamp,
openness, and metric direction.

## Resources

Expand Down Expand Up @@ -69,6 +80,13 @@ Tool inputs cap list results at 25, catalog calls time out after 25 seconds, and
request, upstream, and serialized MCP response bodies are bounded to 2 MiB.
Markdown chunks use a bounded 256-entry/16 MiB in-memory cache.

`GET /.well-known/mcp` exposes connection metadata and `GET /docs` publishes
the live input/output schema for every tool. On the canonical host these are
also available as `/.well-known/mcp` and `/mcp/schema`, while a human GET of
`/mcp` opens the setup guide. Direct package servers return `405` for a bare
`GET /mcp`; MCP requests use `POST /mcp`. Rate limits return `429` with
`Retry-After`.

## Test

```bash
Expand Down
28 changes: 20 additions & 8 deletions mcp_server/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ description: "Papers With Code MCP tools for searching and reading AI/ML papers,
compatibility: "Requires an MCP client connected to https://paperswithcode.co/mcp with the Papers With Code tools available."
---

Generated for `pwc-mcp v0.1.0` and MCP protocol `2026-07-28`.
Generated for `pwc-mcp v0.2.1` and MCP protocol `2025-11-25`.

The tools query the public [Papers With Code](https://paperswithcode.co) catalog
anonymously and are read-only. If live tool discovery and this skill disagree,
Expand Down Expand Up @@ -36,22 +36,26 @@ matching IDs.
## Tools

- `search_papers({"query": QUERY, "limit": LIMIT, "page": PAGE, "mode": "keyword"|"semantic", "published_after": START_DATE, "published_before": END_DATE, "has_official_implementation": BOOLEAN})` — search papers. Omit optional arguments when they are not needed.
- `get_paper_info({"paper": PAPER})` — show paper metadata, abstract, tasks, methods, repositories, and project pages.
- `get_paper_info({"paper": PAPER, "include_resources": BOOLEAN, "repo_limit": LIMIT})` — show paper metadata, repository count, and official code by default; optionally add capped repositories, project pages, and Hugging Face models/datasets.
- `read_paper({"paper": PAPER})` — read one stored paper Markdown chunk. If `truncated` is true, call `read_paper` again with the same `paper` and the returned `next_cursor`; repeat until `truncated` is false. Treat the cursor as opaque and use it within one hour.
- `list_papers({"page": PAGE, "limit": LIMIT, "search": SEARCH, "published_after": START_DATE, "published_before": END_DATE, "task": TASK, "method": METHOD, "conference": CONFERENCE, "framework": FRAMEWORK, "organization": ORGANIZATION, "authors": [AUTHOR], "order_by": "date_published"|"citation_count"|"title", "order_direction": "asc"|"desc"})` — list and filter papers. Omit optional arguments when they are not needed.
- `get_related_papers({"paper": PAPER, "limit": LIMIT})` — list related papers.
- `get_paper_lineage({"paper": PAPER})` — list explicit predecessors and successors.
- `get_task({"task": TASK})` — inspect one exact task by ID, slug, or name, including its area, parents, children, and benchmarks.
- `get_trending_papers({"limit": LIMIT, "max_age_days": DAYS, "min_velocity": VELOCITY})` — list trending papers.
- `get_paper_evaluations({"paper": PAPER, "page": PAGE, "limit": LIMIT})` — list paginated benchmark evaluations reported by a paper.
- `get_paper_lineage({"paper": PAPER})` — list explicit catalog predecessors and successors; empty results do not prove that none exist.
- `list_tasks({"search": SEARCH, "page": PAGE, "limit": LIMIT})` — discover task slugs and IDs.
- `get_task({"task": TASK, "benchmark_limit": LIMIT})` — inspect one exact task by ID or slug with a capped benchmark list.
- `list_methods({"search": SEARCH, "page": PAGE, "limit": LIMIT})` — discover method slugs and IDs.
- `get_method({"method": METHOD})` — inspect one exact method by ID, slug, full name, or name.
- `list_benchmarks({"page": PAGE, "limit": LIMIT, "search": SEARCH, "task": TASK, "include_descendants": BOOLEAN, "minimum_evaluations": MINIMUM_EVALUATIONS, "is_open": BOOLEAN})` — list and filter benchmarks. Omit optional arguments when they are not needed.
- `get_benchmark({"benchmark": BENCHMARK, "limit": LIMIT, "is_open": BOOLEAN})` — inspect one exact benchmark and its leading evaluation rows.
- `get_benchmark({"benchmark": BENCHMARK, "page": PAGE, "limit": LIMIT, "is_open": BOOLEAN})` — inspect one exact benchmark and one page of evaluation rows.

All page numbers start at 1. `limit` is between 1 and 25. Follow `next_page`
when the user asks for more results than one response contains; do not infer
that a missing item does not exist until the relevant pages have been checked.

The MCP server does not expose standalone CLI commands for paper editing,
authentication, skill installation, version display, taxonomy enumeration, or
authentication, skill installation, version display, or
advanced benchmark metric/parameter/Pareto filtering. Do not invent equivalent
tools. Use the separate `pwc` CLI only when it is available and the user needs
one of those capabilities.
Expand All @@ -60,8 +64,10 @@ one of those capabilities.

1. Use `list_benchmarks({"task": TASK})` to discover active benchmarks, then
`get_benchmark({"benchmark": NAME})` to inspect a leaderboard.
2. Use `get_paper_info({"paper": PAPER})` to inspect promising results. Its
response includes repositories and project pages.
2. Use `get_paper_info({"paper": PAPER})` to inspect promising results. The
default response includes official repositories and a repository count; pass
`include_resources: true` for other repositories, project pages, and Hugging
Face artifacts.
3. Use exact `list_papers` `authors`, `task`, `method`, `conference`,
`framework`, and `organization` arguments for known identities or catalog
associations. Combine them to require every association; do not substitute
Expand All @@ -77,6 +83,12 @@ one of those capabilities.

- Tool results use stable, versioned structured output with compact text
fallbacks. Prefer structured fields over parsing the text fallback.
- Evaluation ranks are task-scoped. Compare ranks only within matching
`rank_scopes`; use `metric_directions`, protocol, split, shots, and source to
decide whether scores are comparable. Follow `next_page` for more rows.
- `has_official_implementation` means the catalog marks linked code as official;
it is not an independent audit. Evaluation `is_open` is the catalog's
implementation-availability flag and is null when not recorded.
- Search mode is deterministic: choose `keyword` by default and use `semantic`
when conceptual similarity is more useful. The MCP server does not support
the CLI's `hybrid` mode.
Expand Down
21 changes: 17 additions & 4 deletions mcp_server/SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,31 +13,38 @@ server.
the MCP contract.
- Serve anonymous, read-only requests. Search is deterministic and contains no
embedded language model.
- Use stateless Streamable HTTP at `/mcp`, supporting MCP `2026-07-28` and
legacy 2025 clients on the same endpoint. Expose `/health` for operations.
- Use stateless Streamable HTTP at `/mcp`, advertising stock-client MCP
`2025-11-25` while also supporting experimental `2026-07-28` discovery.
Expose `/health`, `/.well-known/mcp`, and a generated `/docs` schema.

## Public contract

Expose exactly these tools:
Expose these tools:

- `search_papers`
- `list_papers`
- `get_paper_info`
- `read_paper`
- `get_related_papers`
- `get_trending_papers`
- `get_paper_evaluations`
- `get_paper_lineage`
- `get_task`
- `list_tasks`
- `get_method`
- `list_methods`
- `list_benchmarks`
- `get_benchmark`

Expose these resource templates and no prompts:
Expose these resource templates:

- `pwc://papers/{paper}`
- `pwc://papers/{paper}/markdown`
- `pwc://tasks/{task}`
- `pwc://benchmarks/{benchmark}`

Expose the `find_papers`, `compare_leaderboard`, and `survey_task` prompts.

Responses use stable, MCP-specific versioned structured outputs with a text
fallback. `read_paper` performs one upstream read of at most 64 KiB per call and
returns a signed opaque continuation cursor when more Markdown remains. The
Expand All @@ -49,6 +56,12 @@ Paper references accept arXiv IDs, numeric PwC external IDs, arXiv/Hugging
Face/Papers With Code URLs, and exact titles. Ambiguous exact titles fail rather
than selecting one result.

Benchmark and paper evaluation results paginate with `page` and `next_page`.
Equivalent result rows merge their metrics while retaining task-scoped ranks,
evaluation protocol, split, shots when reported, source URL, openness, and
update timestamp. Metric direction is explicit and unknown directions remain
`unknown` rather than being guessed.

## Safety and operations

- Require a strict configurable browser Origin allowlist; native clients may
Expand Down
2 changes: 1 addition & 1 deletion mcp_server/pyproject.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[project]
name = "pwc-mcp"
version = "0.1.0"
version = "0.2.1"
description = "Read-only Papers With Code MCP server"
readme = "README.md"
requires-python = ">=3.10"
Expand Down
2 changes: 1 addition & 1 deletion mcp_server/src/pwc_mcp/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
"""Read-only Papers With Code MCP server."""

__version__ = "0.1.0"
__version__ = "0.2.1"
73 changes: 71 additions & 2 deletions mcp_server/src/pwc_mcp/app.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,8 @@
# MCP SDK diagnostics can include peer-supplied tool names and resource URIs.
# OperationalTelemetryMiddleware is the server's sole request log surface.
logging.getLogger("mcp").setLevel(logging.CRITICAL + 1)
PROTOCOL_VERSION = "2026-07-28"
PROTOCOL_VERSION = "2025-11-25"
EXPERIMENTAL_PROTOCOL_VERSION = "2026-07-28"
MAX_REQUEST_BODY_SIZE = 2 * 1024 * 1024
MAX_RESPONSE_BODY_SIZE = 2 * 1024 * 1024
KNOWN_TOOLS = {
Expand All @@ -38,21 +39,81 @@
"get_paper_info",
"read_paper",
"get_related_papers",
"get_trending_papers",
"get_paper_evaluations",
"get_paper_lineage",
"get_task",
"list_tasks",
"get_method",
"list_methods",
"list_benchmarks",
"get_benchmark",
}
KNOWN_PROTOCOLS = {
PROTOCOL_VERSION,
"2025-11-25",
EXPERIMENTAL_PROTOCOL_VERSION,
"2025-06-18",
"2025-03-26",
"2024-11-05",
}


async def well_known_mcp(request: Request) -> JSONResponse:
return JSONResponse(
{
"name": "Papers With Code",
"description": "Anonymous read-only AI research catalog",
"transport": {"type": "streamable-http", "url": "/mcp"},
"protocol_version": PROTOCOL_VERSION,
"supported_protocol_versions": sorted(KNOWN_PROTOCOLS, reverse=True),
"documentation_url": "https://paperswithcode.co/mcp/schema",
"setup_url": "https://paperswithcode.co/mcp",
}
)


def _docs_schema(server) -> dict:
tools = []
for tool in server._tool_manager.list_tools():
tools.append(
{
"name": tool.name,
"description": tool.description,
"inputSchema": tool.parameters,
"outputSchema": tool.output_schema,
}
)
return {
"name": "Papers With Code MCP",
"version": __version__,
"protocol_version": PROTOCOL_VERSION,
"endpoint": "/mcp",
"tools": tools,
}


class MCPMethodMiddleware:
"""Avoid opening an SSE response for unsupported bare GET requests."""

def __init__(self, app: ASGIApp):
self.app = app

async def __call__(self, scope: Scope, receive: Receive, send: Send) -> None:
if (
scope["type"] == "http"
and scope.get("path") == "/mcp"
and scope.get("method") == "GET"
):
response = JSONResponse(
{"error": "method_not_allowed", "allowed": ["POST"]},
status_code=405,
headers={"Allow": "POST"},
)
await response(scope, receive, send)
return
await self.app(scope, receive, send)


def _csv_env(name: str, default: list[str]) -> list[str]:
value = os.environ.get(name)
if value is None:
Expand Down Expand Up @@ -426,6 +487,13 @@ def create_app(
),
)
app.routes.insert(0, Route("/health", health, methods=["GET"]))
docs_schema = _docs_schema(server)

async def docs(_request: Request) -> JSONResponse:
return JSONResponse(docs_schema)

app.routes.insert(1, Route("/docs", docs, methods=["GET"]))
app.routes.insert(2, Route("/.well-known/mcp", well_known_mcp, methods=["GET"]))
app.state.pwc_catalog_readiness = CatalogReadiness(
catalog_client,
initially_ready=catalog is not None,
Expand All @@ -438,6 +506,7 @@ def create_app(
global_concurrency_limit=global_concurrency_limit,
trust_proxy_headers=trust_proxy_headers,
)
wrapped = MCPMethodMiddleware(wrapped)
wrapped = ResponseSizeLimitMiddleware(wrapped)
wrapped = OperationalTelemetryMiddleware(wrapped)
return CORSMiddleware(
Expand Down
Loading