A Cloudflare Worker that caches litellm's model pricing data and provides:
- Web UI — filterable, sortable table of all LLM model prices
- REST API — query models by provider, mode, capabilities, cost, and context window
- MCP server — native remote MCP endpoint at
/mcp - OpenAPI spec — at
/openapi.jsonfor integration with LLM tools and MCP clients
Data is refreshed automatically every 6 hours via cron trigger. Zero ongoing cost on Cloudflare's free tier.
Production deployment:
https://llm-prices.generality.org/
Requires Node 26 or later for the toolchain (nvm use reads .nvmrc); npm install refuses older versions. The Worker itself runs on workerd.
# Install dependencies
npm install
# Run locally
npm run devOn first run, visit http://localhost:8787 — the table will be empty until data is loaded. Trigger a data refresh by calling the scheduled handler (wrangler dev supports this via the dashboard).
The repository is Generality-Labs/llm-prices. The Worker and its MODEL_PRICES KV namespace live in the Generality Labs Cloudflare account; wrangler.toml declares the production environment (account, namespace id, custom domain, six-hour refresh schedule). The top level of wrangler.toml is local development and tests only.
Pull requests and pushes to main run type-checking, the unit and worker-runtime tests, a wrangler deploy --dry-run bundle check, and the pre-commit stack (Biome, zizmor, actionlint, mdformat) through the shared worker-ci workflow.
Every push to main then deploys production through the shared worker-deploy workflow and smoke-tests GET /health. Two repository secrets are required:
CLOUDFLARE_API_TOKEN— an account-owned token with Workers Editor at the Workers product scope, plus Zone > Zone > Read and Zone > Workers Routes > Write ongenerality.org(wrangler resolves the zone for the custom domain on every deploy).CLOUDFLARE_ACCOUNT_ID— the Generality Labs account id.
Set the HEALTH_URL variable on the production GitHub environment to https://llm-prices.generality.org/health to enable the post-deploy check.
Template pin. .copier-answers.yml's _commit records the release tag of the template this repo is aligned with (currently v1.1.0); running uvx copier update moves it forward.
Cloudflare Workers Builds is no longer used. Disconnect the Builds connection on the llm-prices Worker (Settings > Builds) before merging a change that introduces the named environments: Builds runs a bare npx wrangler deploy, which now resolves the local-only top-level config and would create a stray llm-prices-dev Worker instead of deploying production.
Keep the personal-account Worker at https://llm-prices.llm-prices.workers.dev running with its KV namespace and six-hour refresh schedule until existing clients have migrated. Older installations of Inspect Costs Plugin use that endpoint, and their HTTP client does not follow redirects. The old endpoint must continue returning pricing responses directly. Merging the plugin's URL update alone does not update existing installations.
Authenticate with the Generality Labs Cloudflare account using npx wrangler login, or set CLOUDFLARE_API_TOKEN in your shell. Push the Worker's secrets with npm run secrets from a gitignored .dev.vars.production containing REFRESH_SECRET=... (see .dev.vars.example), then deploy:
npm ci && npm run typecheck && npm test && npm run deploynpm run deploy runs wrangler deploy --env production, the named environment in wrangler.toml that carries the real KV namespace, route, and cron trigger. A bare wrangler deploy (no --env) reads the top-level, local-only configuration instead and would create a separate llm-prices-dev Worker rather than touching production.
The Worker refreshes pricing automatically every six hours. To populate a new deployment immediately, push the secret with npm run secrets, which reads a gitignored .dev.vars.production file containing:
REFRESH_SECRET=your-refresh-secretThen run:
npm run refreshnpm run refresh reads the same .dev.vars.production file (falling back to the older .env location) and sends the secret in an Authorization header to https://llm-prices.generality.org/api/refresh.
Returns {"status":"ok"}. Used by the post-deploy smoke test.
Query parameters:
| Param | Description |
|---|---|
q |
Search model name or provider; supports wildcards like gpt-*-codex and multi-term queries like claude sonnet |
provider |
Filter by provider (e.g. openai, anthropic) |
mode |
Filter by mode (chat, embedding, completion, etc.) |
supports |
Comma-separated capabilities (vision, function_calling, reasoning, prompt_caching) |
max_input_cost |
Max input cost per token |
min_context |
Minimum context window (tokens) |
sort |
Sort field (e.g. input_cost_per_token, max_input_tokens) |
order |
asc or desc |
limit |
Results per page (default 100) |
offset |
Pagination offset |
List all available providers.
List all available model modes.
Cache metadata (last update time).
Export Inspect-compatible model pricing as JSON or YAML.
Query parameters:
| Param | Description |
|---|---|
model |
Inspect model name. Repeat the parameter to request multiple models. |
models |
Comma-separated Inspect model names. Alternative to repeated model parameters. |
format |
json or yaml (default json). Use yaml for --model-cost-config. |
The response format matches Inspect's ModelCost object shape:
inputoutputinput_cache_writeinput_cache_read
All values are returned in dollars per million tokens.
Examples:
curl "https://llm-prices.generality.org/api/inspect-costs?model=openai/gpt-4o&model=anthropic/claude-sonnet-4-5&format=yaml" -o pricing.yamlinspect eval ctf.py --model-cost-config pricing.yaml --cost-limit 2.00curl "https://llm-prices.generality.org/api/inspect-costs?models=openai/gpt-4o,google/gemini-2.5-pro,openrouter/gryphe/mythomax-l2-13b&format=json" -o pricing.jsonIf you want a Claude, GPT, or Gemini model but are not sure which exact model key to use, search the catalog first and pick the model yourself. You can also add sort (for example sort=key&order=desc or sort=input_cost_per_token&order=asc) to make the list easier to scan:
curl "https://llm-prices.generality.org/api/models?provider=anthropic&q=claude&sort=key&order=desc"curl "https://llm-prices.generality.org/api/models?provider=openai&q=gpt&sort=key&order=desc"curl "https://llm-prices.generality.org/api/models?provider=gemini&q=gemini&sort=key&order=desc"Then request Inspect-formatted pricing for the exact model you selected:
curl "https://llm-prices.generality.org/api/inspect-costs?model=anthropic/claude-sonnet-4-5&format=yaml" -o pricing.yamlIf you want to confirm which cached dataset key was matched, add debug=1:
curl "https://llm-prices.generality.org/api/inspect-costs?model=anthropic/claude-sonnet-4-5&format=yaml&debug=1" -o pricing-debug.yamlProvider naming notes:
- Use Inspect-style provider prefixes such as
openai,anthropic,google,openrouter,groq,ollama,bedrock,azureai,cf,fireworks,together, andperplexity. - The service maps common Inspect names to LiteLLM dataset keys where they differ, for example:
google/...->gemini/...azureai/...->azure_ai/...cf/...->cloudflare/@cf/...groq/...->xai/...is not applied;groq/...resolves against Groq dataset keys, whilegrok/...resolves against xAI keysfireworks/...->fireworks_ai/...together/...->together_ai/...
- If a model name cannot be resolved, the API returns a
400with the unresolved model names and the candidate dataset keys it tried.
OpenAPI 3.1 spec for tool/MCP integration.
Remote MCP server endpoint exposed over Streamable HTTP.
Available tools:
search_models— search and filter models using the same provider/mode/query/capability/cost/context filters as the REST APIexport_inspect_costs— export Inspect-compatible model cost config for one or more Inspect model names as JSON or YAMLlist_providers— list all known providerslist_modes— list all known model modesget_metadata— return the last refresh timestamp and total model count
For clients that support remote MCP directly, use:
https://llm-prices.generality.org/mcp
For clients that only support local stdio MCP, bridge with mcp-remote:
{
"mcpServers": {
"llm-prices": {
"command": "npx",
"args": [
"mcp-remote",
"https://llm-prices.generality.org/mcp"
]
}
}
}The /openapi.json endpoint can be used directly with clients that support OpenAPI-based tool import.
Example: find the cheapest chat models with vision support and 100K+ context:
GET /api/models?mode=chat&supports=vision&min_context=100000&sort=input_cost_per_token&order=asc&limit=10
MIT