API keys and the REST API
Create an API key in the dashboard, save your registered server's scores from CI and run its evaluations with the REST API, its endpoints, limits and errors.
On this page
The REST API at https://api.forecall.dev saves and reads the scores of the servers you
registered in the dashboard, and runs their evaluations.
Its OpenAPI description is at
api.forecall.dev/v1/openapi.json.
API keys
Owners and admins create keys in the dashboard under API keys. A key is for one purpose:
- CI: saves lints from your build. It can do nothing else.
- Scripts: for your own scripts. It can save lints too, and is the key for evaluations.
- Agents: for your AI agent's use of the Failure KB through
mcp.forecall.dev. It looks up and reports failures and cannot call this REST API.
The key is shown once, when it is created; Forecall keeps only its hash. If you lose it, revoke it
and create another. Send it in every request. GET /v1/me answers with the key's organization and
kind:
curl -H "Authorization: Bearer $FORECALL_API_KEY" https://api.forecall.dev/v1/me
Save a lint from CI
POST /v1/servers/{id}/lint takes a server's tools/list as its body, in any of the
accepted formats, up to 1 MiB and 200 tools. {id} is the server's
id, the last part of its page's URL in the dashboard (/servers/{id}). Add ?label= to name the
version, up to 50 characters.
As in the dashboard, a version is added only when the tools differ from the latest version, and
each call saves a lint. The answer has the report, the same as the linter's, and what changed from
the version before (changes, or null when there is no scored version before).
In GitHub Actions, with the key in a secret:
- run: npx forecall@0.2 dump -o tools.json -- node dist/server.js
- run: >
curl -fsS -X POST --data-binary @tools.json
-H "Authorization: Bearer ${{ secrets.FORECALL_API_KEY }}"
"https://api.forecall.dev/v1/servers/${{ vars.FORECALL_SERVER_ID }}/lint?label=${{ github.sha }}"
GET /v1/servers/{id}/lint/latest reads the latest lint of the latest version.
Evaluations
An evaluation asks models which of your server's tools they would call for each case, an utterance written for a tool, and checks the arguments. It comes with the Pro and Team plans and needs a Scripts key. Evaluate your server with models explains the cases, the metrics and the units, and how to run evaluations in the dashboard.
POST /v1/servers/{id}/evaluations/estimate answers with the units evaluating the server's latest
version would take, and takes none. The body is optional JSON:
| Field | What it is | Default |
|---|---|---|
models |
The models, as vendor/name: anthropic/claude-opus-5-5, anthropic/claude-sonnet-5-5, openai/gpt-6-astra, openai/gpt-6.1-sol (google/gemini-3.8-flash is paused for now) |
anthropic/claude-sonnet-5-5, openai/gpt-6.1-sol |
distractors |
Bundled public servers whose tools the models see too: bhived, context7, debugbase, filesystem, firecrawl, memory, notion, playwright |
None |
lang |
The utterances' language: en, or ja on Team |
en |
Each tool gets 8 cases on Pro and 20 on Team, with a few that call for no tool and a few aimed at
tools that are easily confused. Writing the cases is counted too, unless an evaluation of the same
version and settings has them already; a model's results the organization already has are not
counted again. The answer also shows this month's units left and the credits (balance).
POST /v1/servers/{id}/evaluations takes the same body, with an optional name, takes the
estimate's units and starts the evaluation. It answers 202 with the evaluation's id, the units
of each model's run, and a Location to follow. When the month's units and the credits fall short,
it answers 402 and starts nothing. If the evaluation fails, the units of the cases it did not
finish come back as credits, valid for six months.
curl -fsS -X POST -H "Authorization: Bearer $FORECALL_API_KEY" \
-d '{"models": ["anthropic/claude-sonnet-5-5", "openai/gpt-6.1-sol"]}' \
"https://api.forecall.dev/v1/servers/$SERVER_ID/evaluations"
GET /v1/evaluations/{id} shows an evaluation: its status (queued, running, done or
failed), its cases, and for each model the cases answered so far and, once its run is done, the
metrics: selectionAcc (the expected tool called), argValidity (valid arguments when it was),
falseCallRate (a tool called when none should be) and confusion (the tools called instead).
Limits
Each key may make a number of requests a minute, set by the organization's plan; past it, the API
answers 429 with Retry-After. Saving a lint uses one unit from the organization's allowance for
the month (UTC), counted for each kind of key. When the allowance and the organization's credits
are used up, saving answers 402 and saves nothing. An evaluation takes the units of its estimate
from the allowance of Scripts keys. Reading and estimating use no units. The dashboard's
API keys page shows this month's use.
Errors
Every error has the same shape, {"error": {"code": "…", "message": "…"}}, with detail when the
input says where it went wrong (a path, the keys to rename, the names used twice):
| Status | Codes |
|---|---|
| 401 | unauthorized: no key, a wrong one, or a revoked one |
| 402 | quota_exceeded, plan_has_no_evaluations, payment_failed (while the organization's last payment has failed: fix the card in the dashboard's Plan and billing) |
| 404 | not_found: no such server or evaluation in the key's organization, no version or no lint yet |
| 413 | input_too_large |
| 422 | invalid_json, unrecognized_shape, invalid_tool, too_many_tools, snake_case_keys, duplicate_names, invalid_label, invalid_request, unknown_models, unknown_distractors, lang_not_in_plan |
| 429 | rate_limited |