# API keys and the REST API

Create an API key in the dashboard, save your registered server's scores from CI and run its evaluations with the REST API, its endpoints, limits and errors.

The REST API at `https://api.forecall.dev` saves and reads the scores of the servers you
[registered in the dashboard](https://forecall.dev/en/docs/dashboard), and runs their evaluations.
Its OpenAPI description is at
[api.forecall.dev/v1/openapi.json](https://api.forecall.dev/v1/openapi.json).

## API keys

Owners and admins create keys in the dashboard under **API keys**. A key is for one purpose:

- **CI**: saves lints from your build. It can do nothing else.
- **Scripts**: for your own scripts. It can save lints too, and is the key for
  [evaluations](#evaluations).
- **Agents**: for your AI agent's use of the Failure KB through `mcp.forecall.dev`. It looks up
  and reports failures and cannot call this REST API.

The key is shown once, when it is created; Forecall keeps only its hash. If you lose it, revoke it
and create another. Send it in every request. `GET /v1/me` answers with the key's organization and
kind:

```sh
curl -H "Authorization: Bearer $FORECALL_API_KEY" https://api.forecall.dev/v1/me
```

## Save a lint from CI

`POST /v1/servers/{id}/lint` takes a server's `tools/list` as its body, in any of the
[accepted formats](https://forecall.dev/en/docs/tools-list#formats), up to 1 MiB and 200 tools. `{id}` is the server's
id, the last part of its page's URL in the dashboard (`/servers/{id}`). Add `?label=` to name the
version, up to 50 characters.

As in the dashboard, a version is added only when the tools differ from the latest version, and
each call saves a lint. The answer has the report, the same as the linter's, and what changed from
the version before (`changes`, or `null` when there is no scored version before).

In GitHub Actions, with the key in a secret:

```yaml
- run: npx forecall@0.2 dump -o tools.json -- node dist/server.js
- run: >
    curl -fsS -X POST --data-binary @tools.json
    -H "Authorization: Bearer ${{ secrets.FORECALL_API_KEY }}"
    "https://api.forecall.dev/v1/servers/${{ vars.FORECALL_SERVER_ID }}/lint?label=${{ github.sha }}"
```

`GET /v1/servers/{id}/lint/latest` reads the latest lint of the latest version.

## Evaluations

An evaluation asks models which of your server's tools they would call for each case, an
utterance written for a tool, and checks the arguments. It comes with the Pro and Team plans and
needs a **Scripts** key. [Evaluate your server with models](https://forecall.dev/en/docs/evaluations) explains the
cases, the metrics and the units, and how to run evaluations in the dashboard.

`POST /v1/servers/{id}/evaluations/estimate` answers with the units evaluating the server's latest
version would take, and takes none. The body is optional JSON:

| Field | What it is | Default |
|---|---|---|
| `models` | The models, as `vendor/name`: `anthropic/claude-opus-5-5`, `anthropic/claude-sonnet-5-5`, `openai/gpt-6-astra`, `openai/gpt-6.1-sol` (`google/gemini-3.8-flash` is paused for now) | `anthropic/claude-sonnet-5-5`, `openai/gpt-6.1-sol` |
| `distractors` | Bundled public servers whose tools the models see too: `bhived`, `context7`, `debugbase`, `filesystem`, `firecrawl`, `memory`, `notion`, `playwright` | None |
| `lang` | The utterances' language: `en`, or `ja` on Team | `en` |

Each tool gets 8 cases on Pro and 20 on Team, with a few that call for no tool and a few aimed at
tools that are easily confused. Writing the cases is counted too, unless an evaluation of the same
version and settings has them already; a model's results the organization already has are not
counted again. The answer also shows this month's units left and the credits (`balance`).

`POST /v1/servers/{id}/evaluations` takes the same body, with an optional `name`, takes the
estimate's units and starts the evaluation. It answers `202` with the evaluation's `id`, the units
of each model's run, and a `Location` to follow. When the month's units and the credits fall short,
it answers `402` and starts nothing. If the evaluation fails, the units of the cases it did not
finish come back as credits, valid for six months.

```sh
curl -fsS -X POST -H "Authorization: Bearer $FORECALL_API_KEY" \
  -d '{"models": ["anthropic/claude-sonnet-5-5", "openai/gpt-6.1-sol"]}' \
  "https://api.forecall.dev/v1/servers/$SERVER_ID/evaluations"
```

`GET /v1/evaluations/{id}` shows an evaluation: its status (`queued`, `running`, `done` or
`failed`), its cases, and for each model the cases answered so far and, once its run is done, the
metrics: `selectionAcc` (the expected tool called), `argValidity` (valid arguments when it was),
`falseCallRate` (a tool called when none should be) and `confusion` (the tools called instead).

## Limits

Each key may make a number of requests a minute, set by the organization's plan; past it, the API
answers `429` with `Retry-After`. Saving a lint uses one unit from the organization's allowance for
the month (UTC), counted for each kind of key. When the allowance and the organization's credits
are used up, saving answers `402` and saves nothing. An evaluation takes the units of its estimate
from the allowance of Scripts keys. Reading and estimating use no units. The dashboard's
**API keys** page shows this month's use.

## Errors

Every error has the same shape, `{"error": {"code": "…", "message": "…"}}`, with `detail` when the
input says where it went wrong (a path, the keys to rename, the names used twice):

| Status | Codes |
|---|---|
| 401 | `unauthorized`: no key, a wrong one, or a revoked one |
| 402 | `quota_exceeded`, `plan_has_no_evaluations`, `payment_failed` (while the organization's last payment has failed: fix the card in the dashboard's Plan and billing) |
| 404 | `not_found`: no such server or evaluation in the key's organization, no version or no lint yet |
| 413 | `input_too_large` |
| 422 | `invalid_json`, `unrecognized_shape`, `invalid_tool`, `too_many_tools`, `snake_case_keys`, `duplicate_names`, `invalid_label`, `invalid_request`, `unknown_models`, `unknown_distractors`, `lang_not_in_plan` |
| 429 | `rate_limited` |
