Skip to content

MCP server

SignalFlag's MCP server lets an agent operate the platform on your behalf. That might be a coding agent in your repository, a chat client asking about results, or a step in your CI pipeline.

There is nothing to install and no API key to create. Your client prompts you to sign in through your browser on first use.

LLM-friendly docs

Pair the MCP server with our LLM-friendly documentation, published as plain markdown following the llms.txt convention:

  • /llms.txt, a curated index of the docs
  • /llms-full.txt, the full corpus as a single file
  • Any page is also available as raw markdown by appending .md to its URL (for example /setup/cli.md)

Point your agent at these alongside the MCP server for full context on the platform.

Reads, writes, and approvals

Your agent can read anything you can read, and it can set things up on its own. That covers creating projects and branches, registering builds and experiences, assembling test suites, pushing metrics configs, and building dashboards.

Your agent cannot launch a batch, rerun a job, cancel a run, archive anything, delete anything, or change a project's agent instructions. Those six actions produce a draft that you review and submit in the web app. The agent never executes them, no matter how it is prompted. See Approving consequential actions.

Every write and every draft is recorded against your account and can be listed at any time.

For the full picture of what an agent can do with the server, see Agent Capabilities.

Connect

Every client needs the same one thing, the server URL:

Text
https://bff.resim.ai/mcp

If you use Cursor or VS Code, one click does it:

Add to Cursor Add to VS Code

Your editor asks you to confirm, then prompts for sign-in on first use. The manual steps for every client are below.

Claude Code

Bash
claude mcp add --transport http signalflag https://bff.resim.ai/mcp

Your browser opens a consent page, then SignalFlag sign-in. Tokens refresh on their own.

If you set up SignalFlag with the signalskills plugin, it registers this server for you. See Set Up SignalFlag with an AI Coding Agent.

claude.ai

Go to Settings, then Connectors, then Add custom connector, and enter the server URL.

Use this client for asking questions about results rather than working in a repository. It can render charts inline.

Claude Desktop

Open Settings, then Developer, then Edit Config, and add the signalflag entry to your mcpServers:

JSON
{
  "mcpServers": {
    "signalflag": {
      "url": "https://bff.resim.ai/mcp"
    }
  }
}

Restart Claude Desktop. You will be prompted to authenticate through your browser on first use.

Cursor

Open Settings, then MCP, and click Add new global MCP server:

JSON
{
  "mcpServers": {
    "signalflag": {
      "url": "https://bff.resim.ai/mcp"
    }
  }
}

Cursor's callbacks are already registered.

CI and other machine credentials

A pipeline cannot open a browser, so it uses a machine credential. Ask us for one bound to your organization, then send the token as a bearer header:

Text
Authorization: Bearer <token>

Machine credentials get reads, writes and drafts, with two differences:

  • A draft created by a machine names a reviewer, and that person submits it in the web app. Your pipeline can prepare a run, but a human still starts it.
  • Machine credentials cannot see or steer your browser tabs.

Scopes

On first connection you consent to three scopes:

Scope Allows
read Listing and fetching anything your account can already see
write Setting up entities: projects, experiences, suites, metrics configs, dashboards
ui:steer Moving a view in a SignalFlag browser tab you have open, and showing you a toast

ui:steer lets an agent navigate a tab you already have open. It cannot submit anything, each tab has a steering toggle you control, and every command is recorded. It is still rolling out, so it may not be available to your organization yet. See Agent Capabilities.

Your first run

Ask your agent something about your tests and it should just work. There is nothing to configure and no first command for you to run.

If you want somewhere to start, pick the row that matches what you have today.

Where you are Try asking
Nothing set up yet "What projects do I have?" then "Create a project called nav-regression."
A project, no experiences "Register the scenarios in s3://my-bucket/nav/ as experiences and tag them nav."
Experiences and a build, nothing run "Build a test suite from the nav experiences and draft a run against my latest build."
Tests already running "How did my last batch do?" or "Why did the failing tests in it fail?"
Metrics you want to change "Add a metric tracking median goal distance per test, and validate it before pushing."

Ask the first question whatever stage you are at. It confirms the connection works and tells you what your agent can see.

Creating a project, registering experiences and pushing a metrics config all happen immediately. Running the suite does not. That comes back as a draft with a link, and nothing runs until you submit it.

If your results already live on disk, such as sim runs, test reports or recorded logs, the Set Up SignalFlag with an AI Coding Agent guide walks your agent through getting them in.

Rate limits and budgets

Limit Value
Calls per user 120 per minute
Calls per organization 600 per minute
Data queries per organization 200 per day
Log and batch analyses per organization 100 per day

Exceeding a per-minute limit returns HTTP 429 and a JSON-RPC error. Back off and retry after a short delay.

The per-day budgets cover the tools that run real queries or call a model. Exhausting one returns a budget_exhausted error rather than failing quietly.

Troubleshooting

Errors come back with a code your agent can act on.

Code Meaning
not_found The id does not exist, or is not in your organization
invalid_argument The arguments do not satisfy the tool's schema
insufficient_permission Your account cannot perform this operation
rate_limited Slow down and retry
budget_exhausted A daily query or analysis budget is spent
membership_required Your account is not a member of an organization yet

A missing authorization scope comes back as HTTP 403 with a WWW-Authenticate header, which compliant clients use to prompt you for consent again.

Two things to try when a session goes badly rather than errors outright.

If your agent is guessing at project names or calling tools at random, it skipped orienting itself. Tell it to run whoami first, then ask again.

If it reports that it cannot write anything, check whoami for your permissions. An account with read-only permissions sees every read tool and no others.