Set Up SignalFlag with an AI Coding Agent
An AI coding agent can help you set up tests in SignalFlag quickly with minimal tokens to onboard and use our platform to its ideal capacity with useful graphs and metrics. You point it at a project that already produces results (simulation runs, test reports, recorded robot logs, data files, model evaluations), and it does the setup work for you. The SignalFlag skills give your agent these functions:
- Reviews what your runs produce.
- Asks a few questions in plain language.
- Writes the code that sends your results to SignalFlag, and picks charts that fit your data.
- Sends a small trial first, checks that the numbers match your files, then sends the real results and gives you the links.
You don't need to know how SignalFlag is organized to start. This page covers installing the SignalFlag skills, connecting your agent to your SignalFlag account, and what to ask for.
Reading time: about 6 minutes.
Setup time: 15 to 45 minutes, depending on how many questions you choose to answer.
Before you start
- An AI coding agent that supports skills and remote MCP servers, such as Claude Code, installed and signed in.
- A SignalFlag account, and ideally the name of the SignalFlag project your results should go into. If you don't remember the name, say so. The agent notes it in the plan and looks it up after you approve, when it signs in.
- Python 3 on your machine. The agent creates a separate Python environment for the upload code, so your project's own packages aren't touched.
- A project with results on disk, such as a folder of runs, a test suite you run locally, recorded logs, or CSV or Parquet files.
Step 1: Install the SignalFlag skills
Skills teach your agent how to set up and use SignalFlag. They live in the signalskills repository and cover getting results you already have into SignalFlag. The MCP server also ships its own skills for running tests, diagnosing failures and other day-to-day work, listed in Agent Capabilities.
In Claude Code, the skills come as a plugin, which also adds the connection to SignalFlag.
- Open Claude Code in a terminal.
-
Type these two commands, one at a time:
Claude Code/plugin marketplace add resim-ai/signalskills /plugin install signalskills@signalskills -
Close Claude Code and open it again, so the plugin loads.
To check it worked, type /plugin. You should see signalskills in the list, enabled.
-
Clone the repository:
Shellgit clone https://github.com/resim-ai/signalskills.git ~/signalskills -
Link the skills into
~/.agents/skills, the shared skills folder that Codex, Cursor, Devin CLI and Devin Desktop all read:Shellmkdir -p ~/.agents/skills for s in ~/signalskills/skills/signalflag-*; do ln -s "$s" ~/.agents/skills/; doneLinking rather than copying lets a later
git pullupdate the skills in place. If your agent reads a different folder (see the table below these steps), put that folder in place of~/.agents/skillsin both lines. -
Restart your agent so it picks up the new skills.
| Agent | Skills folder in your home directory |
|---|---|
| Codex | ~/.agents/skills |
| Cursor | ~/.agents/skills or ~/.cursor/skills |
| Devin CLI | ~/.agents/skills or ~/.config/devin/skills |
| Devin Desktop | ~/.agents/skills |
| Gemini CLI | ~/.gemini/skills |
| GitHub Copilot CLI | ~/.copilot/skills |
| OpenCode | ~/.config/opencode/skills |
| Windsurf | ~/.codeium/windsurf/skills |
Devin in the cloud reads skills from the repositories it works in, not from your home directory, so this setup doesn't apply to it.
Step 2: Connect your agent to your SignalFlag account
The agent uses the SignalFlag MCP server to read your results back (statuses, numbers, charts) and confirm everything arrived correctly. The server can also set things up, such as projects and metrics configs, but launching, rerunning, canceling, archiving and deleting always come back to you as a draft to approve in the web app. See Approving consequential actions.
The plugin added this connection in Step 1.
- In Claude Code, type
/mcp. - Choose
signalflagfrom the list and select Authenticate. - Your browser opens. Sign in with your SignalFlag account and approve.
- Back in Claude Code,
signalflagshould now show as connected.
-
Add the SignalFlag MCP server to your agent's MCP configuration, using this URL:
Texthttps://bff.resim.ai/mcpThe MCP server guide has the configuration for several agents, including one-click setup for Cursor and VS Code. Any server name works, since the skills find the server by its tools.
-
Start the sign-in from your agent. Your browser opens. Sign in with your SignalFlag account and approve.
You only need to do this once.
Tip
Separately, when the agent first sends results, it may need you to sign in once more for uploads. It will post a link in the chat. Open it, approve, and the agent carries on by itself. If your agent can't run commands in the background, it asks you to run the sign-in command yourself instead. You never need to share a password with your agent.
Step 3: Start with your project
- In a terminal, go to the folder of the project that has your results.
- Start your agent there.
- Ask for what you want in your own words. See the examples in the next section.
You don't need to name the skills or use any special command, because the agent recognizes the request and starts the setup.
Example prompts
Pick the one closest to your situation, and change the details to match yours.
| Your situation | Try asking |
|---|---|
| You run simulations that save a folder per run | "Get our simulation runs into SignalFlag. They're in runs/." |
| You have a test suite (for example pytest) you run locally or in CI | "Make my pytest run report its results to SignalFlag." |
| You have recorded robot logs (for example ROS bags or MCAP files) | "Put these recordings into SignalFlag and show me how the robot did on each one." |
| You have data files (CSV, Parquet, HDF5) | "Upload our drive telemetry to SignalFlag so we can explore it later." |
| You train models and evaluate checkpoints | "Hook my checkpoint evaluations up to SignalFlag and rank the checkpoints." |
| You want results to upload every time your test script runs | "Wire SignalFlag into run_suite.py so every run reports automatically." |
| You replay recorded camera or sensor data through your perception stack | "Wire SignalFlag into our perception replay harness. I want to see what got worse per class and range, and scrub to the missed frames." |
| You compare model versions on a labeled photo set | "Evaluate each new detector version on our photo set in SignalFlag and show me where it got better or worse." |
| Your results already upload, but something looks wrong | "Our tests already upload to SignalFlag but some metrics show errors. Am I organizing this right?" |
| It's already set up and you want to improve something | "Tune the controller gain and show me the progress in SignalFlag." |
Once your results are in SignalFlag, the MCP Server guide has prompts for running, diagnosing and comparing tests.
It helps to mention:
- Your checks after a run: "I check whether it reached the goal and how far off the path it went."
- Who will use the results: "My lead wants to see whether we're getting better over time." The agent sizes the charts to the reader, with a few headline numbers for a lead and more depth for an engineer debugging.
- Different kinds of tests: "These are regression tests, and these are known failures we're tracking", or "same stack, but on sim drives and real drives". The agent keeps each kind in its own test suite so comparisons stay like for like.
- Anything off-limits: "Don't change
run_tests.sh."
FAQ
How does setup work after I ask?
- The agent reads your project and your results before asking anything.
- It asks how involved you want to be.
- "You decide": the agent makes the choices from what it finds, and only asks which SignalFlag project to use. Before anything is sent, it stops once and shows you everything it chose. This suits people new to SignalFlag.
- "Walk me through it": the agent asks one question at a time, each with a few options and its recommendation. This suits people who want to understand and shape each decision.
- It writes a short plan in your project, named after your type of test, for example
docs/signalflag/route-test-brief.md. It describes your tests, how they map into SignalFlag, and which charts and pass/fail checks it will set up. You can read and edit it. - It reads your data and proposes metrics and charts. It adds checks that flag a bad run and point to the likely cause, puts video or a replay first where your runs have them, and keeps everything else available for digging in later. It shows you the plan and waits for your go. Nothing is installed and nothing is signed in to before that.
- It builds the upload code in your project, in its own Python environment.
- It sends a small trial first to a separate test area in SignalFlag, and checks that the numbers there match your files.
- It sends your real results and gives you links to the results page and the dashboard.
If your project already sends results to SignalFlag, the agent reads your existing setup first, keeps what you built, answers your question, and fixes only what's wrong.
Which files does the agent add to my project?
- A plan file for each type of test,
docs/signalflag/<test type>-test-brief.md. - Upload code, usually a small script or a few files, plus a settings file in
.resim/metrics/. - A separate Python environment (for example
.venv-signalflag/), which it adds to.gitignore.
It only changes your existing code if you agree to it, and says exactly what it changed. Nothing is committed for you, so review the changes and commit them when you're happy.
How do I keep the skills up to date?
We improve the skills regularly.
Type these in Claude Code, then close and reopen it:
/plugin marketplace update signalskills
/plugin update signalskills@signalskills
Type /plugin to see which version you have.
Pull the latest version into your clone, then restart your agent:
git -C ~/signalskills pull
If you copied the skill folders instead of linking them, copy them again after pulling.
Which SignalFlag terms will I see?
| Word | Meaning |
|---|---|
| Project | Your team's space in SignalFlag. Your admin usually sets it up. |
| Branch | A named line of results that get charted together over time, for example all the rounds of one test suite. The agent suggests a name. |
| Batch | One upload, for example one round of your test suite. |
| Test | One item inside a batch, for example one scenario, one route or one seed. Keeping the same name across rounds lets SignalFlag compare them. |
| Metric | A number, chart or video shown for a test or batch, sometimes with a pass/fail check attached. |
| Dashboard | Charts across many batches, for example "how is this getting better over time". |
The SignalFlag server doesn't show up in my agent
Restart your agent after installing the skills or adding the server. In Claude Code, run /plugin and check that signalskills is enabled, then look for signalflag in /mcp. In other agents, check that the server URL in your MCP configuration is https://bff.resim.ai/mcp. The MCP Server troubleshooting section explains the error codes the server returns.
The sign-in link expired
Links last about 15 minutes. Tell the agent "the login link expired" and it will post a new one.
The dashboard looks empty or out of date
Dashboards refresh on their own and can take a few minutes after an upload. Open it again a little later. With only one round of results, trend charts show a single point until the next round.
The agent asks about a run it can't identify
If a run is missing information (for example which model checkpoint it used), the agent leaves it out rather than guessing. Tell it what the run was, or let it skip that run.
I see extra test areas in my SignalFlag project
During setup the agent sends trial uploads to separate areas whose names end in -lowres or -scratch, so your real results stay clean. You can archive them once you're happy.
I'm not sure about something the agent chose
Ask it, for example "Why did you pick that branch name?" or "Change the warning limit to 5 cm." Every choice is written in the plan file (docs/signalflag/<test type>-test-brief.md), so you can also edit it there and ask the agent to update the setup.
Getting help
Contact your SignalFlag representative or our support team, and include what you asked the agent and what happened. If your agent can export the conversation, attach that too. In Claude Code, /export saves it to a file.