Agent Capabilities
The SignalFlag MCP server gives your agent 126 tools. You do not need to learn them.
This page covers what you can ask an agent to do, which tools it will reach for, and where it has to stop and wait for you.
Setting up the connection
This page assumes your agent is already connected. For Claude Code, Cursor, the claude.ai connector and CI credentials, see MCP Server.
Three kinds of tool
Every tool falls into one of three classes, and the class decides whether your agent acts or asks.
| Class | What happens | Covers |
|---|---|---|
| Read | Runs immediately | 77 tools. Listing and fetching anything, log contents, metrics, analyses, quota |
| Write | Runs immediately, and is recorded | 42 tools. Setting things up: projects, branches, builds, systems, experiences, tags, test suites, metrics configs, dashboards, notes |
| Draft | Never runs | 6 tools. Launching, rerunning, canceling, archiving, deleting, and changing agent instructions |
One more sits outside the three: steer_ui_session moves a view in a browser tab you have
open, and is covered under Working alongside you in the app.
Your agent can create an experience or push a metrics config on its own. It cannot launch a batch, rerun a job, cancel a run, archive an experience or delete anything. Those produce a draft that you review and submit in the web app.
Every write and every draft is recorded against the account whose token the agent is using, and you can list that history at any time.
Skills
Skills are task instructions the server ships to your agent, so you do not have to explain
how SignalFlag works every time. Your agent fetches one with get_skill after orienting
itself, without being asked. See MCP Server for what that looks
like from your side.
| Skill | For |
|---|---|
onboard |
Getting a repository from nothing to a first drafted batch |
run-tests |
Choosing a build and suite and drafting a run |
diagnose-batch |
Working a failing batch back to a root cause |
regression-check |
Comparing a branch against main |
author-metrics-config |
Writing or fixing a metrics config |
dashboards |
Building and refreshing dashboards |
stage-experiences |
Bulk-registering and tagging experiences |
containerize-sim |
Packaging a simulator into a build image |
fleet |
Checking hardware agents and queues |
drafts |
Working with actions that are waiting for a human |
ui-sessions |
Working alongside you in an open browser tab |
Capabilities by task
Getting oriented
Ask what projects exist, what changed recently, or where something lives.
list_projects, get_project, list_branches, list_builds, get_build,
list_systems, get_system.
get_project also returns your project's agent instructions, which is how you give every
agent working on a project the same standing context.
Setting up a project to test
Ask it to register experiences from a bucket, tag them, and assemble a suite.
upsert_experience, validate_experience_location, upsert_experience_tag,
add_experience_tags, create_test_suite, revise_test_suite, register_build,
upsert_system, create_branch.
These run without asking you. They are upsert-by-name, so re-running the same request updates rather than duplicating, which makes them safe to repeat when a setup script is half-finished.
Running tests
Ask it to run a suite against a build.
draft_launch_batch, then you submit.
Your agent assembles the run, validates the ids, and hands you a link. You open it, see a prefilled form with a banner naming the agent and what it intends to do, change anything you like, and submit. The agent can then read back what you actually ran.
Seeing how a run went
Ask how the last batch did, or whether a suite is passing.
list_batches, get_batch, list_jobs, get_job, get_metrics_summary,
get_metric_detail, get_metric_chart, list_batch_runs, get_batch_usage.
get_batch accepts a wait, so an agent can watch a running batch to completion and
summarize it when it lands rather than polling in a loop.
In claude.ai and Claude Desktop, get_metric_chart returns a chart your agent can render
inline. In a terminal-based coding agent it will describe the figure instead.
Diagnosing a failure
Ask why a test failed.
list_batch_errors, get_job, read_log, list_logs, get_log, list_events,
get_event, request_log_analysis, get_log_analysis, get_batch_analysis.
Error codes distinguish problems in your code and data from problems on our side, so an
agent can tell whether to fix something or simply rerun. read_log reads a window of a log
with search and tailing, so a large log does not have to be pulled down whole.
Catching regressions
Ask whether a branch broke anything relative to main.
get_batch_suggestions finds the right baseline to compare against, then compare_batches
returns the tests whose status changed.
Authoring metrics and dashboards
Ask it to add a metric, fix a failing one, or build a dashboard.
get_metrics_config, get_metrics_config_schema, list_topics, get_topic_schema,
preview_topic_data, preview_metric, validate_metrics_config, push_metrics_config,
list_chart_templates, list_metrics_sets, validate_status_query, list_dashboards,
get_dashboard, upsert_dashboard, refresh_dashboard.
Your agent can validate and preview a metric against real data before pushing anything, so ask it to preview before it pushes.
Asking questions of your data
Ask a question that no existing metric answers.
query_emissions runs read-only SQL over your emitted data. preview_topic_data samples
a topic when you need to see its shape first.
These cost money to run, so they are budgeted. See Limits below.
Capturing what you learned
Ask it to write down what it found so the next investigation starts further along.
save_org_note, list_org_notes, get_org_note.
Notes an agent writes are marked unreviewed until a person confirms them, and agents treat unreviewed notes with suspicion. Confirm the ones you want relied on.
Watching your fleet
Ask about hardware agents and queue depth.
list_agents, get_agent, get_pool_labels.
Working alongside you in the app
An agent acting for you can see the SignalFlag tabs you have open, and nothing else. Once it can see a tab, it can move that tab around so the two of you are looking at the same thing.
Ask for it in the terms you would use with a person.
| Try asking | What happens |
|---|---|
| "What am I looking at?" | It reads the open tab and answers from the same data you can see |
| "Take me to the failing job" | The tab navigates to that page |
| "Show me the metrics tab" | The tab switches tabs where the page has them |
| "Tell me when this batch finishes" | It waits for the page to change, then shows you a toast |
list_ui_sessions, get_ui_session and steer_ui_session are the tools behind those.
Look for the pill at the bottom left of the app, labeled "Agent steering on" or "Agent steering off". It is on by default, so an agent you are already talking to can move your view. Switch it off and the server refuses commands for that tab. The setting is per tab and survives a reload, so turning it off in one tab does not affect another and does not quietly come back.
Four things bound it. Steering only moves a view; forms, launches and deletions stay behind your own click. Every command shows a toast naming the agent that sent it. An agent sees only the tabs belonging to the account whose token it holds. And machine credentials cannot see your tabs at all, so a CI pipeline can never look at your browser.
Auditing what agents have done
Ask what the agents working in your organization have actually changed.
list_agent_actions returns the full trail with who requested each item, recorded against
the account whose token the agent was using. It covers the writes that happened without
asking you as well as the proposals that waited. get_draft returns the state of any single
proposal.
Approving consequential actions
Six actions never happen on their own. Your agent proposes them, and you approve or decline them in the web app.
| Your agent proposes | Tool |
|---|---|
| Launching a build against experiences, tags or a test suite | draft_launch_batch |
| Rerunning a whole batch, or named jobs within it | draft_rerun |
| Canceling a running batch | draft_cancel |
| Archiving entities of one type | draft_archive |
| Permanently deleting one entity | draft_delete |
| Replacing a project's agent instructions | draft_update_instructions |
Everything else an agent can write, it simply writes. There is no permission that turns this off and no override for trusted agents.
Approving one
You get a notification in the web app and a link. Opening it takes you to the form for that action, already filled in, with a banner naming who proposed it and what it will do.
Change anything you want to change, then submit. The action runs as if you had filled in the form yourself, and your agent can see what you changed and carries on against what you actually ran.
If you do not want it, close it. A proposal you never submit expires after seven days. An agent can also withdraw its own, so one you were about to look at may occasionally disappear.
Proposals from CI
A pipeline running on a machine credential can prepare work but cannot approve it. When a machine proposes something it names a reviewer, and that person submits it in the web app. Your pipeline can register the build, assemble the suite and propose the run, and a human decides whether to spend the compute.
Why one is waiting, or is not
| State | Meaning |
|---|---|
draft |
Waiting for a person |
submitting |
Someone has submitted it and it is being carried out |
submitted |
Done, the action ran |
discarded |
Withdrawn by the agent, or declined |
expired |
Nobody acted within seven days |
stale |
Submitting it would no longer do anything |
A proposal goes stale when the world moves on, such as a cancel for a batch that has since finished.
Budgets
Two of the groups above cost money to run, so they are budgeted per organization per day:
metrics compute through query_emissions, preview_metric and preview_topic_data, and
token usage through request_log_analysis and request_batch_analysis. An agent that
exhausts one gets a clear error rather than a silent failure.
See MCP Server for the current limits.
When no tool fits
Two escape hatches cover cases the tool surface does not reach yet. call_customer_api
makes a read-only call to the public REST API, and query_graphql runs a read-only GraphQL
query. Both are available only over MCP. If you find yourself relying on them, tell us which
tool is missing.