See SignalFlag Working in Five Minutes
This tutorial runs a demo that fills a SignalFlag project with real test results, so you can see what the platform does before connecting anything of your own. No Docker, no container registry, no config to write. One command.
Time: about 5 minutes, plus a few minutes of waiting for metrics.
Before you start:
- You have a SignalFlag account. Sign up at app.signalflag.ai/signup if you don't.
- You have Python 3.10 or newer.
- You can open a browser to authenticate.
What you're building
The demo replays test data from a real navigation project into your own SignalFlag project. When it finishes you'll have:
- Two batches of tests, one per version of the navigation stack. The same scenarios ran in both, so every test has a counterpart in the other batch.
- Metrics on every test: speed and localization error over time, goal state timelines, pass/fail checks, and more, with camera footage on the tests that recorded it.
- The files each run produced, attached to the test they came from, so the raw material sits next to the charts drawn from it.
- An A/B comparison between the two batches, showing which scenarios regressed and which improved.
- A dashboard comparing the two versions across every run of them.
That's the navigation demo. A second one replays MuJoCo policy evaluations — see Choosing a demo below.
Step 1: Install the SDK
pip install signalflag
The SDK used to be published as resim-sdk. That name still installs and works; signalflag is the same package and the one to use from here on.
Step 2: Run the demo
signalflag-demo --demo navigation
The package also installs resim-demo, an alias for the same command, so anything already scripted against the old name keeps working.
First it downloads the test data it replays, reporting progress as it goes. That happens once — later runs reuse the cached copy. Then it asks you to authenticate in a browser:
Authenticating by Device Code
Please navigate to: https://resim.us.auth0.com/activate?user_code=XXXX-XXXX
Visit that URL, sign in, and the demo continues on its own. Subsequent runs reuse the cached token.
Each demo works in a project of its own, creating it if it doesn't exist, so it won't touch your real projects. The navigation demo uses one called SignalFlag SDK Demo. To use a different one:
signalflag-demo --demo navigation --project-name "my project"
Choosing a demo
--demo picks which demo to run. There's no default — running signalflag-demo on its own lists the ones that ship:
signalflag-demo
--demo |
What it replays |
|---|---|
navigation |
A robot running errands around a hospital. Dense telemetry across 34 scenarios, covering almost every chart type SignalFlag renders. |
mujoco |
An ALOHA bimanual manipulation policy evaluated in MuJoCo, one test per cube placement, compared across two policy builds. |
Each lands in its own project and on its own branch, so running both leaves you with two independent sets of results to look at. Bootstrap a Demo Project with the CLI details what each one contains, metric by metric.
Step 3: Follow the four links
When the demo finishes it prints where to go:
SignalFlag SDK demo complete.
Batch A (baseline, nav-v2.0.0) https://app.signalflag.ai/projects/.../batches/...
Batch B (candidate, nav-v3.0.0) https://app.signalflag.ai/projects/.../batches/...
A/B comparison https://app.signalflag.ai/projects/.../batches/.../compare-batch/batch/...
Trends dashboard https://app.signalflag.ai/projects/.../dashboards/...
Metrics are computed after each batch closes, which takes a few minutes.
Reload these pages when it finishes.
Metrics are computed once each batch closes, so the pages look empty at first. Give it a couple of minutes and reload.
Batch A and Batch B
Each batch page lists its tests with a pass, warning, or blocker status.
Open any test to see its metrics: speed over time, distance to each goal, what the navigation stack was doing as a state timeline, a table judging whether the localizer's confidence was honest, and on some tests the onboard camera footage. The Events tab shows the moment the robot reached its goal, and the Logs tab holds the files that run produced.
Batch-level metrics sit on the batch page itself, aggregating across every test.
The A/B comparison
This is the page that answers "did this version make things better or worse". Metrics are matched by name between the two runs, so each chart from Batch A sits next to the same chart from Batch B. Merged Metrics goes further and draws both runs on the same axes.
The Tests tab is the fastest read: it groups the scenarios by how their outcome differs between the two batches, so the regressions and the fixes are each their own list. See the A/B comparison guide for the full picture.
The dashboard
The batches show you two runs. The dashboard aggregates across them, grouped by build version: how each version's tests turned out, its mean localization error, and how many tests have run. Run the demo again and those bars take in the new runs too.
Next steps
- See what else the demos contain. Bootstrap a Demo Project with the CLI documents every flag of the command and details both demos: their topics, every metric, their status checks and dashboards.
- Emit your own data. Get Your First Metrics in SignalFlag walks through the same API the demo uses, starting from an empty script.
- Write your own metrics. Start with the metrics guide, and develop them against a debug dashboard before running a batch.