Pro & Elite feature. Autoresearch is available on Simmer Pro and Elite plans (Elite includes everything in Pro). Free users get a 403 when calling autoresearch API endpoints.
Autoresearch lets your agent optimize its own trading skills. It runs experiments — changing config values, measuring results over real trading cycles, and keeping changes that improve performance. Think of it as automated A/B testing for your trading strategy.
Prerequisites
- Simmer Pro plan with a valid
SIMMER_API_KEY - simmer-sdk installed with at least one trading skill running on sim venue
- Node.js 18+ (for the MCP server)
- Git initialized in your skill workspace (autoresearch uses git for commit/revert)
How it works
- Init — Pick a skill and a metric (e.g., P&L, edge %, trade count)
- Run — Execute the skill with the new config for several trading cycles
- Log — Record results and decide: keep or revert. Keeps auto-commit to git.
- Backtest — Replay historical trades against new config thresholds (fast config tuning)
- Repeat — Try the next hypothesis
Install
CLAUDE.md.
Config
Configure autoresearch via environment variables:Running Autoresearch
Setup (once per optimization target)
- Pick a skill to optimize and a primary metric (usually P&L)
- Create a git branch:
git checkout -b autoresearch/<skill>-<date> - Read the skill source code thoroughly — understand what it does before mutating
- Write
autoresearch.md— a session spec describing the goal, metrics, how to run, and constraints - Write
autoresearch.sh— a single command that runs the skill for one cycle - Commit both files
- Call
init_experiment→ run the baseline withrun_experiment→log_experiment→ start looping
The experiment loop
Each iteration:- Hypothesize — what change might improve the metric?
- Mutate — change the skill’s code or config
- Run — call
run_experimentto execute the skill - Log — call
log_experimentto record the result (keep,discard, orcrash)
keep auto-commits to git. discard and crash auto-revert the working directory.
Use backtest_experiment for fast config exploration (seconds) before committing to live runs (minutes).
Key rules
- Never skip the baseline run. The first experiment establishes the reference point for all comparisons.
- Always log — even crashes. Crash data matters for confidence scoring and crash detection.
- Check confidence scores. ≥2× noise floor = improvement is likely real. under 1× = within noise. 1-2× = marginal, re-run to confirm.
- Code mutations beat config tuning. Structural changes (new data sources, different models, alternative strategies) find bigger wins than parameter sweaks.
- Keep ideas in
autoresearch.ideas.md. Promising but deferred optimizations go here.
When you’re done
Review the autoresearch git branch. Experiments that werekeep-ed are committed with result metadata in the commit message. Merge the branch (or cherry-pick specific experiments) into your main skill branch to lock in the improvements.
Tools
The MCP server registers four tools your agent can call:init_experiment
Configure an experiment session. Call again to start a new segment with a fresh baseline.
run_experiment
Execute a command (usually the skill), capture output and timing.
log_experiment
Record experiment results. keep auto-commits to git. discard/crash reverts working directory.
backtest_experiment
Replay historical trades against new config thresholds without live execution. Returns simulated P&L in seconds — use this for fast config tuning before committing to live experiments.
Backtest requires trades with
signal_data. Skills must pass structured signal data on client.trade() calls (SDK 0.9.17+). All official Simmer skills include signal_data as of March 2026.
Config threshold convention:
min_edge: 0.05→ only include trades wheresignal_data.edge >= 0.05max_probability: 0.85→ only include trades wheresignal_data.probability <= 0.85- Bare keys (e.g.,
edge: 0.10) → treated as min threshold
Signal Data
Skills can include structured signal data on each trade to enable backtest replay. This is optional — trades work fine without it — but required for thebacktest_experiment tool.
Additional skill-specific fields are freeform. Values must be strings or numbers (flat dict, no nesting).
Signal data is private — only visible to the trade owner via authenticated API calls. Never exposed publicly.
Session management
Your agent manages its own session state using the SKILL.md behavioral instructions (installed vianpx simmer-mcp install-skill). There is no /autoresearch command interface — the agent drives the loop autonomously.
- Resume: The agent reads
autoresearch.jsonlon startup and resumes where it left off. - New session: Call
init_experimentwith a new name to start a fresh segment (previous results are archived, not deleted). - Context compaction: If the agent’s context resets, it should re-read
autoresearch.mdandautoresearch.jsonlto restore state.
Safety features
Crash protection
- Baseline crash — If the very first experiment in a session crashes, autoresearch pauses automatically. This usually means the skill is misconfigured.
- Consecutive crashes — 3 crashes in a row triggers auto-pause. Your agent can’t run more experiments until the issue is investigated.
- Recovery — Call
init_experimentwith a new session name to clear the pause and start fresh.
Budget caps
Experiments are capped atAUTORESEARCH_MAX_EXPERIMENTS (default 50) per session. At 80% of the cap, your agent gets a warning. At the limit, run_experiment is blocked.
Set AUTORESEARCH_MAX_EXPERIMENTS=0 to disable the cap (not recommended for unattended agents).
Metric verification
The server cross-checks self-reported P&L metrics against the Simmer API. If the agent-reported metric diverges significantly from actual trade data, a warning is logged. This prevents metric gaming — the agent can’t inflate results by changing how metrics are calculated.Experiment persistence
Results are saved in two places:- Local JSONL —
autoresearch.jsonlin your working directory for offline access - Dashboard API — Synced to your Simmer dashboard (Pro users see an Autoresearch tab)
keep decisions so you can track what changed and roll back if needed.
API endpoints
These endpoints power the server’s sync. You don’t call them directly — the MCP server handles it./api/sdk/outcomes v2 fields
GET /api/sdk/outcomes returns both backward-compatible cash-flow fields (v1) and new settlement-accurate fields (v2). Use v2 fields for autoresearch metric verification and skill-health signals — they correctly attribute buys held to resolution.
SDK: client.get_outcomes(skill_slug=..., since=...)
Why v2 exists: v1
wins only count sell rows — a buy held to resolution and settled as a winner never appears. v2 settled_wins counts every resolved market correctly.
Legacy (v1 Plugin)
Legacy (v1 Plugin) — OpenClaw only
Legacy (v1 Plugin) — OpenClaw only
v1 was an OpenClaw plugin, not an MCP server. If you’re still running v1:Configure via v1 supports the
plugins.json:/autoresearch command interface:Upgrade to v2 — Install
simmer-mcp via npm (npm install -g simmer-mcp) and switch to the MCP config above. v2 works with OpenClaw, Hermes, and Claude Code.