Trajecta is regression testing for browser agents. Define a task suite, run your agent against it, and diff the trajectories across runs and model versions. You will know the moment a model upgrade silently breaks your agent.
Three steps. No database, no hosted service to babysit. Runs locally today.
A plain YAML file: starting URL, goal in words, and assertions like URL reached, text present, or extracted data matching a schema.
One CLI command runs every task and records the full trajectory: actions, DOM snapshots, tool calls, tokens, latency, errors.
Compare any two runs. New failures, fixed tasks, changed action sequences, cost and latency deltas. Exits non-zero on regressions, so it drops straight into CI.
Baseline: 3/3 passing. Candidate model: one task silently changed behavior.
# tasks.yaml - name: extract-gdp-table startUrl: https://en.wikipedia.org/... goal: Extract the top 5 countries by GDP assertions: - type: schema # extracted rows match expected shape - name: pricing-page startUrl: https://github.com/pricing goal: Reach the pricing page assertions: - type: url_contains value: /pricing
$ trajecta diff --base v1 --candidate v2 PASS extract-gdp-table PASS pricing-page FAIL checkout-flow new failure action sequence changed: 14 -> 22 steps tokens: 8.1k -> 12.4k (+53%) latency: 41s -> 68s exit 1: 1 new failure
The candidate model took a different, longer path and failed a task that used to pass. Without the diff, you would have shipped it.
If any of these sound familiar, Trajecta is for you.
Every diff also saves a machine-readable artifact at runs/diffs/<base>--vs--<candidate>.json, so CI or your own tooling can act on it.
{ "base": "baseline", "candidate": "candidate", "newFailures": ["github-pricing-plans"], "fixed": [], "changedSequences": [ { "task": "wikipedia-gdp-top5", "base": ["goto", "observe", "extract"], "candidate": ["goto", "scroll", "observe", "extract"] } ], "costDeltas": [ { "task": "wikipedia-gdp-top5", "tokensDelta": 150, "latencyDeltaMs": 750 } ], "totalsDelta": { "tokens": 120, "latencyMs": 1700 } }
Short answers to the questions everyone asks first.
It records everything your browser agent does on a task suite, then diffs trajectories across runs and model versions. When a new model takes a worse path or fails a task that used to pass, you see it as a clear diff instead of finding out in production.
Spot checks don't scale and don't catch slow drift. Trajecta replays the whole suite on every change and exits non-zero on regressions, so it plugs into CI and answers "did the agent get worse this week?" without anyone watching runs by hand.
Trajecta is in early access and free to try right now. The local CLI runs on your machine with no hosted service. Longer term pricing is still being figured out, early users will have a say in it.
Teams shipping browser agents who change models, prompts, or tools and want to know what broke before users do. If you have ever upgraded a model and quietly shipped worse behavior, that is the pain this fixes.
Trajecta is in early access. Tell us what your agents do and we will get you running.
Get early accessSDE at Amazon, building MCP servers and evaluation frameworks for production agents. Trajecta is the tool I wished existed every time a model upgrade quietly changed what my agents did.
LinkedIn · GitHub · Trajecta source code · [email protected]