Skip to content

Set up Evals

An Eval suite holds test inputs and checks for an agent. Open Evals in the project to create a suite and its fixtures.

  1. Choose New suite and name it.
  2. Enter the Blueprint ID to identify the agent. A release-gate suite requires it.
  3. Choose the Target, concurrency and timeout. Start with the draft and one fixture.
  4. Add default checks that every fixture should meet. A text check is a useful first check.
  5. Save the suite. Expand Show fixtures, then choose Add fixture.
  6. Name the fixture and enter its Input. A plain message is enough for a first fixture.
  7. Add any fixture-specific checks and save it. These join the suite’s default checks.

Tool mocks (JSON) supplies test outputs for tool calls. The form warns that a tool call without a mock fails the fixture. Use fictional data in fixtures and mock outputs.

The form also offers judges, trajectory matching and similarity scoring. Judges add model calls. Similarity needs an expected output and an OpenAI key. Trajectory matching needs captured tool calls.

Use Capture from run to open the capture form when a recorded trace should become a fixture. Review captured customer content before keeping it as test data.

Run opens a confirmation with the fixture count, model choice and optional published version. Scroll through the model choices while Run suite and Cancel stay visible. Keep Blueprint default to use the agent’s configured model, or select an available model for this run. The dialog warns that fixtures and judges use real model calls charged to your model keys.

Choose Run suite to submit the run. The confirmation stays open while it starts, then opens the admitted run. If submission fails, the dialog shows an error and keeps your choices so you can submit again. The run detail and history update automatically when the run finishes, including while you use another browser tab. Check its status and results before deciding whether the suite passed.

The page has Run history and Compare controls. Compare remains disabled until there are at least two recorded runs.

These defaults come from the limits catalog. Organization overrides can change the values marked Yes.

Limit Default Support can change it
Fixtures an eval suite may hold. A run executes all of them 25 Yes
LLM judges one eval fixture may run (suite defaults plus the fixture’s own) 5 Yes
Agentic steps one eval judge may take 8 Yes
Longest time one eval fixture may run before it is stopped 10 min Yes
Eval fixtures run in parallel 16 No
Stored traces one scoring call may score 100 Yes