Set up Evals
An Eval suite holds test inputs and checks for an agent. Open Evals in the project to create a suite and its fixtures.
Create a suite
Section titled “Create a suite”- Choose New suite and name it.
- Enter the Blueprint ID to identify the agent. A release-gate suite requires it.
- Choose the Target, concurrency and timeout. Start with the draft and one fixture.
- Add default checks that every fixture should meet. A text check is a useful first check.
- Save the suite. Expand Show fixtures, then choose Add fixture.
- Name the fixture and enter its Input. A plain message is enough for a first fixture.
- Add any fixture-specific checks and save it. These join the suite’s default checks.
Tool mocks (JSON) supplies test outputs for tool calls. The form warns that a tool call without a mock fails the fixture. Use fictional data in fixtures and mock outputs.
Choose optional scoring
Section titled “Choose optional scoring”The form also offers judges, trajectory matching and similarity scoring. Judges add model calls. Similarity needs an expected output and an OpenAI key. Trajectory matching needs captured tool calls.
Use Capture from run to open the capture form when a recorded trace should become a fixture. Review captured customer content before keeping it as test data.
Run status
Section titled “Run status”Run opens a confirmation with the fixture count, model choice and optional published version. Scroll through the model choices while Run suite and Cancel stay visible. Keep Blueprint default to use the agent’s configured model, or select an available model for this run. The dialog warns that fixtures and judges use real model calls charged to your model keys.
Choose Run suite to submit the run. The confirmation stays open while it starts, then opens the admitted run. If submission fails, the dialog shows an error and keeps your choices so you can submit again. The run detail and history update automatically when the run finishes, including while you use another browser tab. Check its status and results before deciding whether the suite passed.
The page has Run history and Compare controls. Compare remains disabled until there are at least two recorded runs.
Limits
Section titled “Limits”These defaults come from the limits catalog. Organization overrides can change the values marked Yes.
| Limit | Default | Support can change it |
|---|---|---|
| Fixtures an eval suite may hold. A run executes all of them | 25 | Yes |
| LLM judges one eval fixture may run (suite defaults plus the fixture’s own) | 5 | Yes |
| Agentic steps one eval judge may take | 8 | Yes |
| Longest time one eval fixture may run before it is stopped | 10 min | Yes |
| Eval fixtures run in parallel | 16 | No |
| Stored traces one scoring call may score | 100 | Yes |