Skip to main content
Tool-calling evaluations are useful when the main risk is model I/O discipline: protocol support, structured outputs, tool-call formatting, retries, and whether the harness can turn task context into tool-aware model messages.

Current Runtime Path

Use harnesses that explicitly support tool-aware model protocols: Confirm available benchmark and harness ids with Supported Components before documenting or launching a run.

Tool-Use Run

What To Check