Skip to main content
FrontierScience evaluates expert-level scientific research tasks and uses a judge model for scoring.

Runtime Status

Run Example

Common Parameters

Outputs

Per-task details are written to results/frontierscience/<model>/<run>/details/. Aggregate metrics are written to summary.md.