BenchFlow documentation
Start a real local evaluation in five minutes, then explore the full runtime.
Start in 5 minutes
Run a real SkillsBench task locally with Docker and Codex.
Getting started
Run your first scored evaluation locally with Docker.
Authentication
Use subscription logins, API keys, and provider credentials.
Concepts
Tasks, agents, environments, rollouts, scenes, and verifiers.
Run any benchmark
Native, translated, and as-is paths to one result contract.
Running evaluations
Run one task, local batches, configs, and skill comparisons.
Sandboxes
Start with Docker; choose a cloud backend only when needed.
Authoring tasks
Create and validate a native task.md package.
Native task.md format
Complete authoring guide for BenchFlow's native task package.
Task package standard
The versioned contract for native task packages.
Environment plane
Provision versioned, stateful services independently from tasks.
Continue runs
Continue timed-out agent work with record-replay.
LLM-as-judge
Score subjective deliverables with weighted rubric criteria.
Rubric review
Audit finished rollouts for reward hacking and task quality.
Skill evals
Measure whether a skill improves agent performance.
Use cases
Multi-role, multi-turn, skill-generation, and stateful patterns.
Progressive disclosure
Drive multi-round runs with deterministic verifier feedback.
Sandbox hardening
How BenchFlow protects verifiers and hidden task data.
External agents
Run agents from the public BenchFlow agents repository.
Trajectory upload
Contribute redacted trajectory captures without a BenchFlow account.
CLI reference
BenchFlow command-line reference.
Python API
BenchFlow Python SDK reference.