Work in progress · accepting task contributions

Are AI agents good physicists?

FrontierPhysics: Benchmark how AI agents do frontier physics research.

A merged task earns 4 points, a review earns 1. At 12 you are a co-author.

What a task looks like

A native BenchFlow task.md package. The prompt describes an outcome and never names a skill. The oracle must pass with reward 1.0 before any agent runs.

Prompts and oracle logic are human-authored.

Read the contributor guide
tasks/<task-id>/
  task.md            # prompt + metadata
  environment/
    Dockerfile       # frozen environment
    skills/          # mentor skills
  oracle/
    solve.sh         # must reach reward 1.0
  verifier/
    test.sh
    test_outputs.py  # checks the science

Example tasks

Each one comes from research a contributor had already done.

What you get

Earn 12 points, become a co-author

Co-authorship on the FrontierPhysics paper and dataset. Get there with 3 authored tasks, by reviewing, or any mix that adds up.

+4
A task you authored is merged
+1
A task you reviewed is merged
12
Co-authorship on the paper and dataset