BenchFlow builds the environments AI agents learn in.
A frontier environment lab for AI agents. We ship SkillsBench, ClawsBench, PostTrain, and the runtime that runs them.
What we ship
01· BenchmarkSkillsBench
Do procedural skills make agents better at real work? 86 tasks, 11 domains.
02· EnvironmentClawsBench
Five wire-compatible workplaces. Capability and safety, scored separately.
03· ArenaPostTrain
Contribute an environment. We post-train and score what generalizes.
04· ScienceFrontierPhysics
Evaluating agents for end-to-end frontier physics research.
05· RuntimeRuntime
Run a real task locally: Docker, an ACP agent, an independent verifier.
Ecosystem
- Mar 7Founders, Inc. · San Francisco
Skillathon — the first Agent Skills hackathon
Builders wrote skills and tasks across five tracks to extend SkillsBench. Organized with Sundial; 200+ at Founders, Inc.
luma.com/1khurilu ↗ - May 26CAIS · San Jose
Agent Skills ’26 workshop
First workshop on agent skills. Speakers: Dawn Song, Ross Taylor, Kanav Garg (DeepMind), Yu Su. Live SkillsBench design challenge.
agentskills-workshop.org ↗ - May 27San Francisco
SkillsBench 1.0 Launch party
Presented by Google DeepMind — the afterparty to the CAIS workshop. We announced SkillsBench 1.0: 100+ expert-curated tasks, built with Kaggle. Talks from Xiangyi Li, Wenbo Chen, Han Lee, and Kaggle.
luma.com/deepmind-634c ↗ - May 30AGI House · San Francisco
Google DeepMind Enterprise Build Day
Co-hosted at AGI House. Full-day build on the post–I/O agentic stack — Gemini, Antigravity, Managed Agents — with DeepMind in the room.
agihouse.org/events/gemini-build-day ↗ - Aug 15–16San Francisco
re:AGENT — End to End Agentic Science
Two-day build weekend on the infrastructure scientific agents still need: datasets, tools, and evaluation. Tracks for AI scientists, literature-scale analysis, and biological design.
luma.com/g6org075 ↗