About
A frontier environment lab for AI agents.
We build the environments AI agents need to learn and be evaluated on real computer work, not static prompts. Based in the Bay Area.

Xiangyi Li
Founder
Coding model inference & inference at Tesla

Bingran You
MTS
PhD in quantum physics at Berkeley

Carrie Chen
MTS
CS @ Cornell · host @ Cornell RL research seminar
BenchFlow is a data lab. We build stateful environments with services, files, tools, verifiers, traces, and replay, so agents can learn from whole workflows.
We ship SkillsBench for procedural skills, ClawsBench for simulated workplaces, Robo Use for embodied agents, PostTrain for environment-side post-training, and the runtime that runs them. Hugging Face, Fireworks, and Prime Intellect provide compute for PostTrain. We sit on the OpenEnv technical committee.
Started in late 2024. SkillsBench has 1,800+ GitHub stars and 260+ citations on Semantic Scholar, and the Discord has 1,700+ members. We work with frontier labs, the agent skills community, and academic partners through Agent Skills ’26 at CAIS, PostTrain, and re:AGENT.