News
- Research
How we improved SkillsBench v1.1 scores by 69.4% using env0
env0 provides high-fidelity simulated Google Workspace and Slack services, enabling safe post-training experiments that improve Qwen3.5-9B through limited SFT on curated task trajectories.

- Update
GenesisBench: can language intelligence create physical intelligence
Ongoing research: a benchmark evaluating whether coding agents can turn language intelligence into physical intelligence.

- Workshop
Agent Skills ’26 workshop accepted at CAIS
First workshop on agent skills, in San Jose at the ACM CAIS conference. Speakers: Dawn Song, Ross Taylor (General Reasoning), Kanav Garg (DeepMind), Yu Su.

- Release
ClawsBench paper on arXiv
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces. 6 models, 4 harnesses, 7,224 trials.

- Hackathon
Agent Skills Hackathon II at AGI House
Second hackathon in the Agent Skills series, hosted at AGI House. 600+ total participants across both events.

- Hackathon
Agent Skills Hackathon I at Founders, Inc.
First public hackathon in the Agent Skills series. Held at Founders, Inc. in San Francisco.

- Release
SkillsBench paper on arXiv
SkillsBench: Benchmarking How Well Skills Work for AI Agents. 84 tasks, 7 models, 3 harnesses.
