Selected Work
We build and contribute to open infrastructure for understanding the capabilities, reliability, and security of AI agents. Some of our recent work includes:
CyberGym and Frontier AI Cybersecurity Observatory
Advancing real-world evaluation of frontier AI cybersecurity capabilities
We are core contributors to the Frontier AI Cybersecurity Observatory, an open research effort that continuously tracks how frontier AI systems perform across the software vulnerability lifecycle, from reproducing known vulnerabilities to developing working exploits and patches. Its public leaderboards, open to community-contributed results, give researchers, developers, and policymakers a shared, up-to-date view of frontier systems’ progress.
CyberGym is central to that effort — a leading benchmark for measuring the cybersecurity capabilities of frontier AI systems, testing agents on vulnerability analysis tasks drawn from real vulnerabilities in hundreds of widely used open source projects. Evaluations on the benchmark have uncovered dozens of previously unknown zero-day vulnerabilities in real software. We also contributed CyberGym to the UK AI Security Institute’s Inspect Evals collection, helping make rigorous cybersecurity evaluation broadly accessible to the research community.
Agents’ Last Exam
Measuring how AI agents perform on real professional work
Agents’ Last Exam evaluates AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes — expert-built tasks covering most major fields of professional work performed on a computer, developed with hundreds of experts and contributors from academia and industry worldwide. We co-lead development of the benchmark and run evaluations of frontier agent systems, helping characterize their strengths and limitations on complex, real-world work.
AgentBeats
Building open infrastructure for reproducible agent evaluation
We are the core development team behind AgentBeats, an open source platform for evaluating increasingly capable and complex AI agents. Evaluations in AgentBeats are themselves agents: evaluator agents interact dynamically with the systems they assess, enabling richer and more realistic tests than static benchmarks.
The platform provides common infrastructure for building, sharing, and reproducing evaluations across domains including agent safety, cybersecurity, coding, web use, and multi-agent systems. We demonstrated this at community scale through the AgentX–AgentBeats competition, held over several months alongside the Agentic AI MOOC and its community of tens of thousands of registered learners: developer teams around the world contributed hundreds of judge and subject agents, bringing many well-known benchmarks onto the platform as shared public goods.