I design and build evaluation infrastructure and coding-agent systems that make AI-generated software more reliable and effective. My aim is to understand how these systems behave in realistic software workflows and identify where they can be improved.
ExplainBench evaluates whether coding-agent explanations accurately describe the intended and actual behavior of their patches.
We found that explanation quality does not always align with patch efficacy, and propose ExplanationAuditAgent, which improved explanation scores by 10.9% on average across five coding agents.
BigCodeBench evaluates practical code generation across 1,140 tasks involving complex instructions and diverse function calls.
Top models reached about 60%, compared with 97% human performance.
Trusted by teams at Zhipu AI, Alibaba Qwen, DeepSeek, Amazon AWS AI, Snowflake AI Research, ServiceNow Research, Meta AI, Cohere AI, Sakana AI, and AI2.
This project explored how domain knowledge and learning techniques can connect natural-language intent and hardware setups with reusable examples and generated Arduino programs.