Yonghong Zhang
PhD Candidate in Economics & Business, Universidad Autónoma de Madrid
I am an applied economist working on causal inference for climate policy, and on whether AI agents can carry out empirical research reliably rather than merely describe it.
My empirical work studies China's emissions-trading pilots, asking whether carbon markets deliver environmental co-benefits beyond CO₂. It uses difference-in-differences and spatial econometrics on a 2004–2022 provincial panel.
My methodological work asks a narrower question. When an LLM agent reports a causal estimate, what was actually measured: the model, the scaffold around it, or the scorer? I build benchmarks that separate the three, and auditing tools that check identification assumptions against evidence rather than assertion.
I am supervised by Ricardo Correia and Isabel María Parra Oller, and expect to complete the PhD in early 2028. I am happy to hear about research visits and collaborations.
Awards & News
Awards
- 1st place Best Defender, AgentBeats × Lambda Security Arena (UC Berkeley RDI × Lambda), with Team MateFin. A 91.4% defender win rate across 21 held-out adversarial runs.
- 1st place Web Agent Track (tie), AgentX–AgentBeats Phase 1 at UC Berkeley RDI, a Google DeepMind-sponsored track, with Team MateFin.
- 2nd place OpenEnv Custom Track, AgentX–AgentBeats at UC Berkeley RDI, with Team MateFin.
- Accésit Best Poster, UAM Doctoral Week, as the representative poster for the Economics and Business doctoral programme.
News
- 2026 Selected for the Machine Learning Summer School (MLSS 2026), Max Planck Institute for Intelligent Systems and the University of Tübingen.
- 2026 Accepted presentations at the Wolpertinger Conference (Roma Tre University) and ICFAB (University of Surrey, Sustainable Finance track).
Papers
-
CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows
Benchmarks frontier LLMs on causal-inference workflows, pairing method and design interpretation with execution-grounded coefficient recovery.
Under review
-
The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean
Shows how agent benchmarks can mistakenly measure scaffolds and scorers rather than model capability.
Under review
-
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
Introduces ARGUS, a bounded LLM auditing system for difference-in-differences identification assumptions, using an assumption–implication–evidence rubric.
Under review
-
Do Carbon Markets Deliver Environmental Co-Benefits? Evidence from China's ETS Pilots
Difference-in-differences with province and year fixed effects on a 2004–2022 panel, with spatial-spillover analysis.
Working paper
Projects
-
ComtradeBench / OpenEnv
An execution-grounded benchmark for reliable LLM tool use under adversarial and unreliable trade-data API conditions. Demo
-
OfficeQA Agent
A document-grounded retrieval and reasoning agent for U.S. Treasury Bulletin question answering.
-
Climate Claim Classifier
Structured and auditable classification of climate-related claims.
-
PhD Knowledge Base Starter
A lightweight LLM-powered research knowledge-base workflow for Claude Code and Obsidian.