Postdoctoral researcher at the Halıcıoğlu Data Science Institute, UC San Diego, working on resource-efficient agents for science and engineering.
AI agents can already generate work at scale. Turning that into real acceleration in science and engineering means accounting for what each step actually costs. I build the agents, the evaluations, and the benchmarks around that constraint. That instinct comes from particle physics, where every measurement is genuinely expensive and progress depends as much on engineering as on science.
🌐 yuema137.github.io · Research · Projects · Publications · Talks · Blog · CV
- Cost-aware agentic systems — closed-loop systems that propose, implement, and evaluate experiments under real compute budgets
- Scientific agent evaluation — representing decision trajectories and action spaces, so agents can be judged on outcomes and on what they spent
- ML for scientific discovery — neural processes and multi-fidelity surrogates for expensive simulation
- Statistical inference at scale — Bayesian and likelihood methods for rare-event searches, and the data systems behind them
Three layers of the same problem: a system that carries out the research, a methodology for representing what any such system did, and a benchmark that asks the question inside one field.
| Project | What it is | Role |
|---|---|---|
| SIDERIUS | A closed-loop agentic system for autonomous research under real compute budgets | Sole architect |
| SciTra | A methodology for representing what a scientific agent did — decision trajectories and action spaces — so runs are comparable across agents, tasks, and fields, plus the infrastructure that implements it. Not tied to a single benchmark or field. 50+ scientists, 20+ institutions, with BenchFlow | Founder & lead |
| FrontierPhysics | A benchmark in one field: how AI agents carry out frontier physics research iteratively. With the BenchFlow team | Core team |
Experimental particle physics, with research experience across the XENONnT, LEGEND, and KamLAND experiments — statistical inference, detector data analysis, signal processing, simulation, and scientific software infrastructure.
I am particularly interested in building open-source tools that make AI agents more reliable, measurable, and useful for scientific discovery.




