MohammadHossein Rezaei
I am a Machine Learning Research Scientist at Scale AI, working on post-training and evaluation. Most of my recent work is on using rubrics as a reward signal in open-ended domains: eliciting criteria online as a policy improves, how policies reward-hack the rubric verifiers, and how to post-train without a verifier at all. I also work on building benchmarks. Recently, I've been working on how to measure progress in recursive self-improvement.
Before Scale, I was a research intern at Stanford, working on benchmarking physical-social norm understanding. I earned my B.S. in Computer Science at the University of Arizona, where I worked on making encoder models robust against negation.