MohammadHossein Rezaei

MohammadHossein Rezaei

I am a Machine Learning Research Scientist at Scale AI, working on post-training and evaluation. Most of my recent work is on using rubrics as a reward signal in open-ended domains: eliciting criteria online as a policy improves, how policies reward-hack the rubric verifiers, and how to post-train without a verifier at all. I also work on building benchmarks.

Before Scale, I was a research intern at Stanford, working on benchmarking physical-social norm understanding. I earned my B.S. in Computer Science at the University of Arizona, where I worked on making encoder models robust against negation.

>