MohammadHossein Rezaei
I am a Machine Learning Research Scientist at Scale AI, working on post-training and evaluation. Most of my recent work is on using rubrics as a reward signal in open-ended domains: eliciting criteria online as a policy improves, how policies reward-hack the rubric verifiers, and how to post-train without a verifier at all. I also work on building benchmarks.
Before Scale, I was a research intern at Stanford, working on benchmarking physical-social norm understanding. I earned my B.S. in Computer Science at the University of Arizona, where I worked on making encoder models robust against negation.
Selected publications
Rubric-Guided Self-DistillationarXiv 2026
EgoNormia: Benchmarking Physical-Social Norm UnderstandingACL Findings 2025