MohammadHossein Rezaei

prof_pic.jpg
Email: mhrezaei@arizona.edu

I am a Machine Learning Research Engineer at Scale AI Scale AI where I work on post-training and evaluation of LLMs. I worked on OnlineRubrics, an approach for post-training LLMs with evolving rubrics to improve alignment in tasks without verifiable ground-truth.

I earned a B.S. in Computer Science from the University of Arizona UArizona. I was a member of the Computational Language Understanding (CLU) Lab, advised by Eduardo Blanco, where I worked on making SLMs more robust against negation by further pre-training and paraphrasing in affirmative terms.

Previously, I was a research intern at Stanford University Stanford in the SALT Lab advised by Diyi Yang. There, I co-created EgoNormia, a benchmark for evaluating physical-social norm understanding in vision-language models.

selected publications

  1. Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers
    MohammadHossein RezaeiAnas MahmoudZihao Wang, Utkarsh Tyagi, Advait Gosai, Razvan-Gabriel Dumitru, Aakash Sabharwal, Bing Liu, and Yunzhong He
    Jun 2026
  2. Reward Hacking in Rubric-Based Reinforcement Learning
    Anas MahmoudMohammadHossein RezaeiZihao WangAnisha GunjalBing Liu, and Yunzhong He
    May 2026
  3. Online Rubrics Elicitation from Pairwise Comparisons
    In Proceedings of the 43rd International Conference on Machine Learning (ICML), Jul 2026
  4. EgoNormia: Benchmarking Physical-Social Norm Understanding
    MohammadHossein Rezaei*Yicheng Fu*Phil Cuvin*Caleb ZiemsYanzhe ZhangHao Zhu, and Diyi Yang
    In Findings of the Association for Computational Linguistics: ACL 2025, Jul 2025