Draft:Evan Frick
Submission declined on 17 November 2025 by Pythoncoder (talk).
Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
|
Evan Frick
[edit]Evan Frick (born July 17, 2002) is an American machine learning researcher and computer scientist known for his contributions to the evaluation and post-training of large language models (LLMs). He is a researcher at Arena.ai and a contributor to the LMSYS Org (Large Model Systems Organization), specifically working on the Chatbot Arena benchmarking platform.[1]
Education
[edit]Frick attended the University of California, Berkeley, where he earned a B.A. in Computer Science (2024) and an M.S. in Electrical Engineering and Computer Science (2025). During his graduate studies, he conducted research on reward modeling and reinforcement learning from human feedback (RLHF) under the supervision of Jiantao Jiao and in collaboration with Ion Stoica.[2]
Career and research
[edit]Frick’s research focuses on the "post-training" phase of AI development, specifically how to align model outputs with human preferences. As a member of the LMSYS Org team, he co-developed Arena-Hard and the Prompt-to-Leaderboard (P2L) system, which aim to provide more nuanced evaluations than traditional static benchmarks.[3]
He has held positions at:
Arena.ai: Member of Technical Staff focusing on preference modeling. Nexusflow: Machine Learning Engineer, where he contributed to the development of the "Athene" series of open-weights models. Berkeley AI Research (BAIR): Research contributor to the Starling-7B project, a model series utilizing Reinforcement Learning from AI Feedback (RLAIF). His work has been published in machine learning conferences including the International Conference on Machine Learning (ICML) and the International Conference on Learning Representations (ICLR).[4]
Selected publications
[edit]Frick, E., et al. (2025). "Prompt-to-Leaderboard: Prompt-Adaptive LLM Evaluations." Proceedings of the 42nd ICML. Zhu, B., Frick, E., et al. (2024). "Starling-7B: Improving Helpfulness and Harmlessness with RLAIF." Conference on Language Modeling (COLM). Lambert, N., ... Frick, E., et al. (2024). "How to Evaluate Reward Models for RLHF." ICLR 2024.
References
[edit]- ↑ Norton, Rosa (2025-05-06). "As companies pour billions into AI, a ranking system by UC Berkeley students has all eyes on it". Berkeley News. Retrieved 2024-11-17.
- ↑ "Reward Modeling for Human Preferences". EECS at UC Berkeley. Retrieved 2024-11-17.
- ↑ "How Early Access to NVIDIA GB200 Systems Helped LMArena Build a Model to Evaluate LLMs". NVIDIA Technical Blog. 2025-06-18.
- ↑ Frick, Evan; Chen, Connor; et al. (2025). "Prompt-to-Leaderboard: Prompt-Adaptive LLM Evaluations". Proceedings of the 42nd International Conference on Machine Learning. 267. PMLR: 17672–17689.
External links
[edit]Official website Evan Frick publications indexed by Google Scholar

LLM-generated pages with certain obvious signs of being machine generated may be deleted without notice.
Instead, only summarize in your own words a range of independent, reliable, published sources that discuss the subject.
See the advice page on large language models for more information.