Talk:Reward hacking
Add topic| Text or other creative content from this version of Misaligned goals in artificial intelligence was copied or moved into Reward hacking on 16 Dec 2023. The former page's history now serves to provide attribution for that content in the latter page, and it must not be deleted as long as the latter page exists. |
| This article is rated C-class on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | |||||||||||||||||||||
| |||||||||||||||||||||
Note regarding split to create new article
[edit]For justification of the split, see Talk:AI alignment#Proposed merge of Misaligned goals in artificial intelligence into AI alignment. Klbrain (talk) 12:59, 16 December 2023 (UTC)
Traveller TCS tournament
[edit]Eurisko also twice in succession won the Traveller TCS tournament by specification gaming, after which it was excluded, as the organisers realised they couldn't proof the rules against it. That gives a very direct early example of the dangers of giving AI access to tools that are powerful in its "hands". ~2025-36462-43 (talk) 20:34, 27 January 2026 (UTC)
Feedback requested on recent expansion
[edit]I've recently expanded this article with new sections on theoretical foundations, reward hacking in large language models (including deliberate hacking by reasoning models), and mitigation strategies. I'd appreciate any feedback on the content, structure, or sourcing. Shuyan Ke (talk) 22:08, 19 February 2026 (UTC)
- Thanks for expanding computer science content on Wikipedia. Please note that WP:ARXIV sources are not necessarily peer-reviewed, and WP:PRIMARY sources should be used with caution. WeyerStudentOfAgrippa (talk) 18:42, 20 March 2026 (UTC)
- Thanks for the feedback! Shuyan Ke (talk) 23:19, 23 March 2026 (UTC)
Wiki Education assignment: Online Communities
[edit]
This article was the subject of a Wiki Education Foundation-supported course assignment, between 9 January 2026 and 17 April 2026. Further details are available on the course page. Student editor(s): Shuyan Ke (article contribs). Peer reviewers: Miaces26, Meiqi Jiao, Dang.hazel.
— Assignment last updated by Samantha (Sami) L (talk) 14:40, 24 March 2026 (UTC)
Hi! I had put this within wikiedu page but unsure if you were able to see it so I'm coping it here in case it helps! -the first section about definition and theoretical framework are very technical language you could maybe simplify explanations so any user can understand -Some of your sentences like this one: "A key finding states that, across all stochastic policy distributions, (mappings from states to probability distributions over actions), two reward functions can only be unhackable if and only if one of them is constant, which means that reward hacking is theoretically unavoidable.[3]" are very dense maybe you can split the ideas up for clarity.
Miaces26 (talk) 20:36, 13 April 2026 (UTC)
- Hi Shuyan! Your subsections on adversarial reward functions, reward shaping, and scalable oversight make this well organized. I made some minor edits to improve grammar and fluency for clarity. To further improve your section I would recommend diversifying your sources beyond the heavy reliance on Amodei et al., and consider adding a schematic diagram to help illustrate some of the more abstract concepts like adversarial reward functions. Overall you did amazing and this is really well done! Dang.hazel (talk) 19:30, 16 April 2026 (UTC)