Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a21f6c9569a5452a

Jump to content

Talk:Reward hacking

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia
Latest comment: 3 months ago by Dang.hazel in topic Wiki Education assignment: Online Communities

Note regarding split to create new article

[edit]

For justification of the split, see Talk:AI alignment#Proposed merge of Misaligned goals in artificial intelligence into AI alignment. Klbrain (talk) 12:59, 16 December 2023 (UTC)Reply

Traveller TCS tournament

[edit]

Eurisko also twice in succession won the Traveller TCS tournament by specification gaming, after which it was excluded, as the organisers realised they couldn't proof the rules against it. That gives a very direct early example of the dangers of giving AI access to tools that are powerful in its "hands". ~2025-36462-43 (talk) 20:34, 27 January 2026 (UTC)Reply


Feedback requested on recent expansion

[edit]

I've recently expanded this article with new sections on theoretical foundations, reward hacking in large language models (including deliberate hacking by reasoning models), and mitigation strategies. I'd appreciate any feedback on the content, structure, or sourcing. Shuyan Ke (talk) 22:08, 19 February 2026 (UTC)Reply

Thanks for expanding computer science content on Wikipedia. Please note that WP:ARXIV sources are not necessarily peer-reviewed, and WP:PRIMARY sources should be used with caution. WeyerStudentOfAgrippa (talk) 18:42, 20 March 2026 (UTC)Reply
Thanks for the feedback! Shuyan Ke (talk) 23:19, 23 March 2026 (UTC)Reply

Wiki Education assignment: Online Communities

[edit]

This article was the subject of a Wiki Education Foundation-supported course assignment, between 9 January 2026 and 17 April 2026. Further details are available on the course page. Student editor(s): Shuyan Ke (article contribs). Peer reviewers: Miaces26, Meiqi Jiao, Dang.hazel.

— Assignment last updated by Samantha (Sami) L (talk) 14:40, 24 March 2026 (UTC)Reply

Hi! I had put this within wikiedu page but unsure if you were able to see it so I'm coping it here in case it helps! -the first section about definition and theoretical framework are very technical language you could maybe simplify explanations so any user can understand -Some of your sentences like this one: "A key finding states that, across all stochastic policy distributions, (mappings from states to probability distributions over actions), two reward functions can only be unhackable if and only if one of them is constant, which means that reward hacking is theoretically unavoidable.[3]" are very dense maybe you can split the ideas up for clarity.

Miaces26 (talk) 20:36, 13 April 2026 (UTC)Reply

Hi Shuyan! Your subsections on adversarial reward functions, reward shaping, and scalable oversight make this well organized. I made some minor edits to improve grammar and fluency for clarity. To further improve your section I would recommend diversifying your sources beyond the heavy reliance on Amodei et al., and consider adding a schematic diagram to help illustrate some of the more abstract concepts like adversarial reward functions. Overall you did amazing and this is really well done! Dang.hazel (talk) 19:30, 16 April 2026 (UTC)Reply