Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

Jump to content

Center for Human-Compatible Artificial Intelligence

From Wikipedia, the free encyclopedia
Center for Human-Compatible Artificial Intelligence
Formation2016; 10 years ago (2016)
HeadquartersBerkeley, California
Director
Stuart J. Russell
Executive director
Mark Nitzberg
Parent organization
University of California, Berkeley
Websitehumancompatible.ai

The Center for Human-Compatible Artificial Intelligence (CHAI) is a research center at the University of California, Berkeley focusing on artificial intelligence (AI) safety and value alignment. The center was founded in August 2016 by a group of academics led by Berkeley computer science professor Stuart J. Russell.[1][2]

Organization and funding

[edit]

CHAI's faculty affiliates have included Russell, Pieter Abbeel and Anca Dragan from Berkeley, Bart Selman and Joseph Halpern from Cornell,[3] Michael Wellman and Satinder Singh Baveja from the University of Michigan, and Tom Griffiths and Tania Lombrozo from Princeton.[4]

In 2016, the Open Philanthropy Project (OpenPhil) recommended a grant of $5,555,550 over five years to UC Berkeley to support CHAI's launch.[5] OpenPhil subsequently recommended $200,000 in support for CHAI over two years in November 2019.[6] In January 2021, it recommended a five-year renewal grant of $11,355,246.[7] In April 2020, OpenPhil separately recommended a grant of up to $50,000 to the World Economic Forum for a Global AI Council workshop co-developed with CHAI.[8]

Research

[edit]

The center studies how AI systems can pursue specified objectives in ways that conflict with human intentions.[2] CHAI's approach to AI safety research focuses on value alignment strategies, particularly inverse reinforcement learning, in which an AI system infers human preferences from observing human behavior.[9]

In 2016, Dylan Hadfield-Menell, Dragan, Abbeel and Russell introduced cooperative inverse reinforcement learning (CIRL). The framework models a human and a robot as cooperating to maximize the human's reward, while the robot initially does not know the reward function. It allows the human to teach and the robot to seek information, rather than treating learning solely as observation of human behavior.[10]

The center has also worked on modeling human-machine interaction in scenarios where intelligent machines have an "off-switch" that they are capable of overriding.[11] The 2017 paper "The Off-Switch Game" analyzes how uncertainty about an action's value can give a robot an incentive to allow a human to switch it off; the human's decision provides information about that value.[12] In Human Compatible (2019), Russell explains that the model's results depend on assumptions about human behavior; allowing for human error can reduce the robot's incentive to defer to a person.[13]

See also

[edit]

References

[edit]
  1. ↑ Norris, Jeffrey (2016-08-29). "UC Berkeley launches Center for Human-Compatible Artificial Intelligence". Retrieved 2026-09-27.
  2. 1 2 Solon, Olivia (2016-08-30). "The rise of robots: forget evil AI – the real risk is far more insidious". The Guardian. Retrieved 2026-09-27.
  3. ↑ Cornell University. "Human-Compatible AI". Retrieved Dec 27, 2019.
  4. ↑ Center for Human-Compatible Artificial Intelligence. "People". Retrieved 2026-09-27.
  5. ↑ Open Philanthropy Project (August 2016). "UC Berkeley — Center for Human-Compatible AI (2016)". Archived from the original on 2023-12-12. Retrieved 2026-09-27.
  6. ↑ Open Philanthropy Project (Nov 2019). "UC Berkeley — Center for Human-Compatible AI (2019)". Archived from the original on 2023-09-25. Retrieved Dec 27, 2019.
  7. ↑ "UC Berkeley — Center for Human-Compatible Artificial Intelligence (2021)". Open Philanthropy. January 2021. Archived from the original on 2024-01-02. Retrieved 2026-09-27.
  8. ↑ "World Economic Forum — Global AI Council Workshop". Open Philanthropy. April 2020. Archived from the original on 2023-09-01. Retrieved 2023-09-01.
  9. ↑ Conn, Ariel (2016-08-31). "New Center for Human-Compatible AI". Future of Life Institute. Retrieved 2026-09-27.
  10. ↑ Hadfield-Menell, Dylan; Dragan, Anca; Abbeel, Pieter; Russell, Stuart (2016). "Cooperative Inverse Reinforcement Learning". Advances in Neural Information Processing Systems. Vol. 29.
  11. ↑ Bridge, Mark (2017-06-10). "Making robots less confident could prevent them taking over". The Times.
  12. ↑ Hadfield-Menell, Dylan; Dragan, Anca; Abbeel, Pieter; Russell, Stuart (2017). "The Off-Switch Game". Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence. pp. 220–227. doi:10.24963/ijcai.2017/32.
  13. ↑ Russell, Stuart (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking. pp. 199–200. ISBN 978-0-525-55861-3.
[edit]