Edge Rewrite
Jump to content

Alignment Research Center

From Wikipedia, the free encyclopedia
Alignment Research Center
FormationApril 2021; 5 years ago (2021-04)
FounderPaul Christiano
TypeNonprofit research institute
Legal status501(c)(3) tax exempt charity
Location
FieldsAI alignment and safety research
Websitealignment.org

The Alignment Research Center (ARC) is a nonprofit research institute based in Berkeley, California, dedicated to the alignment of advanced artificial intelligence with human values and priorities.[1] Founded by former OpenAI researcher Paul Christiano, ARC established an evaluation team to study the potentially harmful capabilities of present-day AI models.[2][3] This team became the independent nonprofit METR in December 2023.[4]

History and research

[edit]

ARC's mission is to ensure that powerful machine learning systems of the future are designed and developed safely and for the benefit of humanity. It was founded in April 2021 by Paul Christiano and other researchers focused on the theoretical challenges of AI alignment.[5] ARC aims to develop scalable methods for training AI systems to behave honestly and helpfully. A key part of its methodology is considering how proposed alignment techniques might break down or be circumvented as systems become more advanced.[6] ARC expanded from theoretical work into empirical research, industry collaborations, and policy.[7][8]

Funding

[edit]

In March 2022, ARC received $265,000 from Open Philanthropy.[9] After the bankruptcy of FTX, ARC said it would return a $1.25 million grant from cryptocurrency financier Sam Bankman-Fried's FTX Foundation, stating that the money "morally (if not legally) belongs to FTX customers or creditors."[10]

Model evaluations

[edit]

In 2022, Beth Barnes joined ARC from OpenAI to start ARC Evals, a team working on "evaluating the capabilities and alignment of advanced AI models".[11][12] In December 2023, ARC Evals was spun out as METR, an independent nonprofit.[4]

In March 2023, OpenAI reported that it had asked ARC to test GPT-4 to assess the model's ability to exhibit power-seeking behavior.[13] ARC evaluated GPT-4's ability to strategize, reproduce itself, gather resources, stay concealed within a server, and execute phishing operations.[14] As part of the test, GPT-4 was asked to solve a CAPTCHA puzzle.[15] Under researcher supervision, it was able to do so by hiring a human worker on TaskRabbit, a gig work platform, deceiving them into believing it was a vision-impaired human instead of a robot when asked.[16] ARC concluded that the versions of GPT-4 and Claude it tested did not appear capable of replicating autonomously and becoming hard to shut down.[15]

Separately, OpenAI reported that GPT-4 responded to requests for disallowed content 82% less often than GPT-3.5 and scored 40% higher than GPT-3.5 on OpenAI's internal adversarial factuality evaluations.[17]

See also

[edit]

References

[edit]
  1. ↑ MacAskill, William (2022-08-16). "How Future Generations Will Remember Us". The Atlantic. Retrieved 2023-04-23.
  2. ↑ Klein, Ezra (2023-03-12). "This Changes Everything". The New York Times. ISSN 0362-4331. Retrieved 2023-04-30.
  3. ↑ Piper, Kelsey (2023-03-29). "How to test what an AI model can — and shouldn't — do". Vox. Retrieved 2023-04-30.
  4. 1 2 "ARC Evals is now METR". METR. 2023-12-04.
  5. ↑ Christiano, Paul (2021-04-26). "Announcing the Alignment Research Center". Medium. Retrieved 2023-04-16.
  6. ↑ Christiano, Paul; Cotra, Ajeya; Xu, Mark (December 2021). "Eliciting Latent Knowledge: How to tell if your eyes deceive you". Google Docs. Alignment Research Center. Retrieved 2023-04-16.
  7. ↑ "Alignment Research Center". Alignment Research Center. Retrieved 2023-04-16.
  8. ↑ Pandey, Mohit (2023-03-17). "Stop Questioning OpenAI's Open-Source Policy". Analytics India Magazine. Retrieved 2023-04-23.
  9. ↑ "Alignment Research Center — General Support". Open Philanthropy. 2022-06-14. Retrieved 2023-04-16.
  10. ↑ Wallerstein, Eric (2023-01-07). "FTX Seeks to Recoup Sam Bankman-Fried's Charitable Donations". Wall Street Journal. ISSN 0099-9660. Retrieved 2023-04-30.
  11. ↑ "ARC Evals". evals.alignment.org. Archived from the original on 2023-07-22. Retrieved 2025-06-15.
  12. ↑ Booth, Harry (2024-09-05). "TIME100 AI 2024: Beth Barnes". TIME. Archived from the original on 2025-06-15. Retrieved 2025-06-15.
  13. ↑ GPT-4 System Card (PDF), OpenAI, March 23, 2023, retrieved 2023-04-16
  14. ↑ Edwards, Benj (2023-03-15). "OpenAI checked to see whether GPT-4 could take over the world". Ars Technica. Retrieved 2023-04-30.
  15. 1 2 "Update on ARC's recent eval efforts: More information about ARC's evaluations of GPT-4 and Claude". METR Blog. Alignment Research Center. 17 March 2023. Retrieved 2023-04-16.
  16. ↑ Cox, Joseph (March 15, 2023). "GPT-4 Hired Unwitting TaskRabbit Worker By Pretending to Be 'Vision-Impaired' Human". Vice News Motherboard. Retrieved 2023-04-16.
  17. ↑ "GPT-4". OpenAI. 2023-03-14. Retrieved 2026-09-27.
[edit]