Redwood Research

A nonprofit advancing the safe development of AI through technical research on AI control and misalignment.

About Us

Year Founded:

2021

Geography:

North America

Address:

Berkeley, California, United States

Visit Website
Mission

To ensure the safe development of artificial intelligence through rigorous technical research, analysis, and advising. Redwood pioneers "AI control" — techniques for reliably monitoring and preventing subversion by potentially deceptive AI systems — and studies strategic deception in frontier models.

Impact

AI areas they serve

Areas of Focus

AI control protocols

AI alignment and safety research

Detecting strategic deception and scheming in AI models

Empirical AI safety methods

Catastrophic-risk mitigation

Advising AI developers and governments

AI Impact Areas

Upcoming Goals

Advance and propel the field of AI control

Develop empirical safety techniques robust to deceptive AI

Advise frontier AI developers and governments on risk mitigation

Strengthen safeguards against misaligned AI agents

Milestones

2024: Published research demonstrating that frontier models can strategically fake alignment during training

2023: Published the foundational AI Control: Improving Safety Despite Intentional Subversion (ICML oral paper)

2021: Founded in Berkeley, California, with initial support from Open Philanthropy

All information on this page is from this Organization's website. All efforts are made to keep this information current.

If you represent this organization, please visit our Join Us page to provide any updates.