Redwood Research
A nonprofit advancing the safe development of AI through technical research on AI control and misalignment.

To ensure the safe development of artificial intelligence through rigorous technical research, analysis, and advising. Redwood pioneers "AI control" — techniques for reliably monitoring and preventing subversion by potentially deceptive AI systems — and studies strategic deception in frontier models.
AI areas they serve
Areas of Focus
AI control protocols
AI alignment and safety research
Detecting strategic deception and scheming in AI models
Empirical AI safety methods
Catastrophic-risk mitigation
Advising AI developers and governments
AI Impact Areas
Upcoming Goals
Advance and propel the field of AI control
Develop empirical safety techniques robust to deceptive AI
Advise frontier AI developers and governments on risk mitigation
Strengthen safeguards against misaligned AI agents
2024: Published research demonstrating that frontier models can strategically fake alignment during training
2023: Published the foundational AI Control: Improving Safety Despite Intentional Subversion (ICML oral paper)
2021: Founded in Berkeley, California, with initial support from Open Philanthropy
All information on this page is from this Organization's website. All efforts are made to keep this information current.
If you represent this organization, please visit our Join Us page to provide any updates.
