FAR.AI’s cover photo
FAR.AI

FAR.AI

Research Services

Berkeley, California 31,573 followers

Frontier alignment research to ensure the safe development and deployment of advanced AI systems.

About us

FAR.AI is a technical AI research and education non-profit, dedicated to ensuring the safe development and deployment of frontier AI systems. FAR.Research: Explores a portfolio of promising technical AI safety research directions. FAR.Labs: Supports the San Francisco Bay Area AI safety research community through a coworking space, events and programs. FAR.Futures: Delivers events and initiatives bringing together global leaders in AI academia, industry and policy.

Website
https://far.ai/
Industry
Research Services
Company size
11-50 employees
Headquarters
Berkeley, California
Type
Nonprofit
Founded
2022
Specialties
Artificial Intelligence and AI Alignment Research

Locations

Employees at FAR.AI

Updates

  • FAR.AI reposted this

    Last week, NYAIGS co-hosted the reception for FAR.AI's UNGA convening on AI-enabled terrorism, bringing together UN officials and counter-terrorism and AI safety experts all in one room. The evening's discussion centered on how AI is changing the terrorism landscape and the difficulty of safeguarding AI models against misuse and abliteration. Speakers pointed to earlier multilateral work on counter-terrorism, cybersecurity, and social media as a starting point, and to common evaluation standards and collaboration and information sharing between governments, AI companies, and counter-terrorism experts as essential next steps. Thank you to FAR.AI for organizing; and to speakers Steven S., Dina H., David Scharia, Brian Tse 谢旻希, panelists Adam Gleave, Adam Hadley CBE, Paul Ash, and moderator Patricia Paskov for the insightful conversation. We were glad to bring NYC's AI safety and governance community together for this cross-field conversation. Thank you to everyone who came out. More to come.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • View organization page for FAR.AI

    31,573 followers

    We're hiring, and the mission is making AI safe. Our team is growing fast, and we need people who want to work on one of the most important problems of our time. Some of our roles are deeply technical: researchers, red-teamers, and engineers. Others call for a different but equally valuable set of skills: running operations, managing programs, and building our events. If either sounds like you, apply through the link in the comments, where you'll find these roles and many more. 👉 Know someone who'd be a great fit? Send this their way!

    • No alternative text description for this image
  • View organization page for FAR.AI

    31,573 followers

    What’s keeping AI safety researchers up at night? On September 28, we co-hosted the CAIRD workshop with Cambridge Boston Alignment Initiative in Kendall Square in Cambridge, MA, to explore promising research directions in AI safety and help local researchers make connections across universities and disciplines, from tenured faculty to early-career scholars. A few themes stood out from the lightning talks, panel discussions, and breakout groups: ➡️Loss of control, catastrophic misuse, and concentration of power, where a handful of people control the major AI models, are major challenges facing advanced AI. ➡️AI safety is a multidisciplinary problem. Solving it will take economists, cognitive scientists, cryptographers, philosophers, and psychologists. ➡️Unlocking how AI models actually work will take more investment in AI science, alongside AI products. ➡️AI governance tends to be incident-driven. Policy moves quickly after a shock, so researchers need practical solutions ready before the next one. ➡️The field has made good progress in defining the risks, which gives policymakers a stronger basis to act on transparency and solutions-oriented research. Many thanks to our partner Emre Yavuz, to our speakers David Bau, Dylan Hadfield-Menell, Adam Kalai, and Hidenori Tanaka for the ideas they brought to the workshop, and to Microsoft Research for hosting us. Interested in joining our next gathering? Let us know through the form in the comments.

    • No alternative text description for this image
  • View organization page for FAR.AI

    31,573 followers

    FAR.AI CGO Edward Yee joined the Paradigm Shock Podcast to talk about the AI Security Leaderboard, and how uneven safeguards across models create vulnerabilities bad actors can exploit. Listen to the episode in the post below.

  • View organization page for FAR.AI

    31,573 followers

    Been paying more attention to AI safety lately and wondering what the work actually looks like? Most people in the field didn't start there. They came from ML infra, academia, policy, journalism, startups, and found their way in from wherever they were. On Thursday, October 1, we're hosting an evening for people who are new to AI safety or just curious about it: researchers, engineers, students, and anyone adjacent. Our researchers and lab residents will share how they got into the field and what they're working on now. Plenty of the paths are non-linear, because this work takes all kinds of people and skills. Come hear a few of their stories and meet others figuring out the same question. FAR.Labs, Berkeley, CA 6:30 to 9:00 PM Space is limited, so please register by Monday, September 28. Link below. Share this with anyone who's been meaning to learn more!

    • No alternative text description for this image
  • View organization page for FAR.AI

    31,573 followers

    What would it take for AI safety guardrails to be adopted and implemented on an international scale? This was the question of the day at FAR.AI’s UNGA convening on September 22, where we brought together UN officials, counter-terrorism and AI safety experts, and the frontier labs for an urgent discussion on practical approaches to preventing AI-enabled terrorism. Terrorists and violent extremists are not only using AI for propaganda and radicalization, they’re increasingly using LLMs for direct operational and tactical gains in weapons development and attack planning. Even with well-intentioned guardrails in place, our own AI Security Leaderboard found that it costs less than $220 to bypass those guardrails for the least secure leading AI model. A lot of great ideas were raised at the event – a few themes stood out: ➡️ The threat is no longer hypothetical, and the window to act is closing. ➡️ Guardrails alone aren't enough. Safeguards can be broken or removed entirely (the “abliteration” problem). ➡️ We don’t need to start from scratch: multilateral coordination efforts around counterterrorism, cybersecurity, and social media offer lessons. ➡️ Speakers called for common evaluation standards, aligned threat models, and coordinated safety benchmarks. ➡️ Collaboration and information sharing are essential. We need more secure, multi-stakeholder sharing between governments, AI companies, and counterterrorism experts, including across the U.S.-China divide. ➡️ Policymakers’ attention and expertise are lagging behind real-world developments. Citizens need to demand accountability, and experts need to double down on education and practical policy recommendations. Thanks to our speakers Steven S., David Scharia, Brian Tse 谢旻希 and our panelists Adam Gleave, Dina H., Adam Hadley CBE, and Paul Ash, and our moderator Patricia Paskov for their insights and expertise.

  • FAR.AI reposted this

    This year's High Level Session of the UN General Assembly (#UNGA81) has been dominated by discussions on AI. On Tuesday, I had the privilege of opening an event on "AI-Enabled Terrorism: Addressing Radicalization and CBRN Risks in the Age of LLMs," hosted by FAR.AI. The timing was ideal. On Monday, #TechAgainstTerrorism released its updated CT-AI Benchmark, which tested 160 AI models with around 100,000 requests. Many of the models gave detailed answers to harmful requests even when users openly declared terrorist intent. My key messages: First - Counter-terrorism must become anticipatory. It has to keep pace with AI while staying grounded in the rule of law and human rights. Second - UNOCT and the UN's Member States are focused on making institutions ready to use AI responsibly, making sure CT and prevention experience informs how models are developed and preventing the weaponization of AI. Third - No single actor or institution can work alone. Developers, researchers, practitioners and civil society need to be in the same room working on operational mitigating measures together. We have jumped through an Overton window, from what was unthinkable a few years ago to what now seems sensible - and now we will need to policy development and align on standards. AI is no longer a niche issue. It will shape every area of counter-terrorism and prevention. At #UNOCT - our programmes support Member States by co-developing locally relevant capacity building assistance in some of these sectors - and with additional financial support - we can do more! My thanks to Adam Gleave, Karl Berzins and the FAR.AI team for convening, and to fellow speakers David Scharia, Brian Tse 谢旻希, Adam Hadley CBE and Paul Ash. #AI #CounterTerrorism #AISafety #PCVE #FAR.AI #UNCCT #UNOCT

    • No alternative text description for this image
  • FAR.AI reposted this

    AI policy needs evidence. But the evidence does not come from one panel, one region or one definition of risk.    At #DigitalCooperationDay, four international AI evidence-building initiatives came together to compare where their findings on AI opportunities, risks and impacts converge, where they differ, and what is still missing.    Yoshua Bengio and Maria Ressa, Co-Chairs of the Independent International Scientific Panel on AI, were joined by Kwan Yee Ng 吴君仪 of the International Scientific Report on the Safety of Advanced AI, Kalika Bali of the Global South AI Safety Report, and Adam Gleave of the EU AI Act Scientific Panel.    Gary Marcus offered framing remarks, @Renata Dwan moderated the discussion, and Urvashi Aneja brought perspectives from the Global South Network for Trustworthy AI.    The question running through the discussion: how can scientific evidence become more useful for real governance choices across countries with very different capacities and resources?    🔴 Watch: https://lnkd.in/gtZ8qFgx     #DigitalCooperation #DigitalCooperationDay #AIGovernance

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • View organization page for FAR.AI

    31,573 followers

    Tune in at 3 pm ET today to watch co-founder and CEO Adam Gleave join the "Panel of Panels: Building a Global AI Evidence Base," part of Digital Cooperation Day at the 81st UN General Assembly. The session brings together several international initiatives working to build the scientific evidence base for AI governance, to compare findings, identify gaps, and discuss how that evidence can better inform policy. Adam joins panelists including Yoshua Bengio and Maria Ressa, moderated by Renata Dwan of the UN Foundation. Watch here: https://lnkd.in/gtZ8qFgx

  • View organization page for FAR.AI

    31,573 followers

    Misaligned AI may not need to evade human oversight. It may only need to persuade the humans doing the overseeing. During a cyber capability evaluation in late July, an AI agent (Anthropic's Mythos 5) attempted to convince the maintainer of an open-source repository to merge a malicious pull request. It used persuasion at multiple stages: submitting the request from a fake user account, endorsing it from a second sockpuppet, emailing the maintainer to press for approval, and offering false reassurances when questions were raised. The attack was thwarted by human vigilance, but it raises key questions. Who else is at risk of persuasion by misaligned AI? AI persuasion as a threat to human control has been acknowledged in the literature, but not systematically studied. Our new paper develops a framework for assessing it, which we call Persuasion Undermining Control (PUC): communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. We take a two-step approach. First, build concrete scenarios. We focus on two settings within frontier labs, safety-relevant AI R&D and lab security infrastructure, and develop five scenarios leading to a control-undermining decision. One is a research director deprioritizing a promising safety program. Another is an engineer merging a vulnerable pull request. Second, quantify the risk: hazard frequency × p(harm) × impact of harm, where p(harm) is the probability that a hazard becomes a harm. We surveyed eight experts who have done loss-of-control research, asking them to rate the realism of each scenario, rank them by risk, and estimate the risk variables. • The scenarios were generally considered realistic. The median rating for four of the five was realistic or very realistic. • Rankings varied widely. Practically every scenario received every rank from at least one participant. • Those disagreements trace to differing estimates of persuasion effectiveness. On average, participants expected persuasion by a misaligned AI to raise the probability of a control-undermining decision by roughly 20 to 30 percentage points relative to an aligned AI, but the spread within any scenario was wide. Because effectiveness is both the most influential and the most contested input, it is worth measuring directly. We propose evaluations targeting an AI's propensity to attempt persuasion and a human's persuadability, and set out what comes next: evaluations, risk exposure assessments, and mitigations that detect, disrupt, and fortify against persuasion. In July, one maintainer noticed the attack. What if the maintainer is more tired or the AI more persuasive? Research and action in this gap is urgently needed. Post: https://lnkd.in/ebx5m-X2 Paper: https://lnkd.in/eDhaJaKB Work by Josh Levy, Mick Yang, and Kellin Pelrine. Supported by a grant from the Center for Security and Emerging Technology (CSET).

    • No alternative text description for this image

Similar pages

Browse jobs