FAR.AI is hiring a Research Lead to develop and lead our work on pre-training safety, shaping models’ capabilities and internal representations at their source, rather than trying to fix them after the fact.
Our initial focus is capability control: removing harmful capabilities while preserving benign ones. We see this as a promising way to prevent misuse of open-weight models in areas such as CBRN and cyber by removing offensive capabilities, and reducing loss-of-control risks by removing knowledge of oversight mechanisms. We will validate approaches like pre-training data filtering at scale, drive adoption of successful methods, and explore techniques such as gradient routing and unlearning..
We are scaling methods like Deep Ignorance by over an order of magnitude (>100B parameter models with >1T tokens). You will direct this work, partner with our red team to stress-test the resulting models, and analyze how well the methods scale to frontier systems.
Our research directions include:
You'll build and lead the team, set its research direction, mentor Members of Technical Staff to scale your vision, and remain hands-on enough to write code and run experiments yourself. This role offers high autonomy in an impact-driven environment, pursuing empirically grounded, scalable ML safety research.
FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.
We’re structured to support that work from early research through real-world adoption:
Independent by design. We can pursue what's most impactful based on our theory of change and share what we find publicly.
A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments.
Serious infrastructure for ambitious research. A dedicated engineering team runs our compute cluster and experiment-scaling stack, so researchers spend their time on research instead of on infra.
Setting the standard. Our events convene key decision makers; our red-team works with frontier developers and governments; and our communications inform the public. Together, this drives adoption and sets the new standard in safety.
Since our founding in July 2022, we've grown to 50+ staff, published 40+ academic papers, and convened leading AI safety events. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026, and ICLR, and features in the Financial Times, Nature News, Wired Magazine and MIT Technology Review. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments including the EU AI Office and publish the AI Security Leaderboard based on our red-teaming expertise. We help steer and grow the AI safety field through developing research roadmaps with renowned researchers such as Yoshua Bengio; running FAR.Labs, an AI safety-focused co-working space in Berkeley housing 40+ members; and supporting the community through targeted grants to technical researchers.
We explore promising research directions in AI safety and scale up only those showing a high potential for impact. When an approach proves effective, we develop it into a minimum viable demonstration and work with AI developers and governments to support real-world adoption.
Our recent and ongoing research includes:
Adversarial Robustness: working to rigorously solve security problems through building a science of security and robustness for AI, from demonstrating superhuman systems can be vulnerable, to scaling laws for robustness and jailbreaking constitutional classifiers.
Mechanistic Interpretability: finding issues with Sparse Autoencoders, probing deception using AmongUs, understanding learned planning in SokoBan, and interpretable data attribution.
Red-teaming: conducting pre- and post-release adversarial evaluations of frontier models (e.g. Claude 4 Opus, ChatGPT Agent, GPT-5); developing novel attacks to support this work.
Evals: developing evaluations for new threat models, e.g. persuasion and tampering risks, and launching a new research agenda on eval awareness
Mitigating AI deception: studying when lie detectors induce honesty or evasion, and developing approaches to deception and sandbagging.
Applied Interpretability: using interpretability to tackle concrete safety problems (better probes, backdoor detection, deception monitoring), aiming for fast feedback loops, often in collaboration with our other pods.
Research Leads define and own a research workstream end-to-end. Day-to-day, that means:
This role would be a great fit if you:
This role would be a poor fit if you:
To be a strong candidate for the Research Lead - Pre-Training Safety role, you likely:
It is preferable if you:
If you are missing key leadership experience or are earlier in your career, we encourage you to consider the open Research Scientist pathway and invite you to contribute to one of our existing agendas.
We're also open to more senior versions of this role; simply apply or reach out to talent@far.ai.
If based in the USA or Singapore, you will be an employee of FAR.AI (501(c)(3) research non-profit / non-profit CLG). Outside the USA or Singapore, you will be employed via an EOR organisation on behalf of FAR.AI or as a contractor.
If you have any questions about the role, please do get in touch at talent@far.ai.
If you have any questions about the role, feel free to contact us at talent@far.ai. Otherwise, if you don't have questions, the best way to ensure a proper review of your skills and qualifications is by applying directly via the application form. Please don't email us to share your resume (it won't have any impact on our decision). Thank you!
Keep listings trustworthy for everyone.