About Neolithic
Neolithic is a nonprofit startup in San Francisco working to reduce catastrophic risk from AI. We build open-source, agentic tools that automate and scale AI safety research: tooling for control experiments, safety evaluations, safety training datasets, and infrastructure for research agents. Neolithic was founded in 2026 by Leo McKee-Reid and Sevan Hayrapet.
Why this role matters
Capabilities research is increasingly automated. Safety research is still largely done by hand. Many of the field's bottlenecks are engineering problems that shared tooling could solve. As one of our first hires, you will set the technical direction, the culture and the hiring bar for everyone who follows.
What you'll do
- Scope, build, iterate on and maintain tools as production software.
- Work with researchers at safety organisations as design partners, removing the bottlenecks in their workflows.
- Shape technical direction, culture and hiring, and take on whatever is most needed.
Projects in your first few months
- Rogue Internal Deployment Arena — a mock AI lab, with realistic internal tooling, permissions and oversight, where AI control researchers can test monitors and control protocols against agents attempting rogue internal deployments.
- Critical Infrastructure Cyber Ranges — simulated critical-infrastructure environments for evaluating and training cyber classifiers against realistic attack and benign activity in high-risk settings.
- Agent Swarm Infrastructure — infrastructure for running many AI agents in parallel, built so every action can be logged, inspected and monitored at scale.
- Multi-agent transcript analysis — tools for analysing transcripts from multi-agent AI control experiments, so researchers can find and understand the critical moments across thousands of agent interactions.
Who we're looking for
We're looking for people who fit one of three profiles:
- Agents engineer — brings deep hands-on experience building agents and the infrastructure around them.
- Technical generalist — ships fast, builds product and does the user research to know what to build.
- Researcher turned builder — brings AI safety research experience and is moving toward heavy engineering work.
Whichever you are, we expect production engineering experience, hands-on work with LLMs or agents, fluent use of frontier AI tools, high agency, and a real motivation to reduce risk from AI.
It's a plus if you maintain a popular open-source project, have done safety research (including fellowships such as MATS or LASR), have been an early engineer at a startup, or have built serious agent or evaluation pipelines. Non-traditional backgrounds are welcome.
What you get
- High counterfactual impact on a neglected problem.
- A path to senior roles as the team grows.
- Flexible hours, no management layers and unlimited PTO.
- Benefits through our fiscal sponsor, BERI, and US visa sponsorship.
Interview process
- 15-minute call with Leo
- 3-hour work test
- 50-minute interview
- Paid in-person work trial