San Francisco, CA · On-site · Full-time Compensation: $150,000–$250,000 base + profit sharing (total cash ~$250K–$450K) + equity
About the Company
Our client builds the training data and evaluation infrastructure that frontier AI labs use to improve their models, partnering with leading labs to design high-signal datasets and run rigorous evaluations that go beyond static benchmarks. It's a small, early team where individual contributors have direct impact on how the next generation of models learns and improves. The founding team comes from top quant firms, big tech, and leading AI research labs.
Founded 2025 · 11–50 people · Industry: AI / ML
The Role
Prove that the client's data works — design and run training experiments that isolate the impact of its datasets on model behavior (SFT and RL post-training), and turn the results into defensible evidence for partner labs.
Tech stack: LLM post-training (SFT, RL).
What you'll be doing
- Run controlled SFT and RL experiments to measure dataset impact on model performance
- Quantify lift across capabilities including reasoning, tool use, long-horizon tasks, and domain-specific workflows
- Communicate findings with partner labs to drive sales
- Work with internal delivery leads to iterate on data quality based on experimental results
Requirements
- Strong familiarity with LLM training and evaluation methodologies
- Able to design lightweight experiments and extract actionable insights from messy results
- Comfortable working across multiple domains including finance, software engineering, and policy
Nice to Haves
- Has run controlled post-training experiments end to end and can point to a specific data intervention that shifted model behavior measurably
- Comfortable reading messy experimental results without needing clean data to find signal
- Strong quantitative instincts paired with SWE ability — can actually ship the experiment, not just design it
- Has worked adjacent to or inside frontier labs or eval orgs, with a real sense of what "high-signal data" means
Why Join
- Frontier-lab leverage: your work directly shapes the datasets leading AI labs use to train next-gen models
- Strong backing and pedigree: a founding team from top quant firms, big tech, and leading AI research labs
- High cash comp: base plus profit sharing pushes total cash to ~$250K–$450K, with equity on top
- Build, don't theorize: experimental, high-leverage IC work at the edge of model development
Details
- Location — San Francisco, CA
- Work policy — On-site
- Compensation — $150,000–$250,000 base + profit sharing (total cash ~$250K–$450K) + equity
- Visa sponsorship — Not available
- Employment type — Full-time