
OpenAI
Mid
Close the capability overhang between power users and average consumers — build evals, training data, and reward signals that make ChatGPT more capable for everyone.
OpenAI's Personal AGI North Stars team is hiring a Research Engineer/Scientist to improve model capability and behavior across tool-use, feature discovery, connectors, and instruction following, deployed to millions of ChatGPT/API users. You'll own a research agenda, build evals to track capability improvements, and work across a large ML codebase to translate behavioral bottlenecks into training data and reward signal changes. This interview probes ML research judgment, eval design rigor, and hands-on ability to debug and iterate in a complex research stack.
Practice this interview
Free · a live voice mock calibrated to this exact role
What this interview tests
- Identifying and closing model capability bottlenecks (tool-use, instruction following, feature discovery)
- Building robust evals that track real capability/behavior improvements
- Translating eval findings into training data and reward signal changes
- Debugging and iterating within a large, complex ML research codebase
- Balancing independent research ownership with cross-team collaboration
Common question themes
Walk through how you'd design an eval for a specific model capability gap
Describe translating an eval result into a concrete training data or reward signal change
Tell me about debugging a subtle behavior issue in a large ML codebase you didn't build
How do you avoid an eval being gamed or failing to generalize
How would you prioritize a research agenda that touches tool-use, connectors, and instruction following simultaneously
How candidates describe it
Real Research Engineer/Research Scientist interview stories — retold from candidates' public write-ups, with sources.
Amazon · Applied Scientist (L4)OfferAmazon Applied Scientist Interview Experience: Alexa Speech Team, 2021
A redirected recruiter call turned into an Amazon Applied Scientist loop with the Alexa Speech team: a phone screen, a split five-round virtual onsite across two teams, a bar raiser, and an added ML-breadth round, ending in an offer with a downlevel from L5 to L4.
Interviewed July 2021 · Boston, MA (Remote)
Google · L5 Machine Learning EngineerOfferGoogle L5 machine learning engineer interview: phone screen skipped, four technical rounds, and an offer
A candidate applying for an L5 machine learning engineer role at Google had the phone screen skipped due to a referral and prior tenure at the company, then went through four technical and design rounds plus a behavioral round before receiving an L5 offer. The loop was one leg of a broader search that produced offers from several companies in the same cycle.
Interviewed April 2022 · Remote
Related interviews

OpenAI
Senior
Forward Deployed Engineer

OpenAI
Senior
Data Scientist, Identity

OpenAI
Senior
AI Deployment Engineer

Notion
New grad
Software Engineer, New Grad (AI)

Amazon
Senior
Research Software Engineer, Calibration, MQS Center for Quantum Computing

Databricks
Senior