All jobs
    Cursor logo

    Software Engineer, RL Data

    Original title · English translation pending

    CursorSan Francisco, United StatesPosted 2026-09-15Last seen in source Website
    Location
    San Francisco, United States
    Employment
    FullTime
    Workplace
    OnSite
    EngineeringAI
    Apply on the company site

    Role description

    Original description · English translation pending

    Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

    Software Engineer, Reinforcement learning

    • SpaceXAI is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.

    About the role

    • As a Software Engineer on the RL Data team at SpaceXAI, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.

    What you’ll do

    • Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.
    • Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.
    • Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
    • Partnering with research on whether a dataset is actually teaching the thing we think it is.

    You may be a fit if

    • You write careful, fast code and have strong software engineering fundamentals.
    • You like setting tasks: breaking a fuzzy capability into something concrete you can measure.
    • You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.
    • You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.

    Listing sourced from ashby. wwshemi does not process applications.