Search Jobvertise Jobs
Jobvertise

Research Engineer AI Alignment & Evaluation
Location:
US-CA-San Francisco
Jobcode:
9084cb00-05eb-4333-94cf-cbd73cab2925
Email Job | Report Job

Report this job





Incorrect company
Incorrect location
Job is expired
Job may be a scam
Other







Apply Online
or email this job to apply later

Research Engineer AI Alignment & Evaluation

AI Safety / Research Engineering | San Francisco, CA | Hybrid / In-Person

About the Company

We are representing a high-growth AI research organization working at the intersection of frontier model evaluation, AI safety, and security.

The team develops sophisticated evaluation environments designed to surface undesirable or misaligned model behavior and help leading AI organizations better understand how advanced systems behave under complex, long-horizon conditions.

This is a technically rigorous environment for engineers who are interested in AI alignment, agent behavior, model evaluation, and building systems that help make increasingly capable AI more reliable and controllable.

The Role

This is an opportunity to join a small, highly technical team as a Research Engineer with significant end-to-end ownership.

You will independently design and build evaluation environments that test frontier AI systems for subtle forms of undesirable behavior. You will own the full lifecycle of each environment, from initial concept and failure-mode identification through implementation, grader development, testing, measurement, and refinement.

A significant part of the role involves working directly with advanced LLM agents: prompting them to perform technical tasks, reviewing their output, identifying subtle errors, and making judgment calls where current models still fall short.

The role is ideal for a strong software engineer or technical researcher who enjoys ambiguous problems, learns new domains quickly, and is deeply interested in AI alignment and security.

What You'll Do

  • Design and build complex evaluation environments for frontier AI models.
  • Own evaluation projects end to end, including ideation, implementation, testing, grading, measurement, and iteration.
  • Investigate potential model failure modes and identify ways advanced agents may exploit or circumvent intended constraints.
  • Develop and improve software infrastructure used to isolate, reproduce, and evaluate model behavior.
  • Work extensively with LLM-based agents to accelerate implementation and research workflows.
  • Review agent-generated work critically and identify subtle technical or conceptual errors.
  • Build long-horizon tasks that operate near the edge of current model capabilities.
  • Apply strong qualitative judgment when evaluating behavior that cannot be captured through simple automated metrics.
  • Rapidly learn unfamiliar technical domains as required by individual evaluation environments.
  • Share findings, lessons, and technical context with a highly collaborative research and engineering team.

What We're Looking For

  • 1+ years of experience in software engineering, machine learning engineering, technical research, or a closely related field.
  • Strong traditional software engineering fundamentals.
  • Proficiency with Python(link removed)>
  • Strong interest in AI alignment, AI safety, or AI security.
  • Ability to reason carefully about complex systems and ambiguous failure modes.
  • Strong conceptual judgment and the ability to think through how an autonomous agent may interpret or exploit a task.
  • Ability to learn new technical domains quickly.
  • Experience using LLMs or AI agents effectively as part of technical workflows.
  • Strong ability to assess whether agent-generated work is correct, including when errors are subtle.
  • Comfortable taking full ownership of technically demanding projects with limited oversight.
  • High standards for quality, execution, and accountability.

Nice to Have

  • Experience building evaluation frameworks, benchmarks, simulation environments, or agent-based systems.
  • Exposure to frontier language models or autonomous agent workflows.
  • Background in AI safety, alignment research, adversarial testing, or security.
  • Experience designing tasks that require multi-step or long-horizon reasoning.
  • Research experience involving model behavior, reward hacking, robustness, or control mechanisms.

Why This Role Is Exciting

  • Own technically challenging research environments from concept through final evaluation.
  • Work directly with state-of-the-art AI systems and agentic workflows.
  • Tackle problems at the frontier of AI safety, model behavior, and alignment.
  • Join a small technical team where individual work has meaningful visibility and impact.
  • Operate with substantial autonomy while receiving frequent technical feedback.
  • Build expertise across a wide range of domains rather than working within a narrow product surface.
  • Contribute to work focused on understanding and mitigating undesirable AI behavior rather than simply increasing model capabilities.

Work Model

  • Full-time position.
  • San Francisco-based role with regular in-office collaboration expected.
  • Flexibility around hybrid working arrangements.
  • Open to candidates willing to relocate.
  • Visa transfers and new visa sponsorship may be available.
  • Work is highly ownership-driven, with emphasis on the quality of what you ship.

Confidential details removed: salary, client name, founder names, exact address, company links, investor names, funding details, exact team size, founding year, and highly identifiable wording.

W3 Sourcing

Apply Online
or email this job to apply later


 
Search millions of jobs

Jobseekers
Employers
Company

Jobs by Title | Resumes by Title | Top Job Searches
Privacy | Terms of Use


* Free services are subject to limitations