F
Figma

Director, Research - AI Evals

San Francisco, CA • New York, NY • United StatesEnglishContractorFull timeFigma is growing our team of passionate creatives…
StrategyOperationsProductEngineeringAI & Data
Want to know how this fits you?

Give Aptiora your real career evidence. We’ll compare it with this opportunity without inventing anything.

Find my fit
At a glance
LocationSan Francisco, CA • New York, NY • United States
Work styleNot specified
TypeContractor
ScheduleFull time
RemunerationFigma is growing our team of passionate creatives and builders on a…
Start dateIf you're excited to shape the future of design and…
DeadlineNot stated
Required languagesNot stated
The opportunity

About the role

The ideal candidate brings deep, hands-on experience evaluating AI/LLM-powered products — blending human evaluation with automated, model-based approaches — along with the product instinct and communication skills to make evaluation genuinely useful. Partnering with Product, Design, Engineering, and Data Science, you'll sit upstream of nearly every AI shipping decision at Figma and directly shape the quality of features used by millions of people.

Stay in Aptiora while you decide.

Responsibilities, requirements, fit, evidence and preparation are organized here. The original posting stays available for final verification.

Conditions

Conditions

  • This is a full time role that can be held from one of our US hubs or remotely in the United States.

Eligibility

Exempt employees are eligible for employer‑provided paid flexible PTO in addition to flexible paid sick leave. Figma also offers sales incentive compensation for most sales roles and an annual bonus plan for eligible non-sales roles.

How to apply

  • What you'll do at Figma: Own AI evaluation methods and operations for Figma's AI-powered experiences — define quality dimensions, design how we measure them, and turn results into decision-ready signal Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars, combining human evaluation with automated/model-based approaches (e.g., LLM-as-judge) where appropriate Partner with engineering to stand up repeatable, reproducible evaluation pipelines and regression testing, so evaluation is a routine part of how AI features are built and shipped Produce clear readouts and dashboards that let stakeholders confidently make go/no-go and prioritization decisions Socialize a shared definition of quality so evaluation standards are adopted across teams rather than re-invented — and advocate for evaluation as a strategic partner in the product process Manage a small team to execute our AI evals in partnership with contractors, internal staff, and/or LLMs We'd love to hear from you if you have: 10+ years of experience in product, research, applied research, or a closely related field, including 2+ years of management experience Direct, hands-on experience owning the evaluation of AI/LLM-powered products Expertise designing and running AI evaluation — human evaluation programs, rubric and benchmark/golden-dataset construction, inter-rater reliability — and sound judgment about when and how to apply automated/model-based approaches (e.g., LLM-as-judge), including their limitations Strength across both qualitative and quantitative methods, comfort with data and metrics, and the ability to reason about model behavior Demonstrated success in identifying the riskiest assumptions behind an ambiguous quality question, prioritizing them, and designing right-sized evaluation to build confidence A proven track record of gaining buy-in from executive and cross-disciplinary stakeholders — transcending methodology to articulate a larger user story and the "so what" to inspire action While it's not required, it's an added plus if you also have: Experience building or co-building automated evaluation pipelines and regression testing in partnership with engineering, or familiarity with eval tooling (e.g., Braintrust, LangSmith, DeepEval, or equivalents) Experience standing up a new function, practice, or discipline from scratch 2+ years in product design, user-centric product management, data science, product development, and/or front-end engineering A familiarity and depth of experience using Figma's products At Figma, one of our values is Grow as you go.
  • If you’re excited about this role but your past experience doesn’t align perfectly with the points outlined in the job description, we encourage you to apply anyways.
  • We will work to ensure individuals with disabilities are provided reasonable accommodation to apply for a role, participate in the interview process, perform essential job functions, and receive other benefits and privileges of employment.
  • By applying for this job, the candidate acknowledges and agrees that any personal data contained in their application or supporting materials will be processed in accordance with Figma's Candidate Privacy Notice.
The organization

About Figma

Verified

We have limited verified information about this company. You can still explore its active opportunities and check the official source.

161active opportunities
15observed locations
3work modes observed
AITechnologyHealthEducationFinancePublic ImpactMobilityEnergy
Keep exploring Aptiora

If this one isn’t right.

Related live opportunities you can compare without restarting your search.

Source & verification

Direct employer source

Aptiora keeps the source available for trust and final verification while the working experience stays here.

✓ Checked 21 minutes ago