Skip to main content
NEUN
Back to Careers

App

Principal AI Data & Evaluation Strategist

NEW
United StatesContractorGlobal

šŸ’° USD 312,000 - 312,000/yr

šŸ“Š ExecutivešŸ  Remote
ActivePosted within the last 30 days

Job Description

[AI-summarized by JobStash]

You will serve as a consultative expert and thought partner alongside program managers, providing strategic guidance during project discovery, solution design, data planning, evaluation-framework creation, quality-rubric development, and pilot execution. You'll work directly with stakeholders, engineering teams, researchers, and program leadership to define project requirements before execution begins. This is not an operational-delivery role; you'll bring subject-matter expertise, challenge assumptions, identify risks early, and ensure projects are designed for quality, scalability, and measurable outcomes. It is a remote, flexible engagement contributing to frontier AI work.

Requirements

  • ā—10+ years in AI, data science, machine learning, research, or AI operations
  • ā—Experience designing large-scale AI data programs
  • ā—Experience with LLM evaluation, human feedback systems, AI benchmarking, and dataset development
  • ā—Strong written English and the ability to communicate complex ideas clearly
  • ā—Comfortable working independently in a remote environment
  • ā—Based in the United States

Responsibilities

  • ā—Lead project discovery and translate business goals into AI data requirements
  • ā—Define project success metrics and identify risks and failure modes early
  • ā—Design training and evaluation datasets and define taxonomy structures
  • ā—Recommend data-sourcing methodologies and establish ground-truth standards
  • ā—Identify coverage gaps and edge cases
  • ā—Develop human-evaluation frameworks and quality-measurement methodologies
  • ā—Establish acceptance criteria and design benchmarking approaches
  • ā—Build calibration mechanisms
  • ā—Challenge assumptions that may impact project quality
  • ā—Recommend industry best practices and guide teams through ambiguity
  • ā—Support executive and client reviews

Tech Stack

AI Data Creationcalibrationdata scienceDataset DevelopmentData sourcingGround Truth StandardsHuman EvaluationHuman feedbackLLM evaluationmodel benchmarkingproject:Braintrust
Expired
Search