App
Principal AI Data & Evaluation Strategist
NEWUnited StatesContractorGlobal
š° USD 312,000 - 312,000/yr
š Executiveš Remote
ActivePosted within the last 30 days
Job Description
[AI-summarized by JobStash]
You will serve as a consultative expert and thought partner alongside program managers, providing strategic guidance during project discovery, solution design, data planning, evaluation-framework creation, quality-rubric development, and pilot execution. You'll work directly with stakeholders, engineering teams, researchers, and program leadership to define project requirements before execution begins. This is not an operational-delivery role; you'll bring subject-matter expertise, challenge assumptions, identify risks early, and ensure projects are designed for quality, scalability, and measurable outcomes. It is a remote, flexible engagement contributing to frontier AI work.
Requirements
- ā10+ years in AI, data science, machine learning, research, or AI operations
- āExperience designing large-scale AI data programs
- āExperience with LLM evaluation, human feedback systems, AI benchmarking, and dataset development
- āStrong written English and the ability to communicate complex ideas clearly
- āComfortable working independently in a remote environment
- āBased in the United States
Responsibilities
- āLead project discovery and translate business goals into AI data requirements
- āDefine project success metrics and identify risks and failure modes early
- āDesign training and evaluation datasets and define taxonomy structures
- āRecommend data-sourcing methodologies and establish ground-truth standards
- āIdentify coverage gaps and edge cases
- āDevelop human-evaluation frameworks and quality-measurement methodologies
- āEstablish acceptance criteria and design benchmarking approaches
- āBuild calibration mechanisms
- āChallenge assumptions that may impact project quality
- āRecommend industry best practices and guide teams through ambiguity
- āSupport executive and client reviews
Tech Stack
AI Data Creationcalibrationdata scienceDataset DevelopmentData sourcingGround Truth StandardsHuman EvaluationHuman feedbackLLM evaluationmodel benchmarkingproject:Braintrust