TQUSI0980_6003 - AI Evaluation Engineer

Job Type: Contract

Work Mode: Remote

The project: 


We are building a GenAI regulatory science co-pilot for regulatory submission teams. It works like a virtual Health Authority reviewer. It extracts and assesses regulatory risks, maps submission content to relevant regulatory guidance and precedent, checks that key messages are supported by the underlying clinical data, and anticipates likely Health Authority questions. This system directly informs regulatory submission strategy, so its outputs must be accurate, grounded in cited sources, and traceable. We are hiring an experienced

contractor to build and run the evaluations for the system’s core reasoning engines, working with regulatory affairs subject-matter experts (SMEs) and the product engineering team.


What you will do

  • Build gold sets with SMEs: annotate regulatory and clinical documents with the expected outputs (for example, extracted risks, relevant guidance and precedent, and whether the data supports each key message). 
  • Write the annotation guidelines, run adjudication, and measure inter-rater agreement.
  • Evaluate extraction: score risk extraction for precision, recall, and completeness, and compare successive engine versions against a baseline.
  • Evaluate retrieval and ranking: measure retrieval recall (Recall@K), ranking quality (MRR, nDCG), and question coverage when the engine retrieves regulatory guidance, precedent, and prior Health Authority interactions.
  • Evaluate grounding and traceability: build code-based checks that every finding traces through evidence and document back to a valid source and citation.
  • Detect fabricated guidance, fabricated precedent, and unsupported regulatory claims.
  • Evaluate reasoning with LLM judges: design and validate LLM-as-judge rubrics for regulatory relevance, evidence sufficiency, risk assessment, and explanation quality, calibrated against SME judgment.
  • Evaluate each pipeline stage: create evaluation datasets and baselines for each stage of a multi-stage engine (evidence structuring, question coverage, retrieval, ranking, mapping, grounding), plus an end-to-end review with SMEs.
  • Report and harden: turn failure analysis into clear readouts for the product team, and package the evaluations as regression suites that run in CI/CD.


Required experience

  • 5+ years of hands-on work in AI/ML evaluation, NLP, or applied data science, including recent delivery on LLM, RAG, or agent systems.
  • Has built evaluation frameworks and gold datasets for information extraction and retrieval/ranking systems. Please share examples.
  • Strong grounding in evaluation metrics and statistics: precision/recall/F1,
  • Recall@K, MRR, nDCG, inter-rater reliability (e.g., Cohen’s kappa or
  • Krippendorff’s alpha), and confidence intervals for small evaluation sets.
  • Experience designing and validating LLM-as-judge systems against human expert ratings.
  • Experience running annotation projects with domain experts: guidelines, adjudication, and quality control.
  • Strong Python; experience with evaluation harnesses (e.g., Ragas, DeepEval, promptfoo, Inspect, or in-house equivalents), LLM APIs, and Git/CI/CD.
  •  Can deliver against fixed milestones with little onboarding and work independently.
  • Fluent professional English, written and spoken.


Nice to have

  • Experience with regulatory affairs, regulatory intelligence, or submissions:
  • FDA/EMA guidance, CTD structure (e.g., Modules 2.5 and 2.7), Health Authority questions, and labeling.
  • Experience evaluating AI in regulated (GxP) environments.
  • Familiarity with clinical trial outputs (CSRs, TLFs, statistical outputs) and drug development.
  • Experience evaluating document-comparison, claim-verification, or fact-checking systems.
  • Publications or open-source work in AI evaluation, information retrieval, or biomedical NLP.


Want To
WORK FOR YOU?

GET THE QUOTE

Want To
WORK WITH US?

CAREER