This job has conditions that may affect your decision to apply
English is required · C2 level
Description
Senior-level AI Quality Engineer role at Accenture Netherlands focused on evaluating agentic AI systems — autonomous agents, multi-step workflows, and RAG applications. Combines hands-on automation of quality evaluation (LLM-as-a-Judge, schema validation, CI/CD test harnesses) with red teaming, adversarial security testing, and production monitoring. Involves consulting and stakeholder alignment to translate business risks into measurable quality criteria and release recommendations.
Responsibilities
Automate evaluation of multi-turn conversations and long-horizon tasks using metrics for reasoning accuracy, evidence grounding, plan execution, latency, token usage, and recovery
Apply schema validation, tracing, and calibrated LLM-as-a-Judge frameworks to detect inaccurate tool calls, loops, and misuse
Implement filters and golden datasets to test harmful content, policy violations, hallucinations, bias, and brand drift
Use human-in-the-loop review for high-risk edge cases
Run red-team, prompt-injection, jailbreak, privacy, and PII-leak tests across agent workflows
Analyze failures in model logic and API integrations to identify root causes and controls
Embed test harnesses in CI/CD and monitor production telemetry for regressions, behavioral drift, and recovery
Maintain evaluation data lineage, versioning, curation, and golden responses
Translate business risks into measurable quality criteria, practical controls, and release recommendations
Align stakeholders on user journeys, edge cases, thresholds, mitigations, testability, observability, governance, and evaluation strategy
Communicate findings and residual risks for go/no-go decisions; coach teams and share reusable assets and standards
Requirements
Bachelor's or Master's in Computer Science, Data Science, or a related quantitative field
3+ years in machine learning QA, AI evaluation engineering, or MLOps engineering
Strong Python proficiency and experience with AI evaluation tools (MLflow, Galileo, PyRIT)
Experience evaluating autonomous agents, multi-step workflows, or RAG applications
Deep knowledge of prompt engineering, vector databases, tool schema design, and execution telemetry
Hands-on experience in adversarial testing, threat modeling, or model safety guardrails
Strong communication, technical leadership, and growth mindset
Dutch B2+ for collaboration with Dutch-speaking clients and stakeholders
English C2
TMAP or ISTQB Certification (bonus)
Contributions to open-source projects or engineering communities (bonus)
Conditions
4x9 Workweek Option, Hybrid Working, FlexDay
Competitive Base Salary, MyBenefits Budget, Annual Bonus, Net Allowance, Employee Share Purchase Plan
Collective Health Insurance Scheme, Mental Health Support, Well-being Hub
26 Vacation Days, Special Leave Types (parental leave, partner leave, bereavement leave, gender affirmation leave), Culture Days
NS Business Card, Mobility Budget (electric bike, car lease, or public transport reimbursement)
Performance Achievement Program, Learning & Development Platform, Guidance from People Leads