DataArt
Senior AI Evaluation Engineer with LangGraph
remote • ARM • senior • Не указана
# Senior AI Evaluation Engineer with LangGraph
ArmeniaBulgariaCyprusGeorgiaKazakhstanLatviaPolandRomaniaSerbiaUkraine
## Technology stack
LangGraph, AWS AgentCore Evaluations, Python, LLM as Judge Frameworks, CI/CD Pipelines, OpenTelemetry, Agent Based Systems, Automated Evaluation Frameworks, Large Language Models, A/B Testing, Shadow Mode Validation, Deployment Automation, Quality Monitoring Tools
## Project overview
The project focuses on establishing a comprehensive evaluation and quality assurance platform for agent based AI systems. The solution provides automated testing, deployment validation, production feedback integration, and continuous quality monitoring to ensure high standards for agent performance and user experience.
## Team
Medium team (10-20 people)
You will work with AI engineers, ML engineers, platform engineers, software developers, and product stakeholders in a collaborative environment focused on quality, reliability, observability, and continuous improvement. The team is responsible for defining evaluation standards and operationalizing AI quality across enterprise platforms.
## Position overview
We are looking for a Senior AI Evaluation Engineer to design and implement enterprise grade evaluation frameworks for agent based AI systems. In this role, you will build automated quality assessment capabilities, define deployment gate strategies, and create scalable evaluation methodologies that improve the reliability, accuracy, and safety of AI driven solutions throughout the development lifecycle.
## Responsibilities
- Design and develop build time evaluation frameworks for LangGraph based agent systems
- Create automated test harnesses for graph level and node level validation
- Design evaluation strategies that combine deterministic grading and LLM as judge methodologies
- Implement multi layer evaluation frameworks covering tool selection accuracy, execution trajectory quality, reasoning effectiveness, and output quality
- Develop multi turn conversation simulations and context retention scoring mechanisms
- Implement multi trial reliability testing methodologies, including pass at k and pass power k approaches
- Design and maintain CI/CD deployment gates that validate quality metrics before production releases
- Build staging validation, shadow mode comparison, and controlled rollout evaluation workflows
- Integrate production evaluation feedback into build time testing frameworks to improve quality coverage
- Collaborate with platform and engineering teams to establish evaluation standards, thresholds, and governance practices
- Analyze evaluation results and provide recommendations for improving agent reliability and performance
- Contribute to technical documentation, testing standards, and quality engineering best practices
## Requirements
- 4+ years of experience building automated testing frameworks, evaluation platforms, or quality assurance solutions for machine learning, large language model, or agent based systems
- Hands on experience designing multi layer evaluation frameworks that combine deterministic and LLM as judge grading approaches
- Experience implementing automated quality gates that can block deployments based on predefined metric thresholds
- Experience working with LangGraph or a comparable agent orchestration framework
- Strong understanding of agent behavior evaluation, workflow validation, and AI quality measurement techniques
- Experience designing scalable testing and validation processes for production AI systems
- Strong Python development experience
- Knowledge of CI/CD practices, deployment automation, and release governance
- Experience analyzing evaluation data and translating findings into platform improvements
- Strong communication and collaboration skills
## Nice to have
- Hands on experience with AWS AgentCore Evaluations, including evaluation execution and custom evaluator development
- Experience designing and operating shadow mode or canary deployment strategies for machine learning or AI systems
- Experience creating automated feedback loops that convert production incidents into regression test scenarios
- Knowledge of production observability, monitoring, and evaluation pipelines
- Experience with enterprise AI governance and quality assurance programs
- Understanding of agent observability and telemetry driven quality improvement processes
## FAQ for Candidates
Work on global projects, grow your career in a supportive, flexible, and innovative tech environment. We help cover the cost of IT certifications and provide access to top-tier courses and learning platforms. View current openings and take the next step with us.
### What does DataArt do, and which industries does it serve?
DataArt is a global software engineering company that helps businesses build powerful data, analytics, and AI solutions. We work with clients across a range of industries — including Finance, Healthcare & Life Sciences, Consumer Goods & Retail, Travel, Media & Entertainment, Mobility, and Manufacturing.
Learn more about what we do [here](https://www.dataart.com/company/about-us).
### What's the work-life balance like at DataArt? Does DataArt offer remote work options?
DataArt supports flexible work formats to help you find the balance that works best for you. You can choose to work from the office, go hybrid, or stay fully remote—each option comes with equal opportunities for growth. We’ll help set you up with secure access and the equipment you need. With 46 remote and onsite official locations, you can join us from almost anywhere in the world.
Learn more about how we work [here](https://www.dataart.team/career/workplace-options).
### What is DataArt's company culture like?
At DataArt, we put people first—fostering a culture built on trust, flexibility, and professional growth. We believe in open communication, mutual respect, and the freedom to choose how and where you work. Diversity, equity, and inclusion are core to our values, and we actively support a workplace where everyone can thrive. From mental health support to global sustainability efforts, we aim to create a healthy, empowering environment for all.
[Read more about our culture and values](https://www.dataart.team/career/culture).
### What is the typical career path at DataArt?
There's no one-size-fits-all career path at DataArt—your growth is yours to shape. Whether you want to deepen your technical skills, move into management or sales, or even switch professions entirely, you'll have the support to do it. With exam fees fully covered, you can also earn professional certifications, like AWS, Azure, or Google Cloud. With tools like the Professional Development Map, the Talent Lab, and access to expert mentoring, we help you build your desired career.
Learn how we support both [newcomers starting their careers](https://www.dataart.team/start-career) and experienced [professionals looking to grow further](https://www.dataart.team/advance-career).
### What can I expect in terms of compensation and benefits at DataArt?
DataArt offers competitive compensation along with a range of thoughtful benefits that support your well-being, growth, and daily comfort. You’ll get flexible vacation and sick leave, mental health programs, and access to a corporate laptop or BYOD option. We also offer bonuses for referrals, parental leave, and smooth exit and return-to-work options.
[Read more about how things work at DataArt](https://www.dataart.team/career/culture).
### How does DataArt approach employee retention and turnover?
At DataArt, we focus on creating an environment where people want to stay and grow. With 95% of employees recommending us to a friend on Glassdoor and a 100% CEO approval rating, we’re proud of the trust we’ve built. From mentoring programs and professional development services to support and conflict resolution programs, we invest in our colleagues' growth and wellbeing—because when people feel valued, they stick around.
Learn more about our company >[here](https://www.dataart.team/).
### What is the interview process like, and how can I prepare?
Our interview process is designed to be thorough, transparent, and supportive. It starts with a CV review, followed by an HR interview to discuss your background and goals. You'll then go through communication and technical assessments, where we evaluate your English skills and knowledge of relevant technologies. If all goes well, you’ll meet someone from the project team to learn more about the work, and we’ll guide you every step of the way—including helping you prep for any client interviews.
Read more and get tips for each stage [here](https://www.dataart.team/career/how-to-get-in).
### What skills and experience does DataArt look for in candidates?
DataArt looks for candidates with strong technical skills, analytical thinking, and a commitment to continuous learning. We value experience in relevant technologies and industries, adaptability, and excellent communication skills. If you’re a junior or just starting out, don’t worry—our mentorship programs and training will support your growth from day one. The most important thing is your willingness to learn and develop professionally in our collaborative, people-first environment.
[Read more about development in DataArt.](https://www.dataart.team/start-career)
### How does DataArt support learning and professional development?
At DataArt, continuous learning is part of our culture. You’ll have access to expert mentors, leadership guidance, and thousands of courses from top platforms like Udemy and LinkedIn Learning. You can shape your own path—whether it’s growing within your role or switching to a new one.
[Explore how we support your growth.](https://www.dataart.team/start-career)