Independent evaluation for AI-decision making for Automated Vehicles
Project PREP0005052 · NIST sponsor Zeid Kootbally
Overview
The Measurement Science for Automated Vehicles (MSAV) project at NIST is seeking a candidate to
support its effort on independent evaluation of AI-enabled decision-making in automated driving
systems. This effort develops a measurement science framework and an open-source toolkit that
evaluate ADS decision quality as a black box, across four progressive tiers (real-time safety monitoring,
predictive outcome analysis, decision comparison against reference baselines, and systematic weakness
diagnosis), without access to proprietary algorithms. The candidate will focus on two core areas: 1)
simulation engineering for scenario-based ADS testing in CARLA and 2) software integration for the
open-source evaluation toolkit and Evaluation Gateway.
Qualifications
- MS required (PhD preferred) in Computer Science, Robotics, AI/Machine Learning, or related engineering fields.
- Strong programming experience in Python and C++.
- Experience with autonomous vehicle simulation environments (CARLA, SUMO, or similar) and scenario description languages (OpenSCENARIO).
- Experience building software toolkits, APIs, and standardized interfaces; familiarity with containerized deployment (Docker).
- Knowledge of autonomous vehicle systems architecture and behavioral planning concepts.
- Experience with ROS 2 on Linux systems.
- Experience with version control software and workflow (Git/GitHub/GitLab).
- Familiarity with data modeling, schema design, and validation methodologies.
Research Proposal
Key responsibilities will include but are not limited to:
Simulation Engineering
- Manage the CARLA simulation infrastructure and build the scenario execution pipeline that runs OpenSCENARIO 2.0 test cases.
- Configure multi-agent traffic behavior using layered modeling (scripted, reactive car- following models such as IDM and MOBIL, stochastic, and learned agents) for large-scale evaluation runs.
- Integrate hardware-in-the-loop testing and optimize performance for large-scale multi- agent test runs.
- Support integration of open-source automated driving stacks (for example, Autoware) to test against diverse architectures.
Software Integration
- Develop the open-source evaluation toolkit and design the Evaluation Gateway API and standardized interfaces.
- Implement the standardized behavioral data logging schema that defines the observable ADS output required for interoperable evaluation.
- Build the containerized deployment pipeline (for example, Docker) and the Cryptographic Hashing Message Schema enabling any ADS to connect as a black-box client without exposing source code, model weights, or training data.
- Produce automated reporting templates and technical documentation supporting reproducible, auditable evaluation.