Toward AI Standardization: A Triadic Human-AI Collaboration Framework for Multi-Level Autonomous Mobility
1


Abstract

The goal of the current study is to introduce a triadic human–AI collaboration framework that could be applied in transportation systems such as automated vehicles, micromobility systems, and vehicle teleoperation. Previous standards (e.g., SAE Levels of Automation) have focused on defining automation levels based on who controls the vehicle. However, it is still not clear how human users and AI should collaborate in real-time, especially in dynamic driving contexts where roles can shift frequently. To fill the gap, this study proposed a triadic human-AI collaboration framework with three AI roles (i.e., Advisor, Co-Pilot, and Guardian) that can dynamically adapt to human needs based on real-time data, such as mental states and environmental conditions. The Advisor AI offers informational support without direct intervention; the Co-Pilot AI provides partial intervention when needed, with the goal of sharing control with humans; the Guardian AI performs emergency overrides if necessary. The use cases for these AI roles in the context of micromobility devices (i.e., e-scooters) are presented to demonstrate how these roles can influence user preferences and trust. Overall, the study takes a first step toward a universal role-based collaborative framework for AI standardization and explores how AI technologies can be embedded in future transportation systems while considering human interactions.

Human-Centered AI; Triadic Framework; Human-AI Collaboration; Human-AI Teaming; Vehicle Automation; AI Standardization; AI Agents

1 Introduction↩︎

In recent years, the automotive industry has been rapidly evolving. Many manufacturers and mobility providers now equip vehicles with automated systems, such as Tesla’s Full Self-Driving (FSD) system, which allows vehicles to be partially or fully driven by machines instead of human drivers [1]. With the rapid development of artificial intelligence (AI), these automated systems are becoming more capable of managing driving tasks [2] across diverse environments, ranging from urban to rural areas. In addition, these automated systems extend beyond traditional automated vehicles to include micromobility devices (e.g., e-scooters, e-bikes, and e-wheelchairs), as well as teleoperated vehicles (e.g., robotaxis like Waymo or Tesla Robocabs) [3], [4].

With AI systems increasingly integrated into these transportation modes and collaborating with humans in real-time, there is a need to redefine and clarify human–AI collaboration roles. For example, during a semi-autonomous vehicle takeover request prompted by an unexpected construction zone, AI can help drivers decide how to maneuver the vehicle, or provide partial torque on the steering wheel to assist drivers in steering through a shared-control mechanism. Traditional levels of automation classification, such as Driving Automation Levels (SAE International’s J3016 standard) [1] ranging from Level 0 (no automation) to Level 5 (full autonomy), focus on determining who (i.e., human or machine) is in control [5], [6]. However, these static levels of automation categories do not specify how humans and AI should coordinate in continuous and dynamic driving environments where human–AI interaction patterns shift frequently. In practice, the extent to which automation level is used can vary within a single trip; for example, drivers may rely on Level 3 semi-autonomous functions on the highway and then revert to manual control in dense urban traffic or work zones. Thus, applying a single static classification may fail to accurately represent the dynamic shifts during a trip.

This mismatch creates practical problems for end-users. Human operators may have difficulties in understanding their own roles as well as the AI’s. This means AIs may perform many tasks without keeping drivers in the loop, leading to a lack of system transparency and reduced user trust in the AI; in contrast, AIs may fail to provide adequate support or clear feedback, resulting in user frustration and decreased satisfaction, further lowering user trust [7]. This can be particularly important for vulnerable populations, such as older adults who may be experiencing cognitive and physical declines or people with disabilities [8]. Thus, there is a need to standardize AI roles in human-AI interaction and clarify how human drivers should partner with the AI.

The goal of this study is to propose and demonstrate the application of a novel triadic human-AI collaboration framework. This framework builds upon existing human-automation interaction models [9], [10] but introduces a novel integration of real-time data, AI system roles, and human input in intelligent mobility systems. It clarifies the roles of the human operator, the AI system, and the data input, which may inform future AI standardization in human-AI interactions.

2 Related Work↩︎

2.1 Intelligent Mobility↩︎

Surface transportation has grown rapidly with the advancement of technology. Automated vehicles (AVs) featuring capabilities such as adaptive cruise control, lane-keeping, or conditionally automated driving have become one of the most common vehicles driven on public roads. Based on SAE’s Levels of Automation, these vehicles are classified into six levels, from Level 0 (no automation) to Level 5 (fully autonomous) [1]. One significant distinction between Levels 2 and 3 is the role of human drivers, which shifts from actively monitoring and controlling vehicles (Level 2) to a more passive role in which they only intervene when necessary (Level 3). Today, most AVs on the road are still semi-autonomous and require human drivers to resume manual control at any moment, due to system failures or immediate road events such as missing lane markers or entering a construction zone [11]. This transition period, often called a takeover, involves a signal response phase and a post-takeover phase, which includes: 1) perceiving and processing signals for a takeover request, 2) processing information from the driving and vehicle environment to make timely decisions, and 3) executing the maneuvering plan [12].

While more AVs aim to be fully autonomous at Level 5, other applications of AV technology are also being explored. One example is teleoperated driving, where a remote human oversees or partially controls a vehicle when needed. This approach allows remote operators to monitor multiple vehicles simultaneously and handle edge cases, while routine driving tasks are delegated to vehicles’ onboard automated systems[4], [13]. This form of transportation has been used in applications such as robotaxis, autonomous truck fleets, and vehicles operating in hazardous areas. For example, robotaxis can handle most driving tasks themselves; however, when experiencing edge cases, they still need remote human operators to do one of the following: 1) remote driving, where operators directly control the vehicle or share control between the operator and the vehicle, or 2) remote assistance, where operators provide higher-level guidance, such as confirming decisions during uncertain conditions to the vehicle system. Similar to the in-vehicle takeover process, remote operators need time to gain situation awareness and make timely decisions after receiving the remote intervention request.

Micromobility (e.g., e-scooters, e-bikes, e-wheelchairs) has also benefited from automated systems. These transportation modes, which serve the last-mile needs of short-distance travel (e.g., less than 3 miles)[14], [15], have recently become one of the major travel options in the U.S. [16], [17]. Rather than relying on fully autonomous systems, micromobility services typically preserve active rider involvement while receiving targeted system support [18], [19]. For example, when users face unexpected scenarios on the road, such as sudden obstacles, this system could send out a warning to alert them about the potential hazard, offer maneuver support (e.g., assisting steering or braking) while the riders perform the action, or directly intervene in the vehicle to handle critical scenarios. Although the role transitions differ across critical contexts, i.e., rider-to-system in micromobility and system-to-human in vehicle takeovers or teleoperation, each of these scenarios requires clear, real-time collaboration between humans and increasingly intelligent automated systems to maximize safety and user experience [20], [21].

With advances in AI, automated systems are becoming more capable, allowing mobility technologies to increasingly assist with, or even manage, the driving task on the road. Such systems represent a form of embodied AI, which can interpret environmental and human inputs (e.g., via advanced sensors and driver biometrics) and deliver real-time, context-aware assistance [22], [23]. The challenges are now more centered around how to integrate AI in real-time without undermining human oversight or, conversely, overwhelming the user. There is a need for role-based collaboration frameworks that detail “who does what” - whether it is a driver in a semi-autonomous car, a remote teleoperator controlling a robotaxi, or a micromobility rider.

2.2 Human-AI Collaboration Paradigms and Frameworks↩︎

There is a growing emphasis on human-AI teaming in human-AI collaboration research (e.g., [24]). In this context, AI is considered a co-worker or partner in driving tasks, rather than an assistive tool [25]. As such, humans and AI work together to combine their capabilities to achieve shared goals. This is different from traditional human-automation interaction, where humans may be out of the loop and only intervene when the automation requires a human takeover [26].

Over the years, various frameworks for classifying and implementing human-automation collaboration in vehicles have been proposed, such as supervisory control, shared control, and layered autonomy. Specifically, for supervisory (or traded) control, a human’s primary role is to monitor and intervene without continuously acting [27]. This aligns with Sheridan’s levels of automation model. Here, one agent, either the human or the AI, is active at a time. In this arrangement, humans delegate tasks to AI and only supervise the system [28]. This setting applies to SAE Level 2-3 automated driving where human drivers may occasionally intervene and take over.

Shared control indicates that humans and automation simultaneously contribute to the control of the vehicle. For example, the driver holds the steering wheel while an AI co-pilot system applies torque to guide lane-keeping [29]. In practice, this shared-control paradigm requires both the driver and AI to be actively involved in perceiving the environment and acting on the vehicle, compensating for the other as needed.

As for layered autonomy, the control architecture is mapped into multiple levels, each with different scopes of authority [27]. For example, in teleoperated driving, remote operators typically perform two main roles: higher-level trip planning and lower-level direct control. In the human-AI collaboration paradigm, AI may manage the path and speed (upper layer) while the human operator handles immediate maneuvers (lower layer), or vice versa. The division of responsibilities may depend on factors such as environmental conditions, AI system reliability, operator cognitive workload, and/or specific safety requirements. Another example is Toyota’s concept of “parallel autonomy” (Toyota Guardian system) [30], [31], in which human drivers manage the direction control, while an autonomous safety system continuously monitors the environment and can intervene and override human input when an immediate danger is detected.

While shared control, supervisory control, and layered autonomy provide useful paradigms for allocating responsibilities, these frameworks are limited in their capacity to comprehensively describe human-AI collaboration roles in dynamic and real-time mobility environments. Moreover, they often lack the specificity to distinguish how particular AI roles engage with human users and adapt to different vehicle types, environmental contexts, and task requirements. As AI becomes more embedded in diverse transportation modes, a more flexible and integrated framework is needed to describe how humans and AI should collaborate moment to moment.

2.3 Existing Automation Taxonomies and AI Standardization↩︎

Current industry standards and guidelines, such as ISO 26262 (Functional Safety for Road Vehicles) [32] and UL 4600 (Standard for Safety for the Evaluation of Autonomous Products) [33], are mainly concerned with functional safety and hazard analysis. They provide frameworks for mitigating risks such as hardware malfunctions, sensor failures, or software bugs. However, these standards may not specify how an AI system should inform a user in partial automation scenarios or how the AI should escalate from gentle alerts to direct interventions (e.g., forced braking). In other words, the complexities of dynamic human-AI collaboration are not yet fully addressed in these standards and guidelines.

A similar limitation can be found in SAE International’s J3016 standard, which defines six levels of driving automation as described in Section II.A. While J3016 clarifies who is primarily responsible for controlling the vehicle, it does not detail the real-time interplay between humans and AI systems [5], [6]. For example, AI systems are expected to support drivers during takeover actions in semi-autonomous vehicles (e.g., Levels 2–3), where human drivers need to resume manual control. However, during these transitions, drivers often experience reduced vigilance and situation awareness, highlighting the need for AI to adapt to the driver’s real-time mental state and surrounding conditions. A current challenge to smooth transitions is the lack of clarity about roles between human drivers and AI, making it uncertain which specific tasks each party should handle. This could be especially problematic in scenarios where automation levels shift within a single routine trip, such as switching from conditional automation (i.e., SAE Level 3) on highways to manual control in urban areas. In contrast to sudden takeovers, these transitions driven by changing driving contexts (e.g., switching from highway automation to urban manual control) require proactive coordination between human and AI systems. This ambiguity can lead to critical issues such as limited understanding of the automation state, miscalibrated trust (e.g., over- or under-reliance), and inconsistent or sub-optimal decision-making.

While more emerging AI standards are currently under development, there is a need to move toward a human-centered AI approach[34] that includes explicit frameworks that go beyond static “levels of control” to define evolving roles within continuous human-AI interactions.

Table 1: Triadic AI Roles in Human-AI Collaboration: Control Scopes, Key Functions, and Adaptive Behaviors Across Mobility Domains
Dimension Advisor Co-Pilot Guardian
Control Scopes
- Monitors driving conditions
- Offers purely informational alerts,
prompts, and suggestions
- Escalates alerts if the driver is
inattentive or if hazards become critical
- Dynamically adjusts steering, braking,
or speed when necessary
- Allows driver override at any time
- Provides active assistance
(steering nudges, speed modulation)
to maintain stability and safety
- Assumes full control
in life-threatening situations
- Overrides erroneous driver
input if it poses severe danger
- Typically operates quietly in the
background
Key Functions
changes, environmental conditions)
- Context-aware guidance (traffic, weather,
upcoming maneuvers)
- Driver state monitoring (physiological
monitoring) for soft reminders (fatigue,
break suggestions)
merges, brakes)
- Partial automation (lane-centering,
adaptive cruise, speed assist)
- Driver state monitoring (eye tracking,
grip strength) for real-time assistance
collision risk \(>\) \(95\%\))
- Immediate intervention (evasive steering,
emergency braking)
- Stabilization (correct skids, hydroplaning)
Features
state (distraction) or environment severity
- Scales back non-critical notifications to
avoid alert fatigue
inattentive, the environment is complex
(dense traffic, inclement weather),
or user input is insufficient
- Reduces involvement when the driver
re-engages, passing control back smoothly
- Provides post-intervention transparency
by explaining actions afterward
- Deactivates once the critical event has
passed
(Automotive)
- “Sharp curve in \(600\) feet”
- “We’ve been driving two hours;
want to find a rest stop?”
- “You seem distracted; I can handle speed
until you’re ready.”
- Automatically slows if the driver misses
a stop sign.
collision avoided - you’re safe now.”
- “Steering correction applied:
roads are icy.”
- (Post-crisis) “I took control because a
collision was imminent. Are you alright?”
(Micromobility:
e-Scooters)
crosswalk.”
- “Battery low in \(2\) miles - consider a
charging stop.”
narrow lanes - shall I?”
- “I’ll limit speed
slightly to ensure control.”
emergency brake if a sudden obstacle
appears and the rider fails to react.
- To riders: “Braked to prevent a collision.”
(Teleoperation)
real-time sensor data to the teleoperator:
“Vehicle behind you is approaching fast;
consider a lane change soon.”
if the remote driver is momentarily
overwhelmed. “I’m providing a gentle
steering assistance - confirm to proceed.”
a crash is imminent, it triggers an
emergency maneuver (e.g., remote override
of throttle/brakes).
- (Post-event) “Override executed due to
collision risk - vehicle is stable.”
Emotional
Support
feeling?”)
- Encourages breaks if fatigue is detected
- Reduces isolation: chatty voice-based
interaction if driver prefers
next few minutes while you calm down.”)
- Lighthearted banter or reassurance to
de-stress driver
- Encourages safer decisions (seatbelt
checks, speed moderation)
- Post-event reassurance (“You’re safe
now; do you want to pull over and rest?”)
- Explains why override happened
(reduces confusion or distrust)
Cases (SAE
Levels)
reminders, hazard beeps)
- Level 2-3: More robust situation
prompts (traffic, route info, takeover
requests)
- Level 4-5: Emphasizes user preferences
& comfort info
cruise, partial maneuvers during driver
inattention
- Level 4-5: Optional co-control based on
user preferences (e.g., scenic routing,
speed adjustment, or allowing the user to
“practice” driving)
might only intervene when a collision
is imminent
- Level 4-5: Full override in rare system
failures (e.g., sensor malfunction,
user incapacitation)

3 The Triadic Framework for Human–AI Collaboration↩︎

To fill the identified gap, this study developed a triadic framework of AI roles that can switch fluidly in real time, aligned with the dynamic and adaptive nature of human-AI collaboration in complex transportation environments. We conducted a broad literature review of frameworks and role classifications in human-centered AI, human-AI teaming, human-automation interaction, and human-centered embodied AI across a wide range of application domains, as well as classical human factors models such as Sheridan’s Levels of Automation and Endsley’s Situation Awareness framework [35], [36]. From the review, we classified key functions that AI may offer in a real-time driving environment and proposed a triadic human-AI collaboration framework with three distinct and non-hierarchical AI roles in terms of collaboration boundaries and communication strategies: Advisor, Co-Pilot, and Guardian AI.

3.0.0.1 Advisor AI

This AI role provides the human user (driver, rider, or remote operator) with continuous informational support, such as context-aware alerts or informational suggestions, without directly intervening in vehicle control. This AI role can prevent driver complacency or confusion by providing timely updates, keeping the drivers informed and engaged.

3.0.0.2 Co-Pilot AI

Co-Pilot AI shares partial control with the human user by offering dynamic action support that adapts to human action, such as adjusting steering, braking, or speed when user input is insufficient, attention is compromised, or external conditions become complex. This AI role reduces driver workload without removing the driver from the control loop.

3.0.0.3 Guardian AI

This AI role serves as a continuous safety governor, autonomously taking full control only to prevent collisions or perform emergency maneuvers. It can even override driver inputs when erroneous actions and/or severe danger are predicted.

These three AI roles are not locked to a fixed setting; instead, they transition dynamically according to user state and situational demands. For example, if repeated alerts from Advisor AI are ignored, or if the system detects driver distraction/fatigue, it may shift to Co-Pilot AI. Similarly, in scenarios where partial assistance cannot adequately prioritize safety, Guardian AI can step in with a full override. Guardian AI can revert to either Advisor or Co-Pilot mode as well, based on user needs and environmental context. For instance, experienced drivers may prefer to remain highly engaged, or e-scooter riders may prefer manual riding for entertainment purposes [37], [38]. Table 1 consolidates and extends the earlier preprint version of this framework by detailing the three AI roles (Advisor, Co-Pilot, and Guardian) in human-AI collaboration, including their control scope, key functions, adaptive features, and practical applications across diverse mobility domains including automated vehicles, micromobility devices, and teleoperation scenarios.

4 Integrating Real-Time Data from Humans and the Environment↩︎

Figure 1: System architecture illustrating the scenario of a human interacting with an intelligent mobility system.

Fig. 1 illustrated the system architecture of our proposed triadic framework, highlighting how real-time human and environmental data is used to inform adaptive role transitions among AI agents for human-AI collaboration. As illustrated in Fig. 1, the framework begins with human-centered sensing systems, which include cameras (e.g., RGB-D cameras) and sensors (e.g., eye tracking devices), to continuously acquire multi-modal data streams from both the user and the surrounding environment. These data streams are then processed by specialized multi-modal AI models. For example, image processing models are applied to visual data captured from eye-tracking systems and RGB-D cameras, enabling human gaze analysis and environmental perception. Graph neural networks are employed to model skeletal structures derived from wearable sensor data, facilitating the interpretation of human postures and motion dynamics. Furthermore, advanced temporal modeling techniques are utilized to process high-dimensional time-series data in order to capture underlying physiological patterns and temporal dependencies. Building upon these multi-modal AI models, the triadic framework is implemented via AI agents, each representing a distinct collaboration role: the Advisor AI, the Co-Pilot AI, and the Guardian AI, which interpret and respond to real-time user and environment context. Through this three-role structure, we may further design intelligent human-machine interfaces incorporating multiple sensory channels (e.g., visual, auditory, and tactile modalities) to convey information from AI agents. These interfaces facilitate intuitive and adaptive interactions between the user and each AI role, particularly within complex and dynamic transportation environments. The details are described as follows.

4.1 Human-Centered Sensing Systems↩︎

The transition of AI’s role, as discussed in Session 3, depends on the user’s state, which can be inferred through multiple layers of human physical and physiological data, each requiring different types of sensors for accurate detection. At a fundamental level, macro-level indicators such as posture and gestures provide interpretable human behavioral data [39], [40]. These can be captured using vision-based devices, including RGB cameras, depth sensors, and wearable motion trackers, which analyze body movements and seating positions [41] to infer users’ physical intention and engagement. On a more subtle level, gaze patterns and facial expressions are associated with users’ cognitive states [42], [43]. These cues, though less overt than body movements, can be detected using eye-tracking and other vision-based devices [44]. These sensors monitor human attention expressions, determining the user’s level of focus and potential distraction. At a deeper physiological level, “under-the-skin” signals such as heart rate variability, skin conductance, and other biometric markers offer assessments of mental states. These indicators can be measured using wearable sensors (e.g., smartwatches, chest straps, or capacitive seat sensors) to detect stress, fatigue, drowsiness, heightened alertness, and other mental states [45], [46].

In this triadic framework, the AI roles can shift based on user state and environment context. Similarly, human’s operational role, whether as a driver, rider, or remote operator, influences their expectations of AI involvement and the type of data needed for effective collaboration.

For drivers, whether assisted by automation or not, humans are still responsible for vehicle control according to current regulations. AI, in this case, can cycle among all three AI roles depending on need. It draws on macro-level behavioral data (e.g., gestures for commands [39]), micro-level cognitive data (e.g., gaze, blink rate for focus [47]), and physiological data (e.g., heart rate variability for fatigue detection [48]) to tailor feedback, issue timely warnings, share partial steering control, or execute a full override of braking. This multi-role adaptability provides safe and effective driving support.

For riders, who delegate driving responsibilities to either an AI system or another human, AI’s role centers on safety and comfort. To provide a smooth and reassuring ride, AI needs to actively communicate with riders with status updates, helping them maintain situation awareness despite not being in control. The AI should dynamically adjust the vehicle’s behavior based on riders’ preferences. This is achieved by leveraging engagement and preference data, such as facial expressions, physiological signals (e.g., heart rate, skin conductance), and verbal and non-verbal cues [49]. Through continuous interpretation of these signals, the AI can tailor its actions and interactions to foster trust and enhance the overall riding experience.

For remote operators who may oversee multiple AI-driven vehicles, AI plays a key role in managing information flow and decision support. Unlike drivers, remote operators rely on real-time data streams of vehicle telemetry (e.g., speed, system status) and sensor feeds (e.g., camera, LiDAR data) from the environments, as well as network metrics such as latency or bandwidth that affect command timing [50]. Using behavioral (e.g., body orientation for vehicle focus [51]), cognitive (e.g., pupil dilation for cognitive load [52]), and physiological indicators (e.g., galvanic skin response for alertness [53]), the AI filters critical alerts, prioritizes tasks, and suggests or executes interventions for synchronizing both human and system actions.

Figure 2: Application of three AI roles in critical micromobility scenarios.

4.2 Multi-Modal AI for Heterogeneous Data Processing↩︎

Human-AI collaboration can be further enhanced by an integrated framework that acquires and fuses heterogeneous data streams from both human-centered sensing and environmental sensing systems. While Section 4.1 focused on human-centered data (e.g., physiological signals), environmental sensors such as radar and LiDAR provide complementary information that further enhances the AI’s understanding of the surrounding content. Radar can be used to detect object velocities and distances, providing crucial information about moving obstacles [54][56]. LiDAR, on the other hand, offers high-resolution three-dimensional maps of the surrounding environment for precise detection of objects, road boundaries, and free space [55]. Combining these environmental sensors with human-centered data allows for a comprehensive understanding of both the user and the user’s surrounding environment.

To effectively process the diverse data formats collected from human-centered and environmental sensors, several advanced AI models can be leveraged. For image data, cutting-edge models such as Vision Transformers (ViTs) [57], Convolutional Neural Network (CNN) [58], and ResNet [59] are used to capture spatial and semantic features of the environment. In parallel, depth sensor data representing human skeletal motion are transformed into graph structures, where joints are represented as nodes, and their anatomical connections as edges. Graph neural networks (GNNs) are employed to capture these spatial and temporal relationships that include complex human postures, gestures, and movement sequences [60], [61]. Continuous time-series signals collected from wearable sensors require sophisticated signal processing approaches. Temporal modeling techniques, including Transformer [62], temporal convolutional networks (TCNs) [63], and advanced time-series decomposition approaches [64], [65], can be applied to capture the underlying physiological rhythms and anomalies. Furthermore, radar and LiDAR data can be processed effectively using models such as PointNet [66] or VoxelNet [67], which are well-suited for handling three-dimensional point clouds. Collectively, these multi-modal AI models support the robust integration of human-centered and environmental data and establish the technical foundation for the triadic AI roles.

4.3 Triadic Framework for Real-Time Collaboration↩︎

Building on the foundational multi-modal AI models, we implement the triadic framework to function in complex transportation environments. As stated in Section 3, these role-specific AI agents vary in autonomy and user engagement yet draw on the same integrated data streams to determine when and how to act.

The Advisor AI agent uses the underlying AI models to provide data-driven insights, alerts, and contextual observations. This agent serves as an analytical companion to provide timely information and suggestions while leaving decision-making authority and final control with the user.

The Co-Pilot AI agent represents a shared-control partner that assists in decision-making and low-level actuation. The Co-Pilot AI continuously engages with the user for optimal actions, highlighting potential risks and adapting steering, braking, or speed in real-time based on both environmental changes and user feedback. This layer provides a human-AI partnership in which the agent reduces human cognitive and physical workload without removing the human from the control loop.

The Guardian AI agent acts as a safety governor, guided by predictive modeling and learned behavioral patterns. When multi-modal data indicate an immediate hazard, the Guardian AI agent actively initiates or directs actions on behalf of the user, with minimal need for human input.

The integration of multi-modal data into the triadic AI framework enables various levels of support, from passive advisory functions to fully autonomous guidance. Furthermore, given that each role may change in real time, it is essential to design intelligent human-machine interfaces (HMIs) that effectively communicate these transitions so that users can understand AI behavior and enhance appropriate levels of trust and user acceptance.

Figure 3: The example of the triadic AI framework in eight HMIs (e.g., Advisor AI in the auditory format, Co-Pilot AI in the visual format, and Guardian AI in the tactile format).

5 Application Scenarios and Case Studies↩︎

In this section, we present an online study focusing on micromobility (i.e., e-scooters) as a preliminary validation of the Triadic Human-AI Collaboration Framework[68]. While integrating AI-driven systems is a potential solution [18], [19] to mitigate safety issues caused by the increasing usage of e-scooters [69], the precise role of AI in micromobility has not yet been systematically investigated. To address this gap, we conducted an online survey [68] to compare the effect of three AI roles (i.e., Advisor, Co-Pilot, and Guardian, Fig. 2) on user preference. An online approach allowed a large and demographically varied sample to be obtained in a relatively short period, which can provide early feasibility data before subsequent in-lab studies for objective evaluation.

The three AI agents were presented through human-machine interfaces (HMIs) in three modalities, visual, auditory, and tactile, similar to previous driving studies (e.g., [12], [70]). To increase the variations of the HMIs, each visual, auditory, and tactile HMI was further developed into multiple types (Fig. 3), including three visual types: AR glasses, control panel display, and road projection; two auditory types: informative (e.g., “Exit ahead”) and conversational (e.g., “We are entering a new road.”) voice assistance; and three tactile types: handlebar, footpad, and helmet.

The survey was conducted on a crowdsourcing platform, Prolific, to compare user preferences for the three AI agents and three modalities (nine role-by-modality combinations, see Fig. 3) illustrated with eight HMI types. A total of \(473\) valid responses (mean age = \(46.29\)) were collected. To systematically evaluate user preference, the questionnaire, which included usefulness and satisfaction ratings, was adopted [71]. The results revealed that no statistically significant differences were found among the three AI roles, regardless of usefulness or satisfaction ratings. In terms of the HMI modality, auditory stimuli offered greater usefulness and satisfaction compared to visual and tactile options. Within the three visual HMIs, the control panel was the least useful compared to AR glasses and road projection. However, regarding satisfaction, AR glasses were the least preferred. The findings for auditory HMIs demonstrate that informative assistance was more useful and satisfying than conversational assistance. Among the comparisons of the three tactile HMIs, results suggest that incorporating displays on handlebars was associated with higher usefulness and satisfaction than footpads or helmets.

Overall, the findings of this study provide early evaluations of the Triadic Human-AI Collaboration Framework for user preferences, which could inform the design of follow-up studies on 1) user preferences for other transportation modes, such as automated vehicles or teleoperated vehicles, and 2) objective validation of the framework in controlled experiments or real-world trials; and (3) development of regulatory and standards frameworks that address human-AI collaboration across vehicle types, including personal, shared, and commercial mobility systems.

6 Discussion and Future Work↩︎

6.1 Conceptual Contributions↩︎

This study proposes a triadic human-AI collaboration framework (Advisor, Co-Pilot, and Guardian) to fill gaps in the literature and specify how responsibility shifts among these AI roles across transportation domains. The level of AI involvement should be determined by real-time sensing of both the human state and environmental conditions. One core contribution of this study is to define AI roles that map directly onto the diverse tasks required in intelligent mobility. While earlier frameworks (e.g., SAE Levels of Automation) emphasize who is in control, this triadic approach clarifies how control may flow among distinct AI roles, reducing role ambiguity between AI and drivers, riders, or remote operators.

6.2 Implications for Industry and Standards↩︎

The role-based framework proposed in this study has important implications for ongoing and future AI standardization efforts. For example, industry guidelines (e.g., ISO 26262 and UL 4600, or new AI guidelines that become available in the near future) could incorporate “AI role protocols” to specify the boundaries for triggering the movement of the system from Advisor to Co-Pilot mode, or how Guardian overrides need to be documented and communicated to the users. Standardized role definitions may facilitate audits, certification processes, and legal or liability considerations. Explicit definitions of how AI roles respond to real-time sensing can also facilitate interoperability across automakers, micromobility platforms, and teleoperated services while helping users consistently interpret system actions and safety overrides. Additionally, the framework highlights the need to address potential ethical and legal challenges, such as determining liability in cases where a Guardian AI override leads to unintended consequences, as well as ensuring interoperability with existing regulatory and technical standards.

6.3 Communication Strategies: What/Action-Focused (Directive) vs. Why-Focused (Explanatory)↩︎

An additional dimension that may complement the triadic framework proposed in this study is the AI communication strategy, that is, how AI presents information to humans, either through direct alerts or commands where AI tells the human what to do (or what the AI is doing) without additional context. For example, the car may issue a sharp warning: “Brake now!” when it detects an immediate hazard. On the other hand, explanatory AI may prioritize communicating “why” through a brief justification or context for the request, such as “Take over, construction zone ahead.” Direct alerts have the advantage of being quick and unambiguous, while explanatory communication may increase the transparency of AI systems. How the two communication strategies interact with the three AI roles under different circumstances (e.g., mild hazard vs. immediate threat) requires further research.

6.4 Adaptive HMIs as the Interface for AI Roles↩︎

Because each of the three AI roles involves different levels of interaction, adaptive human-machine interfaces (HMIs) play a critical role in conveying which AI agent is active at any given moment. For example, Advisor AI may rely on minimal auditory alerts, while Co-Pilot might highlight shared steering controls in the user interface (UI). Similarly, with different AI communication strategies, either What/Action-Focused (Directive) or Why-Focused (Explanatory), the modality used to convey HMIs may vary. The examples shown in Table I are in a conversational format, which primarily relies on the auditory channel. In real-life time-critical events, users may need multimodal HMIs (i.e., combined visual, auditory, and/or tactile feedback) to quickly perceive information from AI. Finally, for people with different perceptual or cognitive abilities, HMIs should also be adaptive to effectively communicate with human users. For example, older adults with general sensory decline may experience fluctuating day-to-day cognitive functioning. In this case, for effective interaction, AI may need an adaptive HMI that adjust its parameters (e.g., intensity and duration) and modalities in response to users’ real-time states and environmental conditions.

6.5 Broader Application and Empirical Validation↩︎

Although this study primarily addresses automated vehicles, micromobility, and vehicle teleoperation, the triadic framework may be extended to other high-stakes domains. For example, healthcare robotics could use Advisor AI to suggest adjustments or flag potential issues, Co-Pilot AI to share partial control with a medical practitioner, and Guardian AI to intervene when a critical error is detected. Further work is needed to test the framework’s feasibility across these applications. Even within the surface transportation domain, although conceptually robust, the triadic framework has yet to be validated for real-world effectiveness through simulations, in-lab controlled experiments, or field studies, particularly in relation to critical factors such as safety, trust, and operators’ workload. Increasing the complexity of testing conditions may yield more realistic and generalizable insights, especially when combined with measurable indicators such as reaction time, level of human intervention, or physiological responses in human-AI teaming scenarios. Building on this foundation, future research may examine the effects of the three AI roles on driving, riding, or operating performance; trust in automation and AI; perceived usefulness and satisfaction across diverse participant groups, including novices, expert drivers, older adults, or individuals with disabilities.

7 Conclusion↩︎

This study introduces a Triadic Human-AI Collaboration Framework, featuring three distinct AI roles—Advisor, Co-Pilot, and Guardian, to enhance AI standardization and real-time cooperation in transportation systems. The framework establishes a dynamic approach to AI involvement by adapting to human states and environmental conditions. Further empirical validation, through controlled experiments and real-world trials, is necessary to refine its applicability across diverse domains. Overall, this study lays the groundwork for both future AI standardization and human-AI collaboration in AI-driven transportation systems.

References↩︎

[1]
SAE, SAE Levels of Driving Automation - Refined for Clarity and International Audience.” 2021.
[2]
Y. Yang and M. Y. Kim, Promoting Sustainable Transportation: How People Trust and Accept Autonomous Vehicles—Focusing on the Different Levels of Collaboration Between Human Drivers and Artificial Intelligence—An Empirical Study with Partial Least Squares Structural Equation Mod,” Sustainability, vol. 17, no. 1, p. 125, 2025, doi: 10.3390/su17010125.
[3]
L. E. Juanicó, Conceptual Proposal for Enhancing Autonomous Taxi Systems: A Focus on Tesla’s Cybercab,” Conicet (Argentinean National Council of Scientific Researches), 2024.
[4]
F. Tener and J. Lanir, Design of a High-Level Guidance User Interface for Teleoperation of Autonomous Vehicles,” in ACM international conference proceeding series, 2023, pp. 287–290, doi: 10.1145/3581961.3609845.
[5]
L. N. Boyle, N. van Nes, K. Bengler, and J. D. Lee, Workshop to Go Beyond Levels of Automation,” 16th International Conference on Automotive User Interfaces and Interactive Vehicular Applications, AutomotiveUI 2024 - Adjunct Conference Proceedings, pp. 247–248, Sep. 2024, doi: 10.1145/3641308.3677402.
[6]
T. Inagaki and T. B. Sheridan, A critique of the SAE conditional driving automation definition, and analyses of options for improvement,” Cognition, Technology and Work, vol. 21, no. 4, pp. 569–578, Nov. 2019, doi: 10.1007/s10111-018-0471-5.
[7]
J. D. Lee and K. A. See, Trust in automation: Designing for appropriate reliance,” Human Factors, vol. 46, no. 1, pp. 50–80, 2004, doi: 10.1518/hfes.46.1.50_30392.
[8]
G. Huang, Y. H. Hung, R. W. Proctor, and B. J. Pitts, Age is more than just a number: The relationship among age, non-chronological age factors, self-perceived driving abilities, and autonomous vehicle acceptance,” Accident Analysis and Prevention, vol. 178, no. September 2022, p. 106850, Dec. 2022, doi: 10.1016/j.aap.2022.106850.
[9]
T. B. Sheridan and R. Parasuraman, “Human-automation interaction,” Reviews of human factors and ergonomics, vol. 1, no. 1, pp. 89–129, 2005.
[10]
R. Parasuraman, T. B. Sheridan, and C. D. Wickens, “A model for types and levels of human interaction with automation,” IEEE Transactions on systems, man, and cybernetics-Part A: Systems and Humans, vol. 30, no. 3, pp. 286–297, 2000.
[11]
A. D. McDonald et al., Toward Computational Simulations of Behavior During Automated Driving Takeovers: A Review of the Empirical and Modeling Literatures,” vol. 61. SAGE Publications Inc., pp. 642–688, Jun. 2019, doi: 10.1177/0018720819829572.
[12]
G. Huang and B. J. Pitts, Takeover requests for automated driving: The effects of signal direction, lead time, and modality on takeover performance,” Accident Analysis and Prevention, vol. 165, no. October 2021, p. 106534, Feb. 2022, doi: 10.1016/j.aap.2021.106534.
[13]
C. Kettwich, A. Schrank, H. Avsar, and M. Oehl, A Helping Human Hand: Relevant Scenarios for the Remote Operation of Highly Automated Vehicles in Public Transport,” Applied Sciences (Switzerland), vol. 12, no. 9, p. 4350, 2022, doi: 10.3390/app12094350.
[14]
D. Meroux, A. Broaddus, C. Telenko, and H. Wen Chan, “How should vehicle miles traveled displaced by e-scooter trips be calculated?” Transportation research record, vol. 2677, no. 1, pp. 356–368, 2023.
[15]
L. Gebhardt, S. Ehrenberger, C. Wolf, and R. Cyganski, “Can shared e-scooters reduce CO2 emissions by substituting car trips in germany?” Transportation Research Part D: Transport and Environment, vol. 109, p. 103328, 2022.
[16]
M. Jafarzadehfadaki and V. P. Sisiopiku, “Embracing urban micromobility: A comparative study of e-scooter adoption in washington, DC, miami, and los angeles,” Urban Science, vol. 8, no. 2, p. 71, 2024.
[17]
M. Lee, J. Y. Chow, G. Yoon, and B. Y. He, “Forecasting e-scooter substitution of direct and access trips by mode and distance,” Transportation research part D: transport and environment, vol. 96, p. 102892, 2021.
[18]
E. Comission. D.-G. for Mobility and Transport, White paper on transport: Roadmap to a single european transport area: Towards a competitive and resource-efficient transport system. Publications office of the European Union, 2011.
[19]
M. Gerla, E.-K. Lee, G. Pau, and U. Lee, “Internet of vehicles: From intelligent grid to autonomous cars and vehicular clouds,” in 2014 IEEE world forum on internet of things (WF-IoT), 2014, pp. 241–246.
[20]
M. Fan, X. Yang, T. Yu, Q. V. Liao, and J. Zhao, “Human-ai collaboration for UX evaluation: Effects of explanation and synchronization,” Proceedings of the ACM on human-computer interaction, vol. 6, no. CSCW1, pp. 1–32, 2022.
[21]
J. Rezwana and M. L. Maher, “Understanding user perceptions, collaborative experience and user engagement in different human-AI interaction designs for co-creative systems,” in Proceedings of the 14th conference on creativity and cognition, 2022, pp. 38–48.
[22]
F. De Santis, Autonomous mobile robots: configuration of an automated inspection system,” PhD thesis, Politecnico di Torino, 2023.
[23]
V. Vanniyakulasingam, Autonomous mission configuration on Spot from Boston Dynamics,” PhD thesis, Politecnico di Torino, 2023.
[24]
J. Rezwana and M. L. Maher, “Designing creative AI partners with COFI: A framework for modeling interaction in human-AI co-creative systems,” ACM Transactions on Computer-Human Interaction, vol. 30, no. 5, pp. 1–28, 2023.
[25]
S. Berretta, A. Tausch, G. Ontrup, B. Gilles, C. Peifer, and A. Kluge, “Defining human-AI teaming the human-centered way: A scoping review and network analysis,” Frontiers in Artificial Intelligence, vol. 6, p. 1250725, 2023.
[26]
R. Parasuraman and V. Riley, “Humans and automation: Use, misuse, disuse, abuse,” Human factors, vol. 39, no. 2, pp. 230–253, 1997.
[27]
J. Sarabia, M. Marcano, J. Pérez, A. Zubizarreta, and S. Diaz, “A review of shared control in automated vehicles: System evaluation,” Frontiers in Control Engineering, vol. 3, p. 1058923, 2023.
[28]
J. C. de Winter, S. M. Petermeijer, and D. A. Abbink, “Shared control versus traded control in driving: A debate around automation pitfalls,” Ergonomics, vol. 66, no. 10, pp. 1494–1520, 2023.
[29]
D. A. Abbink, M. Mulder, and E. R. Boer, “Haptic shared control: Smoothly shifting control authority?” Cognition, Technology & Work, vol. 14, pp. 19–28, 2012.
[30]
F. Naser et al., “A parallel autonomy research platform,” in 2017 IEEE intelligent vehicles symposium (IV), 2017, pp. 933–940.
[31]
J. Clifford, Toyota Guardian autonomous driving technology amplifies human car control.” 2019, Accessed: Feb. 15, 2025. [Online]. Available: https://mag.toyota.co.uk/toyota-guardian-autonomous-driving-technology-amplifies-human-car-control/.
[32]
ISO, ISO 26262-9:2018 - Road vehicles Functional safety.” pp. 1–29, 2018, Accessed: Feb. 25, 2025. [Online]. Available: https://www.iso.org/standard/68391.html.
[33]
UL Standards & Engagement, ANSI/UL 4600 Standard for Safety for the Evaluation of Autonomous Products.” 2023, Accessed: Feb. 25, 2025. [Online]. Available: https://www.shopulstandards.com/ProductDetail.aspx?UniqueKey=28341.
[34]
B. Shneiderman, Human-centered AI. Oxford University Press, 2022.
[35]
T. Sheridan, “Human and computer control of undersea teleoperators,” Man-Machine Systems Laboratory Report, 1978.
[36]
M. R. Endsley, “Toward a theory of situation awareness in dynamic systems,” Human factors, vol. 37, no. 1, pp. 32–64, 1995.
[37]
K. Wang, X. Qian, D. T. Fitch, Y. Lee, J. Malik, and G. Circella, “What travel modes do shared e-scooters displace? A review of recent research findings,” Transport Reviews, vol. 43, no. 1, pp. 5–31, 2023.
[38]
Z. Christoforou, A. de Bortoli, C. Gioldasis, and R. Seidowsky, “Who is using e-scooters and how? Evidence from paris,” Transportation research part D: transport and environment, vol. 92, p. 102708, 2021.
[39]
A. S. Mahomed and A. K. Saha, “Driver posture recognition: A review,” IEEE Access, 2024.
[40]
W. A. Espericueta Luna, Y. J. Wu, Y. Luo, B. Hu, and N. Gravina, “Using AI-powered video feedback to improve ergonomics: An analog experiment,” Journal of Organizational Behavior Management, pp. 1–27, 2025.
[41]
P. Meda, A. V. Contreras, W.-H. Lo, G. Huang, and Y. Luo, “Insights for the future of car rental and ridesharing: Driving behavior across different levels of automation,” Mineta Transportation Institute, 2025.
[42]
P. Prasse, D. R. Reich, S. Makowski, T. Scheffer, and L. A. Jäger, “Improving cognitive-state analysis from eye gaze with synthetic eye-movement data,” Computers & Graphics, vol. 119, p. 103901, 2024.
[43]
M. Lohani, B. R. Payne, and D. L. Strayer, “A review of psychophysiological measures to assess cognitive states in real-world driving,” Frontiers in human neuroscience, vol. 13, p. 57, 2019.
[44]
Y. Luo, Y. Chen, and B. Hu, “Multisensory evaluation of human-robot interaction in retail stores-the effect of mobile cobots on individuals’ physical and neurophysiological responses,” in Companion of the 2023 ACM/IEEE international conference on human-robot interaction, 2023, pp. 403–406.
[45]
K. A. Rimes, K. Lievesley, and T. Chalder, “Stress vulnerability in adolescents with chronic fatigue syndrome: Experimental study investigating heart rate variability and skin conductance responses,” Journal of Child Psychology and Psychiatry, vol. 58, no. 7, pp. 851–858, 2017.
[46]
L. T. Smith et al., “Using resting state heart rate variability and skin conductance response to detect depression in adults,” in 2020 42nd annual international conference of the IEEE engineering in medicine & biology society (EMBC), 2020, pp. 5004–5007.
[47]
R. Gavas et al., “Blink rate variability: A marker of sustained attention during a visual task,” in Adjunct proceedings of the 2020 ACM international joint conference on pervasive and ubiquitous computing and proceedings of the 2020 ACM international symposium on wearable computers, 2020, pp. 450–455.
[48]
K. Lu, A. S. Dahlman, J. Karlsson, and S. Candefjord, “Detecting driver fatigue using heart rate variability: A systematic review,” Accident Analysis & Prevention, vol. 178, p. 106830, 2022.
[49]
B. Hu et al., “Exploring the effect of human-drone communication modality on safety and balance control in virtual construction environments,” Ergonomics, pp. 1–14, 2024.
[50]
F. M. Ortiz, M. Sammarco, L. H. M. Costa, and M. Detyniecki, “Vehicle telematics via exteroceptive sensors: A survey,” arXiv preprint arXiv:2008.12632, 2020.
[51]
M. Raza, Z. Chen, S.-U. Rehman, P. Wang, and P. Bao, “Appearance based pedestrians’ head pose and body orientation estimation using deep learning,” Neurocomputing, vol. 272, pp. 647–659, 2018.
[52]
H. Zheng, Y. Luo, B. Hu, and W. C. Giang, “A comparison of workload demands imposed by different types of distracted walking tasks and its effect on gait,” in Proceedings of the human factors and ergonomics society annual meeting, 2020, vol. 64, pp. 1713–1717.
[53]
N. A. Nawawi, R. Sudirman, and U. U. Sheikh, “Drowsiness detection using galvanic skin response and electro-occulograph,” in Journal of physics: Conference series, 2023, vol. 2622, p. 012004.
[54]
S. Yao et al., “WaterScenes: A multi-task 4D radar-camera fusion dataset and benchmarks for autonomous driving on water surfaces,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 11, pp. 16584–16598, 2024.
[55]
M. Nawaz, J. K.-T. Tang, K. Bibi, S. Xiao, H.-P. Ho, and W. Yuan, “Robust cognitive capability in autonomous driving using sensor fusion techniques: A survey,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 5, pp. 3228–3243, May 2024.
[56]
X. Gao, S. Roy, and G. Xing, MIMO-SAR: A hierarchical high-resolution imaging algorithm for mmWave FMCW radar in autonomous driving,” IEEE Transactions on Vehicular Technology, vol. 70, no. 8, pp. 7322–7334, 2021.
[57]
A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020.
[58]
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
[59]
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
[60]
G. Liu, R. Xie, S.-H. Fang, H.-C. Wu, and K. Yan, “Novel human-posture recognition system based on advanced graph convolutional network using skeletal data,” IEEE Journal of Selected Areas in Sensors, vol. 1, pp. 224–236, 2024.
[61]
G. Liu et al., “Automatic human posture recognition using Kinect sensors by advanced graph convolutional network,” in Proceedings of IEEE international symposium on broadband multimedia systems and broadcasting (BMSB), 2022, pp. 01–07.
[62]
A. Vaswani et al., “Attention is all you need,” in Advances in neural information processing systems, 2017, pp. 5998–6008.
[63]
C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks for action segmentation and detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 156–165.
[64]
K. Yan et al., “Novel subject-dependent human-posture recognition approach using tensor regression,” IEEE Sensors Journal, vol. 25, no. 1, pp. 1041–1053, 2025.
[65]
S.-Y. Chang, H.-C. Wu, and G. Liu, “Robust multichannel decorrelation via tensor einstein product,” IEEE Transactions on Signal Processing, vol. 73, pp. 275–291, 2025.
[66]
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 652–660.
[67]
Y. Zhou and O. Tuzel, “Voxelnet: End-to-end learning for point cloud based 3d object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4490–4499.
[68]
W.-H. Lo and G. Huang, “A multi-modal human–AI collaboration framework for e-scooters: Evaluating AI roles in user preference,” in Proceedings of the human factors and ergonomics society annual meeting, 2025 (accepted for publication).
[69]
Q. Ma, H. Yang, A. Mayhue, Y. Sun, Z. Huang, and Y. Ma, “E-scooter safety: The riding risk analysis based on mobile sensing data,” Accident Analysis & Prevention, vol. 151, p. 105954, 2021.
[70]
P. Bazilinskyy, S. M. Petermeijer, V. Petrovych, D. Dodou, and J. C. de Winter, “Take-over requests in highly automated driving: A crowdsourcing survey on auditory, vibrotactile, and visual displays,” Transportation research part F: traffic psychology and behaviour, vol. 56, pp. 82–98, 2018.
[71]
J. D. Van Der Laan, A. Heino, and D. De Waard, “A simple procedure for the assessment of acceptance of advanced transport telematics,” Transportation Research Part C: Emerging Technologies, vol. 5, no. 1, pp. 1–10, 1997.

  1. This work was supported, in part, by the U.S. Department of Transportation (USDOT) University Transportation Center - Mineta Consortium for Equitable, Efficient, and Sustainable Transportation (grant number: 69A3552348328).↩︎