What Types of Human-AI Teams Exist?


1 Introduction↩︎

Human-AI teaming as a field has grown exponentially popular in recent years, and broadly refers to a team consisting of one or more humans and one or more AI working together interdependently towards a shared goal [1]. To do so, each entity must act as a team member, and possess unique and complementary capabilities. In other words, human-AI teaming aims to capitalise on the complementary strengths of both humans and AI, with the goal of improving overall team performance, across several applications. For example, it has been studied in numerous domains, ranging from safety-critical systems such as healthcare (e.g. [2]), work-based settings (e.g. [3]), and recreational settings such as video games (e.g. [4]).

Given this explosion of interest, there have been recent reviews on the human-AI teaming literature, such as [1]. However, such reviews reveal many human-AI teaming papers focus heavily on the ‘AI’ technical aspect, rather than on the ‘team’ aspect. This represents an important gap in understanding, as teaming is the concept that separates human-AI teaming from other related terminologies, such as human-AI collaboration [5] and AI decision-support tools [6]. Further, knowing who is involved in a team is only one component of what distinguishes and explains a team [7], and in particular overlooks which holistic characteristics explain currently studied human-AI teams.

This leads to two problems. Firstly, it is difficult to extract concrete examples of when an AI is no longer a ‘mere tool’ and instead perceived as a teammate from existing reviews. In turn, the specificity of the term is difficult to explain and use to separate different human and AI contexts. Secondly, and more crucially, it is not clear what types of teams are studied within human-AI teaming from a teaming perspective. In particular, it is unclear what types of tasks are pursued by humans and AI, and in what ways are team members organised to allow for collaborative working. This is important to consider, as the wide range of application domains referenced earlier makes it unclear what these disparate environments have in common, if anything. Overall, it has become unclear what is specific about human-AI teaming, and in turn what specifically has been studied under this umbrella definition. In turn, it is difficult to understand what is cohesive about this line of research, and how effectively insights can be shared across papers.

To resolve these issues in definitional uncertainty, recent reviews have called for more research from a teaming perspective, rather than one that solely focuses on the AI involved [1]. Analysing existing papers through such a teaming perspective, and consequently an interdisciplinary lens, is an initial step in this process. Furthermore, as existing definitions for human-AI teams were originally inspired from psychological theories on teaming (e.g. [8]), a natural starting point for understanding human-AI teams from a teaming perspective is to apply psychological taxonomies of teams to human-AI teams. In doing so, we may better understand and categorise what kinds of teams are studied within human-AI teaming, as well as understand how human-AI teams are unique from all-human teams.

In this paper, we apply psychological taxonomies of teaming to understand studies conducted within human-AI teaming research. A scoping review revealed 53 experimental papers on human-AI teaming, which were categorised based on the taxonomies provided by [7] via a deductive content analysis [9]. Five main types of team were found, each distinct and with unique combinations of team level characteristics. Consequently, whilst there is a large body of work on human-AI teaming, the teams studied are not inherently interchangeable, despite sharing the same overarching definition. In turn, the ability to synthesise and confidently transfer key insights between similar studies is reduced. The findings also reveal differing implicit assumptions between papers on what makes a human-AI team a ‘team,’ which requires further clarification.

There are therefore three main contributions of this work. Firstly, we contribute an initial taxonomy that classifies types of human-AI teams studied in the literature. Secondly, we contribute a discussion on issues arising from the current approach to classifying all studies under the same definition of human-AI teaming, despite this covering a wide and disparate number of teaming types. Finally, we conclude the work with guidance on how to report human-AI teams in future papers to aid in clarity and specificity. The checklist provided aims to inspire reflection on what types of teams we as a field are interested in pursuing, as well as more concrete reasons as to why.

2 Background↩︎

2.1 Human-AI teaming↩︎

Several papers have, as part of their work, offered a definition of human-AI teaming. For example, [10] defines them as “a mixed entity of two or more subjects (i.e. at least one each being human or AI), who interact interdependently and perform shared tasks to achieve the same valued goals.” From a recent literature review and synthesis of literature, a more comprehensive definition has been proposed by [1]. Bolded text emphasises the key concepts underpinning the definition:

“Human-AI teaming is a process between one or more human(s) and one or more (partially) autonomous AI system(s) acting as team members with unique and complementary capabilities, who work interdependently toward a common goal. The team members’ roles are dynamically adapting throughout the collaboration, requiring coordination and mutual communication to meet each other’s and the task’s requirements. For this, a mutual sharing of intents, shared situational awareness and developing shared mental models are necessary, as well as trust within the team.”

As can be seen from these definitions, despite being composed of multiple concepts, they remain abstract and high level in nature, leaving room for a wide range of researcher interpretation. In other words, it is unclear how practically useful such definitions are for a researcher interested in identifying and designing a human-AI team. For example, how would one use ‘acting as team members with unique and complementary capabilities’ to distinguish an AI teammate from a mere AI tool? What does it mean to share a goal with an AI? Is the type of common goal important, and at what level of abstraction should the goal be considered (e.g. task level, team level)? Overall, there is not a clear sense for where the ‘team’ component exists in these definitions, outside of a ‘perception’ that an AI is a teammate. Consequently, how such perceptions would manifest, or be measured reliably at face value, is unclear.

It is therefore unsurprising that literature reviews have highlighted many papers do not have a clear foundation of what the teaming concept refers to, and it was very rare to see connections between psychological literature on teaming and human-AI teaming papers [1]. This is interesting, as human teams are the inspiration behind the definition itself; O’Neill et al., 2022 [11] for example is commonly referenced as a definitional paper on human-AI teaming (despite itself referring to human-autonomy teaming), which links its conception of teams to [8]. In turn, this paper sources its definition directly from psychological literature on teaming [12].

This ambiguity has led to a very wide range of the types of teams being studied, all under the same definition. For example, [13] considers how a human and an AI would work together to allow permitted personnel into a restricted building via the facial recognition. Conversely, [14] considers how two human teammates would work with an AI to perform reconnaissance photography as part of an aerial-based mission. Both studies refer to these setups as human-AI teams, however their similarities outside of the presence of an AI are limited. The level of role adaption, number of teammates, unique capabilities each possess, and the extent to which an AI is acting as a teammate for example, varies significantly between settings.

Overall, it is difficult to understand what a human-AI team would look like from reading the overarching definitions, and given the wide range of applications studied, looking at any individual studies would likely be similarly confusing. There is therefore a need to understand what exactly is being studied and identify types of teams studied, from a teaming perspective. Indeed, [1] explicitly calls in their discussion for a “holistic approach involving multiple disciplines,” and for research to bring “the teaming idea, and established theories and empirical research from human-human teaming, into the field.” Consequently, in this work we apply existing psychological taxonomies of teams to human-AI teams, to better understand what types of human-AI teams are currently studied.

2.2 Teaming taxonomies↩︎

Within the psychological literature on teaming there have been several taxonomies over the years, creating a large yet overlapping field to draw insights from. To address this, [7] provided a synthesis of the main taxonomies into a combined model, which accounts for both task level (i.e. what actions the team performs) and team level characteristics (i.e. the overall composition of the team). We use the synthesis taxonomies provided by [7] to identify different types of teams studied within human-AI teaming, for two further reasons. Firstly, the team level characteristics were designed to be discrete categories, allowing for easy clustering of differences between papers. Secondly, the taxonomies were created by the same authors that inspired the original human-AI teaming definition (i.e. human autonomy teaming, where [8] references [12]). As such, they offer an applicable lens to view human-AI teams, with discrete categories allowing for easy creation of teaming types. The taxonomies are overviewed here, and linked to previous human-AI teaming literature where relevant.

2.2.1 Task level characteristics↩︎

Task level characteristics are the types of work a team member can engage in, and were specifically designed in this taxonomy to be mutually exclusive and exhaustive categories. As explained by [7] (Pages 107-111), the following task level characteristics can be performed by members of a team:

  • Managing others: Directing, supervising, or overseeing the work of others in an authoritative role. Intends to encourage productivity amongst subordinates. Does not include managing processes, which are a type of problem-solving

  • Advising others: Providing consultative professional support (e.g. expert assistance or advice) where the advisor lacks authority over the advisee

  • Human Service: Providing a good or service to another party via social interaction. Intends to bring satisfaction to a client/customer

  • Negotiation: Two or more parties in conflict seeking to resolve differences and reach agreement via social interaction. Does not include collaborative efforts to reach a common objective/goal, which is a type of problem-solving

  • Psychomotor action: Technical and/or motor functioning requiring psychological processing to perform calculated or elaborate movements (e.g. manipulation/use of a product/machine, or a task achieved by engaging in psychomotor action)

  • Defined problem-solving: Problem solving tasks with predetermined or conclusive solutions or correct answers (e.g. yes or no, option A, B, or C). Involves choosing between two or more options rather than generating a new, unique solution

  • Ill-defined problem-solving: Problem solving tasks lacking predetermined or conclusive solutions or correct answers (e.g. planning, knowledge generation)

Each task type can be completed by one or more members of a team, and a team member may engage in multiple task types within a given teaming situation. For example, a team mate may be in charge of classifying an object (defined problem-solving), as well as providing a suggestion for an action based on this classification to another teammate who is in charge of the task (advising others). Therefore, it is possible for two teams to have the same types of tasks as part of the work, whilst on the surface appearing very different. For example, working with an AI teammate to score goals in a video game has the same task types as working with an AI to pilot an aircraft (defined problem-solving, ill-defined problem-solving, and psychomotor action). The inverse is also true; two teams may appear similar in setup, but in fact require very different tasks. For example, as part of disaster response simulations it is common to navigate difficult terrain to locate and rescue victims. However, depending on the team setup this will affect the specific tasks involved. For example, some rescue operations will involve team members to physically rescue victims (i.e. psychomotor action), however some teams may be purely strategic in allocation of rescue attempts, thereby being exclusive to defined and ill-defined problem-solving.

Consequently, whilst task level analysis is useful to describe and understand what types of work are engaged in and by whom, they are not useful for overall team classifications. Instead, team-level characteristics can be considered, explained in the following subsection.

2.2.2 Team-level characteristics↩︎

Team-level characteristics present a holistic understanding of a team, used to describe a team at a specific moment in time, as the makeup and structure may naturally shift or change during its life. As explained by [7] (Pages 115-119), teams can be categorised using the following variables, each with their own discrete categories.

Task interdependence is how much the outcomes of team members are influenced by/depend on the actions of others. It has been noted as an important factor in Human-AI Teaming (e.g. [15], [16]), and one that separates it from simple tool use (e.g. [10], [17]). There are four types: Pooled, where each team member contributes to the outcome without interacting with other group members; Sequential, where one team member must act before another can act; Reciprocal, a one-on-one format of back-and-forth style interaction between team members (but not multiple members at once); and intensive, where all team members interact as a unit to jointly collaborate.

Role structure is the extent to which roles are fundamentally different/not interchangeable versus each person is capable of performing every component. This is not referring to team members who simply perform different roles; the key difference is whether all team members could perform all roles. This may not always be true, such as in the case of a role requiring specific skills or knowledge, such as a software engineer in a team of researchers. In the case of human-AI teaming, humans are typically trained in the task at hand, and therefore have expertise/specialisation an AI does not have access to. There are consequently two types: Functional, where team members perform fundamentally different roles due to their level of expertise/specialisation; and divisional, where team members perform a specific part of the overall task, but are capable of performing each part.

Leadership structure refers to the pattern/distribution of leadership functions, such as choosing a direction and aligning goals among team members. There are four types: External manager, where leadership roles are performed by someone outside of the team; Designated, where leadership roles are performed by one member of the team, across time and tasks; Temporary, where leadership roles are rotated across members of the team, but not held at the same time; and distributed, where leadership roles are performed by multiple members of the team simultaneously.

Communication structure is the pattern/flow of communication and information sharing among team members, both verbal and non-verbal. It has also been noted as an important factor in human-AI teaming (e.g. [18], [19]), as how humans and AI communicate can impact the effectiveness of the team’s performance. There are three types: Hub-and-wheel, where communication passes through a central team member, often but not always the team leader, before dissemination to all other team members; Chain, where communication passes ‘up and down’ based on a hierarchical structure, such as rank or leadership position; and star, where communication passes freely to and from all team members with no central point of contact or hierarchical structure.

Physical distribution is the spatial location of team members in relation to one another. Whilst recent attention has been on the increased use of virtual teams aided by technology, this is separate than AI performing within the team itself. There are three types: Colocated, where team members are close enough to easily communicate face-to-face; Distributed, where team members are far enough that most communication is computer-mediated (e.g. e-mail or video); and mixed, where some team members are colocated and some are distributed.

Finally, team life span is the length of time a team exists as a functional, active unit. This is not the same as the time taken to perform a complete ‘cycle’ of the main task of the team, but rather how long the team is functional. As such, how often the team meets and who is in the team may change over time, but this does not affect the team lifespan itself. Life span is one way that human-AI teaming is different than simple AI tool use, as teams are typically considered to be intentionally formed (e.g. [20], [21]) and exist for more than one interaction (e.g. [22], [23]). There are two types: Ad hoc, where teams brought together to address a specific event before disbanding (e.g. emergency response team); and long term, where teams exist for an extended period of time (e.g. a management team that exists for as long as its organisation exists).

The characteristics can be considered in isolation, or combined to create a holistic description of the team. For example, a team could have intensive task interdependence, divisional roles, designated leadership structure, hub-and-wheel communication, mixed physical distribution, and an ad hoc lifespan. This may describe a team of volunteers and medics assembled to rescue victims from a natural disaster, where there is a central communicator coordinating different members.

Overall, the team level characteristics outlined here are designed to explain how teams operate as a whole, rather than what team members perform together and individually. This holistic account better highlights what is unique about a given team, and is consequently useful for considering how Human-AI Teams behave a unit. However, given the nature of AI and its increasing role in teaming environments, it is also useful to understand what types of work humans and AI are being asked to perform. Therefore, in this paper, we apply the above taxonomies to experimental studies of Human-AI Teaming, to reveal how teaming is currently conceptualised.

3 Methods↩︎

We performed a scoping review of experimental studies on Human-AI Teaming, in order to classify current work into the taxonomies by [7] described above.

3.1 Paper screening↩︎

A scoping review was conducted to collect experimental studies on Human-AI Teaming, as these reviews are useful for clarifying concepts and examining how research is being conducted in a field [24]. For reporting the methodology and results, we followed the PRISMA extension guidance for scoping reviews [25].

The steps taken to identify papers eligible for inclusion are shown in Figure 1.

Figure 1: Paper screening process following the PRISMA method.

To find the initial papers, the ACM Digital Library, IEEE, and Web of Science were searched. Eligible papers were those dating anytime before April 2025 and written in English. The search term [“human ai team*”] was used, and papers specifically discussing Human-AI teaming were collected, and screened for those that describe experimental studies.

The screening process was conducted by the first author, after deliberation with the co-author to create the inclusion/exclusion criteria. The initial search resulted in 204 papers from ACM, 25 from IEEE, and 323 from Web of Science, for a total of 552 papers. After removing conference proceeding summary documents, citations, books, and duplicates, this was reduced to 366. The titles of the papers and any keywords were then screened to find those specific to Human-AI Teaming (i.e. containing the phrase ‘human-AI team*’). Papers that did not use human-AI teaming in the title or keywords were excluded, leaving 288 papers. Papers unrelated to Human-AI Teaming were removed, as well as papers on a similar yet distinct topic (e.g. human-AI collaboration), leaving 261 papers.

These papers were then read by the first author to assess if they studied Human-AI Teaming. Five papers only used the phrase in the keyword and nowhere else in the text, and four papers were inaccessible, and so were removed, leaving 252 papers. Many papers used the phrase human-AI teaming to refer only to the performance of a system that included humans and AI, and provided no further information about the setup. These papers were consequently removed, leaving 141 papers. Finally, papers that were specific to building AI systems were removed, leaving 119 papers.

Of these papers, 57 involved an experimental design. However, four involved no AI in the experimental design (Papers 146, 151, 166 and 176), and one paper involved two studies in the same paper (Paper [17]). Therefore, the analysed dataset included 53 papers, and led to the analysis of 54 studies.

3.2 Data Analysis↩︎

Once papers were collected, the next step was to broadly map the relevant literature by applying qualitative analysis [26]. As we were interested in going beyond a narrative description of experimental studies, a deductive content analysis was selected for paper analysis [9]. This means applying an existing codebook to a dataset. We used the task-level and team-level characteristic taxonomies presented by [7] (described in Section 2.2 previously) to code the types of experimental studies present in the field.

To do so, information including the number of humans and AI present, the tasks and roles assigned to humans and AI, and the team goal, were extracted. In other words, information about what humans and AI were asked to do, in pursuit of what goal, within the human-AI team, was extracted for each paper. Further, to aid in describing the studies, descriptives such as intended application domain and experimental environment were also extracted.

As the taxonomies were built for human only teams, it was necessary to make some minor reinterpretations to apply the team-level characteristics to AI. The task-level characteristics mapped well to the types of tasks performed by humans and AI, apart from psychomotor action. As interacting with an AI involves interacting with some form of computer, this could be interpreted as always involving technical functioning to achieve an outcome. This would reduce the usefulness of the category, and so we instead chose to consider psychomotor actions only when a team member required interaction with something other than the computer used to communicate with an AI (e.g. controlling an aircraft, moving an avatar around a map). However, the team-level characteristics of physical distribution and team-life span proved more challenging to apply. As an AI is usually situated within a computer, it was difficult to interpret the physical distribution of the team. Similarly, not all papers specified how long the team should typically last for, as there were many experimental designs. Therefore, physical distribution was interpreted to mean whether the human and AI were in the same room/environment, and team life span was interpreted as whether the task at hand, if conducted in a real world scenario, was likely to be performed once (i.e. ad hoc), or part of a longer term working arrangement (i.e. long term).

Using the task-level and team-level characteristic taxonomies, a deductive content analysis was performed by the first author. This involved assigning any relevant task-level characteristic to the extracted human and AI roles, and one discrete label from each category of the team level characteristics to summarise team being studied. Once complete, clusters of types of human teams were created by using the team-level characteristics and isolating those with the same configurations. Team-level rather than task-level characteristics were used as the team level contained mutually exclusive, discrete labels, allowing for easy separation of studies. However, task level characteristics of humans and AI in each cluster were extracted in order to aid in describing each cluster, along with the typical types of task, application domains, and number of human and AI members in the team. Further, as the categories of physical distribution and team life span proved difficult to interpret in the context of human-AI teaming, and were reinterpreted from their original meaning, these were not used to cluster the studies. However, in clusters where there was a high percentage of a certain type of these two categories, it is mentioned for purposes of description only.

In total, seven clusters with at least two studies were identified, of which five had at least three studies, and three had at least six. Consequently, eight studies did not form a cluster, meaning they represented unique combinations of team level characteristics. For brevity, we only report clusters with at least three studies, however the full analysis can be found in the appendix.

4 Results↩︎

4.1 Sampled Papers Overview↩︎

In total, there were 54 experimental studies on Human-AI Teaming in the literature, of which the majority were published in the last two years. The publication rate is shown in Figure 2.

Figure 2: The publication frequency of papers that use the term Human-AI Teaming in the title or keywords. The blue line represents experimental studies, of which 54/57 are analysed in this paper.

An overview of the experimental papers is provided in Table 1. Figures 3, 4 & 5 highlight the trends in application domains, study environments, team makeups, and task types.

Figure 3: The type of application domain (left) and type of experimental environment (right) for the experimental studies.
Figure 4: The type of tasks (left) and type of team goals (right) present in the experimental studies.
Figure 5: The number of humans and AI present in the experimental studies. Note that four papers contained experimental conditions with 1 human 2 AI as well as 2 humans 1 AI; these were split to create an overall number of 58.
Table 1: A summary of the papers analysed, in terms of application domain, team makeup, study environment, team goal, and tasks undertaken.
Paper Application Type Team Makeup Environment Goal Task Task Type
[27] Experimental Game 1 human 1 AI CAJA identify defective vs non-defective boxes classifying objects Object Classification
[28] Cybersecurity/Security 1 human 1 AI Cybersecurity game detect cybersecurity threats detect cybersecurity threats Security
[29] Disaster Response 3 humans 1 AI Minecraft rescue survivors rescue survivors Search and Rescue
[30] Disaster Response 1 human 1 AI Simulation rescue hostages rescue hostages Search and Rescue
[31] Cybersecurity/Security 1 human 1 AI Simulation detect cybersecurity threats detect cybersecurity threats Security
[32] Cybersecurity/Security 1 human 1 AI Simulation detect deepfakes deepfake detection Object Classification
[33] Game 1 human 1 AI Defend the Pass win the game kill all monsters game Teamwork Games
[34] Puzzles 3-4 humans 1 AI Puzzle win the game solve pattern recognition puzzle Teamwork Games
[35] Classification 1 human 1 AI Experiment annotate faces face detection Object Classification
[17] Classification 1 human 1 AI Experiment classify objects appraise house prices Object Classification
[17] Classification 1 human 1 AI Experiment classify objects image classification Object Classification
[36] Experimental Game 1 human 1 AI Experimental games win the game bowling; maze; hide and seek Teamwork Games
[37] Game 1 human 2 AI Guess the Word win the game guess the word game Teamwork Games
[38] Classification 1 human 1 AI Experiment classify objects classifying objects Object Classification
[39] Game 1 human 1 AI Ticket to Ride win the game connect cities game Teamwork Games
[40] Classification 1 human 1 AI Experiment classify sentiment of reviews sentiment analysis Object Classification
[16] Game 1 human 1 AI Rocket League score the most goals score in rocket league Teamwork Games
[4] Game 1 human 1 AI Rocket League score the most goals score in rocket league Teamwork Games
[41] Shopping 1 human 2 AI Simulation fulfil orders collect products in a supermarket Teamwork Games
[42] Game 1 human 1 AI Codenames win the game codenames game Teamwork Games
[43] Aviation/Military/Space 1 human 5 AI Simulation teach drone how to move drone maneovering Aerialbased
[2] Healthcare 4 humans 1 AI Simulation diagnose and provide treatment diagnosing patients Workbased Activity
[44] Aviation/Military/Space 1 human 1 AI Moon landing game safely land the space craft landing a space craft game Space
[45] Disaster Response 1 human 1 AI Simulation rescue victims rescue victims Search and Rescue
[46] Game 1 human 1 AI Overcooked deliver soup cooking game Teamwork Games
[19] Game 1 human 1 AI Arma III collect as many crates as possible Arma III crate collection game Teamwork Games
[47] Puzzles 4 humans 1 AI Experiment score as high as possible multiple choice quiz Quiz
[48] Aviation/Military/Space 1 human 1 AI Simulation perform successful maintenance spaceship maintenance Space
[21] Game 1 human 1 AI Hanabi win the game Hanabi Teamwork Games
[49] Work 1 human 1 AI Overcooked win the game cooking game Teamwork Games
[13] Cybersecurity/Security 1 human 1 AI Simulation allow entry for authorised personnel only facial recognition Object Classification
[50] Education 6-8 humans 1 AI Experiment complete coursework management coursework Workbased Activity
[51] Classification 1 human 1 AI Experiment classify objects classifying objects Object Classification
[52] Aviation/Military/Space 2 humans 1 AI Simulation safely land the plane emergency aircraft Aerialbased
[3] Education 1 human 1 AI Simulation design a quiz design a quiz Workbased Activity
[53] Education 1 human 1 AI Experiment track thesis progression diary for thesis progression Workbased Activity
[54] Puzzles 2-3 humans 1 AI Simulation solve problems problem solving; creativity task Workbased Activity
[55] Puzzles 3-4 humans 1 AI Quiz win the game multiple choice quiz Quiz
[56] Game 1 human 5-7 AI Netrek win the game destroy enemies in Netrek Teamwork Games
[57] Disaster Response 1 human 1 AI Simulation rescue victims rescue victims Search and Rescue
[58] Aviation/Military/Space 2 humans 1 AI Arma III eliminate targets search and destroy Aerialbased
[59] Classification 1 human 1 AI Experiment categorise birds correctly classifying objects Object Classification
[60] Education 2-5 humans 1 AI Coursework build space craft and fly it successfully build spacecraft Space
[14] Aviation/Military/Space 1 human 2 AI; 2 human 1 AI Simulation complete surveillance reconnaissance photography Aerialbased
[61] Education 1 human 1 AI Packet Tracer learn about network connections Learn about computer networking Workbased Activity
[62] Aviation/Military/Space 1 human 2 AI; 2 human 1 AI Simulation complete surveillance reconnaissance photography Aerialbased
[63] Experimental Game 1 human 1 AI CAJA identify defective vs non-defective boxes classifying objects Object Classification
[64] Game 1 human 2 AI; 2 human 1 AI Rocket League win the game score in rocket league Teamwork Games
[65] Education 1 human 1 AI Experiment identify cell lineages cell lineage Workbased Activity
[66] Game 1 human 5-7 AI Netrek win the game destroy enemies in Netrek Teamwork Games
[67] Disaster Response 1 human 2 AI; 2 human 1 AI NeoCITIES deploy rescources for disaster relief deploy resources for disaster relief Search and Rescue
[68] Work 1 human 1 AI Experiment create a new fitness app decide app functionalities Workbased Activity
[69] Classification 1 human 1 AI Experiment classify objects classifying objects Object Classification
[70] Cybersecurity/Security 2 humans 1 AI Simulation correctly identify suspicious luggage airport security Xray Security

There are some strong areas of focus within Human-AI teaming experimental studies. For example, gaming applications were a common domain studies wished to provide insight into (10 papers), alongside classification (8 papers) and aviation/military/space (7 papers). Other common applications include puzzles (such as multiple choice quizzes; 6 papers), education (6 papers), disaster response (such as search and rescue; 5 papers), and cybersecurity/security (5 papers). Experiment/simulation setups were a common way to study Human-AI teaming (29 papers), as well as commercial gaming environments (17 papers) such as Overcooked (2 papers) and Rocket League (3 papers). Experimental games/platforms included bespoke environments made for the purposes of the study (6 papers), and education refers to studies conducted in an educational setting (e.g. coursework; 2 papers).

The types of tasks and team goals similarly followed the trends in application domain; teamwork games (16 papers) typically involved the goal of winning the game (20 papers), and object classification tasks (12 papers) were typically associated with the goal of classifying objects (16 papers). Tasks and goals were not fully aligned in all cases for several reasons; for example, some security tasks were associated to object classification (e.g. facial recognition), and some teamwork games were related to goals such as rescuing victims.

Finally, there was a strong preference for Human-AI teams to involve one human and one AI (34 papers). The next largest combination was two humans with one AI (7 papers), followed by one human with two AI (6 papers) and more than three humans with one AI (6 papers). Interestingly, only one side of the team would vary in number; if there were multiple humans, there would only be one AI, and vice versa.

Overall, the most common way Human-AI teaming has been experimentally studied is one human with one AI, in a gaming application and environment, involving teamwork games where the goal is to win the game. Other common setups included classifying objects within simulated environments.

4.2 Task Level Characteristics↩︎

Within the tasks studied in experimental studies, there were a variety of roles humans and AI were assigned. These are shown in Table 2.

Table 2: An overview of the human and AI roles found within experimental studies of human-AI teams.
Paper human role AI role
[27] accept or reject recommendation; make final decision recommends label; has access to hidden information
[28] create filter rule, or accept recommendation; make final decision recommends filter
[29] adhere or disregard guidance, ask for further guidance; make final decision advise
[30] control or observe drone; search for hostages; rescue hostages; make final decision detect hostages; mark map locations; navigate environment
[31] review information; address detected threat; ask for help from AI; make final decision; discuss evidence with AI log suspicious activity; discuss evidence with human; provide contextual information to human
[32] rate confidence in deepfake; modify answer based on model prediction; make final decision predict deepfake
[33] place teammate destroy enemies
[34] solve puzzle provide clue
[35] accept or reject recommendation; edit recommendation; make final decision recommend box boundary
[17] provide initial prediction; accept or adjust AI prediction; classify images; make final decision provide recommendation
[17] provide initial prediction; accept or adjust AI prediction; classify images; make final decision provide recommendation
[36] provide positive or negative feedback on AI behaviour accept feedback from human
[37] discuss clues; provide clue to AI guess the word from clues given by humans; provide clue to AI; discuss clues
[38] make final decision; consider AI prediction provide initial recommendation
[39] build connections between cities; consider AI’s prediction AI infers opponent’s actions
[40] assign sentiment; consider AI prediction; make final decision provide confidence score; provide recommendation
[16] score goals score goals
[4] score goals score goals
[41] collect items from aisles; deliver items to AI provide order information
[42] generate a clue; consider AI suggestion provide clue suggestions; advise
[43] demonstrate movements to AI; correct AI when necessary follow demonstration
[2] diagnose patient; select treatment; consult AI provide diagnostic suggestions; ventilate patient automatically
[44] control speed; control rotation; follow path; complete n-back test control speed; control rotation; calculate initial path; provide confidence score; ask for help
[45] assess victims; aid victims; replace robot battery; carry victim to safety clear rubble; enter buildings; carry victim to safety
[46] select strategy for agent; adapt to agents strategy; follow plan follow strategy
[19] collect crates collect crates
[47] make initial decision; decide if to invoke AI; make final group decision provide suggestion
[48] accept or reject recommendation recommend procedure
[21] provide hint; discard a card; play a card provide hint; discard a card; play a card
[49] request items; prepare food prepare food; respond to requests
[13] compare face to record; grant or deny entry; accept or reject recommendation provide recommendation
[50] track project progress; generate ideas; summarise sources; draft interviews; draft reports; prepare presentation answer prompts
[51] make final decision; consider AI prediction suggest label; provide confidence score
[52] follow known procedures; consider AI suggestion; land plane at airport suggest alternate airports
[3] review AI suggestions; edit AI suggestions; create questions create question suggestions
[53] define problems; report progress; discuss findings; suggest new problems analyse data; model behaviour
[54] solve puzzle; discuss with others offer suggestions
[55] make initial decision; consider AI decision; make final group decision suggest answer
[56] protect planets; eliminate enemies protect planets; eliminate enemies; capture planets
[57] follow AI guide human to victims
[58] ground; surveillance aerial; move; destroy enemies; update surveillance
[59] assign AI attributes guide creation of attributes
[60] build rocket; consider AI suggestions provide feedback; provide progress
[14] take photos; control aircraft; provide target airspeed and altitude; create flight plan; provide waypoint information; confirm photos create flight plan; provide waypoint information; confirm photos
[61] power devices; move cables device configuration
[62] take photos; control aircraft; provide target airspeed and altitude; create flight plan; provide waypoint information; confirm photos create flight plan; provide waypoint information; confirm photos
[63] accept or reject recommendation; make final decision recommends label; has access to hidden information
[64] score goals score goals
[65] assign 64-cell embryo; consider AI predictions provide predictions
[66] protect planets; eliminate enemies protect planets; eliminate enemies; capture planets
[67] allocate resources; police response; fire response hazmat response; fire response
[68] designer; discuss app features software developer; discuss app features
[69] make initial decision; consider AI decision; make final decision make recommendations
[70] identify luggage discuss capabilities; identify luggage

There is a large variety in roles assigned to both humans and AI, though a common role for human teammates was to accept/reject an AI recommendation to make a final decision (18 studies), whilst a common role for AI was to provide a recommendation/advise a human teammate (15 studies). The types of tasks are classified in the task level characteristic taxonomy from [7] in Table 3, and summarised in Figure 6.

Table 3: A table showing the task level characteristics of both humans and AI in experimental papers. Note that papers 146, 151, 166 and 176 contain no AI.
1-7 (l)9-14 Human Teammate(s) AI Teammates(s)
Paper Defined problem solving Ill-defined problem solving Psychomotor action Managing others Advising others Human Service Defined problem solving Ill-defined problem solving Psychomotor action Managing others Advising others Human service
1-7 (l)9-14 [27]
[28]
[29]
[30]
[31]
[32]
[33]
[34]
[35]
[17]
[17]
[36]
[37]
[38]
[39]
[40]
[16]
[4]
[41]
[42]
[43]
[2]
[44]
[45]
[46]
[19]
[47]
[48]
[21]
[49]
[13]
[50]
[51]
[52]
[3]
[53]
[54]
[55]
[56]
[57]
[58]
[59]
[60]
[14]
[61]
[62]
[63]
[64]
[65]
[66]
[67]
[68]
[69]
[70]
1-7 (l)9-14
Figure 6: A bar chart showing the different task level characteristics for each paper, separated by human and AI teammates.

The majority of studies involve defined problem-solving for both humans (52 studies) and AI (47 studies), and a sizeable portion involve ill-defined problem-solving (human roles 29 studies, AI roles 23 studies) and psychomotor action (human roles 22 studies, AI roles 17 studies). Human service was very rare for both humans and AI (one instance each), and there were no instances of negotiation as there were no team dynamics in which humans and AI intentionally held opposing views. The main difference between roles was in managing vs advising others, with the majority of human roles involving managing (20 studies) rather than advising (4 studies), and the opposite trend for AI (2 managing vs 32 advising studies). Overall, the majority of studies on Human-AI teaming involve humans and AI engaged in defined and ill-defined problem solving, psychomotor action, where the human takes a managing role whilst the AI takes an advisory role.

4.3 Team Level Characteristics↩︎

Experimental studies varied in terms of the types of team level characteristics, shown in Table 4 and summarised in Figure 7.

Table 4: A table showing the team level characteristics of the papers studied. Note that 166 and 176 are empty as these experiments contained no AI and only one person.
Paper Task interdependence Role structure Leadership structure Communication structure Physical distribution Team life span
[27] Sequential Functional Designated Chain Colocated Long Term
[28] Sequential Functional Designated Chain Colocated Long Term
[29] Intensive Functional Distributed Star Mixed Ad hoc
[30] Reciprocal Functional Designated Chain Colocated Ad hoc
[31] Reciprocal Functional Designated Chain Colocated Long Term
[32] Sequential Functional Designated Chain Colocated Long Term
[33] Sequential Functional Designated Chain Colocated Ad hoc
[34] Intensive Functional Distributed Star Mixed Ad hoc
[35] Sequential Functional Designated Chain Colocated Long Term
[17] Sequential Functional Designated Chain Colocated Long Term
[17] Sequential Functional Designated Chain Colocated Long Term
[36] Sequential Functional Designated Chain Colocated Ad hoc
[37] Intensive Functional Distributed Hub-and-Wheel Colocated Ad hoc
[38] Sequential Functional Designated Chain Colocated Long Term
[39] Reciprocal Functional Designated Chain Colocated Ad hoc
[40] Sequential Functional Designated Chain Colocated Long Term
[16] Reciprocal Divisional Distributed Star Colocated Ad hoc
[4] Reciprocal Divisional Distributed Star Colocated Ad hoc
[41] Sequential Functional Designated Hub-and-Wheel Colocated Long Term
[42] Sequential Divisional Designated Chain Colocated Ad hoc
[43] Intensive Functional Designated Hub-and-Wheel Colocated Ad hoc
[2] Intensive Functional Distributed Star Colocated Ad hoc
[44] Reciprocal Divisional Distributed Star Colocated Ad hoc
[45] Reciprocal Functional Distributed Star Colocated Ad hoc
[46] Reciprocal Functional Designated Chain Colocated Ad hoc
[19] Reciprocal Divisional Distributed Star Colocated Ad hoc
[47] Sequential Functional Distributed Star Colocated Ad hoc
[48] Reciprocal Divisional Designated Chain Colocated Ad hoc
[21] Sequential Divisional Temporary Star Colocated Ad hoc
[49] Reciprocal Functional Designated Chain Colocated Ad hoc
[13] Sequential Functional Designated Chain Colocated Long Term
[50] Intensive Divisional Distributed Star Mixed Long Term
[51] Sequential Functional Designated Chain Colocated Ad hoc
[52] Intensive Functional Temporary Star Colocated Ad hoc
[3] Reciprocal Divisional Designated Chain Colocated Long Term
[53] Sequential Functional Designated Chain Colocated Long Term
[54] Intensive Divisional Temporary Star Distributed Ad hoc
[55] Intensive Divisional Distributed Star Mixed Ad hoc
[56] Intensive Functional Distributed Star Colocated Ad hoc
[57] Sequential Functional Designated Chain Distributed Ad hoc
[58] Intensive Functional Distributed Star Distributed Ad hoc
[59] Sequential Functional Designated Chain Colocated Long Term
[60] Intensive Functional Distributed Star Colocated Ad hoc
[14] Intensive Functional Distributed Star Distributed Ad hoc
[61] Reciprocal Functional Designated Chain Colocated Ad hoc
[62] Intensive Functional Distributed Star Distributed Ad hoc
[63] Sequential Functional Designated Chain Colocated Long Term
[64] Intensive Divisional Distributed Star Colocated Ad hoc
[65] Sequential Functional Designated Chain Colocated Long Term
[66] Intensive Functional Distributed Star Colocated Ad hoc
[67] Intensive Functional Distributed Star Distributed Ad hoc
[68] Reciprocal Functional Distributed Star Distributed Long Term
[69] Sequential Functional Designated Chain Colocated Long Term
[70] Intensive Functional Distributed Star Colocated Long Term
Figure 7: A tree diagram illustrating how each of the six team level characteristics break down.

The majority of discrete categories were present in the experimental studies, except for pooled interdependence, and external manager leadership structure. This is likely because all Human-AI Teaming scenarios involved a leader who was present inside the team, and there was always a level of interaction between team members.

As can be seen, some categories were more common than others, such as a preference for functional role structures (i.e. the team members do distinctly different roles that are not interchangeable). Given that Human-AI Teaming is commonly associated with the idea that humans and AI have complementary strengths and weaknesses (e.g. [1], [22]), this is understandable. However, some categories within characteristics were more varied. For example, task interdependence was fairly evenly split between sequential, reciprocal, and intensive. Consequently, there were a varied number of combinations of team characteristics present in the literature. The most common combinations are explored in the following section.

4.4 The Sub-Types of Human-AI Teams↩︎

The five largest clusters of Human-AI Teams are described here, accounting for 41 of 53 papers (77%). They are presented in decreasing size order.

4.4.1 Team Type 1: AI Assistant↩︎

The largest cluster consists of 18 studies (34%), and are teams where a singular AI assists a singular human with a team task. The characteristics of this cluster are summarised in Figure 8.

Figure 8: A summary of the AI Assistant sub-type of Human-AI Teaming.

Teams of this kind have a sequential task interdependence, with functional roles, designated leadership, and chain communication, where humans and AI are primarily colocated (17/18) and primarily have a long term team lifespan (14/18). In other words, a human and AI take a turn in completing their part of the task, where their roles are not interchangeable, as the AI assistant has access to more information via model training and the human is the only one allowed to make a final decision. Further, one team member remains in charge throughout, and they communicate within a hierarchical structure.

For example, in Paper [38], the team is required to classify a series of images including easily recognisable objects (e.g. church, golf ball) and challenging dog breeds (e.g. old English sheepdog, Australian terrier). To do so, an AI model recommends a classification decision, then the human also considers a recommendation before making the final team decision. Consequently, the team members work in sequence (AI prediction followed by human decision) where it is not possible for the human and AI to swap roles. Further, the human is designated in charge of the team decision, meaning the AI communicates its decision ‘up’ the chain to its leader. The team members are somewhat colocated, in that they are both working on a digital task, and the team lifespan could be considered long term if the task of classification was part of the team’s ‘job.’

The types of tasks humans are likely to engage in within this cluster include defined problem-solving (17/18), and managing others (typically their AI teammate, 8/18), whereas AI team members are likely to engage in defined problem-solving (17/18) and advising others (i.e. their human teammate, 15/18). For example, in Paper [17], the team is asked to appraise a series of real estate descriptions, whereby the human first estimates the house price, followed by seeing the AI’s prediction, and subsequently ‘adjusting’ the AI’s prediction to create the final prediction. Consequently, the human and AI are both involved in defined problem-solving, however the AI is advising the human about the prediction, and the human is managing the AI by altering its output. This is slightly in contrast to the example explained from Paper [38], whereby there is no managing of the AI by the human, as they make their own judgement after seeing the AI’s prediction.

Common application domains where this type of team are studied include classification tasks (8/18), experimental games (3/18), and cybersecurity/security (3/18). Common tasks similarly involve object classification (11/18), security (2/18), teamwork games (2/18), and workbased activities (2/18).

4.4.2 Team Type 2: Ad hoc dependency teams↩︎

The second largest cluster consists of 11 studies (20%), and are teams where there are at least three members, with at least one team member being an AI, performing different roles together. The characteristics of this cluster are summarised in Figure 9.

Figure 9: A summary of the Ad hoc dependency team sub-type of Human-AI Teaming.

Teams of this kind have an intensive task interdependence, with functional roles, distributed leadership, and star communication, where humans and AI are primarily colocated (5/11) or mixed distribution (4/11) and primarily have an ad hoc team lifespan (10/11). In other words, a mix of at least three humans and AI work as a unit, where their roles are not interchangeable, leadership is shared simultaneously between team members, and they communicate freely without the need for a central point of contact. They typically come together to perform a specific task before disbanding, but may or may not be physically located close to one another during. Overall, the most common combination was three team members (5/11 studies), with comparable preference for multiple humans (6/11) versus multiple AI (5/11).

For example, Paper [14] explores an aviation/military type of environment involving a team of two humans and one AI: a navigator (AI), pilot (human), and photographer (human). Together, they performed reconnaissance photography of targets during missions. Consequently, the team members work alongside each other intensely, where it is not possible to swap roles. Further, leadership is distributed in that all three must lead their own portion of the task, leading to a star communication where all members communicate freely with one another. The members would come together to perform the reconnaissance, before disbanding once the task is complete.

The types of tasks both humans and AI are likely to engage in within this cluster include defined problem-solving (11 humans vs 10 AI), ill-defined problem-solving (9 humans vs 8 AI), and psychomotor action (7 humans vs 6 AI). For example, in Paper [56] a human works alongside 5-7 AI in the game Netrek to destroy enemy ships (human and AI role) and conquer planets (AI role only). Consequently, both humans and AI engage in defined problem-solving to destroy targets, as well as ill-defined problem-solving in order to strategise during engagements with enemies. Both also involve psychomotor action, in that humans and AI control ships as part of the activity.

Common application domains for this team type include aviation/military/space (3/11 papers), puzzles (2/11), and disaster response (2/11). Common tasks similarly include aerialbased (3/11), teamwork games (3/11), and search and rescue (2/11).

4.4.3 Team Type 3: Ad hoc forced dependency teams↩︎

The third largest cluster consists of six studies (11%), and are teams with a singular human and singular AI, similar to Cluster 1. The characteristics of this cluster are summarised in Figure 10.

Figure 10: A summary of the Ad hoc forced dependency team sub-type of Human-AI Teaming.

Teams of this kind have a reciprocal task interdependence, with functional roles, designated leadership, and chain communication, where humans and AI are colocated and primarily have an ad hoc team lifespan (5/6). In other words, a human and AI engage in a back and forth interaction, where their roles are not interchangeable, one team member remains in charge throughout, and they communicate within a hierarchical structure.

This is therefore different than Cluster 1 (AI assistant), as teams here involve multiple stages of interacting with one another, and teams are typically formed to perform a specific task before disbanding. They are consequently ‘forced’ to interact with one another to achieve the team goal, typically because not all actions can be completed by one person, or one side of the team has access to information the other does not. This is in contrast to Cluster 1, where both teammates are technically capable of completing the task on their own, but are theorised to perform better in combination. The team type is also similar to Cluster 2 (ad hoc dependency teams), however only contains two teammates, thereby changing the interdependence, leadership, and communication style.

For example, in Paper [46], the team is required to work together in order to successfully deliver food orders in the game Overcooked. Each team member has access to specific elements of the cooking process (e.g. chopping, frying), meaning they have different roles to fill. The human is the designated leader, as the AI is instructed on what elements of the cooking is required, consequently creating a chain style of communication. The human and AI exist within the same simulated environment, and they come together to complete the level before disbanding.

The types of tasks humans are likely to engage in within this cluster include defined problem-solving (6/6), psychomotor action (4/6), and managing others (i.e. their AI teammate, 4/6), whereas AI are likely to engage in defined problem-solving (5/6). For example, Paper [30] explores a combat search and rescue simulation where an AI controls a drone that surveys the area looking for hostages, which the human teammate can control whether the system is fully autonomous or not. Consequently, both the human and AI are involved in defined problem-solving (finding hostages) and psychomotor action (controlling the drone), but the human is also involved in managing others (i.e. the AI drone) and ill-defined problem-solving (when/how to alter the drone’s navigation) , whilst the AI advises the human. This is therefore different from Cluster 1 again, as there is more likely to be psychomotor action involved.

Application domains were varied for this team type, including games (2/6), as well as disaster response, work, education, and cybersecurity/security (1 each). Common tasks primarily involved teamwork games (3/6).

4.4.4 Team Type 4: Paired Equanimity↩︎

The fourth largest cluster consists of four studies (7%), and are teams with two teammates that perform the same role together. The characteristics of this cluster are summarised in Figure 11.

Figure 11: A summary of the paired equanimity team sub-type of Human-AI Teaming.

Teams of this kind have a reciprocal task interdependence, with divisional roles, distributed leadership, and star communication, where humans and AI are colocated and have an ad hoc team lifespan. In other words, a human and AI engage in a back and forth interaction, where they both perform the same role, leadership is shared simultaneously, and they communicate freely. The teammates come together to perform a specific task before disbanding, are are colocated during the task. This is therefore different than Cluster 3 (Ad hoc forced dependency) because here the team members share leadership and roles, which in turn changes the communication style to a star rather than a chain.

For example, in Paper [16] a human and AI work together to score goals in the game Rocket League. Consequently, the team members work reciprocally (i.e. passing the ball back and forth) whilst performing the same role of controlling a team vehicle. Further, leadership is distributed as both teammates can take control of the ball, leading to a star communication style. The members come together to perform the game before disbanding, whilst both being colocated during the match.

The types of tasks humans and AI engage with overlap, as they are both likely to engage in defined problem-solving (4/4), ill-defined problem-solving (4/4), and psychomotor action (4/4). For example, in Paper [19], a human and AI work together in the combat simulation game Arma III to collect as many boxes, in numerical order, as possible, in eight minutes. Consequently, both human and AI engage in the same defined and ill-defined problem-solving (i.e. locating boxes whilst coordinating actions) as well as psychomotor actions, by navigating the environment.

Common application domains where this type of team are studied games (3/4) and aviation/military/space (1/4). Similarly, the common tasks include teamwork games (3/4) and space settings (1/4).

4.4.5 Team Type 5: Group Equanimity↩︎

The fifth largest cluster consists of three studies (6%), and are teams where there are at least three members, with at least one team member being an AI, perform the same role together. The characteristics of this cluster are summarised in Figure 12.

Figure 12: A summary of the Group Equanimity team sub-type of Human-AI Teaming.

Teams of this kind have an intensive task interdependence, with divisional roles, distributed leadership, and star communication, where humans and AI are likely mixed distribution (2/3 studies) and have an ad hoc team lifespan (2/3). In other words, a mix of at least three humans and AI work as a unit, where they all perform the same role, leadership is shared simultaneously, and they communicate freely. The teammates come together to perform a specific task before disbanding, are are a mixture of physically colocated and distributed during the task. Overall, the typical team makeup was between 1-8 humans with 1-2 AI, with a slight preference for multiple humans (2/3) rather than multiple AI (1/3).

Consequently, this team type is similar to Cluster 2 (ad hoc dependency teams), however the roles being performed are the same/interchangeable. It is also similar to Cluster 4 (paired equanimity), except there are at least three team members, which subsequently changes the task interdependence.

For example, in Paper [55], 3-4 humans work with an AI to answer a series of quiz questions. Each member first individually guesses the answer, before they each show their answer to the team to reach a consensus. Consequently, the team members work alongside each other intensely, where the task of deciding the answer is jointly undertaken. Further, leadership is distributed across team members to reach a consensus, leading to a star communication. The members come together to answer the quiz, before disbanding.

The types of tasks humans are likely to engage with include defined problem-solving (3/3) and ill-defined problem-solving (2/3), and AI are likely to engage in defined problem-solving (2/3) and advising others (2/3). For example, in Paper [64] a game of Rocket League is played between three team members (2 humans, 1 AI), who work together to score goals. This experimental setup is therefore different than Cluster 4 (paired equanimity), as there are three teammates rather than two. For this experiment, both humans and AI engage with defined and ill-defined problem-solving (i.e. coordinating the scoring of goals), as well as psychomotor action by moving vehicles to interact with the ball.

Common application domains include puzzles (2/3), and education (1/3). Common task types include teamwork games, puzzles, and workbased activities.

5 Discussion↩︎

In this paper, the types of human-AI team currently studied in the literature was explored, to understand what the term means for the field. To do so, 53 experimental papers exploring human-AI teaming were analysed, revealing strong preferences for teams consisting of 1 human and 1 AI, in the areas of games, object classification, and aviation/military/space. The studies were categorised using the task and team level characteristic taxonomies provided by [7]. Both humans and AI were most likely to be involved in defined problem-solving tasks, however humans were also more likely to be involved in managing others, whilst AI were more likely to be involved in advising others. In terms of team level characteristics, five main clusters were observed, representing unique combinations. These are discussed below.

5.1 Unique Subtypes of Human-AI Teaming Exist↩︎

The five main clusters found from the analysis are summarised in Table 5. Several insights can be made about the human-AI teaming literature by considering these clusters. For example, the main areas of interest to the research community can be seen as either falling into Cluster 1 (AI Assistant) or Cluster 2 (Ad hoc Dependency), accounting for 55% of papers (29/53). In other words, the majority of studies focus on teaming where an AI assists a human with a task, or a group of at least three team members with different roles working closely together.

Table 5: A summary of the five main clusters of human-AI teaming sub-types.
Cluster AI Assistant Ad hoc Dependency Ad hoc Forced Dependency Paired Equanimity Group Equanimity
Team Makeup 2 3+ 2 2 3+
Task Interdependence sequential intensive reciprocal reciprocal intensive
Role Structure functional functional functional divisional divisional
Leadership Structure designated distributed designated distributed distributed
Communication Structure chain star chain star star
Physical Distribution colocated/distributed colocated/mixed/distributed colocated colocated colocated/mixed
Team Life Span long term/ad hoc ad hoc/long term ad hoc/long term ad hoc ad hoc/long term
Most Common Application classification aviation/military/space games games puzzles
Most Common Task object classification aerialbased teamwork games teamwork games teamwork games
Number of Studies 18 11 6 4 3

However, we can also see areas that have not received as much attention from the literature. Considering there were six dimensions of team level characteristics, of which four were used here to cluster studies, the number of potential combinations far surpasses those seen here. Several papers did not form a cluster as the combination of characteristics was unique. Paper [43] for example looked at a team of 1 human with 5 AI drones, where the human provided demonstrations of movement and corrected the AI where appropriate. Consequently, this represents a team with intensive task interdependence, functional roles, designated leadership, hub-and-wheel communication, colocation, and ad hoc lifespan. Furthermore, other team types were not seen at all, such as any involving a pooled task interdependence, or an external manager leadership structure. There is no reason to assume these types of teams cannot exist, or would not be of interest to study for a specific context. Therefore, the application of the teaming taxonomies to experimental studies can reveal gaps in study, and interesting areas for future research.

5.2 Implications for Human-AI Teaming Research↩︎

As can be seen, there are strong overlaps in the types of team level characteristics observed between clusters, such as a slight preference for distributed leadership and star communication structures. However, each cluster is unique in its combination of characteristics, creating inherent differences between them. It is therefore important to consider each one as unique, whilst still falling under the broader umbrella term of ‘a human-AI team.’

However, doing so presents an issue for the field of human-AI teaming. Despite their overlaps, each cluster remains unique. For example, Cluster 1 (AI Assistant) and Cluster 4 (Paired Equanimity) share almost no team level characteristics in common, yet both are ‘human-AI teams.’ Consequently, any insights gained from studying a human-AI team with the characteristics of Cluster 1 are unlikely to transfer to a study with the team characteristics of Cluster 4. If, for example, a paper using a Cluster 4 team makes an insight into how to manage leadership roles in a human-AI team (such as [16]), it would be difficult to transfer this insight on distributed leadership to Cluster 1, with its designated leadership (such as [38]).

Currently, authors would be required to read each individual study to assess its similarity to the author’s conception of a human-AI team. There has been a lack of taxonomic language that could aid in this comparison; that is, knowledge that team-level characteristics exist, have been previously defined, and can be used to separate different kinds of teams. Further, because of this lack of previous language, it may prove challenging for authors to extract this information, depending on the level of information about team level characteristics provided within papers to date. Alternatively, authors may seek meta reviews which synthesise key insights across papers to create general guidance. However, the nuances between team types are likely lost in this process, reducing the usefulness and applicability of reviews. In doing so, authors may find it difficult to understand what guidance applies to their type of human-AI team, and what does not.

Therefore, it is important to not treat all human-AI teaming work as inherently interchangeable, due to differences in team level characteristics. This is not the same argument as remaining mindful of contexts such as application domain, which has been argued previously (e.g. calls for wider sociotechnical perspectives; [1]). As seen in the analysis, each cluster contained a range of application domains and task environments, though there were some preferences in certain clusters. Furthermore, it is important not to consider a team’s makeup as the only important distinction between two studies on human-AI teams. Three of the five main clusters involved 1 human working with 1 AI, however the ways in which they work together vary greatly. For example, whilst Cluster 3 (Ad hoc Forced Dependency) and Cluster 4 (Paired Equanimity) both involve 1 human and 1 AI, they are not similar in terms of team level characteristics. As a consequence, the ways in which these humans and AI work together will be markedly different. Therefore, we argue that care should be taken when synthesising research across studies, so that important nuances are not overlooked.

5.2.1 Wider Reflections on Applying Teaming Taxonomies to Human-AI Teaming↩︎

The analysis highlights different interpretations of the human-AI teaming definition in experimental work, particularly in terms of interdependence, complementary capabilities, and perceptions of AI as a teammate. We are not the first to highlight differences in definitions, however others have focused primarily on the outcomes of teams. For example, [71] considers differences in teaming situational awareness, whilst [67] considers how teaming composition impacts team performance, and [17] considers the impact of complementarity on team performance. Instead, we focus specifically on the words used when describing teams, and what this tells us about how researchers actively interpret the definition. Each difference in interpretation consequently highlights ambiguities in the definition that require resolution in order to truly synthesise the field, explored below.

Firstly, in human-AI teaming definitions there is an emphasis on interdependence (e.g. [1]). Similarly, interdependence is also important in team level characteristics taxonomies (e.g. [7]), however here the interdependence is further specified into four sub-types (pooled, sequential, reciprocal, and intensive). Currently, there has not been a discussion in the literature on what interdependence means in practice, and if the type of interdependence is important for human-AI teams. This highlights a need for a more nuanced conversation around what is meant by interdependence for a given research context.

Secondly, human-AI teaming definitions asks us to consider that humans and AI have “unique and complementary capabilities” [1]. However, as was seen in the results, some team types involve teams where roles are the same between humans and AI (i.e. divisional role structures). In particular, Clusters 4 and 5 involve divisional roles between humans and AI, seemingly contradicting the requirement that humans and AI have ‘unique’ capabilities. This implies complementarity is either not an important component of human-AI teaming in all contexts, or certain types of teams studied currently do not align well under the human-AI teaming definition.

The question of whether all experimental studies analysed here align under current conceptions of human-AI teaming is brought into starker contrast for the final requirement that AI is perceived as a teammate. It is also one of the most important components of the definition, as it separates teaming from other similar terms such as human-AI interaction or human-AI tool use. By considering the analysis performed here, interpreting this requirement can be challenging. For example, for Cluster 1, and in some ways Cluster 2, it is difficult to articulate how AI in this team is more than a simple tool. In Cluster 1, the task level characteristics show that the human remains in control throughout, and is the only one to make a final decision. Further, the human is technically capable of completing the entire task on their own, with the assumption that AI is there to enhance the performance or efficiency of the task. Therefore, it is unclear how to describe the AI in these teams as being more than a tool, or conversely, explain what a tool has to be capable of doing before it is no longer considered a tool. Even considering the team level characteristics cannot aid in distinguishing tools from teammates. For example, divisional roles may imply a level of similarity between humans and AI, thereby implying a perception of being a teammate. However, functional roles were very common in experimental studies, and further can indicate that an AI provides a necessary skillset to the team.

Overall, there are issues in applying existing human-AI teaming definitions to practical implementations and concrete examples. These difficulties stem from two areas. Firstly, as part of the human-AI teaming definition, we are asked to consider an AI as if it was a human. This is because the definitions have been built from human-only teaming literature, consequently leaving us to evaluate AI based on human characteristics. In other words, human-AI teaming, by definition, attempts to ascribe human characteristics to a computer. This is akin to assigning teaming taxonomies to technologies such as a vehicle, an automatic light, or a calculator. By taking some generous and interpretive steps, one could consider these technologies to have teaming qualities such as interdependence. However, is this a helpful framing? Or is the desire to understand human-AI teams within a framework built primarily for human-only teams too restrictive?

Secondly, the wide range of teaming examples observed in this study highlights the need to answer whether the human-AI teaming definition is specific enough for purpose, or in some cases, overused in inappropriate settings. If it is important that human-AI teams involve AI that are uniquely capable and perceived as teammates, then many of the studies analysed would not fit under this definition. For example, all of Cluster 1 would be removed for being a simple tool rather than a teammate, and Clusters 4 and 5 would similarly be removed for lack of unique functional roles. This would remove 25 studies, or 46% of all studies analysed. Therefore, the main question for the field becomes: when is a human-AI team a team, and what characteristics are truly emblematic of this?

To synthesise and advance the human-AI teaming literature, there is a need to consider more carefully what aspects of teaming should be built into its definition. In particular, which aspects of teaming are feasible to transfer from human-only to human-AI teams, and what characteristics are unique to human-AI teams, is yet to be determined. Synthesis is particularly important to achieve as safety-critical applications are increasingly interested in human-AI teaming, particularly due to recent advancements in agentic AI systems and their conversational capabilities. Clearer taxonomies and definitions are becoming increasingly essential, as they form the basis for three critical processes: (1) safety analysis and assurance, where defining the scope of the subject of interest is key [72]; (2) regulatory frameworks, such as determining how to oversee a collaborative team that goes beyond individually regulated components (e.g. a qualified clinician working with an AI medical device), and; (3) the allocation of moral responsibility and legal liability [73], which mitigates the risk of ‘responsibility gaps’ [74] and the ‘problem of many hands’ [75]. However, whilst important to consider, it is unlikely that taxonomies and definitions built solely from human-only teams will be sufficient to capture all nuances of human-AI teams, particularly for safety-critical contexts. For example, the distinction from tool to teammate, and how to interpret physical distribution, are unlikely resolvable from psychological literature alone. Therefore, there is a need to develop taxonomies of human-AI teams that build on human-only team literature, but also extends and explains what is unique about having an AI teammate.

5.3 A Checklist for Reporting Human-AI Teams↩︎

Given the presence of multiple sub-types of human-AI teams present in the literature, it is important that future work more accurately describes what is being studied under the global human-AI teaming definition. The following checklist is provided as an initial way to help authors consider the specifics of the human-AI teams they are studying. When designing and reporting human-AI teaming studies, we suggest the following questions be considered, shown in Table 6.

Table 6: Questions to guide the specification of the type of human-AI team to be studied or reported.
Question Type Question
General Team Characteristics How many human and AI team members are in the team?
What is the intended application domain (e.g. healthcare, aviation, work offices)
What is the overall team goal?
Task Level Characteristics What roles do human teammates perform?
What roles do AI teammates perform?
How do human and AI teammates interact?
Team Level Characteristics What is the task interdependence? (pooled, sequential, reciprocal, intensive)
What is the role structure? (functional, divisional)
What is the leadership structure? (external, designated, temporary, distributed)
What is the communication structure? (hub-and-wheel, chain, star)
Is physical distribution an important component?
What is the team life span? (ad hoc, long term)

In doing so, authors are encouraged to understand their human-AI team more specifically, in turn allowing easier comparison to other work and aiding the identification of overlaps as well as gaps in understanding. It is not intended to be an exhaustive list, but rather a systematic basis for scoping and framing the conversation around what is unique about human-AI teaming research and its sub-types, and encourage more transparent experimental reporting. We also invite authors to consider what other questions are important to consider when describing a human-AI team, to build upon this initial step.

5.4 Limitations & Future Work↩︎

There are a few limitations to the work undertaken in this study. Firstly, only papers that explicitly included the term “human-AI team*" in the title or keywords were included in the sample. It is possible papers were missed that did not include this specific phrase (such as human agent teaming, or human autonomy teaming), or did not do so in the title or keywords. However, a large number of papers were still collected, and can still provide a scoping overview of the current experimental literature.

Secondly, an established taxonomy was used to conduct the content analysis, to reduce the risk of bias in interpreting teaming characteristics. However, the taxonomy required slight interpretations to apply to human-AI teams, which could have introduced errors in coding. To enhance transparency, we have included the dataset of all papers analysed, as well as their assigned categories. We encourage readers to review the data and consider how the taxonomic labels apply to the papers in the set, as well as other literature in the field.

In terms of future work, the current study highlighted the existence of multiple sub-types of human-AI teams, all studied under the same term. Some combinations of team characteristics are more commonly studied than others, indicating potential areas of future research in these understudied teams. For example, future work could look into characteristics such as pooled task interdependence or external manager leadership, and how these influence human-AI teaming.

However, it is also important to note the sub-types are unique and not interchangeable. Consequently, there is a need for more clarity in how human-AI teams are described in future research. We encourage readers to reflect on what types of human-AI teams they are interested in, and what characteristics of the team are important for fellow readers to know in order to understand the team. To this end, to aid in effective research communication, we also encourage readers to use the proposed checklist as a baseline for how to report a human-AI team in experimental work.

6 Conclusion↩︎

To further advance the field of human-AI teaming, there is a need for more precise language to describe the types of teams that are studied. We propose an initial checklist on key features to report and consider when designing a team, and encourage authors to reflect more broadly on what types of teaming they are interested in. In doing so, we hope there will be a better synthesis of human-AI teaming literature, as well as new avenues in exploring currently understudied combinations of team level characteristics.

7 Acknowledgements↩︎

This work was supported by the Centre for Assuring Autonomy, a partnership between Lloyd’s Register Foundation and the University of York.

8 Appendix↩︎

Table 7: All clusters of human-AI teams found in the analysis.
Cluster Paper Application Type Task Type Team Makeup Task interdependence Role structure Leadership structure Communication structure Physical distribution Team life span
1 [64] Generic/Theoretical Teamwork Games 1 human 2 AI; 2 human 1 AI Intensive Divisional Distributed Star Colocated Ad hoc
1 [55] Generic/Theoretical Quiz 3-4 humans 1 AI Intensive Divisional Distributed Star Mixed Ad hoc
1 [50] Education Workbased Activity 6-8 humans 1 AI Intensive Divisional Distributed Star Mixed Long Term
2 [54] Misc. Teamwork Workbased Activity 2-3 humans 1 AI Intensive Divisional Temporary Star Distributed Ad hoc
3 [43] Aviation/Military/Space Aerialbased 1 human 5 AI Intensive Functional Designated Hub-and-Wheel Colocated Ad hoc
4 [37] Game Teamwork Games 1 human 2 AI Intensive Functional Distributed Hub-and-Wheel Colocated Ad hoc
5 [56] Game Teamwork Games 1 human 5-7 AI Intensive Functional Distributed Star Colocated Ad hoc
5 [66] Generic/Theoretical Teamwork Games 1 human 5-7 AI Intensive Functional Distributed Star Colocated Ad hoc
5 [14] Aviation/Military/Space Aerialbased 1 human 2 AI; 2 human 1 AI Intensive Functional Distributed Star Distributed Ad hoc
5 [62] Aviation/Military/Space Aerialbased 1 human 2 AI; 2 human 1 AI Intensive Functional Distributed Star Distributed Ad hoc
5 [2] Healthcare Workbased Activity 4 humans 1 AI Intensive Functional Distributed Star Colocated Ad hoc
5 [60] Education Space 2-5 humans 1 AI Intensive Functional Distributed Star Colocated Ad hoc
5 [70] Cybersecurity/Security Security 2 humans 1 AI Intensive Functional Distributed Star Colocated Long Term
5 [58] Aviation/Military/Space Aerialbased 2 humans 1 AI Intensive Functional Distributed Star Distributed Ad hoc
5 [67] Disaster Response Search and Rescue 1 human 2 AI; 2 human 1 AI Intensive Functional Distributed Star Distributed Ad hoc
5 [29] Disaster Response Search and Rescue 3 humans 1 AI Intensive Functional Distributed Star Mixed Ad hoc
5 [34] Generic/Theoretical Teamwork Games 3-4 humans 1 AI Intensive Functional Distributed Star Mixed Ad hoc
6 [52] Aviation/Military/Space Aerialbased 2 humans 1 AI Intensive Functional Temporary Star Colocated Ad hoc
7 [48] Aviation/Military/Space Space 1 human 1 AI Reciprocal Divisional Designated Chain Colocated Ad hoc
7 [3] Education Workbased Activity 1 human 1 AI Reciprocal Divisional Designated Chain Colocated Long Term
8 [44] Aviation/Military/Space Space 1 human 1 AI Reciprocal Divisional Distributed Star Colocated Ad hoc
8 [16] Game Teamwork Games 1 human 1 AI Reciprocal Divisional Distributed Star Colocated Ad hoc
8 [4] Game Teamwork Games 1 human 1 AI Reciprocal Divisional Distributed Star Colocated Ad hoc
8 [19] Game Teamwork Games 1 human 1 AI Reciprocal Divisional Distributed Star Colocated Ad hoc
9 [30] Disaster Response Search and Rescue 1 human 1 AI Reciprocal Functional Designated Chain Colocated Ad hoc
9 [39] Game Teamwork Games 1 human 1 AI Reciprocal Functional Designated Chain Colocated Ad hoc
9 [46] Game Teamwork Games 1 human 1 AI Reciprocal Functional Designated Chain Colocated Ad hoc
9 [49] Work Teamwork Games 1 human 1 AI Reciprocal Functional Designated Chain Colocated Ad hoc
9 [61] Education Workbased Activity 1 human 1 AI Reciprocal Functional Designated Chain Colocated Ad hoc
9 [31] Cybersecurity/Security Security 1 human 1 AI Reciprocal Functional Designated Chain Colocated Long Term
10 [45] Disaster Response Search and Rescue 1 human 1 AI Reciprocal Functional Distributed Star Colocated Ad hoc
10 [68] Work Workbased Activity 1 human 1 AI Reciprocal Functional Distributed Star Distributed Long Term
11 [42] Game Teamwork Games 1 human 1 AI Sequential Divisional Designated Chain Colocated Ad hoc
12 [21] Game Teamwork Games 1 human 1 AI Sequential Divisional Temporary Star Colocated Ad hoc
13 [33] Game Teamwork Games 1 human 1 AI Sequential Functional Designated Chain Colocated Ad hoc
13 [53] Education Workbased Activity 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [36] Experimental Game Teamwork Games 1 human 1 AI Sequential Functional Designated Chain Colocated Ad hoc
13 [35] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [17] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [17] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [59] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [28] Cybersecurity/Security Security 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [51] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Ad hoc
13 [38] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [40] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [69] Classification Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [32] Cybersecurity/Security Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [27] Experimental Game Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [63] Experimental Game Object Classification 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [13] Cybersecurity/Security Security 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [65] Education Workbased Activity 1 human 1 AI Sequential Functional Designated Chain Colocated Long Term
13 [57] Disaster Response Search and Rescue 1 human 1 AI Sequential Functional Designated Chain Distributed Ad hoc
14 [41] Shopping Teamwork Games 1 human 2 AI Sequential Functional Designated Hub-and-Wheel Colocated Long Term
15 [47] Misc. Teamwork Quiz 4 humans 1 AI Sequential Functional Distributed Star Colocated Ad hoc

References↩︎

[1]
S. Berretta, A. Tausch, G. Ontrup, B. Gilles, C. Peifer, and A. Kluge, “Defining human-AI teaming the human-centered way: A scoping review and network analysis,” Frontiers in Artificial Intelligence, vol. 6, p. 1250725, 2023.
[2]
N. Bienefeld, M. Kolbe, G. Camen, D. Huser, and P. K. Buehler, “Human-AI teaming: Leveraging transactive memory and speaking up for enhanced team effectiveness,” Frontiers in Psychology, vol. 14, p. 1208019, 2023.
[3]
X. Lu, S. Fan, J. Houghton, L. Wang, and X. Wang, “ReadingQuizMaker: A human-NLP collaborative system that supports instructors to design high-quality reading quiz questions,” Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. p. 1–18, 2023.
[4]
C. Flathmann, B. G. Schelble, P. J. Rosopa, N. J. McNeese, R. Mallick, and K. C. Madathil, “Examining the impact of varying levels of AI teammate influence on human-AI teams,” International Journal of Human-Computer Studies, vol. 177, p. 103061, 2023.
[5]
D. Wang et al., “From human-human collaboration to human-AI collaboration: Designing AI systems that can work together with people,” Extended abstracts of the 2020 CHI conference on human factors in computing systems, pp. p. 1–6, 2020.
[6]
A. Čartolovni, A. Tomičić, and E. L. Mosler, “Ethical, legal, and social considerations of AI-based medical decision-support tools: A scoping review,” International Journal of Medical Informatics, vol. 161, p. 104738, 2022.
[7]
J. L. Wildman, A. L. Thayer, M. A. Rosen, E. Salas, J. E. Mathieu, and S. R. Rayne, “Task types and team-level attributes: Synthesis of team classification literature,” Human resource development review, vol. 11, no. 1, pp. 97–129, 2012.
[8]
N. J. McNeese, M. Demir, N. J. Cooke, and C. Myers, “Teaming with a synthetic teammate: Insights into human-autonomy teaming,” Human factors, vol. 60, no. 2, pp. 262–273, 2018.
[9]
K. Krippendorff, Content analysis: An introduction to its methodology. Sage publications, 2018.
[10]
R. Zhang, N. J. McNeese, G. Freeman, and G. Musick, “" an ideal human" expectations of AI teammates in human-AI teaming,” Proceedings of the ACM on Human-Computer Interaction, vol. 4, no. CSCW3, pp. 1–25, 2021.
[11]
T. O’neill, N. McNeese, A. Barron, and B. Schelble, “Human–autonomy teaming: A review and analysis of the empirical literature,” Human factors, vol. 64, no. 5, pp. 904–938, 2022.
[12]
E. Salas, N. J. Cooke, and M. A. Rosen, “On teams, teamwork, and team performance: Discoveries and developments,” Human factors, vol. 50, no. 3, pp. 540–547, 2008.
[13]
A. Smith, H. P. van Wagoner, K. Keplinger, and C. Celebi, “Navigating AI convergence in human–artificial intelligence teams: A signaling theory approach,” Journal of Organizational Behavior, 2025.
[14]
W. Duan et al., “Understanding the evolvement of trust over time within human-AI teams,” Proceedings of the ACM on Human-Computer Interaction, vol. 8, no. CSCW2, pp. 1–31, 2024.
[15]
Q. Gao, W. Xu, M. Shen, and Z. Gao, “Agent teaming situation awareness (ATSA): A situation awareness framework for human-AI teaming,” arXiv preprint arXiv:2308.16785, 2023.
[16]
C. Flathmann, W. Duan, N. J. Mcneese, A. Hauptman, and R. Zhang, “Empirically understanding the potential impacts and process of social influence in human-AI teams,” Proceedings of the ACM on Human-Computer Interaction, vol. 8, no. CSCW1, pp. 1–32, 2024.
[17]
P. Hemmer, M. Schemmer, N. Kühl, M. Vössing, and G. Satzger, “Complementarity in human-AI collaboration: Concept, sources, and evidence,” European Journal of Information Systems, vol. 34, no. 6, pp. 979–1002, 2025.
[18]
G. Cabour, A. Morales-Forero, É. Ledoux, and S. Bassetto, “An explanation space to align user studies with the technical development of explainable AI,” AI & SOCIETY, vol. 38, no. 2, pp. 869–887, 2023.
[19]
R. Zhang, W. Duan, C. Flathmann, N. McNeese, G. Freeman, and A. Williams, “Investigating AI teammate communication strategies and their impact in human-AI teams for effective teamwork,” Proceedings of the ACM on Human-Computer Interaction, vol. 7, no. CSCW2, pp. 1–31, 2023.
[20]
M. Loper and V. Sitterle, “Evolving lvc to include evaluation of human-ai teaming dynamics,” 2023 Winter Simulation Conference (WSC), pp. p. 2506–2517, 2023.
[21]
C. Attig, P. Wollstadt, T. Schrills, T. Franke, and C. B. Wiebel-Herboth, “More than task performance: Developing new criteria for successful human-AI teaming using the cooperative card game hanabi,” Extended abstracts of the chi conference on human factors in computing systems, pp. p. 1–11, 2024.
[22]
M. J. McGrath, A. Duenser, J. Lacey, and C. Paris, “Collaborative human-AI trust (CHAI-t): A process framework for active management of trust in human-AI collaboration,” Computers in Human Behavior: Artificial Humans, p. 100200, 2025.
[23]
C. Flathmann, B. G. Schelble, R. Zhang, and N. J. McNeese, “Modeling and guiding the creation of ethical human-AI teams,” Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp. p. 469–479, 2021.
[24]
Z. Munn, M. D. Peters, C. Stern, C. Tufanaru, A. McArthur, and E. Aromataris, “Systematic review or scoping review? Guidance for authors when choosing between a systematic or scoping review approach,” BMC medical research methodology, vol. 18, pp. 1–7, 2018.
[25]
A. C. Tricco et al., “PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation,” Annals of internal medicine, vol. 169, no. 7, pp. 467–473, 2018.
[26]
H. Arksey and L. O’malley, “Scoping studies: Towards a methodological framework,” International journal of social research methodology, vol. 8, no. 1, pp. 19–32, 2005.
[27]
G. Bansal, B. Nushi, E. Kamar, D. Weld, W. Lasecki, and E. Horvitz, “A case for backward compatibility for human-ai teams,” arXiv preprint arXiv:1906.01148, 2019.
[28]
R. Olla, E. Hand, S. J. Louis, R. Houmanfar, and S. Sengupta, “A cybersecurity game to probe human-AI teaming,” 2024 IEEE Conference on Games (CoG), pp. p. 1–5, 2024.
[29]
A. Amresh, N. Cooke, and A. Fouse, “A minecraft based simulated task environment for human AI teaming,” Proceedings of the 23rd ACM international conference on intelligent virtual agents, pp. p. 1–3, 2023.
[30]
J. Schwalb, V. Menon, N. Tenhundfeld, K. Weger, B. Mesmer, and S. Gholston, “A study of drone-based AI for enhanced human-AI trust and informed decision making in human-AI interactive virtual environments,” 2022 IEEE 3rd International Conference on Human-Machine Systems (ICHMS), pp. p. 1–6, 2022.
[31]
S. Tariq, M. B. Chhetri, S. Nepal, and C. Paris, “A2C: A modular multi-stage collaborative decision framework for human–AI teams,” Expert systems with applications, vol. 282, p. 127318, 2025.
[32]
E. Josephs, C. Fosco, and A. Oliva, “Artifact magnification on deepfake videos increases human detection and subjective confidence,” arXiv preprint arXiv:2304.04733, 2023.
[33]
C. Ong, K. McGee, and T. L. Chuah, “Closing the human-AI team-mate gap: How changes to displayed information impact player behavior towards computer teammates,” Proceedings of the 24th Australian Computer-Human Interaction Conference, pp. p. 433–439, 2012.
[34]
J. Zvelebilova, S. Savage, and C. Riedl, “Collective attention in human-AI teams,” arXiv preprint arXiv:2407.17489, 2024.
[35]
C. Xu, K.-C. Lien, and T. Höllerer, “Comparing zealous and restrained ai recommendations in a real-world human-ai collaboration task,” Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. p. 1–15, 2023.
[36]
L. Zhang, Z. Ji, and B. Chen, “Crew: Facilitating human-ai teaming research,” arXiv preprint arXiv:2408.00170, 2024.
[37]
I. Munyaka, Z. Ashktorab, C. Dugan, J. Johnson, and Q. Pan, “Decision making strategies and team efficacy in human-AI teams,” Proceedings of the ACM on Human-Computer Interaction, vol. 7, no. CSCW1, pp. 1–24, 2023.
[38]
S. A. Mahmood, Z. Lu, and M. Yin, “Designing behavior-aware AI to improve the human-AI team performance in AI-assisted decision making,” International Joint Conferences on Artificial Intelligence Organization, 2024.
[39]
J. Newn, R. Singh, F. Allison, P. Madumal, E. Velloso, and F. Vetere, “Designing interactions with intention-aware gaze-enabled artificial agents,” IFIP Conference on Human-Computer Interaction, pp. p. 255–281, 2019.
[40]
G. Bansal et al., “Does the whole exceed its parts? The effect of ai explanations on complementary team performance,” Proceedings of the 2021 CHI conference on human factors in computing systems, pp. p. 1–16, 2021.
[41]
C. C. Jorge, C. M. Jonker, and M. L. Tielman, “How should an AI trust its human teammates? Exploring possible cues of artificial trust,” ACM Transactions on Interactive Intelligent Systems, vol. 14, no. 1, pp. 1–26, 2024.
[42]
M. Sidji, W. Smith, and M. J. Rogerson, “Human-AI collaboration in cooperative games: A study of playing codenames with an LLM assistant,” Proceedings of the ACM on Human-Computer Interaction, vol. 8, no. CHI PLAY, pp. 1–25, 2024.
[43]
M. S. Islam et al., “Human-AI collaboration in real-world complex environment with reinforcement learning,” Neural Computing and Applications, vol. 37, no. 23, pp. 18957–18987, 2025.
[44]
K. Momose, R. Mehta, J. Moukpe, T. R. Weekes, and T. C. Eskridge, “Human-AI teamwork interface design using patterns of interactions,” International Journal of Human–Computer Interaction, vol. 41, no. 11, pp. 7112–7134, 2025.
[45]
M. P. Schadd, T. A. Schoonderwoerd, K. van den Bosch, O. H. Visker, and T. Haije, ‘I’m afraid i can’t do that, dave’; getting to know your buddies in a human–agent team,” Systems, vol. 10, no. 1, p. 15, 2022.
[46]
S. Bhambri, M. Verma, U. Biswas, A. Murthy, and S. Kambhampati, “Incorporating human flexibility through reward preferences in human-AI teaming,” arXiv preprint arXiv:2312.14292, 2023.
[47]
W. Ye, F. Bullo, N. Friedkin, and A. K. Singh, “Modeling human-ai team decision making,” arXiv preprint arXiv:2201.02759, 2022.
[48]
M. Li, A. V. Kamaraj, and J. D. Lee, “Modeling trust dimensions and dynamics in human-agent conversation: A trajectory epistemic network analysis approach,” International Journal of Human–Computer Interaction, vol. 40, no. 14, pp. 3571–3582, 2024.
[49]
S. Zhang et al., “Mutual theory of mind in human-ai collaboration: An empirical study with llm-driven ai agents in a real-time shared workspace task,” arXiv preprint arXiv:2409.08811, 2024.
[50]
M. Darban, “Navigating virtual teams in generative AI-led learning: The moderation of team perceived virtuality,” Education and Information Technologies, vol. 29, no. 17, pp. 23225–23248, 2024.
[51]
V. Babbar, U. Bhatt, and A. Weller, “On the utility of prediction sets in human-ai teams,” arXiv preprint arXiv:2205.01411, 2022.
[52]
J. Würfel, A. Papenfuß, and M. Wies, “Operationalizing ai explainability using interpretability cues in the cockpit: Insights from user-centered development of the intelligent pilot advisory system (IPAS),” International Conference on Human-Computer Interaction, pp. p. 297–315, 2024.
[53]
L. P. Prieto Santos et al., “Single-case learning analytics: Feasibility of a human-centered analytics approach to support doctoral education,” JUCS-Journal of Universal Computer Science, vol. 29, no. 9, pp. 1033–1068, 2023.
[54]
A. M. Harris-Watson, L. E. Larson, N. Lauharatanahirun, L. A. DeChurch, and N. S. Contractor, “Social perception in human-AI teams: Warmth and competence predict receptivity to AI teammates,” Computers in Human Behavior, vol. 145, p. 107765, 2023.
[55]
F. Milella, C. Natali, T. Scantamburlo, A. Campagner, and F. Cabitza, “The impact of gender and personality in human-AI teaming: The case of collaborative question answering,” IFIP Conference on Human-Computer Interaction, pp. p. 329–349, 2023.
[56]
R. Mallick, C. Flathmann, C. Lancaster, A. Hauptman, N. McNeese, and G. Freeman, “The pursuit of happiness: The power and influence of AI teammate emotion in human-AI teamwork,” Behaviour & Information Technology, vol. 43, no. 14, pp. 3436–3460, 2024.
[57]
M. Zhao, R. Simmons, and H. Admoni, “The role of adaptation in collective human–AI teaming,” Topics in cognitive science, vol. 17, no. 2, pp. 291–323, 2025.
[58]
B. G. Schelble et al., “Towards ethical AI: Empirically investigating dimensions of AI ethics, trust repair, and performance in human-AI teaming,” Human Factors, vol. 66, no. 4, pp. 1037–1055, 2024.
[59]
S. Jia, Z. Li, N. Chen, and J. Zhang, “Towards visual explainable active learning for zero-shot classification,” IEEE Transactions on Visualization and Computer Graphics, vol. 28, no. 1, pp. 791–801, 2021.
[60]
R. Marrone, A. Zamecnik, S. Joksimovic, J. Johnson, and M. De Laat, “Understanding student perceptions of artificial intelligence as a teammate,” Technology, Knowledge and Learning, vol. 30, no. 3, pp. 1847–1869, 2025.
[61]
A. I. Hauptman, B. G. Schelble, W. Duan, C. Flathmann, and N. J. McNeese, “Understanding the influence of AI autonomy on AI explainability levels in human-AI teams using a mixed methods approach,” Cognition, Technology & Work, vol. 26, no. 3, pp. 435–455, 2024.
[62]
W. Duan et al., “Understanding the processes of trust and distrust contagion in human–AI teams: A qualitative approach,” Computers in Human Behavior, vol. 165, p. 108560, 2025.
[63]
G. Bansal, B. Nushi, E. Kamar, D. S. Weld, W. S. Lasecki, and E. Horvitz, “Updates in human-ai teams: Understanding and addressing the performance/compatibility tradeoff,” Proceedings of the AAAI conference on artificial intelligence, vol. 33, no. 1, pp. p. 2429–2437, 2019.
[64]
R. Zhang, W. Duan, C. Flathmann, N. McNeese, B. Knijnenburg, and G. Freeman, “Verbal vs. Visual: How humans perceive and collaborate with AI teammates using different communication modalities in various human-AI team compositions,” Proceedings of the ACM on human-computer Interaction, vol. 8, no. CSCW2, pp. 1–34, 2024.
[65]
J. Hong, R. Maciejewski, A. Trubuil, and T. Isenberg, “Visualizing and comparing machine learning predictions to improve human-AI teaming on the example of cell lineage,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 4, pp. 1956–1969, 2023.
[66]
R. Mallick, C. Flathmann, W. Duan, B. G. Schelble, and N. J. McNeese, “What you say vs what you do: Utilizing positive emotional expressions to relay AI teammate intent within human–AI teams,” International Journal of Human-Computer Studies, vol. 192, p. 103355, 2024.
[67]
N. J. McNeese, B. G. Schelble, L. B. Canonico, and M. Demir, “Who/what is my teammate? Team composition considerations in human–AI teaming,” IEEE Transactions on Human-Machine Systems, vol. 51, no. 4, pp. 288–299, 2021.
[68]
E. Georganta and A.-S. Ulfert, “Would you trust an AI team member? Team trust in human–AI teams,” Journal of occupational and organizational psychology, vol. 97, no. 3, pp. 1212–1241, 2024.
[69]
Q. Zhang, M. L. Lee, and S. Carter, “You complete me: Human-ai teams and complementary expertise,” Proceedings of the 2022 CHI conference on human factors in computing systems, pp. p. 1–28, 2022.
[70]
T. Erengin, R. Briker, and S. B. de Jong, “You, me, and the AI: The role of third-party human teammates for trust formation toward AI teammates,” Journal of Organizational Behavior, 2024.
[71]
B. Lou, T. Lu, T. Raghu, and Y. Zhang, “Unraveling human-AI teaming: A review and outlook,” arXiv preprint arXiv:2504.05755, 2025.
[72]
I. Habli et al., “The big argument for ai safety cases,” arXiv preprint arXiv:2503.11705, 2025.
[73]
Z. Porter et al., “Unravelling responsibility for AI,” Journal of Responsible Technology, p. 100124, 2025.
[74]
A. Matthias, “The responsibility gap: Ascribing responsibility for the actions of learning automata,” Ethics and information technology, vol. 6, no. 3, pp. 175–183, 2004.
[75]
L. Royakkers and S. D. Zwart, Moral responsibility and the problem of many hands. Routledge, 2015.