Align AI to Dynamic Human-AI Workflows

Valerie Chen\(^{1}\), Cleotilde Gonzalez\(^{1}\), Anita Williams Woolley\(^{1}\), Michael Lee\(^{2}\), Tongshuang Wu\(^{1}\), Vincent Conitzer\(^{1}\), Aarti Singh\(^{1}\)
\(^{1}\)Carnegie Mellon University \(^{2}\)University of California, Irvine


1 Introduction↩︎

Consider a developer using an AI coding assistant. The developer initially relies on suggestions provided by the AI assistant to move quickly, but after encountering incorrect outputs, the developer becomes more cautious, preferring explanations and verifying results more carefully. As the assistant proves reliable again, their reliance increases and supervision might reduce.

This example of human-AI collaboration reflects not only changing preferences but also evolving trust and reliance shaped by a user’s interaction history with an AI system. It also highlights how the AI assistant serves not just to emulate human-like behavior; rather, humans and AI should coordinate, adapting roles to leverage complementary strengths [1]. Building AI systems that can work effectively with people, therefore, requires accounting for these dynamics. Despite these interaction-driven dynamics, most alignment techniques for AI models focus on emulating or mimicking humans using static preferences or demonstrations (e.g., as in preference-based RL [2], [3], dueling bandits [4], behavior cloning [5][7]). This paradigm persists in modern alignment methods, which primarily target conversational behavior and verbal outputs through preference optimization [8][11]. For example, reinforcement learning from human feedback (RLHF) optimizes models using stand-alone comparisons [12].

As agentic AI systems are increasingly taking actions in the world and integrating into human workflows [13][18], this shift highlights a fundamental gap in how alignment is currently defined and evaluated. It is not evident that the kinds of static and emulative alignment for conversational behavior and verbal outputs will prove effective or appropriate for more agentic systems [19]. In this paper, we argue that this gap requires moving towards methods that capture longitudinal interactions (Figure 1). Doing so raises a set of open challenges: what does it mean for AI systems to complement human collaborators and how can such behaviors be optimized in practice?

Figure 1: Rethinking AI Alignment for Human-AI Workflows. Current alignment paradigms optimize models using static preferences to emulate human outputs. This paper argues for shifting from snapshot evaluation and optimization toward a dynamic, complementary view of alignment, where the goal is to optimize the behavior of human-AI systems over sequences of interaction.

Because these challenges are not well understood solely from a machine learning perspective, we draw on insights synthesized from a cross-disciplinary workshop spanning machine learning and the social sciences. First, we draw on prior social science literature on human-human collaboration as a foundation for understanding how trust and complementarity emerge in multi-agent interaction. We then discuss how human-AI collaboration extends these foundations but under new conditions: AI systems amplify asymmetries in capability and information in ways that are often difficult for users to interpret, while lacking the social and institutional cues that support trust and accountability in human teams. AI systems can also take on new roles that reshape how joint work is structured. Finally, we identify new challenges for alignment in interaction that are unique to human-AI collaboration.

In this paper, we argue for moving beyond optimizing snapshot model performance to optimizing trajectories of human-AI systems over time, drawing on interdisciplinary perspectives to better characterize the dynamics of human-AI interaction. While we introduce foundational considerations for aligning AI systems in human workflows, many open questions remain. To chart concrete paths forward, we identify key conceptual, translational, and evaluation challenges, and conclude with takeaways for the human-AI alignment community.

2 Background↩︎

Table 1: Standard alignment methods optimize static objectives and omit key properties of collaborative interaction.Here, \(x\) denotes an input, \(y\) a model output, \(\pi_\theta\) a learned policy, and \(r_\phi\) a learned reward model. We contrast these methods with a collaborative objective over trajectories \(\tau\), where reward depends on interaction history \(\tau_{<t}\) and joint human-AI outcomes.
Method Objective
Behavior cloning [5][7] \(\max_{\pi_\theta} \log \pi_\theta(y^H \mid x)\)
[0.4pt/1.5pt]
Preference-based alignment [2], [9], [11], [12] \(\max_{\pi_\theta} \mathbb{E}_{y \sim \pi_\theta(\cdot \mid x)} [r_\phi(x,y)]\)
[0.4pt/1.5pt]
Dueling bandits [4] \(\max_{\pi_\theta} \mathbb{E}_{(y,y') \sim \pi_\theta} \left[\mathbf{1}(y \succ y')\right]\)
[0.4pt/1.5pt]
Collaborative alignment \(\max_{\pi_\theta} \mathbb{E}_{\tau}\left[\sum_t r^*(x_t, a_t^H,a_t^{AI};\tau_{<t})\right]\)
[0.4pt/1.5pt]

2.1 The Limitations of Current Alignment Approaches↩︎

We first review the dominant paradigm underlying many modern alignment approaches. Let \(x \in \mathcal{X}\) denote an input, such as a prompt or context, and let \(y \in \mathcal{Y}\) denote a model output. Alignment is often operationalized through a preference function \(r^*(x,y)\) that assigns a scalar value reflecting human judgment of the output.

In practice, \(r^*\) is not directly observed, but is approximated using human feedback, such as ratings or pairwise comparisons over outputs. This yields a learned reward model \(\hat{r}(x,y)\), which is then used to optimize a policy \(\pi_\theta(y \mid x)\): \[\max_{\pi_\theta} \;\mathbb{E}_{x \sim \mathcal{D}, \, y \sim \pi_\theta(\cdot \mid x)} \left[ \hat{r}(x, y) \right].\]

This formulation has been productive for improving individual model outputs, but it embeds several assumptions that become limiting in real-world human-AI systems (see Table 1). First, preferences are treated as static, without incorporating interaction history. Second, alignment is defined at the level of individual outputs rather than as a collaborative process. Third, the objective is specified in terms of a reward over model outputs alone, rather than a joint reward over human-AI interaction. As a result, the human is primarily represented as a source of preference labels, while the AI system is optimized to emulate or satisfy an inferred preference function.

2.2 Formalizing Alignment for Collaborative Settings↩︎

To address the limitations of current approaches, we move from isolated inputs and outputs to modeling sequential interaction between agents. Let the interaction trajectory be \[\tau = (x_1, a_1^A, a_1^B, x_2, a_2^A, a_2^B, \dots, x_T,a_T^A, a_T^B),\] where \(x_t\) denotes the observable interaction context at time \(t\), and \(a_t^A, a_t^B\) are the actions taken by two collaborators. At each step, actions are conditioned on the preceding interaction history \(\tau_{<t}\) as well as latent collaborator state \(z_t\), such as goals, beliefs, expertise, or mental models that may not be directly observable [20], [21]. Rather than optimizing for a fixed input \(x\), this formulation models collaboration as an evolving process in which behavior depends on both observable interaction context and unobserved internal state over time.

In this setting, reward is also defined as a function of how the interaction unfolds over time: \[r_t = r^*(x_t, a_t^A, a_t^B; \tau_{<t}).\]

We can also capture complementarity between collaborators, which arises when \[r^*(x_t,a_t^A, a_t^B; \tau_{<t}) > \max \big( r^*(x_t,a_t^A, \emptyset; \tau_{<t}), \; r^*(x_t,\emptyset, a_t^B; \tau_{<t}) \big).\] This formulation captures both human-human and human-AI collaboration.1 In human-human settings, both agents adapt to each other through interaction, but their behaviors are not directly controllable or optimized by a system designer. In human-AI settings, however, one of the agents is a parameterized system with policy \(\pi_\theta\), where: \[a_t^{AI} \sim \pi_\theta(\cdot \mid \tau_{<t}),\] and the policy \(\pi_\theta\) can be optimized.

In human-AI collaboration, the AI system is a parameterized system with policy \(\pi_\theta\) that can be optimized: \[\max_{\pi_\theta} \;\mathbb{E}_{\tau \sim \pi_\theta} \left[ \sum_{t=1}^T r^*(x_t,a_t^H, a_t^{AI}; \tau_{<t}) \right],\] where the trajectory distribution depends jointly on human and AI actions. We concretize our discussion using a well-known example: consider a developer working with an AI coding assistant. The human actions \(a_t^H\) may include prompting, verifying outputs, or editing generated code, while AI actions \(a_t^{AI}\) may include generating code or expressing uncertainty before proceeding.

We observe that optimizing \(\pi_\theta\) is not simply a matter of maximizing immediate user preference because an action that is preferred in the moment may still produce poor downstream trajectories, for example if it increases overreliance on the AI system [25][27]. Conversely, actions such as expressing uncertainty or requesting clarification may improve the long-run human-AI workflow even if they are locally less preferred. For AI-assisted coding, the relevant outcome is therefore not only whether users accept generated code, but whether the joint workflow produces correct, maintainable code while preserving appropriate human oversight [28].

3 Interdisciplinary Perspectives on Human-AI Alignment↩︎

To make progress on building AI systems that can take actions that complement users in the long term, we argue that this research agenda cannot be addressed by the machine learning community alone. It requires integrating complementary perspectives from fields such as human-computer interaction, cognitive science, and organizational psychology, which have long studied how humans collaborate over time and are beginning to shed light on the nuances of human-AI collaboration [29][32]. We therefore convened an interdisciplinary workshop to answer these questions.

3.1 Methodology↩︎

The theme of the workshop convening was creating flexible Human-AI teams to achieve complementarity, which called for participants interested in the interdisciplinary study of how to design and deploy AI systems in ways that are dynamically aligned with human values, robust to unexpected behavior, and safe even under failure modes. We now discuss the details of the workshop held in September 2025:

Participants. Participation was determined through an application process intended to select a diverse group of attendees in terms of disciplinary background, methodological expertise, career stage, and application domain. In total, 70 participants attended the workshop, representing 43 institutions across 8 countries. Participants included professors, postdoctoral researchers, industry research scientists, and graduate students. Their backgrounds were roughly split evenly between machine learning and social sciences.

Pre-workshop survey. Prior to the workshop, all participants completed a survey used to identify discussion topics for in-workshop activities. The survey focused on two dimensions: future directions and barriers to progress.

Workshop procedure. The workshop took place over two days. The first half focused on tutorials, talks, and posters, establishing shared context across disciplines. The second half focused on structured brainstorming and discussion. Each topic was led by two experts with complementary expertise spanning machine learning and social or decision sciences. Participants were divided into six groups with roughly balanced disciplinary representation, and groups rotated across topics to ensure broad participation.

Presentation of findings. The findings presented in this paper are synthesized from the second day of the workshop, including topic discussions, structured rotations, in-depth brainstorming sessions, and written summaries from facilitators and participants. Across these activities, several recurring themes emerged. Participants repeatedly pointed to insights from human-human collaboration as a foundation for understanding effective interaction, which we draw on to identify key principles. They also highlighted ways in which these principles change in human-AI settings. Our findings are summarized in Figure 2.

3.2 Lessons from Human-Human Interaction↩︎

We use human-human collaboration as the starting point because social science has long studied how humans coordinate with one another. The workshop discussion repeatedly returned to two bodies of work that are especially relevant for dynamic and complementary alignment: research on trust as an evolving relational state, and research on team cognition as a mechanism for coordinating complementary expertise [33], [34].

Lesson 1: A user’s trust in their collaborator is multidimensional and changes over interactions. In collaborative settings, individuals must decide whether and how much to rely on a partner, despite not fully knowing their partner’s full competence, intentions, or future behavior. Trust is commonly referred to as a willingness to accept vulnerability under these types of uncertainty [35]. Prior literature shows that trust is multidimensional, comprising beliefs about a collaborator’s competence, benevolence, and integrity [36][38]. Perceived competence primarily determines whether individuals delegate tasks or accept a partner’s outputs, while perceived benevolence and integrity shape whether they are willing to share information in the collaborative process [39]. Behavioral experiments show individuals will typically continue to rely on partners who are capable but occasionally make mistakes, but will rapidly withdraw when they infer deception or misaligned intent [40][43].

Trust also evolves through interactions and is updated based on observed behavior, feedback, and repair attempts [36], [41], [44]. Computational and cognitive models characterize trust as emerging through repeated interaction, reinforcement, expectation formation, and reciprocal adaptation [45][47]. Longitudinal studies of team interactions show that early impressions quickly give way to experience, with consistent performance strengthening reliance and failures triggering reassessment [40].

Figure 2: From Human-Human Collaboration to Human-AI Alignment. We draw on literature studying human teamwork and collaboration and identify ways in which existing principles change in human-AI settings.

Lesson 2: Collaboration depends on complementary expertise, but realizing these gains requires active coordination. While trust governs when and how individuals rely on a collaborator over time, a separate challenge is identifying how collaborators can effectively coordinate and divide labor to leverage complementary expertise. This process requires coordination [48][51] and, in practice, groups often fail to do this well. For example, classic studies of group decision-making show that teams tend to focus on information that is already shared among members, while neglecting unique, and potentially critical, information held by individuals [52], [53]. Related work shows that coordination may lead to production blocking (i.e., because group members cannot contribute simultaneously, some ideas are delayed, forgotten, or never voiced) [54], [55], evaluation apprehension (i.e., one’s hesitation to share ideas that may be judged) [54], [55], and premature consensus (i.e., early convergence without full exploration) [56].

Effective teams overcome these challenges by developing shared structures that guide reliance. One key mechanism is transactive memory, or a shared understanding of “who knows what” within the group [57], [58]. This allows members to route questions, divide labor, and access expertise without redundant effort, enabling more efficient coordination [59][61]. For example, in software teams, developers often specialize in different components of a system, and effective collaboration depends on knowing which teammate to consult or defer to for a given issue [62], [63]. A complementary mechanism is the development of shared mental models, which align members’ expectations about tasks, roles, and procedures [64][66]. These shared expectations reduce the need for constant communication and enable smoother coordination, particularly in dynamic or high-stakes settings [34], [67], [68]. In aviation and military teams, shared mental models help crew members anticipate handoffs, monitor threats, and adapt under time pressure [69], [70]; in medical trauma teams, shared expectations about roles and protocols help specialists coordinate rapidly while still adapting when the case violates routine assumptions [71].

3.3 Revisiting Lessons in Human-AI Collaboration↩︎

We revisit lessons from human-human collaboration to examine how these principles are reshaped in the context of human-AI collaboration.

Extension 1: Trust must be formed under amplified and less interpretable asymmetries. Human-AI workflows amplify the challenge of reasoning about uncertainty in collaboration because asymmetries between collaborators are larger and qualitatively different. AI systems may differ from humans in speed, scale, memory, statistical pattern recognition, confidence calibration, and vulnerability to distribution shift [16], [18], [72], [73]. While users often recognize that AI systems behave differently from human collaborators, prior work shows that they struggle to form accurate mental models of system behavior and to calibrate reliance appropriately [74]. For example, explanations and decision aids are intended to make model behavior more transparent, yet they do not reliably help users determine when a system will succeed or fail, and can lead to over-reliance or under-utilization [75][78].

More broadly, the absence of familiar social and institutional cues makes AI systems harder to interpret as collaborators. For example, studies of human-AI interaction show that users often form incomplete or incorrect mental models of system behavior, leading them to misinterpret outputs or rely on inappropriate heuristics when deciding whether to trust the system [79], [80]. In practice, this can manifest as users treating confidence signals or fluent outputs as reliable evidence of correctness, even when these signals are difficult to calibrate across tasks and interaction histories [79], [81]. At the same time, AI systems lack the reputational histories, role expectations, and accountability structures that help people contextualize errors in human collaborators, making unexpected failures harder to interpret or recover from [25], [26], [82], [83]. As a result, efficient outputs can mask uncertainty, brittleness, or dataset-specific artifacts that users cannot easily anticipate or inspect.

Extension 2: Realizing complementarity requires new forms of coordination with AI. As in human-human collaboration, complementarity in human-AI collaboration requires deciding which tasks should be performed by the human, which should be performed by the AI system and how uncertainty should be communicated [23], [84], [85]. AI systems also change the coordination problem because they can shape both the work product and the process by which people coordinate around it. In production roles, AI can draft, retrieve, summarize, classify, write code, and check work, changing the distribution of cognitive labor across a workflow [13], [14], [86], [87]. In coordination roles, AI has the potential to maintain shared representations of evidence, assumptions, model limitations, data provenance, and decision rationales, making it easier for participants to see what has been considered and what remains uncertain [17], [88][90]. These roles parallel transactive memory and shared mental models in human teams, but they require explicit design [29][31], [72], [90].

Figure 3: Bridging Barriers to Dynamic Human-AI Alignment. Participants in the workshop noted three core barriers to move from static, output-centric views of alignment toward a trajectory-level perspective and recommended a set of takeaways.

4 Toward Dynamic, Complementary Alignment↩︎

To identify paths toward progress, we first examine participants’ perceived barriers to advancing alignment for human-AI collaboration. We then discuss existing building blocks from the machine learning community that can help address these challenges, allowing us to synthesize a set of takeaways and future directions for the human-AI alignment community, summarized in Figure 3.

4.1 Barriers to Progress↩︎

Conceptual barriers. The most common barrier identified by participants was that there is insufficient theory of human-AI interaction. One contributing factor towards this might be that insights currently exist across human-computer interaction, cognitive science, and organizational psychology, which is also exemplified through the diversity of participants from computer science and social science who attended the workshop. Compared to more established areas such as social and organizational psychology, which have decades of theory on human teamwork and coordination, research on human-AI interaction is relatively nascent [33], [34], [91]. Participants also noted that this challenge is compounded by the rapid evolution of AI systems themselves.

Translation barriers. A second cluster of barriers reflected challenges in grounding human-AI alignment research in real-world settings. Participants emphasized that understanding effective human-AI collaboration often requires studying systems in deployment to surface issues that were previously not present in model-centered evaluation [92], [93]. While there are increasingly many examples of such systems in practice, these deployments are frequently concentrated in private or industry settings, where interaction data are not publicly accessible. Participants also highlighted gaps in interdisciplinary collaboration between theoretical and application-grounded perspectives [94].

Evaluation barriers. A third set of barriers fell under the broad umbrella of shared evaluation infrastructure. Participants highlighted the lack of standardized metrics, limited access to high-quality interaction datasets and interactive testbeds, and the scarcity of longitudinal studies as major obstacles to progress. While growing efforts have emphasized documentation and accountability practices for datasets and models [88], [89], [95], participants noted that comparable infrastructure for studying human-AI interaction remains underdeveloped.

4.2 Existing Building Blocks in ML Literature↩︎

The barriers identified by workshop participants do not imply that the machine learning community is starting from scratch:

  • Multi-Agent Reinforcement Learning. Multi-agent RL studies how interacting agents learn to coordinate, communicate, and divide responsibilities under shared or mixed objectives [21], [96], [97]. Emerging work on collaborative interaction optimization extends this framing to language-model systems that proactively elicit intent and optimize long-term interaction success [98]. These directions introduce important concepts for modeling complementary expertise and interaction dynamics, but most existing work assumes simulated agents, well-specified objectives, or bounded interaction settings, leaving many challenges of real-world human-AI collaboration unresolved.

  • Preference Learning. Modern alignment methods learn from human preferences over model outputs, enabling systems to optimize for helpfulness, harmlessness, and instruction following [8], [10][12]. More recent extensions, including multi-turn RLHF, expand these objectives from isolated responses to trajectories of interaction [99]. However, these approaches still primarily optimize conversational preferences and assistant behavior, rather than the broader dynamics of collaboration that emerge in human-AI workflows over time.

  • Human Data Requirements. Recent work increasingly curates datasets and benchmarks from real user interactions, including large-scale chatbot logs, software engineering tasks derived from GitHub issues, and coding-agent interaction traces [100][104]. These resources make evaluation more grounded in deployment data, but they still primarily capture task distributions, conversation traces, or agent capabilities, rather than directly validating whether model behavior improves longitudinal human-AI collaboration in real workflows.

4.3 Directions for the human-AI alignment community↩︎

We posit that the central challenge of building towards dynamic, complementary alignment is not the absence of relevant methods, but rather the absence of formulations that ground them in human modeling and social science insights:

Addressing conceptual barriers. Prior work on human teamwork shows that collaboration depends on evolving trust, complementary expertise, coordination, and shared representations developed over interaction. These observations suggest that existing ML methods may need to be reformulated around richer models of collaborative interaction. For example, a fruitful direction is to extend preference learning methods beyond explicit judgments over isolated outputs toward interaction-level properties such as calibrated reliance, coordination, and adaptation over time. This also raises new challenges around supervision and data collection, including whether such signals should be elicited directly from users or inferred implicitly from interaction behavior and workflow trajectories.

Similarly, existing multi-agent reinforcement learning formulations often optimize shared task reward without explicitly modeling complementary expertise or evolving shared understanding between collaborators. One promising direction is to incorporate ideas from transactive memory and shared mental models into these formulations by learning representations over collaborator capabilities, responsibilities, and uncertainty that evolve throughout interaction. At the same time, applying these methods to human-AI collaboration introduces practical challenges, as reinforcement learning methods are often highly data-intensive, while human interaction data are expensive to collect and constrained by deployment settings.

Addressing translation barriers. Many of the interaction phenomena discussed earlier only emerge through sustained interaction in realistic workflows. We argue that progress requires closer collaboration across communities that study different aspects of human-AI systems. Machine learning research contributes methods and evaluation frameworks, HCI and cognitive science contribute theories of interaction and human behavior, organizational research contributes insights about coordination and teamwork in practice, and industry deployments provide access to real-world workflows where these dynamics can be observed at scale. Creating stronger mechanisms for these perspectives to inform one another can ground future alignment research more directly in realistic collaborative settings.

Addressing evaluation barriers. Recent datasets and benchmarks increasingly draw from real user interactions, including chatbot logs, software engineering workflows, and agent trajectories, representing important progress toward grounding evaluation in deployment data. Prior work from social and decision sciences suggests that effective collaboration is often reflected not only in whether a task is completed successfully, but in how collaborators adapt to one another over time, recover from failures, allocate responsibilities, and develop shared understanding throughout interaction. Future evaluation frameworks should therefore explore metrics that are more directly grounded in theories of collaboration to better understand when strong model performance translates into effective human-AI collaboration in practice, along with characterizing the effect of different training paradigms.

5 Alternative Views↩︎

We discuss potential counterarguments and objections to dynamic human-AI alignment:

Prioritize alignment for autonomous AI systems. A first alternative view is that static alignment of AI systems is limited but indispensable. This view would argue that the community primarily relies on snapshot evaluations because they provide scalability and reproducibility, not because the community ignores the value of long-term interaction and complementarity [105]. A related perspective is that alignment should focus on dynamic settings in which AI systems may operate autonomously over long time horizons or primarily interact with other AI systems. Our position is that human-AI alignment does not replace static alignment or other alignment objectives around autonomous AI systems. Rather, it complements them by focusing on an increasingly important class of real-world workflows in which humans and AI systems jointly perform work over extended periods.

Complementarity is a workflow design problem, not an alignment problem. A second alternative view is that research on human-AI complementarity belongs primarily to HCI, human factors, or organizational design. As AI systems are deployed in agentic and interactive contexts, we argue that this boundary is increasingly difficult to maintain. In these growing deployments, the model behavior can alter what users attend to, when they verify, how they delegate, and how work is reorganized [17], [18]. Recent studies demonstrate how human-AI teams fail to outperform the best human-only or AI-only baseline unless the collaboration is deliberately structured [24], [106]. Thus, achieving robust complementary collaboration may require not only redesigning the surrounding workflow, but also fundamentally rethinking how models are trained, optimized, and evaluated for interactive human-AI settings.

Optimizing for human-AI alignment creates safety concerns. A final view is that optimizing human-AI interactions creates new risks. If an AI system is rewarded for downstream workflow outcomes, it may shape user beliefs, trust, attention, or reliance in ways that users did not explicitly endorse [25], [42], [43]. For example, an AI system might learn to withhold information or induce skepticism by nudging a user’s behavior toward an ulterior objective under the justification of improving long-run performance [107]. Dynamic alignment therefore introduces fundamental research challenges that should be treated as central to the broader alignment agenda rather than as reasons to avoid studying these systems altogether. In our view, the existence of these risks strengthens, rather than weakens, the case for rigorous academic study.

6 Conclusion↩︎

We propose a shift from static, emulative alignment toward interactive and complementary alignment. We contrast standard alignment objectives with a trajectory-level formulation in which human and model behavior co-evolve, and ground this perspective in insights from an interdisciplinary workshop. First, we draw on social science accounts of human collaboration, where we identify key dynamics such as evolving trust and coordination, and show how these are extended and transformed in human-AI settings. We further highlight how human-AI systems introduce new asymmetries, uncertainty, and coordination challenges. Finally, we distill barriers and takeaways for the human-AI alignment community, outlining implications for designing and evaluating human-AI workflows.

References↩︎

[1]
C. Gonzalez and H. Heidari, “A cognitive approach to human–ai complementarity in dynamic decision-making,” Nat. Rev. Psychol., vol. 4, no. 12, pp. 808–822, Oct. 2025.
[2]
C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz, “A survey of preference-based reinforcement learning methods,” Journal of Machine Learning Research, vol. 18, no. 136, pp. 1–46, 2017.
[3]
Y. Xu, R. Wang, L. Yang, A. Singh, and A. Dubrawski, “Preference-based reinforcement learning with finite-time guarantees,” Advances in Neural Information Processing Systems, vol. 33, pp. 18784–18794, 2020.
[4]
Y. Yue and T. Joachims, “Interactively optimizing information retrieval systems as a dueling bandits problem,” in Proceedings of the 26th annual international conference on machine learning, 2009, pp. 1201–1208, doi: 10.1145/1553374.1553527.
[5]
D. Pomerleau, ALVINN: An autonomous land vehicle in a neural network,” in Advances in neural information processing systems, 1989, pp. 305–313.
[6]
B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and Autonomous Systems, vol. 57, no. 5, pp. 469–483, 2009.
[7]
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics (aiSTATS), 2011, pp. 627–635.
[8]
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Advances in neural information processing systems, 2017.
[9]
D. M. Ziegler et al., “Fine-tuning language models from human preferences,” arXiv preprint arXiv:1909.08593, 2019.
[10]
N. Stiennon et al., “Learning to summarize with human feedback,” in Advances in neural information processing systems, 2020, vol. 33, pp. 3008–3021.
[11]
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” in Advances in neural information processing systems, 2023.
[12]
L. Ouyang et al., “Training language models to follow instructions with human feedback,” Advances in neural information processing systems, vol. 35, pp. 27730–27744, 2022.
[13]
M. H. Jarrahi, “Artificial intelligence and the future of work: Human-ai symbiosis in organizational decision making,” Business Horizons, vol. 61, no. 4, pp. 577–586, 2018.
[14]
A. Maedche et al., ai-based digital assistants,” Business & Information Systems Engineering, vol. 61, no. 4, pp. 535–544, 2019.
[15]
K. C. Kellogg, M. A. Valentine, and A. Christin, “Algorithms at work: The new contested terrain of control,” Academy of Management Annals, vol. 14, no. 1, pp. 366–410, 2020.
[16]
B. Shneiderman, “Human-centered artificial intelligence: Reliable, safe & trustworthy,” International Journal of Human–Computer Interaction, vol. 36, no. 6, pp. 495–504, 2020.
[17]
E. Horvitz, “Principles of mixed-initiative user interfaces,” in Proceedings of the SIGCHI conference on human factors in computing systems, 1999, pp. 159–166, doi: 10.1145/302979.303030.
[18]
R. Parasuraman, T. B. Sheridan, and C. D. Wickens, “A model for types and levels of human interaction with automation,” IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans, vol. 30, no. 3, pp. 286–297, 2000, doi: 10.1109/3468.844354.
[19]
S. Z. Shen et al., “Scaling collaborative effort with agents,” in Findings of the association for computational linguistics: ACL 2026, 2026, pp. 12254–12271.
[20]
N. C. Rabinowitz, F. Perbet, H. F. Song, C. Zhang, S. M. A. Eslami, and M. Botvinick, “Machine theory of mind,” in Proceedings of the 35th international conference on machine learning, 2018, vol. 80, pp. 4218–4227.
[21]
M. Carroll et al., “On the utility of learning about humans for human-ai coordination,” in Advances in neural information processing systems, 2019, vol. 32.
[22]
G. Bansal, B. Nushi, E. Kamar, D. S. Weld, W. S. Lasecki, and E. Horvitz, “Does the whole exceed its parts? The effect of ai explanations on complementary team performance,” in Proceedings of the SIGCHI conference on human factors in computing systems, 2021, pp. 1–16.
[23]
C. Rastogi, L. Leqi, K. Holstein, and H. Heidari, “A taxonomy of human and ML strengths in decision-making to investigate human-ML complementarity,” in Proceedings of the AAai conference on human computation and crowdsourcing, 2023, vol. 11, pp. 127–139.
[24]
M. Vaccaro, A. Almaatouq, and T. W. Malone, “When combinations of humans and ai are useful: A systematic review and meta-analysis,” Nature Human Behaviour, vol. 8, no. 12, pp. 2293–2303, 2024.
[25]
R. Parasuraman and V. Riley, “Humans and automation: Use, misuse, disuse, abuse,” Human Factors, vol. 39, no. 2, pp. 230–253, 1997.
[26]
J. D. Lee and K. A. See, “Trust in automation: Designing for appropriate reliance,” Human Factors, vol. 46, no. 1, pp. 50–80, 2004.
[27]
E. L. Bjork and R. A. Bjork, “Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning,” Psychology and the real world: Essays illustrating fundamental contributions to society, vol. 2, no. 59–68, pp. 56–64, 2011.
[28]
H. Mozannar et al., Expert Certification“The RealHumanEval: Evaluating large language models abilities to support programmers,” Transactions on Machine Learning Research, 2025, [Online]. Available: https://openreview.net/forum?id=hGaWq5Buj7.
[29]
L. A. Suchman, Plans and situated actions: The problem of human-machine communication. Cambridge University Press, 1987.
[30]
J. Hollan, E. Hutchins, and D. Kirsh, “Distributed cognition: Toward a new foundation for human-computer interaction research,” ACM Transactions on Computer-Human Interaction, vol. 7, no. 2, pp. 174–196, 2000, doi: 10.1145/353485.353487.
[31]
S. Amershi et al., “Guidelines for human-ai interaction,” in Proceedings of the 2019 CHI conference on human factors in computing systems, 2019, pp. 1–13, doi: 10.1145/3290605.3300233.
[32]
C. Gonzalez et al., “Toward a science of human–ai teaming for decision making: A complementarity framework,” PNAS Nexus, vol. 5, no. 3, p. pgag030, Mar. 2026, doi: 10.1093/pnasnexus/pgag030.
[33]
S. W. J. Kozlowski and D. R. Ilgen, “Enhancing the effectiveness of work groups and teams,” Psychological Science in the Public Interest, vol. 7, no. 3, pp. 77–124, 2006.
[34]
M. A. Marks, J. E. Mathieu, and S. J. Zaccaro, “A temporally based framework and taxonomy of team processes,” Academy of Management Review, vol. 26, no. 3, pp. 356–376, 2001.
[35]
D. M. Rousseau, S. B. Sitkin, R. S. Burt, and C. Camerer, “Not so different after all: A cross-discipline view of trust,” Academy of Management Review, vol. 23, no. 3, pp. 393–404, 1998.
[36]
R. C. Mayer, J. H. Davis, and F. D. Schoorman, “An integrative model of organizational trust,” Academy of Management Review, vol. 20, no. 3, pp. 709–734, 1995.
[37]
F. D. Schoorman, R. C. Mayer, and J. H. Davis, “An integrative model of organizational trust: Past, present, and future,” Academy of Management review, vol. 32. Academy of Management Briarcliff Manor, NY 10510, pp. 344–354, 2007.
[38]
A. Legood, L. Van Der Werff, A. Lee, D. Den Hartog, and D. Van Knippenberg, “A critical review of the conceptualization, operationalization, and empirical literature on cognition-based and affect-based trust,” Journal of Management Studies, vol. 60, no. 2, pp. 495–537, 2023.
[39]
K. T. Dirks and D. L. Ferrin, “The role of trust in organizational settings,” Organization Science, vol. 12, no. 4, pp. 450–467, 2001.
[40]
D. J. McAllister, R. J. Lewicki, and S. Chaturvedi, “Trust in developing relationships: From theory to measurement.” in Academy of management proceedings, 2006, vol. 2006, pp. G1–G6.
[41]
R. J. Lewicki and C. Wiethoff, “Trust, trust development, and trust repair,” The handbook of conflict resolution: Theory and practice, vol. 1, no. 1, pp. 86–107, 2000.
[42]
B. J. Dietvorst, J. P. Simmons, and C. Massey, “Algorithm aversion: People erroneously avoid algorithms after seeing them err,” Journal of Experimental Psychology: General, vol. 144, no. 1, pp. 114–126, 2015.
[43]
J. M. Logg, J. A. Minson, and D. A. Moore, “Algorithm appreciation: People prefer algorithmic to human judgment,” Organizational Behavior and Human Decision Processes, vol. 151, pp. 90–103, 2019, doi: 10.1016/j.obhdp.2018.12.005.
[44]
M. Yu, M. Saleem, and C. Gonzalez, “Developing trust: First impressions and experience,” Journal of Economic Psychology, vol. 43, pp. 16–29, 2014.
[45]
C. Gonzalez, N. Ben-Asher, J. M. Martin, and V. Dutt, “A cognitive model of dynamic cooperation with varied interdependency information,” Cognitive science, vol. 39, no. 3, pp. 457–495, 2015.
[46]
I. Juvina, M. Saleem, J. M. Martin, C. Gonzalez, and C. Lebiere, “Reciprocal trust mediates deep transfer of learning between games of strategic interaction,” Organizational Behavior and Human Decision Processes, vol. 120, no. 2, pp. 206–215, 2013.
[47]
J. L. Harman, J. O’Donovan, T. Abdelzaher, and C. Gonzalez, “Dynamics of human trust in recommender systems,” in Proceedings of the 8th ACM conference on recommender systems, 2014, pp. 305–308.
[48]
I. D. Steiner, Group process and productivity. New York: Academic Press, 1972.
[49]
J. R. Hackman, Leading teams: Setting the stage for great performances. Boston, MA: Harvard Business School Press, 2002.
[50]
T. W. Malone and K. Crowston, “The interdisciplinary study of coordination,” ACM Computing Surveys, vol. 26, no. 1, pp. 87–119, 1994.
[51]
K. Crowston, “A coordination theory approach to organizational process design,” Organization Science, vol. 8, no. 2, pp. 157–175, 1997.
[52]
G. Stasser and W. Titus, “Pooling of unshared information in group decision making: Biased information sampling during discussion,” Journal of Personality and Social Psychology, vol. 48, no. 6, pp. 1467–1478, 1985.
[53]
G. M. Wittenbaum, G. Stasser, and C. J. Merry, “Tacit coordination in anticipation of small group task completion,” Journal of Experimental Social Psychology, vol. 32, no. 2, pp. 129–152, 1996.
[54]
M. Diehl and W. Stroebe, “Productivity loss in brainstorming groups: Toward the solution of a riddle,” Journal of Personality and Social Psychology, vol. 53, no. 3, pp. 497–509, 1987.
[55]
B. Mullen, C. Johnson, and E. Salas, “Productivity loss in brainstorming groups: A meta-analytic integration,” Basic and Applied Social Psychology, vol. 12, no. 1, pp. 3–23, 1991.
[56]
I. L. Janis, Groupthink, 2nd ed. Boston, MA: Houghton Mifflin, 1982.
[57]
D. M. Wegner, “Transactive memory: A contemporary analysis of the group mind,” in Theories of group behavior, B. Mullen and G. R. Goethals, Eds. Springer, 1987, pp. 185–208.
[58]
K. Lewis, “Measuring transactive memory systems in the field: Scale development and validation,” Journal of Applied Psychology, vol. 88, no. 4, pp. 587–604, 2003.
[59]
D. W. Liang, R. Moreland, and L. Argote, “Group versus individual training and group performance: The mediating role of transactive memory,” Personality and Social Psychology Bulletin, vol. 21, no. 4, pp. 384–393, 1995.
[60]
Y. Ren and L. Argote, “Transactive memory systems 1985–2010: An integrative framework of key dimensions, antecedents, and consequences,” Academy of Management Annals, vol. 5, no. 1, pp. 189–229, 2011.
[61]
L. Argote and Y. Ren, “Transactive memory systems: A microfoundation of dynamic capabilities,” Journal of Management Studies, vol. 49, no. 8, pp. 1375–1382, 2012.
[62]
S. Faraj and L. Sproull, “Coordinating expertise in software development teams,” Management Science, vol. 46, no. 12, pp. 1554–1568, 2000.
[63]
J. D. Herbsleb and A. Mockus, “An empirical study of speed and communication in globally distributed software development,” IEEE Transactions on Software Engineering, vol. 29, no. 6, pp. 481–494, 2003.
[64]
R. Klimoski and S. Mohammed, “Team mental model: Construct or metaphor?” Journal of Management, vol. 20, no. 2, pp. 403–437, 1994.
[65]
J. E. Mathieu, T. S. Heffner, G. F. Goodwin, E. Salas, and J. A. Cannon-Bowers, “The influence of shared mental models on team process and performance,” Journal of Applied Psychology, vol. 85, no. 2, pp. 273–283, 2000.
[66]
J. A. Cannon-Bowers and E. Salas, “Reflections on shared cognition,” Journal of Organizational Behavior, vol. 22, no. 2, pp. 195–202, 2001.
[67]
A. C. Edmondson, “Psychological safety and learning behavior in work teams,” Administrative Science Quarterly, vol. 44, no. 2, pp. 350–383, 1999.
[68]
A. W. Woolley, C. F. Chabris, A. Pentland, N. Hashmi, and T. W. Malone, “Evidence for a collective intelligence factor in the performance of human groups,” Science, vol. 330, no. 6004, pp. 686–688, 2010.
[69]
G. F. Goodwin, N. Blacksmith, and M. R. Coats, “The science of teams in the military: Contributions from over 60 years of research.” American Psychologist, vol. 73, no. 4, p. 322, 2018.
[70]
E. Salas, N. J. Cooke, and M. A. Rosen, “On teams, teamwork, and team performance: Discoveries and developments,” Human Factors, vol. 50, no. 3, pp. 540–547, 2008.
[71]
S. Faraj and Y. Xiao, “Coordination in fast-response organizations,” Management Science, vol. 52, no. 8, pp. 1155–1169, 2006.
[72]
A. Dafoe, Y. Bachrach, G. Hadfield, E. Horvitz, K. Larson, and T. Graepel, “Cooperative ai: Machines must learn to find common ground,” Nature, vol. 593, no. 7857, pp. 33–36, 2021.
[73]
L. Bainbridge, “Ironies of automation,” Automatica, vol. 19, no. 6, pp. 775–779, 1983, doi: 10.1016/0005-1098(83)90046-8.
[74]
D. A. Norman, “The problem with automation: Inappropriate feedback and interaction, not over-automation,” Philosophical Transactions of the Royal Society of London. B, Biological Sciences, vol. 327, no. 1241, pp. 585–593, 1990, doi: 10.1098/rstb.1990.0101.
[75]
G. Bansal, B. Nushi, E. Kamar, W. S. Lasecki, D. S. Weld, and E. Horvitz, “Beyond accuracy: The role of mental models in human-ai team performance,” Proceedings of the AAai Conference on Human Computation and Crowdsourcing, vol. 7, pp. 2–11, 2019.
[76]
Z. Buçinca, M. B. Malaya, and K. Z. Gajos, “To trust or to think: Cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making,” Proceedings of the ACM on Human-computer Interaction, vol. 5, no. CSCW1, pp. 1–21, 2021.
[77]
T. Miller, “Explanation in artificial intelligence: Insights from the social sciences,” Artificial Intelligence, vol. 267, pp. 1–38, 2019.
[78]
V. Chen, Q. V. Liao, J. Wortman Vaughan, and G. Bansal, “Understanding the role of human intuition on reliance in human-AI decision-making with explanations,” Proc. ACM Hum.-Comput. Interact., vol. 7, no. CSCW2, Oct. 2023, doi: 10.1145/3610219.
[79]
P. K. Kahr, G. Rooks, M. C. Willemsen, and C. C. Snijders, “Understanding trust and reliance development in ai advice: Assessing model accuracy, model explanations, and experiences from previous interactions,” ACM Transactions on Interactive Intelligent Systems, vol. 14, no. 4, pp. 1–30, 2024.
[80]
S. Dhuliawala, V. Zouhar, M. El-Assady, and M. Sachan, “A diachronic perspective on user trust in ai under uncertainty,” in Proceedings of the 2023 conference on empirical methods in natural language processing, 2023, pp. 5567–5580.
[81]
Z. Li and M. Steyvers, “Learning to trust: How humans mentally recalibrate ai confidence signals,” arXiv preprint arXiv:2603.22634, 2026.
[82]
K. A. Hoff and M. Bashir, “Trust in automation: Integrating empirical evidence on factors that influence trust,” Human Factors, vol. 57, no. 3, pp. 407–434, 2015.
[83]
E. Glikson and A. W. Woolley, “Human trust in artificial intelligence: Review of empirical research,” Academy of Management Annals, vol. 14, no. 2, pp. 627–660, 2020.
[84]
H. Mozannar and D. Sontag, “Consistent estimators for learning to defer to an expert,” Proceedings of the International Conference on Machine Learning, pp. 7076–7087, 2020.
[85]
P. Hemmer, M. Schemmer, N. Kühl, M. Vössing, and G. Satzger, “Complementarity in human-ai collaboration: Concept, sources, and evidence,” European Journal of Information Systems, vol. 34, no. 6, pp. 979–1002, 2025.
[86]
S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer, “The impact of ai on developer productivity: Evidence from github copilot,” arXiv preprint arXiv:2302.06590, 2023.
[87]
A. Ziegler et al., “Measuring github copilot’s impact on productivity,” Communications of the ACM, vol. 67, no. 3, pp. 54–63, 2024.
[88]
M. Mitchell et al., “Model cards for model reporting,” in Proceedings of the conference on fairness, accountability, and transparency, 2019, pp. 220–229.
[89]
M. Pushkarna, A. Zaldivar, and O. Kjartansson, “Data cards: Purposeful and transparent dataset documentation for responsible ai,” in Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, 2022, pp. 1776–1826.
[90]
G. Klein, P. J. Feltovich, J. M. Bradshaw, and D. D. Woods, “Common ground and coordination in joint activity,” Organizational Simulation, vol. 53, pp. 139–184, 2005.
[91]
S. M. Fiore and T. J. Wiltshire, “Technology as teammate: Examining the role of external cognition in support of team cognitive processes,” Frontiers in Psychology, vol. 7, p. 1531, 2016.
[92]
K. Holstein, J. Wortman Vaughan, H. Daumé III, M. Dudik, and H. Wallach, “Improving fairness in machine learning systems: What do industry practitioners need?” in Proceedings of the 2019 CHI conference on human factors in computing systems, 2019, pp. 1–16, doi: 10.1145/3290605.3300830.
[93]
N. Sambasivan, S. Kapania, H. Highfill, D. Akrong, P. Paritosh, and L. M. Aroyo, Everyone wants to do the model work, not the data work: Data cascades in high-stakes ai,” in Proceedings of the 2021 CHI conference on human factors in computing systems, 2021, pp. 1–15, doi: 10.1145/3411764.3445518.
[94]
J. Grudin, “Ai and HCI: Two fields divided by a common focus,” ai magazine, vol. 30, no. 4, pp. 48–48, 2009.
[95]
T. Gebru et al., “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021, doi: 10.1145/3458723.
[96]
J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson, “Learning to communicate with deep multi-agent reinforcement learning,” in Advances in neural information processing systems, 2016.
[97]
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in Advances in neural information processing systems, 2017.
[98]
S. Wu et al., “CollabLLM: From passive responders to active collaborators,” in International conference on machine learning, 2025.
[99]
K. Lu et al., “Multi-turn reinforcement learning from preference human feedback,” arXiv preprint arXiv:2405.08448, 2024.
[100]
L. Zheng et al., “LMSYS-chat-1M: A large-scale real-world LLM conversation dataset,” in International conference on learning representations, 2024.
[101]
W. Zhao, X. Ren, J. Hessel, et al., “WildChat: 1M ChatGPT interaction logs in the wild,” arXiv preprint arXiv:2405.01470, 2024.
[102]
C. E. Jimenez et al., “SWE-bench: Can language models resolve real-world GitHub issues?” in International conference on learning representations, 2024.
[103]
J. Baumann, V. Padmakumar, X. Li, J. Yang, D. Yang, and S. Koyejo, “SWE-chat: Coding agent interactions from real users in the wild,” arXiv preprint arXiv:2604.20779, 2026.
[104]
V. Chen et al., “How can we assess human-agent interactions? Case studies in software agent design,” arXiv preprint arXiv:2510.09801, 2025.
[105]
D. Kiela et al., “Dynabench: Rethinking benchmarking in NLP,” in Proceedings of the 2021 conference of the north american chapter of the association for computational linguistics: Human language technologies, 2021, pp. 4110–4124, doi: 10.18653/v1/2021.naacl-main.324.
[106]
G. Bansal, B. Nushi, E. Kamar, E. Horvitz, and D. S. Weld, “Is the most accurate ai the best teammate? Optimizing ai for teamwork,” Proceedings of the AAai Conference on Artificial Intelligence, vol. 35, no. 13, pp. 11405–11414, 2021.
[107]
S. Emmons, C. Oesterheld, V. Conitzer, and S. Russell, “Observation interference in partially observable assistance games,” arXiv preprint arXiv:2412.17797, 2024.

  1. We note that prior work on human-AI collaboration only considers the one-time-step notion of complementarity [22][24].↩︎