Thought Experiments for Conceptual Work: A New Application of a (Very) Old Method


Abstract

In this paper, we propose thought experiments (TEs) as a crucial method for Human-Computer Interaction (HCI) researchers to engage in conceptual work. As an interdisciplinary field, HCI often uses concepts as fundamental building blocks for larger theories. However, the conceptual commitments we make in this process carry normative consequences. TEs are a well-established philosophical method, whereby a hypothetical but tractable scenario logically progresses to a conclusion. We outline TEs as an interrogative method that brings conceptualizations to their normative implications through logical moves. We illustrate the value of thought experiments through two examples: (1) original thought experiments to critique stakeholders in Value-Sensitive Design and (2) Helen Nissenbaum’s use of thought experiments to generate contextual integrity. We discuss how TEs precisely anticipate the potential harms of technologies, allowing HCI to operationalize current calls for increased scrutiny of research ethics and broader impacts.

<ccs2012> <concept> <concept_id>10003120.10003121.10003126</concept_id> <concept_desc>Human-centered computing HCI theory, concepts and models</concept_desc> <concept_significance>500</concept_significance> </concept> </ccs2012>

1 Introduction↩︎

Human-Computer Interaction (HCI) research relies on concepts. From the idea of “the user” to what constitutes “interaction,” concepts serve as guideposts for HCI—anchoring abstract thinking [1]; guiding methods [2]; and informing design [3]. Over time, concepts evolve through reasoned debate within the field [4][6]. For example, “the user” is a central concept in HCI and initially described a single person directly interacting with a system. Now the concept encompasses generalized and complex notions of how people use technology [4], [7], [8].

Adopting a particular definition for some concept constitutes making a conceptual commitment—decisions about the underlying constructs that influence theoretical and empirical investigations. These commitments can be explicit, such as formulating the user as a single individual interacting with a system [9]. More often than not, however, conceptual commitments are unstated and underpin ongoing work. For example, [10] found that the concept of the “human” in human-centered machine learning is multi-faceted but often implicitly defined by researchers. Indeed, ample work in HCI challenges the conceptual commitments we make in our practice [10], e.g., [11], [12][14].

Conceptual commitments carry normative consequences. Normative consequences are moral or values-based outcomes that prescribe what is good, desirable, or what we ought to do [15], [16]. We can describe ought-ness in terms of logical consequence (e.g., because I commit to accessibility, I “ought” to use 14pt font), or as moral claims about what is right or wrong (e.g., one ought not to steal)1. For HCI, normative consequences are the politics, values, and ethical outcomes of our decision-making [17].

If left unchecked, the normative consequences of conceptual commitments can contribute to tangible harm. For example, work in feminism [18][20] and critical race theory [11], [21] for HCI describes how our concepts of “who matters” are deeply steeped in societal biases and impact our designs. Such work historically examines the societal and practical consequences of HCI, and then seeks to identify the conceptual commitments that give rise to them.

Our paper starts from an inverse perspective: what if we could evaluate the costs of conceptual commitments in HCI (perhaps a priori), and logically inspect the practical and societal consequences that the commitment causes? Concepts are powerful building blocks in HCI. Therefore, a failure to understand how our conceptual commitments bear on research decisions risks allowing unethical practices to go unchecked. By unpacking our concepts—the granular components of HCI—we can understand where our normative senses and subsequent research practices come from. The field needs an interrogative methodology for these conceptual commitments: precise reasoning that takes a conceptual commitment to its logical (and often normative) consequence [22].

We contribute thought experiments for HCI: a method for HCI researchers to deductively identify conceptual commitments and their normative consequences. Thought experiments are a subclass of philosophical arguments [23] that rely on building hypothetical scenarios to explore a logical problem. They are a tool for deductive reasoning; when using a thought experiment, the author(s) precisely toggle experimental variables and explicitly work through logical steps to arrive at a conclusion. This method advances our understanding of how normative consequences relate to the conceptual commitments underlying our theories. In other words, thought experiments answer the question: if we adopt this particular conceptual commitment, what normative consequences necessarily follow?

In this paper, we first outline the necessary components of a thought experiment. Then, we introduce evaluative criteria for assessing the rigor of a thought experiment—tractability, validity, and soundness. We demonstrate three practical applications of thought experiments: (1) interrogating the Value-Sensitive Design formulation of stakeholder, (2) testing alternative formulations of stakeholder, and (3) generating formulations necessary for contextual integrity. We illustrate the application of thought experiments by using them to interrogate stakeholderism in HCI, a concept with implications for human-centeredness, justice, and other normative criteria [10], [24], [25]. The formulation of stakeholder in VSD determines who gets considered during the design process and, thus, what constitutes a problem. We also show how thought experiments can be generative for HCI by charting Helen Nissenbaum’s [26] use of thought experiments to develop contextual integrity, a well-known theory of privacy used in HCI [27][29]. We discuss how researchers and practitioners can use thought experiments to further HCI’s goals, such as more precise broader impact statements and ethics critiques in peer review. In short, we offer thought experiments as a reliable, rigorous, and precise method for researchers to interrogate the normative consequences of conceptual commitments.

2 Related Work↩︎

2.1 Background and Overview of Thought Experiments↩︎

Thought experiments have a rich history in diverse disciplines, from Achilles and the Tortoise [30] to Schrödinger’s cat [31]. Scientists have equally used thought experiments to support propositions or to refute them. Notable examples include Galilieo’s Falling Bodies Problem [32], Einstein’s Magnet and Conductor Experiment [33], [34], and Newton’s Bucket [35], [36]. In moral philosophy, thought experiments are often used to interrogate moral stances rather than scientific propositions. The Trolley Problem2 interrogates utilitarianism [37] (albeit a simplified version of the position), and the Violinist Argument similarly interrogates a utilitarian perspective regarding abortion rights [38]. Thought experiments have been heavily leveraged in computational philosophy and related fields. For example, in 1714, Leibniz proposed the Mill Argument to refute the claim that purely material things like machines can think or perceive [39]. This perspective later evolved into computationalism in artificial intelligence, the idea that mental states and computational states are analogous [40]. Over 200 years after Leibniz, John Searle proposed the Chinese Room thought experiment (elaborated in Figure 2) to refute the computationalists of his time [41].

Given this diversity of uses, philosophers have not yet coalesced on a single definition of thought experiments. In the 1990s, the power of thought experiments to create knowledge was widely disputed. [42]—a noted empiricist—argued that thought experiments are a specific subclass of argument with unique explanatory power but no epistemic power. Norton writes, “if thought experiments produce new information about the world it is because they bring to light the hidden consequences of and relations between facts that we already know” [p. 3].

Similarly, thought experiments explain the moral or physical consequences of our intuitions but are not a method for assessing the truth of those initial intuitions. Alternatively, [32] argues that thought experiments can moderate epistemological power by scoping what is logically available. Poetically, [43] called thought experiments “a priori science that happens in the laboratory of the mind” [p. 1] claiming that thought experiments hold the same epistemic power as physical ones.

In this paper, we adopt Norton’s Argument View that thought experiments are a particular type of argument designed to explain hidden logical consequences [23], [34], [44]. It is important to note that this deviates from colloquial definitions of thought experiments as speculative games. We frame thought experiments in alignment with philosophy, their discipline of origin. Thought experiments are a formal interrogative method that reveals the implications of conceptualizations. In this way, thought experiments build on complementary prior work in HCI, as described below.

2.2 Complementary Work in HCI↩︎

2.2.1 Conceptual Work↩︎

Like many social sciences, HCI welcomes both empirical and theoretical contributions (Figure 1). Empirical contributions generate quantitative or qualitative descriptions of phenomena in the world. Meanwhile, theoretical contributions give us architectures of concepts (e.g., frameworks, taxonomies, and models) for understanding our findings. Together, the two inform each other [46].

This synergy between theory and empiricism lies on an invisible foundation of conceptual work. As articulated by [47], concepts are the building blocks of theory. For example, articulating how we conceptualize the human in human-centered machine learning (HCML) is central to building robust theories of data ethics [10], [48]. Therefore, conceptual work must precede theoretical or empirical work.

A useful concept X has boundary conditions such that we know with some confidence that Y is an instance of X, but Z is not [49]. Conceptual work is the disambiguation of concepts through definitions, comparisons, or distinctions. [2] describe conceptual work as the “problematizing” stage of HCI research, noting that conceptual problems are second-order: they do not pertain to the world directly nor to our observations or experiences thereof. They argue that papers that propose theories are a form of conceptual contribution. We build on this position of theory generation as a form of conceptual work by shifting focus one step deeper: clarifying and testing consequences of the concepts themselves. HCI has an abundance of theories, many of which are translated from other disciplines, such as feminist theory [18], critical race theory [11], and stakeholder theory [25]. Therefore, the underlying concepts—feminism, power, stakeholder—must be similarly translated. HCI adopts these concepts from their original domains without necessarily interrogating the consequences of conceptual borrowing. We argue that this lack of conceptual methods contributes to a gap in normative reasoning in HCI.

We propose thought experiments as a necessary conceptual method in HCI because they elucidate the edge cases of our conceptual commitments. Here, a weak concept is one where absurd logical conclusions are permissible [49], [50]; the conceptual commitment is not robust enough to avoid preposterous but logically valid conclusions. In other words, if a conceptualization allows for obviously absurd conclusions, it indicates there may be a fundamental flaw in the concept itself. In this paper, we specifically show how weak conceptualizations of the stakeholder allow for absurd design choices, such as including clear bad actors (e.g., for-profit prison owners) in the design process. The upshot here is not that researchers are making these choices; it is that they do not have rigorous concepts to justify current practices. The normative implications of our conceptual commitments do not align with our normative intuitions.

2.2.2 Normative Reasoning↩︎

In philosophy, normative reasoning is the dimension of thinking concerned with “oughts” [51]. This is distinct from normative reasoning in social and behavioral sciences, which focuses on behaviors that are normalized through societal adoption (e.g., hetero-normativity [52]). Throughout this paper, we use normative in the philosophical sense.

Normative reasoning in philosophy encompasses ethical and moral reasoning. While the distinction between ethics and morality is a long-standing discourse, this paper adopts the stance that ethics defines right and wrong in an environment [15]. Meanwhile, morality defines right and wrong for an individual [53]. The intersection of ethics and morals is the sense of right and wrong that informs our actions. This paper focuses on normative statements (i.e., “oughts”) that lie in that intersection. Specifically, we focus on the implicit oughts that inform design decisions.

HCI has a rich history of ethical and moral reasoning. Prior work has explored ethics—how we ought to do research within our environments—in terms of operationalizing “isms.” For example, consequentialism involves thinking of harms and benefits [54], [55], while abolitionism involves interrogating the fundamental purpose behind institutions [21], [56], [57]. This important work lies in the space between theory and empiricism; it involves translating theory into calls to action for the field and its artifacts. Meanwhile, morality in HCI is discussed as researcher positionality. Recent work has called for increased researcher positionality statements [58][60] to surface how our moral positions impact research decisions.

In this paper, we take the stance that conceptual positions carry normative consequences. When we make a decision about a concept, we also bring implications for what is morally permissible. Yet, these consequences are difficult to anticipate without the proper methods. We propose thought experiments as an accessible method for HCI practitioners to engage with normative reasoning that often goes unchecked in HCI.

2.2.3 Distinction from Speculative Methods↩︎

HCI researchers and practitioners may be inclined to draw parallels between thought experiments and speculative methods, given that the latter also involves reasoning about hypothetical scenarios. Indeed, some prior work has claimed that speculative methods such as design fiction can be conceived of as thought experiments e.g., [61], [62][64]. However, we caution against such comparisons. Thought experiments and speculative methods leverage hypotheticals in entirely different ways for entirely different purposes.

Speculative methods like futuring [65], design fiction [66], and design probing [67] invite us to envision alternative realities and futures, deriving inspiration from storytelling and fiction [68]. They often focus on being provocative [69], critical [70], or radically imaginative [66]. By contrast, thought experiments use hypotheticals to create logical guardrails for testing a proposition deductively [23], [32], [35]. This logical interrogation, originating from a separate tradition rooted in philosophy, allows us to precisely evaluate propositions, including propositions about concepts.

To illustrate this difference, we can look to the futures cone: a conical projection of preferable, probable, plausible, and possible futures [71], [72]. In speculative methods, participants are tasked with traversing the time axis and building hypothetical worlds [73]. In thought experiments, there is no concept or categorization of futures. The author picks an imagined scenario that captures the relevant dimensions, regardless of its likelihood as a possible future. Once the author posits a hypothetical, there is no world-building, only deducing. The crucial contribution of a thought experiment is the logical conclusion one makes, not the hypothetical world one built. While superficial similarities may exist, looking beneath the surface reveals that speculative methods and thought experiments are highly dissimilar in their purposes and intellectual origins.

2.3 Our cases↩︎

2.3.1 Stakeholderism in HCI↩︎

In 1984, Freeman introduced the stakeholder concept in Organizational Sciences as a new way of framing ideal corporate strategy for firms [74]. Stakeholder theory asserts that the firm is a network of value-creation obligations for those who can affect or be affected by the realization of an organization’s purpose (the wide definition) or those without whose support the organization would not exist (the narrow definition) [75]. In the past 40 years, stakeholder theory has become widely accepted and iterated on in organizational research. For example, stakeholder theory has grounded corporate social responsibility [75], business ethics [76], and non-financial value creation [77]. Moreover, stakeholder theory transformed how managers and corporations are evaluated.  [78] notes that, under stakeholder theory, firms should be evaluated by “the ability of its managers to create sufficient wealth, value, or satisfaction for those who belong to each stakeholder group."

In HCI and CSCW, stakeholder theory has been heavily adapted for systems and applications. In this sense, the firm is a system or application, and a stakeholder is someone who should be considered in the design process. Similar to organizational research, the term stakeholder carries moral weight; when we fail to include stakeholders in the design process, we are skirting a moral obligation to do right by them [24], [79], [80]. Researchers agree that it is imperative to better understand the term stakeholder [25], [81], [82]. Previous work has shown that excluding specific stakeholder considerations from the design process can harm vulnerable populations. For example, [83] found that including parents in predictive child-welfare algorithm design fosters a richer understanding of the algorithm’s potential shortfalls. [10] found that predictive mental health researchers have not coalesced on a conceptualization of stakeholders; data subjects are conceptualized as a myriad of different roles, from patients to social media users. This discord can cause dramatic differences in data handling procedures and limit human agency in sociotechnical systems [84].

This paper uses thought experiments to interrogate a widely accepted conceptualization of stakeholders in HCI through Value-Sensitive Design (VSD). VSD states that understanding relevant community stakeholders’ motivations, values, and goals is an essential first step of the design process [85]. Subsequently, VSD conceptualizes stakeholders as anyone who either (1) interacts with a system or output or (2) is affected by a system or output. In VSD, (1) is a direct stakeholder, while (2) is an indirect stakeholder. However, previous work has noted limitations of this conceptualization.  [4] note that stakeholder-system relationships do not follow straightforward patterns. [86] note that researcher positionality is often understated in VSD research papers. Furthermore, identifying “key” stakeholders relies on researcher rationale that is often overlooked [87]. We build on these critiques by proposing various expansions of VSD and alternative conceptualizations. We use thought experiments to demonstrate that these alternatives better achieve the moral imperatives of stakeholderism in HCI.

2.3.2 Privacy and Contextual Integrity in HCI↩︎

In 2004, Helen Nissenbaum presented Contextual Integrity (CI) as a tool for understanding privacy [26]. CI asserts that adequate privacy is provided when an information flow conforms to both implicit and explicit norms surrounding information sharing. CI has since been formalized into five parameters: (1) the information subject; (2) the sender; (3) the recipient; (4) the information type; and (5) the transmission principles that represent the norms of information flow [88]. A contextual integrity analysis requires a practitioner to define each of these parameters in relation to the context they are assessing. Notably, CI is a large departure from previous privacy theories, such as privacy as control [89], secrecy [90], or information security [91].

While not originally published in HCI venues, Nissenbaum and her peers have inspired radical change to both HCI research and privacy policy surrounding data sharing. For example, [29] leveraged contextual integrity and its metaphors to guide ethical decision-making for research using big data. Policy-oriented institutions, such as the Federal Trade Commisssion, have cited Nissenbaum’s work as grounds for new laws around data privacy [92]. Moreover, CI has sparked numerous open questions. [59] ask if contextual integrity can be applied to online mental health communities that may not want to be discovered. [27] ask if contextual integrity is still relevant to inferential data.

We elaborate on Nissenbaum’s methodology for generating contextual integrity. We highlight the thought experiments that Nissenbaum and her peers used throughout their iterations of CI. By formalizing Nissenbaum’s implicit use of thought experiments into our explicit structure, we show that thought experiments can be conceptually generative—giving us formulations that serve as the building blocks for larger theories.

3 How to Conduct a Thought Experiment↩︎

In philosophy, there is ample discourse around the different taxonomies, implications, and specifics of thought experiments [23], e.g. [32], [35], [93]. These discussions rely on a shared understanding of thought experiments to guide a proper philosophical analysis [94]. In this section, we explain common features of building and evaluating a thought experiment. We propose a definition of thought experiments that is both true to the centuries of philosophy research and applicable to HCI: thought experiments are a type of argument that (1) posit a hypothetical and (2) invoke experimental conditions as a means to interrogate the normative implications of conceptual commitments [23]. Thought experiments require both the author and the reader to adopt a philosophical stance—a space between the intuitive and empirical where experimental hypotheticals are crucial reasoning devices [45].

3.1 Components of Thought Experiments↩︎

As [35] articulates, engaging with a thought experiment involves these general steps:

  1. Postulating a proposition or counterfactual

  2. Visualizing some situation that we have set up in the imagination

  3. Carrying out an operation to see what happens

  4. Drawing a conclusion

Each of these steps relies on a rhetorical component:

  1. Proposition. A claim or stance the author puts forth to motivate the thought experiment [42].

  2. Analogy. The imagined hypothetical that acts as a physical setting for the subsequent “experiment." [95]

  3. Experimental Variable. An experimental dimension the author varies to further their argument [32].

  4. Conclusion. A claim about the initial proposition that logically follows from the thought experiment.

A complete thought experiment adopts a stance and reasons about the implications of changing specific dimensions within a hypothetical world. Thus, thought experiments take a proposition to its logical conclusion. In this paper, we present thought experiments in both prose and formal logic for comprehensiveness. However, expertise in formal logic or the technical logistics of philosophical argumentation is not necessary to produce a high-quality thought experiment. Many famous scientists have presented seminal thought experiments in prose only (e.g., John Searle’s Chinese Room Thought Experiment [41]).

3.1.1 Example.↩︎

In 1980, [41] proposed the Chinese Room thought experiment as a refutation of his computationalist contemporaries. At the time, AI hype was causing optimism and panic. By the late 1970s, some AI researchers claimed that computers already understood at least some natural language [96]. To this day, Searle’s Chinese Room is heralded as a seminal use of thought experiments in computational fields [97]. In this section, we walk through Searle’s Chinese Room [41] to exemplify how to formally conduct this common thought experiment.

Proposition. To start, we seek to refute a computational conceptualization of the mind: the human mind is a program being run on the brain in the same way that software is a program that gets run on a computational processing unit. In popular discourse, this implies the existence of Strong AI, a computer that, given unlimited resources, could replicate every task of the human brain. More formally,

Proof.

  1. The mind is a program run on an extremely complex machine (i.e, the brain) (Definition of Mind)

  2. Given the right resources, a computer can perform every task the human mind can(Implication)

 ◻

Analogy. Imagine we, a native English speaker who knows no Chinese, are asked to converse with a native Chinese speaker who knows no English. However, we are in a room (henceforth, The Chinese Room) that has buckets of Chinese symbols and a book of instructions for how to manipulate the symbols (see Figure 2).

Suppose our conversant is outside this room, passing us questions written in Chinese. By properly following the instruction manual, we can manipulate this input and produce a sequence of Chinese characters that correctly answers their question. We can infinitely continue this process of engaging in “conversation.” However, Searle notes that this conversation does not demand that we (the human) understand Chinese [41]. Under this computational representation of conversation, a human can exist in the Chinese Room and never learn a single word of Chinese.

Experimental Variable. Suppose instead of placing ourselves in the Chinese Room, we place an unlimited processor in the room. Note the parallels in our initial analogy to a computer-based situation. Instead of a room with buckets of symbols, there is a database of Chinese characters; instead of an instruction manual, the processor is given a program for transforming Chinese. Essentially, the Chinese Room is now a computer. We can then picture the conversation task as a Chinese speaker passing in a question input and receiving an answer output.

Conclusion. In our original analogy, we do not arrive at a greater understanding of Chinese. In fact, we could answer an infinite number of questions from the conversant and not understand a single word of Chinese. However, we would be simulating conversation by accurately answering Chinese questions.

This conclusion holds in our computer-based case as well. If we do not understand Chinese based on implementing the appropriate program for speaking Chinese, neither does any other digital computer solely on that basis. Note that the computer, by our conceptualization of the mind, does not have anything the human does not have.

At this point, we have reached a logical contradiction: there is a state of the mind (i.e., understanding Chinese) that a functionally described machine, either human or  computational, could never have. This contradicts our original conceptualization of the human mind as a sophisticated program run on the brain. By definition, programs are input manipulation tasks. However, the Chinese Room experiment demonstrates that one could manipulate symbols for eternity and never reach understanding. Therefore, there exists a human mental state that is impossible to describe under a computational conceptualization of the mind. Subsequently, the computational conceptualization of the mind does not hold. In formal logic, the Chinese Room argues:

Proof.

  1. \(mind = program(Brain)\)(Definition of Mind)

  2. \(\forall \text{ tasks } t,\ mind \text{ can complete }t \rightarrow computer \text{ can complete }t\) (Implication)

  3. \(Let\ t = \text{understanding Chinese}\)

  4. \(mind \text{ can complete }t\)

  5. \(computer \text{ cannot complete }t\)

  6. \(\exists \text{ task } t \text{ such that } mind \text{ can complete }t \land computer \text{ cannot complete }t\) (Contradiction of 2)

  7. \(mind \neq program(Brain)\)

 ◻

3.1.2 Scoping Thought Experiments.↩︎

As with empirical experiments, thought experiments have guardrails that scope the realm of imagined possibility. In speculative design methods, these guardrails have been described as the “futures cone”: a conical projection of preferable, probable, plausible, and possible outcomes [72]. Similarly, the author must define and defend a reasonable realm of possibility in thought experiments. For physics thought experiments, such as Galileo’s Falling Bodies problem, the realm of possibility is all previously proven physics principles. In Searle’s Chinese Room, one can only communicate by passing notes off Chinese characters. These guardrails help make “instinctive knowledge" [95] (i.e., intuitions) a permissible explanatory device. Because the author has built a world with specific dimensions, they have scoped what is logically available to the reader. This technique allows readers to interrogate how their intuitions interact with this hypothetical scenario. The realm of possibility in thought experiments is crucial because it allows readers to rely on their intuitions while also interrogating them.

3.2 Evaluating a Thought Experiment↩︎

Like most logical arguments, a high-quality thought experiment must engage in high-quality reasoning. Building on formative work about thought experiments [35], [44], [95], we describe three evaluative criteria for thought experiments: tractability, validity, and soundness. Tractability evaluates whether the hypothetical is appropriate for the argument being made. Meanwhile, validity and soundness evaluate the logical moves themselves.

3.2.1 Evaluating the Analogy.↩︎

Transforming a problem into a new solution space necessitates creating a tractable version of the initial problem statement. However, this transformation creates the potential for faulty representations that strong-arm or strawman a problem statement into an inappropriate solution space. [12] describe this in HCI as “solving a computationally tractable solution to the problem rather than the problem itself.” Thought experiments suffer from a similar problem: authors can build analogies that do not parallel the real dilemma being interrogated. We clarify and explain how to evaluate thought experiments by contrasting two principles—tractability and realizability [95].

Tractability describes whether the imagined scenario properly interrogates the proposition without begging for a trivial logical argument. Tractability is crucial for useful thought experiments with tangible implications. However, it is important to distinguish between tractability versus realizability when evaluating thought experiments. While the imagined scenario should be relevant to the real world, it does not have to perfectly mirror a real-world experimental setup. Examining tractability in thought experiments is a common mechanism for critically evaluating them. For example, critics of the Chinese Room have noted that modern computer programs can interact with the real world, which could potentially lead to understanding (see “The Robot Reply" [40], [98], [99]). Therefore, a computer in a windowless room is not tractable to modern programs, so the thought experiment’s conclusion is irrelevant to the original proposition. Searle has responded to these critics by expanding his thought experiment to include real-world interaction and still reaching similar conclusions [100]. Note that avid criticism in philosophy is often the sign of a provocative argument rather than a highly flawed one, so debates like these do not indicate that Searle’s argument should be rejected.

Realizability—the ability for a thought experiment to be replicated in a world free of physical, ethical, time, or financial limitations—is not necessary for a useful thought experiment. [95] argues that physical unrealizability need not be a defining condition of thought experiments. Unrealizable amounts of money or fictitious characters are not reasons why thought experiments would be rejected; preposterous exaggerations may capture a concept well. Mach describes this as the core of taking a philosophical stance: sitting in the space between logic and instinctive knowledge. Therefore, thought experiments should not be evaluated on how well they would replicate in a condition-less world. Rather, they should be evaluated on how well they parallel certain ideas and give them understandable constructs. For example, Searle represents the idea of computational omniscience with a program that can fluently converse in Chinese.

For an exaggerated example of a thought experiment analogy that makes a conclusion trivially true, we could change the task in Searle’s Chinese Room from conversing in Chinese to a non-language-based task. Suppose instead we change the input to several digits \(n\) and the output to the \(n\)-th digit of Pi. The human in the room is given a book with all of the digits of Pi (note that while physically impossible because Pi is infinite, this is logically plausible). In essence, instead of Searle’s Chinese Room we have Searle’s Pi Room. In this scenario, there is no deeper level of understanding that is available to the human. It is not the case that a human can reach a mental state in this task that a computer cannot. Therefore, we would no longer reach a logical contradiction from our initial conceptualization of Strong AI.

Searle’s Pi Room is a poorly designed thought experiment because the task of finding a digit of Pi is an inherently computational task. Therefore, the conclusion that a computer and human would complete it analogously is trivial. There is no notion of “understanding Pi" in the same way that there is a notion of”understanding Chinese." Therefore, this analogy is untractable; humans obviously complete a myriad of tasks that are non-computational in nature, such as learning a new language. Searle’s Pi Room does not parallel the real-world dilemma of computationalism in relevant ways, which, in turn, causes it to rely on trivial logic.

3.2.2 Evaluating The Argument.↩︎

Recall that we adopt Norton’s view that thought experiments are a specific type of argument. Therefore, they are subject to the same type of evaluation as other reasoning methods, like proofs and deductive logic. Particularly in philosophy, evaluating logical arguments has a rich history with two evaluation criteria [22], [101], [102]: (1) validity and (2) soundness.

Validity describes the quality of each step in the argument; an argument is valid if the truth of the premises logically guarantees the truth of the conclusion. For example, the Chinese room experiment takes a common logical form, reductio ad absurdum. Searle posits: if Strong AI is true, then any human mind state can be achieved by a computational program. Through a thought experiment, he reaches the logical contradiction that human understanding cannot be simulated by a computational system, whether that system is human or mechanic. In other words, Searle is able to refute the conceptualization of Strong AI by showing that accepting it would lead us to an absurd conclusion. Thought experiments, given their hypothetical nature, provide the tools needed to set up an illustrative case that might not occur in the real world. These cases are often edge cases that are difficult to reason about in informal ways. By taking logical steps and ending up at an absurd conclusion, Searle demonstrates a flaw in the initial conceptualization.

An invalid thought experiment means that the author has used an unjustified logical move to progress the thought experiment. For an exaggerated example, Searle’s logic would be invalid if he had used the claim (1) “I could run a program for Chinese without thereby coming to understand Chinese" to deduce (2)”therefore, no computer can converse in Chinese." Jumping from claim (1) to claim (2) would require that Searle establishes an equivalence between “understanding" and”conversing." However, this claim would undermine the initial setup that one can converse in Chinese if given the right tools. Therefore, this logical jump would make the argument invalid.

Soundness describes whether the initial premises are, in fact, true. Critiquing the soundness of a thought experiment necessitates critiquing whether the initial premises accurately represent the conceptualization being interrogated. Thought experiments rely on formalizing arguments in the aether into logical representations of them.

We identify thought experiments where the initial premise is a conceptual commitment made by translating theory into HCI or a commitment made through a design. In this respect, soundness in thought experiments describes whether the initial premise is a fair, logical representation of the concept or situation. For example, critics of Searle argue that he takes a reductivist conceptualization of computationalism (see “The Intuition Reply" [103], [104]). These critiques are valuable because they allow us to consider the parallels between the thought experiment and popular discourse about concepts. Searle responded to these critics by noting that many prominent figures, like Herbert Simon, already claim that we have machines that can think [105]. This response highlights that Searle’s thought experiment accurately represents at least one flavor of the computationalist position he is interrogating. Therefore, Searle’s Chinese Room is argued to be fairly sound. This example shows that soundness critiques can illuminate useful discussions about the status quo and the pragmatic implications of conceptualizations in popular discourse.

3.2.3 Normative Consequences of Thought Experiments↩︎

It is important to recognize that using thought experiments carries its own normative consequences. As with any research method, engaging with thought experiments presumes a particular epistemic stance. We adopt [44]’s view that thought experiments do not transcend empiricism. For example, replacing an interview study with a thought experiment that speculates, “What would these stakeholders think?” is epistemically invalid. Researchers cannot gain knowledge about people’s perceptions through thought experiments. Thought experiments are designed to answer conceptual research questions that inform theoretical and empirical ones.

Moreover, thought experiments do not create normative guarantees. An argument presented as a thought experiment may be logically valid and sound but have a problematic conclusion. For example, Iris Young notes that the outcome of Rawls’s famous Veil of Ignorance thought experiment is an unacceptable formulation of justice, as it discounts systemic oppression [106]. Put concisely, thought experiments do not necessarily guarantee ethical conclusions. Instead, thought experiments make the implicit both explicit and precise to ground further deliberation.

4 Interrogating the VSD Formulation of Stakeholders↩︎

In this section, we present several thought experiments to interrogate the concept of stakeholders in VSD. First proposed by [85], Value-Sensitive Design (VSD) is a methodology that asserts system designers should understand values early in the design process to build usable and ethical systems. Recall from our Related Work (Section [sec:related-work]) that, under VSD, stakeholders are either direct or indirect based on their relationship to a system or output.

As VSD researchers have noted, this is an extremely general conceptualization of stakeholders. In theory, every human on the planet could be considered an indirect stakeholder [86], [107]. Therefore, as VSD methods have become more popular, researchers have augmented VSD’s initial stakeholder definition, specifying that key stakeholders are people “who are or will be significantly implicated by the technology.” [108]

Since stakeholder prioritization is left to the researcher’s discretion in VSD, it is important to have robust methods that unpack researcher intuitions (see Figure 3). Unchecked intuitions about who gets to be included in the design process have caused severe harm, often to communities that are already marginalized [10], [83], [109], [110].

In this section, we ask: What are the normative consequences of committing to the stakeholderism established in VSD?

We use thought experiments to demonstrate that the stakeholder concept in VSD creates room for problematic normative judgments that do not align with previous work [111].

4.1 Experiment Setup↩︎

We begin by setting up a tractable situation where stakeholderness has significant normative implications.

Proposition. Broken down into premises, VSD states:

Proof.

  1. \(P \in \{stakeholders\} \iff (P \text{ interacts with } T \lor P \text{ is affected by } T)\) (VSD-definition of stakeholder)

  2. \(P \text{'s perspective should be valued } \iff P \in \{stakeholders\}\) (VSD-implication of stakeholder inclusion)

 ◻

Analogy. Suppose we have a technical system \(T\), designed to predict a defendant’s recidivism probability in court. \(T\)’s primary use is for judges to determine proper sentencing for a convicted defendant. We use the social role of person \(P\) as our experimental variable and represent the set of all stakeholders of \(T\) as \(\{stakeholders\}\). We present the first few thought experiments in both prose and formal logic. Recall that formal logic is not necessary for thought experiments. We demonstrate a prose-only thought experiment in Section [sec:dal-te].

4.2 Basic Example↩︎

To start, we provide a straightforward thought experiment that shows VSD’s conceptualization of stakeholder inclusion justifies consulting the primary user (i.e., judges) in the design process.

Experimental Variable. Let \(P\) be a judge who will use \(T\). Given \(P\)’s job as a judge, we know that \(P\) will interact with \(T.\) Therefore, \(P\) meets VSD’s criteria for stakeholder inclusion.

Conclusion. Because \(P\) is a stakeholder of \(T\), \(P\)’s perspective should be valued in designing \(T.\) Written formally,

Proof.

  1. \(P := \text{judge}\)(Given)

  2. \(P \in \{stakeholders\} \iff (P \text{ interacts with } T \lor P \text{ is affected by } T)\) (VSD-definition of stakeholder)

  3. \(P \text{'s perspective should be valued } \iff P \in \{stakeholders\}\) (VSD-implication of stakeholder inclusion)

  4. \(P \text{ interacts with } T\) (From 1)

  5. \(P \in \{stakeholders\}\) (From 2, 4)

  6. \(P\)(From 3, 5)

 ◻

This thought experiment demonstrates that the VSD conceptualization of a stakeholder justifies consulting judges during the design process. However, we can start to reveal more convoluted normative consequences by altering \(P\)’s role.

4.3 Conceptualization is Over-Inclusive↩︎

Next, we show that VSD deviates from normative prescriptions (i.e., calls for what ought to be) by being over-inclusive. Said another way, we use a thought experiment to show that people who, based on prior work [111][113], should be intentionally excluded in the design process are justified stakeholders under VSD. This over-inclusiveness implies that the conceptualization does not align with what ought to be the case. In other words, VSD allows for absurd stakeholder choices. The following thought experiment demonstrates the normative consequence of said commitment:

Let \(P\) be a for-profit prison owner. Given \(P\)’s social role, \(P\) will be heavily affected by \(T.\) For example, if \(T\) over-predicts recidivism, more offenders could be sent to prison, and \(P\)’s capital would increase. Therefore, \(P\) meets VSD’s criteria for stakeholder inclusion. Because \(P\) is a stakeholder of \(T\), \(P\)’s perspective should be valued in designing \(T.\)

Proof.

  1. \(P := \text{for-profit prison owner}\)(Given)

  2. \(P \in \{stakeholders\} \iff (P \text{ interacts with } T \lor P \text{ is affected by } T)\) (VSD-definition of stakeholder)

  3. \(P \text{'s perspective should be valued } \iff P \in \{stakeholders\}\) (VSD-implication of stakeholder inclusion)

  4. \(P \text{ is affected by } T\) (From 1)

  5. \(P \in \{stakeholders\}\) (From 2, 4)

  6. \(P\)(From 3, 5)

 ◻

This thought experiment demonstrates that a for-profit prison owner is a stakeholder. Researchers could justifiably incorporate \(P\)’s perspective into the design process. Moreover, researchers have no normative justification to exclude \(P\) from being a stakeholder. However, valuing \(P\)’s perspective in the design process would be inconceivable. Most would balk at including a for-profit prison owner in a values-sensitive exchange, as prioritizing capital interests has historically caused significant harm and exploitation in prison systems [111], [113]. Note that our current conceptualization of stakeholders has no inclusion or exclusion criteria concerning purely capital interests. Essentially, VSD is over-inclusive; monied claims can justify stakeholdership. However, nothing in the conceptualization makes this hierarchy morally reprehensible.

In this thought experiment, we demonstrate that the VSD concept of a stakeholder is not robust enough to guide the decisions that researchers need to make. In effect, there is a misalignment between the concept’s normative implications and the field’s normative intuitions. The problem here is not that a designer will now think it is moral to include a for-profit prison owner as a stakeholder. Rather, the term “stakeholder” does not properly capture our intuitions on what ought to be. Therefore, the outcome of \(P\) as a for-profit prison owner reveals that we need to bound the conceptualization of stakeholders to separate moral obligations from financial ones.

4.4 Conceptualization is Under-Inclusive↩︎

  Next, we show how VSD simultaneously narrows the definition of stakeholder. Similar to the above thought experiment, we show that VSD’s conceptualization of stakeholders is bounded implicitly rather than explicitly. By committing to VSD’s definition, we draw a boundary that excludes specific stakeholders from consideration.

Let \(P\) be a legitimately elected legislator who votes on whether \(T\) gets used in the judicial system. Suppose, because of their role as a legislator rather than a judge, \(P\) will never interact with \(T.\) Therefore, \(P\) is a stakeholder of \(T\) if and only if \(P\) is affected by \(T.\) Yet legislators, we posit, are not likely to be meaningfully affected by the recidivism prediction system. Therefore, by VSD, \(P\) is not a stakeholder of \(T\). Accepting VSD’s conceptualization of stakeholders precludes us from needing or seeking to consult legislators when designing systems—even with major public policy implications.

We pause to address a potential question: are legislators not affected by the system? For one, legislators are answerable to their constituents—and unhappy constituents are not likely to re-elect them for a second term. Therefore, if a system affects the legislator’s constituents, it involves the legislator by extension. But then, should our calculus change if the legislator has been elected for a life term? If so, we commit ourselves to providing a different justification for why the life-term legislator is affected by the system. Maybe the legislator is a stakeholder because they could one day be subjected to the system’s outputs. But then, is everyone in society a stakeholder? And does stakeholderness scale with the degree to which one is likely to be affected by the system? In that case, a corrupt legislator might be considered a more central stakeholder than a law-abiding one, by virtue of being more likely to be affected by the system.

The normative consequence here can be further elucidated by leveraging one of thought experiments’ major strengths, enabling us to go beyond what is empirically available to the realm of what is logically available. Suppose we elaborate \(P\)’s role such that \(P\) has a relationship to technology \(T\) that does not allow any interaction or effect.

We introduce the divine alien legislator (DAL): a being legitimately elected by the people of Earth to make policy decisions for them. The divine alien legislator lives on a faraway planet several light-years away and, therefore, its livelihood cannot be affected by any system on Earth. We can take this unaffectedness one step further: the DAL is indifferent to what happens on Earth and, therefore, has no personal stake in their own decisions. However, they can hear and provide their divine perspective on any policy questions, including ones relating to the deployment of technological systems. By many reasonable normative accounts, we might conclude that a legitimately elected divine alien legislator is a stakeholder, too—and that we should, therefore, consider their perspective. And yet, accepting VSD’s conceptualization precludes us from reaching this conclusion. In this case, our commitment is that legitimacy does not matter for stakeholdership. The only attributes that matter are the interaction and effect between technology \(T\) and person \(P\). Therefore, if we argue that someone is a stakeholder, we have to contend it by virtue of the system’s impact on them; we cannot appeal to legitimacy.

The DAL thought experiment enables us to consider a perfect case that does not exist in the real world—that of a legislator who is not (indeed, ) affected by a system. In doing so, we can bypass arguments about whether legislators are or are not affected by a system to reveal that VSD emphasizes interaction and effect to exclude other potentially influential factors, such as legitimacy. We can then directly ask whether this normative consequence is one we wish to accept.

5 Testing Alternative Stakeholder Formulations↩︎

The above thought experiments reveal problematic normative consequences of the VSD conceptualization of a stakeholder. Next, we provide two reformulations of a stakeholder and validate these conceptualizations with thought experiments. We continue with the same experimental setup as before, considering a recidivism-risk assessment technology \(T\) and experimenting with the social role of person \(P.\) However, to better align the normative implication of the conceptualization, we adjust our definition of stakeholder inclusion and the implications of being a stakeholder.

5.1 Further Bounding Stakeholders↩︎

Reflecting on \(P\) as a for-profit prison owner or a divine alien legislator, we need a boundary on stakeholder inclusion that excludes those with solely capital interests but includes those who have legitimate influence over \(T\). Fortunately, the stakeholder concept is not original to this paper nor HCI. Therefore, we can turn to previous attempts to bound the concept in Organizational Sciences and use thought experiments to uncover the utility of those boundaries.

Taken from [115], stakeholder legitimacy asks: Does the system have some moral or contractual obligation to the stakeholder? [114] expand this notion of stakeholder legitimacy to create a taxonomy of stakeholders that incorporates three salient dimensions: (1) Legitimacy, (2) Power, and (3) Urgency. They establish that being a stakeholder implies one has at least one of these three dimensions. Moreover, these attributes allow us to create a hierarchy of stakeholder considerations when overlaid on top of each other (see Figure 4). For example, if \(P\) were a researcher on a tight paper deadline studying recidivism predictive systems, they probably have significant urgency but no power or legitimacy. [114] call stakeholders who only have urgency “mosquitoes buzzing in the ears" and note they should not be prioritized over those with legitimacy or power.

We demonstrate a thought experiment with a legitimacy-centered formulation of stakeholders. In scoping whose perspective we value to legitimate stakeholders only, we can appropriately narrow our conceptualization to the following premises for stakeholderism:

Proof.

  1. \(P \in \{legitimate\ stakeholders\} \iff (T \text{ has a moral or contractual obligation to } P)\) (Definition)

  2. \(P \text{'s perspective should be valued } \iff P \in \{legitimate\ stakeholders\}\) (Implication)

 ◻

Now, we can use a thought experiment to test whether this new conceptualization aligns with normative prescriptions. Returning to \(P\) as a for-profit prison owner, we have previously established that \(P\)’s business would be heavily affected by \(T.\) However, \(T\) has no moral obligation to \(P\) beyond the basic moral obligation that \(T\) has to all people, such as protecting basic human rights. Moreover, \(T\) has no contractual obligation to \(P.\) For example, \(P\) is not a client paying for \(T.\) Therefore, \(P\)’s perspective should not be valued in the design process. In formal argumentation:

Proof.

  1. \(P := \text{for-profit prison owner}\)(Given)

  2. \(P \in \{legitimate\ stakeholders\} \iff (T \text{ has a moral or contractual obligation to } P)\) (Definition)

  3. \(P \text{'s perspective should be valued } \iff P \in \{legitimate\ stakeholders\}\) (Implication)

  4. \(T \text{ has no obligation to }P\)(Given)

  5. \(P \notin \{legitimate\ stakeholders\}\)(From 2,4)

  6. \(P\)(From 3, 5)

 ◻

The above thought experiment concludes that a designer has no moral imperative to value the perspective of the for-profit prison owner. Practitioners may consider the stances of those who do not have legitimacy, but now we have a vocabulary and logical argument to evaluate the normative consequences of doing so.

Let us consider the other outstanding case: the divine alien legislator. We have previously established that \(P\) neither interacts with nor is affected by \(T.\) However, \(P\) was legitimately elected by the people to create and iterate upon policy, including policy surrounding the use of \(T\). In fact, \(P\) has been ethically bestowed full decision-making power on whether \(T\) can be used for its intended purpose. Given this scenario, \(P\) is a legitimate stakeholder. Therefore, \(P\)’s perspective should be valued.

Proof.

  1. \(P := \text{divine alien legislator}\)(Given)

  2. \(P \in \{legitimate\ stakeholders\} \iff T \text{ has a moral or contractual obligation to } P\) (Definition)

  3. \(P \text{'s perspective should be valued } \iff P \in \{legitimate\ stakeholders\}\) (Implication)

  4. \(P \text{ was legitimately elected by the people for the explicit purpose of creating policy}\)(Given)

  5. \(P \in \{legitimate\ stakeholders\}\)(From 2, 4)

  6. \(P\)(From 3, 5)

 ◻

By introducing a dimension of legitimacy, we have bounded our conceptualization of stakeholders to solve the under-inclusivity problem revealed when \(P\) is a divine alien legislator. The above thought experiments are examples of how to include bounding attributes into a conceptualization and then further test its normative consequences.

5.2 Adding an Evaluative Normative Theory↩︎

Recall that VSD leaves the onus on the designer to determine key stakeholders—those whose perspectives should be prioritized during the design process. However, HCI has not yet coalesced on robust ways to identify and justify the designer’s choices. Therefore, it is imperative we have methods that allow us to normatively judge stakeholder inclusion reasoning. Critical Race Theory (CRT) for HCI [11] states that researchers are obligated to prioritize the perspectives of systemically marginalized or vulnerable populations (see Figure 5). By justifying the decision to prioritize marginalized voices in design processes, we can ground the following normative position: empirical methods that elucidate the perspectives of minority groups should be preferred to those that forego or minimize those perspectives. We can incorporate CRT into a thought experiment that adds an evaluative layer to stakeholder inclusion decisions.

Let \(P\) be a previously convicted criminal in a minority racial group. Returning to our VSD conceptualization of stakeholder, \(P\) is a stakeholder of \(T\) if and only if \(P\) interacts with \(T\) or is affected by \(T.\) \(P\) is clearly affected by \(T\) as \(T\)’s output directly determines \(P\)’s sentencing. However, what happens if \(P\)’s perspective conflicts with the for-profit prison owner who, under this conceptualization, is also a stakeholder? We need a normative framework to justify our decision on stakeholder hierarchy. We can adopt Critical Race Theory, which states that key stakeholders are those who are part of systemically marginalized communities within \(T\)’s problem space. By this premise, \(P\)’s perspective should be prioritized over other stakeholders because \(P\) is part of a marginalized community. More formally:

Proof.

  1. \(P := \text{previously convicted criminal in a minority racial group}\)(Given)

  2. \(P \in \{stakeholders\} \iff (P \text{ interacts with } T \lor P \text{ is affected by } T)\) (VSD-definition of stakeholder)

  3. \(P \text{'s perspective should be valued } \iff P \in \{stakeholders\}\) (VSD-implication of stakeholder inclusion)

  4. \(P\text{'s perspective should be prioritized if } P \text{ is a systemically marginalized stakeholder}\)(By CRT)

  5. \(P \text{ is affected by } T\) (Given)

  6. \(P \in \{stakeholders\}\) (From 2, 5)

  7. \(P\)(From 3, 6)

  8. \(P\)(From 1, 4)

 ◻

Through this thought experiment, we have justified the decision to prioritize marginalized voices in design processes. Furthermore, this justification explicates the normative framework–Critical Race Theory–that we are adopting. Readers can explicitly understand our philosophical stance rather than tacitly reasoning about it. Incorporating CRT is normative insofar as it describes a moral obligation; when we elaborate on the concept of a key stakeholder with Critical Race Theory, we establish a moral hierarchy of perspectives (see Figure 5). This allows designers to justify design decisions that they made for stakeholder inclusion.

5.3 Evaluating our Thought Experiment↩︎

Next, we evaluate these thought experiments based on our criteria outlined in Section [sec:how-to].

5.3.1 Analogy.↩︎

All of our thought experiments rely on the same setup: a recidivism-risk predictive system, \(T\), to be used by judges in courts. This setup parallels emerging technologies. In 2016, [112] audited COMPAS, a recidivism-risk system used in Wisconsin courts. They found the system disproportionately predicted future criminal activity along racial lines.

Furthermore, the experimental variables we chose have parallels to stakeholders that designers must consider. In our case, judges represent the group of intended users; the for-profit prison owner represents stakeholders who primarily have capital interests in the system. While the divine alien legislator is an absurd hypothetical, legislators are often asked to dictate policies that they will never be directly affected by. For example, male legislators can vote on current legislation around women’s health. We chose to use the divine alien legislator rather than a real-world example because it focuses the reader’s attention on a stakeholder’s level of interaction and effect. While our thought experiment may not be realizable, it is tractable.

5.3.2 Argument.↩︎

Moreover, thought experiments should be assessed in terms of logical validity and soundness.

Validity. We present our thought experiments in prose and formal logic to highlight their logical validity. We rely on the logical implications from the initial premises to reach our conclusions. For example, VSD states that \(P\) is a stakeholder if and only if they interact with technology \(T\) or are affected by \(T\). We use this if and only if to determine that judges are stakeholders in our experiment and, therefore, judges’ perspectives should be valued.

Soundness. Our logical representation of VSD stakeholders is heavily informed by foundational VSD literature [85]. Therefore, we argue that our formulation is sound. Note that VSD relies heavily on practitioners identifying key stakeholders (See Figure 3). VSD stakeholder theory demands breadth for this practitioner’s discretion. However, we show that this conceptual breadth leads to normative breadth; VSD does not ground normative standards on who should be prioritized in design. Even though we take a broad representation of VSD’s position on stakeholder theory, we believe it accurately represents the level of practitioner discretion within VSD methodologies.

6 Generating Formulations and Concepts↩︎

The previous sections have demonstrated an interrogative application of thought experiments—we take a widely accepted conceptualization of stakeholders, and then interrogate and iterate upon it. In this section, we ask:

Can we use thought experiments to speculate about a concept that does not have a clear formulation, but is the center of emerging discussions?

Recall that a conceptualization serves as an initial proposition to interrogate. Therefore, thought experiments may also be generative in defining and clarifying new concepts in HCI and related areas. In this section, we present thought experiments where the goal is to generate a concept or conceptualization rather than interrogate one.

We highlight Helen Nissenbaum’s [26], [88], [116] use of thought experiments as a conceptually generative method for HCI. Specifically, we formalize the thought experiments used by Nissenbaum and her peers to create the building blocks of contextual integrity. Contextual integrity asserts that privacy is about situational information norms rather than simply keeping certain information secret. Contextual integrity has been a staple framework in the HCI community for considering moral obligations surrounding personal data and information [27], [29], [59].

While Nissenbaum often used thought experiments as an explanatory device, our thought experiment structure from Section [sec:how-to] elucidates the formulations produced through this conceptual work. There is an ongoing discourse about applying contextual integrity to emerging machine learning methods in HCI and related human-centered areas [59]. Below, we apply our thought experiment structure to three thought experiments used by Nissenbaum and her peers. We illustrate how these thought experiments were generative as they created appropriate conceptual boundaries for the normative arguments Nissenbaum sought to make. Notably, these thought experiments were successful as contextual integrity has since informed radical changes in privacy policy [92].

6.1 Basic Example↩︎

One further feature is key to understanding what we mean here by “contexts,” for not only are they characterized by roles and norms but also by certain ends, or values. In the case of health care, an onlooker (say, from another planet) observing a typical health care setting of a hospital, will be unable to make proper sense of the goings-on without appreciating the underlying purpose behind it, that is, alleviating illness and promoting health. Although settling the exact nature of the ends and values for any given context is not a simple matter—even in the case of health care, which is relatively robust—the central point is that the roles and norms of a context make sense, largely, in relation to them. [88, p. 3]

Crucial to the notion of contextual integrity, is the concept of a “context." However, [26] notes that context is not well-defined in most disciplines, including HCI [14]. Therefore, we must generate a formulation of context before establishing contextual integrity. We take the thought experiment from  [88] and demonstrate how it proves generative.

Analogy. Imagine we have hospital \(H\) as a representation of a typical healthcare setting. Like most hospitals, \(H\)’s core goals are to alleviate illness and promote health. Now, suppose person \(P\) is observing \(H\).

Experimental Variable. Let the experimental variable be \(P\)’s understanding of \(H\)’s goals. Therefore, we have the following cases:

  1. Let \(P\) be a visiting doctor. Therefore, \(P\) has a deep and internalized understanding of \(H\)’s goals. In fact, \(P\) has agreed to the Hippocratic Oath, and, therefore, has deliberately aligned their values to the values of \(H\). In this case, we can conclude that \(P\) fully appreciates the purpose of \(H\).

  2. Let \(P\) be an average person who stops and observes the hospital through a window. While \(P\) is seeing everything a doctor may see, \(P\) may not have the same goals as a doctor. However, \(P\) has been educated on the goals of hospitals and the risks of a world without health care. \(P\) has probably been sick before and needed to go to the hospital themselves. In this case, \(P\) may not professionally align with \(H\)’s goals but certainly appreciates them.

  3. Let \(P\) be an alien from a planet with no sickness and, therefore, no hospitals. \(P\) is given a looking glass where they can observe all of the ongoings of \(H\). In this case, \(P\)’s understanding of the situation is distinctly different from a human understanding because they do not have an appreciation for the underlying purpose of a hospital.

Conclusion. This thought experiment suggests that context requires a relational conceptualization. Rather than an objective set of roles and norms (i.e., the jobs and tasks of a hospital), there is also a level of values-understanding that contributes to context. In this sense, context is a set of roles and norms that make sense given the values of the settings. [88] use this idea that contexts are distinctive social settings with underlying ends to formalize the notion that contexts contain norms. The violation of these contextual norms contributes to a violation of privacy under contextual integrity.

6.2 Generating a Formulation Based on A Counterfactual↩︎

A second consideration is the compatibility of notice-and-consent with the paradigm of a competitive free market, which allows sellers and buyers to trade goods at prices the market determines. Ideally, buyers have access to the information necessary to make free and rational purchasing decisions. Because personal information may be conceived as part of the price of online exchange, all is deemed well if buyers are informed of a seller’s practices collecting and using personal information and are allowed freely to decide if the price is right...A deeper ethical question is whether individuals indeed freely choose to transact–accept an offer, visit a website, make a purchase, participate in a social network–given how these choices are framed as well as what the costs are for choosing not to do so. While it may seem that individuals freely choose to pay the informational price, the price of not engaging socially, commercially, and financially may in fact be exacting enough to call into question how freely these choices are made.  [116, p. 34]

In this thought experiment, we start with the following conceptualization:

Proposition. All exchanges of personal information for online resources (i.e., information flows) are an instantiation of a free market exchange.

Analogy. Suppose we have person \(P\) with personal information \(P_{i}\) using online resource \(O\) in a free market setting. In essence, \(P\) is a buyer looking to exchange \(P_{i}\) for use of \(O\).

Experimental Variable. Let the experimental variable be the consequence to \(P\) of rejecting the exchange of \(P_{i}\) for \(O\). Therefore, we have the following cases:

  1. Let \(P_{i}\) be \(P\)’s credit card information and \(O\) be an online store. In this case, if \(P\) rejects the exchange, \(P\) cannot purchase an item from \(O\). This consequence seems reasonable as it mirrors an off-line setting; in order to buy most things an individual has to give their credit card information to the storekeeper.

  2. Let \(P_{i}\) be \(P\)’s cookie activity and \(O\) be an informational website on Medicare, a government-funded healthcare plan. In this case, if \(P\) rejects the exchange, \(P\) risks not getting necessary public health information.

  3. Let \(P_{i}\) be all of \(P\)’s activity on social media platform \(O\). Moreover, \(P\) uses \(O\) specifically to connect with recovery support groups. Now the consequence is extreme; if \(P\) rejects the exchange then they lose access to a necessary support group for their livelihood.

Conclusion. From these cases, we can conclude that there exists an information flow that is not a free exchange. In Case 2, \(P\) no longer has access to a public information resource. In Case 3, \(P\) no longer has access to a necessary support group. This thought experiment suggests that the privacy paradigm captured by notice-and-consent is based on a faulty presumption that information exchanges mirror free-market exchanges of goods. Now we can ask: Which dimensions would a new conceptualization of privacy need to include? Our experimental variable was the consequence of \(P\) rejecting an information exchange. We captured various cases by changing the values of \(P_{i}\), \(O\), and \(P\)’s intended use of \(O\). This suggests that a new conceptualization of privacy would need to be bounded by the following dimensions: information type, recipient type, and use norms. In fact, Nissenabum incorporates these three dimensions as a subset of the five parameters of contextual integrity: (1) the information subject; (2) the sender; (3) the recipient; (4) the information type; and (5) the transmission principles which represent the norms of information flow [28, p. 3]. A contextual integrity analysis requires a practitioner to define each of these parameters in relation to the context they are assessing.

6.3 Generating a Concept Based on an Emerging Event↩︎

Remember the hubbub over Google Street View in Europe? Germans, in particular, objected to the photo-taking cars. Many people, using the standard privacy paradigm, were like, "What’s the problem? You’re standing out in the street? It’s public!" But Nissenbaum argues that the reason some people were upset is that reciprocity was a key part of the informational arrangement. If I’m out in the street, I can see who can see me, and know what’s happening. If Google’s car buzzes by, I haven’t agreed to that encounter. Ergo, privacy violation. [92, p. 2]

In the early 2010s, Google deployed numerous labeled cars with cameras to take photos for Google Street View, a new feature of Google Maps3. Subsequently, a German woman sued Google for taking pictures of her house, arguing that Googled violated her privacy [117]. However, the court ruled that taking photos from the street is legal, as streets are public property rather than private. While this ruling was based on the dichotomy of public vs. private, we ask: What concept is being violated here that is not violated in similar situations?

Analogy. Suppose \(P\) is an average German citizen on a German street.

Experimental Variable. Our experimental variable here is visibility. We present the following cases:

  1. Suppose we are in a world without cameras and \(P\) is out taking a walk. In this case, \(P\) is only visible to the other people on the street. Notably, if someone on the street can see \(P\), \(P\) can likely see them too.

  2. Suppose \(P\) is on a walk and ends up in the background of a tourist photo taken by tourist \(T\). \(T\) may show the photo to their friends and maybe even post it on social media. In this case, \(P\) cannot see the audience of \(T\)’s friends that will end up seeing the photo.

  3. Suppose \(P\) walks in front of the Google Street View car and, therefore, is in a photo on Google Maps. Similar to Case 2, \(P\) cannot see the audience that can see \(P\). However, this metaphorical audience is now all Google Maps users who use the Google Street View feature on this German street. Notably, this is potentially orders of magnitude larger than the audience in Case 2.

Conclusion. This thought experiment demonstrates that the incorporation of cameras into society changes an individual’s visibility. Moreover, the addition of the Google Street Car adds a dimension of scale to the scenario. This combination of visibility and scale is what Nissenbaum uses to formulate the concept of reciprocity: “is it possible for subjects to see those who see them[118, p. 229]. While reciprocity does not need to be equivalent, technology causes it to fracture significantly. Nissenbaum incorporates this concept of reciprocity into a larger privacy framework by thinking of the (bi)-directional norms of information flows [88]. For example, there are no norms around information reciprocity between a patient and their doctor; doctors should not share their personal medical state with their patients. However, friends typically exchange phone numbers in a bidirectional manner. Reciprocity allows us to view information flows as exchanges with direction and scale, rather than a simple transfer from one person to the other. Notably, even though the German courts ruled in Google’s favor, Google placed a voluntary 10-year halt on its Street View program in Germany after German public outcry at the court’s decision [119].

The above examples highlight how thought experiments can be conceptually generative in addition to being interrogative. In each thought experiment, the conclusions lead us to a new formulation that is essential to building out contextual integrity. For example, understanding that visibility also has a component of scale (the size of the resulting audience) allows us to generate the concept of reciprocity. Future work in HCI can use this generative capacity of thought experiments to build out new concepts surrounding emerging topics. For example, tangential to data privacy, recent work has focused on theorizing data as labor to roadmap future research directions but still lays out a myriad of open questions [120]. We suggest that thought experiments can precisely build out foundational concepts surrounding data contribution, labor, and leverage in the same way Nissenbaum does for data privacy.

7 Discussion↩︎

We introduce thought experiments as a method for interrogating the normative consequences of conceptual commitments that HCI researchers make. This section discusses the utility of thought experiments to expand HCI’s methodological toolbox. We discuss concepts beyond stakeholders that could benefit from further application of our method. Finally, we identify key opportunities for researchers, reviewers, and critics to apply thought experiments.

7.1 Progressing the Methodology of HCI↩︎

Conventional methods and theories in HCI are well suited to answering certain questions: Which of these two design variants will balance best among the desired criteria [121]? What are the processes by which users come to adopt (or reject) a certain software [122]? However, as technology becomes part of our everyday lives, recent HCI work has begun engaging with moral questions: What would it look like to adopt normative lenses, such as feminist theory [18], Critical Race Theory [11], or FAccT4 frameworks [123], to achieve a moral standard?

 While these theories in HCI allow us to justify moral stances, HCI desperately needs methods that interrogate how practitioners take on implicit normative stances. For instance, how we conceptualize “the user" has consequences in focusing our research, design, and policy attention on certain actors and unintentionally ignoring others [4]. These conceptualizations also have an impact on who matters, not just for the field of HCI but for the much broader space of compassionate design in sociotechnical systems. The above thought experiments illustrate an analogous point for the concept of stakeholders: that different formulations of the concept carry different implicit commitments with different normative consequences.

We take the stance that understanding these unstated normative consequences is crucial to understanding what makes harm possible in sociotechnical systems. For example, what makes it possible for certain stakeholder groups, such as data subjects, to be consistently under-considered in predictive systems [120]? Our thought experiments demonstrate that VSD presumes a broad definition of stakeholderism. This methodological limitation has moral implications for whose voice gets considered in the design process. By recentering stakeholderism around morally-laden criteria, such as legitimacy or marginalization, we can make these implicit normative positions explicit.

Our work is in conversation with previous scholarship surrounding marginalization and harms. For example, [112] ask: What makes it possible for predictive systems to propagate human bias? [124] asks: What makes it possible for systems of research to propagate the silencing and exclusion of certain voices? While thought experiments are not a silver bullet to answering these questions, they illuminate where conceptual commitments are laden with values. Our thought experiments show that uninterrogated conceptual commitments can make the morally suspect seem benign, despite researchers’ intentions.

Bringing these normative implications to light is essential to progressing HCI towards its ethical aspirations. For example, [59] found that by articulating ethical tensions in predictive mental health, they revealed the methodological gaps that amplified the normative problems plaguing the field. [125] posit that community-based norms are the foundation for research ethics. We use thought experiments to unpack how certain formulations can contribute to community-wide values misalignments, such as misidentifying key stakeholders. We show that further thought experiments can reformulate concepts to reduce misalignment. In other words, thought experiments answer: What makes it possible for systems to be so misaligned? Interrogating these questions allows us to validate potential conceptual solutions.

Our proposal to progress the methodology of HCI presumes that practitioners are equipped to use thought experiments. We believe this presumption is justified, as there is a rich history of domain experts (i.e., physicists, computationalists, and mathematicians) engaging in thought experiments without strict philosophical training. However, there may be a learning curve to the widespread adoption of thought experiments in HCI and related research venues (CHI, CSCW, FAccT, etc.). Future efforts could leverage pre-existing infrastructure, such as educational workshops at conferences and symposia, to engage with external experts. The authors of this paper are committed to fostering those efforts where necessary, as we believe in sustained and interdisciplinary efforts to increase the quality of ethical reasoning in research.

7.2 Future Subjects for Thought Experiments↩︎

We use the concept of stakeholders in HCI as an instance of conceptual work that has led to uninterrogated normative consequences. This section highlights other concepts that could benefit from thought experiments. Future work could apply our thought experiments to analogous problem areas or leverage the method outlined in this paper to innovate new thought experiments.

Stakeholder inclusion has become a specific problem in Artificial Intelligence (AI) pipelines, leading to “participatory AI" methods [126]. Similar to VSD, these methods make conceptual commitments about who is a participant and the nature of participating. Properly conceptualizing AI participation is crucial in the current zeitgeist. Many communities have taken drastic action as they have been excluded from participating in AI, despite their data contributions. For example, fanfiction communities are poisoning generative AI models [127], and screenwriters are on strike because of job threats from ChatGPT [128]. Our work on interrogating conceptualizations of inclusion aligns with current calls for reconceptualizing data contributors as laborers [120] and AI advocates as participants [129]. Future work could adapt our thought experiments to the rapidly emerging field of generative AI. For example, if we changed our technology \(T\) to a generative AI system, what are the normative consequences of appealing to stakeholder legitimacy?

Moving away from questions about participant inclusion, thought experiments may be useful when generating idealistic concepts. For example, Helen Nissenbaum’s work highlights the methods gap that thought experiments solve; her work was not originally published in HCI venues, despite being heavily cited in HCI papers today [27], [29], [59]. Speculatively, this is because Nissenbaum used non-traditional methods for HCI at the time (i.e., thought experiments). When we have a subjective concept, such as privacy, how can we reach a formal and robust conceptualization? Nissenbaum’s work frequently uses thought experiments to demonstrate how viewing privacy as contextual integrity demands that information flows cannot be evaluated in a vacuum. We argue that HCI does not have a formal way of evaluating these hypothetical but tractable logical arguments. This gap can potentially lead to further exclusion of necessary work, such as Nissenbaum’s. Future work could continue Nissenbaum’s conceptual inquiry into privacy. For example, researchers could build on [27] and ask: What are the normative consequences of committing to contextual integrity in inferential systems? By answering these questions, HCI practitioners can precisely link value-laden conceptual decisions to their implications, allowing us to build and design more ethical systems.

7.3 Thought Experiments as a Tool For...↩︎

Thought experiments originated from philosophy. However, they have a rich history of being used by scientists to answer the philosophical questions of their field. For example, Newton and Galileo were physicists; Pythagoras was a mathematician; Alan Turing was a computer scientist. While some discussions about thought experiments are technical [32], [44], thought experiments have never been a technical method reserved for philosophy experts. In fact, HCI researchers are uniquely qualified to engage with thought experiments, given our expertise lies at the intersection of theoretical and technical domains. We propose thought experiments to build on HCI researcher expertise while addressing a methodological gap in the field. his section outlines key opportunities for HCI researchers to engage with thought experiments as tools for supplementing their other work.

7.3.1 Innovators in HCI↩︎

Innovation in HCI is often motivated by either (1) anticipating user needs or (2) reflecting on failures to meet user needs. [8] articulated the theory of sociotechnical gaps by identifying failed applications. Thought experiments can similarly identify moments for creative problem-solving and beneficial reconceptualizations. In this paper, we demonstrate how HCI can use thought experiments to build appropriate guardrails around design motivations. This mirrors the success of frameworks and taxonomies, which often serve as comprehensive lists of the status quo to guide future design [2]. Thought experiments iterate on the concepts that underlie these tools. For example, we show asking “what are the consequences of stakeholderness?” raises the question “who is a stakeholder?” Moreover, we show how to identify concepts that need to be strengthened. Recall our thought experiment about including a for-profit prison owner as a stakeholder. Here, we appeal to the hypothetical nature of thought experiments to identify conceptual weakness. However, it raises a larger gap within innovative HCI—how can we innovate with weak concepts? Innovation in HCI relies on appropriately scoping how we motivate design resources. Flawed conceptualizations risk misappropriating our attention to the point where we may miss the mark entirely. Therefore, we propose thought experiments as a way of scoping design motivation and, subsequently, foster more focused innovation in HCI.

7.3.2 Critics in HCI↩︎

Second, thought experiments can ground criticism around emerging technologies, not only in academia [109], [110] but also in the numerous reports of harms caused by unchecked technologies. With thought experiments, we can focus our interrogation on specific dimensions, such as stakeholder power, and bring normative consequences to the forefront of the conversation. For example, our divine alien legislator demonstrates the importance of considering stakeholder legitimacy. Without considering legitimacy, we risk prioritizing capital interests, such as the for-profit prison owner, in our design processes. Especially with the advent of AI technology, there is a growing need to rigorously critique large platform technologies. Moreover, many critical theorists argue that we should examine how technology practices intersect with identity, class, and power [11], [70], [130]. Thought experiments serve as a complement to the important and ongoing work in criticism. We hope that future critiques use thought experiments to iterate on other prominent concepts in social technology.

7.3.3 Reviewers in HCI↩︎

Thought experiments could also be applied in the peer review process. [131] recently called for normative implications to be an evaluative criterion during peer review. They argue that authors have an obligation to evaluate the potential negative consequences of their work; reviewers, in turn, have an obligation to scrutinize said consequences when considering the merits of the work. Thought experiments can be incorporated into the review process as a mechanism for authors and reviewers to explore, defend, or critique broader implications of potential publications. For example, authors who study non-traditional communities may use thought experiments to justify broader stakeholder inclusion. Reviewers may use a thought experiment to critique an author’s assumptions around stakeholderness and moral hierarchy, thereby giving a reviewer grounds to reject a paper on flawed normative reasoning. Thought experiments open the door to a transparent and precise dialogue surrounding the positive and negative normative consequences of research projects.

Our work establishes a method to reason about potential broader impacts. Thought experiments tie moral intuition to logical arguments. HCI researchers are often asked to consider their empirical findings alongside their informed moral speculations without clear scaffolding. Venues such as NeurIPS, EMNLP, and the NSF have recently encouraged publications to have broader impact statements. However, prior work has found a large information disparity in these statements [48], [132]. Thought experiments can precisely explicate the potential consequences of implicit conceptual commitments. For example, Machine Learning (ML) practitioners could use thought experiments to formally address how they conceptualize data subjects in their pipelines. This application could help address known ethical tensions, such as the privacy implications of inferring mental health states from social media data [59]; compensation for data subjects who are doing labor [120]; and copyright implications for non-traditional data sources, such as online fanfiction [133].

Thought experiments invite rules into the conversation to focus the reader toward specific dimensions. This makes thought experiments interrogative rather than speculative – the experiment takes on a stance (conceptual commitment) and seeks to precisely understand its normative implications. Consequently, using thought experiments creates a formal argument about potential harm. Thought experiments are a precise method that relies heavily on logical steps.

7.4 Limitations↩︎

This paper proposes thought experiments as a conceptual tool for HCI researchers. While this use of thought experiments aligns with previous philosophical work on the method [45], it is only one potential use of thought experiments for deliberate thinking. As mentioned, thought experiments have been used as explanatory devices for moral positions (see The Violinist Argument [38]) or quasi-experiments for the natural sciences (see Schrödinger’s Cat [31]). Future work could expand our initial proposal of thought experiments and adapt them to solve non-conceptual problems in HCI.

We also adopt Norton’s view that thought experiments are a unique type of argument [34]. Although this view is well-accepted, others have argued it is too conservative. For example, [32] directly responds to Norton saying that thought experiments generate new knowledge rather than simply revealing certain truths. This debate is ongoing in philosophy; future work could adopt Gendler’s (or others’) stance to reveal new and exciting uses for thought experiments in HCI.

8 Conclusion↩︎

This paper proposes thought experiments as a method for conceptual work. We leverage the rich philosophical history of thought experiments to highlight how they can logically take conceptual stances to their normative consequences. We demonstrate the use of thought experiments as an interrogative tool for borrowed concepts. We validate thought experiments’ value to HCI by proposing various thought experiments around conceptualizations of stakeholders. Thought experiments raise the potential for HCI researchers to better reflect on ethics, broader impacts, and their own positionalities. We suggest that if thought experiments were heavily used in HCI, the field would be better equipped to meet current moral imperatives.

References↩︎

[1]
Herbert A Simon.1973. . Artificial intelligence4, 3-4(1973), 181–201.
[2]
Antti Oulasvirta Kasper Hornbæk.2016. . In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems(San Jose, California, USA) (CHI ’16). Association for Computing Machinery, New York, NY, USA, 4956–4967.
[3]
Evan Caragay, Katherine Xiong, Jonathan Zong, and Daniel Jackson.2024. . In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–16.
[4]
Eric P S Baumer Jed R Brubaker.2017. . In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems(Denver, Colorado, USA) (CHI ’17). Association for Computing Machinery, New York, NY, USA, 6291–6303.
[5]
Lucy Suchman.1993. . Comput. Support. Coop. Work2, 3(Sept.1993), 177–190.
[6]
Steve Harrison, Phoebe Sengers, and Deborah Tatar.2011. . Interact. Comput.23, 5(Sept.2011), 385–392.
[7]
Geoff Cooper John Bowers.1995. . In The social and interactional dimensions of human-computer interfaces. Cambridge University Press, USA, 48–66.
[8]
Mark S Ackerman.2000. . Human–Computer Interaction15, 2-3(Sept.2000), 179–203.
[9]
Stuart K Card, Thomas P Moran, and Allen Newell.2018. The Psychology of Human-Computer Interaction. CRC Press.
[10]
Stevie Chancellor, Eric P S Baumer, and Munmun De Choudhury.2019. . Proc. ACM Hum. Comput. Interact.3, CSCW(Nov.2019), 1–32.
[11]
Ihudiya Finda Ogbonnaya-Ogburu, Angela D R Smith, Alexandra To, and Kentaro Toyama.2020. . In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–16.
[12]
Eric P S Baumer M Six Silberman.2011. . In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems(Vancouver BC Canada). ACM, New York, NY, USA.
[13]
Phoebe Sengers, Kirsten Boehner, Michael Mateas, and Geri Gay.2008. . Personal and Ubiquitous Computing12, 5(June2008), 347–358. https://doi.org/10.1007/s00779-007-0161-4.
[14]
Paul Dourish.2004. . Personal and ubiquitous computing8(2004), 19–30.
[15]
John Dewey James Hayden Tufts.2022. Ethics. DigiCat.
[16]
Terrell Ward Bynum.2000. . ACM SIGCAS Comput. Soc.30, 2(June2000), 6–13.
[17]
Katie Shilton.2018. . Foundations and Trends® in Human–Computer Interaction12, 2(2018), 107–171.
[18]
Shaowen Bardzell.2010. . In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI). ACM, Atlanta, GA, 1301–1310.
[19]
Shaowen Bardzell Jeffrey Bardzell.2011. . In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems(Vancouver, BC, Canada) (CHI ’11). Association for Computing Machinery, New York, NY, USA, 675–684.
[20]
Shaowen Bardzell.2018. . ACM Transactions on Computer-Human Interaction25, 1(Feb.2018), 6:1–6:24. https://doi.org/10.1145/3127359.
[21]
Veronica Abebe, Gagik Amaryan, Marina Beshai, Ilene, Ali Ekin Gurgen, Wendy Ho, Naaji R Hylton, Daniel Kim, Christy Lee, Carina Lewandowski, Katherine T Miller, Lindsey A Moore, Rachel Sylwester, Ethan Thai, Frelicia N Tucker, Toussaint Webb, Dorothy Zhao, Haicheng Charles Zhao, and Janet Vertesi.2022. . In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA) (CHI EA ’22, Article 1). Association for Computing Machinery, New York, NY, USA, 1–12.
[22]
Jaakko Hintikka, Ilpo Halonen, and Arto Mutanen.2002. . In Studies in Logic and Practical Reasoning, Dov M Gabbay, Ralph H Johnson, Hans Jürgen Ohlbach, and John Woods(Eds.). Vol. 1. Elsevier, 295–337.
[23]
John D Norton.1996. Can. J. Philos.26, 3(1996), 333–366.
[24]
Terrence Neumann, Maria De-Arteaga, and Sina Fazelpour.2022. . In 2022 ACM Conference on Fairness, Accountability, and Transparency(Seoul, Republic of Korea) (FAccT ’22). Association for Computing Machinery, New York, NY, USA, 1504–1515.
[25]
Advait Deshpande Helen Sharp.2022. . In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society(Oxford, United Kingdom) (AIES ’22). Association for Computing Machinery, New York, NY, USA, 227–236.
[26]
H Nissenbaum.2004. . Wash Law Rev.(2004).
[27]
Patrick Skeba Eric P S Baumer.2020. . Proc. ACM Hum.-Comput. Interact.4, CSCW2(Oct.2020), 1–22.
[28]
Sebastian Benthall, Seda Gürses, and Helen Nissenbaum.2017. . Foundations and Trends® in Privacy and Security2, 1(2017), 1–69.
[29]
Michael Zimmer.2018. . Social Media+ Society4, 2(2018), 2056305118768300.
[30]
W C Salmon.1999. Zeno’s Paradoxes. Hackett Publishing, Cambridge, MA.
[31]
E Schrödinger.1935. . Naturwissenschaften23, 48(Nov.1935), 807–812.
[32]
Tamar Szabó Gendler.1998. . Br. J. Philos. Sci.49, 3(Sept.1998), 397–424.
[33]
J John Stachel, David C Cassidy, and Robert Schulmann.1987. The Collected Papers of Albert Einstein volume I. the earlv years: 1879-1902. http://wilsonquarterly-legacy-attachments.s3.amazonaws.com/full-issues/social-mobility-in-america.pdf. .
[34]
John Norton.1991. . Thought experiments in science and philosophy129(1991).
[35]
James Robert Brown.1995. Thought experiments.
[36]
Isaac Newton N W Chittenden.1850. Newton’s Principia: The Mathematical Principles of Natural Philosophy. Geo. P. Putnam.
[37]
Judith Jarvis Thomson.1985. . Yale Law J.94, 6(May1985), 1395.
[38]
J J Thomson.2004. . Ethics(2004).
[39]
Gottfried Wilhelm Leibniz.1714. Monadology. (1714).
[40]
Georges Rey.1986. . Philos. Stud.50, 2(1986), 169–185.
[41]
John R Searle.1980. . Behav. Brain Sci.3, 3(Sept.1980), 417–424.
[42]
John D Norton.2004. Philos. Sci.71, 5(Dec.2004), 1139–1151.
[43]
James Robert Brown.2011. The laboratory of the mind: Thought experiments in the natural sciences (2 ed.). Routledge, London, England.
[44]
John D Norton.2002. . (2002).
[45]
Martin Kornberger Saku Mantere.2020. . Organization Theory1, 3(July2020), 2631787720942524.
[46]
Y Rogers.2004. . Annual review of information science and technology(2004).
[47]
Peter P Kirschenmann.1982. . In Scientific Philosophy Today: Essays in Honor of Mario Bunge, Joseph Agassi Robert S Cohen(Eds.). Springer Netherlands, Dordrecht, 85–98.
[48]
Leah Hope Ajmani, Stevie Chancellor, Bijal Mehta, Casey Fiesler, Michael Zimmer, and Munmun De Choudhury.2023. . In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency(Chicago, IL, USA) (FAccT ’23). Association for Computing Machinery, New York, NY, USA, 1311–1323.
[49]
Barry Bozeman Mary K Feeney.2007. . Adm. Soc.39, 6(Oct.2007), 719–739.
[50]
Gabriel Abend.2019. . Sociol. Theor.37, 3(Sept.2019), 209–233.
[51]
Bryan A Sisk, Jessica Mozersky, Alison L Antes, and James M DuBois.2020. . Am. J. Bioeth.20, 4(May2020), 62–70.
[52]
Katy Weathington Jed R Brubaker.2023. . Proc. ACM Hum.-Comput. Interact.7, CSCW1(April2023), 1–26.
[53]
Larry R Churchill.1982. . J. Higher Educ.53, 3(May1982), 296–306.
[54]
Su Lin Blodgett, Q Vera Liao, Alexandra Olteanu, Rada Mihalcea, Michael Muller, Morgan Klaus Scheuerman, Chenhao Tan, and Qian Yang.2022. . In CHI Conference on Human Factors in Computing Systems Extended Abstracts. ACM, New York, NY, USA.
[55]
Renee Shelby, Shalaleh Rismani, Kathryn Henne, Ajung Moon, Negar Rostamzadeh, Paul Nicholas, N’mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk.2023. . In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. ACM, New York, NY, USA.
[56]
Calvin A Liang, Sean A Munson, and Julie A Kientz.2021. . ACM Trans. Comput.-Hum. Interact.28, 2(April2021), 1–47.
[57]
Gopinaath Kannabiran.2023. . Interactions30, 1(Jan.2023), 19–21.
[58]
Shruthi Sai Chivukula, Aiza Hasib, Ziqing Li, Jingle Chen, and Colin M Gray.2021. . In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems(Yokohama, Japan) (CHI ’21, Article 295). Association for Computing Machinery, New York, NY, USA, 1–13.
[59]
Stevie Chancellor, Michael L Birnbaum, Eric D Caine, Vincent M B Silenzio, and Munmun De Choudhury.2019. . In Proceedings of the Conference on Fairness, Accountability, and Transparency(Atlanta, GA, USA) (FAT* ’19). Association for Computing Machinery, New York, NY, USA, 79–88.
[60]
Amanda M Williams Lilly Irani.2010. . In CHI ’10 Extended Abstracts on Human Factors in Computing Systems. ACM, New York, NY, USA.
[61]
Eric P S Baumer, Mark Blythe, and Theresa Jean Tanenbaum.2020. . In Proceedings of the 2020 ACM Designing Interactive Systems Conference(Eindhoven, Netherlands) (DIS ’20). Association for Computing Machinery, New York, NY, USA, 1901–1913.
[62]
Laura Barendregt Nora S Vaage.2021. . She Ji: The Journal of Design, Economics, and Innovation7, 3(Sept.2021), 374–402.
[63]
Mark Blythe Enrique Encinas.2018. . Foundations and Trends® in Human–Computer Interaction12, 1(2018), 1–105.
[64]
Julian Bleecker.2009. Design fiction: A Short Essay on Design, Science, Fact, and Fiction. 561–578 pages.
[65]
Christina N Harrington, Shamika Klassen, and Yolanda A Rankin.2022. . In CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA) (CHI ’22, Article 450). Association for Computing Machinery, New York, NY, USA, 1–10.
[66]
Paul Coulton, Joseph Galen Lindley, Miriam Sturdee, and Michael Stead.2017. . In Proceedings of Research through Design Conference 2017. eprints.lancs.ac.uk, GBR, 16.
[67]
Renee Noortman, Britta F Schulte, Paul Marshall, Saskia Bakker, and Anna L Cox.2019. . In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems(Glasgow, Scotland Uk) (CHI ’19, Paper 422). Association for Computing Machinery, New York, NY, USA, 1–14.
[68]
Daniel M Russell Svetlana Yarosh.2018. Interactions25, 2(Feb.2018), 36–40.
[69]
Kirsten E Bray, Christina Harrington, Andrea G Parker, N’deye Diakhate, and Jennifer Roberts.2022. . In CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA) (CHI ’22, Article 452). Association for Computing Machinery, New York, NY, USA, 1–13.
[70]
James Pierce, Phoebe Sengers, Tad Hirsch, Tom Jenkins, William Gaver, and Carl DiSalvo.2015. . In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems(Seoul, Republic of Korea) (CHI ’15). Association for Computing Machinery, New York, NY, USA, 2083–2092.
[71]
T Hancock C Bezold.1994. . Healthc. Forum J.37, 2(1994), 23–29.
[72]
Tjark Gall, Flore Vallet, and Bernard Yannou.2022. . Futures143(Oct.2022), 103024.
[73]
Anthony Dunne Fiona Raby.2013. Speculative Everything: Design, Fiction, and Social Dreaming. MIT Press.
[74]
RE Freeman.1984. . (1984).
[75]
Sergiy D Dmytriyev, R Edward Freeman, and Jacob Hörisch.2021. . J. Manag. Stud.58, 6(Sept.2021), 1441–1470.
[76]
Kevin Gibson.2000. . J. Bus. Ethics26, 3(2000), 245–257.
[77]
David Hatherly, Ronald K Mitchell, J Robert Mitchell, and Jae Hwan Lee.2020. . Business & Society59, 2(Feb.2020), 322–350.
[78]
Max E Clarkson.1995. . AMRO20, 1(Jan.1995), 92–117.
[79]
Haiyi Zhu, Bowen Yu, Aaron Halfaker, and Loren Terveen.2018. . Proc. ACM Hum.-Comput. Interact.2, CSCW(Nov.2018), 1–23.
[80]
Stevie Chancellor.2023. . Commun. ACM66, 3(Feb.2023), 78–85.
[81]
Angie Zhang, Alexander Boltz, Jonathan Lynn, Chun-Wei Wang, and Min Kyung Lee.2023. . In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg, Germany) (CHI ’23, Article 859). Association for Computing Machinery, New York, NY, USA, 1–19.
[82]
Angie Zhang, Olympia Walker, Kaci Nguyen, Jiajun Dai, Anqing Chen, and Min Kyung Lee.2023. . Proc. ACM Hum.-Comput. Interact.7, CSCW1(April2023), 1–32.
[83]
L Stapleton, M H Lee, D Qing, M Wright, and others.2022. . 2022 ACM Conference(2022).
[84]
Jeffrey Bardzell Shaowen Bardzell.2015. . Aarhus Ser. Hum. Centered Comput.1, 1(Oct.2015), 12.
[85]
Batya Friedman.1996. . Interactions3, 6(Dec.1996), 16–23.
[86]
Alan Borning Michael Muller.2012. . In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems(Austin, Texas, USA) (CHI ’12). Association for Computing Machinery, New York, NY, USA, 1125–1134.
[87]
Daisy Yoo.2017. . In Proceedings of the 2017 ACM Conference Companion Publication on Designing Interactive Systems(DIS ’17 Companion). ACM, New York, NY, USA, 280–284.
[88]
A Barth, A Datta, J C Mitchell, and H Nissenbaum.2006. . In 2006 IEEE Symposium on Security and Privacy (S&P’06)(Berkeley/Oakland, CA). IEEE.
[89]
Richard A Posner.1978. . Buffalo Law Rev.28(1978), 1.
[90]
Carol Warren Barbara Laslett.1977. . J. Soc. Issues33, 3(July1977), 43–51.
[91]
Paul B Thompson.2001. . Ethics Inf. Technol.3, 1(2001), 13–19.
[92]
Alexis C Madrigal.2012. . The Atlantic(March2012).
[93]
Roy A Sorensen.1992. Thought Experiments. New York: Oxford University Press.
[94]
Sören Häggqvist.2009. . Can. J. Philos.39, 1(2009), 55–76.
[95]
Ernst Mach.1893. The Science of Mechanics: A Critical and Historical Exposition of Its Principles. Open court publishing Company.
[96]
Herbert A Simon.1977. Artificial intelligence systems that understand. https://www.ijcai.org/Proceedings/77-2/Papers/096.pdf. .
[97]
John Preston Mark Bishop.2002. Views into the Chinese Room: New Essays on Searle and Artificial Intelligence. Oxford University Press.
[98]
Dairon Rodrı́guez, Jorge Hermosillo, and Bruno Lara.2012. . Minds Mach.22, 1(Feb.2012), 25–34.
[99]
Daniel Dennett.1980. . Behav. Brain Sci.3, 3(Sept.1980), 428–430.
[100]
M Bishop J Preston.2001. In Essays on Searle’s Chinese Room Argument, M Bishop J Preston(Eds.). Oxford University Press.
[101]
Colin R Caret Ole T Hjortland.2015. Foundations of Logical Consequence. OUP Oxford.
[102]
John W Lloyd.2012. Foundations of Logic Programming. Springer Science & Business Media.
[103]
Ned Block.1995. Behav. Brain Sci.18, 2(June1995), 272–287.
[104]
Mark D Sprevak.2007. . Br. J. Philos. Sci.58, 4(Dec.2007), 755–776.
[105]
John Searle.1984. . Brains Minds Media(1984), 28–41.
[106]
Iris M Young.1981. . Soc. Theory Pract.7, 3(1981), 279–302.
[107]
Janet Davis, Lisa P Nathan, and Others.2015. . Handbook of ethics, values, and technological design: Sources, theory, values and application domains(2015), 11–40.
[108]
Batya Friedman, Peter H Kahn, Alan Borning, and Alina Huldtgren.2013. . In Early engagement and new technologies: Opening up the laboratory, Neelke Doorn, Daan Schuurbiers, Ibo van de Poel, and Michael E Gorman(Eds.). Springer Netherlands, Dordrecht, 55–95.
[109]
M A DeVito, A M Walker, and J R Fernandez.2021. . of the ACM on Human-Computer …(2021).
[110]
H F Cheng, L Stapleton, A Kawakami, and others.2022. . Proceedings of the(2022).
[111]
Michael A Hallett.2006. Private Prisons in America: A Critical Race Perspective. University of Illinois Press.
[112]
Julia Angwin, Jeff Larson, Lauren Kirchner, and Surya Mattu.2016. Machine Bias. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing. .
[113]
Michael Cohen.2015. . The Washington Post(April2015).
[114]
Ronald K Mitchell, Bradley R Agle, and Donna J Wood.1997. . Acad. Manage. Rev.22, 4(1997), 853–886.
[115]
Robert Phillips.2003. . Bus. Ethics Q.13, 1(Jan.2003), 25–41.
[116]
Helen Nissenbaum.2011. . Daedalus140, 4(Oct.2011), 32–48.
[117]
Caroline McCarthy.2011. German court rules Google Street View is legal. https://www.cnet.com/culture/german-court-rules-google-street-view-is-legal/. .
[118]
Helen Nissenbaum.2019. . Theor. Inq. Law20, 1(March2019), 221–256.
[119]
Aggi Cantrill Stephanie Bodoni.2023. . Bloomberg News(July2023).
[120]
Hanlin Li, Nicholas Vincent, Stevie Chancellor, and Brent Hecht.2023. . In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency(Chicago, IL, USA) (FAccT ’23). Association for Computing Machinery, New York, NY, USA, 1151–1161.
[121]
Don Norman.2013. The Design of Everyday Things: Revised and Expanded Edition. Basic Books.
[122]
Jonathan Grudin.1988. . In Proceedings of the 1988 ACM conference on Computer-supported cooperative work(Portland, Oregon, USA) (CSCW ’88). Association for Computing Machinery, New York, NY, USA, 85–93.
[123]
Bruno Lepri, Nuria Oliver, Emmanuel Letouzé, Alex Pentland, and Patrick Vinck.2018. . Philos. Technol.31, 4(Dec.2018), 611–627.
[124]
Miranda Fricker.2007. . (June2007).
[125]
Amy S Bruckman, Casey Fiesler, Jeff Hancock, and Cosmin Munteanu.2017. . In Companion of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing(Portland, Oregon, USA) (CSCW ’17 Companion). Association for Computing Machinery, New York, NY, USA, 113–115.
[126]
Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang.2021. . NeurIPS(2021).
[127]
Amanda Silberling.2023. . TechCrunch(June2023).
[128]
Alissa Wilkinson.2023. The looming threat of AI to Hollywood, and why it should matter to you. .
[129]
Organizers Of Queerinai others.2023. . In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency(Chicago, IL, USA) (FAccT ’23). Association for Computing Machinery, New York, NY, USA, 1882–1895.
[130]
André Brock.2018. . New Media & Society20, 3(March2018), 1012–1030.
[131]
Brent Hecht, Lauren Wilcox, Jeffrey P Bigham, Johannes Schöning, Ehsan Hoque, Jason Ernst, Yonatan Bisk, Luigi De Russis, Lana Yarosh, Bushra Anjum, Danish Contractor, and Cathy Wu.2021. . (Dec.2021).  [cs.CY].
[132]
Priyanka Nanayakkara, Jessica Hullman, and Nicholas Diakopoulos.2021. . In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society(Virtual Event, USA) (AIES ’21). Association for Computing Machinery, New York, NY, USA, 795–806.
[133]
Casey Fiesler, Cliff Lampe, and Amy S Bruckman.2016. . In Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing(San Francisco, California, USA) (CSCW ’16). Association for Computing Machinery, New York, NY, USA, 1450–1461.

  1. Given the subject of this paper, many of our thought experiments elucidate moral claims but thought experiments can interrogate normative ones.↩︎

  2. https://neal.fun/absurd-trolley-problems/↩︎

  3. https://www.google.com/streetview/how-it-works/↩︎

  4. Fairness, Accountability, and Transparency↩︎