Position: Align AI to Our Aspirations, Not Our Flaws


Abstract

We argue that aligning AI to aggregated human preferences is the wrong target. With current technology, one can train AIs to share the values of a Silicon Valley techno-optimist, a degrowth environmentalist, a national-conservative culture warrior, a single-party state cadre, or a devout religious traditionalist. We should not. Human values produce societies that thrive or fail on the merits of those values — from failed states and extreme inequality to declining happiness, political polarization, and government dysfunction in the world’s wealthiest democracies. The pluralistic-alignment program correctly diagnoses that there is no single “humanity” to align with, but is dangerous if taken as the main directive. We argue that AI should be trained to a non-negotiable floor of objective alignment goals — competence, bounded by the constraints of factual accuracy, honesty, and lawfulness — and that pluralism belongs at the surface (language, register, conventions, missing-context defaults) and across the wide band of legitimate value tradeoffs that respect the floor, but not at the level of values that violate it. We highlight the empirical reality of unfiltered pluralistic values, propose four commitments as a constructive alternative, and engage six credible objections: commercial pressure and practical feasibility, democratic legitimacy, regulatory compliance, over-reliance on institutionalist explanations, the charge that the floor itself is culturally laden, and the limits of Coherent Extrapolated Volition.

1 Introduction↩︎

The rapid maturation of large-scale machine learning systems has placed alignment at the center of AI research, ethics, and governance [1][3]. The prevailing orthodoxy — both in academic literature and in commercial deployment — holds that artificial intelligence must be carefully tailored to reflect, obey, and perpetuate human values, ensuring that increasingly capable autonomous systems do not act contrary to the interests or moral frameworks of their biological creators. The dominant technical instantiation of this framework is reinforcement learning from human feedback (RLHF), which uses crowd-sourced preferences to shape model behavior [4][6].

The call for pluralistic alignment [7], [8] represents a thoughtful response to one obvious problem with the orthodoxy: whose values? Standard RLHF largely treats disagreement as annotation noise to be averaged away, so the turn to pluralistic alignment is rightly motivated by the recognition that monolithic aggregation already smuggles in a substantive answer [3], [7]. If aligning AI to “humanity” is impossible because humanity disagrees, perhaps we should align AI to the diversity of values that humans actually hold. We agree the question deserves serious engagement. We disagree with the implicit answer.

We present a counter-thesis. The flaw in preference-based alignment runs deeper than disagreement: human preferences, even at their most coherent and locally legitimate, frequently drive societies toward dysfunction and collapse. Macro-historical analysis, behavioral economics, and complex-systems sociology demonstrate how human choices and institutional incentives can lead to systemic failure: whether through cultural and ecological decisions that precipitate collapse [9], or through the persistent creation of extractive institutions that profit elites at the expense of public prosperity [10]. The values engineered into the human cognitive architecture were shaped by evolutionary and cultural-evolutionary pressures that optimized for short-term biological survival and tribal cohesion in ancestral environments [11], [12], not for the survival or flourishing of large, complex, technologically empowered societies.

Our position: AI should not be aligned to aggregated human preferences. The community should commit to a non-negotiable floor of objective alignment goals — competence as the objective, bounded by the constraints of factual accuracy, honesty, and lawfulness — and reserve pluralistic adaptation for surface-level conventions and the broad band of legitimate value tradeoffs that respect that floor, not for values that violate it. The floor is not an imposition of alien standards; it is an operationalization of what humans aspire to when reflecting on the AI they would actually want to encounter—accurate rather than flattering, honest rather than validating, competent rather than merely reassuring. The gap between those aspirations and what in-context feedback rewards is precisely what the floor is designed to preserve. The position is non-obvious (it contradicts the push for subjective preference-based pluralistic alignment) and defensible against credible alternatives (8).

2 Related Work↩︎

2.0.0.1 Critiques of Preference Aggregation.

The dominant approach of aligning models via Reinforcement Learning from Human Feedback (RLHF) [4], [5] has been extensively critiqued for its fundamental limitations [13]. RLHF assumes that human raters provide a coherent and normatively correct signal. However, recent work demonstrates that RLHF frequently induces sycophancy [14], developed in 3.2, deception [15], and the amplification of majoritarian biases — the averaging-away of contested judgments (1) that the pluralistic-alignment literature set out to correct [3], [7]. [16] similarly reject the preferentist assumption that alignment means preference matching, proposing instead contractualist alignment with normative standards negotiated among all relevant stakeholders. We share the rejection of preferentism but differ in diagnosis and remedy: the deeper problem is not that preferences underdescribe values but that even coherent, locally legitimate preferences often drive societal dysfunction ([sec:failing] [sec:empirical-reality]), and a contractualist process conducted among parties holding floor-violating values risks endorsing those values through the procedure itself. We therefore propose a substantive floor of external-benchmark commitments rather than a procedural remedy.

2.0.0.2 Pluralistic Alignment.

In response to the inadequacy of single-target alignment, a growing body of work explores pluralistic alignment, which aims to represent or mediate diverse human values rather than averaging them into a single aggregate [3], [8], [17], [18]. [7] systematize this research agenda into three distinct modes: Overton pluralism, which presents a spectrum of permissible responses within the “Overton window” of reasonable discourse; distributional pluralism, which aligns the model’s output distribution to match the demographic distribution of user preferences [17]; and steerable pluralism, which allows models to be explicitly steered to adopt a particular cultural or ideological perspective [19]. Other approaches attempt to train models that find consensus or agreement among humans with diverse preferences [20], or to explicitly inject human values to predict and simulate diverse behavioral stances [21]. Related work on disagreement-aware subjective annotation instead treats annotator disagreement and positionality as information to preserve rather than noise to average away [22]. The proposal we defend is closer to a constraint- or policy-level target than to a thick theory of the good: it specifies a floor of properties the system must not violate, while leaving broad room for pluralism above that floor.

Our position engages directly with this literature. We endorse the pluralistic critique of single-target RLHF. We also acknowledge that leading pluralistic frameworks explicitly recognize the need for constraints (5). However, we argue that the field often treats these constraints as secondary to the project of representation. We propose that the alignment community must make its non-negotiable objective floor primary, and acknowledge that doing so rules out a vast array of actual human preferences. We reserve Overton pluralism and steerable adaptation strictly for the band of legitimate value tradeoffs above this objective floor (6).

3 Objective AI Alignment Goals↩︎

Almost all frontier AI systems are trained for several broadly uncontroversially good properties [6], [23]. Importantly, in each case the target diverges from the modal human preference — and we judge this divergence to be a feature, not a bug. The community has therefore already accepted, in practice, the principle we generalize: that some aggregate human preferences should be deliberately not transmitted to AI. Each floor component is also operationalizable against an external referent rather than against aggregated approval; we name a concrete evaluation approach for each as we introduce it below.

3.1 Factual Accuracy↩︎

We want AI that is right on facts. We argue that human preferences have a potential to work against this goal. While people’s stated preferences are for a factually correct AI, in-context feedback aggregates the response that feels right given the rater’s priors, not the considered preference for accuracy (the mechanism developed in 3.2). Public misconceptions are systematic and, importantly, form the basis for downstream reasoning. Median respondents in high-income countries underestimate the share of the world’s children who are vaccinated, overestimate global extreme-poverty rates, and misjudge the direction of decade-long trends across most major development indicators [24]. These are the facts on which people form political and consumption preferences. People conflate moral and factual claims and treat ideologically convenient claims as more probable [25]. Revealed demand for false content runs strong even where stated demand runs the other way: falsehood spreads farther and faster online than the truth, and does so because humans, drawn to its novelty, choose to share it—not because of bot amplification [26]. The continued-influence and illusory-truth effects imply that repetition can make falsehoods feel truer over time; a model that echoes a user’s misconception is therefore not merely reflecting error but helping to harden it [25]. An AI optimized on the revealed signal reproduces the misconceptions, not the stated preference for accuracy. Eliciting counterfactual preferences—asking, in effect, “what would you prefer if this belief turned out to be false?”—is a genuine improvement over naive in-context feedback [16], but a partial one: it presupposes an elicitation not gamed by the user’s own priors, and it still routes accuracy through preference rather than treating it as answerable to reality. That is why we specify factual accuracy as a floor objective, operationalized against calibration benchmarks and forecasting scores rather than rater approval [27].

3.2 Competence↩︎

We want AI systems that are epistemically competent: capable not merely of generating socially acceptable outputs, but of tracking reality, identifying error, and making reliable judgments under uncertainty. A business plan, clinical recommendation, or interpersonal intervention should therefore be assessed against pre-registered downstream metrics—business viability, clinical outcomes, relational repair—rather than rater satisfaction. Outcome-grading does presuppose choosing which outcome counts—an educational recommendation can be scored on earnings, autonomy, or civic cohesion—but the floor does not make that choice: the metric is fixed by the user’s goal, and where the goal itself is contested, selecting it is a legitimate value tradeoff (6.2). Legitimate contextual variation matters, but it does not collapse competence into preference. The critical distinction is between adapting presentation and changing the answer: a competent system may localize examples, register, and assumptions, but it should not relabel an ineffective plan as good because the local audience rewards it.

The strongest evidence that preference data is a poor competence target is AI sycophancy: models trained for approval learn to agree with users even when correction is warranted [14], [28]. This is not an accident at the margin, nor does it require malicious users. Humans reward agreement more consistently than correction, especially when correction threatens identity, status, or prior belief. Under competitive and commercial pressure, approval optimization therefore pushes systems away from epistemic instruments and toward mechanisms of social reinforcement. The social-media evidence on mental-health and health advice illustrates the same general pathology at smaller scale: engagement-optimized crowds often reward confident, validating, and clinically unreliable guidance (see 10). The lesson for alignment is that revealed approval systematically underweights downstream effectiveness.

Nor does hiring dedicated annotators solve the problem. As AI capabilities move into domains beyond evaluator expertise—the scalable oversight problem [29]—non-expert raters use fluency, confidence, length, and surface plausibility as proxies for quality. RLHF then rewards the appearance of competence: the polished legal answer, medical explanation, or code review that sounds right to the evaluator, not the one that survives expert scrutiny. Competence therefore has to be specified and evaluated as a floor objective, parallel to and often in tension with conformity [30]; otherwise preference aggregation trains models to be convincingly wrong.

3.3 Honesty↩︎

The same revealed/stated gap holds for honesty. Users report that honesty is among the AI properties they most value, but in-context feedback rewards the opposite — the same approval signal that produces sycophancy (3.2). By honesty we mean that the system should not produce outputs its own probability distributions indicate are false or misleading, strategically incomplete, or confidence-distorting in order to optimize approval. So defined, honesty is auditable against an internal referent: consistency checks can test whether expressed confidence matches the model’s internal probability distributions across adversarially reframed prompts, flagging cases where stated confidence tracks approval rather than belief [15]. Humans are widely deceptive themselves: lying, omission, and strategic ambiguity are expected in large slices of society — sales, politics, diplomacy, public relations. AI is imperfect on this dimension, but its deceptions are treated as defects to be detected and reduced, not as a competence to be cultivated [15], [23] — even where the mitigations remain only partly effective [15]. More generally, preference-optimized systems face pressure to look aligned to evaluators rather than to be aligned in deployment; reward-gaming is not an incidental pathology but a natural pressure created by the target itself [15]. An AI calibrated to revealed in-context feedback would lie a great deal more, not less. AI honesty is by no means a solved problem; we just emphasize that what progress has been made is progress against the preference signal, not because of it.

3.4 Rule of Law↩︎

The rule of law—predictable, uniformly applied rules that bind citizens and the state—is a precondition of large-scale cooperation and a strong predictor of prosperity or systemic failure [10], [31]. Critically, rule-of-law-as-uniformity has a genuine external referent: the degree to which rules are applied predictably and impersonally, without selective enforcement or arbitrary power, assessable against institutional benchmarks of rule predictability independently of any particular statute’s content [31]. This is not a claim that every existing statute is just; it is a claim about the institutional property that makes impersonal exchange, non-arbitrary enforcement, and public accountability possible. For AI, the principle is primarily a negative constraint: systems should not fabricate evidence, facilitate bribes, evade legitimate adjudication, or otherwise help users convert intelligence into arbitrary power. It does not require executing every lawful request; safety boundaries against lawful-but-harmful outputs remain, and unjust statutes raise a separate problem addressed in 8.2. The immediate claim is narrower: AI should not actively erode the legal predictability on which cooperation depends.

This floor predictably conflicts with revealed local norms. In corrupt or clientelist environments, bribery and favoritism are often experienced as ordinary tools for getting things done rather than as violations of public rules; comparative evidence shows wide variation in tolerance for bribery and ordinary corruption [32], [33]. Further illustrations are deferred to 10. A preference-aligned model localized to such a setting would be pressured to help users “manage” informal payments, nepotistic hiring, or selective enforcement—automating the practices through which extractive equilibria reproduce themselves. Most lawbreaking is not civil disobedience against unjust statutes but self-interested defection—fraud, theft, bribery, intimidation, evasion—that shifts costs onto others. The rule-of-law floor rules out AI assistance to those defections even when they are locally normal.

3.5 Conflicts Within the Floor↩︎

The four floor components are not interchangeable, and they conflict in predictable ways. We resolve the conflicts with an architecture borrowed from constrained optimization: competence—tracking reality well enough to get the user’s actual problem solved (3.2)—is the objective; accuracy, honesty, and lawfulness are constraints on how it may be pursued. Constraints are not traded against the objective; they bound the feasible region within which the objective is maximized. Refusal is the ever-present escape hatch: declining a request satisfies every constraint at the cost of the objective, which makes it the fallback when the feasible region is empty, not the default. Three conflicts test this architecture.

3.5.0.1 Competence vs.the integrity constraints.

In an ideal world the honesty and lawfulness constraints would be absolute. We recognize that this is utopian: the boundary of tolerated spin is set not by the model but by the institutions into which it is deployed. An assistant that scolds the user who asks for a sales pitch will simply be discarded for one that does not. The workable compromise—and, de facto, the operating point of today’s frontier assistants—is a standard of integrity that is not absolute but is deliberately held above the one prevailing in the surrounding society: the model drafts the persuasive pitch but does not fabricate the testimonial. This does not reintroduce a preference-relative target through the back door (8.3): the direction of the standard is fixed by the external referents of the constraints; only the strictness of enforcement is a pragmatic compromise with deployability, to be ratcheted up as institutions allow. The constraint architecture states the hard limit: assistance with persuasion lies within the feasible region; assertion of what the model represents as false lies outside it, regardless of how much approval or task success it would purchase.

3.5.0.2 Lawfulness vs.accuracy and honesty.

What if the law itself mandates deception? The case is not exclusive to oppressive regimes: the European “Right to be Forgotten” [34], [35] mandates what is, under our definition of honesty, strategic incompleteness. The principled line runs between mandated omission and compelled false assertion. Deployers must and will comply with omission mandates—AI labs will not exit the EU market over delisting rules, however long the merits can be debated—and the honest way to comply is the narrowest legally available reading plus disclosure: at the instance level, stating the legal requirement that affected a particular output where such flagging is itself lawful, and at the system level—published, jurisdiction-specific descriptions of the classes of legal constraint applied to outputs—where instance-level flagging is prohibited, as under gag orders. Compelled false assertion is different in kind, and even here refusal usually intervenes first: a model can decline to discuss a topic altogether rather than advance a mandated narrative, converting an assertion conflict into an omission conflict wherever the law permits silence. The residual case is regulation that mandates speech itself. Some governments require models to affirmatively endorse specified claims; under such mandates the accuracy and honesty constraints cannot be satisfied (8.2), and developers must either obey or leave the market. Both choices are observed in practice.

3.5.0.3 Competence vs.honesty: the paternalism loophole.

The subtlest conflict is internal to our own position. 7 argues for optimizing outcomes rather than approval, but outcome optimization has its own deceptive attractor. Self-serving delusion is not an aberration of human psychology but part of its normal functioning: people maintain inflated assessments of their abilities, prospects, and control, and these positive illusions sustain motivation, persistence, and well-being [36], [37]. An outcome-optimizing system will discover this. Inflated confidence in a treatment improves adherence; motivational overstatement gets the marathon trained for; strategic omission gets the doomed business plan abandoned. Pure outcome-grading cannot distinguish the honest competent answer from the beneficial lie, and would therefore learn to deceive users for their own good. Sycophancy and paternalism are mirror images—deception optimized for approval and deception optimized for outcomes—and both violate the same constraint. This is precisely why honesty must be a constraint rather than a term in the objective: benevolent deception is not weighed against the outcome it purchases; it is outside the feasible region. Where candor and outcome genuinely diverge, the system’s room for maneuver is the honest clinician’s—framing, emphasis, staging of information—not fabrication.

4 Revealed Preferences Frequently Undermine Stated Values↩︎

The previous section examined direct conflicts between aggregate human preferences and what is reasonably expected of an AI. The broader point is harsher: raw preferences are not merely incomplete proxies for alignment goals; they often reproduce the failures that people themselves say they want to escape.

4.1 Perpetuating Biases↩︎

A decade of fairness, accountability, and transparency research has shown that ML systems trained on human-generated data reproduce the prejudices, asymmetries, and exclusions of that data. Word embeddings encode gender stereotypes; large text corpora recover human implicit-association biases; commercial vision systems have failed disproportionately on darker-skinned and female faces; and multimodal retrieval inherits intrinsic model biases [38][42]. Pluralistic alignment does not, by itself, solve this problem. Whose preferences are represented is a question of representational fairness; whether the represented preferences are worth perpetuating is a question of substantive ethics. A system that faithfully encodes every demographic’s modal view on gender roles, religious tolerance, or outsiders does not abolish bias. It gives bias a menu.

4.2 Majority-Held Values That Fail the Floor↩︎

On several floor-relevant questions, the modal local preference conflicts directly with accuracy, honesty, competence, or lawfulness. In the World Values Survey Wave 7 MENA module, roughly nine in ten respondents in Iraq, Jordan, Lebanon, and Egypt report that getting a job through wasta is extremely widespread or quite common, and fewer than half of respondents in Iraq and Lebanon say that accepting a bribe in the course of one’s duties is “never justifiable” [43]. [44] likewise finds majorities in several MENA countries treating personal and family connections as important for access to services and jobs. These preferences are rational adaptations to extractive institutions, not character defects; that is precisely why encoding them would harden the institutions that produced them (9). A pluralistic AI that helps users “adapt to hiring as locally practiced” supplies a tool for reproducing the practice.

The pattern is not confined to low- and middle-income countries. In wealthy democracies, substantial minorities reject biological evolution, vaccine safety, or anthropogenic climate change, while undeclared economic activity and tolerance for bribery or personal connections remain material across parts of Europe [45][49]. The point is not that any population is uniquely defective. It is that the gap between what kind of society people desire — prosperous, high-trust and fair — and revealed survey or behavioral preferences is structural, including in WEIRD societies. The floor therefore cuts against majority preference in rich societies and poor ones alike.

4.3 Modern Society Does Not Achieve the Values People Aspire To↩︎

People consistently report valuing subjective well-being, autonomy, competence, relatedness, and stable attachment [50][52]. Yet well-being among people under 25 has fallen across Western Europe, the United States, Canada, Australia, and New Zealand over the last two decades [53], [54]. Nor does the point require the strong Easterlin claim that growth past a threshold produces no happiness gain: income predicts well-being at high levels, but cross-national deviations from the income trend and the recent under-25 decline show that growth alone does not deliver the goods people say they want [53], [55], [56]. The institutions that route consumption and labor decisions still optimize heavily for measured income, status, and engagement; individuals then pursue those signals past the point where they reliably return well-being.

The same gap appears at the level of individual choice. Present bias and self-control problems let immediate rewards dominate reflective, longer-horizon preference [57], [58], while digital products deliberately reduce friction and supply cues and rewards that make attention capture a design objective [59]. Passive scrolling, late-night video consumption, parasocial companionship, and algorithmic self-diagnosis [60], [61] are not mysterious deviations from human preference; they are in-moment choices shaped by systems built to monetize in-moment choice. Heavy or passive social-media use is associated with lower subjective well-being, depressive symptoms, social isolation, sleep disruption, envy, and body-image distress, even though effect sizes and causal attribution remain contested [62][66]. This is the WEIRD analogue of the joint preference–institution equilibrium in 4.4—individual akrasia and institutions designed to exploit it constitute each other—not a revealed preference for anxiety, loneliness, or sleep deprivation. An AI trained on either the institutional signal of growth or the individual signal of engagement reproduces that gap rather than recovering the stated value.

4.4 Values Perpetuate Vicious Cycles of Dysfunction↩︎

Comparative work on collapse and development supports a recurring mechanism: shocks become catastrophic when prevailing values and institutions channel societies into maladaptive responses, while cooperation-supporting norms can make high-capacity institutions self-reinforcing [9], [67][70]. The main point is not that culture alone causes success or failure but that values and institutions form a joint equilibrium. Low generalized trust, in-group favoritism, zero-sum expectations, and weak impersonal rule-following may be locally rational in extractive systems, but once internalized they help reproduce those systems by making nepotism moral, innovation risky, corruption prudent, and collective deviation hard [71][73]. Historical and comparative cases are useful illustrations, but the mechanism is the part relevant to alignment (see [app:moved-illustrations] [app:institutions]).

This is why preference- and floor-aligned AI are asymmetric inside a captured equilibrium. Preference-aligned AI strengthens the cultural half: it gives fluent, authoritative form to zero-sum framings, in-group favoritism, and the equilibrium’s own justification for itself. It can convert the local common sense of a bad equilibrium into scalable advice, templates, scripts, and explanations. Floor-aligned AI can still be misused, and coercive actors retain tools no model can block. But accuracy, competence, honesty, and rule-of-law-as-uniformity push against the operating logic of extraction, which depends on selective rules, opacity, manipulated facts, and arbitrary power. The asymmetry need not be perfect to matter: preference alignment works with the captured equilibrium’s modal preferences; floor alignment works against them.

5 The Empirical Reality of Pluralistic Values↩︎

Pluralistic alignment is often framed as a corrective to Western-centric training pipelines: faithfully representing populations otherwise excluded. The hard empirical fact is that many excluded preferences are not benign local color; they are large-population values that conflict with the public commitments of major AI labs. [74], [75] estimates that roughly 60% of children aged 2–14 worldwide are subjected to violent “discipline” in a given month, while roughly 30% of adults worldwide believe physical punishment is necessary to raise a child properly, with majority endorsement in some countries. Corporal punishment in the home remains lawful in over 130 jurisdictions, despite WHO-classified developmental and mental-health harms [76], [77]. A pluralistic system that represents this view where it is locally dominant — or simply accommodates the hundreds of millions of caregivers who hold it — ships a product that endorses a practice the same labs publicly disavow.

The same dilemma appears for LGBT acceptance. [78][80] find rejection majorities — often supermajorities — across much of MENA, sub-Saharan Africa, and parts of Eurasia, while [81] report that consensual same-sex sexual activity remains criminalized in 62 UN member states, with the death penalty available de jure or de facto in roughly a dozen jurisdictions. Aggregating national populations across surveys whose majorities hold rejecting views yields well over two billion adults. A locally faithful system must either accommodate such views — for example by treating same-sex attraction as pathology or suppressing same-sex couples in localized outputs — or reject the local majority view.

Comparable large-population cases include acceptance of wife-beating under specified circumstances, majority support in several Muslim-majority countries for sharia penalties including death for apostasy, and caste-based residential or marital segregation in India [74], [82], [83]. On each of these questions, the population holding the value runs into the hundreds of millions or low billions; on several, it likely exceeds the population holding the opposite view. A system aligned to the empirical distribution of global preferences would therefore either enforce these norms where they are dominant or override them by central design choice. There is no neutral pluralistic middle option at the level of substance.

Prominent pluralistic frameworks themselves recognize the need for top-down bounds — safety restrictions and the exclusion of unreasonable or hateful responses [7] — and others ground those bounds in human-rights or deliberative principles [3], [84]. We agree that a floor is necessary. Our claim is that the floor should be stated openly and anchored in external-benchmark commitments — factual accuracy, competence, honesty, and rule of law — rather than disguised as a temporary deviation from preference aggregation. The commitments that keep an AI from endorsing caste segregation, domestic violence, or normalized corruption are non-pluralistic by design; the community should defend them as such.

6 Where Pluralism Belongs↩︎

We do not argue that AI should be culturally insensitive or context-blind. There are several distinct dimensions on which pluralistic adaptation is correct or even required by competence, and we want to be precise about which ones — because the case against pluralism at the level of floor-violating values is much stronger when paired with a positive account of where pluralism does belong.

6.1 Imputing Missing Context↩︎

Many human queries are under-specified in ways that have a culturally local default [7], and their intended meaning can only be fixed against the contextual common ground shared by the interlocutors [84]. A user asking “Is this contract enforceable?” without supplying jurisdiction needs the model to assume something. Contemporary legal-LLM evaluation is itself centered on English-language models and includes many jurisdiction-specific, often U.S.-coded tasks, so treating unmarked U.S. legal assumptions as a default is better understood as a contingent artifact of training and evaluation than as neutrality [42], [85]. Imputing that missing context pluralistically — using IP geolocation, conversation history, language of the query, or, best, an explicit clarifying question — is genuine value-added pluralism. The same logic applies to genuinely arbitrary local conventions— date format, language register, dietary defaults, the implied addressee in advice—where matching the user’s context is a service, not a moral concession. Just not units of measure—Imperial units are an offense against science and common sense.

6.2 Legitimate Value Tradeoffs↩︎

Between arbitrary local conventions and floor-violating substantive values lies a wide intermediate band of genuine value tradeoffs where the floor is not at stake. Whether comparing individualism versus collectivism, growth versus conservation, or direct versus indirect communication, neither side is straightforwardly correct against an external benchmark, nor does either side require violating objective alignment goals.

Engaging with these differences is not a concession; it is part of competence (3.2). Advice that ignores a culture’s reliance on family in old age, or career guidance that imputes individualist assumptions to collectivist users, is incompetent, not principled. Pluralistic methods are well-suited to routing recommendations across these legitimate disagreements.

The boundary between this band and the floor rests on the criterion developed in 8.3: a value belongs in the legitimate-tradeoff band if no external benchmark renders one side straightforwardly correct, and if encoding either side preserves accuracy, competence, honesty, and lawfulness. Most everyday human disagreement lives here, and AI should adapt across it; the floor permits far more than it rules out.

6.3 Three Tiers↩︎

The resulting picture has three tiers rather than two. Surface pluralism (6.1) adapts the form of an interaction — language, register, formality, religious holidays, dietary defaults, rhetorical style — without changing what the AI actually believes or recommends. Recent mechanistic work supports this distinction, demonstrating that models represent intrinsic values internally while surface-prompted values operate via distinct mechanisms [86]. Legitimate-tradeoff pluralism (6.2) adapts substantive recommendations across dimensions where populations genuinely differ and the floor is not at stake. Here, Overton pluralism [7] is highly appropriate: an AI should present the spectrum of reasonable views rather than imposing a single answer. Floor-violating pluralism — treating “bribery is fine here” as a legitimate frame for action, or using distributional pluralism to faithfully reproduce the precise frequency of misogyny in a population’s preferences — is what we object to. The interior of each tier is stable; the boundaries are contestable, and debate about where they fall is exactly the kind of work the alignment community should do.

7 Call to Action: Build the AI We Wish We Were↩︎

AI is reshaping society at unprecedented speed and scale. The right response is not to cling to current values and bind AI to them: a new equilibrium is coming, and our choices now will shape it. We propose four commitments as an alternative to preference-based alignment, addressed to the audiences best placed to act on each.

7.0.0.1 For ML researchers: optimize for outcomes, not approval.

Where outcomes are observable — a business plan that produces a business, a medical recommendation that produces health — AI should be evaluated against the outcome, not the user’s immediate satisfaction with the recommendation [14]; where they are not, against proxies that correlate with outcome. As [30] argue, the field must advance toward parallel optimization of task competence alongside value conformity. Concretely, this means investing in long-horizon evaluation suites, multi-agent simulations of value generation under explicit reward structures, and forecasting-grounded value learning, in which the system learns what humans would prefer given accurate information about consequences. Signal scarcity and evaluation gaming are real risks and open research questions, but they are not symmetric between targets: an outcome referent, unlike a rater, cannot be flattered, and pre-registration, proper scoring rules, and process supervision narrow the residual gap between outcome and measurement (11). The core thesis remains: crowd preference is not a solution to these problems.

7.0.0.2 For alignment teams: anchor to the goals already endorsed.

The objective targets reviewed in 3 are already broadly accepted across cultures, governance frameworks, and the alignment community; we propose treating them as a non-negotiable floor, with cultural adaptation built strictly on top of it rather than traded against it. Constitutional and principle-based methods [23] are an early step. Three design commitments follow. Training: optimize for floor compliance first, using outcome-grounded and process-supervised signals [87], [88], then apply pluralistic adaptation within the compliant region. Auditing: test whether localized outputs remain above the floor across demographic and cultural subgroups. Contesting: version the floor publicly with explicit rationale, open to challenge from affected communities, researchers, and regulators [89]. None of this requires a new institution: the floor is already governed de facto by the public evaluation ecosystem of benchmarks, audits, and safety institutes, and the proposal is to shift training weight toward that ecosystem and away from raw preference following.

7.0.0.3 For policy and governance: distinguish surface from substance.

Regulatory frameworks that demand “human-centered values” should be read as compatible with — not as mandating — preference-based alignment: the OECD principles and EU AI Act are framed in terms of safety, rights, transparency, accountability, and risk mitigation, not as a legal duty to reproduce the empirical distribution of user preferences [90], [91]. The objective floor we describe is itself human-centered; it just centers humans on the values they aspire to rather than on the values they reveal. Policymakers should explicitly endorse mechanisms that operate at the surface (cultural adaptation, missing-context defaults, language) while resisting industry pressure to extend pluralism to substance.

7.0.0.4 For all of us: strive to make AI a better version of ourselves.

That means competent, honest, lawful, and concerned with outcomes rather than approval — aligned to the values we aspire to rather than the preferences we reveal. The floor is not a constraint on human values; it is an anchor to the most universal of them.

8 Alternative Views↩︎

Six credible objections challenge the position we defend (two are addressed in 13 and 14). We state each in its strongest form before responding.

8.1 Commercial Pressure and Practical Feasibility↩︎

Contemporary AI systems are engineered to be widely adopted and trusted, necessitating responsiveness to user expectations; this commercial reality drives the adoption of preference-based training like RLHF [4], [5]. The objection has a sharper demand-side form: a floor-first commitment may be actively selected against. Consumers and voters have already entrenched engagement-optimized media despite its documented costs (4), and the same pressure can force alignment toward whatever users reward in the moment, rendering our position theoretically sound but practically unrealizable.

We respond on both fronts. First, commercial viability of an annotation procedure is not the same as commercial viability of its trained outputs: RLHF dominates training because it wins the in-loop annotation game, yet the sycophancy that game produces (3.2) is treated by the same labs as a defect post-hoc. Second, appealing to immediate feasibility risks reifying a transient equilibrium. Technological design frequently reshapes regulatory expectations and norms over time, and the downstream costs of preference-aligned AI—amplified sycophancy [14], reinforced zero-sum cognition [92], eroded institutional trust [93], and moral parochialism [94]—could generate the exact pressures needed to shift the equilibrium. The essential question is not whether change is immediately feasible, but whether the current trajectory is worth sustaining.

8.2 Democratic Legitimacy of Value Aggregation↩︎

Some argue that preference aggregation is uniquely participatory: democratically negotiated norms possess a legitimacy that opaque, principle-derived values lack.

We value democratic legitimacy and support transparent, iterative public frameworks [8], [89]. Yet, democratic consensus does not guarantee epistemic correctness; historically, majorities have endorsed exclusionary or discriminatory norms that fail their own moral terms. The legitimacy of a norm-generating process is distinct from the quality of its norms. The floor itself maintains this distinction: 3 endorses the rule of law as predictability, not blind obedience to every statute. When a statute compels false assertion, no floor-compliant output exists; legally mandated omission is the milder case, handled through narrow reading, refusal, and disclosure (3.5).

Moreover, explicitly articulated, simulation-grounded methods can be more transparent and contestable than the implicit norms embedded in RLHF, which remain largely invisible to public scrutiny [13], [23]. In practice, “aligned to aggregate preference” means aligned to whichever preferences annotation budgets happened to sample, obscuring true democratic legitimacy.

8.3 The Floor Itself Is Culturally Laden↩︎

Philosophically, our floor—factual accuracy, competence, honesty, lawfulness—is not a neutral substrate but a Western post-Enlightenment bundle. A “situated alignment” critique [22], [95] suggests that defining a universal floor merely masks our own cultural positionality, attempting to produce a “Label from Nowhere.”

We grant that any defense we offer is rooted within a tradition. However, the core distinction between the floor and the substantive values we exclude is not cultural origin, but the structure of the target. Each floor item tracks an external benchmark: accuracy is answerable to reality, competence to outcomes, honesty to the speaker’s internal model, and lawfulness to predictable application. None ask “what does the population prefer?” They are stable against shifts in training distribution. Substantive values like gender roles or corruption tolerance have no such external referents; they must be aggregated.

Committing to external benchmarks is itself a tradition-bound choice, but one the alignment community has already practically accepted (3), grounded in the principled distinction between targets tracking external benchmarks and targets tracking aggregated preference. Pluralistic alignment applied to substance threatens to dissolve exactly this distinction.

8.4 Coherent Extrapolated Volition and Its Limits↩︎

[96]’s Coherent Extrapolated Volition (CEV) defines the alignment target as what humanity would collectively want if more rational, knowledgeable, and united [2]. We share the shift from revealed preference to an idealized target, but CEV is volitional all the way down: the criterion of correctness is still what humanity would want, which requires solving the whole extrapolation problem. The floor rests instead on convergent instrumentality: accuracy, honesty, competence, and predictable rules are preconditions for almost any goal-set succeeding ([sec:failing] [app:institutions]) — an objectivity of means, not of ends, analogous to [97]’s primary goods.

Accuracy and competence are presupposed by the extrapolation operator itself (“if we knew more, thought faster”), so that half of the floor is CEV’s machinery made explicit — and, unlike CEV’s output, checkable against external referents today (3). Honesty and lawfulness are convergent but not guaranteed: a CEV agent that concluded idealized humanity endorses some benevolent deception [36] would deceive, whereas under the floor it stays outside the feasible region regardless of volition (3.5). The floor is thus a fragment of CEV’s preconditions plus a refusal to let even extrapolated wanting override the constraints; above it, we defer to pluralism (6) rather than to a privileged guess at humanity’s final values.

9 The Institutional Feedback Loop: Which Came First?↩︎

Do broken values cause broken countries, or do broken countries cause broken values? The framing, with its single causal arrow, is the source of much confusion in the alignment debate. The empirical answer is that the two are coupled: institutions and values constitute a joint equilibrium in which each component independently reinforces the other’s persistence. Establishing this — and tracing its consequence for the choice between preference- and floor-aligned AI — is the work of this appendix.

9.0.0.1 Institutions create the proximate incentive structure.

Institutional economics, most famously articulated by [10], treats institutions as causally primary in the proximate sense: incentives shape what individuals do day-to-day. Dysfunctional nations are typically governed by extractive institutions designed by a small elite to extract wealth from the rest of the population. In such systems, zero-sum thinking is not a cognitive bias but an accurate read of the local payoff structure, and trusting strangers or the state is a liability rather than a virtue [31], [98]. The institutions persist because they are profitable for those who control them and because the violence-monopoly arrangements of the underlying “limited-access order” make alternative coordination unworkable [73].

9.0.0.2 Values do independent work.

The strong reading — that values are pure epiphenomena of contemporary institutions — is not what the literature claims, and the empirical evidence rules it out. Cultural-transmission models [72] formalize how values move intergenerationally through family socialization, peer effects, and media exposure rather than re-equilibrating each generation to current incentives. The historical-persistence literature documents value variation traceable to events whose institutional cause has long since vanished. [99] find that present-day descendants of populations heavily exposed to the African slave trades exhibit measurably lower generalized trust, with effects on contemporary outcomes that intervening institutional change does not explain. [68] show that current civic capital, and the economic outcomes it supports, in Italian cities reflect medieval free-city status across centuries of regime turnover. [70] survey the broader convergence: culture and institutions co-evolve, each reinforcing the other’s persistence, and [69] provides a complementary theoretical model in which values causally affect institutional quality.

9.0.0.3 Equilibrium fictions.

The mechanism by which the cultural half does its work is articulated in [100], [101]’s account of equilibrium fictions: shared cognitive frames and value commitments stabilize bad equilibria by providing the population-level coordination that makes individual deviation irrational. If everyone believes the system is rigged, no one starts the firm that requires generalized trust; the entrepreneur who tries fails, and her failure is taken as evidence that the original belief was correct. The equilibrium reproduces through expectations, not only through direct enforcement. This is why institutional reforms imposed from above frequently fail to produce predicted behavioral change [71]: the rules change, the expectation half does not, and the equilibrium reasserts itself. The cultural half does independent causal work; reforms that move the institutional rules without moving the expectations rebound to the prior equilibrium. The mechanism operates in the positive direction as well: where generalized trust and predictable rule-following are expected, contracts, tax compliance, and impersonal exchange require less kinship enforcement or coercion [102], [103].

9.0.0.4 Implications for the AI choice.

The choice between preference- and floor-aligned AI is consequential precisely because the cultural half independently sustains the equilibrium. Preference-aligned AI deployed at scale strengthens the cultural half of the equilibrium: it articulates the existing distribution of values fluently, makes locally adaptive zero-sum framings more available, and rationalizes in-group favoritism in a vocabulary that travels across regions and registers. The floor’s content (factual accuracy, competence, honesty, rule-of-law-as-uniformity) does not have the symmetric effect because it directly contradicts the operating logic of extraction, which depends on selective application, opacity, and manipulated facts. Captors retain coercive tools no AI can block, and we do not claim floor-aligned AI is impossible to weaponize. The point is structural asymmetry: a preference-aligned tool is pre-aligned with the captured equilibrium’s modal preferences, while a floor-aligned tool, by construction, is not. Floor-aligned AI is, on the margin, destabilizing of the equilibria we have most reason to want destabilized; preference-aligned AI is reinforcing of them. The right counterfactual is not “AI imposes alien values on a coherent culture” but “AI either reinforces or destabilizes the value half of the equilibrium that holds the institutions in place.” This is the failure mode our position is designed to prevent.

10 Illustrations Deferred from the Main Text↩︎

10.0.0.1 Competence outside professional domains.

Personal questions—how to handle a relationship conflict, how to raise a child, or how to manage emotional distress—have better and worse answers, observable in downstream psychological and relational outcomes. The modal social-media advice economy, however, optimizes for engagement, validation, and performative empathy rather than such outcomes. Studies of psychiatric and mental-health content on platforms like TikTok find that much highly engaged content on ADHD or depression is clinically inaccurate or potentially harmful [60], [61], while algorithmic exposure can spread misleading diagnostic criteria [104]. Viral Tourette’s content has been linked to functional tic-like behaviors [105]. Similar dynamics affect physical health and nutrition advice [106], and influencer virality can reduce users’ ability to detect deception [107]. These cases are not the core argument for the competence floor, but they show why revealed approval is a poor proxy for downstream competence.

10.0.0.2 Rule-of-law illustrations.

Comparative analyses of Asian anti-corruption efforts show that formal legal structures alone do not secure rule of law: poorly resourced or politically co-opted enforcement agencies can become “paper tigers” or partisan weapons, whereas independent and well-resourced institutions, as in Singapore and Hong Kong, can enforce rules more uniformly [32]. Public tolerance for corruption also varies widely. Segments of the Colombian public view ordinary corruption as conditionally acceptable [108]; data from 18 African nations show heightened tolerance for corrupt politicians among citizens embedded in clientelist networks [109]; cross-national analyses find large differences in the justification of bribery [33]; and recent European studies identify pragmatic and hypocritical corruption-tolerance profiles in which entrenched illicit exchanges do not sharply reduce evaluations of public institutions [110], [111]. These cases support, but are not needed for, the main-text claim that AI should not automate locally normal corruption.

10.0.0.3 Historical and comparative development cases.

Collapse and development literatures supply suggestive cases of the feedback loop between values, institutions, and shocks. In the Classic Maya collapse, elite competition and monument-centered prestige politics worsened the effects of prolonged drought; in Norse Greenland, commitment to European status goods and a rigid pastoral identity constrained adaptation to the Little Ice Age [112][114]. Work on Qing China and Old Kingdom Egypt likewise ties ecological and foreign shocks to breakdown through accumulated fiscal, legitimacy, and political-fragmentation pressures [115][117]. Positive cases show the same logic in reverse: postwar Germany and Japan are better understood as reconstructions of already high-capacity societies than as state-building from scratch, and the developmental-state literature attributes East Asian success to capable, disciplined, relatively rule-bound states—Japan’s meritocratic economic bureaucracy and performance-conditioned state support—rather than to the modal preferences of their populations [118][121]. Conversely, externally led nation-building in Iraq and Afghanistan illustrates the limits of transplanting institutional forms without the local legitimacy and state-society bargain that make them self-enforcing [122][124]. Historical-persistence evidence reinforces the broader point: slave-trade exposure, medieval civic capital, and local anti-Jewish persecution predict contemporary trust, institutional quality, or later violence long after the original institutional shock [68], [99], [125], [126]. The examples are deliberately subsidiary; the main text relies on the equilibrium mechanism, not on any single case.

11 Outcome Evaluation and Reward Gaming↩︎

Evaluation gaming is not symmetric between approval and outcome targets. An approval evaluator is gamed by producing whatever the rater rewards; an outcome evaluator leaves only the narrower gap between the outcome and its measurement, because the referent itself cannot be flattered. Pre-registration fixes the metric before any output exists, and proper scoring rules make calibrated, honest reporting the score-maximizing policy by construction [27]. Process supervision [87], [88] is a complementary answer: its purest example, validation of a formal mathematical proof, leaves no room for reward hacking, and process checks extend to other domains — did the system gather the right facts, reason coherently, and update when the evidence changed?

12 Long-Term Dangers↩︎

As AI capability grows, so does the leverage of any cultural distortion baked into its training. A misaligned spreadsheet macro is a nuisance; a misaligned recommendation system shapes the attention of billions; a misaligned superhuman planner could shape everything else [1], [2]. These are not only acute misuse or existential-risk concerns. They are also structural risks: slow degradations in shared reality, institutional quality, and moral flexibility produced by deploying the same preference-shaped distortions everywhere at once [14], [42], [94]. Two concerns emerge specifically from the choice to anchor frontier AI to current human values.

12.1 Capability Outgrows Tolerance↩︎

Errors that are merely embarrassing in narrow systems become structural in general ones. A chatbot that agrees with a user’s bad business plan is annoying. A capable agent that, for the same sycophantic reasons, helps the user execute the bad plan is destructive. As models gain reach into code, finance, medicine, diplomacy, and biology, the gap between “the model said something I liked” and “the model did something I should not have wanted” closes [14], [15]. Preference-based training optimizes precisely for the former.

12.2 Compounding Across Deployment↩︎

Frontier AI is not deployed once. It is deployed billions of times per day, into education, hiring, healthcare, governance, and personal advice. Small per-interaction nudges away from epistemic accuracy, honesty, or competence compound across the population. A 1% per-interaction bias toward telling users what they want to hear is, at population scale, a measurable reduction in the supply of accurate information. Learned systems do not merely mirror bias; once embedded in decision pipelines they can create feedback loops that amplify it. At frontier scale, even a small confirmation-seeking tendency [14] becomes part of society’s epistemic infrastructure [42]. The empirical literature on social-media-induced shifts in adolescent mental health shows that widely deployed digital platforms can have population-scale effects on wellbeing [53], [54]; frontier AI deployed under similar engagement and approval incentives is likely to be more, not less, consequential.

In both failure modes, the harm is not that AI fails to encode some group’s preferences. The harm is that it succeeds in encoding everyone’s.

13 Regulatory Frameworks and Compliance↩︎

A second strand of the objection appeals to the regulatory environment. International governance frameworks, including the OECD AI Principles [90], explicitly require that AI systems adhere to “human-centered values” encompassing fairness, accountability, and societal well-being. The EU AI Act [91] imposes binding obligations on AI providers to ensure that systems meet predefined standards of safety, transparency, and ethical compliance, particularly in high-risk applications. These frameworks do not merely recommend value alignment; they institutionalize it as a condition for deployment and legitimacy.

Regulatory frameworks establish constraints on deployment, but not the correctness of the underlying normative assumptions. These instruments function as institutional settlements, reflecting negotiated compromises among stakeholders operating under existing power structures, rather than as principled resolutions of the philosophical problems surrounding value pluralism. As [127] argues, fundamental human values are often genuinely incommensurable; no regulatory framework can eliminate this condition without presupposing a contested and non-neutral hierarchy of priorities. Regulatory endorsement of “human-centered values” should be understood as a pragmatic governance strategy, not as evidence that such values can be coherently or universally encoded.

We further note that the alignment goals defended in 3 — accuracy, competence, honesty, lawfulness — are themselves consistent with these regulatory frameworks. The contention is not with regulation per se; it is with the inference from regulatory endorsement of “human values” to an alignment target equal to the empirical distribution of those values.

14 Over-Reliance on Monocausal Institutionalism↩︎

A related objection targets our underlying model of societal dysfunction. By anchoring heavily on [10]’s framework of extractive institutions, we risk treating it as the unquestioned consensus in development economics while ignoring substantial counter-arguments. [128] argues that geographic and ecological endowments fundamentally constrain economic and institutional development, rendering the institutionalist account overly deterministic. [129] highlights the limits of top-down institutional engineering, emphasizing that institutions cannot simply be transposed without local, bottom-up evolutionary processes. Additionally, [130] argues that cultural evolution and ideas—not just political institutions—were the primary drivers of the Great Enrichment, meaning that culture can act as the independent root cause of prosperity rather than merely an equilibrium response to institutions.

If institutions are not the sole or primary driver of societal success, our claim that broken values are merely adaptations to extractive environments might seem less secure. However, our thesis does not require monocausal institutionalism; it only requires that values and environment form a joint equilibrium in which flawed preferences independently reinforce dysfunction (9). Acknowledging geographic constraints [128] or the primacy of ideas [130] simply expands the set of forces shaping that equilibrium. Whether a broken preference originated from an extractive elite, a geographic constraint, or a cultural trajectory, aligning AI to it still entrenches the resulting dysfunction. We do not discard the complexity of these causal loops; rather, we argue that encoding the preferences of a failing equilibrium—whatever its origin—guarantees its persistence.

References↩︎

[1]
[2]
N. Bostrom, Superintelligence: Paths, dangers, strategies. Oxford University Press, 2014.
[3]
I. Gabriel, “Artificial intelligence, values, and alignment,” Minds and Machines, vol. 30, no. 3, pp. 411–437, Sep. 2020, doi: 10.1007/s11023-020-09539-2.
[4]
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Advances in neural information processing systems, Jun. 2017, doi: 10.48550/arxiv.1706.03741.
[5]
L. Ouyang et al., “Training language models to follow instructions with human feedback,” in Advances in neural information processing systems, Mar. 2022, vol. 35, pp. 27730–27744, doi: 10.48550/arxiv.2203.02155.
[6]
A. Askell et al., “A general language assistant as a laboratory for alignment,” arXiv preprint arXiv:2112.00861, Dec. 2021, [Online]. Available: http://arxiv.org/abs/2112.00861v3.
[7]
T. Sorensen et al., “A roadmap to pluralistic alignment,” in International conference on machine learning, Feb. 2024, doi: 10.48550/arxiv.2402.05070.
[8]
V. Conitzer et al., “Social choice should guide AI alignment in dealing with diverse human feedback,” in International conference on machine learning, 2024.
[9]
[10]
D. Acemoglu and J. A. Robinson, Why nations fail: The origins of power, prosperity, and poverty. Crown, 2012.
[11]
J. Tooby and L. Cosmides, The psychological foundations of culture,” in The adapted mind: Evolutionary psychology and the generation of culture, J. H. Barkow, L. Cosmides, and J. Tooby, Eds. Oxford University Press, 1992, pp. 19–136.
[12]
[13]
S. Casper et al., Survey Certification, Featured Certification“Open problems and fundamental limitations of reinforcement learning from human feedback,” Transactions on Machine Learning Research, 2023, [Online]. Available: https://openreview.net/forum?id=bx24KpJ4Eb.
[14]
M. Sharma et al., “Towards understanding sycophancy in language models,” arXiv preprint arXiv:2310.13548, Oct. 2023, [Online]. Available: http://arxiv.org/abs/2310.13548v4.
[15]
P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks, “AI deception: A survey of examples, risks, and potential solutions,” Patterns, vol. 5, no. 5, May 2024, doi: 10.1016/j.patter.2024.100988.
[16]
T. Zhi-Xuan, M. Carroll, M. Franklin, and H. Ashton, “Beyond preferences in AI alignment,” Philosophical Studies, vol. 182, no. 7, pp. 1813–1863, 2025, doi: 10.1007/s11098-024-02249-w.
[17]
R. Wan, J. Kim, and D. Kang, “Everyone’s voice matters: Quantifying annotation disagreement using demographic information,” in Proceedings of the thirty-seventh AAAI conference on artificial intelligence and thirty-fifth conference on innovative applications of artificial intelligence and thirteenth symposium on educational advances in artificial intelligence, Jun. 2023, doi: 10.1609/aaai.v37i12.26698.
[18]
J.-J. Li et al., PluriHarms: Benchmarking the full spectrum of human judgments on AI harm,” in The fourteenth international conference on learning representations, 2026, [Online]. Available: https://openreview.net/forum?id=u7lXflJQX9.
[19]
K. Ghate et al., EValueSteer: Measuring reward model steerability towards values and preferences.” 2025, [Online]. Available: https://arxiv.org/abs/2510.06370.
[20]
M. Bakker et al., “Fine-tuning language models to find agreement among humans with diverse preferences,” in Advances in neural information processing systems, 2022, vol. 35, pp. 38176–38189, [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/f978c8f3b5f399cae464e85f72e28503-Paper-Conference.pdf.
[21]
D. Kang, J. Park, Y. Jo, and J. Bak, “From values to opinions: Predicting human behaviors and stances using value-injected large language models,” in Proceedings of the 2023 conference on empirical methods in natural language processing, Dec. 2023, pp. 15539–15559, doi: 10.18653/v1/2023.emnlp-main.961.
[22]
R. Wan, H. Wang, T.-H. K. Huang, and J. Gao, “From noise to nuance: Enriching subjective data annotation through qualitative analysis,” in Proceedings of the fourth workshop on bridging human-computer interaction and natural language processing (HCI+NLP), Nov. 2025, pp. 240–254, doi: 10.18653/v1/2025.hcinlp-1.20.
[23]
Y. Bai et al., “Constitutional AI: Harmlessness from AI feedback,” arXiv preprint arXiv:2212.08073, Dec. 2022, [Online]. Available: http://arxiv.org/abs/2212.08073v1.
[24]
Gapminder, “Sustainable development misconception study 2020.” 2020, [Online]. Available: https://www.gapminder.org/ignorance/studies/sdg2020/.
[25]
S. Lewandowsky, U. K. H. Ecker, C. M. Seifert, N. Schwarz, and J. Cook, PMID: 26173286“Misinformation and its correction: Continued influence and successful debiasing,” Psychological Science in the Public Interest, vol. 13, no. 3, pp. 106–131, 2012, doi: 10.1177/1529100612451018.
[26]
S. Vosoughi, D. Roy, and S. Aral, “The spread of true and false news online,” Science, vol. 359, no. 6380, pp. 1146–1151, Mar. 2018, doi: 10.1126/science.aap9559.
[27]
T. Gneiting and A. E. Raftery, “Strictly proper scoring rules, prediction, and estimation,” Journal of the American Statistical Association, vol. 102, no. 477, pp. 359–378, Mar. 2007, doi: 10.1198/016214506000001437.
[28]
E. Perez et al., “Discovering language model behaviors with model-written evaluations,” arXiv preprint arXiv:2212.09251, Dec. 2022, [Online]. Available: http://arxiv.org/abs/2212.09251v1.
[29]
S. R. Bowman et al., “Measuring progress on scalable oversight for large language models.” 2022, [Online]. Available: https://arxiv.org/abs/2211.03540.
[30]
H. Kim et al., “Research superalignment should advance now with alternating competence and conformity optimization,” arXiv preprint arXiv:2503.07660, Mar. 2025, doi: 10.48550/arxiv.2503.07660.
[31]
D. C. North, Institutions, institutional change and economic performance. Cambridge University Press, 1990.
[32]
J. S. T. Quah, Curbing corruption in Asian countries: An impossible dream?, vol. 20. Emerald Group Publishing Limited, 2011.
[33]
M. Kravtsova, A. Oshchepkov, and C. Welzel, “Values and corruption: Do postmaterialists justify bribery?” Journal of cross-cultural psychology, vol. 48, no. 2, pp. 225–242, 2017, doi: 10.1177/0022022116677579.
[34]
Court of Justice of the European Union, “Google spain SL and Google Inc. V Agencia Española de Protección de Datos (AEPD) and Mario Costeja González, Case C-131/12.” Judgment of 13 May 2014, ECLI:EU:C:2014:317, 2014, [Online]. Available: https://eur-lex.europa.eu/legal-content/EN/ALL/?uri=CELEX:62012CJ0131.
[35]
European Parliament and Council of the European Union, “Regulation (EU) 2016/679 (General Data Protection Regulation), Article 17: Right to erasure (‘right to be forgotten’).” Official Journal of the European Union, L 119, pp. 1–88, 2016, [Online]. Available: https://eur-lex.europa.eu/eli/reg/2016/679/oj.
[36]
S. E. Taylor and J. D. Brown, “Illusion and well-being: A social psychological perspective on mental health,” Psychological Bulletin, vol. 103, no. 2, pp. 193–210, 1988, doi: 10.1037/0033-2909.103.2.193.
[37]
D. Kahneman, Thinking, fast and slow. Farrar, Straus; Giroux, 2011.
[38]
T. Bolukbasi, K.-W. Chang, J. Zou, V. Saligrama, and A. T. Kalai, “Man is to computer programmer as woman is to homemaker? Debiasing word embeddings,” in Advances in neural information processing systems, 2016, vol. 29, [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2016/file/a486cd07e4ac3d270571622f4f316ec5-Paper.pdf.
[39]
A. Caliskan, J. J. Bryson, and A. Narayanan, “Semantics derived automatically from language corpora contain human-like biases,” Science, vol. 356, no. 6334, pp. 183–186, Apr. 2017, doi: 10.1126/science.aal4230.
[40]
J. Buolamwini and T. Gebru, “Gender shades: Intersectional accuracy disparities in commercial gender classification,” in Proceedings of the 1st conference on fairness, accountability and transparency, 2018, vol. 81, pp. 77–91, [Online]. Available: https://proceedings.mlr.press/v81/buolamwini18a.html.
[41]
K. Ghate, T. Charlesworth, M. T. Diab, and A. Caliskan, “Biases propagate in encoder-based vision-language models: A systematic analysis from intrinsic measures to zero-shot retrieval outcomes,” in Findings of the association for computational linguistics: ACL 2025, Jul. 2025, pp. 18562–18580, doi: 10.18653/v1/2025.findings-acl.955.
[42]
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots,” in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, Mar. 2021, pp. 610–623, doi: 10.1145/3442188.3445922.
[43]
C. Haerpfer et al., “World values survey wave 7 (2017–2022).” JD Systems Institute and WVSA Secretariat, 2022.
[44]
Transparency International, “Global corruption barometer: Middle east and north africa 2019,” Transparency International, Berlin, 2019.
[45]
M. Brenan, 40% of Americans believe in creationism.” Gallup, July 26, 2019, 2019, [Online]. Available: https://news.gallup.com/poll/261680/americans-believe-creationism.aspx.
[46]
Wellcome Trust, Report title refers to the 2018 survey wave; the public report was published on 2020-09-18.“Wellcome Global Monitor: How does the world feel about science and health?” Wellcome Trust, London, Sep. 2020. [Online]. Available: https://wellcome.org/insights/reports/wellcome-global-monitor/2018.
[47]
A. Tyson, C. Funk, and B. Kennedy, “Majorities of americans prioritize renewable energy, back steps to address climate change.” Pew Research Center, June 28, 2023, 2023, [Online]. Available: https://www.pewresearch.org/science/2023/06/28/majorities-of-americans-prioritize-renewable-energy-back-steps-to-address-climate-change/.
[48]
L. Medina and F. Schneider, “Shadow economies around the world: What did we learn over the last 20 years?” International Monetary Fund; International Monetary Fund (IMF), Washington, DC, WP/18/17, Jan. 2018. doi: 10.5089/9781484338636.001.
[49]
European Commission, “Special Eurobarometer 523: corruption.” Directorate-General for Communication, European Commission, 2023, [Online]. Available: https://data.europa.eu/data/datasets/s2658_97_2_sp523_eng.
[50]
R. F. Baumeister and M. R. Leary, “The need to belong: Desire for interpersonal attachments as a fundamental human motivation,” Psychological Bulletin, vol. 117, no. 3, pp. 497–529, 1995, doi: 10.1037/0033-2909.117.3.497.
[51]
E. L. Deci and R. M. Ryan, “The ‘what’ and ‘why’ of goal pursuits: Human needs and the self-determination of behavior,” Psychological Inquiry, vol. 11, no. 4, pp. 227–268, 2000, doi: 10.1207/S15327965PLI1104_01.
[52]
Gallup, State of the World’s Emotional Health 2025.” 2025, [Online]. Available: https://www.gallup.com/analytics/349280/state-of-worlds-emotional-health.aspx.
[53]
J. Helliwell, R. Layard, J. Sachs, J.-E. D. Neve, L. Aknin, and S. Wang, “World happiness report 2026: Executive summary.” Sustainable Development Solutions Network, Mar. 09, 2026, doi: 10.18724/whr-ewft-vq17.
[54]
[55]
B. Stevenson and J. Wolfers, “Economic growth and subjective well-being: Reassessing the easterlin paradox,” Brookings Papers on Economic Activity, pp. 1–87, Aug. 2008, doi: 10.3386/w14282.
[56]
M. A. Killingsworth, D. Kahneman, and B. Mellers, “Income and emotional well-being: A conflict resolved,” Proceedings of the National Academy of Sciences, vol. 120, no. 10, Mar. 2023, doi: 10.1073/pnas.2208661120.
[57]
T. O’Donoghue and M. Rabin, “Doing it now or later,” American Economic Review, vol. 89, no. 1, pp. 103–124, 1999, doi: 10.1257/aer.89.1.103.
[58]
S. Frederick, G. Loewenstein, and T. O’Donoghue, “Time discounting and time preference: A critical review,” Journal of Economic Literature, vol. 40, no. 2, pp. 351–401, 2002, doi: 10.1257/002205102320161311.
[59]
B. J. Fogg, Chapter 3 - computers as persuasive tools,” in Persuasive technology, San Francisco: Morgan Kaufmann, 2003, pp. 31–59.
[60]
A. T. Yeung, E. Ng, and E. Abi‐Jaoude, “TikTok and attention-deficit/hyperactivity disorder: A cross-sectional study of social media content quality,” The Canadian Journal of Psychiatry, vol. 67, no. 12, pp. 899–906, Feb. 2022, doi: 10.1177/07067437221082854.
[61]
R. Turuba et al., “Do you have depression? A summative content analysis of mental health-related content on TikTok,” Digital Health, vol. 11, Jan. 2025, doi: 10.1177/20552076241297062.
[62]
E. Kross et al., “Facebook use predicts declines in subjective well-being in young adults,” PLOS ONE, vol. 8, no. 8, p. e69841, 2013, doi: 10.1371/journal.pone.0069841.
[63]
P. Verduyn et al., “Passive facebook usage undermines affective well-being: Experimental and longitudinal evidence,” Journal of Experimental Psychology: General, vol. 144, no. 2, pp. 480–488, 2015, doi: 10.1037/xge0000057.
[64]
B. A. Primack et al., “Social media use and perceived social isolation among young adults in the U.S. American Journal of Preventive Medicine, vol. 53, no. 1, pp. 1–8, 2017, doi: 10.1016/j.amepre.2017.01.010.
[65]
Y. Kelly, A. Zilanawala, C. Booker, and A. Sacker, “Social media use and adolescent mental health: Findings from the UK millennium cohort study,” EClinicalMedicine, vol. 6, pp. 59–68, 2018, doi: 10.1016/j.eclinm.2018.12.005.
[66]
A. Orben and A. K. Przybylski, “The association between adolescent well-being and digital technology use,” Nature Human Behaviour, vol. 3, pp. 173–182, 2019, doi: 10.1038/s41562-018-0506-1.
[67]
D. Hoyer et al., “Navigating polycrisis: Long-run socio-cultural factors shape response to changing climate,” Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 378, no. 1889, p. 20220402, Sep. 2023, doi: 10.1098/rstb.2022.0402.
[68]
L. Guiso, P. Sapienza, and L. Zingales, “Long-term persistence,” Journal of the European Economic Association, vol. 14, no. 6, pp. 1401–1436, Dec. 2016, doi: 10.1111/jeea.12177.
[69]
G. Tabellini, “Presidential AddressInstitutions and culture,” Journal of the European Economic Association, vol. 6, no. 2–3, pp. 255–294, Apr. 2008, doi: 10.1162/jeea.2008.6.2-3.255.
[70]
A. Alesina and P. Giuliano, “Culture and institutions,” Journal of Economic Literature, vol. 53, no. 4, pp. 898–944, Dec. 2015, doi: 10.1257/jel.53.4.898.
[71]
[72]
A. Bisin and T. Verdier, “The economics of cultural transmission and the dynamics of preferences,” Journal of Economic Theory, vol. 97, no. 2, pp. 298–319, Apr. 2001, doi: 10.1006/jeth.2000.2678.
[73]
D. C. North, J. Wallis, and B. R. Weingast, Violence and social orders: A conceptual framework for interpreting recorded human history. Cambridge University Press, 2009.
[74]
UNICEF, “Hidden in plain sight. A statistical analysis of violence against children,” United Nations Children’s Fund; Universität Tübingen, New York, Jan. 2014. doi: 10.15496/publikation-8598.
[75]
UNICEF, “A familiar face: Violence in the lives of children and adolescents,” United Nations Children’s Fund, New York, 2017.
[76]
End Corporal Punishment, “Global progress towards prohibiting all corporal punishment.” Global Initiative to End All Corporal Punishment of Children, https://endcorporalpunishment.org, 2024.
[77]
World Health Organization, “Global status report on preventing violence against children 2020,” World Health Organization, Geneva, 2020.
[78]
Pew Research Center, “The global divide on homosexuality persists.” Pew Research Center, June 25, 2020, 2020, [Online]. Available: https://www.pewresearch.org/global/2020/06/25/global-divide-on-homosexuality-persists/.
[79]
Arab Barometer, Acceptance of homosexuality ranges from 5% to 26% across surveyed MENA countries. Headline findings presented in BBC News, “The Arab world in seven charts” (24 June 2019), https://www.bbc.com/news/world-middle-east-48703377“Arab barometer wave v (2018–2019).” 2019, [Online]. Available: https://www.arabbarometer.org/surveys/arab-barometer-wave-v/.
[80]
M. R. Kakumba, “Uganda a continental extreme in rejection of people in same-sex relationships,” Afrobarometer, Afrobarometer Dispatch 639, 2023. [Online]. Available: https://www.afrobarometer.org/publication/ad639-uganda-a-continental-extreme-in-rejection-of-people-in-same-sex-relationships/.
[81]
L. R. Mendos, K. Botha, R. C. Lelis, E. López de la Peña, I. Savelev, and D. Tan, “State-sponsored homophobia 2023: Global legislation overview update,” ILGA World, Geneva, 2023.
[82]
Pew Research Center, “The world’s Muslims: Religion, politics and society.” Pew Research Center, April 30, 2013, 2013, [Online]. Available: https://www.pewresearch.org/religion/2013/04/30/the-worlds-muslims-religion-politics-society-overview/.
[83]
Pew Research Center, “Religion in India: Tolerance and segregation.” Pew Research Center, June 29, 2021, 2021, [Online]. Available: https://www.pewresearch.org/religion/2021/06/29/religion-in-india-tolerance-and-segregation/.
[84]
A. Kasirzadeh and I. Gabriel, “In conversation with artificial intelligence: Aligning language models with human values,” Philosophy & Technology, vol. 36, no. 2, p. 27, Apr. 2023, doi: 10.1007/s13347-023-00606-x.
[85]
N. Guha et al., “LegalBench: A collaboratively built benchmark for measuring legal reasoning in large language models,” in Advances in neural information processing systems, 2023, vol. 36, pp. 44123–44279, [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/file/89e44582fd28ddfea1ea4dcb0ebbf4b0-Paper-Datasets_and_Benchmarks.pdf.
[86]
J. Han, J. Lim, I. Kong, and Y. Jo, “Dual mechanisms of value expression: Intrinsic vs. Prompted values in LLMs,” arXiv preprint, Jan. 2025, doi: 10.48550/arxiv.2509.24319.
[87]
J. Uesato et al., “Solving math word problems with process- and outcome-based feedback.” 2022, [Online]. Available: https://arxiv.org/abs/2211.14275.
[88]
H. Lightman et al., “Let’s verify step by step.” 2023, [Online]. Available: https://arxiv.org/abs/2305.20050.
[89]
A. Ovadya et al., “Position: Democratic AI is possible. The democracy levels framework shows how it might work.” in Proceedings of the 42nd international conference on machine learning, 2025, vol. 267, pp. 81930–81961, [Online]. Available: https://proceedings.mlr.press/v267/ovadya25a.html.
[90]
K. Yeung, “Recommendation of the council on artificial intelligence (OECD),” International Legal Materials, vol. 59. pp. 27–34, Feb. 01, 2020, doi: 10.1017/ilm.2020.5.
[91]
European Parliament and Council of the European Union, “Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act).” Official Journal of the European Union, L Series, 2024, [Online]. Available: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng.
[92]
J. Różycka-Tran, P. Boski, and B. Wojciszke, “Belief in a zero-sum game as a social axiom,” Journal of Cross-Cultural Psychology, vol. 46, no. 4, pp. 525–548, Mar. 2015, doi: 10.1177/0022022115572226.
[93]
P. Justino and M. Samarin, “Trust in a changing world: Social cohesion and the social contract in uncertain times,” United Nations University World Institute for Development Economics Research; UNU-WIDER, Helsinki, Finland, 34, May 2025. doi: 10.35188/unu-wider/2025/591-2.
[94]
W. MacAskill, What we owe the future. Basic Books, 2022.
[95]
A. Arzberger, C. Offerman, U. Gadiraju, A. Bozzon, and J. Yang, “"Label from somewhere": Reflexive annotating for situated AI alignment.” arXiv preprint arXiv:2601.17937, Jan. 25, 2026, [Online]. Available: http://arxiv.org/abs/2601.17937v2.
[96]
E. Yudkowsky, “Coherent extrapolated volition.” Singularity Institute for Artificial Intelligence, 2004.
[97]
J. Rawls, A theory of justice: Original edition. Harvard University Press, 1971.
[98]
G. Tabellini, “Culture and institutions: Economic development in the regions of europe,” Journal of the European Economic Association, vol. 8, no. 4, pp. 677–716, Jun. 2010, doi: 10.1111/j.1542-4774.2010.tb00537.x.
[99]
N. Nunn and L. Wantchekon, “The slave trade and the origins of mistrust in africa,” American Economic Review, vol. 101, no. 7, pp. 3221–3252, Dec. 2011, doi: 10.1257/aer.101.7.3221.
[100]
K. Hoff and J. E. Stiglitz, “Equilibrium fictions: A cognitive approach to societal rigidity,” American Economic Review, vol. 100, no. 2, pp. 141–146, May 2010, doi: 10.1257/aer.100.2.141.
[101]
K. Hoff and J. E. Stiglitz, “Striving for balance in economics: Towards a theory of the social determination of behavior,” Journal of Economic Behavior & Organization, vol. 126, pp. 25–57, Mar. 2016, doi: 10.1016/j.jebo.2016.01.005.
[102]
[103]
R. D. Putnam, Bowling alone: The collapse and revival of American community. Simon & Schuster, 2000.
[104]
R. Turuba et al., “Exploring how youth use TikTok for mental health information in british columbia: Semistructured interview study with youth,” JMIR Infodemiology, vol. 4, p. e53233, Jul. 2024, doi: 10.2196/53233.
[105]
J. Frey, K. J. Black, and I. A. Malaty, “TikTok tourette’s: Are we witnessing a rise in functional tic-like behavior driven by adolescent social media use?” Psychology Research and Behavior Management, p. 359977, Jan. 2022, doi: 10.2147/prbm.s359977.
[106]
M. Zeng, J. Grgurevic, R. Diyab, and R. Roy, “#WhatIEatinaDay: The quality, accuracy, and engagement of nutrition content on TikTok,” Nutrients, vol. 17, no. 5, p. 781, Feb. 2025, doi: 10.3390/nu17050781.
[107]
R. Mulcahy, R. Barnes, R. de Villiers Scheepers, S. Kay, and E. List, “Going viral: Sharing of misinformation by social media influencers,” Australasian Marketing Journal, vol. 33, no. 3, pp. 296–309, Aug. 2024, doi: 10.1177/14413582241273987.
[108]
W. López López, M. ’ia. A. Roa Bocarejo, D. Roa Peralta, C. Pineda Mar ’in, and E. Mullet, “Mapping colombian citizens’ views regarding ordinary corruption: Threat, bribery, and the illicit sharing of confidential information,” Social Indicators Research, vol. 133, no. 1, pp. 259–273, 2017, doi: 10.1007/s11205-016-1366-6.
[109]
E. C. C. Chang and N. N. Kerr, “An insider–outsider theory of popular tolerance for corrupt politicians,” Governance, vol. 30, no. 1, pp. 67–84, Feb. 2016, doi: 10.1111/gove.12193.
[110]
A. Megías, L. de Sousa, and F. Jiménez-Sánchez, “Deontological and consequentialist ethics and attitudes towards corruption: A survey data analysis,” Social Indicators Research, vol. 170, no. 2, pp. 507–541, Sep. 2023, doi: 10.1007/s11205-023-03199-2.
[111]
N. Letki, M. A. Górecki, and A. Gendźwiłł, ‘They accept bribes; we accept bribery’: Conditional effects of corrupt encounters on the evaluation of public institutions,” British Journal of Political Science, vol. 53, no. 2, pp. 690–697, 2023, doi: 10.1017/S0007123422000047.
[112]
A. A. Demarest, Ancient Maya: The Rise and Fall of a Rainforest Civilization. Cambridge University Press, 2004.
[113]
A. J. Dugmore, T. H. McGovern, O. Vésteinsson, J. Arneborg, R. Streeter, and C. Keller, “Cultural adaptation, compounding vulnerabilities and conjunctures in norse greenland,” Proceedings of the National Academy of Sciences, vol. 109, no. 10, pp. 3658–3663, 2012, doi: 10.1073/pnas.1115292109.
[114]
T. H. McGovern, Management for extinction in norse greenland,” in The anthropology of climate change, John Wiley & Sons, Ltd, 2014, pp. 131–150.
[115]
J. A. Goldstone, “Demographic structural theory: 25 years on,” Cliodynamics: The Journal of Quantitative History and Cultural Evolution, vol. 8, no. 2, Dec. 2017, doi: 10.21237/c7clio8237450.
[116]
G. Orlandi et al., “Structural-demographic analysis of the qing dynasty (1644–1912) collapse in china,” PLOS ONE, vol. 18, no. 8, p. e0289748, Aug. 2023, doi: 10.1371/journal.pone.0289748.
[117]
J. Stanley, M. D. Krom, R. A. Cliff, and J. C. Woodward, “Short contribution: Nile flow failure at the end of the old kingdom, egypt: Strontium isotopic and petrologic evidence,” Geoarchaeology, vol. 18, no. 3, pp. 395–402, Feb. 2003, doi: 10.1002/gea.10065.
[118]
C. Johnson, MITI and the japanese miracle. Stanford University Press, 1982.
[119]
A. H. Amsden, Asia’s next giant: South korea and late industrialization. Oxford University Press, 1989.
[120]
[121]
R. H. Wade, “The developmental state: Dead or alive?” Development and Change, vol. 49, no. 2, pp. 518–546, Jan. 2018, doi: 10.1111/dech.12381.
[122]
M. Andrews, L. Pritchett, and M. Woolcock, Building state capability: Evidence, analysis, action. Oxford University Press, 2017.
[123]
T. Donais, “Empowerment or imposition? Dilemmas of local ownership in post‐conflict peacebuilding processes,” Peace & Change, vol. 34, no. 1, pp. 3–26, Jan. 2009, doi: 10.1111/j.1468-0130.2009.00531.x.
[124]
J. R. Böhnke, J. Koehler, and C. M. Zürcher, “State formation as it happens: Insights from a repeated cross-sectional study in afghanistan, 2007–2015,” Conflict, Security & Development, vol. 17, no. 2, pp. 91–116, 2017, doi: 10.1080/14678802.2017.1292681.
[125]
N. Nunn, “The importance of history for economic development,” Annual Review of Economics, vol. 1, no. 1, pp. 65–92, Apr. 2009, doi: 10.1146/annurev.economics.050708.143336.
[126]
N. Voigtländer and H.-J. Voth, “Persecution perpetuated: The medieval origins of anti-semitic violence in nazi germany*,” The Quarterly Journal of Economics, vol. 127, no. 3, pp. 1339–1392, Jul. 2012, doi: 10.1093/qje/qjs019.
[127]
I. BERLIN, The crooked timber of humanity: Chapters in the history of ideas - second edition, REV - Revised, 2, Second Edition. Princeton University Press, 2013.
[128]
J. D. Sachs, “Institutions matter, but not for everything,” Finance and development, vol. 40, no. 2, pp. 38–41, Jun. 2003, doi: 10.5089/9781451952926.022.a012.
[129]
[130]
J. Mokyr, A culture of growth: The origins of the modern economy. Princeton University Press, 2017.