Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols


Abstract

As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures at scale. We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led). Analyzing 4,323 governance participation records, we combine LLM-assisted coding, topic modeling, and multi-layer network analysis to examine how institutional design shapes thematic priorities and community structure. We find that while governance form influences substantive focus, both regimes exhibit comparable levels of participation inequality and community fragmentation. Discourse alignment is denser in the permissionless setting, suggesting that open governance may foster greater thematic convergence despite decentralized participation. These findings illustrate how LLM-assisted methods can advance the empirical study of technology governance, with implications for designing more equitable agentic AI standards. All data and code are openly available.

1 Introduction↩︎

Interoperability is an inherited characteristic of decentralized applications [1]. As AI agent protocols proliferate, the governance of their interoperability standards remains empirically underexamined. Who controls the rules by which autonomous agents discover, negotiate, and coordinate across organizational boundaries? This question sits at the intersection of artificial intelligence and institutional design [2], [3], yet the public discourse through which such standards are negotiated has been difficult to study at scale. Prior work relies predominantly on manual coding with fixed categories, limiting both thematic discovery and structural analysis across large text corpora.

We address this gap with an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis. We validate the pipeline on two standards: ERC-8004 and Google A2A. The two protocols address the same technical problem of AI agentic communication across systems, and their governance processes are both public on GitHub or Forums. Therefore, comparing them isolates the effect of governance form on who participates and what they discuss.

ERC-8004, which belongs to Ethereum Improvement Proposals (EIP), defines a smart-contract interface standard for on-chain AI agent identity. It “proposes to use blockchains to discover, choose, and interact with agents across organizational boundaries without pre-existing trust” [4]. Its lifecycle follows the stages defined in EIP-1, and is advanced by rough consensus among forum participants and EIP editors, with no formal votes and permissionless implementations [5]. Google A2A introduces “An open protocol enabling communication and interoperability between opaque agentic applications” [6]. It was initiated by Google and donated to the Linux Foundation in June 2025, now governed by an eight-seat Technical Steering Committee (TSC) composed entirely of corporate representatives [7], [8].

Comparing the two cases, we ask:

RQ: Compared to corporate hierarchy, does the governance structure of a permissionless DAO really achieve a higher degree of decentralization?

We decompose this into three sub-questions:

  • RQ1 (Decision Architecture): How do the formal decision procedures, entry rights, and authority structures of the two regimes differ?

  • RQ2 (Discourse Composition): How does governance form shape the topical and argumentative composition of participation discourse?

  • RQ3 (Relational Networks): How does governance form shape co-participation ties, discourse-level consensus and conflict structures, and actor–topic divisions of collaborative labor?

To answer these questions, we first reconstruct the two governance mechanisms from their official documents and visualize them as decision-flow diagrams. We then collect 4,323 governance participation records from their public repositories and apply an LLM-assisted annotation pipeline to label each record’s argumentative function, stance, and stakeholder affiliation [9]. On top of this annotation, we run two topic-discovery methods: unsupervised BERTopic [10], and LLM-inductive Thematic-LM [11]. Moreover, we run three network analyses: co-participation social network analysis (SNA) [12], discourse network analysis [13], and socio-semantic bipartite networks [14]. The methodological triad mirrors the three sub-questions.

Three findings emerge. First, the two cases instantiate structurally opposite decision architectures: ERC-8004 advances by rough consensus with permissionless deployment, while A2A vests binding authority in a corporate technical steering committee. Second, governance form shapes the thematic focus of deliberation: DAO governance concentrates discourse on security mechanisms and protocol principles, whereas corporate governance distributes deliberation across engineering-execution workstreams. Third, both networks exhibit comparable levels of participation inequality and community fragmentation; discourse congruence is nonetheless denser in ERC-8004, reflecting tighter within-community consensus formation.

This paper makes three contributions. Methodologically, the architecture can generalize to any large-scale governance discourse analysis, offering a computational lens on the black box of public text negotiation, communication and coordination. Empirically, it provides the first matched-case, multi-method comparison of DAO and corporate governance of agentic standardization with text corpora at scale. Theoretically, it pioneers the fundamental problem of who controls the future AI agentic infrastructure, with the interdisciplinary standing of institutional design and AI.

2 Related Work↩︎

Three streams of literature converge on this study. For a more specific comparison, see Appendix 13.

Decentralized governance via blockchain. First, Harvey provides a comprehensive overview of the technologies, theories, and practices inherent to decentralized finance (DeFi) and blockchain technology [15]. Beck et al. [16] provided the foundational Information System framework mapping blockchain governance along decision rights, accountability and incentives dimensions. Ziolkowski et al.  [17] identified six major governance challenges specific to the blockchain system, and half of them are unique to this technology. Kiayias and Lazos [18] systematised the field with a SoK review of blockchain governance mechanisms. More recently, Ellinger et al. [19] explored the balanced polycentric governance of digital commons. Reineke et al. [20] developed an integrative theoretical framework, showing the evolution of decentralization concepts. Sunyaev et al. [21] called for purposeful, design-oriented decentralization rather than ending at ideology. Motea and Oba [22] interrogated the democratic legitimacy of blockchain governance structures.

Governance of decentralized versus corporate structure. Murray et al. [23] examined how smart contracts and DAOs alter agency costs in corporate contracting, providing leading opinions. Lumineau et al. [24] highlights tacitness of transactions and social implications of this technology. Rahman et al. [25] and Hunt et al. [26] analyzed the dynamics and power accumulation of platforms. Hui and Tucker [27] address AI along with a decentralized ecosystem, proposing an innovative governance framework. However, these studies rely on theoretical or interview-based methods; none of them use governance-participation data or computational methods to test structural difference between DAO and corporate governance in the same domain.

Computational studies of cooperative work. Researchers have long examined coordination in open, online communities. Mockus et al. [28] documented participation inequality in Apache and Mozilla. Im et al. [29] analyzed Wikipedia’s Requests for Comments (RfC), a deliberation mechanism structurally analogous to EIP rough consensus, and revealed a persistent imbalance between deliberation and resolution. Germonprez et al. [30] catalogued structural realities of contemporary open-source projects, including pervasive corporate engagement; Li et al. [31] examined code-of-conduct conversations on GitHub as a window into informal governance norms in open-source repositories. Kulakowski and Frasincar [32] introduced CryptoBERT, a domain-adapted BERT variant pre-trained on 3.2 million cryptocurrency social-media posts, establishing a specialized embedding backbone for blockchain-native corpora. More recently, Wu et al. [33] applied sentiment and discourse analysis to six DAO forums; Stine and Agarwal [34] proposed comparative discourse analysis via topic models; Qiao et al. [11] introduced Thematic-LM for LLM-assisted inductive thematic analysis of large corpora; Ao et al. [12] used social network analysis on on-chain Aave data to show voting-power concentration; Leifeld [13] proposed discourse network analysis concentrating on participants’ stance; Roth and Cointet [14] connects semantic analysis with social topology. Wang el al.  [35] found the threat of power concentration via scaled empirical analysis of scale SnapShot data. Özdemir Sönmez et al. [36] quantified voting power concentration and participation apathy across DAO governance models, especially in token-based organizations. Chen et al. [37] applied a quasi-experimental PSM-DID design to 98,000 Steemit users, finding that governance token ownership boosts curation effort but reduces creation novelty.

3 Methodology↩︎

We selected MiniMax-M2.5 [38] as the LLM backbone for its reasoning capability and low cost. The model is trained on complex real-world environments and achieves 80.2% on SWE-Bench Verified; frontier models of this capability tier have been shown to complete complex multi-turn analytical tasks with high reliability [39], making them suitable for assigning nuanced governance-process labels such as Argument Type and Consensus Signal. For design of the comparative case, see Appendix 7.

3.1 Data Collection and LLM Annotation↩︎

We collect governance participation records from the public repositories. For ERC-8004, data originate from two sources: (1) 113 posts from the Ethereum Magician forum thread [4]; and (2) 36 GitHub records from the nine pull requests that directly modified ERCS/erc-8004.md [40]. For Google A2A, data originate from three streams in the a2aproject/A2A repository [6]: (1) 3,104 issue and issue-comment records, (2) 1,955 pull requests and review-comment records, and (3) 822 GitHub Discussion records.

The raw records totaled 6,030. We removed the following data: records with fewer than 20 characters of body text, which primarily are CI notifications, merge-conflict markers, and bot-generated status messages; and records attributed to verified bot accounts. After such filtering, 4,323 records are retained (ERC-8004: 142; Google A2A: 4,181). SHA-256 checksums for all raw data fields are provided in the authors’ GitHub repository 1. See also Appendix 8.

The retained records was annotated with four categorical fields:

  • Stakeholder Institution: Google / MetaMask / Ethereum Foundation / Coinbase / Independent / Unknown

  • Argument Type: Technical / Governance-Principle / Economic / Process / Off-topic

  • Stance: Support / Oppose / Modify / Neutral / Off-topic

  • Consensus Signal: Adopted / Rejected / Pending / N/A

The institution affiliations of the top 109 contributors were also manually reviewed. The cascade is detailed in Appendix 11.

3.2 Discourse Composition Analysis↩︎

To characterize the topical and argumentative composition of the two governance discourses, we apply three methods of increasing inductive depth.

3.2.1 Supervised argument typing↩︎

Treating the LLM-assigned argument_type as the record’s communicative function, we test cross-case independence with a chi-square test [41] and within-ERC-8004 temporal change across three consecutive two-month phases spanning submission to mainnet deployment (phase boundaries in Appendix 9).

3.2.2 BERTopic comparative discourse analysis↩︎

Following Stine and Agarwal [34], we fit BERTopic [10] jointly on the combined corpus. Texts are embedded with all-MiniLM-L6-v2 [42], reduced via UMAP (\(n_{\text{neighbors}}=15\), cosine, seed\(=42\)), and clustered with HDBSCAN (min cluster size \(10\)) into \(K=19\) topics plus a noise class. For each topic \(t\) and case \(c\), the within-case share is \[p_t^{(c)} = n_t^{(c)}/n^{(c)}.\] Cross-case divergence is measured by Jensen–Shannon divergence [43]: \[\mathrm{JSD}(p,q) = \tfrac{1}{2}\mathrm{KL}(p,m) + \tfrac{1}{2}\mathrm{KL}(q,m)\] with \[\mathrm{KL}(p, m) = \sum_{i} p(i) \log \frac{p(i)}{m(i)}, \qquad m = \tfrac{1}{2}(p+q),\] and \(\mathrm{JSD}\!=\!0\) denotes identical distributions, \(\mathrm{JSD}\!=\!1\) disjoint.

To assess whether ERC-8004’s topic concentration is an artefact of a general-purpose embedding model rather than a genuine structural property, we also embed the 142 ERC-8004 records with CryptoBERT, a model further pre-trained on cryptocurrency social-media posts (StockTwits, Reddit, Telegram, Twitter) [32].

3.2.3 Thematic-LM inductive themes↩︎

To complement the embedding based view with human-interpretable labels, we apply the method of Thematic-LM [11], a four-stage multi-agent pipeline: (1) open coding assigns a short code to each record; (2) aggregation groups 300 sampled codes into 14 raw clusters; (3) codebook review merges and refines to 19 themes (T01–T19); (4) theme assignment labels every record with its best-fitting theme, or Unclassified when confidence is insufficient. We again report per-theme shares and JSD between cases.

3.3 Relational Network Analysis↩︎

We construct three complementary networks for each case, each layer adding discursive information to the previous. All graphs use actors as nodes; what changes is the semantic content of an edge.

3.3.1 Co-participation network (SNA)↩︎

Following Ao et al. [12], an undirected edge connects two contributors who both posted to the same discussion thread (forum topic, GitHub issue, pull request, or discussion); edge multiplicity counts co-occurrences. We report six structural measures:

  • Density: the realized fraction of possible ties. Computed by \[\rho = \frac{2E}{N(N-1)}\] where \(E\) denotes the number of edges and \(N\) denotes the number of nodes.

  • Degree Gini: inequality of interaction counts (\(d_i\) is node degree). Computed by \[G = \frac{\sum_{i,j}|d_i - d_j|}{2N\sum_i d_i}\] where \(d_i\) is the number of edges of node \(i\). \(G = 0\) implies perfectly equal participation; \(G \to 1\) implies that interactions are concentrated in a small elite.

  • Components and giant-component ratio: fragmentation into disconnected threads. Computed by \[\mathrm{GCR} = N_{\max}/N.\]

  • Newman–Girvan modularity: reported for institution partitions and Louvain partitions [44] [45]. Computed by \[Q = \frac{1}{2m} \sum_{i,j} \left[ A_{ij} - \frac{k_i k_j}{2m} \right] \delta(c_i, c_j)\] where \(m\) is the total number of edges, \(A_{ij}\) is the adjacency matrix entry, \(k_i\) is the degree of node \(i\), and \(\delta(c_i, c_j) = 1\) if nodes \(i\) and \(j\) belong to the same community, 0 otherwise.

  • Core–periphery: Borgatti–Everett coreness [46]. Computed by \[\rho_{\mathrm{BE}} = \mathrm{corr}(A,\,\Delta), \qquad \Delta_{ij} = \delta_i \cdot \delta_j\] where \(\delta_i \in \{0,1\}\) is the coreness label of node \(i\). Statistical significance is evaluated by comparing the observed \(\rho_{\mathrm{BE}}\) against a null distribution of random graphs that preserve the empirical degree sequence (configuration model), following the \(q\)-\(s\) test of Kojaku and Masuda [47].

  • Betweenness centrality and network efficiency: identifies governance brokers—actors who mediate between otherwise disconnected groups: \[b(v) = \sum_{s \neq v \neq t} \frac{\sigma_{st}(v)}{\sigma_{st}}\] where \(\sigma_{st}\) is the total number of shortest \(s\)\(t\) paths and \(\sigma_{st}(v)\) the number that pass through \(v\).

  • Network efficiency: mean normalized harmonic centrality, measured by \[\bar{h} = \frac{1}{n(n-1)}\sum_{u \neq v} d(u,v)^{-1}\] where \(d(u,v)^{-1}=0\) for unreachable pairs [48].

3.3.2 Discourse Network Analysis (DNA)↩︎

Following Leifeld [13], we construct the co-participation layer with stance-aware edges that capture shared discursive positions. Let \(A\) be the actor set and \(T=\{T_{01},\ldots,T_{19}\}\) the Thematic-LM codebook. For each record we encode stance into Support\(=+1\), Modify\(=+0.5\), Neutral\(=0\), Oppose\(=-1\) and obtain an actor–theme stance matrix \(M\) similar to the pseudo-matrix below:

Theme A Theme B Theme C
Alice \(+1\) \(-1\)
Bob \(+1\) \(+0.5\) \(0\)
Carol \(-1\) \(+1\)

Two networks are projected from \(M\): the congruence network \(G^+\) connects actors who take same-sign stances on at least one shared theme (Alice and Bob agree on Theme A), and the conflict network \(G^-\) connects actors who take strictly opposite-sign stances (Alice and Bob conflict on Theme B); edge weights count the number of themes satisfying each criterion. We report density, Louvain modularity and betweenness centrality \(b(v)\) of the top-5 discourse brokers.

3.3.3 Socio-semantic bipartite network↩︎

Following Roth and Cointet [14], we construct a two-mode (bipartite) network \(\mathcal{B} = (A \cup T,\, E)\), where \(A\) is the set of actors, \(T\) the set of Thematic-LM themes, and an edge \((a, t) \in E\) exists whenever actor \(a\) authored at least one record assigned to theme \(t\). The weight \(B_{at}\) equals the number of such records, forming a non-negative integer matrix \(B \in \mathbb{Z}_{\ge 0}^{|A|\times|T|}\).

Two one-mode projections are derived from \(B\). The actor–actor projection \(W^A = BB^\top\) links any two actors by the number of themes they co-discussed (self-loops removed); the theme–theme projection \(W^T = B^\top B\) links any two themes by the number of actors who engaged in both.

Per-actor topic diversity is measured by Shannon entropy: \[H(a) = -\sum_{t=1}^{|T|} \hat{p}_{at} \log_2 \hat{p}_{at}, \qquad \hat{p}_{at} = \frac{B_{at}}{\textstyle\sum_{t'} B_{at'},} \label{eq:entropy}\tag{1}\] where \(H(a)=0\) for a pure specialist (all posts in one theme) and \(H(a)=\log_2|T|\) for a perfect generalist (uniform spread across all \(|T|\) themes).

We further characterize the distribution of \(\{H(a)\}_{a \in A}\) via its Gini coefficient. The higher the Gini, the more thematic breadth is concentrated in a few actors. We also measure per-theme actor concentration as the Gini of the column vector \(\{B_{at}\}_{a \in A}\) for each \(t\).

The thematic overlap coefficient is measured by \[\Omega = \frac{|T_1 \cap T_2|}{\min(|T_1|, |T_2|)}, \label{eq:overlap}\tag{2}\] where \(T_c\) denotes the active theme set of case \(c \in \{1,2\}\), quantifies cross-case thematic alignment; \(\Omega = 1\) means one community’s thematic space is a subset of the other’s.

Actor-set sizes across the three filtering stages are detailed in Appendix 12.

4 Results↩︎

4.1 Decision Architectures↩︎

Figure 1: Left: ERC-8004 governance decision flow. Right: A2A governance structure and decision flow.

ERC-8004 and Google A2A represent contrasting governance arche-types, as detailed in figure 1. ERC-8004 remains at the stage of “draft” when the data was fetched (Mar 2026), nevertheless, its three canonical registry contracts were deployed on Ethereum mainnet on January 29, 2026 [49]. Proposals are discussed through the Ethereum Magicians forum and are recorded on GitHub. Ownership of Google A2A transitioned from Google-controlled to Linux Foundation governance, and Contested decisions result in a GitVote [50]. For more details and pseudocode, see Appendix 10.

4.2 Discourse Composition↩︎

4.2.1 Supervised argument typing↩︎

Figure 2: Left: argument-type distribution of both cases. Right: ERC-8004 argument-type distribution across three two-month phases. Technical arguments dominate in both regimes, but corporate governance carries roughly double the procedural coordination burden.
Figure 3: Stance \times argument-type cross-tabulation (% within stance row).

Figure 2 compares the LLM-assigned argument type distributions. Both communities are predominantly Technical (74.3% in ERC-8004 vs.% in A2A), but A2A devotes almost twice the share to Process arguments (25.4% vs.%). The cross-case difference is significant with a small effect size (\(\chi^2(3) = 52.88\), \(p < .001\), Cramér’s \(V = .103\)): both cases remain technically grounded, but corporate governance carries substantively heavier coordination overhead.

Within ERC-8004, argument composition shifts significantly across three two-month phases (\(\chi^2(6)=25.32\), \(p<.001\), \(V=.315\)). Technical arguments dominate Phases 1–2 (\(\geq\)​80%); in Phase 3 Process discussion surges to 53% as deliberation moves from substantive design to editorial ratification. This two-stage pattern is consistent with rough-consensus norms in open standards bodies. Small-\(n\) caution applies to Phases 2–3 (\(n=11\) and \(n=17\)).

Figure 3 cross-tabulates stance with argument type. Support and Modify stances are dominated by Technical arguments in both cases (ERC Support 73%; A2A Support 87%): substantive engagement is predominantly technical regardless of governance form. Governance-Principle arguments, particularly under Oppose, are more visible in ERC-8004, reflecting the principled deliberation characteristic of the EIP process. The sharpest cross-case contrast appears in the Neutral row: 44% of A2A Neutral records are Process in nature, against 26% in ERC-8004—corporate governance generates procedural commentary even among non-committal participants.

4.2.2 BERTopic↩︎

Figure 4: BERTopic cross-case divergence. ERC-8004 concentrates on foundational agent architecture, while A2A distributes deliberation across engineering-execution workstreams.

BERTopic identifies 19 topics plus a noise class over the combined corpus. The global Jensen–Shannon divergence between the two cases’ topic distributions is \(\mathrm{JSD}_{\text{BERTopic}}=0.288\), a moderate but meaningful structural separation (Figure 4). ERC-8004 is strikingly concentrated: 67.6% of its records fall into Topic 0 (agent, agents, for, of), the broad agent-discourse cluster, versus 21.2% for A2A. Implementation-heavy topics—Task/Message management (8.2%), JSON/proto spec (7.7%), PR-contribution workflows (7.2%), SDK samples (5.9%)—carry substantial A2A weight but zero ERC-8004 records, confirming that low-level engineering is absent from the EIP forum. The only topic with a higher ERC share than A2A is Topic 9 (suggestion, sap, linkedin; 2.8% vs.%), signalling corporate voices (SAP, LinkedIn) entering the EIP thread.

CryptoBERT on ERC data shows more precise results. It divides the "agent" topic to T0 (Onchain, Reputation) and T2 (Trust, Feedback), and the total percentage is 73.2%, which aligns with the 67.6% of pure BERTopic (figure 5).

Figure 5: CryptoBERT topic frequency for ERC-8004.

4.2.3 Thematic-LM↩︎

The Thematic-LM pipeline yields a 19-theme codebook (T01–T19). Figure 6 overlays per-case record share (diverging bars) with per-theme actor participation rate (lines with markers). Table 1 reports the six themes with the largest cross-case divergence. The global Jensen–Shannon divergence is \(\mathrm{JSD}=0.216\).

The most discriminating theme is T08 Trust & Security Mechanisms. It covers agent-trust scoring, reputation schemes and on-chain credential verification—the foundational problem of establishing trustworthy agency. By contrast, A2A spreads deliberation across T06 Documentation & Examples, T18 Clarifications & Information Requests, T07 Community Collaboration & Contributions, and T01 Protocol Specification & Versioning. A2A additionally covers three themes: T09 (Transport & Protocol Mechanisms), T14 (Project Governance & Process), and T16 (Streaming & Real-time Communication). T09 and T16 are purely engineering-execution concerns entirely absent from the EIP forum; T14’s ERC-8004 records were authored by automated review bots and thus fall outside the actor network.

The two methods agree on the direction and approximate magnitude of divergence, with BERTopic’s \(\mathrm{JSD}=0.288\) slightly exceeding Thematic-LM’s \(\mathrm{JSD}=0.216\). The gap is expected: Thematic-LM compresses divergence by absorbing related concerns into shared conceptual themes, while BERTopic retains finer-grained lexical distinctions (ERC’s “trustless”, “reputation” versus A2A’s “samples”, “proto”). Together the two estimates bracket the structural divergence and confirm its robustness: the two governance processes occupy meaningfully different but overlapping discourse spaces, with ERC-8004 concentrated on the constitutive layer of the design space (“what to build and why”) and A2A on executive ones (how to build, document, and ship it).

Table 1: Thematic-LM themes with largest cross-case divergence. The result is aligned with figure [fig:bertopic-divergence].
ID Theme ERC % A2A % \(\Delta\)
T08 Trust & Security 34.5 4.0 \(+30.5\)
T01 Protocol Specification 13.4 7.9 \(+5.5\)
T06 Documentation & Examples 2.1 10.8 \(-8.7\)
T05 SDK Development 0.7 4.6 \(-3.9\)
T15 Tooling & Automation 0.7 4.5 \(-3.8\)
T09 Transport Mechanisms 0.0 3.2 \(-3.2\)
Figure 6: Thematic-LM overlaid chart. Bars show per-case record share; lines show actor participation rate. The DAO concentrates deliberation on security and trust; the corporate project distributes effort across documentation, tooling, and engineering-execution themes.

4.3 Relational Networks↩︎

4.3.1 Co-participation network (SNA): core and periphery↩︎

Table 2 reports structural metrics for the co-participation networks (ERC-8004: \(N=67\); Google A2A: \(N=771\)). For network visualization, see 7.

Figure 7: Your network diagram caption here.

Both communities exhibit a high degree of inequality. The Gini coefficient of degree is 0.804 for ERC-8004 and 0.779 for A2A, indicating that in both cases, a small elite drives the majority of interactions. The top-3 contributors account for 32.3% of all ERC-8004 interactions and 14.9% of A2A interactions. In A2A, two out of three of Google employees, and another is from Microsoft. In ERC-8004, the three contributors are: MarcoMetaMask (Marco De Rossi), who obviously by name is from Metamask; spengrah (Spencer Graham), cofounder of Hats Protocol; and pcarranzav (Pablo Carranza Vélez), who is from The Graph, a decentralized protocol for indexing and querying blockchain data. Their formal organizational mandate is indeterminate from public records, but their persistent participation accrues reputational standing.

Both networks are structurally fragmented, and neither shows statistically significant core-periphery structure. Louvain community detection recovers 46 communities for ERC-8004 and 358 for A2A, closely matching the component counts and confirming that participation organizes around parallel threads rather than a coherent deliberative body. Institution labels also do not predict interaction structure in either case. The giant component covers 32.8% of ERC-8004 nodes and 53.4% of A2A nodes; 333 of 771 A2A participants (43.2%) are complete isolates with no co-participation edges. Density indicates that the vast majority of possible co-participation pairs are unrealized. Under the Borgatti-Everett test [46], ERC-8004 returns \(p = .095\) and A2A returns \(p = 1.000\); neither crosses the \(\alpha = .05\) threshold.

Table 2: Governance Network Structural Metrics. Despite structurally opposite decision architectures, both networks exhibit comparably high participation inequality: a small elite drives most interactions regardless of governance form.
Metric ERC-8004 Google A2A
Nodes 67 771
Edges 65 1,230
Density 0.029 0.004
Gini (degree) 0.804 0.779
Top-3 degree share 32.3% 14.9%
Components 43 346
Giant component ratio 0.328 0.534
Modularity—institution \(-\)0.059 \(-\)0.034
Modularity—Louvain 0.425 0.473
Louvain communities 46 358
CP \(p\)-value (BE test) 0.095 1.000
CP significant (\(p<.05\)) No No
Top-1 betweenness centrality 0.069 0.136
Betweenness Gini 0.931 0.979
Top-3 betweenness share 70.8% 48.5%
Network efficiency \(\bar{h}\) 0.050 0.110

4.3.2 Discourse network: congruence and conflict↩︎

The stance layer (Table 3) reveals structures that are invisible in the co-participation graph.

ERC-8004 achieves denser within-community consensus, while A2A generates far greater absolute conflict volume. Congruence density is \(0.148\) for ERC-8004 versus \(0.082\) for A2A: within the tighter EIP community, participants more often share positions. Conflict edges in A2A are 34\(\times\) larger in absolute count (2,531 vs.), which is not only an effect of actor numbers, but also from the broader range of technical positions supported by a multi-vendor engineering project.

The top-3 betweenness share in the congruence network drops to \(34.5\%\) for ERC-8004 and \(12.2\%\) for A2A, compared with \(70.8\%\) and \(48.5\%\) in the co-participation network (Table 2). Discourse brokerage is more distributed than structural interaction brokerage in both cases; yet ERC-8004’s congruence network remains markedly more concentrated, consistent with a compact EIP core mediating contested positions on trust and protocol security.

Table 3: Discourse Network Analysis metrics. The DAO achieves denser within-community agreement; the corporate regime generates a much larger volume of conflict edges, reflecting the broader technical surface of a multi-vendor engineering project.
Metric ERC-8004 Google A2A
Actors 66 710
Active themes 17 / 19 19 / 19
Congruence edges 318 20,638
Congruence density 0.148 0.082
Congruence modularity (Louvain) 0.2886 0.2453
Top-1 betweenness centrality 0.061 0.018
Betweenness Gini 0.859 0.889
Top-3 betweenness share 34.5% 12.2%
Conflict edges 74 2,531
Mean actor theme diversity 1.470 2.211

4.3.3 Socio-semantic bipartite network: thematic concentration of discursive labor↩︎

The actor–theme bipartite layer (Table 4, Figure 8) reveals where each governance community directs its deliberative effort. Peripheral participation is the structural baseline in both cases: median actor Shannon entropy \(H = 0\) in both ERC-8004 and A2A, meaning the majority of contributors engage a single theme across their entire record [28]. This is not a governance-specific feature but an intrinsic property of large-scale deliberative systems.

Figure 8: Per-actor Shannon entropy over Thematic-LM themes. H=0 marks pure specialists; the H=0 bin dominates both distributions.
Table 4: Socio-semantic bipartite network metrics. Both regimes share the same thematic space, but governance form drives which themes absorb deliberative effort and how tightly specialized those subgroups become.
Metric ERC-8004 Google A2A
Actors 66 710
Active themes 17/19 19/19
Actor–actor projection edges 645 59,007
Mean actor entropy \(H\) 0.348 0.617
Median actor entropy \(H\) 0 0
Gini\((H)\) 0.773 0.707
Max \(H\) 2.664 3.834
Mean theme Gini (actor concentration) 0.085 0.453
Most-discussed theme T08 T06
Thematic overlap \(\Omega\) 1.000

What differs across governance forms is which themes absorb deliberative effort, and how broadly the most active participants range. Mean actor entropy is \(0.348\) for ERC-8004 and \(0.617\) for A2A: A2A’s core contributors span roughly twice as many themes on average. Gini\((H)\) remains high in both cases (\(0.773\) vs.\(0.707\)), confirming that thematic breadth is itself concentrated in a small number of generalist actors in both communities.

The cross-case divergence is sharpest at the theme level. In ERC-8004, \(34.5\%\) of actors participated in T08 (Trust & Security Mechanisms), and \(13.1\%\) participated in T01 (Protocol Specification & Versioning). For A2A, the pattern inverts for engineering-execution themes: Documentation & Examples (T06) engaged \(10.8\%\) of A2A actors but only \(2.1\%\) of ERC-8004 actors. Mean theme actor-concentration Gini is \(0.453\) for A2A against \(0.085\) for ERC-8004, indicating that each A2A theme attracts a narrower, more dedicated subgroup. The thematic overlap coefficient \(\Omega = 1.0\): all 16 active ERC-8004 themes reappear in A2A.

5 Discussion↩︎

5.1 Decentralization More a Design than a Fact↩︎

Decentralized governance is appealing for its idealistic design of distributed authority and trustless coordination. In practice, however, the ERC-8004 process exhibits a centralization paradox.

Routine decisions are not subject to voting. The power to decide, for example, which proposals are “ready”, which objections are “substantive”, which positions count as “rough consensus”, accumulates in the hands of a few minorities, mirroring the participation inequality patterns repeatedly documented in voluntary online cooperation [28], [30].

More fundamentally, no formal hierarchy is provided to distribute the power of decision. One possible reason for this is that the small scale of community size does not request such hierarchy; however, one direct drawback is that people tend to engage in groupthink, resulting in the denser discourse congruence.

These two features compound: a certain minority holds the voice. The decentralization promised by the EIP architecture exists at the level of entry rights; in practice, routine decision authority concentrates around whoever has the time and reputation to remain in the room.

5.2 Role of Open Source and the Limitations↩︎

A natural objection is that the convergence we observe stems not from governance form but from open-source-society (OSS) norms shared by both cases. First, the topical divergence between the two cases (\(\mathrm{JSD}=0.288\) on BERTopic, \(0.216\) on Thematic-LM) cannot be reduced to OSS norms; it is governance-driven, reflecting the constitutive–regulative distinction between proposing a standard (what the protocol is) and shipping an implementation (how it is built). Second, ERC-8004’s discourse-congruence density is nearly twice that of A2A, the opposite of what an OSS-scale-only account would predict. The cautious framing that follows is therefore layered: open-source publication establishes a baseline of high participation inequality and fragmented community structure; governance form then redirects which themes this skewed participation engages and which actors accrue informal authority.

Here are two more possible limitations. First, ERC-8004 contributes 142 records against A2A’s 4,181, so themes of low frequency carry wide confidence intervals. Second, A2A’s TSC meetings, internal Google design reviews, and partner negotiations occur outside the public repository. The structural concentration we observe in A2A may therefore understate deliberation among its core members.

5.3 Who Controls Future Directions?↩︎

The answer is, in both cases, a small group of elites, despite the different constitutions. In ERC-8004, they are those who stay in the room; for A2A, corporation representatives are the gatekeepers. The deeper risk of the DAO model is not that elites emerge—they emerge in any sustained deliberation [28], [30]—but that, absent formal accountability structures, their influence operates through reputational authority that is opaque to outsiders and difficult for newcomers to contest.

Moreover, the governance form determines what problems communities treat as worth deliberating about. ERC-8004’s deliberation is dominated by trust and security, while A2A also spreads its deliberation across engineering and execution issues. The concentration differences of DAO and corporate discourse reflect a value distribution that propagates into deployed systems: communities that deliberate about accountability will architect for accountability; those that deliberate about velocity will architect for velocity [3]. As agentic AI scales from research prototypes to critical infrastructure, these upstream deliberative choices become societal choices.

Three implications follow for those positioned to act on these findings. Standards bodies should pair open entry rights with procedural accountability mechanisms, because open access alone does not redistribute deliberative authority. Protocol architects should treat the deliberative agenda itself as a design artifact: topics absent from the docket become embedded defaults, so a periodic audit of what is not discussed is as essential as auditing what is. Practitioners and policymakers adopting either standard should consult its deliberation record alongside its specification, because the values that shaped the rules are rarely fully recoverable from the rules themselves.

6 Data Expansion and Robustness Verification↩︎

To verify that main-text findings survive annotator choice and case scope, we expanded the DAO corpus from a single ERC-8004 thread to a 34-ERC agent-standardization cluster, re-annotated both cases with multiple independent models, and compared results against the original MiniMax-M2.5 labels.

6.1 Design↩︎

Data Expansion. Beyond ERC-8004, we added 33 contemporaneous (\(\ge\)​2025-08) agent-standardization ERCs from the Ethereum Magicians forum, totaling 34 unique ERC standards: 8001, 8004, 8033, 8041, 8107, 8118, 8122, 8126, 8150, 8160, 8162, 8165, 8166, 8171, 8181, 8183, 8184, 8196, 8203, 8210, 8217, 8220, 8226, 8239, 8240, 8242, 8257, 8259, 8263, 8264, 8273, 8274, 8275, 8294. Among these, ERC-8183 (Agentic Commerce), ERC-8274 (AI Inference Proof Verification), and ERC-8210 (Agent Assurance) are the three most deliberatively active additions. The expanded ERC corpus contains 1,664 annotated records, up from 142 in the main text. The A2A corpus was reused.

Cross-Model Annotation. Three models (DeepSeek-V4-Flash, GLM-4-Plus, and Moonshot-v1-auto) independently annotated every record on all five fields (stakeholder_institution, argument_type, stance, consensus_signal, key_point). Majority voting (2 out of 3) yielded consensus labels for 1,664 ERC and 4,058 A2A records. Together with the original MiniMax-M2.5, four independent annotators provide inter-annotator agreement estimates.

Table 5: Three-model cross-consensus agreement rates (majority voting across 3 rounds within each model, then 2-of-3 across models). “No maj.” = all three models disagree on that field.
ERC (\(N{=}1{,}641\)) A2A (\(N{=}3{,}761\))
2-4 (lr)5-7 Field 3/3 2/3 No maj. 3/3 2/3 No maj.
Argument type 68.0% 30.0% 2.0% 67.5% 30.7% 1.9%
Stance 58.8% 39.1% 2.1% 56.3% 40.3% 3.4%
Consensus signal 59.7% 37.8% 2.5% 61.2% 37.3% 1.5%

Cross-Round Annotation. Each of the three models independently re-annotated every record three times on the three core fields (argument_type, stance, consensus_signal). Majority vote across rounds within each model produced a per-model consensus; majority vote across models then produced the final cross-consensus. Per-field cross-consensus agreement rates are reported in Table 5. When all three models disagree for a given field (no majority; \(<2\%\)\(3.5\%\) of records), the label is assigned by plurality with an effectively arbitrary tie-break and is treated as low-confidence in subsequent analyses.

6.2 Robustness Verification↩︎

Cross-model reliability. Four-model Fleiss’ \(\kappa\) was computed (Table 6). All three fields reach Moderate agreement in both cases, with argument type—the core field driving the \(\chi^2\), BERTopic, and Thematic-LM analyses—at \(\kappa=0.545\) (ERC) and \(\kappa=0.529\) (A2A). The strongest pairwise agreement occurs between GLM-4-Plus and Moonshot-v1-auto (argument type \(\kappa=0.671\) ERC, \(0.555\) A2A), suggesting these two models share similar interpretive biases relative to DeepSeek-V4-Flash and MiniMax-M2.5.

Table 6: Four-model Fleiss’ \(\kappa\) (MiniMax-M2.5, DeepSeek-V4-Flash, GLM-4-Plus, Moonshot-v1-auto). ERC \(N=144\), A2A \(N=3{,}760\).
Field ERC A2A
Argument type 0.545 0.529
Stance 0.579 0.530
Consensus signal 0.485 0.483

Cross-round self-consistency. Within-model test-retest reliability (Table 7) reveals clear model-level differences. GLM-4-Plus and Moonshot-v1-auto achieve almost perfect self-consistency (Fleiss’ \(\kappa>0.81\) on all three fields in both cases), while DeepSeek-V4-Flash yields Moderate-to-Substantial values (\(\kappa=0.49\)\(0.63\)). The gap confirms that annotator model choice dominates stochastic variation: two models are near-deterministic in their labeling behavior, while one exhibits higher intrinsic variance [51]. Notably, Moonshot-v1-auto, a generic auto-routing model with no special reasoning configuration, outperforms both alternatives in self-consistency, challenging the intuition that reasoning models are inherently more reliable annotators.

Table 7: Within-model cross-round Fleiss’ \(\kappa\) (3 rounds, majority vote). ERC \(N{=}1{,}664\), A2A \(N{\approx}3{,}845\).
ERC A2A
2-4 (lr)5-7 Model AT St CS AT St CS
Moonshot-v1-auto 0.978 0.961 0.945 0.976 0.963 0.938
GLM-4-Plus 0.925 0.904 0.861 0.910 0.874 0.815
DeepSeek-V4-Flash 0.634 0.554 0.507 0.565 0.543 0.490
AT = argument type; St = stance; CS = consensus signal.

6.3 Substantive Replication↩︎

We re-ran the annotation pipeline with three additional models (Moonshot-v1-auto, GLM-4-Plus, DeepSeek-V4-Flash; ICR \(\kappa\) \(\approx\) 0.7–0.9 for the top two, § A.1) and replicated the Thematic-LM discourse analysis [11] using Moonshot-v1-auto as the LLM backbone, yielding a 12-theme codebook with 96.6% coverage. Table 8 reports the key combined metrics.

Table 8: Replication metrics: multi-model consensus annotation + Moonshot Thematic-LM codebook.
Metric ERC (cluster) A2A (re-anno.)
DNA actors 194 713
DNA congruence density 0.403 0.252
DNA polarization index 0.056 0.740
Giant component ratio 0.917 0.285
Mean actor entropy \(H\) 0.734 0.511
Dominant theme (% record) Compliance (31.1%) Documentation (31.4%)
Thematic JSD 0.092

Three of four main-text findings replicate: (1) Technical argument types dominate regardless of annotator, with A2A receiving roughly twice the Process share; (2) both networks remain steeply unequal (betweenness Gini \(\approx\)​0.8); (3) ERC discourse congruence density exceeds A2A (0.403 vs.). Thematic content also replicates: ERC concentrates on constitutive themes (compliance, standards, verification), A2A on executive themes (documentation, coordination, community). However, network connectivity reverses: the expanded ERC network coalesces into one dominant component (GCR 0.328\(\rightarrow\)​0.917), while the re-annotated A2A network fragments further (GCR 0.534\(\rightarrow\)​0.285). The cross-case thematic JSD of 0.092 is lower than the main-text BERTopic value (0.288), partly reflecting Moonshot’s coarser 12-theme granularity. We interpret the connectivity reversal as an observability differential: permissionless DAO governance externalizes deliberation into a genuinely interconnected public record, whereas corporate governance internalizes decisive coordination channels (TSC calls, private forums), leaving the public GitHub trace persistently fragmented. Participation concentration is comparable across both governance forms, but observability is not.

7 Comparative Case Study Design↩︎

We adopt a comparative case study design to examine how governance structure shapes participation dynamics in AI protocol standardization [52]. Interoperability is an inherited characteristic of decentralized applications (dApps) in DeFi [1], and ERC-8004 as a protocol specifically designed for AI agents extends this feature. Google A2A, on the corporation’s side, realized similar functions in different contexts. Moreover, they originated within the same year (2025). Nevertheless, they differ sharply in governance architecture. ERC-8004 belongs to Ethereum, a permissionless and community-driven open-source DAO, and Google A2A is now governed by TSC. Therefore, the two protocols constitute a theoretically matched pair. Holding the technical domain constant while varying the governance form enables structured comparison of participation patterns, discourse composition, and network analysis.

8 Data Provenance↩︎

All raw data files are versioned with SHA-256 checksums stored in data/raw/CHECKSUMS.json in the project repository. Table 9 lists the file names, record counts, and collection dates for the primary raw data files used in this study.

Table 9: Raw Data Files and Record Counts
File Records Collected
forum_posts.json 113 2026-03
github_comments_filtered.json 36 2026-03
a2a_issues.json 3,104 2026-03
a2a_prs.json 1,955 2026-03
a2a_discussions.json 822 2026-03
Total (raw) 6,030
Total (retained) 4,323

9 ERC-8004 Lifecycle Phase Boundaries↩︎

Topic analysis within ERC-8004 was conducted across three consecutive two-month phases spanning the full proposal lifecycle (2025-08-13 to 2026-01-29):

  • Phase 1 (Aug 13 – Oct 13, 2025): Initial submission and early community review. \(n = 100\) records.

  • Phase 2 (Oct 13 – Dec 13, 2025): Sustained review period prior to Last Call designation. \(n = 11\) records.

  • Phase 3 (Dec 13, 2025 – Feb 13, 2026): Last Call through Final ratification (mainnet deployment 2026-01-29). \(n = 17\) records.

10 Decision Architectures↩︎

Figure 9: EIP Lifecycle
Figure 10: A2A Decision Process (TSC Governance)

There are 3 types of EIPs: standard track EIP, Meta EIP and Informational EIP. The standards track EIP can be broken down to 4 categories: core, networking, interface and ERC [5]. All EIPs, including ERC-8004 in our case, undergo 5 stages: idea, draft, review, last call and final. EIPs updated continuously are assigned to the special stage of “living”. All EIPs are decided by rough consensus: Core EIPs by Ethereum core developers at the meeting called AllCoreDevs, and others by common developers and users.

This reflects a structural feature of ERCs as application-layer specifications: unlike Core EIPs, which modify the Ethereum protocol itself and requires hard-fork activation when deployment, ERCs define smart-contract interface standards that any party may implement independently, regardless of the proposal’s formal lifecycle status [5].

Google A2A launched under google/A2A in April 2025 and was donated to the Linux Foundation in June 2025 [8], migrating to the vendor-neutral a2aproject/A2A organization. Governance authority rests with an eight-seat Technical Steering Committee (TSC), comprising representatives each from Google, Microsoft, Cisco, AWS, Salesforce, SAP, IBM, and ServiceNow. Independent contributors cannot join the TSC during an 18-month startup phase [7]. Routine pull requests are merged upon maintainer approval without a formal vote [7]; contested specification changes escalate to a GitVote [50], a GitHub-native ballot in which only TSC members cast binding votes at a 51% threshold.

For pseudocodes, see Algorithm 9 and 10.

11 Institution Label Provenance Cascade↩︎

Institution labels were assigned using the following three-tier cascade, applied in priority order:

  1. Manual investigation (high confidence): A systematic review of GitHub profiles, LinkedIn pages, personal websites, and EIP commit metadata was conducted for the top 109 contributors across both cases (all 71 ERC-8004 participants plus the top 38 A2A contributors by post count). This review yielded 40 institutional upgrades, replacing LLM-inferred labels with verified affiliations. The original LLM-inferred label is preserved in a separate institution_lm field for all records to enable sensitivity checks.

  2. LLM inference (low confidence): For the remaining 517 authors, institution was inferred by MiniMax-M2.5 from contextual signals in the record text (e.g., repository ownership, self-identification, email patterns where visible).

12 Actor Filtering Stages↩︎

The three network analyses operate on successively filtered actor sets (Table 10). Starting from the annotated records (ERC: 142 records / 71 contributors; A2A: 4,181 / 778), the co-participation network drops four ERC and seven A2A bot/short-text residuals, yielding \(N=67\) and \(N=771\) respectively. The DNA and socio-semantic analyses further (a) inner-join against the Thematic-LM coded_records.json (12 records lose a theme assignment) and (b) exclude records whose stance is Off-topic or Unclassified, dropping actors whose only contributions fall into those categories. This produces the final stance-bearing actor sets of \(N=66\) (ERC-8004) and \(N=710\) (Google A2A). All SNA statistics use \(N=67\) / \(N=771\); all DNA and socio-semantic statistics use \(N=66\) / \(N=710\).

Table 10: Actor and record counts through the three filtering stages.
ERC-8004 Google A2A
Records Actors Records Actors
Annotated (retained) 142 71 4,181 778
SNA (bot & short-text filter) 130 67 4,230 771
DNA / Socio-semantic 126 66 3,759 710

13 Related Work Table↩︎

Table ¿tbl:tab:litreview? positions sixteen representative works into three classes by paper type, each evaluated by criteria appropriate to that class. Panel (a) lists foundational works that originated the methods or concepts we inherit; panel (b) lists perspective papers that advance theoretical claims about DAO and corporate governance without conducting their own empirical tests; panel (c) lists empirical and methodological studies are the key methodological works evaluated additionally by openness. This study extends panel (c) with all data and codes open.

Paper Year Core Contribution Inheritance for this Study
Russell [53] ‘Rough Consensus’ & the Internet–OSI Standards War 2006 “Rough consensus” as a standardization mode Frames ERC-8004’s decision rule
Roth & Cointet [14] Social and Semantic Coevolution in Knowledge Networks 2010 Coevolution of social and semantic ties Root of our network–discourse layer
Leifeld [13] Discourse Network Analysis of Policy Change 2013 Discourse Network Analysis (DNA) DNA layer used in §Methods
Beck et al. [16] Governance in the Blockchain Economy 2018 IS framework: rights / accountability / incentives Defines the decentralization axis
Grootendorst [10] BERTopic: Neural Topic Modeling 2022 Neural topic modeling with c-TF-IDF Used in our topic-discovery pipeline
Paper Year Central Claim Domain Backing
Murray et al. [23] Contracting in the Smart Era 2021 Smart contracts and DAOs reshape agency cost DAO\(\times\)Corp Theory
Lumineau et al. [24] Blockchain Governance: A New Way of Organizing 2021 Blockchain as a new mode of organizing DAO\(\times\)Corp Theory
Hui & Tucker [27] Decentralization, Blockchain, AI 2025 Decentralization is the right frame for AI governance AI-Proto Conceptual
Reineke et al. [20] Decentralization: A Revolution or a Mirage? 2025 Decentralization is contested between revolution and mirage DAO\(\times\)Corp Review
Sunyaev et al. [21] From Ideology to Design: Purposeful Decentralization 2026 Decentralization should be design-driven, not ideological DAO\(\times\)Corp Conceptual
Positioning of this study relative to prior work, organized by paper type. Each panel applies an evaluation rubric appropriate to its class; the synthesis line at the bottom describes how this study bridges all three.
Paper Year Method Setting Data Code Comparison
Stine & Agarwal [34] Comparative Discourse Analysis via Topic Models 2020 Comparative topic models Online political discourse \(\times\) \(\times\) \(✔\)
Qiao et al. [11] Thematic-LM: LLM-Based Thematic Analysis 2025 LLM-aided thematic analysis Generic corpora \(\times\) \(\times\) \(\times\)
Ao et al. [12] Is DeFi Actually Decentralized? 2023 Social network analysis Aave (DAO) \(✔\) \(✔\) \(\times\)
Leifeld [13] Discourse Network Analysis of German Pension Politics 2013 Discourse network analysis German pension policy \(\times\) \(\times\) \(\times\)
Roth & Cointet [14] Social and Semantic Coevolution in Knowledge Networks 2010 Socio-semantic network analysis Knowledge networks \(\times\) \(\times\) \(\times\)

(a) Foundational / Seminal Works

(b) Perspective / Opinion Papers

References↩︎

[1]
C. R. Harvey and D. Rabetti, “International business and decentralized finance,” Journal of International Business Studies, vol. 55, no. 4, pp. 840–863, 2024, doi: 10.1057/s41267-024-00705-7.
[2]
J. Evans, B. Bratton, and B. Agüera y Arcas, “Agentic AI and the next intelligence explosion,” Science, vol. 391, no. 6791, 2026, doi: 10.1126/science.aeg1895.
[3]
World Economic Forum and Accenture, “Organizational transformation in the age of AI: How organizations maximize AI’s potential.” World Economic Forum, 2026, [Online]. Available: https://www.weforum.org/publications/organizational-transformation-in-the-age-of-ai-how-organizations-maximize-ais-potential/.
[4]
Ethereum Magicians, Accessed: Mar. 10, 2026ERC-8004: Trustless agents.” Ethereum Magicians Forum, topic 25098, Aug. 2025, [Online]. Available: https://ethereum-magicians.org/t/erc-8004-trustless-agents/25098.
[5]
Ethereum Foundation, Accessed: Mar. 28, 2026EIP-1: EIP purpose and guidelines.” Ethereum Improvement Proposals, Oct. 2015, [Online]. Available: https://eips.ethereum.org/EIPS/eip-1.
[6]
A2A Authors, Accessed: Mar. 28, 2026Agent2Agent (A2A) Protocol.” GitHub, Feb. 2025, [Online]. Available: https://github.com/a2aproject/A2A.
[7]
A2A Authors, Accessed: Mar. 28, 2026GOVERNANCE.md.” GitHub, 2025, [Online]. Available: https://github.com/a2aproject/A2A/blob/main/GOVERNANCE.md.
[8]
Google Cloud, Accessed: Mar. 28, 2026“Google Cloud donates A2A to Linux Foundation.” Google Developers Blog, Jun. 2025, [Online]. Available: https://developers.googleblog.com/en/google-cloud-donates-a2a-to-linux-foundation/.
[9]
N. A. Carlson and V. Burbano, “The use of LLMs to annotate data in management research: Foundational guidelines and warnings,” Strategic Management Journal, vol. 47, no. 3, pp. 699–725, 2026, doi: https://doi.org/10.1002/smj.70023.
[10]
M. Grootendorst, BERTopic: Neural topic modeling with a class-based TF-IDF procedure.” arXiv:2203.05794, Mar. 2022.
[11]
T. Qiao, C. Walker, C. Cunningham, and Y. S. Koh, “Thematic-LM: A LLM-based multi-agent system for large-scale thematic analysis,” in Proceedings of the ACM on web conference 2025, 2025, pp. 649–658, doi: 10.1145/3696410.3714595.
[12]
Z. Ao, L. W. Cong, Gergely. Horvath, and L. Zhang, “Is decentralized finance actually decentralized? A social network analysis of the Aave protocol on the Ethereum blockchain.” arXiv:2206.08401 [econ.GN], Dec. 2023.
[13]
P. Leifeld, “Reconceptualizing major policy change in the advocacy coalition framework: A discourse network analysis of german pension politics,” Policy Studies Journal, vol. 41, no. 1, pp. 169–198, 2013, doi: https://doi.org/10.1111/psj.12007.
[14]
C. Roth and J.-P. Cointet, Dynamics of Social Networks“Social and semantic coevolution in knowledge networks,” Social Networks, vol. 32, no. 1, pp. 16–29, 2010, doi: https://doi.org/10.1016/j.socnet.2009.04.005.
[15]
C. R. Harvey, A. Ramachandran, and J. Santoro, DeFi and the future of finance. Hoboken, NJ: Wiley, 2021.
[16]
R. Beck, C. Mueller-Bloch, and J. King, “Governance in the blockchain economy: A framework and research agenda,” Journal of the Association for Information Systems, vol. 19, pp. 1020–1034, Oct. 2018, doi: 10.17705/1jais.00518.
[17]
R. Ziolkowski, G. Miscione, and G. Schwabe, “Decision problems in blockchain governance: Old wine in new bottles or walking in someone else’s shoes?” Journal of Management Information Systems, vol. 37, no. 2, pp. 316–348, 2020, doi: 10.1080/07421222.2020.1759974.
[18]
A. Kiayias and P. Lazos, “SoK: Blockchain governance,” in Proceedings of the 4th ACM conference on advances in financial technologies, 2023, pp. 61–73, doi: 10.1145/3558535.3559794.
[19]
E. W. Ellinger, R. W. Gregory, T. Mini, T. Widjaja, and O. Henfridsson, “Skin in the game: The transformational potential of decentralized autonomous organizations,” MIS quarterly, vol. 48, no. 1, pp. 245–272, 2024.
[20]
P. Reineke, R. Katila, and K. M. Eisenhardt, “Decentralization in organizations: A revolution or a mirage?” Academy of Management Annals, vol. 19, no. 1, pp. 298–342, 2025, doi: 10.5465/annals.2022.0206.
[21]
A. Sunyaev, M. Avital, and M. C. Lacity, “From ideology to design: Toward purposeful decentralization of information systems governance,” Journal of the Association for Information Systems, vol. 41, no. 1, 2026, doi: 10.1177/02683962261430919.
[22]
M. Motea and P. Oba, “Who governs the ledger: Rethinking blockchain governance through democratic innovation,” Humanities and Social Sciences Communications, vol. 13, no. 1, p. 770, 2026, doi: 10.1057/s41599-026-06980-z.
[23]
A. Murray, S. Kuban, M. Josefy, and J. Anderson, “Contracting in the smart era: The implications of blockchain and decentralized autonomous organizations for contracting and corporate governance,” Academy of Management Perspectives, vol. 35, no. 4, pp. 622–641, 2021, doi: 10.5465/amp.2018.0066.
[24]
F. Lumineau, W. Wang, and O. Schilke, “Blockchain governance—a new way of organizing collaborations?” Organization Science, vol. 32, no. 2, pp. 500–521, Mar. 2021, doi: 10.1287/orsc.2020.1379.
[25]
H. A. Rahman, A. Karunakaran, and L. D. Cameron, “Taming platform power: Taking accountability into account in the management of platforms,” Academy of Management Annals, vol. 18, no. 1, pp. 251–294, 2024, doi: 10.5465/annals.2022.0090.
[26]
R. A. Hunt, D. M. Townsend, R. Nugent, J. J. Simpson, M. Stallkamp, and E. Bozdag, “Digital battlegrounds: The power dynamics and governance of contemporary platforms,” Academy of Management Annals, vol. 19, no. 1, 2024, doi: 10.5465/annals.2022.0188.
[27]
X. Hui and C. Tucker, “Decentralization, blockchain, artificial intelligence (AI): Challenges and opportunities,” Journal of Product Innovation Management, vol. 42, no. 5, pp. 947–957, 2025, doi: https://doi.org/10.1111/jpim.12800.
[28]
A. Mockus, R. T. Fielding, and J. D. Herbsleb, “Two case studies of open source software development: Apache and mozilla,” ACM Trans. Softw. Eng. Methodol., vol. 11, no. 3, pp. 309–346, Jul. 2002, doi: 10.1145/567793.567795.
[29]
J. Im, A. X. Zhang, C. J. Schilling, and D. Karger, “Deliberation and resolution on wikipedia: A case study of requests for comments,” vol. 2, no. CSCW, Nov. 2018, doi: 10.1145/3274343.
[30]
M. Germonprez, G. J. P. Link, K. Lumbard, and S. Goggins, “Eight observations and 24 research questions about open source projects: Illuminating new realities,” Proc. ACM Hum.-Comput. Interact., vol. 2, no. CSCW, Nov. 2018, doi: 10.1145/3274326.
[31]
R. Li, P. Pandurangan, H. Frluckaj, and L. Dabbish, “Code of conduct conversations in open source software projects on github,” Proc. ACM Hum.-Comput. Interact., vol. 5, no. CSCW1, Apr. 2021, doi: 10.1145/3449093.
[32]
M. Kulakowski and F. Frasincar, “Sentiment classification of cryptocurrency-related social media posts,” vol. 38, no. 4, pp. 5–9, Jul. 2023, doi: 10.1109/MIS.2023.3283170.
[33]
Y. Quan, X. Wu, W. Deng, and L. Zhang, “Decoding social sentiment in DAO: A comparative analysis of blockchain governance communities,” in 2024 IEEE 24th international conference on software quality, reliability, and security companion (QRS-c), 2024, pp. 216–224, doi: 10.1109/QRS-C63300.2024.00037.
[34]
Z. K. Stine and N. Agarwal, “Comparative discourse analysis using topic models: Contrasting perspectives on china from reddit,” in International conference on social media and society, 2020, pp. 73–84, doi: 10.1145/3400806.3400816.
[35]
Q. Wang, G. Yu, Y. Sai, C. Sun, L. D. Nguyen, and S. Chen, “Understanding DAOs: An empirical study on governance dynamics,” IEEE Transactions on Computational Social Systems, vol. 12, no. 5, pp. 2814–2832, 2025, doi: 10.1109/TCSS.2025.3539889.
[36]
M. Jungnickel, F. Özdemir Sönmez, C. Mulligan, and W. J. Knottenbelt, Just Accepted“DAO governance: Voting power, participation, and controversy - a review and an empirical analysis,” Distrib. Ledger Technol., Nov. 2025, doi: 10.1145/3777416.
[37]
K. Chen, Y. Fan, Y. Fang, and X. (Robert). Luo, “Beyond money: Incentive effects of tokenized ownership on user contribution in DAOs,” Journal of Operations Management, vol. 71, no. 7, pp. 988–1016, 2025, doi: https://doi.org/10.1002/joom.1351.
[38]
MiniMax, Accessed: Apr. 11, 2026MiniMax M2.5: Built for real-world productivity.” MiniMax, Feb. 2026, [Online]. Available: https://www.minimax.io/news/minimax-m25.
[39]
C. Ma et al., “AgentBoard: An analytical evaluation board of multi-turn LLM agents,” in Proceedings of the 38th international conference on neural information processing systems, 2024.
[40]
ERC-8004 Authors, Accessed: Mar. 10, 2026“Erc-8004.md.” GitHub, 2026, [Online]. Available: https://github.com/ethereum/ERCs/blob/c8cd50348b111d13f62e1d71d2b4c8101bd3514f/ERCS/erc-8004.md.
[41]
J. Cohen, Statistical power analysis for the behavioral sciences, 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates, 1988.
[42]
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese BERT-networks,” Jan. 2019, pp. 3973–3983, doi: 10.18653/v1/D19-1410.
[43]
J. Lin, “Divergence measures based on the shannon entropy,” IEEE Transactions on Information Theory, vol. 37, no. 1, pp. 145–151, 1991, doi: 10.1109/18.61115.
[44]
V. D. Blondel, J.-L. Guillaume, R. Lambiotte, and E. Lefebvre, “Fast unfolding of communities in large networks,” Journal of Statistical Mechanics: Theory and Experiment, vol. 2008, no. 10, p. P10008, 2008, doi: 10.1088/1742-5468/2008/10/P10008.
[45]
M. E. J. Newman, “Modularity and community structure in networks,” Proceedings of the National Academy of Sciences, vol. 103, no. 23, pp. 8577–8582, 2006, doi: 10.1073/pnas.0601602103.
[46]
S. P. Borgatti and M. G. Everett, “Models of core/periphery structures,” Social Networks, vol. 21, no. 4, pp. 375–395, 2000, doi: https://doi.org/10.1016/S0378-8733(99)00019-2.
[47]
S. Kojaku and N. Masuda, “A generalised significance test for individual communities in networks,” Scientific Reports, vol. 8, 2018, doi: 10.1038/s41598-018-25560-z.
[48]
V. Latora and M. Marchiori, “Efficient behavior of small-world networks,” Physical Review Letters, vol. 87, no. 19, p. 198701, 2001, doi: 10.1103/PhysRevLett.87.198701.
[49]
Etherscan, Deployed Jan. 29, 2026ERC-8004 identity registry contract (0x8004A169).” Ethereum Mainnet, Jan. 2026, [Online]. Available: https://etherscan.io/address/0x8004A169FB4a3325136EB29fA0ceB6D2e539a432.
[50]
a2aproject,.gitvote.yml.” GitHub, 2025, [Online]. Available: https://github.com/a2aproject/A2A/blob/main/.gitvote.yml.
[51]
J. Ji et al., “Leveraging LLM-based agents for social science research: Insights from citation network simulations,” Humanities and Social Sciences Communications, vol. 13, p. 127, 2026, doi: 10.1057/s41599-025-06193-w.
[52]
R. K. Yin, Case study research and applications: Design and methods, 6th ed. Thousand Oaks, CA: SAGE, 2018.
[53]
A. L. Russell, Rough Consensus and Running Code and the Internet-OSI Standards War,” IEEE Annals of the History of Computing, vol. 28, no. 3, pp. 48–61, Jul. 2006, doi: 10.1109/MAHC.2006.42.

  1. https://github.com/kl41r3/erc8004-a2a-case-study↩︎