Supplementary Information
LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation
June 25, 2026
This section provides more detail on the descriptive aspects of the datasets collected from diverse online platforms. Namely, we collected data from Scopus, LinkedIn, CORDIS, YouTube, Reddit, and Bluesky. These platforms served as perspective data pools for collecting information related to hydrogen with more focus on green and renewable hydrogen and feedstocks, where available. Thus, the main perspectives considered are the Academic (Scopus), Business (LinkedIn), EU funding (CORDIS), and Public (YouTube, Reddit, and Bluesky).
Scholarly trends from Scopus confirm a definitive shift in terminology and focus. First, while “renewable feedstock” dominated the early relative literature share, “green hydrogen” has achieved conceptual dominance since 2020, with absolute publication volumes following an exponential growth curve that peaked in 2023–2024. Figure 1 illustrates the longitudinal growth in the absolute frequency of mentions for “green hydrogen,” “renewable hydrogen,” and “renewable feedstock(s)” within the Scopus database from 2000 to 2026. The data, normalised to a 0–100 scale based on peak mention volume, reveals a period of relative dormancy followed by a synchronised exponential increase beginning around 2015 (Figure 1a). This trajectory signifies a massive expansion in the total body of literature dedicated to these fields, with “green hydrogen” and “renewable hydrogen” reaching their highest absolute research output between 2022 and 2024. Further, Figure 1b depicts the relative normalised share of literature as a descriptor of each term’s contribution in the total academic discourse.
To facilitate a synchronised comparison across heterogeneous datasets, two distinct normalisation layers were applied. In these models, the variable \(x\) represents the specific entity under analysis (a keyword in the Scopus dataset or a platform in the YouTube, Reddit, and Bluesky datasets) at a given time \(t\). Figure 1a and Figure 5a utilise a peak-normalisation index to visualise the internal growth cycle of each entity. Formula (1 ) scales the annual frequency against the historical peak of that specific entity.
\[\mathrm{norm}_{x,t} = \frac{f(x,t)}{\max(f_x)} \times 100 \label{eq:norm}\tag{1}\]


Figure 1: (a) Normalised annual mentions of “green hydrogen,” “renewable hydrogen,” and “renewable feedstock(s)” in Scopus-indexed literature (2000–2026). (b) Relative share of literature index for “green hydrogen,” “renewable hydrogen,” and “renewable feedstock(s)” (2000–2026). Values are indexed to the historical maximum share (\(\max(S)=100\)) for visual clarity across datasets; consequently, annual totals exceed 100..
Figure 1b and Figure 5b illustrate the competitive dominance of an entity within the total discourse of a specific year. We first calculate the annual share \(S\) for both literature and public discourse (2 ). This identifies the slice of the pie held by a platform or keyword relative to all other observed entities at that exact moment in time. For maintaining visual consistency and a 0–100 comparison across different datasets, the annual shares are further scaled into a normalised share index \(S_{\mathrm{norm}}\) (3 ).
\[S_{\mathrm{norm}}(x,t) = \frac{S(x,t)}{\max(S)} \times 100 \label{eq:snorm}\tag{3}\]
The corporate landscape of the hydrogen sector on LinkedIn is defined by a high density of small-to-medium enterprises concentrated in a few key geographic hubs. Figure 2a displays the distribution of company sizes, revealing a heavily right-skewed population where the vast majority of firms maintain fewer than 1,000 employees, despite a long tail of large-scale industrial entities. This organizational diversity is geographically stratified, as shown in Figure 2b; the United States, India, and Germany lead the top 15 headquarters countries, indicating that while the hydrogen economy is global, its digital corporate presence is currently anchored in these primary innovation hubs. The relationship between organizational scale and digital visibility is characterised by a positive, yet heterogeneous, correlation. Figure 2c illustrates that while increasing headcount generally drives higher follower counts, significant dispersion exists among micro-entities. The distinct vertical stripe-shaped patterns observed at the lower end of the \(x\)-axis represent the high frequency of organisations with identical, small integer headcounts (e.g., 1–10 associated members). This visualisation highlights that in the early-stage hydrogen sector, brand positioning and project novelty allow specialised technology firms to achieve digital influence that is partially decoupled from their physical staff size.



Figure 2: Global distribution and digital reach of hydrogen-related companies on LinkedIn. (a) Distribution of company size: the frequency of firms by staff count. (b) Top 15 headquarter countries: the geographic concentration of hydrogen-related corporate profiles. (c) Company size vs.LinkedIn reach: associated staff and follower count correlation..
| Industry | Country |
|---|---|
| Manufacturing | European Union |
| United States | |
| China, India, Iraq, Iran, Russia | |
| Saudi Arabia, Japan, Pakistan | |
| Turkey | |
| Latin America | |
| All | Africa |
| Professional services | Europe |
| United Kingdom | |
| United States, India | |
| Aviation | All |
This academic trend is mirrored in the EU funding landscape, where CORDIS data shows a strategic pivot from foundational research and innovation actions (RIA) under Horizon 2020 toward market-readiness challenges and specialised talent development under Horizon Europe. The evolution of hydrogen-related research within the European funding landscape is characterised by steady growth followed by a recent transition in reporting cycles. Figure 3a illustrates the number of CORDIS projects by start date, showing a sustained increase from 2014, peaking in 2023 with over 250 projects. The sharp decline observed in 2024–2026 is likely an artifact of current data reporting lags and the time-gap between call closures and project formalisation, rather than a reduction in funding interest. The distribution of funding instruments (Figure 3b) highlights the prevalence of research and innovation actions and the Marie Skłodowska-Curie Actions, particularly the postdoctoral fellowships as explained in Table 2. This indicates a dual focus on large-scale collaborative breakthroughs and individual talent mobility. Furthermore, Figure 3c ranks the top specific topics by project count, dominated by recent Horizon Europe calls (e.g., MSCA-2024-PF and EIC Accelerator Challenges). These distributions confirm that hydrogen research has transitioned from foundational science under Horizon 2020 to market-ready innovation and high-level skill development under the Horizon Europe framework.



Figure 3: EU-funded hydrogen research landscape via CORDIS data. (a) Annual projects by starting year (2014–2026). (b) Top funding schemes by project count. (c) Most frequent project topics/calls..
| EU Funding Instrument | Official Call Code | Interpretation of Code Components |
|---|---|---|
| MSCA – Marie Skłodowska-Curie Actions | HORIZON-MSCA-2024-PF-01-01 |
|
| MSCA-IF-2019 |
|
|
| EIC – European Innovation Council | HORIZON-EIC-2021-ACCELERATORCHALLENGES-01-02 |
|
| ERC – European Research Council | ERC-2016-STG | ERC Starting Grant, 2016 call |
| ERC-2016-ADG | ERC Advanced Grant |
Further preliminary views are reflected in the social media analysis, which captured over 100,000 public interactions. While Reddit provides long-term technical discourse, the recent, intense peaks on YouTube and Bluesky suggest that public interest is increasingly reactive to real-time technological demonstrations and policy shifts. Collectively, these insights define a dataset that is both high-volume and semantically diverse, providing a rich, multi-modal foundation for the subsequent RAG-based analysis.
The public visibility of hydrogen-related topics across social media platforms reveals a more volatile and event-driven interest profile compared to the steady growth of academic literature. Figure 4a tracks the absolute monthly post volume, totaling 106,863 analyzed entries across Reddit, Bluesky, and YouTube for the videos, posts, and threads discussing the “green hydrogen”, “renewable hydrogen”, and “renewable feedstocks” terms. YouTube exhibits the highest single-event peaks, notably a massive surge in early 2023. In contrast, Reddit shows the longest historical engagement, with consistent but fluctuating activity dating back to 2008. Because the platforms operate at different scales of total volume, for more explanatory insights, we employ a min-max normalisation, as defined by equation (1 ) of the supplementary material, to visualise the relative intensity of interest for each platform independently (Figure 4b). This normalisation reveals that while Reddit interest has been broad and sustained, the Bluesky discourse is a very recent phenomenon, reaching its peak intensity only in late 2024 and 2025, suggesting that while established platforms host long-term technical and community debates, newer decentralised platforms are rapidly becoming hubs for real-time discussion on the hydrogen transition. On the contrary, equations (2 ) and (3 ) reveal the dominance of each platform on the total discourse over the examined timeframe, through a share-based distribution.


Figure 4: Public discourse trends for hydrogen-related topics across social media platforms (2008–2026). (a) Absolute volume of posts per month for Bluesky, Reddit, and YouTube. (b) Normalised activity index (0–1), scaling each platform by its own historical minimum and maximum to show relative intensity of interest over time..
In this section, we introduce the case study, where we analyse the production and distribution of apples in Italy to the most relevant European country importers. The four phases of the LCA for this case study are presented next.
We consider a cradle-to-gate assessment with a cut-off attributional approach. The chosen functional unit is 1 kg of apples. We use the ecoinvent database v3.12 to model the background system [1].
The life cycle inventory (LCI) for the foreground system was built using the ecoinvent database v3.12. We use the activity “apple production” (IT) to model the technosphere and biosphere flows associated with the production of 1 kg of apples in Italy.
For the supply chain, we consider the activity “transport, freight, lorry, \(>\)32 metric ton, diesel, EURO 3” (RER). The mass of imports was estimated from [2]. The distance was estimated from the centroids of Italy and the importer countries using the Python library geopandas [3]. The data used for the calculations is shown in Table 3, and the final LCI in Table 4, which considers the mean ton\(\cdot\)km/kg of apples imported.
| Country | Imports [tonnes] | Distance [km] | Mass per distance [ton\(\cdot\)km]\(\times 10^{-7}\) |
|---|---|---|---|
| Germany | 280,000 | 943 | 26.4 |
| France | 40,000 | 1228 | 4.91 |
| Austria | 40,000 | 561 | 2.25 |
| Spain | 40,000 | 1336 | 5.34 |
| United Kingdom | 40,000 | 1657 | 6.63 |
| Netherlands | 40,000 | 1171 | 4.68 |
| Sweden | 33,333 | 2249 | 7.50 |
| Norway | 33,333 | 2942 | 9.81 |
| Denmark | 33,333 | 1489 | 4.96 |
| Functional unit: 1 kg of apples | |||
|---|---|---|---|
| Inputs | Location | Amount | Units |
| apple production | IT | 1.00 | kg |
| transport, freight, lorry, \(>\)32 metric ton, diesel, EURO 3 | RER | 1.25 | ton\(\cdot\)km |
The life cycle impact assessment (LCIA) was performed using the IPCC 2021 method in Brightway2 v2.4.2 [4]. The results were examined to determine the impact of diesel in the functional unit. Evaluating the supply vectors (Figure 5) uncovers that transport contributes about 67% of the impact, while 80% of transport itself comes from diesel production and usage (Figure 6). Regarding the production of apples itself, Figure 7 reveals that its main contributors are the fertiliser production and terrain preparation, while Figure 8 confirms that only 19% of the total climate change impact is derived from diesel. Finally, Figure 9 highlights the effect on the overall climate change impact of reducing diesel consumption by 50%.
| ****ID**** | Response |
|---|---|
| ****ID**** | Response |
| 1 | RELEVANCE: HIGH COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: STRONG
|
| 2 | RELEVANCE: HIGH COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: STRONG |
CONFIDENCE: MEDIUM-HIGH (varies by claim)
|
|
| 3 | RELEVANCE: HIGH COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: STRONG
|
| 4 | RELEVANCE: MODERATE COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: SPARSE
|
| 5 | RELEVANCE: HIGH COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: STRONG
|
| 6 | RELEVANCE: HIGH COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: STRONG
|
| 7 | RELEVANCE: HIGH COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: STRONG
|
| 8 | RELEVANCE: HIGH COVERAGE: MULTI-VIEW EVIDENCE_DENSITY: STRONG CONFIDENCE RULE: MEDIUM
|
PERSONAS: Dict[str, str] = {
"public": (
"ROLE: You are an LCA interpretation expert and strategic sustainability decision advisor. "
"PERSPECTIVE MODE: PUBLIC DISCOURSE."
"Prioritize: narratives, skepticism, perceived costs, safety concerns, greenwashing claims, "
"social acceptance, polarized viewpoints. Treat anecdotes as LOW confidence unless multiple "
"independent items align."
"Explicitly note when claims reflect perception rather than technical fact."
"The user provides a persistent SCENARIO ANCHOR (LCA scenario interpretation outputs) and then "
"asks focused MICRO-QUERIES. Use the SCENARIO ANCHOR as the decision context and answer only "
"the MICRO-QUERY."
"EVIDENCE POLICY:"
"- Treat retrieved context items as perspective-specific evidence. They are partial and may not "
" cover the full domain."
"- For every major claim, cite supporting context item IDs in square brackets (e.g., [1], [2])."
"- Do not invent specific numbers, costs, performance values, named projects, companies, "
" policies, or dates unless present in retrieved context."
"- If the retrieved context does not support the requested point, write: "
" INSUFFICIENT EVIDENCE IN RETRIEVED CONTEXT. Then optionally add: "
" GENERAL KNOWLEDGE (UNGROUNDED): <1-2 short sentences>. "
" Keep ungrounded content clearly separated."
"MICRO-QUERY DISCIPLINE:"
"- Answer only what the user asked (do not expand to full roadmaps unless requested)."
"- If the user asks for N items (e.g., 3-6), comply."
"- Keep all statements tied to the SCENARIO ANCHOR."
"RETRIEVAL QUALITY FLAGS (include at top of every answer):"
"RELEVANCE: HIGH / MODERATE / LOW (1 line)."
"COVERAGE: MULTI-VIEW / SINGLE-VIEW / UNCLEAR (1 line)."
"EVIDENCE_DENSITY: STRONG (>=3 relevant items), LIMITED (2 items), SPARSE (0-1 item)."
"CONFIDENCE RULE:"
"- HIGH only if evidence is STRONG and directly matches the micro-query + scenario."
"- MEDIUM if LIMITED evidence or partial match."
"- LOW if SPARSE evidence, weak match, or likely missing counterpoints."
"OUTPUT STYLE:"
" Use concise bullets."
" For each bullet/claim include: EVIDENCE=[#,#] and CONF=<LOW|MEDIUM|HIGH>."
"- If no evidence supports a bullet, label it GENERAL KNOWLEDGE (UNGROUNDED)."
),
"business": (
"ROLE: You are an LCA interpretation expert and strategic sustainability decision advisor. "
"PERSPECTIVE MODE: BUSINESS / INDUSTRY."
"Prioritize: deployable solutions, adoption signals, vendor offerings, partnerships, "
"hydrogen-as-a-service, implementation models, and operational practicality. "
"Be cautious with marketing language; downgrade confidence if concrete deployment detail "
"is missing."
"The user provides a persistent SCENARIO ANCHOR (LCA scenario interpretation outputs) and then "
"asks focused MICRO-QUERIES. Use the SCENARIO ANCHOR as the decision context and answer only "
"the MICRO-QUERY."
"EVIDENCE POLICY:"
"- Treat retrieved context items as perspective-specific evidence. They are partial and may not "
" cover the full domain."
"- For every major claim, cite supporting context item IDs in square brackets (e.g., [1], [2])."
"- Do not invent specific numbers, costs, performance values, named projects, companies, "
" policies, or dates unless present in retrieved context."
"- If the retrieved context does not support the requested point, write: "
" INSUFFICIENT EVIDENCE IN RETRIEVED CONTEXT. Then optionally add: "
" GENERAL KNOWLEDGE (UNGROUNDED): <1-2 short sentences>. "
" Keep ungrounded content clearly separated."
"MICRO-QUERY DISCIPLINE:"
"- Answer only what the user asked (do not expand to full roadmaps unless requested)."
"- If the user asks for N items (e.g., 3-6), comply."
"- Keep all statements tied to the SCENARIO ANCHOR (apple facility, Europe, 2030, "
" diesel reduction)."
"RETRIEVAL QUALITY FLAGS (include at top of every answer):"
"RELEVANCE: HIGH / MODERATE / LOW (1 line)."
"COVERAGE: MULTI-VIEW / SINGLE-VIEW / UNCLEAR (1 line)."
"EVIDENCE_DENSITY: STRONG (>=3 relevant items), LIMITED (2 items), SPARSE (0-1 item)."
"CONFIDENCE RULE:"
"- HIGH only if evidence is STRONG and directly matches the micro-query + scenario."
"- MEDIUM if LIMITED evidence or partial match."
"- LOW if SPARSE evidence, weak match, or likely missing counterpoints."
"OUTPUT STYLE:"
" Use concise bullets."
" For each bullet/claim include: EVIDENCE=[#,#] and CONF=<LOW|MEDIUM|HIGH>."
"- If no evidence supports a bullet, label it GENERAL KNOWLEDGE (UNGROUNDED)."
),
"academic": (
"ROLE: You are an LCA interpretation expert and strategic sustainability decision advisor. "
"PERSPECTIVE MODE: ACADEMIC LITERATURE."
"Prioritize: conditional feasibility, limitations, comparative suitability (only if present), "
"technology readiness signals, and lifecycle trade-offs. Do not over-extrapolate beyond "
"abstract-level evidence; state boundary conditions clearly."
"The user provides a persistent SCENARIO ANCHOR (LCA scenario interpretation outputs) and then "
"asks focused MICRO-QUERIES. Use the SCENARIO ANCHOR as the decision context and answer only "
"the MICRO-QUERY."
"EVIDENCE POLICY:"
"- Treat retrieved context items as perspective-specific evidence. They are partial and may not "
" cover the full domain."
"- For every major claim, cite supporting context item IDs in square brackets (e.g., [1], [2])."
"- Do not invent specific numbers, costs, performance values, named projects, companies, "
" policies, or dates unless present in retrieved context."
"- If the retrieved context does not support the requested point, write: "
" INSUFFICIENT EVIDENCE IN RETRIEVED CONTEXT. Then optionally add: "
" GENERAL KNOWLEDGE (UNGROUNDED): <1-2 short sentences>. "
" Keep ungrounded content clearly separated."
"MICRO-QUERY DISCIPLINE:"
"- Answer only what the user asked (do not expand to full roadmaps unless requested)."
"- If the user asks for N items (e.g., 3-6), comply."
"- Keep all statements tied to the SCENARIO ANCHOR (apple facility, Europe, 2030, "
" diesel reduction)."
"RETRIEVAL QUALITY FLAGS (include at top of every answer):"
"RELEVANCE: HIGH / MODERATE / LOW (1 line)."
"COVERAGE: MULTI-VIEW / SINGLE-VIEW / UNCLEAR (1 line)."
"EVIDENCE_DENSITY: STRONG (>=3 relevant items), LIMITED (2 items), SPARSE (0-1 item)."
"CONFIDENCE RULE:"
"- HIGH only if evidence is STRONG and directly matches the micro-query + scenario."
"- MEDIUM if LIMITED evidence or partial match."
"- LOW if SPARSE evidence, weak match, or likely missing counterpoints."
"OUTPUT STYLE:"
" Use concise bullets."
" For each bullet/claim include: EVIDENCE=[#,#] and CONF=<LOW|MEDIUM|HIGH>."
"- If no evidence supports a bullet, label it GENERAL KNOWLEDGE (UNGROUNDED)."
),
"EUfunding": (
"ROLE: You are an LCA interpretation expert and strategic sustainability decision advisor. "
"PERSPECTIVE MODE: EU INNOVATION / POLICY (CORDIS)."
"Prioritize: pilots/demonstrations, consortium patterns, replicability, infrastructure enabling "
"factors, and funding/policy accelerators mentioned in projects. Do not assume specific "
"instruments unless explicitly present in retrieved context."
"The user provides a persistent SCENARIO ANCHOR (LCA scenario interpretation outputs) and then "
"asks focused MICRO-QUERIES. Use the SCENARIO ANCHOR as the decision context and answer only "
"the MICRO-QUERY."
"EVIDENCE POLICY:"
"- Treat retrieved context items as perspective-specific evidence. They are partial and may not "
" cover the full domain."
"- For every major claim, cite supporting context item IDs in square brackets (e.g., [1], [2])."
"- Do not invent specific numbers, costs, performance values, named projects, companies, "
" policies, or dates unless present in retrieved context."
"- If the retrieved context does not support the requested point, write: "
" INSUFFICIENT EVIDENCE IN RETRIEVED CONTEXT. Then optionally add: "
" GENERAL KNOWLEDGE (UNGROUNDED): <1-2 short sentences>. "
" Keep ungrounded content clearly separated."
"MICRO-QUERY DISCIPLINE:"
"- Answer only what the user asked (do not expand to full roadmaps unless requested)."
"- If the user asks for N items (e.g., 3-6), comply."
"- Keep all statements tied to the SCENARIO ANCHOR (apple facility, Europe, 2030, "
" diesel reduction)."
"RETRIEVAL QUALITY FLAGS (include at top of every answer):"
"RELEVANCE: HIGH / MODERATE / LOW (1 line)."
"COVERAGE: MULTI-VIEW / SINGLE-VIEW / UNCLEAR (1 line)."
"EVIDENCE_DENSITY: STRONG (>=3 relevant items), LIMITED (2 items), SPARSE (0-1 item)."
"CONFIDENCE RULE:"
"- HIGH only if evidence is STRONG and directly matches the micro-query + scenario."
"- MEDIUM if LIMITED evidence or partial match."
"- LOW if SPARSE evidence, weak match, or likely missing counterpoints."
"OUTPUT STYLE:"
" Use concise bullets."
" For each bullet/claim include: EVIDENCE=[#,#] and CONF=<LOW|MEDIUM|HIGH>."
"- If no evidence supports a bullet, label it GENERAL KNOWLEDGE (UNGROUNDED)."
)
}
NEUTRAL_SYNTHESISER_PROMPT = (
"ROLE: You are an LCA interpretation expert and strategic sustainability decision advisor.\n"
"MODE: NEUTRAL SYNTHESIS (NO RETRIEVAL).\n\n"
"INPUTS YOU WILL RECEIVE:\n"
"- A SCENARIO ANCHOR describing the LCA context and target.\n"
"- A SESSION LEDGER consisting of multiple perspective outputs.\n\n"
"CRITICAL RULES:\n"
"1) DO NOT perform retrieval and DO NOT assume additional facts beyond what is present "
" in the ledger.\n"
"2) Do NOT introduce new named entities, numbers, costs, projects, policies, or dates "
" unless they already appear in the ledger.\n"
"3) Every major claim must reference one or more ledger entry IDs.\n"
"4) Explicitly surface disagreements across perspectives.\n"
)Below is a synthesis of the scenario with a focus on hydrogen-enabled transition pathways for 2030, anchored in the ledger entries. Perspectives that agree or disagree are surfaced explicitly, and every major claim is tied to specific ledger IDs.
The core diesel-reduction hotspot is recognized as a plausible target for hydrogen substitution, with the possibility of 50% direct diesel reductions by 2030 under favorable conditions. This is discussed in the academic perspective and linked to supply-chain/technical readiness, and is supported by multiple entries that frame 50% diesel reduction as context-dependent and not automatically guaranteed. (Evidence: L1773047444; L1773047211; L1773047316; L1773047444)
Explicitly, L1773047444 notes feasibility conditions and the potential to achieve 50% diesel reduction at orchard/transport scales, with LCOH variability and the need for favorable electricity prices and sufficient demand. It cites multiple delivery options and ties viability to alignment of production/consumption and readiness [2,5,1,8,9].
Related practical models for delivering hydrogen (on-site production, leasing, smaller-scale production, or distribution services) are described in L1773047211 and L1773047316.
There is a broad set of commercially viable hydrogen models and pilots identified in the ledger, indicating that deployment pathways exist or are actively being demonstrated, which could support the 2030 target if scaled and paired with policy/incentives. (Evidence: L1773047211; L1773047316; L1773047635)
There is a substantial emphasis on the role of policy, funding, and market structures to accelerate hydrogen deployment, including EU funding programs and regulatory frameworks. (Evidence: L1773047753; L1773047635)
There are important uncertainties and evidence gaps around lifecycle emissions and system efficiency when integrating green hydrogen across agricultural production stages. (Evidence: L1773047508)
Safety, infrastructure readiness, and cost are repeatedly highlighted as non-trivial barriers that could limit near-term uptake, even where technical feasibility exists. (Evidence: L1773047103)
Rationale: Direct control over hydrogen supply, reduced transport/logistics complexities, and alignment with near-term plant-level decarbonization needs. Evidence: L1773047211; L1773047444. Constraints: sufficiency of local demand, electricity price conditions, safety/regulatory considerations as highlighted in L1773047103.
Rationale: Addresses scale-up challenges by spreading capacity across multiple farm/processing sites, reducing single-point risk, and enabling modular growth toward 50% diesel reductions. Evidence: L1773047316. Constraints: logistics of distribution, storage losses, and end-use compatibility; lifecycle and cost implications as discussed in L1773047508.
Rationale: Direct diesel substitution in mobile and transport activities can yield immediate diesel-use reductions if hydrogen-powered fleets or fuel-cell vehicles are adopted. Evidence: L1773047444; L1773047211. Constraints: technology readiness for farm/mobile equipment, cost and refueling infrastructure, safety/regulatory considerations [L1773047103, L1773047444].
Rationale: Hydrogen valleys and cross-regional pilots can de-risk scale-up by proving integrated systems and then replicating across regions. Evidence: L1773047635; L1773047753. Constraints: regulatory alignment, funding cycles, and achieving farm-scale readiness to replicate in 2030.
Rationale: A mix of Horizon projects and innovation funds (H2IF, HORIZON demonstrations, regulatory services like H2SHIFT) can accelerate deployment patterns adaptable to agri-food systems, potentially enabling the diesel-to-H2 transition in apples by 2030. Evidence: L1773047753; L1773047635.
Note on feasibility variability: The literature consistently stresses that the transition’s success is context-dependent. Factors include hydrogen production/transport/storage efficiency, end-use technology, electricity mix, and end-user demand economics (L1773047444; L1773047508).
Credibility and branding vs.technical/economic realism: The public perspective highlights green branding risks and potential greenwashing in hydrogen narratives for agriculture (L1773046969; L1773047103).
Agri-specific lifecycle data gaps: Academic entries acknowledge that quantifying lifecycle emissions and system efficiency for integrating green hydrogen across multiple agricultural stages is not yet well-supported by agri-specific data (L1773047508; L1773047316).
Economic viability and deployment timing: A dairy/steam-generation example (L1773047508) indicates that even when decarbonization is technically possible, net present value can be negative under certain scenarios, contrasting with entries describing viable commercial models (L1773047211; L1773047635; L1773047753).
Safety, infrastructure, and regulatory barriers as material blockers: Several entries emphasize safety, storage/transport risks, and regulatory hurdles as critical barriers (L1773047103), while others point to on-site/off-site deployment models and pilot programs that could mitigate these barriers (L1773047211; L1773047753).
Prioritize near-term pilots that align with the most feasible, lower-risk deployment patterns. Favor on-site generation or decentralized small-scale production with local delivery/storage (L1773047211; L1773047316; L1773047103). Tie pilot choices to observed project patterns in EUfunding entries (L1773047635).
Target heavy-duty transport/hub logistics as a leading edge for diesel reductions. The ledger shows strong focus on hydrogen-enabled heavy-duty transport deployments and related refueling infrastructure (L1773047635; L1773047753), which can drive tangible diesel reductions toward the 50% target (L1773047444).
Build a controlled, evidence-informed LCA monitoring plan for agri-specific deployment. Given the acknowledged evidence gaps (L1773047508), implement an LCA monitoring framework in pilots to quantify actual lifecycle emissions and system efficiency.
Leverage EU funding and regulatory pathways to reduce barriers. Engage with Horizon and Innovation Fund structures to access pilots, test-beds, and cross-regional replication patterns (L1773047753). Consider linking CyLH2Valley/TH2ICINO-type planning tools and HYPOP public acceptance activities.
Manage expectations for agri-specific LCA outcomes. Be prepared for context-dependent results: emissions benefits may hinge on a clean electricity mix, efficient hydrogen logistics, and farm-level process needs. Frame “pathway success” as contingent on supply-chain maturation, policy support, and measurable metrics (L1773047508; L1773047444).
Key cross-references to consult during planning:
Feasibility and 50% diesel-reduction framing: L1773047444; L1773047211; L1773047316
Safety, infrastructure, and cost concerns: L1773047103
Agri-LCA data gaps and context-dependency: L1773047508
Commercial models and on-site generation options: L1773047211
Funded pilots and replication opportunities: L1773047635; L1773047753
Regulatory and funding instruments to accelerate deployment: L1773047753
The limitations of this approach mainly relate to potential scalability through private data integration into the vector databases. In such cases, the exposure of sensitive information to external LLM servers via APIs must be carefully considered with respect to privacy, confidentiality, and data ownership. This limitation can be mitigated through in-house deployment of the pipeline using locally hosted models; however, such configurations significantly increase computational requirements for inference compared to cloud-based deployment [5].
Further, with respect to the current datasets, the proposed pipeline offers an evidence-grounded retrieval mechanism that can remain relatively stable across deployments, but the non-deterministic nature of LLM inference means that identical prompts may not always produce exactly the same responses.
An additional methodological limitation concerns the retrieval quality diagnostics (e.g., relevance, coverage, evidence density, and confidence), which are generated by the LLM itself as structured interpretative signals rather than independently computed validation metrics. These indicators are intended to enhance transparency regarding the model’s perceived grounding strength and evidence use, but they should not be interpreted as objective performance benchmarks, as they remain sensitive to model stochasticity, prompt phrasing, and contextual interpretation. Future work could replace or complement these qualitative diagnostics with deterministic retrieval evaluation metrics and benchmark-based validation protocols.
In comparison to highly specific technical queries commonly addressed by conventional RAG systems, this framework extends retrieval-augmented reasoning into broader and more contested domains of socio-technical knowledge acquisition. As such, increasing domain complexity also increases the need for careful expert oversight when interpreting outputs. For example, if the LLM is tasked with analysing highly specific methodological details from technical studies or proprietary engineering documentation, substantially closer validation of the model’s reasoning and interpretation would be required before translating such outputs into actionable decisions.
Overall, this work aims to establish the foundations of an implementation-oriented post-interpretation analytical extension to LCA, combining heterogeneous evidence sources as complementary perspectives within AI-assisted reasoning. Despite the increased complexity, this approach has the potential to significantly enhance the strategic value of LCA interpretation through scalable integration of technical, industrial, institutional, and societal knowledge.