ForecastAgentSearch: Towards a Multi-Expert Agent Search System for Geopolitical Event Forecasting


1 Extended Abstract↩︎

Geopolitical event forecasting aims to anticipate future political, military, and social developments from historical observations and evolving real-world contexts [1][3]. It is important for early warning, risk assessment, policy planning, and strategic decision-making in high-impact scenarios such as international conflicts, diplomatic actions, sanctions, protests, humanitarian crises, and regional instability [1], [4], [5]. This task is especially challenging in regions such as the Middle East, where future events are often shaped by intertwined factors, including regional power dynamics, political alliances, economic pressure, religious and cultural tensions, historical grievances, and multimodal media narratives [6], [7]. Therefore, effective forecasting requires not only temporal reasoning over historical events, but also the ability to identify which evidence and expertise are useful for a specific forecasting task.

Recent advances in large language models (LLMs) have created new opportunities for event forecasting. LLM-based methods can process textual contexts, retrieve relevant historical evidence, and generate predictions through prompting, in-context learning, chain-of-thought reasoning, or retrieval-augmented generation [8][11]. However, most existing approaches still rely on either a single general-purpose predictor or a fixed expert ensemble. A single predictor may follow a dominant reasoning path and overlook alternative geopolitical perspectives, while a fixed expert ensemble may introduce redundant, irrelevant, or costly expert opinions [6], [12], [13]. Such designs are insufficient for complex geopolitical forecasting, where different queries may require different combinations of regional, political, economic, religious, historical, multimodal, and risk-oriented expertise.

Recent forecasting studies suggest two important observations. First, heterogeneous evidence sources may play different functional roles in prediction. For example, multimodal evidence can either highlight salient historical events or provide complementary context beyond textual descriptions [7]. Second, specialized expert models can provide complementary predictive knowledge, but their usefulness is often query-dependent [6], [13]. These observations indicate that reliable geopolitical forecasting requires more than stronger predictive models or larger context windows. It also requires a principled mechanism for deciding which experts should be consulted, how they should be ranked, and how their outputs should be coordinated.

To this end, we introduce ForecastAgentSearch, a system-level formulation that treats complex geopolitical event forecasting as a multi-expert agent search problem. Rather than directly mapping historical contexts to predictions, ForecastAgentSearch introduces an intermediate search and coordination layer over specialized expert agents. In this formulation, expert agents are treated as searchable, rankable, and composable resources, which naturally connects geopolitical forecasting with core problems in information retrieval and agent search [14][16].

Figure 1: Overview of ForecastAgentSearch. Given a geopolitical forecasting query and heterogeneous evidence, the system first identifies task-specific expertise requirements, then searches over a fine-grained expert agent space, retrieves and ranks the most suitable agents, and coordinates their outputs to generate the final forecast with explanations and uncertainty estimates.

As illustrated in Figure 1, ForecastAgentSearch consists of three main stages. The first stage takes a forecasting task and heterogeneous evidence as input, including the forecasting query, historical events, news context, and multimodal evidence. These sources provide complementary signals for prediction: historical events capture temporal actor interactions, news reports provide contextual narratives, and multimodal evidence may reveal additional regional or situational information. Instead of treating all evidence uniformly, ForecastAgentSearch performs task understanding to infer the expertise requirements behind the current query, such as regional locality, political relations, economic pressure, religious or cultural factors, security risks, and historical context.

The second stage searches over a fine-grained expert agent space. Each expert agent is associated with a lightweight profile that describes its region or actor specialization, domain expertise, supported evidence sources, reliability estimate, inference cost, and known limitations. For example, a forecasting query may require experts on Israeli political dynamics, Saudi cultural factors, U.S. policy, oil markets, religious tensions, multimodal conflict evidence, or regional security risks. These profiles allow expert agents to be indexed and retrieved according to both structured metadata and semantic descriptions.

Given a forecasting query, ForecastAgentSearch retrieves and ranks candidate agents according to several task-aware criteria, including relevance to the query, regional or actor locality, historical reliability, inference cost, and complementarity with other selected experts. Rather than consulting all available agents, the system selects a compact set of top-ranked experts that can provide complementary perspectives for the current task. This design reflects the nature of geopolitical forecasting: the usefulness of an expert is highly dependent on the specific event, region, actors, and evidence sources involved.

The final stage coordinates the selected experts to produce the forecast. Each expert can contribute a prediction, an intermediate analysis, supporting evidence, a confidence estimate, or possible risk factors from its own perspective. A coordination module then synthesizes these outputs into the final forecast, together with explanations and uncertainty signals. When experts disagree, the coordinator can compare their evidence, reliability, and domain coverage, rather than simply averaging their predictions. This process resembles a structured think-tank workflow, where multiple specialists contribute different views and a coordinator synthesizes them into a coherent judgment.

The main contribution of this extended abstract is threefold. First, we formulate complex geopolitical event forecasting as an expert-agent search problem, shifting the focus from prediction alone to task-aware expert selection and coordination. Second, we outline ForecastAgentSearch, a system framework that retrieves, ranks, and coordinates specialized agents with different forms of expertise, including regional, political, economic, religious, historical, multimodal, and risk-oriented knowledge. Third, we discuss Middle East event forecasting as a representative testbed for this formulation, due to its complex regional interactions, heterogeneous evidence sources, and diverse analytical perspectives.

Although this extended abstract focuses on the problem formulation and system design, ForecastAgentSearch naturally suggests several evaluation directions. Forecasting quality can be measured by accuracy, Brier score, log score, calibration error, temporal generalization, and uncertainty quality. The quality of agent search can be evaluated by determining whether retrieved experts match the required regions, domains, actors, and evidence sources. Future ablations can compare task-aware agent search with single-predictor forecasting, fixed expert ensembles, all-expert consultation, random expert selection, and variants without reliability-, cost-, or complementarity-aware ranking.

2 Conclusion↩︎

This extended abstract introduces ForecastAgentSearch, which formulates complex geopolitical event forecasting as a multi-expert agent search problem. Instead of relying on a single predictor or a fixed expert ensemble, ForecastAgentSearch treats expert agents as searchable and composable resources, and aims to retrieve, rank, and coordinate suitable experts according to task-specific requirements.

Motivated by recent findings that both evidence utility and expert usefulness are task-dependent, ForecastAgentSearch highlights the need for structured expert profiling, task-aware retrieval, reliability- and cost-aware ranking, complementary team formation, and interpretable aggregation. Middle East event forecasting serves as a meaningful testbed, since it requires reasoning over regional, political, economic, religious, historical, and multimodal factors.

References↩︎

[1]
Liang Zhao.2021. . Comput. Surveys54, 5(2021), 1–37.
[2]
Songgaojun Deng, Maarten de Rijke, and Yue Ning.2024. . In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6459–6469.
[3]
Yunshan Ma, Chenchen Ye, Zijian Wu, Xiang Wang, Yixin Cao, and Tat-Seng Chua.2023. . In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1643–1652.
[4]
Woojeong Jin, Rahul Khanna, Suji Kim, Dong-Ho Lee, Fred Morstatter, Aram Galstyan, and Xiang Ren.2021. . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 4636–4650.
[5]
Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.2024. . Advances in Neural Information Processing Systems37(2024), 50426–50468.
[6]
Haoxuan Li, He Chang, Yunshan Ma, Yi Bin, Yang Yang, See-Kiong Ng, and Tat-Seng Chua.2026. . In WWW.
[7]
Haoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin, Yang Yang, and Tat-Seng Chua.2024. . In MM.
[8]
Ruotong Liao, Xu Jia, Yangzhe Li, Yunpu Ma, and Volker Tresp.2024. . In Findings of the association for computational linguistics: NAACL 2024. 4303–4317.
[9]
Ruilin Luo, Tianle Gu, Haoling Li, Junzhe Li, Zicheng Lin, Jiayi Li, and Yujiu Yang.2024. . arXiv preprint arXiv:2401.06072(2024).
[10]
He Chang, Chenchen Ye, Zhulin Tao, Jie Wu, Zhengmao Yang, Yunshan Ma, Xianglin Huang, and Tat-Seng Chua.2024. . arXiv preprint arXiv:2407.11638(2024).
[11]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al2020. . Advances in neural information processing systems33(2020), 9459–9474.
[12]
Chenchen Ye, Ziniu Hu, Yihe Deng, Zijie Huang, Mingyu Derek Ma, Yanqiao Zhu, and Wei Wang.2024. . arXiv preprint arXiv:2407.01231(2024).
[13]
Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim, and Jiayi Huang.2024. . Authorea Preprints(2024).
[14]
Bin Wu, Arastun Mammadli, Xiaoyu Zhang, and Emine Yilmaz.2026. . arXiv preprint arXiv:2604.22436(2026).
[15]
Zhengliang Shi, Yuhan Wang, Lingyong Yan, Pengjie Ren, Shuaiqiang Wang, Dawei Yin, and Zhaochun Ren.2025. . In Findings of the Association for Computational Linguistics: ACL 2025. 24497–24524.
[16]
Norbert Braunschweiler, Rama Doddipatla, and Tudor-Catalin Zorila.2025. . In Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM). 75–83.