July 07, 2026
Data from Singapore indicated that about 31% of the population had evidence of Helicobacter pylori infection. Persistent H. pylori infection is associated with chronic active gastritis and peptic ulcer disease, and its eradication is key to gastric cancer prevention. However, evidence supporting H. pylori positivity and H. pylori-associated gastritis may be distributed across heterogeneous coded and free-text report fields and may require contextual interpretation of assertion and negation, limiting keyword search, and making manual review difficult to scale. We conducted a retrospective pilot evaluation of the Nimblemind Multi-Agent System (nMAS), a field-name-driven, evidence-linked extraction workflow, using 54 de-identified gastric biopsy pathology reports from a large healthcare system in Singapore. Four clinician-scoped binary fields were evaluated: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. Across 216 feature-case decisions, nMAS correctly classified 213, corresponding to 98.61% overall accuracy. A separately implemented UMA-style MiniMax M2.5 comparator produced similar aggregate and per-field classification metrics. Although predictive performance was similar, nMAS maintained unified report-level outputs with supporting source sentences; the demonstrated contribution is therefore workflow integration and traceability rather than predictive superiority. Under an illustrative, unmeasured scenario, reviewing 1,000 reports at five minutes per manual review versus five seconds per evidence-linked verification would reduce review time from 83.3 to 1.4 staff-hours, corresponding to 81.9 staff-hours and about USD 6,100 in potential staff-time value. Larger multi-institutional studies should evaluate evidence-span correctness, clinician verification time, and generalizability.
clinical NLP, pathology reports, large language models, multi-agent systems, information extraction, Helicobacter pylori, gastric biopsy, H. pylori-associated gastritis
About 31% of the Singapore population was estimated to have evidence of Helicobacter pylori infection, with prevalence increasing with age and varying across ethnic groups [1]. Persistent H. pylori infection is associated with chronic active gastritis and peptic ulcer disease, while eradication is central to gastric cancer prevention [2]–[4]. Reliable identification of biopsy-confirmed H. pylori-positive reports is therefore important for treatment review, eradication follow-up, clinical audit, research cohort assembly, and quality-improvement workflows.
Pathology-based case finding is not a simple keyword-search task. Identical organism terms may appear in affirmative, negated, historical, or ancillary-stain contexts; for example, “Helicobacter organisms are identified” and “No Helicobacter organisms are identified” contain the same target terms but require opposite labels. Relevant evidence may also be distributed across coded specimen fields, specimen labels, diagnoses, microscopic descriptions, and ancillary-test comments, making manual review difficult to scale. Slow manual review can delay treatment review, eradication follow-up, and quality-improvement action, which may limit timely patient care[5]–[7]. At five minutes per report, screening 1,000 candidate pathology reports would require about 83 staff-hours before downstream review or analysis could begin.
Rule-based and supervised extraction systems can improve classification but often require task-specific annotation, maintenance, or retraining as report templates and target schemas change. Large language models provide greater configurability, but extracted labels remain difficult to use for clinical audit or research unless they are linked to supporting source text. These limitations motivate a reusable extraction workflow that combines configurable field definitions, source-text grounding, validation, and clinician-reviewable outputs [8]–[11].
To address this gap, we evaluate the Nimblemind Multi-Agent System (nMAS), a modular, field-name-driven, evidence-linked workflow configured for H. pylori-related feature extraction from gastric biopsy pathology reports. nMAS routes user-specified target fields through tiered extraction modules and returns structured labels with supporting source-text evidence for clinician verification.
In this pilot study, we applied nMAS to 54 de-identified gastric biopsy pathology reports from a large healthcare system in Singapore. We evaluated four clinician-scoped binary fields: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. We assessed whether the workflow could support initial H. pylori case finding while preserving source-linked evidence for clinician verification.
Baseline Clinical and Pathology Information Extraction. Manual review remains the reference standard when interpretation depends on section context, negation, ambiguous wording, and local reporting conventions, but it is difficult to scale. In breast pathology abstraction, Wieneke et al. found that an NLP system still flagged 49.1% of reports for manual review and incorrectly coded 30.8% of reports [12].
Keyword-based retrieval can identify reports containing predefined terms but does not by itself resolve assertion status, section location, or diagnostic context [8], [13].
Rule-based NLP systems extend keyword matching through dictionaries, regular expressions, section parsing, assertion detection, and negation handling. For example, the Mayo Clinical Text Analysis and Knowledge Extraction System (cTAKES) demonstrated modular clinical text processing across electronic health record applications [13], and rule-based methods have also been applied to pathology report extraction [8], [14].
Learning-Based and LLM-Based Medical Abstraction. Supervised and transformer-based methods can model clinical context beyond exact keyword matching. Domain-specific language models such as BioBERT have improved performance across biomedical named-entity recognition, relation extraction, and question-answering tasks [15]. However, supervised clinical information extraction systems typically depend on task-specific annotated data and model development, which can limit their flexibility when extraction targets or output schemas change [8], [16]. This limitation is relevant when clinicians need to define new variables for different audits, research cohorts, or disease-specific workflows.
Large language models (LLMs) have enabled more flexible zero-shot and prompt-based abstraction. UniMedAbstractor (UMA) uses configurable prompt templates and frontier models to structure multiple clinical attributes from real-world data [16]. In pathology, Truhn et al. evaluated GPT-4 for zero-shot extraction from colorectal cancer and glioblastoma histopathology reports [9], while Balasubramanian et al. compared multiple LLMs for structured extraction from breast cancer pathology reports [10]. These studies demonstrate that LLM-based abstraction can be configured for different extraction schemas. However, clinical use still requires reproducible schemas, reliable handling of uncertainty and negation, and mechanisms for identifying unsupported outputs[17], [18].
Current Solutions in Singapore and the Region. Clinical natural language processing (NLP) has also been evaluated in Singapore and other Asian healthcare settings, but directly comparable work on H. pylori gastric biopsy pathology extraction remains limited. In Singapore, Hardjojo et al. developed the Clinical History Extractor for Syndromic Surveillance (CHESS), a rule-based system that extracted infectious-disease symptoms from free-text primary-care electronic medical records and distinguished affirmed, negated, and suspected assertions [19]. Tay et al. evaluated NLP approaches to infer metastatic disease sites from radiology reports in a Singapore oncology setting [20]. These studies support the feasibility of structuring local clinical free text, but they address primary-care surveillance and oncology radiology rather than gastric biopsy pathology.
Regional gastrointestinal studies provide closer comparisons to the present task. In Korea, Song et al. developed an NLP pipeline for extracting gastric disease information from unstructured esophagogastroduodenoscopy reports and linked pathology reports, including categories related to H. pylori-associated gastritis [21]. Bae et al. developed a related pipeline for extracting quality indicators from free-text colonoscopy and pathology reports, demonstrating how automated abstraction can support large-scale gastrointestinal quality monitoring [22]. These studies show that gastrointestinal report extraction is feasible in Asian healthcare settings. However, they primarily evaluate task-specific pipelines and do not address a configurable, field-name driven, evidence-linked multi-agent workflow for Singapore gastric biopsy reports.
Collectively, prior studies show that clinical free-text abstraction and gastrointestinal report extraction are feasible in local and regional healthcare settings [19]–[22]. However, they do not directly address the present operational use case: configurable extraction of H. pylori-related features from Singapore gastric biopsy pathology reports while preserving source-text evidence for clinician verification. This gap matters because biopsy-confirmed H. pylori case finding supports eradication follow-up, treatment-outcome audits, research cohort assembly, and quality-improvement review [5]–[7]. The present study addresses this gap by evaluating nMAS across four clinician-scoped target features: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis.
Dataset Overview. The evaluation dataset consists of 54 de-identified gastric biopsy pathology report records from the Department of Anatomical Pathology, Singapore General Hospital (SGH). SGH is a large tertiary hospital serving a multi-ethnic population [23], [24].
Each report-level record contains de-identified administrative metadata, coded specimen and diagnosis fields, ordering and procedure information, sign-out information, and the full pathology report text. In the source table, these fields include specimen identifier, receive date, de-identified patient fields, TCode and MCode fields, diagnosis report text, order location, procedure code, sign-out field, reference target labels, and reference evidence spans. Direct patient identifiers were removed or replaced with de-identified placeholders before analysis, while clinically relevant wording needed for extraction was preserved.
The reports represent heterogeneous free-text reporting styles across different pathologists, including variation in specimen labels, diagnostic phrasing, organism descriptions, negation patterns, microscopic descriptions, and ancillary test wording. A typical record may include both coded tissue fields, such as ‘T57010 (Gastric biopsy),’’ and diagnostic text, such as ‘Stomach; biopsy. Severe chronic acute antral and body gastritis with mild colonization by Helicobacter pylori.’’ This illustrates why the extraction task requires both structured and unstructured evidence: coded fields help identify the gastric biopsy specimen, while diagnostic text may support H. pylori positivity and H. pylori-associated gastritis.
Table 1 shows representative de-identified examples from the evaluation dataset with their corresponding reference labels. These examples are included to illustrate the evaluation-data format and gold-label definitions. Prompt-level examples, field definitions, and guardrails used to guide extraction are described separately in the Method section.
| Case | Report Excerpt | Reference Target Labels |
|---|---|---|
| A | “Stomach; biopsy. Severe chronic acute antral and body gastritis with mild colonization by Helicobacter pylori.” | Gastric/stomach biopsy: Y; Biopsy: Y; H. pylori positive: Y; H. pylori gastritis: Y. |
| B | “Gastric biopsy: Chronic gastritis. No Helicobacter organisms are identified. Helicobacter organisms are absent.” | Gastric/stomach biopsy: Y; Biopsy: Y; H. pylori positive: N; H. pylori gastritis: N. |
| C | “Gastric biopsy: Helicobacter pylori-associated active chronic gastritis. Helicobacter organisms are identified.” | Gastric/stomach biopsy: Y; Biopsy: Y; H. pylori positive: Y; H. pylori gastritis: Y. |
Target Features. Four binary target features were evaluated. These features were selected through a clinician-informed scoping process. An initial broader set of potential gastric biopsy extraction targets was reviewed with four clinicians, including senior pathologists, to identify the minimum structured information needed for practical H. pylori-related case finding. Gastric/stomach biopsy and biopsy status were included as cohort-definition fields to confirm the anatomical site and specimen type. H. pylori positivity and H. pylori-associated gastritis were included as disease-relevance fields. Together, the four features distinguish relevant gastric biopsy specimens from non-target records and separate organism detection from an explicit diagnostic association between H. pylori and gastritis.
Gastric/Stomach Biopsy: whether the report indicates a gastric or stomach biopsy specimen.
Biopsy: whether the case is a biopsy specimen rather than a non-biopsy specimen type.
H. pylori Positive: whether H. pylori, Helicobacter organisms, or Helicobacter-like organisms are identified.
H. pylori Gastritis: whether the report supports H. pylori-associated gastritis, such as active chronic gastritis explicitly associated with H. pylori.
Reference Standard and Evidence Annotation. Clinician-reviewed reference labels were defined for each report and each of the four target features. Reference labels were determined using the complete de-identified report record, including coded fields, specimen labels, final diagnosis text, microscopic descriptions, organism-related statements, and ancillary stain comments where available. When initial labeling disagreements occurred, they were resolved through adjudication by an additional pathologist, and the final adjudicated labels were used as the reference standard. Reference evidence spans were recorded, where applicable, to document the source text supporting each reference label.
We evaluated the Nimblemind Multi-Agent System (nMAS), a modular clinical information extraction workflow configured for field-name driven extraction from de-identified gastric biopsy pathology reports. The workflow was applied to the four target features defined in Section III-B: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. Fig. 1 summarizes the implemented workflow.
Workflow Overview. As shown in Fig. 1, nMAS processes an input pathology report and a set of user-specified target fields through four stages: request assessment and parsing, complexity-based feature extraction, field-level result aggregation, and source-grounded output validation. The Achievability Agent and Query Parser prepare supported field requests for tiered extraction; FE-MUX then combines the resulting field-level outputs, which are validated against the source report before the final structured output is returned.
Input Standardization and Field Preparation. Before extraction, each report record was converted into a consistent textual representation. Formatting differences involving line breaks, spacing, section separators, and repeated administrative headers were normalized where appropriate. Clinically relevant content was preserved, including coded specimen information, specimen labels, final diagnosis text, microscopic descriptions, organism-related statements, ancillary test results, and negation.
Each target field was paired with an entry in the configurable FIELD_LIBRARY. A field entry specified the target meaning, expected output format, clinical context, in-context demonstrations, and field-specific guardrails. Clinical input
from four clinicians, including senior pathologists, informed the target-field definitions and the clinically relevant distinctions encoded in the demonstrations and guardrails. These components guided interpretation of the requested field at inference
time. The four target fields were configured according to the definitions described in Section [sec:target95features].
Achievability Agent and Query Parsing. The Achievability Agent assessed whether the input contained readable pathology content and whether each requested field was represented in the configured extraction schema. Requests lacking usable report text or a valid field definition were returned as unsupported rather than being assigned an inferred value.
For supported requests, the Query Parser constructed a field-specific extraction request containing the report text, target field, corresponding field definition, and expected output schema. It then assigned the request to an extraction tier according to the configured linguistic and reasoning requirements of the field. Routing did not use the reference labels.
Tiered Feature Extraction and Prompt Design. After query parsing, each supported field request was assigned to one of three extraction routes according to its configured linguistic and reasoning requirements.
Tier 1, Named Entity Recognition Feature Extraction, was used for fields supported primarily by direct lexical or coded evidence. This route consisted of a pre-processor, a natural language processing extraction component, and a post-processor that converted the result into the required format.
Tier 2, Small Language Model Feature Extraction, and Tier 3, Large Language Model Feature Extraction, were used for fields requiring contextual interpretation. Model assignments were determined before the present evaluation through engineering experiments using the same tiered nMAS workflow on a separate pathology information-extraction task. The broader candidate screening included the models ultimately selected for deployment, Qwen2.5-7B-Instruct and DeepSeek-V4-Flash, together with additional candidate models including GLM-5, Gemini 3.1 Pro, DeepSeek V3.2, Kimi K2 Thinking, and Qwen3-Next-80B [25]–[31]. Models were compared based on field-level extraction performance, structured-output reliability, source-text grounding, negation handling, and runtime efficiency.
Following this broader screening, DeepSeek-V4-Flash [26] was evaluated in the deployment workflow and achieved the best overall performance across these criteria. It was therefore selected for Tier 3, whereas Qwen2.5-7B-Instruct [25] was selected for Tier 2. Neither the 54 reports in the present study nor their reference labels were used for model selection or training.
Tier 2 was used for bounded contextual extraction, whereas Tier 3 was used for fields requiring broader contextual interpretation, including assertion-status assessment, negation handling, and the diagnostic association between H. pylori and gastritis. Each language-model route used a pre-processor to combine the report text with the relevant field definition, a model component to extract the requested information, and a post-processor to normalize the result into the required schema.
The extraction prompt used a shared pathology information-extraction instruction instantiated with the report text and requested target field. Field-specific interpretation was supplied through the corresponding FIELD_LIBRARY entry, while
the detailed demonstrations and guardrails are described in the following subsection. The model was required to return structured JSON with the extracted value and a verbatim supporting source sentence.
=’ plus .5em Use only information explicitly present in the report. Do not infer missing facts. Interpret the target feature using the field library, including its context, examples, and guardrails. If the requested value is absent or unclear, omit the feature. Every retained extraction must include a verbatim supporting source sentence, a confidence score, and explanatory notes.
These global guardrails applied to all target fields. They prohibited unsupported inference, required direct source-text evidence for every retained value, preserved conflicting supported mentions for review, and enforced the predefined JSON schema. When evidence appeared in multiple sections, definitive diagnostic interpretation was prioritized over less authoritative content such as clinical history or administrative metadata.
=’ plus .5em For h_pylori_positive, assign Y only when the report affirmatively identifies H. pylori, Helicobacter organisms, or
Helicobacter-like organisms in the gastric specimen. Assign N when the report explicitly states absence or non-identification, such as “No Helicobacter organisms are identified” or a negative H. pylori immunostain. Negation overrides
keyword presence.
For h_pylori_gastritis, assign Y only when the report explicitly links gastritis to H. pylori or Helicobacter organisms, such as “H. pylori-associated active chronic gastritis.” Gastritis alone is
insufficient, and organism positivity alone is insufficient unless the report makes the diagnostic association clear. Assign N when gastritis is present but H. pylori is absent or not diagnostically associated with the gastritis.
Field-Specific In-Context Demonstrations and Guardrails. Each field-specific demonstration paired a short pathology-style excerpt with an expected field value and a guardrail defining the relevant extraction boundary. As illustrated in Table 2, the demonstrations addressed gastric and stomach terminology, common biopsy abbreviations, affirmative and negated organism findings, ancillary-stain evidence, and the requirement for an explicit diagnostic association before assigning H. pylori-associated gastritis. The complete set of demonstrations is provided in Appendix Table 5.
| Target Feature | Pathology-Style Demonstration Excerpt | Expected Field Mapping | Guardrail Demonstrated | ||||
|---|---|---|---|---|---|---|---|
| Gastric Biopsy | “Stomach; biopsy.” | when the specimen is explicitly from stomach or gastric tissue. | Accept stomach and gastric as equivalent site terms. | ||||
| Biopsy | “GASTRIC BX.” | when biopsy is expressed using an accepted abbreviation such as “BX.” | Accept common biopsy abbreviations. | ||||
| H. pylori Positive | “No Helicobacter organisms are identified.” | when the organism statement is explicitly negated. | Negation overrides keyword presence. | ||||
| H. pylori Gastritis | “Helicobacter pylori associated active chronic gastritis.” | when gastritis is explicitly associated with H. pylori. | Require explicit diagnostic association. |
3pt
Feature Extraction Multiplexing. Outputs from the active extraction routes were passed to FE-MUX, which assembled the field-level results into a single report-level structure. The merged representation retained each predicted value together with its source sentence, date, confidence score, and explanatory notes where available. Schema compliance and source-grounding checks were performed by the subsequent validation stage.
Output Validation and Generation. The validation stage assessed both output structure and source grounding. It checked that each result followed the required schema, that the reported source sentence appeared verbatim in the input report, and that the sentence supported the predicted value in context.
Positive labels required affirmative source evidence. Negative labels required explicit absence or negation rather than lack of mention. For H. pylori-associated gastritis, evidence of organism positivity alone was insufficient without diagnostic evidence linking H. pylori to gastritis.
Missing, unsupported, or internally inconsistent outputs were flagged rather than accepted as fully grounded predictions. The generation stage then assembled the validated fields into the final report-level output for clinician review.
External UMA-Style Benchmark. To provide a relevant external comparator, we implemented a Universal Abstraction (UMA)-style benchmark based on the one-attribute prompting strategy described by Wong et al. [16]. UMA was selected because it is a validated, schema-conditioned clinical abstraction framework designed for configurable attribute
extraction without task-specific model training. This made it a closer comparator to the present field-name driven pathology task than a keyword search, a disease-specific rule set, or a supervised model requiring newly annotated training data. The
benchmark used the same 54 report records and the same four target features as the nMAS evaluation. For each report-feature pair, the model received only the Diagnosis (full report) text, the target attribute name, a concise target definition,
field-specific guidance, and positive and negative examples derived from the feature definitions in this manuscript. The prompt deliberately omitted reference labels, nMAS outputs, row-level correctness indicators, and any other gold-label information so
that the benchmark measured independent extraction rather than agreement with known answers or imitation of nMAS behavior.
Each UMA call requested one binary value and source evidence in a compact JSON object:
Figure 2:
.
The benchmark was run using MiniMax M2.5 [32], which was selected as a practical strong-LLM comparator because it was available through the deployment channel used for the broader benchmarking workflow and provided reliable structured JSON responses in preliminary checks.
Reference Standard and Evaluation Metrics. Extraction performance was evaluated using the clinician-reviewed reference labels and evidence annotations described in Section III-C. Each of the 54 reports contributed one prediction for each of the four target fields, yielding 216 feature-case evaluations for nMAS and 216 feature-case evaluations for the UMA-style MiniMax benchmark.
A prediction was counted as correct when its binary value matched the corresponding reference label. Missing, invalid, unsupported, or non-parseable predictions were counted as incorrect rather than being converted into negative labels. Accuracy was calculated separately for each field and across all feature-case evaluations. Positive-class precision, recall, and F1 score were also calculated using the clinician-reviewed labels as the reference standard.
Tier-3 Prompt-Component Ablation. To examine the contribution of field-specific prompt guidance, we conducted a paired ablation experiment on the two Tier 3 fields: H. pylori positivity and H. pylori-associated gastritis. These fields were selected for ablation because they were determined by clinical experts to be the most clinically relevant context-dependent targets for H. pylori-related case finding. Unlike the specimen-related fields, they require interpretation of assertion status, negation, ancillary-stain findings, and the explicit diagnostic association between H. pylori and gastritis.
The full condition used the shared Tier 3 extraction prompt together with the complete field-specific FIELD_LIBRARY entries, whereas the ablated condition omitted the field-specific definitions, aliases, clinical context, guardrails, and
in-context demonstrations.
For this ablation analysis, reports containing exact textual overlap with the in-context demonstrations were excluded to prevent demonstration wording from affecting the comparison between conditions. Both conditions were therefore evaluated on the same remaining 41 reports, yielding 82 paired feature-case decisions per condition.
The nMAS workflow correctly classified 213 of 216 feature-case decisions across 54 de-identified gastric biopsy pathology reports, corresponding to an overall feature-case accuracy of 98.61% (Table 3). Gastric/stomach biopsy identification and biopsy status each achieved 100.00% accuracy. Accuracy was 98.15% for H. pylori positivity and 96.30% for H. pylori-associated gastritis. The three observed errors occurred in these two disease-relevance fields.
| Method | Feature | N | Correct | Acc. | Prec. | Rec. | F1 |
|---|---|---|---|---|---|---|---|
| nMAS | Gastric/stomach biopsy | 54 | 54 | 100.00% | 100.00% | 100.00% | 100.00% |
| nMAS | Biopsy | 54 | 54 | 100.00% | 100.00% | 100.00% | 100.00% |
| nMAS | H. pylori positive | 54 | 53 | 98.15% | 100.00% | 96.43% | 98.18% |
| nMAS | H. pylori gastritis | 54 | 52 | 96.30% | 96.30% | 96.30% | 96.30% |
| UMA + MiniMax M2.5 | Gastric/stomach biopsy | 54 | 54 | 100.00% | 100.00% | 100.00% | 100.00% |
| UMA + MiniMax M2.5 | Biopsy | 54 | 54 | 100.00% | 100.00% | 100.00% | 100.00% |
| UMA + MiniMax M2.5 | H. pylori positive | 54 | 53 | 98.15% | 100.00% | 96.43% | 98.18% |
| UMA + MiniMax M2.5 | H. pylori gastritis | 54 | 52 | 96.30% | 96.30% | 96.30% | 96.30% |
| nMAS overall | Feature-case level | 216 | 213 | 98.61% | 99.38% | 98.77% | 99.08% |
| UMA + MiniMax M2.5 overall | Feature-case level | 216 | 213 | 98.61% | 99.38% | 98.77% | 99.08% |
4pt
The external UMA-style MiniMax M2.5 baseline produced the similar aggregate and per-field performance as nMAS. Both methods correctly classified all gastric/stomach biopsy and biopsy-status decisions and produced the same three errors in the two H. pylori-related fields.
Representative Evidence-Linked Output. In addition to binary feature labels, nMAS returned source-text evidence and supporting metadata for field-level review. The following representative output illustrates the structure returned for an affirmative H. pylori organism finding:
Figure 3:
.
This output format allowed each predicted label to be reviewed together with the source sentence supporting the extraction.
Tier-3 Field-Library Ablation. Table 4 summarizes the controlled field-library ablation on the 41 reports remaining after exclusion of exact textual overlap with the ICL demonstrations. The same report set was used for both the full and ablated conditions.
Across the 82 feature-case decisions per condition, both settings achieved high label-level performance. The full condition correctly classified 81 of 82 decisions, corresponding to 98.78% accuracy and 98.88% positive-class F1. The ablated condition produced the same label-level performance after excluding ICL-overlap reports. Performance for H. pylori positivity was perfect in both conditions. The only remaining error in both settings occurred for H. pylori-associated gastritis in the same report.
Removing field-specific guidance did not change label-level performance in this leakage-controlled pilot ablation; both conditions produced the same overall accuracy, F1 score, and remaining H. pylori-associated gastritis error.
| Cond. | Feature | N | Corr. | Acc. | Prec. | F1 |
|---|---|---|---|---|---|---|
| Full | H. pylori positive | 41 | 41 | 100.00% | 100.00% | 100.00% |
| Full | H. pylori gastritis | 41 | 40 | 97.56% | 100.00% | 97.78% |
| Full | Overall | 82 | 81 | 98.78% | 100.00% | 98.88% |
| Ablated | H. pylori positive | 41 | 41 | 100.00% | 100.00% | 100.00% |
| Ablated | H. pylori gastritis | 41 | 40 | 97.56% | 100.00% | 97.78% |
| Ablated | Overall | 82 | 81 | 98.78% | 100.00% | 98.88% |
2.2pt
Error Analysis. The three observed errors were confined to the two context-dependent disease-relevance fields. They comprised one false positive for H. pylori-associated gastritis and two false negatives from a single report: one for H. pylori positivity and one for H. pylori-associated gastritis. No errors occurred in either gastric/stomach biopsy identification or biopsy status. The same three errors were produced by nMAS and the external UMA-style MiniMax M2.5 baseline.
We highlight the following core findings. First, specimen-related fields were extracted correctly in all 54 reports, reflecting their reliance on explicit specimen labels or coded evidence. Second, all three errors occurred in the more context-dependent H. pylori positivity and H. pylori-associated gastritis fields, indicating that assertion status, negation, ancillary-stain findings, and explicit diagnostic association remain the main technical challenges.
These error patterns also explain why label-level performance alone is insufficient for this use case. In pathology case finding, a reviewer must be able to verify whether a predicted label is supported by the relevant sentence, especially when the same organism terms can appear in affirmative or negated contexts. Similar predictive performance therefore does not establish operational equivalence. UMA outputs required subsequent assembly into report-level results, whereas nMAS maintained one reviewer-facing contract across configuration, extraction, validation, and delivery. Field definitions, validation rules, or extraction models can therefore be revised without changing the clinician-facing output. The contribution is workflow integration and traceability rather than superior performance, consistent with the need for reusable clinical AI infrastructure [11].
The ablation analysis showed identical label-level performance with and without field-specific guidance. This finding should be interpreted in the context of the pilot task: the two Tier 3 H. pylori fields were clinically important but relatively constrained, and many reports contained direct diagnostic wording or explicit negation. The result therefore does not imply that field-specific guardrails and demonstrations are unnecessary in general. Rather, it suggests that the baseline prompt was already sufficient for this small and comparatively straightforward extraction task. For future requests involving larger schemas, more ambiguous clinical concepts, cross-sentence temporal reasoning, unsupported labels, or institution-specific terminology, field-specific guidance is expected to be more important for maintaining reliable and auditable extraction behavior.
To illustrate the possible operational effect in a Singapore context, consider screening 1,000 reports. At five minutes per manual review versus five seconds to verify an evidence-linked output, review time would fall from 83.3 to 1.4 staff-hours, saving 81.9 hours. Using a rounded representative hourly rate of USD 75, informed by Singapore salary benchmarks and the SGD–USD exchange rate, this corresponds to about USD 6,100 in potential staff-time value [33]–[36]. This sensitivity estimate is not a measured saving and excludes variation in implementation costs, report complexity, reviewer seniority, interface design, and the proportion of outputs requiring complete report review.
This pilot study was limited to 54 de-identified reports from a single institution and four binary target fields. It did not prospectively measure latency, implementation effort, clinician review time, or the accuracy and completeness of the returned evidence spans. Performance may also differ across other pathology templates, ancillary-test conventions, multilingual records, scanned documents, addenda, historical findings, and more complex target schemas. Future work should evaluate nMAS on larger, multi-institutional datasets, assess label and evidence-span correctness separately, and prospectively measure clinician review time, usability, and inter-reviewer agreement [18]. Reproducible deployment will additionally require versioned prompts and field definitions, execution logging, and monitoring for unsupported labels, missing evidence, negation errors, parse failures, and model or prompt drift. Until such validation is completed, nMAS should remain an evidence-linked case-finding and review-support workflow rather than a replacement for pathologist review or clinical judgment.
This pilot study evaluated nMAS for evidence-linked extraction of four H. pylori-related fields from 54 de-identified gastric biopsy pathology reports. nMAS achieved 98.61% overall feature-case accuracy and returned structured predictions with supporting source-text evidence. An external UMA-style MiniMax M2.5 comparator achieved the same classification performance, indicating that the present contribution lies in workflow integration, report-level aggregation, validation, and traceability rather than predictive superiority. Larger multi-institutional evaluations, independent evidence-span assessment, prospective clinician-review studies, and component-level ablations are needed before operational use.
The following de-identified report illustrates the source-record structure used for extraction, including coded specimen fields, diagnostic text, and gross-description content.
Specimen ID: 22:AB00001
Receive Date: 2022-01-11
Patient Name: Patient 1
ID: 0000001
Race: A
Sex: B
TCode:
T57010 (Gastric biopsy)
T57010 (Gastric biopsy)
T59604 (Rectal biopsy)
MCode:
GA003 (Severe)
M43000|T57010 (CHRONIC GASTRITIS)
M82110 (Tubular adenoma, NOS)
Diagnosis (full report):
CPOE
CPOE MESSAGE RECEIVED: 11/01/22 1214
HISTOPATHOLOGY: Y
VETTED & ORDER FORM COMPLETED (DR ONLY): Y
SPECIMEN LABEL COMPLETED: Y
ORDERSET:
Routine Specimen:
1. Rectal sigmoid polyp
2. Gastric BX
CLINICAL DIAGNOSIS: anaemia
TIME: 11:56
DIAGNOSIS
(1) Rectosigmoid polyp
TUBULAR ADENOMA WITH LOW GRADE DYSPLASIA.
(2) Stomach; biopsy
SEVERE CHRONIC ACUTE ANTRAL AND BODY GASTRITIS
WITH MILD COLONIZATION BY HELICOBACTER PYLORI.
THERE IS NO INTESTINAL METAPLASIA, DYSPLASIA
OR MALIGNANCY.
GROSS DESCRIPTION
The specimens are received in formalin, labelled
with patient's data and designated as follows.
(A) Rectal sigmoid polyp
It consists of a piece of tissue measuring 0.4 cm
in greatest dimension.
(A1-inked blue; no reserve)
(B) Gastric biopsy
It consists of 3 pieces of tissue measuring
0.1 cm to 0.3 cm in greatest dimension.
(B1-inked yellow; no reserve)
The specimens were fixed in formalin for 6--72 hours.
Order Location: S42A
Procedure:
HT.SPECIALBX2
HP.HE EMBED
Signout: AP-ABC
Table 5 shows the complete set of field-specific ICL demonstrations used to define the four target features. The demonstrations include multiple examples per feature and cover positive wording, negative wording, abbreviations, protocol-based biopsy descriptions, organism identification, ancillary stain evidence, and diagnostic association with gastritis.
| Target Feature | Pathology-Style Demonstration Excerpt | Expected Field Mapping | Guardrail Demonstrated | ||||
|---|---|---|---|---|---|---|---|
| Gastric Biopsy | “Stomach; biopsy.” | when the specimen is explicitly from stomach or gastric tissue. | Accept stomach and gastric as equivalent site terms. | ||||
| Gastric Biopsy | “GASTRIC ANTRUM BIOPSY.” | when the biopsy site is antrum, body, cardia, fundus, or another gastric subsite. | Gastric subsite wording supports gastric biopsy. | ||||
| Gastric Biopsy | “Stomach, Sydney protocol biopsy.” | when Sydney protocol biopsy is described as a stomach/gastric biopsy. | Recognize protocol-based gastric sampling. | ||||
| Biopsy | “Gastric biopsy. It consists of 3 pieces of tissue measuring 0.1 cm to 0.3 cm.” | when the specimen is explicitly described as biopsy tissue. | Use specimen type and gross description. | ||||
| Biopsy | ”GASTRIC BX.” | when biopsy is expressed using an accepted abbreviation such as ‘BX.’’ | Accept common biopsy abbreviations. | ||||
| Biopsy | “Random biopsies taken (Updated Sydney Protocol) antrum, incisura, and body.” | when the procedure states that biopsy samples were taken. | Procedure text can support biopsy status. | ||||
| H. pylori Positive | “MILD COLONIZATION BY HELICOBACTER PYLORI.” | when the report affirmatively states colonization by H. pylori. | Affirmative organism evidence supports positivity. | ||||
| H. pylori Positive | “Helicobacter organisms are identified.” | when Helicobacter organisms are explicitly identified. | Direct organism identification supports positivity. | ||||
| H. pylori Positive | “Helicobacter pylori-like organisms are highlighted by immunohistochemistry.” | when ancillary staining highlights H. pylori-like organisms. | Positive IHC evidence supports positivity. | ||||
| H. pylori Positive | “No Helicobacter organisms are identified.” | when the organism statement is explicitly negated. | Negation overrides keyword presence. | ||||
| H. pylori Positive | “No definite Helicobacter pylori identified.” | when the report states that definite organisms are not identified. | Do not infer positivity from uncertain or negative wording. | ||||
| H. pylori Positive | “No conspicuous Helicobacter pylori organisms are identified, corroborated by negative H. pylori immunostain.” | when negative morphology is supported by negative immunostain. | Negative IHC supports a negative value. | ||||
| H. pylori Gastritis | “Helicobacter pylori associated active chronic gastritis.” | when gastritis is explicitly associated with H. pylori. | Require explicit diagnostic association. | ||||
| H. pylori Gastritis | “Moderate to marked H. pylori associated active chronic gastritis.” | when severity is described together with H. pylori-associated gastritis. | Preserve the association even with severity modifiers. | ||||
| H. pylori Gastritis | “Helicobacter pylori-associated moderate chronic gastritis with focal activity.” | when the diagnosis links H. pylori to chronic gastritis. | Final diagnosis can support the association. | ||||
| H. pylori Gastritis | “Mild chronic gastritis. No Helicobacter organisms are identified.” | when gastritis is present but organisms are explicitly absent. | Gastritis alone is insufficient. | ||||
| H. pylori Gastritis | “Mild-to-moderate chronic gastritis. No definite Helicobacter pylori identified.” | when chronic gastritis is not attributed to H. pylori. | Do not label non-associated gastritis as H. pylori gastritis. | ||||
| H. pylori Gastritis | “Helicobacter organisms are absent.” | when organism absence prevents attribution of gastritis to H. pylori. | Organism absence rules against H. pylori-associated gastritis. |
3pt
=’ plus .5em You are a clinical NLP assistant specialised in parsing Singapore hospital gastrointestinal pathology reports.
Extract the following two H. pylori features from the report: h_pylori_pos and h_pylori_gastritis. Use only the pathology report text. Answer each feature with Y or N. Return only valid JSON
with the keys: h_pylori_pos, verbatim_h_pylori_pos, h_pylori_gastritis, verbatim_h_pylori_gastritis, and reasoning.
=’ plus .5em Use the same baseline Tier-3 prompt shown in Prompt [prompt:tier395ablated]. In addition, apply the field-specific FIELD_LIBRARY entries for h_pylori_positive and h_pylori_gastritis, including their clinical context,
guardrails, and in-context demonstrations. The output key h_pylori_pos corresponds to the field-library entry h_pylori_positive.