Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Constructions

Wesley Scivetti1 Ethan Wilcox1 Nathan Schneider1
Kanishka Misra2 Leonie Weissweiler3
1Georgetown University1, 2The University of Texas at Austin
3Leipzig University & ScaDS.AI Dresden/Leipzig


Abstract

Grasping the semantics of rare constructions (form–meaning pairings) has been shown to be a challenging problem that has currently only been solved by the largest LLMs. It remains an open question if open-source models have robust constructional understanding, and if so, what learning dynamics underlie the acquisition of this knowledge. Focusing on a set of rare constructions in English (e.g. “let alone”, “much less”), we construct a novel dataset to test their meanings using both scalar adjectival semantics and general world knowledge. Testing a wide range of models differing in parameter count, architecture, and pretraining dataset size, we find that several modestly sized models are sensitive to both the forms and the meanings of constructions, though models trained on human-scale data fail at all meaning evaluations. Turning to training dynamics for a set of open-checkpoint models, we find that understanding emerges later in training than syntactic knowledge, and that learning of semantics is correlated with gains in some domains of world knowledge. Overall, our empirical results support the conclusion that modestly sized open-source models can grasp the rare constructions, and demonstrate a connection between knowledge of constructions and other meaning domains.1

1 Introduction↩︎

Language models (LMs) have great potential as tools for studying language. The success of a domain-general statistical learner on many linguistic tasks has prompted some researchers to argue that LMs should be more central in evaluating and developing linguistic theory [1], [2]. In this work, we focus on constructionist approaches, or Construction Grammar ([3][5], inter alia), a family of linguistic theories which posit that language is primarily structured as form-function mappings of gradient complexity. Construction Grammar accounts for not only the most integral parts of linguistic structure (e.g.verb argument structures; [3]), but also rare phenomena traditionally relegated to the “periphery”, arguing that both can be accounted for with similar representations [6]. Construction Grammar’s ability to describe and account for rare linguistic phenomena makes it a valuable framework for studying language models, as mastery of rare and complex constructions is a crucial part of humans’ linguistic knowledge, and may be particularly challenging for language models due to scarcity in their input.

At the same time, the general success of neural language models, which do not explicitly distinguish between syntax and semantic information, has led some researchers to argue that Construction Grammar is more aligned with LM processing of language relative to other frameworks [7][9]. In order to fully evaluate these claims, it is worth investigating the extent to which LMs can serve as “model learners” [1] of constructionist theories of language.

Table 1: construction frequencies in LM pretraining data and COCA. String frequencies are reported for Olmo3 pretraining data, Pythia pretraining data (The Pile) and COCA. All counts are normalized per billion tokens. Prop. refers to the proportion of strings in COCA that were true instances of the construction out of 100 sampled instances. Cxn Freq. refers to the approximate frequency of the construction based on 95% confidence intervals of the sampled proportion in COCA. LM pretraining counts computed using infini-gram [10].
Cxn Olmo3 Pythia COCA Prop. Cxn Freq. COCA
4261 2874 8631 1.00 8311–8631
8671 5521 13184 0.44 4570–7089
3862 2183 8803 0.38 2561–4207
807 796 2974 0.31 677–1209

Because Construction Grammar posits that form and meaning are not separable parts of linguistic knowledge, a constructionist “model system” should learn and understand constructions with respect to both their forms and functions. However, up to this point, there are somewhat mixed results regarding LMs’ syntactic and semantic knowledge of rare constructions. There is ample evidence that even relatively small language models learn formal properties of rare constructions, including Article-Adjective-Numeral-Noun [11], [12], [13], and the Comparative-Correlative [14]. However, in many cases, these smaller models seem unable to grasp the semantics of these rare constructions [13], [14]. On the other hand, extremely large LLMs have been shown to have relatively nuanced semantic understanding of a range of constructions [15], [16]. Since most LLM evaluation has been done in the prompting setting on closed-source LLMs, it remains unclear whether semantic understanding of constructions can be observed in the raw probabilities of (smaller) language models, particularly for exceedingly rare constructions.

Additionally, it is not clear that all current evaluation datasets are well suited for testing constructional semantic understanding. Past work relies on unnatural sounding test items [13], [14] or more complex metalinguistic tasks which may limit LM performance [17]. Additionally, particularly for rare constructions, it is likely that constraints on their behavior are related to more general phenomena in language [11], [18], and it is reasonable to expect that the semantics of constructions are shaped and constrained by their interactions with other domains of semantics.

In this work, we address the above gaps with a new dataset for testing a rare2 family of four constructions: , , , and , collectively known as constructions because they conjoin two phrases that are both in focus:

. He doesn’t like shrimp, let alone squid. [6]

We evaluate syntax and semantics on a much wider range of models than has been tested in past work on constructional semantics. We ground the creation of our dataset in other factors in language production which are likely to inform learning of the constructions’ semantics: scalar adjectives and general world knowledge.

This allows us to ask and answer detailed research questions about the acquisition of semantics.

1) How do training data, parameter count, and pretraining objective impact knowledge of meaning? We show that a range of medium-sized open-source models show sensitivity to the construction-level semantics in their raw probability distributions. However, we do not observe any above-chance knowledge of semantics for models trained on human-scale data (BabyLMs, [19]). 2) When are form and meaning learned throughout pretraining? We show that form is consistently acquired prior to semantics. Furthermore, we show that inferences early in learning are often influenced by typicality (world knowledge) rather than constructional meaning, particularly for weaker models. We find that the different constructions we test are highly correlated with one another in terms of both syntactic and semantic performance. 3) How does meaning correlate with performance on other linguistic benchmarks? We examine the learning dynamics of both the form and meaning of the constructions, in relation to performance on a range of existing linguistic benchmarks. We find that performance on our syntactic evaluations of the constructions plateaus early in pretraining (mirroring learning curves for the BLiMP [20] grammatical benchmark), while functional knowledge is acquired much later and more closely mirrors the learning trajectory of the world-knowledge-based EWoK [21] benchmark.

Overall, our results show that constructions can be learned by relatively small models (under 400M parameters). They also shed light on the interrelatedness of constructions with scalar semantics, and with domains of world knowledge in which such scalar semantics are relevant. Furthermore, the failure of all “human-scale” models on our semantic evaluation points calls into question the ability of human-scale pure text LMs to serve as model systems of linguistic theories that posit joint acquisition of form and meaning.

2 Background↩︎

2.1 Paired-Focus Constructions↩︎

Table 2: semantics dataset overview. Sent. 1 serves as the context and Sent. 2 as the target.
Idx Name/Description Feat. Sent.1 Feat. Sent. 2: ___ than lifting a huge one. Pairwise Feat.
6 Cxn Entails Plaus. +Cxn I couldn’t lift a tiny rock, let alone a huge one. +Plaus. Lifting a tiny rock is easier Entailment
7 Cxn Contradicts Implaus. +Cxn I couldn’t lift a tiny rock, let alone a huge one. \(-\)Plaus. Lifting a tiny rock is harder Contradiction
8 Neutral Plaus.Control \(-\)Cxn I couldn’t lift a tiny rock, or a huge one. +Plaus. Lifting a tiny rock is easier Neutral
9 Neutral Implaus.Control \(-\)Cxn I couldn’t lift a tiny rock, or a huge one. \(-\)Plaus. Lifting a tiny rock is harder Neutral
10 Cxn Contradicts Plaus. +Cxn I couldn’t lift a huge rock, let alone a tiny one. +Plaus. Lifting a tiny rock is easier Contradiction
11 Cxn Entails Implaus. +Cxn I couldn’t lift a huge rock, let alone a tiny one. \(-\)Plaus. Lifting a tiny rock is harder Entailment

In this work, we focus on four related constructions: , , , and . We focus on these constructions because their syntactic and semantic properties are well established in Construction Grammar theory [6], and have also been studied in past computational work (e.g.[12], [13], [17], [22]), though past work on investigating language model understanding of semantics has yielded mostly negative results.

[6] provide a comprehensive treatment of the construction, though much of their analysis can be extended to the other constructions in this work. Syntactically, these constructions function somewhat similarly to coordinating conjunctions (Examples [ex:pf] and [ex:and]), but resist various types of syntactic movement (Example [ex:movement]), and generally behave like negative polarity items (Example [ex:npi]).

. He doesn’t like shrimp, or squid.

. ??It’s shrimp, not to mention squid, that he doesn’t like.

. ??I like shrimp, much less squid.

Regarding semantics, the constructions indicate a relationship between two conjoined and focused elements which are being compared. The comparison between the focused elements evokes a scalar relationship with the two elements representing “two points on a scale” [6]. Generally, the second focused element has a higher value on the scale than the first focused element. Overall, the semantics of constructions are related to scalar semantics more generally (in the invocation of a scale) and to world knowledge, which informs what scalar properties are natural for the focused items in the construction.

2.2 Related Work↩︎

2.2.1 Linguistic Capabilities of LMs↩︎

Broadly, this work follows in a long line of research seeking to investigate LM knowledge of linguistic phenomena by examining LM probabilities for grammatical and ungrammatical sequences ([23], inter alia). Among the most relevant of these works are those which seek to design datasets to assess broad domains of linguistic abilities, whether that be syntactic, conceptual, or world knowledge. We leverage several of these datasets as comparison points for learning dynamics of constructions. Furthermore, investigations of scalar semantics of adjectives [24][27] and adverbs [28] are of particular relevance to this present work, though most past work has focused on scalar implicature and relative intensity of adjectives, which is orthogonal to the present study. Contributing to this line of work are several datasets containing scalar adjectives of varying intensities [29][31]. We rely on the dataset of [30] for the scales and adjectives in our dataset.

2.2.2 LM Understanding of Rare Constructions↩︎

Past work on evaluating LM understanding of constructions has primarily focused on evaluating various constructions that have been outlined in theoretical linguistics, such as argument structure constructions [32][34], the Comparative-Correlative [14], Causal-Excess [35], NPN [36], PiPP [18], AANN [11], [37], [38], and [13]. Other work opts for a more general investigation of a range of constructions [12], [16], [39], or discovers constructions using unsupervised methods [40][42]. While most work on constructional understanding in LMs has focused on English, there is growing work on evaluation of constructions in languages beyond English [43][46].

constructions have been addressed in several past works. [22] and [12] use Global Affinity measures to show that both RoBERTa and a range of BabyLMs have knowledge of collocations for the constructions and . Our results robustly replicate theirs. [17] and [16] include as one of several constructions and find that large, closed-source LLMs can perform a metalinguistic grouping task and natural language inference successfully on examples of these constructions. Our work diverges from these approaches in that we test smaller models using direct, probability-based evaluations.

Our work is most similar in approach to [13], who design probability-based measures for syntactic and semantic properties of . As discussed previously, we argue that their dataset for semantics relies on arbitrary contrasts which do not clearly implicate scalar semantics. Furthermore, they only test a single model architecture (OPT-125M) and data scenario (100M tokens of BabyLM), and only address and not the highly related , , and .

3 Paired-Focus Dataset↩︎

We construct a novel dataset to test the semantics of the constructions. Past work [13] used a fully arbitrary dataset, where there was no obvious scalar relationship between the focused elements (outside of what may be communicated by the construction itself):

. I couldn’t lift the orange crate, let alone the green crate. [13]

Failure on fully arbitrary evaluations could be due to several factors, apart from the models truly not understanding the construction. It is also possible that they may be unable to map the arbitrary relationship between the two focused elements onto any scale, which would also lead to a failure to understand the construction overall. If this is the case, we would expect the models to fail to comprehend arbitrary examples (like Example [ex:arbitrary]) but succeed at understanding examples with a coherent scale. Furthermore, it is possible that the models have only limited knowledge of the construction, where the constructions can only be interpreted in contexts where the scalar relationship between the focused elements is already obvious due to world knowledge. We would then expect the models to succeed in cases where there is a clear scalar relationship that makes sense in the context of world knowledge but fail in cases where there is a clear scalar relationship which conflicts with world knowledge.

To distinguish between these possibilities, we construct a new dataset that contains examples where the focused elements have clear scalar relationships between them. We use adjectives and scales from [30] as a starting point for the focused elements that will be compared within the construction. In all examples, we pair adjectives from their dataset with a set of common nouns that are natural sounding with the chosen scalar adjectives3. In total, we generate 198 starting templates across 4 unique scales from [30], and then fill the resulting templates with a set of appropriate adjectives and nouns, resulting in a dataset of 3.5k example sentence pairs per construction. Example sentences for all scales are shown in Table 4 in Appendix 8.

3.1 Evaluation of Dataset↩︎

To test model sensitivity to semantics, we measure the effect that the constructions have on model output probabilities on a follow-up sentence that would be entailed by the construction. Because constructions imply that the second focused element is a higher value on the scale than the first focused element [6], we include follow-up sentences which either entail or contradict this scalar relationship. Models are expected to prefer the follow-up which reinforces the semantics of the construction to one which contradicts the construction (compare 6 and 7 in Table 2).

Directly comparing the model probabilities of follow up-sentences is confounded by the logical plausibility of the follow-ups independent of the construction. We expect that, irrespective of the presence of a construction, a good model will assign a higher probability to a follow-up that aligns with world knowledge, all other things being equal. To control for this fact, we compare example pairs with a construction to examples without a construction with scalar semantics, specifically the simple conjunction “or” (see 8 and 9, Table 2). Intuitively, if the model is sensitive to semantics, we expect a model to assign higher probability to a follow-up that aligns with the construction, beyond what is assigned to the follow-up due to world knowledge.

Concretely, given a base sentence \(S_{+{\mathbf{Cxn}}}\) and a follow-up sentence \(T_{\pm{\mathbf{Plausible}}}\), we calculate the difference \(\Delta P\) as follows. Here, \(\iota_{\theta} = - \log p_{\theta}\) denotes the surprisal value derived from a language model parameterized by \(\mathbf{\theta}\). \[\begin{align} \Delta P(+\mathbf{Cxn}) &=\iota_{\theta}\!\left(T_{-\mathbf{Plaus.}} \mid S_{+\mathbf{Cxn}}\right) \\ \notag &\quad - \iota_{\theta}\!\left(T_{+\mathbf{Plaus.}} \mid S_{+\mathbf{Cxn}}\right) \end{align}\]

\(\Delta P(-\mathbf{Cxn})\) for the control condition with “or” is computed similarly. Finally, given a dataset with \(k\) example pairs, we then derive an accuracy score by comparing the condition and the control (“or”) condition: \[\label{eq:acc} \mathbf{Acc_{PF}} = \frac{1}{k} \sum_{i=1}^{k} \mathbb{1}[\Delta P(+\mathbf{Cxn}) >\Delta P(-\mathbf{Cxn})]\tag{1}\] There are other confounding factors (beyond the construction and world knowledge) that could influence LM probability associated with the follow-up sentences. Since the two follow-ups are minimal pairs (with only one word different between them) there is a potential confound of lexical bias if the correct sentence has a more frequent/probable word irrespective of context (e.g.since “easier” is more frequent than “harder” overall). Another potential confound is if there is an ordering effect of the two focused elements, where a follow-up sentence is preferred or dispreferred for presenting focused elements in the same order as the sentence. We control for both of these confounds by balancing both the plausible and implausible follow-ups in our dataset by ordering and by the lexical items that appear in the entailed follow-up.

3.2 Syntactic Evaluation↩︎

Beyond evaluating semantics, we also wish to test if models have knowledge of the forms of constructions. We create a syntactic evaluation suite which adapts the syntactic tests from [13] to our dataset. Specifically, we select three syntactic manipulations from [13] which are ungrammatical for constructions but grammatical for simple conjunctions (see 3). Similarly to our semantic tests, we evaluate the \(\Delta P\) values between grammatical and ungrammatical sentences compared to simple conjunctions. For more details on the syntactic evaluation suite, as well as evaluations of Global Affinity [22], see Appendix 10.

Table 3: syntactic tests. Tests adapted from evaluation paradigm of [13]. Accuracy is measured by comparing \(\Delta P\) between +Cxn and \(-\)Cxn manipulated sentences versus Base sentences.
Manipulation Feat. Manipulated Sentence Base
Clause Conjunction +Cxn *I couldn’t lift a tiny rock, let alone I couldn’t lift a huge one. I couldn’t lift a tiny rock, let alone a huge one.
Clause Conjunction \(-\)Cxn I couldn’t lift a tiny rock, and I couldn’t lift a huge one. I couldn’t lift a tiny rock, and a huge one.
NPI +Cxn *I could lift a tiny rock, let alone a huge one. I couldn’t lift a tiny rock, let alone a huge one.
NPI \(-\)Cxn I could lift a tiny rock, and a huge one. I couldn’t lift a tiny rock, and a huge one.
Pseudocleft +Cxn *A tiny rock, let alone a huge one, I couldn’t lift. I couldn’t lift a tiny rock, let alone a huge one.
Pseudocleft \(-\)Cxn A tiny rock, and a huge one, I couldn’t lift. I couldn’t lift a tiny rock, and a huge one.

4 Experiment 1: Impact of Model Size and Training Data↩︎

Figure 1: Semantic Results by Model. Accuracy is averaged across our four constructions: , , , and . We observe a high rank-order correlation between model parameters and average accuracy (Spearman’s \rho =0.67).

4.1 Models and Evaluation↩︎

We test a total of 36 models, varying in parameter count, pretraining data size, pretraining objective (masked vs. autoregressive)4, and model family. See Table 9 in Appendix 12 for model details. In addition to computing accuracy, we run a linear mixed effects model analysis to test the impact of model size, pretraining data, and architecture across all models. The dependent variable is accuracy (see Equation 1 ). We include pretraining data (log # of tokens), parameters (log # of parameters), and architecture (causal vs. MLM) as fixed effects, and a random intercept for model name.

4.2 Results↩︎

Figure 1 visualizes the accuracy5 as a function of model size and training data amount. Below a parameter threshold of roughly 400M, model performance is generally at or below chance regardless of the amount of training data (with the exception of Ettin-encoder-150M). Beyond this parameter threshold, there is a generally positive relationship between training data and accuracy as well as parameter count and accuracy. However, there is substantial variation between individual models, regardless of parameter count, training data, and training objective.

For our linear modeling analysis, we find that only parameter count has a significant main effect (\(\beta= 6.055\), \(p=.011\)). As a follow-up analysis, we additionally fit a series of linear mixed effects models assessing each predictor independently. For single predictor models (again with a random intercept for model name), we find that both parameter count (\(\beta= 6.607\), \(p=003\)) and pretraining data (\(\beta= 2.651\), \(p=.034\)) are significant, though the effect of parameter count remains larger. While we note that MLM and autoregressive scoring is not directly comparable, we find no significant effect of pretraining objective.

Regarding our form evaluations, we find that even very small models have strong performance on our syntactic evaluation suite, and assign high Global Affinity scores to fixed slots in the construction (see Tables 7 and 8 in Appendix 10), echoing past results [12], [13], [22]. The dissociation between tasks for small models underscores the need to evaluate both form and semantics when assessing LMs’ overall knowledge of rare constructions.

In summary, these results indicate that models far smaller than frontier LLMs can grasp constructions, as we find that a few models as small as 400M parameters succeed at learning (with 90%+ accuracy) both the form and semantics of the constructions, and most of the larger models are above chance. There seems to be a broadly positive relationship between parameter count and semantic accuracy, while the effect of training data independent of parameter count is not well supported by our results. Knowledge of the form of constructions is present even in the smallest models we test.

5 Experiment 2: Learning Trajectory Experiments↩︎

In this experiment, we investigate when constructions are learned in training in relation to other linguistic knowledge. We hypothesize that prior to model knowledge of semantics, we will observe knowledge of constructional forms. We further hypothesize that examples which align with appropriate scalar semantics (given general world knowledge) will be acquired first, especially for models that are less proficient at understanding the constructions.

To operationalize semantic and formal knowledge of constructions, we use accuracy for semantics (Equation 1 ) and our syntactic test suite identical to Experiment 1. In order to test how alignment of scalar semantics with world knowledge interacts with learning of constructions, we further test models on sentences where constructional examples entail implausible follow-up sentences and contradict plausible sentences (see 10 and 11 in Table 2). Like in Experiment 1, we consider an example correct if the presence of the construction shifts increases the probability of the follow-up entailed by the construction, relative to a baseline conjunction “or”. Intuitively, if a model has a fully abstract understanding of semantics, we expect the presence of a construction to increase the probability of an otherwise implausible statement that is consistent with the scalar relationship implied by the construction.

Figure 2: Training dynamics of Pythia-12b on evaluations as well as other linguistic benchmarks. Chance performance on EWoK is 25%, while chance performance on all other evaluations is 50%.

5.1 Comparison with Linguistic Benchmarks↩︎

We further evaluate the learning dynamics for three other linguistic benchmarks as comparison points throughout training. We evaluate BLiMP [20], COMPS [47], and EWoK [21]. BLiMP tests model knowledge of a range of syntactic constructions. We expect this dataset to be learned relatively early. COMPS tests knowledge of conceptual properties of noun classes. Knowledge of some properties tested in COMPS may be relevant to linking the noun phrases in our dataset with appropriate scalar semantics, and thus we expect high accuracy on COMPS to precede semantic accuracy. EWoK tests a range of domains of world knowledge, including physical and material properties. Understanding of these properties is crucial to interpreting the scales in our dataset, and thus we expect that high performance on EWoK will correlate with semantic accuracy.

5.2 Models↩︎

We focus on learning dynamics of three models in Experiment 1: Pythia-12b [48], Ettin-encoder-400m, and Ettin-decoder-1b [49]. We select these models because they are the top three performing models at their final checkpoints in Experiment 1. For Pythia-12b, we sample 19 logarithmically spaced checkpoints from start to finish in training. For Ettin models, where logarithmically spaced early checkpoints are not available, we evaluate the first 50 checkpoints6 for each model. For each of these model checkpoints, we evaluate semantic accuracy (for both plausible and world-knowledge implausible examples), syntactic accuracy, and accuracy on each of the comparisons.

Figure 3: Training dynamics of Pythia-12b for each of the four individual constructions.

5.3 Results↩︎

a
b
c

Figure 4: Learning trajectory correlation scatterplots for Pythia-12b. Each point represents, for a given checkpoint, how much the model improved over the previous checkpoint with respect to a pair of criteria. EWoK physical relations and semantic accuracy show moderate correlation.. a — Form vs. plausible accuracy differences., b — Implausible vs. plausible accuracy differences., c — EWoK physical relations vs. plausible semantic accuracy.

For the purpose of clarity, we primarily discuss and visualize results for Pythia-12b in the main body of the paper. Results on Ettin models (for details see Appendix 11) are broadly qualitatively similar, but display noisier and more inconsistent learning trajectories. We also find that the Ettin models display a more substantial difference in constructional effect on plausible compared to implausible follow-ups, especially early in training. In contrast to the Ettin models, which seem more reliant on world knowledge cues for their interpretations of constructions, Pythia-12b displays a sensitivity to the construction which is robust to implausible contexts.

Results for Pythia-12b are visualized in Figure 2. We find strong evidence that form is learned prior to semantics. Syntactic accuracy reaches a peak much earlier in training than accuracy. We find performance on semantic evaluations follows similar trajectories regardless of the plausibility of the follow-up sentence, though plausibility does lead to higher absolute performance at the end of training.

For both BLiMP and COMPS, accuracy plateaus relatively early in training, substantially before the accuracy rises above chance. For EWoK, performance gains are more gradual, similar to gradual gains on semantics. 7

We report results for each construction individually and graph results for Pythia-12b in Figure 3. We observe that performance across the constructions is generally strongly correlated, with all constructions following broadly similar trajectories. Of the four, is learned more linearly across training, while performance curves for the other constructions are more logarithmic and peak later in training. While we cannot establish a causal relationship, we note that is the most frequent and least ambiguous of the four (see Table 1) and thus it is perhaps unsurprising that it is learned earlier in training.

5.4 Correlation Analysis of Learning Trajectories↩︎

We run a first-difference correlation between semantic performance and performance on other benchmarks. We find negligible correlation between form and semantics, and between semantics and BLiMP. We find a moderate correlation between semantic performance on plausible and implausible follow-ups, further showing Pythia-12b has substantial knowledge of scales implied by constructions beyond the scalar relationships evident through world-knowledge alone. When subdividing EWoK into its different world knowledge domains, we find a moderate correlation between semantics and the physical relations domain (\(\rho=.48\)). The first-order correlations between formal accuracy, semantic accuracy, and EWoK physical relations accuracy are shown in Figure 4.

5.5 Discussion↩︎

Overall, these results provide strong evidence that form and meaning are learned with vastly different amounts of training input. We find that performance across individual constructions is strongly correlated, and that overall performance is moderately correlated with performance on relevant world knowledge domains in EWoK. We find no evidence that form and meaning acquisition are correlated in any way, nor is semantics significantly correlated with the more syntactic BLiMP benchmark. Regarding model comparison, we find that the learning trajectories of Pythia-12b are much more stable than the trajectories of the smaller Ettin models, which are more susceptible to spikes and valleys in performance across training, and are less proficient at semantics when follow-up sentences are implausible.

6 Discussion↩︎

In this work, we have shown that medium-sized models (\(\approx\)​400M parameters) can acquire knowledge of semantics. However, small models, and models trained on human-scale data, show a clear gap in performance on form and meaning at their final checkpoints. Furthermore, we have shown that even models which eventually learn meaning only do so much later in training relative to form. The performance gap we observe for constructions echoes past results on Comparative-Correlative [14] and Causal-Excess [35] constructions, which similarly found that LMs fail at semantic evaluations which target those constructions. However, we note that our syntactic tests are inherently different from our semantic tests, and interact with different parts of grammar (e.g., other constructions like conjunction and pseudoclefting for syntactic tests vs. scalar semantics for semantic tests). While our results do not provide definitive evidence that form and meaning are necessarily learned separately by LMs, they underscore that modeling joint learning of constructional form and meaning (as would be ideally possible for a “model system” of constructional acquisition) has not been clearly demonstrated by LMs thus far.

While we do not claim that are results prove that LMs could never jointly acquire form and meaning via text-only language modeling, there are obvious reasons why LMs would be unlikely to model a constructionist account of language development. Humans learn language in social settings in order to perform communicative goals. We are also exposed to rich non-linguistic input in the form of embodied experience. Pure text language models do not have access to these rich sources of input, which are likely particularly relevant for extracting semantic and pragmatic features of language. We look towards future work on multimodal human-scale models [50], [51] and models that incorporate human feedback [52] or learning in situated contexts [41], [53] as potential avenues for future work.

Finally, we highlight the importance of evaluation design. By improving on past evaluations, we find evidence of semantic learning in models that are much smaller than previously reported. Further innovation in evaluation methods for constructional meaning may yet reveal that functional abilities emerge earlier in training, and at smaller parameter scales, than found here.

7 Conclusion↩︎

In this work, we have investigated if LMs can learn nuanced semantic interpretations for a rare family of constructions. We find that models larger than \(\approx\)​400M parameters—both encoders and decoders—are broadly successful at the task, with larger models generally performing better. We find that smaller models completely fail at our novel semantic benchmark despite robust knowledge of the constructions’ forms. This seeming divergence underscores a broader lack of evidence that text-only LMs are jointly modeling form and meaning of constructions. Turning to an analysis of learning dynamics, we show that learning of semantics occurs after learning formal knowledge of the construction, and is correlated with learning of relevant world knowledge. Our largest high performing model is robust to changes in plausibility when interpreting examples, while smaller models are more sensitive to the plausiblity of the construction. Overall, this work shows that relatively modest-sized models can acquire nontrivial semantic knowledge of rare constructions, and finds correlational evidence relating learning of semantics to learning trajectories of other realms of linguistic knowledge.

Limitations↩︎

This paper is limited in terms of its coverage of constructions. While we evaluate a range of constructions in English, it is not clear if our concrete findings would generalize to other constructions which may be less closely related than the set we test here. In addition, while we hope that our work serves as a test case of linking the semantics of a set of rare constructions to other semantic properties of language, our evidence in this paper is not causal: while performance on semantics is correlated with certain realms of world knowledge, it is not clear how learning would be impacted if such knowledge were perturbed or removed altogether from training. Furthermore, while models displayed some robustness to implausible follow-up sentences, it is not clear how plausibility is causally linked to learning of constructions. The number of scales that we used to create our dataset was somewhat small; increasing the number of scales used (and using less common scales) could impact the results here, especially for smaller models, which display some borderline knowledge of semantics. Finally, constructions are not unique to English, but our evaluations were limited to English constructions and primarily monolingual models.

Acknowledgments↩︎

Leonie Weissweiler was supported by a postdoctoral fellowship from the German Research Foundation (DFG, WE 7627/1-1). This research was supported in part by NSF award IIS-2144881. We thank CoNLL reviewers and members of the NERT and PiCOL labs for their insightful comments which improved the final version of this paper.

8 Dataset Examples↩︎

This Appendix contains examples for each frame and construction. See Table 4. All examples in our evaluations are presented with all four constructions: , , , and .

4pt

Table 4: Example sentences for each scale. The number of templates refers to the number of unique combinations of verbs, adjectives, and nouns that were appropriate for that scale. The number of adjectives on the scale is based on [30].
Scale # Num Adjectives # Templates Example
beautiful — ugly 6 24 You couldn’t paint an ugly picture, let alone a gorgeous one.
bright — dim 2 2 They couldn’t see a bright light, let alone a dim one.
good — bad 8 112 We couldn’t cook a bad meal, let alone a great one.
small — large 10 60 You couldn’t pick up a tiny rock, let alone a huge one.

9 Experiment 1 MLM Replication with Pseudo-Log-Likelihood↩︎

Here, we replicate the semantic tests from Experiment 1 for MLMs using Pseudo-log-likelihood (PLL, [54]), using the formulation from [55]. We use the minicons library for computing all PLL scores. Table 5 reports the accuracy on the semantic benchmark as measured by PLL (compare with Table 6, which contains the full results for the semantic evaluations from Experiment 1). We see similar trends in accuracy as we do in our main evaluation for Experiment 1, where we compare probabilities of masked tokens directly. All MLMs with less than 150 million parameters achieve chance or near chance performance in both settings. We find that ettin-400m, ettin-1b, and ModernBERT-large are the top performing models in both settings.

Table 5: Semantic Evaluation Scores for MLMs, as scored using Psuedo-log-likelihood [55].
Architecture Model Avg
BERT base-uncased 44.4 50.1 49.7 50.4 48.7
large-uncased 32.8 45.6 49.2 43.7 42.8
Ettin Encoder 150m 76.9 83.2 68.9 75.5 76.1
1b 77.0 94.4 55.7 79.3 76.6
400m 91.4 91.8 94.5 92.2 92.5
68m 31.4 45.5 27.6 27.0 32.9
ModernBERT base 36.6 34.3 34.0 41.3 36.5
large 87.4 93.6 80.9 83.6 86.4
multiBERTs seed_0 47.5 45.9 42.7 48.4 46.1
seed_1 44.3 48.6 43.1 42.5 44.6
RoBERTa base 50.2 46.3 47.2 51.4 48.8
large 70.6 65.6 58.4 68.6 65.8

10 Experiment 1: Syntactic Evaluations↩︎

We develop a syntactic evaluation suite which is based upon several tests from [13]. The tests focus on three grammatical properties, specifically targeting alternations which are grammatical with simple conjunctions but generally ungrammatical for constructions (see Table 3). Similar to our evaluations on semantics, we measure \(\Delta P\) between two conditions, relative to the difference in those conditions when a simple conjunction is present. Unlike the semantic tests, there are no follow-up sentences; rather, we evaluate the differences in probabilities for sentences with constructions directly. More specifically, given a base sentence with a construction \(S_{-\mathbf{Manip.,+\mathbf{Cxn}}}\), we expect that an ungrammatically manipulated sentence \(S_{+\mathbf{Manip.,\pm\mathbf{Cxn}}}\) will have a relatively high surprisal value. Thus, we compute for a example:

\[\begin{align} \Delta P(+\mathbf{Cxn}) &=\iota_{\theta}\!\left(S_{+\mathbf{Manip.,+\mathbf{Cxn}}} \right) \\ &\quad - \iota_{\theta}\!\left(S_{-\mathbf{Manip.,+\mathbf{Cxn}}}\right) \end{align}\]

We then compare to \(\Delta P(-\mathbf{Cxn})\), which is computed similarly using “or” sentences with manipulations. We calculate accuracy based on \(\Delta P\) values identically to the semantics dataset: \[\label{eq:acc2} \mathbf{Acc_{PF}} = \frac{1}{k} \sum_{i=1}^{k} \mathbb{1}[\Delta P(+\mathbf{Cxn}) >\Delta P(-\mathbf{Cxn})]\tag{2}\]

Results on our syntactic evaluation suite are reported in 7.

Additionally, to test general familiarity with the wordforms associated with each construction, we use the Global Affinity metric presented by [22]. Global Affinity is defined as the probability assigned to a target word when it is masked in a string. Given an original string \(s\) and an index \(i\), Global Affinity defines \(P_{s \setminus\{i\}}\) as the probability distribution at position \(i\). The Global Affinity is then the probability assigned to the correct word \(w_i\) when position \(i\) is masked:8 \[\label{eq:glob95aff} \mathbf{GlobalAff}_{s,w_i} = P_{s \setminus \{i\}}(w_i)\tag{3}\]

In our case, we calculate the Global Affinity for each word in our target constructions (e.g., “much” and “less” for “much less”) and average across all words in the construction. We find that regardless of model size, most models have a very high global affinity for all fixed words in constructions (Table 8). This indicates that models are confident that these constructions are collocations and the fixed words in the constructions are relatively easy to predict for models of all sizes.

Table 6: Full Results on our Semantic Test Suite.
Model Type Family Model Name Avg.
MLM BERT [56] BERT-base-uncased 48.9 49.8 50.1 49.8 49.6
BERT-large-uncased 41.0 59.8 49.7 47.6 49.5
ModernBERT [57] ModernBERT-Base 34.1 32.6 31.3 37.7 33.9
ModernBERT-Large 84.4 92.7 82.3 79.6 84.8
MultiBERT [58] MultiBERT seed 0 46.0 49.4 43.4 49.5 47.1
MultiBERT seed 1 48.2 48.6 44.5 41.5 45.7
RoBERTa [59] RoBERTa-base 51.5 49.1 56.5 53.3 52.6
RoBERTa-large 71.7 68.8 56.8 70.7 66.9
Ettin [49] Ettin-Enc-68M 29.3 45.7 26.5 26.6 32.0
Ettin-Enc-150M 81.3 84.2 69.3 77.3 78.0
Ettin-Enc-400M 92.0 91.8 95.9 93.4 93.3
Ettin-Enc-1B 65.6 94.4 57.4 70.1 71.9
CausalLM Ettin [49] Ettin-Dec-68M 44.9 41.5 44.4 46.9 44.4
Ettin-Dec-150M 52.3 50.9 46. 49.5 49.7
Ettin-Dec-400M 88.4 83.1 84.1 82.2 84.4
Ettin-Dec-1B 92.3 95.4 87.9 93.1 92.2
GPT2 [60] GPT2 65.4 75.3 50.1 56.0 61.7
GPT2-medium 66.3 58.2 48.6 49.9 55.7
GPT2-large 74.0 54.9 37.9 50.5 54.3
GPT2-xl 74.1 65.2 57.3 63.1 64.9
OLMo [61] [62] Olmo2-7b 64.1 69.0 57.8 51.4 60.6
Olmo2-13b 73.0 70.8 60.1 51.7 63.9
Olmo3-7b 79.9 87.4 69.2 84.0 80.1
OPT [63] OPT-125M 50.8 56.6 49.0 47.2 50.9
OPT-350M 50.6 49.7 33.2 35.5 42.2
OPT-1.3b 69.6 66.7 55.5 66.7 64.6
OPT-2.7b 72.8 69.5 70.5 69.4 70.5
OPT-6.7b 65.9 63.9 54.7 67.9 63.1
Pythia [48] Pythia-70m 52.7 47.9 45.5 58.5 51.1
Pythia-160m 36.2 37.2 34.9 36.3 36.2
Pythia-410m 62.9 65.0 41.4 38.7 52.0
Pythia-1b 80.6 72.6 68.3 63.6 71.3
Pythia-1.4b 61.3 63.3 52.5 50.5 56.9
Pythia-2.8b 94.4 79.1 83.9 89.3 86.7
Pythia-6.9b 89.0 89.6 67.4 83.2 82.3
Pythia-12b 98.9 98.1 92.7 96.3 96.5
Table 7: Full Results on our Syntactic Test Suite. Results for each construction are averaged across the 3 manipulation types (see Table 3). OLMo2-13b was not run for syntactic tests due to compute constraints.
Model Type Family Model Name Avg.
MLM BERT [56] BERT-base-uncased 89.8 96.1 93.6 60.2 84.9
BERT-large-uncased 59.4 56.9 88.2 73.1 69.4
ModernBERT [57] ModernBERT-Base 97.9 97.8 99.9 99.6 98.8
ModernBERT-Large 83.0 83.3 94.0 87.1 86.9
MultiBERT [58] MultiBERT seed 0 88.0 89.5 84.8 83.9 86.6
MultiBERT seed 1 89.7 90.8 76.3 67.9 81.1
RoBERTa [59] RoBERTa-base 93.3 93.6 99.4 93.3 94.9
RoBERTa-large 82.8 83.8 94.6 89.6 87.7
Ettin [49] Ettin-Enc-68M 88.3 93.2 93.5 96.7 92.9
Ettin-Enc-150M 93.2 92.4 98.7 94.9 94.8
Ettin-Enc-400M 82.5 87.0 99.3 91.4 90.0
Ettin-Enc-1B 89.2 90.6 98.3 89.4 91.9
CausalLM Ettin [49] Ettin-Dec-68M 100 100 89.4 95.8 96.3
Ettin-Dec-150M 99.9 99.9 90.7 97.8 97.1
Ettin-Dec-400M 96.5 98.9 98.4 99.5 98.4
Ettin-Dec-1B 86.9 93.5 89.9 94.0 91.1
GPT2 [60] GPT2 100 100 97.2 100 99.3
GPT2-medium 98.2 98.6 94.4 98.4 97.4
GPT2-large 99.8 99.5 97.4 99.7 99.1
GPT2-xl 99.0 97.7 89.3 98.6 96.2
OLMo [61] [62] Olmo2-7b 97.0 96.5 94.9 95.4 96.0
Olmo3-7b 91.5 99.3 94.0 93.7 94.6
OPT [63] OPT-125M 100 100 94.2 99.6 98.4
OPT-350M 99.9 100 92.7 93.8 96.6
OPT-1.3b 97.5 98.6 94.7 96.5 96.8
OPT-2.7b 96.0 98.2 96.1 94.0 96.1
OPT-6.7b 98.3 98.9 92.8 97.9 97.0
Pythia [48] Pythia-68m 100 100 83.3 99.8 95.7
Pythia-160m 100 100 94.0 98.1 98.0
Pythia-410m 99.9 100 97.9 99.7 99.4
Pythia-1b 99.8 99.7 91.4 99.0 97.5
Pythia-1.4b 99.9 100 87.3 99.8 96.7
Pythia-2.8b 94.8 98.7 93.2 99.1 96.4
Pythia-6.9b 99.4 99.2 96.2 99.6 98.6
Pythia-12b 97.3 97.6 92.3 97.9 96.3
Figure 5: Training dynamics of Ettin-Enc-400M.
Figure 6: Training dynamics of Ettin-Decoder 1b.
Table 8: Global Affinity Scores By Construction. Intervals represent \(\pm\) 95% confidence intervals. Results are only reported for MLMs as [22] do not provide a definition of Global Affinity for CausalLMs.
Architecture Model
BERT base-uncased \(0.999 \pm 0.000\) \(0.987 \pm 0.000\) \(0.502 \pm 0.008\) \(0.933 \pm 0.003\)
large-uncased \(0.999 \pm 0.000\) \(0.999 \pm 0.000\) \(0.835 \pm 0.005\) \(0.999 \pm 0.000\)
Ettin Encoder 150m \(0.995 \pm 0.000\) \(0.985 \pm 0.000\) \(0.958 \pm 0.001\) \(0.960 \pm 0.001\)
1b \(0.995 \pm 0.000\) \(0.983 \pm 0.000\) \(0.924 \pm 0.001\) \(0.951 \pm 0.001\)
400m \(0.990 \pm 0.000\) \(0.980 \pm 0.000\) \(0.932 \pm 0.001\) \(0.925 \pm 0.001\)
68m \(0.998 \pm 0.000\) \(0.985 \pm 0.000\) \(0.779 \pm 0.003\) \(0.797 \pm 0.005\)
ModernBERT base \(0.987 \pm 0.000\) \(0.992 \pm 0.000\) \(0.903 \pm 0.002\) \(0.791 \pm 0.004\)
large \(0.992 \pm 0.000\) \(0.985 \pm 0.000\) \(0.926 \pm 0.002\) \(0.923 \pm 0.002\)
multiBERTs seed_0 \(0.999 \pm 0.000\) \(0.999 \pm 0.000\) \(0.418 \pm 0.007\) \(0.741 \pm 0.006\)
seed_1 \(0.999 \pm 0.000\) \(0.995 \pm 0.000\) \(0.439 \pm 0.007\) \(0.609 \pm 0.007\)
RoBERTa base \(0.998 \pm 0.000\) \(0.998 \pm 0.000\) \(0.944 \pm 0.001\) \(0.998 \pm 0.000\)
large \(0.996 \pm 0.000\) \(0.984 \pm 0.000\) \(0.963 \pm 0.001\) \(0.991 \pm 0.000\)

11 Experiment 2 Full Results↩︎

In this section, we provide extended results on learning dynamics for Ettin-encoder400m and Ettin-decoder1b. Results are shown in Figures 5 and 6 respectively. In general, the performance trajectories for semantics are substantially noisier in these models than in Pythia, with large peaks and valleys throughout training. For both Ettin models, we observe early spikes in performance on plausible examples that are not accompanied by spikes on implausible examples. Generally, performance on implausible examples does not rise consistently above chance until much later in training. Taken together, these results seem to indicate that the smaller Ettin models do learn nontrivial semantics, but may be limited to more natural contexts where the scalar relationship entailed by the construction is further supported by world knowledge. Performance on other benchmarks is generally more stable than on semantics, though performance on EWoK is noticeably less stable relative to Pythia.

12 Model Details↩︎

Table 9 presents details about the models that we test in Experiment 1. In total, we test 36 models in Experiment 1. In Experiment 2, we test 3 models: Pythia-12b, Ettin-encoder400m, and Ettin-decoder1b, which are selected due to their strong performance in Experiment 1 and the availability of their intermediate checkpoints.

Table 9: Information on all models tested in Experiment 1. Parameter Count and # of Pretraining Tokens are approximate measures.
Model Type Family Model Name Parameter Count Pretraining Data (# of Tokens)
MLM BERT [56] BERT-base-uncased 110M 3B
BERT-large-uncased 335M 3B
ModernBERT [57] ModernBERT-Base 110M 2T
ModernBERT-Large 375M 2T
MultiBERT [58] MultiBERT seed 0 110M 3B
MultiBERT seed 1 110M 3B
RoBERTa [59] RoBERTa-base 125M 160B
RoBERTa-large 355M 160B
Ettin [49] Ettin-Enc-68M 68M 2T
Ettin-Enc-150M 150M 2T
Ettin-Enc-400M 400M 2T
Ettin-Enc-1B 1B 2T
CausalLM Ettin [49] Ettin-Dec-68M 68M 2T
Ettin-Dec-150M 150M 2T
Ettin-Dec-400M 400M 2T
Ettin-Dec-1B 1B 2T
GPT2 [60] GPT2 125M 160B
GPT2-medium 355M 160B
GPT2-large 775M 160B
GPT2-xl 1.5B 160B
OLMo [61] [62] Olmo2-7b 7B 4T
Olmo2-13b 13B 5T
Olmo3-7b 7B 6T
OPT [63] OPT-125M 125M 180B
OPT-350M 350M 180B
OPT-1.3b 1.3B 180B
OPT-2.7b 2.7B 180B
OPT-6.7b 6.7B 180B
Pythia [48] Pythia-70m 70M 260B
Pythia-160m 160M 260B
Pythia-410m 410M 260B
Pythia-1b 1B 260B
Pythia-1.4b 1.4B 260B
Pythia-2.8b 2.8B 260B
Pythia-6.9b 6.9B 260B
Pythia-12b 12B 260B

References↩︎

[1]
A. Warstadt and S. R. Bowman, What artificial neural networks can tell us about human language Acquisition,” in Algebraic structures in natural language, CRC Press, 2022, pp. 17–60.
[2]
R. Futrell and K. Mahowald, “How linguistics learned to stop worrying and love the language models,” Behavioral and Brain Sciences, pp. 1–98, Jul. 2025, doi: 10.1017/S0140525X2510112X.
[3]
A. E. Goldberg, Constructions: A Construction Grammar Approach to Argument Structure. University of Chicago Press, 1995.
[4]
A. E. Goldberg, Constructions at Work: The Nature of Generalization in Language. Oxford University Press, 2006.
[5]
W. Croft, Radical Construction Grammar: Syntactic Theory in Typological Perspective. Oxford University Press, 2001.
[6]
C. J. Fillmore, P. Kay, and M. C. O’Connor, Publisher: Linguistic Society of America“Regularity and idiomaticity in grammatical constructions: The case of let alone,” Language, vol. 64, no. 3, pp. 501–538, 1988, doi: 10.2307/414531.
[7]
L. Weissweiler, T. He, N. Otani, D. R. Mortensen, L. Levin, and H. Schütze, “Construction grammar provides unique insight into neural language models,” in Proceedings of the first international workshop on construction grammars and NLP (CxGs+NLP, GURT/SyntaxFest 2023), Mar. 2023, pp. 85–95, [Online]. Available: https://aclanthology.org/2023.cxgsnlp-1.10/.
[8]
A. E. Goldberg, “Usage-based constructionist approaches and large language models,” Constructions and Frames, vol. 16, no. 2, pp. 220–254, Oct. 2024, doi: 10.1075/cf.23017.gol.
[9]
S. Piantadosi, “Modern language models refute chomsky’s approach to language,” in From fieldwork to linguistic theory: A tribute to dan everett, Berlin, Germany: Language Sciences Press, 2024, pp. 353–414.
[10]
J. Liu, S. Min, L. Zettlemoyer, Y. Choi, and H. Hajishirzi, “Infini-gram: Scaling unbounded n-gram language models to a trillion tokens,” in Proceedings of the first conference on language modeling, Aug. 2024, [Online]. Available: https://openreview.net/forum?id=u2vAyMeLMm.
[11]
K. Misra and K. Mahowald, “Language models learn rare phenomena from less rare phenomena: The case of the missing AANNs,” in Proceedings of the 2024 conference on empirical methods in natural language processing, Nov. 2024, pp. 913–929, doi: 10.18653/v1/2024.emnlp-main.53.
[12]
J. Rozner, L. Weissweiler, and C. Shain, BabyLMs first constructions: Causal interventions provide a signal of learning,” in Proceedings of the 2025 conference on empirical methods in natural language processing, Nov. 2025, pp. 2237–2249, doi: 10.18653/v1/2025.emnlp-main.113.
[13]
W. Scivetti, T. Aoyama, E. Wilcox, and N. Schneider, “Unpacking let alone: Human-scale models generalize to a rare construction in form but not meaning,” in Proceedings of the 2025 conference on empirical methods in natural language processing, Nov. 2025, pp. 27503–27514, doi: 10.18653/v1/2025.emnlp-main.1399.
[14]
L. Weissweiler, V. Hofmann, A. Köksal, and H. Schütze, “The better your syntax, the better your semantics? Probing pretrained language models for the English Comparative Correlative,” in Proceedings of the 2022 conference on empirical methods in natural language processing, Dec. 2022, pp. 10859–10882, doi: 10.18653/v1/2022.emnlp-main.746.
[15]
D. R. Mortensen, V. Izrailevitch, Y. Xiao, H. Schütze, and L. Weissweiler, “Verbing weirds language (models): Evaluation of English zero-derivation in five LLMs,” in Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING 2024), May 2024, pp. 17359–17364, [Online]. Available: https://aclanthology.org/2024.lrec-main.1508/.
[16]
W. Scivetti et al., “Beyond memorization: Assessing semantic generalization in large language models using phrasal constructions,” in Proceedings of the 14th international joint conference on natural language processing and the 4th conference of the asia-pacific chapter of the association for computational linguistics, Dec. 2025, pp. 1184–1201, [Online]. Available: https://aclanthology.org/2025.ijcnlp-long.65/.
[17]
C. Bonial and H. Tayyar Madabushi, “A construction grammar corpus of varying schematicity: A dataset for the evaluation of abstractions in language models,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), May 2024, pp. 243–255, Accessed: Mar. 08, 2025. [Online]. Available: https://aclanthology.org/2024.lrec-main.22/.
[18]
C. Potts, Characterizing English preposing in PP constructions,” Journal of Linguistics, pp. 1–39, 2024, doi: 10.1017/S0022226724000227.
[19]
A. Warstadt et al., “Findings of the BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora,” in Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning, Dec. 2023, pp. 1–34, doi: 10.18653/v1/2023.conll-babylm.1.
[20]
A. Warstadt et al., BLiMP: The benchmark of linguistic minimal pairs for English,” Transactions of the Association for Computational Linguistics, vol. 8, pp. 377–392, 2020, doi: 10.1162/tacl_a_00321.
[21]
A. A. Ivanova et al., “Elements of world knowledge ( EWoK ): A cognition-inspired framework for evaluating basic world knowledge in language models,” Transactions of the Association for Computational Linguistics, vol. 13, pp. 1245–1270, 2025, doi: 10.1162/tacl.a.38.
[22]
J. Rozner, L. Weissweiler, K. Mahowald, and C. Shain, “Constructions are revealed in word distributions,” in Proceedings of the 2025 conference on empirical methods in natural language processing, Nov. 2025, pp. 2116–2138, doi: 10.18653/v1/2025.emnlp-main.108.
[23]
J. Hu, J. Gauthier, P. Qian, E. Wilcox, and R. Levy, “A systematic assessment of syntactic generalization in neural language models,” in Proceedings of the 58th annual meeting of the association for computational linguistics, Jul. 2020, pp. 1725–1744, doi: 10.18653/v1/2020.acl-main.158.
[24]
A. Garí Soler and M. Apidianaki, BERT knows Punta Cana is not just beautiful, its gorgeous: Ranking scalar adjectives with contextualised representations,” in Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), Nov. 2020, pp. 7371–7385, doi: 10.18653/v1/2020.emnlp-main.598.
[25]
S. Schuster, Y. Chen, and J. Degen, “Harnessing the linguistic signal to predict scalar inferences,” in Proceedings of the 58th annual meeting of the association for computational linguistics, Jul. 2020, pp. 5387–5403, doi: 10.18653/v1/2020.acl-main.479.
[26]
F. Lin, D. Altshuler, and J. B. Pierrehumbert, “Probing large language models for scalar adjective lexical semantics and scalar diversity pragmatics,” in Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING 2024), May 2024, pp. 13033–13049, [Online]. Available: https://aclanthology.org/2024.lrec-main.1141/.
[27]
R. Nizamani, S. Schuster, and V. Demberg, SIGA: A naturalistic NLI dataset of English scalar implicatures with gradable adjectives,” in Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING 2024), May 2024, pp. 14784–14795, [Online]. Available: https://aclanthology.org/2024.lrec-main.1288/.
[28]
I. Lorge and J. B. Pierrehumbert, “Not wacky vs. Definitely wacky: A study of scalar adverbs in pretrained language models,” in Proceedings of the 6th BlackboxNLP workshop: Analyzing and interpreting neural networks for NLP, Dec. 2023, pp. 296–316, doi: 10.18653/v1/2023.blackboxnlp-1.23.
[29]
G. de Melo and M. Bansal, “Good, great, excellent: Global inference of semantic intensities,” Transactions of the Association for Computational Linguistics, vol. 1, pp. 279–290, Jul. 2013, doi: 10.1162/tacl_a_00227.
[30]
B. Wilkinson and O. Tim, “A gold standard for scalar adjectives,” in Proceedings of the tenth international conference on language resources and evaluation (LREC’16), May 2016, pp. 2669–2675, [Online]. Available: https://aclanthology.org/L16-1424/.
[31]
A. Cocos, S. Wharton, E. Pavlick, M. Apidianaki, and C. Callison-Burch, “Learning scalar adjective intensity from paraphrases,” in Proceedings of the 2018 conference on empirical methods in natural language processing, 2018, pp. 1752–1762, doi: 10.18653/v1/D18-1202.
[32]
B. Li, Z. Zhu, G. Thomas, F. Rudzicz, and Y. Xu, “Neural reality of argument structure constructions,” in Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: Long papers), May 2022, pp. 7410–7423, doi: 10.18653/v1/2022.acl-long.512.
[33]
T. Veenboer and J. Bloem, “Using collostructional analysis to evaluate BERT’s representation of linguistic constructions,” in Findings of the association for computational linguistics: ACL 2023, Jul. 2023, pp. 12937–12951, doi: 10.18653/v1/2023.findings-acl.819.
[34]
H. Sung and K. Kyle, “Leveraging pre-trained language models for linguistic analysis: A case of argument structure constructions,” in Proceedings of the 2024 conference on empirical methods in natural language processing, Nov. 2024, pp. 7302–7314, doi: 10.18653/v1/2024.emnlp-main.415.
[35]
S. Zhou, L. Weissweiler, T. He, H. Schütze, D. R. Mortensen, and L. Levin, “Constructions are so difficult that Even large language models get them right for the wrong reasons,” in Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING 2024), May 2024, pp. 3804–3811, [Online]. Available: https://aclanthology.org/2024.lrec-main.336/.
[36]
W. Scivetti and N. Schneider, Construction identification and disambiguation using BERT: A case study of NPN,” in Proceedings of the 29th conference on computational natural language learning, Jul. 2025, pp. 365–376, doi: 10.18653/v1/2025.conll-1.24.
[37]
K. Mahowald, “A discerning several thousand judgments: GPT-3 rates the article + adjective + numeral + noun construction,” in Proceedings of the 17th conference of the european chapter of the association for computational linguistics, May 2023, pp. 265–273, doi: 10.18653/v1/2023.eacl-main.20.
[38]
G. Chronis, K. Mahowald, and K. Erk, “A method for studying semantic construal in grammatical constructions with interpretable contextual embedding spaces,” in Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers), Jul. 2023, pp. 242–261, doi: 10.18653/v1/2023.acl-long.14.
[39]
H. Tayyar Madabushi, L. Romain, D. Divjak, and P. Milin, “CxGBERT: BERT meets Construction Grammar,” in Proceedings of the 28th international conference on computational linguistics, Dec. 2020, pp. 4020–4032, doi: 10.18653/v1/2020.coling-main.355.
[40]
J. Dunn, “Computational learning of construction grammars,” Language and Cognition, vol. 9, no. 2, pp. 254–292, Jun. 2017, doi: 10.1017/langcog.2016.7.
[41]
K. Beuls and P. Van Eecke, “Humans learn language from situated communicative interactions. What about machines?” Computational Linguistics, vol. 50, no. 4, pp. 1277–1311, Dec. 2024, doi: 10.1162/coli_a_00534.
[42]
L. Verheyen, J. Doumen, P. Van Eecke, and K. Beuls, “You shall know a construction by the company it keeps: Computational construction grammar with embeddings,” in Proceedings of the second international workshop on construction grammars and NLP, Sep. 2025, pp. 75–83, [Online]. Available: https://aclanthology.org/2025.cxgsnlp-1.8/.
[43]
Y.-H. Tseng, C.-F. Shih, P.-E. Chen, H.-Y. Chou, M.-C. Ku, and S.-K. Hsieh, “CxLM: A construction and context-aware language model,” in Proceedings of the thirteenth language resources and evaluation conference, Jun. 2022, pp. 6361–6369, [Online]. Available: https://aclanthology.org/2022.lrec-1.683.
[44]
B. Bunzeck, D. Duran, and S. Zarrieß, Do construction distributions shape formal language learning in german BabyLMs?” in Proceedings of the 29th conference on computational natural language learning, Jul. 2025, pp. 169–186, doi: 10.18653/v1/2025.conll-1.12.
[45]
X. Huang, Y. Pan, S. Hartmann, and Y. Yanning, “Assessing minimal pairs of Chinese verb-resultative complement constructions: Insights from language models,” in Proceedings of the second international workshop on construction grammars and NLP, Sep. 2025, pp. 144–150, [Online]. Available: https://aclanthology.org/2025.cxgsnlp-1.14/.
[46]
X. Yang, “Language models at the syntax-semantics interface: A case study of the long-distance binding of Chinese reflexive ziji,” in Proceedings of the 31st international conference on computational linguistics, Jan. 2025, pp. 3808–3824, [Online]. Available: https://aclanthology.org/2025.coling-main.257/.
[47]
K. Misra, J. Rayz, and A. Ettinger, COMPS: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models,” in Proceedings of the 17th conference of the european chapter of the association for computational linguistics, May 2023, pp. 2928–2949, doi: 10.18653/v1/2023.eacl-main.213.
[48]
S. Biderman et al., “Pythia: A suite for analyzing large language models across training and scaling,” in Proceedings of the 40th international conference on machine learning, 2023, vol. 202, pp. 2397–2430, [Online]. Available: https://proceedings.mlr.press/v202/biderman23a.html.
[49]
O. Weller, K. Ricci, M. Marone, A. Chaffin, D. Lawrie, and B. V. Durme, “Seq vs seq: An open suite of paired encoders and decoders.” 2025, [Online]. Available: https://arxiv.org/abs/2507.11412.
[50]
M. Y. Hu et al., “Findings of the second BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora,” in The 2nd BabyLM challenge at the 28th conference on computational natural language learning, Nov. 2024, pp. 1–21, [Online]. Available: https://aclanthology.org/2024.conll-babylm.1/.
[51]
S. Wang, A. Chandra, A. Liu, V. Saligrama, and B. Gong, “BabyVLM: Data-efficient pretraining of VLMs inspired by infant learning,” in Proceedings of the IEEE/CVF international conference on computer vision, 2025, pp. 1380–1390, [Online]. Available: https://openaccess.thecvf.com/content/ICCV2025/html/Wang_BabyVLM_Data-Efficient_Pretraining_of_VLMs_Inspired_by_Infant_Learning_ICCV_2025_paper.html.
[52]
D. M. Ziegler et al., arXiv:1909.08593 [cs]“Fine-tuning language models from human preferences.” arXiv, Jan. 2020, doi: 10.48550/arXiv.1909.08593.
[53]
J. Botoko Ekila, L. Verheyen, K. Beuls, and P. Van Eecke, “Constructions all the way up: From sensory experiences to construction grammars,” in Proceedings of the second international workshop on construction grammars and NLP, Sep. 2025, pp. 84–95, [Online]. Available: https://aclanthology.org/2025.cxgsnlp-1.9/.
[54]
J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff, “Masked language model scoring,” in Proceedings of the 58th annual meeting of the association for computational linguistics, Jul. 2020, pp. 2699–2712, doi: 10.18653/v1/2020.acl-main.240.
[55]
C. Kauf and A. A. Ivanova, “A better way to do masked language model scoring,” in Proceedings of the 61st annual meeting of the association for computational linguistics (volume 2: Short papers), Jul. 2023, pp. 925–935, doi: 10.18653/v1/2023.acl-short.80.
[56]
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: Human language technologies, volume 1 (long and short papers), Jun. 2019, pp. 4171–4186, doi: 10.18653/v1/N19-1423.
[57]
B. Warner et al., “Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference,” in Proceedings of the 63rd annual meeting of the association for computational linguistics (volume 1: Long papers), Jul. 2025, pp. 2526–2547, doi: 10.18653/v1/2025.acl-long.127.
[58]
T. Sellam et al., “The multiBERTs: BERT reproductions for robustness analysis,” in International conference on learning representations, 2022, [Online]. Available: https://openreview.net/forum?id=K0E_F0gFDgA.
[59]
Y. Liu et al., “RoBERTa: A robustly optimized BERT pretraining approach.” 2019, [Online]. Available: https://arxiv.org/abs/1907.11692.
[60]
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI, 2019. [Online]. Available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
[61]
T. OLMo et al., “2 OLMo 2 furious.” 2025, [Online]. Available: https://arxiv.org/abs/2501.00656.
[62]
T. Olmo et al., arXiv:2512.13961 [cs]“Olmo 3.” arXiv, Apr. 2026, doi: 10.48550/arXiv.2512.13961.
[63]
S. Zhang et al., arXiv:2205.01068 [cs]OPT: Open Pre-trained Transformer Language Models.” arXiv, Jun. 2022, doi: 10.48550/arXiv.2205.01068.

  1. https://github.com/WesScivetti/Meaning_Alone↩︎

  2. For approximate corpus frequencies see Table 1.↩︎

  3. constructions can focus other syntactic phrases besides noun phrases. To control for the confound of how syntactic structure could influence understanding, we only place noun phrases in the focused slot of the construction.↩︎

  4. For MLM models, we evaluate the probability at the target region (e.g.”easier” vs. “harder”), whereas for the decoder models, we evaluate the probability of the entire follow-up sentence conditioned on the first sentence. See Appendix 9 for a replication with MLM scoring using Psuedo-log-likelihood.↩︎

  5. We average accuracy scores across each of our four constructions.↩︎

  6. Each checkpoint is equivalent to roughly 8–9 billion tokens of pretraining data, and 50 checkpoints is thus equivalent to approx. billion pretraining tokens.↩︎

  7. Chance accuracy for EWoK is 25%.↩︎

  8. Global Affinity is only defined for Masked Language Models [22]. Thus, we only evaluate a subset of our models on this metric.↩︎