In the last years, a number of proofs of the fact that $O_2$ is a multiple context-free grammar (MCFG) were given. Such results can be exploited in the fields of both computational linguistics and of computational algebra. Here, we focus on a recent such proof spelled in terms of factorizations of string tuples, and give a new result with a stronger characterization of such factorizations than in existing theorems.
The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-specific morphosyntactic phenomenon. This article argues instead that the state is a systemic, context-dependent morphosyntactic mechanism that selects grammatical templates across synthetic languages. Within the Template-Based Modular Cognitive framework, taking Riffian as its primary empirical basis, the proposed theory provides a unified explanation for diverse nominal marking patterns traditionally analysed independently and is formalised as a symbolic computational model in which the state is represented by a set-valued function over grammatical templates. A learning algorithm based on finite-set operations acquires and predicts state-dependent grammatical configurations. Beyond nominal morphology, the framework has broader implications for theories of nominal structure and lexical cognition, in particular offering a unified analysis of determiner-noun structure. These results suggest that the state constitutes one instance of a broader class of syntactically conditioned dependencies that also includes agreement and grammatical case.
T. D. Oliveira, L. A. Attilio, M. J. Davila-Fernandezcs.CL
How do large-scale economic transformations shape cultural production? We address this question by combining computational linguistics, econometrics, and formal modelling, using French drama as a well-documented empirical laboratory. Applying latent Dirichlet allocation to a corpus of 1,215 theatrical texts published between 1700 and 1900, we show that aristocratic discourse centred on sovereignty and political authority was gradually displaced by bourgeois and household economic themes as French capitalism developed. Bayesian vector autoregressive models with max-share shock identification suggest a temporal shift in the literary response to economic shocks: bourgeois everyday-life themes reacted to GDP shocks in the eighteenth century, whereas household-economic concerns became responsive only after 1820, amid accelerating industrialisation. A discrete-choice model shows that peer effects among authors and sensitivity to prevailing economic conditions can jointly account for these dynamics. Monte Carlo simulations reproduce the observed historical trajectory with reasonable fidelity. These findings offer a quantitative framework for understanding how economic transformations propagate into cultural production through identifiable social mechanisms, contributing to the study of cultural evolution and the long-run relationship between institutions and literary discourse.
Finite-state transducers (FSTs) are essential for modeling string rewriting in computational linguistics and natural language processing (NLP), particularly for phonological and morphological rewrite rules. Compiling general rewrite rules of the form $A \to B / L \, \_ \, R$, where $A$, $B$, $L$, and $R$ are arbitrary regular languages, is complex due to overlapping matches and context constraints. Traditional methods, such as those by Kaplan and Kay or Karttunen, rely on intricate transducer compositions with auxiliary markers. This paper presents a compact compilation scheme based on the "worsening trick'': generate all legal rewrite candidates, then filter candidates that are worse than another candidate for the same input. Implemented as the built-in rewrite compiler in PyFoma, the construction supports multiple contexts, arbitrary transductions, markup, directed rewriting, weights, and parallel rewriting. The resulting formulas are short and uniform, and where semantics coincide, they reproduce the same rule transducers as earlier approaches while remaining easier to extend. The implementation has been validated against foma on both a substantial collection of rewrite grammars and an automated regression suite covering the major rewrite modalities, with the resulting transducers matching exactly apart from state numbering.