From LLM narratives to parameterized cooperation policies in multi-agent systems
IntroductionLarge language models (LLMs) can generate persuasive narratives that shift agent behavior in multi-agent systems, but deploying raw, unstructured text as an influence mechanism offers no formal guarantees on effectiveness, interpretability, or controllability.MethodsWe introduce the LLM...
IntroductionLarge language models (LLMs) can generate persuasive narratives that shift agent behavior in multi-agent systems, but deploying raw, unstructured text as an influence mechanism offers no formal guarantees on effectiveness, interpretability, or controllability.MethodsWe introduce the LLM Influence Compiler (LIC), a solver–critic pipeline that compiles natural-language cooperation directives into structured, parameterized influence policies defined over a five-field schema: network targeting, narrative theme, intensity, deployment timing, and compiler confidence. Each compiled policy is diffused through a network of numerical agents via an exposure model incorporating fatigue decay, susceptibility heterogeneity, and backlash. Evaluation spans nine controlled experiment blocks comprising more than 200 simulation runs across four topologies (Barabási–Albert, small-world, Erdős–Rényi, modular SBM), four network sizes (n ∈ {80, 160, 320, 640}), five LLM backbones, and two non-stationary perturbations.ResultsCompiled policies raise the mean cooperation rate to 0.826 ± 0.010 (95% bootstrap CI [0.819, 0.834]), an 8.6% relative improvement over the unstructured baseline (0.760 ± 0.005, Mann–Whitney pBonf = 0.040, Cohen's d = 7.90). A controlled decomposition attributes the gain to network-aware targeting (ca. +4.5 pp), dose calibration (+1.9 pp), and an intervention floor (+1.5 pp); the residual contribution of full LLM compilation over a hand-coded rule with the same parameters is statistically indistinguishable from zero under stationary conditions. Under structural non-stationarity (mid-simulation graph rewire), the LLM-mediated re-deployment architecture significantly outperforms a frozen rule (paired t-test p < 0.001, paired Cohen's d = 2.27, 8/8 seeds). Adversarial stress testing confirms that the critic correctly flags 20/20 risky policies as high risk and rejects them, while passing a moderate-baseline policy at medium risk in all 5/5 trials.DiscussionThe compiler abstraction converts an opaque generative process into a decomposable, auditable policy object whose components can be independently attributed, compared, and governed, laying the groundwork for constitutional oversight of LLM-mediated influence in artificial societies.Source: Frontiers AI — Published — Category: Research