Prompt engineering is crucial for deploying LLMs but is poorly understood mathematically. We formalize LLM systems as a class of discrete stochastic dynamical systems to explore prompt engineering through the lens of control theory. We investigate the reachable set of output token sequences $R_y(\mathbf x_0)$ for which there exists a control input sequence $\mathbf u$ for each $\mathbf y \in R_y(\mathbf x_0)$ that steers the LLM to output $\mathbf y$ from initial state sequence $\mathbf x_0$. We offer analytic analysis on the limitations on the controllability of self-attention in terms of reachable set, where we prove an upper bound on the reachable set of outputs $R_y(\mathbf x_0)$ as a function of the singular values of the parameter matrices. We present complementary empirical analysis on the controllability of a panel of LLMs, including Falcon-7b, Llama-7b, and Falcon-40b. Our results demonstrate a lower bound on the reachable set of outputs $R_y(\mathbf x_0)$ w.r.t. initial state sequences $\mathbf x_0$ sampled from the Wikitext dataset. We find that the correct next Wikitext token following sequence $\mathbf x_0$ is reachable over 97% of the time with prompts of $k\leq 10$ tokens. We also establish that the top 75 most likely next tokens, as estimated by the LLM itself, are reachable at least 85% of the time with prompts of $k\leq 10$ tokens. Intriguingly, short prompt sequences can dramatically alter the likelihood of specific outputs, even making the least likely tokens become the most likely ones. This control-centric analysis of LLMs demonstrates the significant and poorly understood role of input sequences in steering output probabilities, offering a foundational perspective for enhancing language model system capabilities.
翻译:摘要:提示工程对于部署大语言模型至关重要,但在数学上仍缺乏清晰理解。我们将大语言模型系统形式化为一类离散随机动力系统,从而通过控制理论的视角探索提示工程。我们研究了输出令牌序列的可达集 $R_y(\mathbf x_0)$,使得对于每个 $\mathbf y \in R_y(\mathbf x_0)$ 存在控制输入序列 $\mathbf u$,能够将大语言模型从初始状态序列 $\mathbf x_0$ 引导至输出 $\mathbf y$。我们针对自注意力机制在可达集意义上的可控性限制提供了解析分析,证明了输出可达集 $R_y(\mathbf x_0)$ 的上界是参数矩阵奇异值的函数。我们还对一组大语言模型(包括 Falcon-7b、Llama-7b 和 Falcon-40b)的可控性进行了补充性实证分析。结果表明,对于从 Wikitext 数据集采样的初始状态序列 $\mathbf x_0$,输出可达集 $R_y(\mathbf x_0)$ 存在下界。我们发现,使用不超过 10 个令牌的提示,序列 $\mathbf x_0$ 后正确的下一个 Wikitext 令牌在超过 97% 的情况下可达。我们还确认,大语言模型自身估计的最可能的 75 个下一个令牌中,使用不超过 10 个令牌的提示时,至少 85% 的情况下可达。引人注目的是,短提示序列能戏剧性地改变特定输出的可能性,甚至使最不可能的令牌变为最可能的。这一基于控制视角的大语言模型分析展示了输入序列在引导输出概率中显著但尚未被充分理解的作用,为增强语言模型系统能力提供了基础性视角。