Large language models (LLMs) exhibit remarkable flexibility in adapting to novel tasks from in-context examples without parameter updates, a capability known as in-context learning (ICL). Prior work has sought to understand this capability by characterizing the algorithms underlying ICL or identifying the circuits that support it. Yet what determines whether a particular task can be effectively learned in context remains unresolved. We address this question by using the LLM's pretrained representation space itself to define the learning problem. We construct a controlled family of binary classification tasks in which labels are determined by linear partitions along different directions in this space. Although all tasks are linearly separable by construction, their in-context learnability varies systematically across directions. We find that successful ICL is accompanied by a geometric reorganization of internal representations that increases task-relevant separability. Causal interventions further show that amplifying activity along a fixed labeling axis is insufficient to induce this reorganization. At the behavioral level, LLM responses are best described by a prototype-like learner operating on representations that are themselves reorganized by context. Together, these findings offer a geometric account of ICL in pretrained LLMs, establish pretrained representational geometry as a constraint on ICL, and quantify the gap between what pretrained representations afford and what in-context learning can exploit.
翻译:暂无翻译