Large language model (LLM) agents rely on reusable skills to solve complex tasks. However, existing skill creation approaches treat skills as isolated and static artifacts, limiting their reusability, reliability, and long-term improvement. We propose MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution), a skill-centric agent framework that lets agents continuously improve their task-solving capability by creating, reusing, and refining skills under a unified lifecycle (creation, memory, management, evaluation, and refinement). Our framework enables agents to create skills on demand, store and reuse them across tasks, organize and select them efficiently, and evaluate them through unit tests and runtime feedback for continuous refinement. We further introduce skill-level memory that accumulates experience for each skill across tasks, enabling more effective reuse and adaptation over time. Experiments on SkillsBench provide initial evidence that lifecycle-managed skills can improve task success, efficiency, reuse, and cross-agent transfer, highlighting the importance of treating skills as long-lived, experience-aware, and testable assets.
翻译:大型语言模型(LLM)智能体依赖可复用技能解决复杂任务。然而,现有技能创建方法将技能视为孤立且静态的产物,限制了其可复用性、可靠性及长期改进能力。我们提出MUSE-Autoskill智能体(基于记忆利用的技能进化),这是一种以技能为中心的智能体框架,通过统一的技能生命周期(创建、记忆、管理、评估与改进),使智能体能够持续提升任务求解能力。该框架支持智能体按需创建技能、跨任务存储与复用技能、高效组织与选择技能,并通过单元测试与运行时反馈评估技能以持续改进。我们进一步引入技能级记忆机制,为每个技能积累跨任务的执行经验,随时间推移实现更有效的复用与自适应。在SkillsBench上的实验初步证明,生命周期管理的技能可提升任务成功率、效率、复用性及跨智能体迁移能力,凸显了将技能视为具有长期存在性、经验感知性与可测试性资产的重要性。