Distributed optimization has experienced a significant surge in interest due to its wide-ranging applications in distributed learning and adaptation. While various scenarios, such as shared-memory, local-memory, and consensus-based approaches, have been extensively studied in isolation, there remains a need for further exploration of their interconnections. This paper specifically concentrates on a scenario where agents collaborate toward a unified mission while potentially having distinct tasks. Each agent's actions can potentially impact other agents through interactions. Within this context, the objective for the agents is to optimize their local parameters based on the aggregate of local reward functions, where only local zeroth-order oracles are available. Notably, the learning process is asynchronous, meaning that agents update and query their zeroth-order oracles asynchronously while communicating with other agents subject to bounded but possibly random communication delays. This paper presents theoretical convergence analyses and establishes a convergence rate for the proposed approach. Furthermore, it addresses the relevant issue of deep learning-based resource allocation in communication networks and conducts numerical experiments in which agents, acting as transmitters, collaboratively train their individual (possibly unique) policies to maximize a common performance metric.
翻译:分布式优化因其在分布式学习与自适应中的广泛应用而备受关注。尽管诸如共享内存、本地内存和基于共识的方法等不同场景已被单独广泛研究,但对其相互关联的进一步探索仍显不足。本文聚焦于智能体协同完成统一任务但可能具有不同目标的场景。在此场景中,每个智能体的行动可能通过相互作用影响其他智能体。智能体的目标是根据局部奖励函数的总和优化其本地参数,且只能获取局部零阶信息。值得注意的是,学习过程是异步的,即智能体异步更新并查询其零阶信息,同时与其他智能体通信,通信延迟有界但可能随机。本文提出了理论收敛性分析,并建立了所提方法的收敛速率。此外,本文探讨了通信网络中基于深度学习的资源分配这一实际问题,并进行了数值实验:智能体(作为发射机)协同训练其各自(可能唯一)的策略,以最大化共同性能指标。