Starlink, as a representative low Earth orbit (LEO) satellite broadband system, makes high-bitrate video streaming possible in regions where terrestrial broadband is unavailable. However, its access links exhibit rapid throughput fluctuations caused by satellite mobility and handovers. Existing learned adaptive bitrate (ABR) algorithms can achieve high average quality of experience (QoE), yet high-bitrate Starlink streaming exposes severe session-level rebuffering that is not captured by average QoE alone. To address it, this paper proposes SafeSABR, a risk-calibrated learned ABR framework for Starlink networks. SafeSABR formulates Starlink ABR as a QoE--severe-risk tradeoff and follows a three-stage design: behavior-cloning pretraining learns a high-QoE ABR prior, risk-calibrated reinforcement learning (RL) fine-tuning reduces severe-tail action tendencies, and a runtime safety auditor uses safe-capacity lower bounds to check policy-requested bitrates before execution. Experiments on real Starlink traces compare SafeSABR with online, prediction-assisted, and learned ABR baselines. Compared with advanced methods, SafeSABR reduces severe-stall sessions from 22.8% to 7.2% and worst-5% session rebuffering from 54.30 s to 22.68 s, with a 1.8% QoE cost. Component analyses further show that risk-calibrated fine-tuning and safe-capacity auditing reduce unsafe bitrate decisions and downstream severe-session rebuffering. These results show that combining risk-calibrated policy learning with decision-aware safe throughput forecasting can move learned ABR toward a safer QoE--severe-risk operating point under volatile Starlink networks.
翻译:Starlink作为典型的低轨卫星宽带系统,使得地面宽带不可用地区的高比特率视频流媒体成为可能。然而,其接入链路因卫星移动和切换而呈现剧烈吞吐量波动。现有学习型自适应比特率(ABR)算法虽能实现高平均体验质量(QoE),但高比特率Starlink流媒体暴露了仅靠平均QoE无法捕捉的严重会话级缓冲现象。为此,本文提出SafeSABR——一种面向Starlink网络的风险校准学习型ABR框架。SafeSABR将Starlink ABR建模为QoE与严重风险之间的权衡,并采用三阶段设计:行为克隆预训练学习高QoE ABR先验,风险校准强化学习(RL)微调减少尾部严重动作倾向,运行时安全审计器在执行前利用安全容量下界检查策略请求的比特率。基于真实Starlink轨迹的实验将SafeSABR与在线型、预测辅助型及学习型ABR基线进行对比。与先进方法相比,SafeSABR将严重卡顿会话比例从22.8%降至7.2%,最差5%会话缓冲时间从54.30秒降至22.68秒,仅付出1.8%的QoE代价。组件分析进一步表明,风险校准微调与安全容量审计可减少不安全比特率决策及下游严重会话缓冲。这些结果表明,将风险校准策略学习与决策感知安全吞吐量预测相结合,能够在高度波动的Starlink网络中将学习型ABR推向更安全的QoE-严重风险运行点。