Starlink, as a representative low Earth orbit (LEO) satellite broadband system, makes high-bitrate video streaming possible in regions where terrestrial broadband is unavailable. However, its access links exhibit rapid throughput fluctuations caused by satellite mobility and handovers. Existing learned adaptive bitrate (ABR) algorithms can achieve high average quality of experience (QoE), yet high-bitrate Starlink streaming exposes severe session-level rebuffering that is not captured by average QoE alone. To address it, this paper proposes SafeSABR, a risk-calibrated learned ABR framework for Starlink networks. SafeSABR formulates Starlink ABR as a QoE--severe-risk tradeoff and follows a three-stage design: behavior-cloning pretraining learns a high-QoE ABR prior, risk-calibrated reinforcement learning (RL) fine-tuning reduces severe-tail action tendencies, and a runtime safety auditor uses safe-capacity lower bounds to check policy-requested bitrates before execution. Experiments on real Starlink traces compare SafeSABR with online, prediction-assisted, and learned ABR baselines. Compared with advanced methods, SafeSABR reduces severe-stall sessions from 22.8% to 7.2% and worst-5% session rebuffering from 54.30 s to 22.68 s, with a 1.8% QoE cost. Component analyses further show that risk-calibrated fine-tuning and safe-capacity auditing reduce unsafe bitrate decisions and downstream severe-session rebuffering. These results show that combining risk-calibrated policy learning with decision-aware safe throughput forecasting can move learned ABR toward a safer QoE--severe-risk operating point under volatile Starlink networks.
翻译:星链作为代表性低地球轨道卫星宽带系统,使地面宽带不可达区域实现高码率视频流成为可能。然而,其接入链路因卫星移动和切换导致吞吐量剧烈波动。现有学习型自适应码率算法虽能实现高平均体验质量,但高码率星链流媒体会暴露未被平均QoE捕获的严重会话级重缓冲问题。为此,本文提出SafeSABR——面向星链网络的风险校准学习型ABR框架。SafeSABR将星链ABR建模为QoE与严重风险间的权衡,采用三阶段设计:行为克隆预训练学习高QoE的ABR先验,风险校准强化学习微调减少严重尾部动作倾向,运行时安全审核器在执行前利用安全容量下界检查策略请求的码率。基于真实星链轨迹的实验将SafeSABR与在线、预测辅助及学习型ABR基线对比。相较于先进方法,SafeSABR将严重卡顿会话比例从22.8%降至7.2%,最差5%会话重缓冲时长从54.30秒降至22.68秒,仅牺牲1.8%的QoE。组件分析进一步表明,风险校准微调与安全容量审核减少了不安全码率决策及下游严重会话重缓冲。这些结果表明,将风险校准策略学习与决策感知的安全吞吐量预测相结合,可在不稳定的星链网络中将学习型ABR推向更安全的QoE-严重风险运行点。