Human animation aims to generate temporally coherent and visually consistent videos over long sequences, yet modeling long-range dependencies while preserving frame quality remains challenging. Inspired by the human ability to leverage past observations for interpreting ongoing actions, we propose FrameCache, a training-free, causality-consistent reference frame framework. FrameCache explicitly converts historical generation results into causal guidance through two complementary mechanisms. First, at the reference level, a novel Screen-Cache-Match (SCM) strategy constructs a dynamic, high-quality reference memory, ensuring motion-consistent appearance guidance to reduce identity drift. Second, at the generative level, a Trajectory-Aware Autoregressive Generation (TAAG) mechanism aligns denoising trajectories across adjacent video chunks. This is achieved through an overlap-aware latent propagation and a dual-domain fusion strategy that seamlessly blends low-frequency structural layouts with high-frequency textural details. Extensive experiments on standard benchmarks demonstrate that FrameCache consistently improves temporal coherence and visual stability while integrating seamlessly with diverse diffusion baselines. Code will be made publicly available.
翻译:人体动画旨在生成长序列中时间连贯且视觉一致的视频,然而在保持帧质量的同时建模长程依赖关系仍具挑战性。受人类利用过往观察来解读当前行为的启发,我们提出FrameCache,一种免训练的因果一致参考帧框架。FrameCache通过两种互补机制将历史生成结果显式转化为因果引导。首先,在参考层面,一种新型的筛选-缓存-匹配(SCM)策略构建动态高质量参考记忆,确保运动一致的外观引导以减少身份漂移。其次,在生成层面,一种轨迹感知的自回归生成(TAAG)机制对齐相邻视频片段间的去噪轨迹。这通过重叠感知的潜在传播以及融合低频结构布局与高频纹理细节的双域融合策略实现。在标准基准上的广泛实验表明,FrameCache在无缝集成多种扩散基线的同时,持续提升了时间连贯性与视觉稳定性。代码将公开提供。