We address the problem of synthesizing multi-view optical illusions: images that change appearance upon a transformation, such as a flip or rotation. We propose a simple, zero-shot method for obtaining these illusions from off-the-shelf text-to-image diffusion models. During the reverse diffusion process, we estimate the noise from different views of a noisy image. We then combine these noise estimates together and denoise the image. A theoretical analysis suggests that this method works precisely for views that can be written as orthogonal transformations, of which permutations are a subset. This leads to the idea of a visual anagram--an image that changes appearance under some rearrangement of pixels. This includes rotations and flips, but also more exotic pixel permutations such as a jigsaw rearrangement. Our approach also naturally extends to illusions with more than two views. We provide both qualitative and quantitative results demonstrating the effectiveness and flexibility of our method. Please see our project webpage for additional visualizations and results: https://dangeng.github.io/visual_anagrams/
翻译:我们研究多视角光学错觉的合成问题:即图像在变换(如翻转或旋转)后呈现不同外观。我们提出一种简单、零样本的方法,通过现成的文本到图像扩散模型获得这些错觉。在逆向扩散过程中,我们从噪声图像的不同视角估计噪声,再将噪声估计进行组合并去噪。理论分析表明,该方法精确适用于可表示为正交变换的视角,而排列变换是其子集。由此提出“视觉幻象图”概念——一种在像素重排下改变外观的图像,包括旋转、翻转以及更复杂的像素排列(如拼图式重排)。我们的方法还自然扩展至支持两个以上视角的错觉。通过定性与定量实验证明该方法的有效性与灵活性。更多可视化结果请参见项目网页:https://dangeng.github.io/visual_anagrams/