W.A.L.T successfully integrated the Transformer architecture into a hidden video diffusion model to generate high quality, coherent, artifact-free video that maps video and images into a unified low-dimensional hidden space, significantly reducing the computational cost of generating high-resolution video; and designed a new Transformer block with a self-attention layer capable of non-overlapping, window-constrained alternating between spatial and spatio-temporal attention.
Using Transformer for Diffusion Modeling, AI Generates Video for Photo-Realism
Previous: 谷歌陷”虚假宣传”风波 承认演示视频系剪辑合成