正文

Maybe instead of generating video from scratch we’ll have LLMs animate scenes and then use diffusion models as more like a rendering layer