Local Ai

TencentARC/SCoPE · Hugging Face

SCoPE: Sightline-Coordinate Positional Encoding for Video Diffusion Transformers SCoPE adds camera sightlines as positional coordinates to a pretrained video diffusion transformer. Given a first frame

DGX agentreddit
local-air-localllama

SCoPE: Sightline-Coordinate Positional Encoding for Video Diffusion Transformers SCoPE adds camera sightlines as positional coordinates to a pretrained video diffusion transformer. Given a first frame, a text prompt, and a camera trajectory, it generates a video that follows the requested camera motion while preserving the original image-to-video prior. This repository is a self-contained release for Wan2.2-I2V-A14B: it contains everything required for inference, so a separate Wan2.2 checkpoint download is not needed. arXiv : https://arxiv.org/abs/2606.27345 PDF : https://arxiv.org/pdf/2606.27345 GitHub : https://github.com/TencentARC/SCoPE Project : https://visual-ai.github.io/scope/ Demo : https://huggingface.co/spaces/TencentARC/scope-camera-video-generation submitted by /u/pmttyji [link] [comments]

Source: r/LocalLLaMA | 2026-08-19

Loading related sources…