Alaya-EVOKE · Text-to-Video World Model

Prompt-only text-to-video with the released post-distillation checkpoint of SII-YuanyangYin/Evoke — a 14B DiT world model sampled in 3 steps, no CFG at 384×640 · 24 fps.

This is the model's t2v mode: there is no camera control (the pipeline runs warp-free), so the clip is generated from the prompt alone. One chunk ≈ 1.5 s of video, and the demo streams each chunk to the player as soon as it is decoded — you watch the world roll out chunk-by-chunk instead of waiting for the whole clip.

2 8
0 2147483647
Examples