Alaya-EVOKE · Text-to-Video World Model
Prompt-only text-to-video with the released post-distillation checkpoint of SII-YuanyangYin/Evoke — a 14B DiT world model sampled in 3 steps, no CFG at 384×640 · 24 fps.
This is the model's t2v mode: there is no camera control (the pipeline runs warp-free),
so the clip is generated from the prompt alone. One chunk ≈ 1.5 s of video, and the demo
streams each chunk to the player as soon as it is decoded — you watch the world roll out
chunk-by-chunk instead of waiting for the whole clip.
2 8
0 2147483647
Examples