WorldDiT: A Unified Architecture for World and Action Modeling
WorldDiT is a pareto frontier robot world-action model architecture without a VLM backbone
We’re releasing WorldDiT, a pareto frontier model architecture bringing world and robot action modeling into one compact diffusion transformer, without relying on a large pretrained VLM action backbone.
WorldDiT in action
Three reasons WorldDiT is different
At 399 million total parameters, WorldDiT lies on the reported LIBERO Pareto frontier for model size and mean success.
Among publicly released methods in the comparison that do not use a large pretrained VLM action backbone, it reports the highest mean success.
WorldDiT unifies future visual world state prediction with continuous robot action generation in a single diffusion transformer.
🤗 Run WorldDiT from Hugging Face with the released checkpoints and complete inference runtime.
📄 Learn more about the architecture, training, and results in our technical report on arXiv.
If you use WorldDiT in your research, feel free to cite the paper.
@article{260723909,
title={{WorldDiT: A Unified Diffusion Architecture for World and Action Modeling}},
author={Sen Wang and R. Gnana Praveen and Bidhan Roy and Marcos Villagra},
year={{2026}},
eprint={{2607.23909}},
archivePrefix={{arXiv}}
}





