Papers
arxiv:2607.28624

PhiZero: A World Model Built Around Physical Language

Published on Jul 30
· Submitted by
taesiri
on Jul 31
Authors:
,
,
,
,

Abstract

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.

Community

os-x-mountain-lion-earth-horizon-cosmos-stock-5120x3202-4003

os-x-mountain-lion-earth-horizon-cosmos-stock-5120x3202-4003

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Hello, PhiZero learns physical language through self-supervised training on large collections of real-world videos. It My Fed Loan follows a reason-then-render approach: first predicting the sequence of future physical state transitions, then rendering those inferred changes into realistic videos. This explicit reasoning improves both interpretability and physical consistency compared to conventional video prediction methods.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.28624
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.28624 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.28624 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.28624 in a Space README.md to link it from this page.

Collections including this paper 3