We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build FracSim, a fracture-aware simulation framework that augments material point method (MPM) simulation with a continuum damage model, producing paired fracture videos and dense, pixel-aligned physical fields at no additional cost beyond standard rendering. FracGen leverages these maps in two ways: it is trained to jointly predict them alongside RGB video, encouraging the model to capture physical state rather than surface appearance; and it is supervised with physics-informed losses that encourage consistency among the predicted maps. As a result, FracGen captures distinct material-specific fracture behavior without expensive test-time simulation or per-scene tuning, while offering fine-grained control over where an object tears, how fast the crack propagates, and how much deformation precedes failure. We further introduce a benchmark for evaluating the physical plausibility of generated fracture video, and show through extensive experiments that FracGen outperforms existing video generation baselines in both physical and visual fidelity.
Pick Intact Objects
Black Bread (Horizontal)
Black Bread (Vertical)
Bread Roll (Horizontal)
Bread Roll (Vertical)
Green Cake (Horizontal)
Green Cake (Vertical)
Kong (Horizontal)
Pig (Horizontal)
Pikachu (Horizontal)
Red Cake (Horizontal)
Red Cake (Vertical)
Steamed Bun (Horizontal)
Steamed Bun (Vertical)
Left: without physics map prediction · Right: with physics map prediction
Left: without physics consistency loss · Right: with physics consistency loss