MoGe-3: Recover Detailed 3D Geometry from One Photo

Recover the visible geometry of a scene from a single photo. MoGe-3 uses sparse 3D refinement to preserve thin structures, small objects, and depth boundaries that a coarse prediction can smooth away.

Frame from the official MoGe-3 video demonstrating fine-detail monocular geometry
Before and after MoGe-3 refinement: straighter railings and finer geometry. MoGe-3

MoGe-3 is a geometry-estimation model developed by researchers at Tsinghua, USTC, and Microsoft Research. It predicts metric depth, point maps, normals, and camera geometry from a photo for depth-aware editing and scene reconstruction.

Why thin structures need 3D refinement

Pixels can be close in an image while belonging to surfaces far apart in 3D. Mixing their features can blur the geometry around railings, plants, and small objects.

MoGe-3's Self-Guided Sparse Refinement lifts an initial point map into a sparse voxel shell. Sparse 3D convolutions use spatial neighbors instead of only image-plane neighbors, and repeated refinement recovers finer structures.

Official MoGe-3 results: single-image geometry with fine-detail refinement. MoGe-3
MoGe-3 lifts a coarse point map into a sparse voxel shell before refining geometry.
MoGe-3 lifts a coarse point map into a sparse voxel shell before refining geometry. MoGe-3

What you can export from a photograph

The MoGe toolchain supports point maps, depth maps, normal maps, and estimated camera field of view. Its inference CLI can save maps alongside GLB and PLY geometry files.

GLB stores image colors as a texture for inspection in a 3D viewer. PLY stores colors on vertices for geometry and point-cloud workflows. Choose the format that matches your editor or processing pipeline.

Run MoGe-3 with an explicit checkpoint

Use Python 3.10 or newer and install the repository dependencies. Select the model with --version v3, then supply your MoGe-3 checkpoint through --pretrained PATH_TO_CKPT.pt.

The command below exports maps and geometry with three refinement steps. Increasing refinement is not the same as increasing input resolution; compare thin structures and boundaries in the output before adding more steps.

moge infer -i photo.jpg \
  --version v3 \
  --pretrained PATH_TO_CKPT.pt \
  --refine_steps 3 \
  --output output \
  --maps --glb --ply

Choose scene geometry or a complete 3D asset

Use MoGe-3 when the arrangement of visible surfaces in a photo is the information you need. Examples include depth-aware compositions, an initial room surface, and geometry inspection.

Use image-to-3D generation when you need a complete product, prop, or character to rotate and reuse. MoGe-3 and asset generation answer different questions, even though both can end with a 3D file.

MoGe-3 vs image-to-3D asset generation

NeedMoGe-3Asset generation
Primary goalRecover visible scene geometryGenerate a complete reusable object
Typical inputA photograph of a scene or objectA clear reference image of the subject
Unseen surfacesNot a measurement of hidden geometryCompleted by the generative model
Useful outputsDepth/point maps, GLB, PLYObject mesh and materials

Frequently asked questions

Is MoGe-3 a TRELLIS 3 release?

No. MoGe-3 is a separate geometry-estimation project. TRELLIS models generate 3D assets; MoGe estimates geometry from visible image evidence.

Which MoGe-3 checkpoint can I use?

The official releases include ViT-L with 370M parameters and ViT-G with 1.25B. Both support metric geometry and normal maps. Download the chosen checkpoint and pass it with --pretrained while selecting --version v3.

Can I run the official MoGe-3 code on macOS?

The current repository explicitly says macOS is unsupported because its FlexGEMM dependency relies on Triton without macOS wheels. Use a compatible environment or the official browser demo.

Is there an online demo or ComfyUI integration?

The project links to the Ruicheng/MoGe-3 Hugging Face demo. ComfyUI v0.37.0, released on September 21, 2026, also added MoGe 3 support.

Official sources

Need a complete object instead?

Use a reference image to generate a 3D asset you can rotate, inspect, and reuse.

Generate a 3D Asset