πŸ’ͺ ElasticFit Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation

Tzu-Hsin Hsieh, Ricardo Marroquim

Delft University of Technology

NeurIPS 2026
ElasticFit teaser

Figure 1. ElasticFit inserts generated 3D objects into pre-existing scenes from natural-language instructions. It combines VLM reasoning, generative adaptation, and geometry- and physics-aware refinement. Red text highlights key fitting constraints.

Given a language instruction and a 3D scene, ElasticFit reasons about where, what, and how to fit a generated object β€” adapting its geometry rigidly, uniformly, or elastically to make it fit the local space.

Abstract

Inserting objects into existing 3D scenes requires more than selecting a plausible location: the inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces.

We introduce ElasticFit, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode β€” rigid placement, uniform scaling, or elastic deformation. These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting.

ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding.


3 Challenges of Fit-Aware 3D Insertion

1
Grounding language in local 3D geometry.

Understanding where the object should be inserted and the spatial constraints of the target region.

Challenge 1 – Grounding language in local 3D geometry
2
Adapting object geometry to available space.

The target object may not naturally fit, requiring its geometryβ€”not just its poseβ€”to adapt to the local scene.

Challenge 2 – Adapting object geometry to available space
3
Ensuring physically plausible insertion.

The adapted object must remain collision-free, properly supported, and physically stable.

Challenge 3 – Ensuring physically plausible insertion

Method

ElasticFit operates in three stages: semantic-spatial grounding via Set-of-Mark VLM queries, scene-conditioned object generation using SDXL + TRELLIS, and geometry- and physics-aware fitting with three adaptation modes refined by PyBullet simulation.

ElasticFit pipeline overview
Figure 2. Overview of ElasticFit. Given a user instruction and an existing 3D scene, ElasticFit infers structured fitting cues G encoding the target region, orientation, constraints, and adaptation behavior. Stage 1 performs semantic-spatial grounding, Stage 2 generates a scene-conditioned object prior, and Stage 3 refines through geometry- and physics-aware fitting.
Three adaptation modes
Figure 3. Visualization of the three adaptation modes in Stage 3. Each row shows the target ghost box, generated initial object, and fitted result. Rigid mode preserves geometry; Uniform mode applies isotropic scaling; Elastic mode applies anisotropic deformation to fit constrained spaces.

Results

69.7%
Spatial Relation Success
91.7%
Support Success Rate
91.7%
Collision-Free Rate
4.18/5
Plausibility Score

Evaluated on 120 tasks across ReplicaCAD and Replica datasets. Improves spatial relation success from 50.8% β†’ 69.7% and support success from 48.3% β†’ 91.7% over the strongest baseline (LayoutGPT).

Qualitative comparison
Figure 4. Qualitative comparison on representative Common Placement Set tasks. Baselines often struggle with physical refinement, ambiguous references, and irregular support geometry. ElasticFit produces more spatially and physically plausible placements.
Common placement set results
Figure 5. Additional qualitative results on the Common Placement Set, showing ElasticFit's ability to handle diverse object types, spatial relations, and support geometries across ReplicaCAD and Replica scenes.
Fit-critical results
Figure 6. Qualitative results on the Fit-Critical Set, demonstrating ElasticFit's elastic deformation mode for objects that must conform to constrained spatial envelopes.

BibTeX

@inproceedings{hsieh2026elasticfit,
  title     = {ElasticFit: Fit-Aware 3D Object Insertion via
               VLM Reasoning and Generative Adaptation},
  author    = {Hsieh, Tzu-Hsin and Marroquim, Ricardo},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}