New AI framework to help machines understand 3D spaces in finer detail

A new framework developed by SUTD researchers could improve how robots, autonomous systems and augmented reality tools recognise objects and navigate complex real-world environments.

Case comparison of three 3D hierarchical semantic segmentation methods (MTHS [Li et al., 2020], DHL [Li et al., 2025], and our ML3DHS) on testing samples of S3DIS-H. L0 and L1 represent the two semantic hierarchies in the indoor point cloud scenes.

For a service robot moving through an office, recognising “furniture” is useful. When the robot has to plan a route or interact with objects, it also needs finer detail, such as whether that furniture is a chair, table or bookcase, and whether an opening is a door or a window. These distinctions can affect how a machine moves, avoids obstacles and decides what it can safely interact with.

Researchers led by Assistant Professor Na Zhao from the Singapore University of Technology and Design (SUTD), with collaborators from Southwest Jiaotong University, have developed a new artificial intelligence (AI) framework that helps machines make sense of three-dimensional (3D) scenes at multiple levels of detail.

Called ML3DHS, the framework is designed for 3D hierarchical semantic segmentation, which assigns several related labels to every point in a 3D scene, from broad categories such as “furniture” to finer labels such as “chair” or “table”. It was presented at the 43rd International Conference on Machine Learning in 2026, under the title “Multi-Label Learning with Contrastive Cluster Self-Supervision for 3D Hierarchical Semantic Segmentation”.

Three-dimensional vision systems are increasingly used to help machines interpret spaces captured by sensors such as LiDAR (Light Detection and Ranging) or depth cameras. These systems are used in embodied intelligence, including robots that navigate buildings, autonomous vehicles that interpret road scenes, and augmented reality applications that overlay digital information onto physical spaces.

However, many existing 3D segmentation models are built to assign a single label to each point in a scene. This “flat” approach can be limiting in real-world settings, where objects often need to be understood at different levels of detail. At a broad level, a model may identify a group of points as “furniture”. At a finer level, it must decide whether those same points belong to a chair, table, sofa or bookcase.

The challenge is that these related levels can compete during training. The broad prediction may be easier to learn, while the finer prediction requires more detailed distinctions. When both levels share too many model components, improving one can affect how well the other learns.

ML3DHS addresses this by allowing the model to share basic 3D information, such as shape and geometry, while using separate components for different levels of detail. It also uses broad predictions to guide finer ones, so recognising a region as furniture can help the model focus on whether it is a chair, table or sofa. The system checks that broad and fine labels remain consistent.

“Each 3D point can carry several related meanings at the same time. A point may belong to furniture at one level and table at another. By treating this as a multi-label learning problem, we can help machines build a more structured representation of 3D scenes.”
Assistant Professor Zhao Na, SUTD

“Each 3D point can carry several related meanings at the same time. A point may belong to furniture at one level and table at another,” said principal investigator, Assistant Professor Na Zhao from SUTD. “By treating this as a multi-label learning problem, we can help machines build a more structured representation of 3D scenes.”

Beyond reconciling multiple label levels, another persistent issue in 3D scene understanding is class imbalance. Common classes such as walls and floors often contain many more points than smaller or rarer objects such as windows, doors, columns and clutter. This can cause models to focus too heavily on dominant classes and miss less common objects that may be important for navigation and interaction. ML3DHS addresses this with an additional training component that helps the model tell object classes apart more clearly, giving rarer objects a stronger learning signal.

The team tested ML3DHS on established indoor and outdoor 3D scene benchmarks, using several AI model architectures. Across these settings, it consistently outperformed previous approaches. On one indoor benchmark, for example, it improved segmentation accuracy by 3.38 percentage points over the best previous method. The gains were especially notable for less common objects such as clutter, windows, doors and columns.

Assistant Professor Zhao added: “For an indoor service robot, broad understanding may tell it that a region contains furniture or an opening. Fine-grained understanding helps it distinguish a door from a window, a chair from a table, and a column from movable clutter. Those distinctions matter when a machine has to choose a safe path or decide which objects it can interact with.”

The researchers note that ML3DHS currently increases model size during training because it uses non-shared decoders and an additional auxiliary branch. Future work could explore lighter models that are more efficient to train and better at recognising extremely rare classes.

Even with these open questions, the framework nevertheless offers a more structured approach to recognising objects across different levels of detail. By giving AI systems a clearer picture of the environments around them, the work could contribute to more reliable machine perception for robots, autonomous systems and interactive 3D technologies.

Published: 30 Sep 2026

Contact details:

8 Somapah Road Singapore 487372

News topics: 
Academic discipline: 
Content type: 
Reference: 

Forty-Third International Conference on Machine Learning