Qualitative examples show that LiDAR and camera modalities degrade differently under adverse conditions
Many of today’s self-driving cars “see” the road through two senses at once. Cameras capture colour, texture, and the outline of a pedestrian, while light detection and ranging (LiDAR) fires laser pulses to measure exactly how far away that pedestrian stands. Both provide complementary information and outperform what either can achieve alone.
This pairing has a weakness, however. Trained mostly on clear daytime footage, the fusion model learns to lean on whichever sensor performs best there and may fail in challenging scenarios. Rain can scatter LiDAR beams for instance, or a night drive leaves cameras battling darkness and glare.
To address this, Assistant Professor Zhao Na from Singapore University of Technology and Design (SUTD) led two studies that built different models. In “CCF: Complementary collaborative fusion for domain generalized multi-modal 3D object detection”, the research team rebalances how a detector weighs its sensors so it holds up in new surroundings. In “PanDA: Unsupervised domain adaptation for multimodal 3D panoptic segmentation in autonomous driving”, they harness unlabelled data from a new environment to help the model adapt once it arrives.
“The real challenge is not simply an increase in noise, but the fact that the reliability of each sensor changes dynamically,” said Assistant Prof Zhao. “Our work teaches the system to generalise beyond the training environment and adjust how much the model relies on each sensor as the available evidence changes.”
Complementary collaborative fusion (CCF) starts with a diagnosis. Detectors of this kind generate candidate objects from each sensor separately, then merge them. Examining a standard detector’s training, the team found LiDAR candidates matched to real objects 37.5 times as often as camera candidates. Conventional evaluations, which report only the fused result, had kept the imbalance hidden.
CCF’s fix constitutes three parts. Each sensor’s candidates get their own training signal, so camera candidates are no longer drowned out. Camera candidates also borrow depth from nearby LiDAR points to sharpen their placement. In addition, during training, complementary patches are masked out of the image and the point cloud, so wherever one sensor’s data is hidden the other’s is kept, forcing the model to learn from both. On nuScenes, an autonomous driving dataset, mean average precision rose by 2.8, 1.3, and 3.2 percentage points on rain, night, and Boston test sets respectively, while preserving daytime performance.
“There is no fixed rule such as ‘rain means camera’ or ‘darkness means LiDAR’,” Assistant Prof Zhao explained. “For each potential object, the model compares the available evidence and learns how much to rely on each sensor.”
PanDA picks up where CCF leaves off with panoptic segmentation, in which every LiDAR point must be labelled by category and, for countable objects such as cars and pedestrians, by individual instance. Instead of hand-labelling footage from a new city or a rainy season, PanDA has the model generate its own provisional labels, then refines them.
In particular, two ideas make this work. The first damages the training data on purpose, blanking out patches inside objects and along their edges from either the image or the LiDAR scan, so the model learns to reconstruct the scene from the surviving sensor and surrounding context. The second cleans up the provisional labels with two experts. A 3D expert uses geometric continuity in the LiDAR data to extend road surfaces and other regions the provisional labels left fragmented. A 2D expert consults large pretrained vision models to relabel objects the model is unsure about.
“The 3D expert can be viewed as a surveyor, while the 2D expert acts as a recogniser,” said Assistant Prof Zhao. “The former repairs incomplete shapes, while the latter corrects classification errors.”
On nuScenes, PanDA lifted panoptic quality over the unadapted baseline by 13.2 percentage points from Boston to Singapore, 8.9 from sunny to rainy weather, and 8.4 from day to night. In the night setting it scored 73.1, higher than two supervised reference models trained with labelled night data, which reached 53.5 and 70.6, respectively.
Assistant Prof Zhao highlighted one caveat: the night domain contains only 602 training frames, constraining the reference models. Before either method reaches a production vehicle, she says, it must be tested on rare events, in other cities, and with other sensor configurations. Its processing speed, reliability, and safety mechanisms must also be validated.
“The next major milestone is to develop perception systems that generalise reliably across different sensor types and configurations. They must also be able to assess which information is trustworthy and make correct perception decisions across diverse scenarios, including rain, nighttime driving, unfamiliar cities, and complex traffic conditions,” she shared.


