Semantic Driven Multi-Camera Pedestrian Detection
Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós, Pablo Carballeira
Introduction
In the current worldwide situation, pedestrian detection has reemerged as a pivotal tool for intelligent video-based systems aiming to solve tasks such as pedestrian tracking, social distancing monitoring or pedestrian mass counting. Automatic people detection is generally considered a solid and mature technology able to operate with nearly human accuracy in generic scenarios garcia2015people ; dollar2012pedestrian ; priscilla2019pedestrian . However, the handling of severe-occlusions is still a major challenge ning2020survey . Occlusions occur due to the projection of the 3D objects onto a 2D image plane. Although recent deep-learning based methods are able to cope with partial occlusions, the detection process fails when only a small part or no part of the person is visible. To cope with severe-occlusions, a potential solution is the use of additional cameras: if they are adequately positioned, the different points of view might allow for disambiguation.
Disambiguation is generally achieved by projecting every camera’s detections on a common reference plane. The ground plane is usually the preferred option as it constitutes a common reference in which people’s height can be disregarded. Per-camera detections can then be combined on the ground plane to refine and complete pedestrian detection. However, there are several challenges to be addressed during this combination or fusion process. Among the striking ones are: the convenience to define common visibility areas where cameras’ views overlap, and how to cope with camera calibration errors and persons’ self-occlusions. See Figure 1 for visual examples of these challenges, which we detail below:
In multi-camera approaches a common strategy is to define an operational area or Area Of Interest on the ground plane. This area represents the overlapping field-of-view of all the involved cameras. It can be used to reduce the impact of calibration errors in the process and to generally ease the fusion of per-camera detections. This area is generally manually defined for each scenario, precluding the automation of the process.
Scene calibration is a well-known task hartley2003multiple which can be performed either manually or using automatic calibration methods based on image cues. In both cases, small perturbations in the calibration process may cause uncertainty in the fusion of the detections on the ground plane. The impact of calibration errors increases with the distance to the camera: generally, calibration is more accurate for pixels belonging to objects close to the camera.
Self-occlusions are caused by the intrinsic three-dimensional nature of people, resulting in the occlusion of some human parts by some others. If the visible parts are different for different cameras and these are used to project a person location on the ground plane, the cameras’ projections will diverge, hindering their fusion.
To cope with these challenges, in this paper we present a multi-camera pedestrian detection method which is driven by semantic information automatically extracted from the 2D images and transferred to the 3D ground plane, and includes the following novel contributions:
A novel approach to globally combine pedestrian detections in a multi-camera scenario by creating connected components in a graph representation of detections.
An height-adaptive optimization algorithm which uses semantic cues to globally refine the location and size of people detections by aggregating information from all the cameras.
The proposed method is applied over an operational area in the ground plane, which is automatically defined by an adaptation of the method described in lopez2018automatic .
Experimental results on public datasets (PETS 2009 web-PETS , EPFL RLC baque2017deep ; web-RLC , EPFL Terrace fleuret2008multicamera ; web-Terrace and EPFL Wildtrack chavdarova-et-al-2018 ; web-wildtrackDataset ) prove that the proposed method: (1) outperforms state-of-the-art monocular pedestrian detectors renNIPS15fasterrcnn ; redmon2018yolov3 , (2) outperforms state-of-the art scene-agnostic multi-camera detection approaches; and (3) results in a performance comparable, and even better, to deep-learning multi-camera detection approaches trained and fine-tuned to the target scenario, while not requiring neither a manually annotated operational area nor a specific training on that scenario.
The rest of the paper is organized as follows: Section 2 reviews the State of the Art, Section 3 describes the proposed method, Section 4 presents and discusses experimental results and Section 5 concludes the paper.
Related Work
Multi-camera people detection faces the combination, fusion and refinement of visual cues from several individual cameras to obtain more people locations. A common pathway in existing approaches starts by defining an operational area, either manually or, as we propose, based on a semantic segmentation. Then, approaches combining detections using a common reference plane, usually follow a three-stages strategy: (1) extract detections on each camera frame, (2) project detections onto the common plane and (3) combine detections and back-project them to the individual views to obtain per-camera people detections. Finally, obtained detections are sometimes post-processed to further refine their localization.
Some approaches peng2015robust ; utasi2013bayesian rely on manually annotated operational areas where evaluation is performed. An advantage of these ad hoc areas is that the impact of camera calibration errors is limited and controlled. Besides, these areas are defined to maximize the overlapping between the field of view of the involved cameras. However, the manual annotation of these operational areas hinders the generalization of people detection approaches. Our previous work in this domain lopez2018automatic resulted in an automatic method for the cooperative extraction of operational areas in scenarios recorded with multiple moving cameras: semantic evidences from different junctures, cameras, and points-of-view are spatio-temporally aligned on a common ground plane and are used to automatically define an operational area or Area of Interest ().
2 Semantic Segmentation
Semantic segmentation is the task of assigning a unique object label to every pixel of an image. During the last years, top performing strategies evolved from the seminal Fully Convolutional Network scheme long2015fully and the use of dilated convolutions YuKoltun2016 . For instance, Zhao et al. zhao2016pyramid proposed to implicitly use contextual information by including relationships between different labels—e.g. an airplane is likely to be on a runway or flying in the sky but not on the water. These relationships reduce the inner complexity of datasets with large sets of labels, generally improving performance. Lately, the development and use of new backbones for feature extraction has benefited the task. Zhang et al. zhang2020resnest proposed a new ResNet modification called ResNeSt that uses channel-wise attention to capture cross-feature interactions and learn diverse object representations. Similarly, Tao et al. tao2020hierarchical proposed the use of a hierarchical attention to combine multi-scale predictions, increasing the performance on the small object instances, as those in PASCAL VOC dataset everingham2010pascal .
3 Monocular people detection
As stated in Section 1, automatic monocular pedestrian detection is considered a mature technology able to obtain accurate results in a broad range of scenarios. Well established object detectors based on CNNs as Faster-RCNN renNIPS15fasterrcnn and YOLOv3 redmon2018yolov3 , have demonstrated their reliability during the last years. Adapting their core schemes, recent approaches have further increased their performance. Specifically, YOLOv3 has been improved by both decreasing the complexity of the model through new architecture designs long2020pp and by efficient model scaling wang2020scaled .
Alternatively, novel detectors—also based on CNNs, have been proposed. Tan et al. tan2020efficientdet proposed a weighted bi-directional feature pyramid network allowing easy and fast multi-scale feature fusion, and obtaining a new family of detectors called EfficientDet, that achieved a new state-of-the-art performance in the COCO dataset lin2014microsoft . Alternatively, Zhu et al. zhu2020deformable proposed to use attention in the form of deformable transformers to also obtain state-of-the-art results.
Nevertheless, even though the most recent works have demonstrated really high performances, in scenarios with severe-occlusions the performance of these algorithms decreases.
4 Projection of per-camera detections
Multi-camera pedestrian detection is fundamentally based on the projection of monocular detections onto a common reference plane. Projection is typically achieved either by using calibrated camera models that relate any 2D image point with a corresponding referenced 3D world direction utasi2013bayesian , or by relying on homographic transformations that project image pixels to a specific 3D plane peng2015robust . In both cases, the ground plane, where people is usually standing on, is chosen as reference for simplicity reasons.
5 Fusion and refinement of per-camera detections
Fusion and refinement approaches can be mainly divided into three different groups depending on how global detections are obtained. The first group encompasses geometrical methods, which combine detections based on the geometrical intersections between image cues. The second group embraces probabilistic methods, that combine detections via optimization frameworks and statistical modeling of the image cues. The third group is composed of solutions based on the ability of deep learning architectures to model occlusions and achieve accurate pedestrian detection at scene level.
Regarding geometrical methods, detections are combined by projecting foreground masks to the ground plane in a multi-view scenario: the intersection of foreground regions leads to pedestrian detection alahi2011sparsity . Accuracy can be increased by projecting the middle vertical axis of pedestrians, leading to a more accurate intersection on the ground plane and, therefore, to a better estimation of the pedestrian’s position kim2006multi . Following the same hypothesis, the use of a space occupancy grid to combine silhouette cues has been proposed: each ground pixel is considered as an occupancy sensor and observations are then used to infer pedestrian detection franco2005fusion . All of these approaches outperform single-camera pedestrian detection algorithms by the use of ground-plane homography projections. Nevertheless, the evaluation of foreground intersections in crowded spaces may lead to the appearance of phantoms or false detections. To handle this problem, the general multi-camera homography framework has been extended by using additional parallel planes to the ground plane delannay2009detection ; khan2009tracking . The intersection of the image cues with these parallel planes is expected to suppress these phantoms. Similarly, parallel planes can be also used to create a full 3D reconstruction of pedestrians, that can then be back-projected to each of the camera views, improving monocular pedestrian detection Aliakbarpour2016 . Finally, Lima et al.lima2021generalizable replicates a preliminar version of the method proposed in this paper, which is available as a preprint lopez2018semantic , with the addition of people re-identification features to guide the fusion of per-camera detections.
Among probabilistic methods, an interesting example is the use of a multi-view model shaped by a Bayesian network to model the relationships between occlusions peng2015robust . Detections are here assumed to be images of either pedestrians or phantoms, the former differentiated from the latter by inference on the network.
Recent approaches are focused on deep learning methods. The combination of CNNs and Conditional Random Fields (CRF) can be used to explicitly model ambiguities in crowded scenes baque2017deep . High-order CRF terms are used to model potential occlusions, providing robust pedestrian detection. Alternatively, multi-view detection can be handled by an end-to-end deep learning method based on an occlusion-aware model for monocular pedestrian detection and a multi-view fusion architecture chavdarova2017deep .
6 Improving detection’s localization
Algorithms in all of these groups require accurate scene calibration: small calibration errors can produce inaccurate projections and back-projections which may contravene key assumptions of the methods. These errors may lead to misaligned detections, hindering their later use. To cope with this problematic, one can rely on an Height-Adaptive Projection (HAP) procedure in which a gradient descent process is used to find both the optimal pedestrian’s height and location on the ground-plane by maximizing the alignment of their back-projections with foreground masks on each camera peng2015robust .
Proposed Pedestrian Detection Method
The proposed method is depicted in Figure 2. First, state-of-the-art algorithms for monocular pedestrian detection and semantic segmentation are used to extract people detections and the semantic cues for each camera respectively. These cues drive the automatic definition of the , and detections outside this area are discarded. Surviving per-camera detections are combined to obtain global 3D detections by establishing rules and constraints on a disconnected graph. These detections are back-projected to their original camera views in order to further refine their location and height estimates.
Monocular Pedestrian Detection: is performed using a state-of-the-art detector. In order to avoid a potential height-bias, we ignore the height and width of the detected bounding boxes, i.e. the pedestrian detection at camera is just represented by the middle point of the base of its bounding box: , in homogeneous coordinateswe use common notation, upper case to denote 3D points/coordinates and lower case to denote 2D camera plane points/coordinates..
Semantic Segmentation: is performed using a state-of-the-art semantic segmentation algorithm. The method is used to label each image pixel for every camera and every frame : , where is one of the pre-trained semantic classes: , where , i.e. floor, building, wall… Figure 3 depicts examples of semantic labels for selected camera frames of the Terrace Dataset fleuret2008multicamera ; web-Terrace .
Projection of People Detections: Let be the homography matrix that transforms points from the image plane of camera to the world ground-plane. The detection of camera , is projected onto the ground plane by:
which corresponds to a 3D point of the ground plane.
2 Pedestrian Semantic Filtering
To obtain a semantic partition of the ground-plane, an adaptation of lopez2018automatic for static-camera scenarios is carried out. We first project every image pixel via . Every projected point inherits the semantic label assigned to :
Thereby, a semantic locus—a ground-plane semantic partition, is obtained for each camera. The extent of each locus is defined by the image support, and missing points inside the locus are completed by nearest-neighbor interpolation.
In order to globally reduce the impact of moving objects and segmentation errors, we propose to temporally aggregate each locus along several frames. In a set of loci obtained for consecutive frames, a given point on the ground plane is labeled with semantic labels, which may be different owing to inaccuracies in the semantic segmentation or to the presence of moving objects. A single temporally-smoothed label is obtained as the mode value of this set. Examples of these per-camera obtained smoothed loci are included in the first four-columns of Figure 4.
We propose to combine these loci to define the . The definition of the is scenario-dependent but can be generalized by defining a set of ground-related semantic classes: floor, grass, pavement, etc. The operational area is obtained as the union of the projected pixels from any camera which are labeled with any class in :
An example of a so-obtained is included in the right-most column of Figure 4.
Detection Filtering
Projected detections lying outside the operational area, , are filtered out and so, discarded for forthcoming stages.
3 Fusion of Multi-Camera Detections
We propose a geometrical approach to combine detections on the ground-plane. Every camera single detection is considered a vertex of a disconnected graph located in the reference plane. Vertices are then joined generating connected components , each representing a joint 3D global detection. The whole fusion process is summarized in Figure 5. The conditions that shall be satisfied to join two vertices or detections, and , are:
That vertices in a connected component are close enough. The - between any two vertices in shall be smaller than a predefined distance : (Figure 5 (a)). may be fixed in the interval between and with no influence in the results. We experimentally set meters to: 1) reduce the computational cost of the final stage (see below) assuming that vertices separated do not belong to the same object and 2) protect against calibration errors, assuming that they are not larger than .
That vertices in a connected component come from different cameras. This condition prevents the joining of two different detections from the same camera which are near in the ground plane. (Figure 5 (b))
To avoid ambiguities, the creation of connected components is performed in order, according to the spatial position of the detections: those with a lower module are combined first.
The outcome of the fusion process for cameras is a set of connected components , each containing detections: , where when a person is occluded or not detected in one or more cameras.
As each connected component is assumed to represent a single person, an initial ground-position of the person is obtained by simply computing the arithmetic mean of all the detections in the connected components (Figure 5 (c)).
4 Semantic-Driven Back-Projection
To obtain correctly positioned detections, i.e. visually precise detections, in each camera, ground-plane detections need to be back-projected to each camera and 2D bounding-boxes enclosing pedestrians need to be outlined based on these projections.
Let be an orthogonal line segment to the ground plane which represents the detected pedestrian and extends from the detection to a 3D point meters above. Using the camera calibration parameters, the segment can be back-projected onto camera . This back-projection defines a 2D line segment , which extends between and (see Figure 6 (a)).
We propose to create 2D bounding-boxes around these back-projected 2D line segments. To this aim, each segment is used as the vertical middle axis of its associated 2D bounding-box . For simplicity, the width of is made proportional to its height. Due to pedestrian self-occlusion, calibration errors and the uncertainty on the pedestrians’ height, this back-projection process results in misaligned bounding-boxes (see Figure 6 (a)), hindering their later use and degrading camera-wise performance.
To handle this problematic, we define an iterative method which aims to globally optimize the alignment between all 3D detections and their respective views or back-projections in all cameras. This method is based on the idea proposed in peng2015robust . While the referenced method is guided by a foreground-segmentation, we instead propose to use a cost-function driven by the set of pedestrian-labeled pixels from the semantic segmentation (e.g. see person label in Figure 3). Next, we detail the full process for the sake of reproducibility.
Method Overview
As a 3D detection , with height , inevitably results in misaligned back-projected 2D detections, the proposed method tries to adapt the 3D detection segment to each camera, generating a set of 3D detection segments, , for each 3D detection and iteratively modifying their positions and height to maximize 2D detections’ alignment with the semantic segmentation masks, while constraining all the segments to have the same final height (as they are all projections of a same pedestrian) and to be located sufficiently close to each other. This process is not performed independently for each 3D detection but jointly and iteratively for all 3D detections. Observe that the joint nature of the optimization problem for all 3D detections is a key step as pedestrian pixels may contain segmentations from more than one pedestrian.
For each 3D segment , the method starts by initializing (i.e., iteration ) the per-camera adapted segments:
Iterative steepest-ascent algorithm
where defines the maximum distance between 3D projections of a single pedestrian, which we set to twice the average width of the human body, i.e. 1 meter, to forestall the effect of nearby pedestrians in the image plane. Performed experiments suggest that variations in value have no significant influence on the results.
where is the distance from to the vertical middle axis of the back-projected bounding box .
At each iteration , the set of camera-adapted segments is moved towards the direction of maximum increment:
The algorithm continues until convergence is reached or the -constrain is violated.
Results
This section addresses the evaluation of the proposed method. To this aim, we first describe the evaluation framework; then, in the ablation studies, we measure the performance improvement of each of the method’s stages; finally, we finish by comparing our approach with alternative state-of-the-art approaches in classic and recent multi-camera datasets.
Results are obtained by evaluating the proposed method over five scenarios extracted from four publicly available multi-camera datasets in which cameras are calibrated and temporally synchronized:
EPFL Terrace fleuret2008multicamera ; web-Terrace : Generally used in the state-of-the-art to evaluate multi-camera approaches. It consists of a 5000 frames sequence per camera showing up to eight people walking on a terrace captured by four different cameras. All the cameras record a close-up view of the scene.
EPFL RLC baque2017deep ; web-RLC : Consists of an indoor sequence of 2000 frames per camera recorded in the EPFL Rolex Learning Center using three static HD cameras with overlapping field of views. All these cameras represent close-up views of the scene.
EPFL Wildtrack chavdarova-et-al-2018 ; web-wildtrackDataset : A challenging multi-camera dataset which has been explicitly designed to evaluate deep learning approaches. It has been recorded with 7 HD cameras with overlapping fields of view. Pedestrian annotations for 400 frames are provided. All of them are used to define the evaluation set used in this paper.
PETS 2009 web-PETS : The most used video sequences from this widely used benchmark dataset have been chosen.
PETS 2009 S2 L1, which contains 795 frames recorded by eight different cameras of a medium density crowd—in this evaluation, we have just selected 4 of these cameras: view 1 (far field view) and views 5, 6 and 8 (close-up views).
PETS 2009 City Center (CC), recorded only using two far-field view cameras with around 1 minute of annotated recording (400 frames per camera).
Table 1 contains a comparative description of these datasets including the type of data and annotations provided, as well as a subjective indication of their complexity for the pedestrian detection task.
Performance Indicators
To obtain quantitative performance statistics according to an experiment-based evaluation criterion the following state-of-the-art performance indicators have been selected: Precision (P), Recall (R), F-Score (F-S), Area Under the Curve (AUC), N-MODA (N-A) and N-MODP (N-P) dollar2009pedestrian ; stiefelhagen2006clear . To globally assess performance, a single value for each statistic and each configuration is provided by averaging per-camera ones.
2 System Setup
A common setup has been used for all the presented results. Faster-RCNN renNIPS15fasterrcnn , YOLOv3 redmon2018yolov3 and EfficientDet-D7 tan2020efficientdet are used as baseline algorithms to obtain monocular pedestrian detections. The three object detectors are pre-trained on the COCO dataset lin2014microsoft and we do not fine-tune nor adapt them to any of the faced scenarios. For the semantic segmentation, the Pyramid Scene Parsing Network (PSP-Net) zhao2016pyramid , pre-trained on the ADE20K dataset zhou2017scene (, has been selected considering a trade-off between performance and efficiency.)
In the Pedestrian Semantic Filtering stage, all frames in each sequence are used for temporal and spatial semantic aggregation, i.e. . For the Semantic-Driven Back-Projection stage, the initial height estimation has been set to an average pedestrian height of m. Besides, for all the datasets, convergence in the iterative steepest-ascent algorithm has been reached before or at the iteration.
3 Results Overview
The evaluation has been performed carrying out two different studies:
The Ablation Studies aim to gauge the impact of the different stages in the performance of the proposed approach. To this end, the following versions of the proposed method are compared:
“Baseline (Faster-RCNN, YOLOv3 and EfficientDet-D7)”, provides reference results of monocamera pedestrian detectors.
“Baseline + Filtering (Filt)” is a simplified version of our method which aims to independently evaluate the effect of the proposed automatic computation obtained by the “Pedestrian Semantic Filtering” stage.
“Baseline + Filtering (Filt) + Fusion (Fus) + Back-Projection (BP)” is the full version of the proposed method, which additionally evaluates the “Fusion of Multi-Camera Detections” and “Semantic-Driven Back-Projection” stages.
Ablation Studies are conducted on four of the described datasets: Terrace, PETS 2009 S2 L1, PETS 2009 CC and RLC.
State-of-the-art Comparison results analyze the proposed method with respect to several non deep-learning state-of-the-art multi-camera pedestrian detectors on the same four scenarios used in the Ablation Studies. Additionally, the method is compared with novel deep-learning methods on the Wildtrack dataset.
4 Ablation Studies
The availability of bounding-box annotations permits to use the classic performance criterion garcia2015people : a detection is considered a TP one if the Intersection Over Union (IoU) with a ground-truth bounding-box is higher than .
Results
Table 2 agglutinates the method’s performance on a per-stage basis. Qualitative examples of automatically generated s and algorithm results are depicted in Figure 7 and Figure 8 respectively. A visual example of the limitations of the Semantic-Driven Back-Projection stage is included in Figure 9.
Discussion
Table 2 shows that filtering-out detections using automatically generated s (Baseline + Filtering) improves the performance of all the baselines for datasets where the ground-plane area does not cover the whole image representation, i.e. datasets containing close-up views of the scene as EPFL Terrace and RLC. In these datasets, our precise s reduce phantom detections obtained by the baseline detectors. Although s are automatically computed, they are more precise (tighter to real scene edges) than those defined in the dataset.
Overall, in the EPFL Terrace dataset, the performance of Faster-RCNN + Filtering improves Faster-RCNN by and in terms of AUC and N-MODA respectively. YOLOv3 + Filtering presents relative increments over YOLOv3 baseline of and for AUC and N-MODA respectively. Finally, EfficientDet + Filtering also overcomes its baseline results by a for N-MODA.
For the EPFL RLC dataset with the proposed , Faster-RCNN is improved by a regarding AUC and by a concerning N-MODA. For YOLOv3, relative increments of a and a in terms of AUC and N-MODA are achieved. EfficientDet gains relative increments of a and a for AUC and N-MODA.
The proposed Filtering stage does not improve baselines’ performance for those datasets in which the ground-plane dominates the scene, i.e. those recorded with far-field view cameras as both scenarios from PETS 2009. In these cases, although the baseline pedestrian detectors may create phantom detections, those lie inside the proposed and no false-pedestrians are suppressed. However, as depicted in Figure 7, the automatically obtained s are larger and more precise than the original operational areas in the datasets, thereby obtaining a more realistic and exhaustive evaluation. Furthermore, observe how the proposed generation method effectively handles multi-class ground partitions as in the PETS 2009 dataset, where the proposed encompasses road, grass, pavement and side-walks classes enabling a high adaptability to unseen scenarios (see Figure 7 right).
Table 2 also shows that the complete method (Baseline + Filtering + Fusion + Back-Projection) notably improves Faster-RCNN baseline’s performance, mainly in scenarios with heavy occlusions, i.e. EPFL Terrace and EPFL RLC (See Table 1 for details). Specifically, for the EPFL Terrace dataset results are relatively increased a , a and a in terms of AUC, F-Score, and N-MODA respectively, whereas relative improvements are of a —in AUC, a —in F-Score terms—and a in N-MODA, for the EPFL RLC dataset.
For YOLOv3 and EfficientDet detectors a similar analysis arises. In scenarios where heavy occlusions are present—EPFL Terrace and RLC datasets, performance is increased. For the EPFL Terrace Dataset, relative increments of a , a and a are obtained when using YOLOv3 detections in terms of AUC, F-Score and N-MODA respectively. In the case of EfficientDet, increments of a , a and a are obtained with respect to the same metrics. For the EPFL RLC dataset, the improvement increases to a , a and a for YOLOv3, whereas a , a and a relative increase is obtained for EfficientDet in terms of AUC, F-Score and N-MODA respectively.
For both PETS scenarios the performance of the EfficientDet mono-camera detector is saturated (97% F-Score). The specific characteristics of this dataset: low pedestrian density over a wide space, low level of occlusions and a high point of view due to cameras being hanged up in streetlights (see Table 1 and Figure 8), turns it in the least complex dataset among those analysed. The generation of new false positive detections and the optimization process related problems ((Figure 9)) may lead to a slight decrease when saturated baselines are used in low complex datasets. Leaving these specific situations aside, the benefits of the proposed method are evident if one accounts for both performance indicators and qualitative results (see Table 2 and Figure 8 respectively): the proposed multi-camera detection approach is able to cope with partial, severe and complete occlusions by combining detections from all the cameras through the proposed semantic-guided process leading to an increase of all the reported metrics.
Focusing specifically on the Semantic-Driven Back-Projection process, results in Figure 8 depict highly tight pedestrian bounding boxes, independently from people’s height, self-occlusions and calibration problems, suggesting that the optimization process is able to automatically adapt bounding-boxes by jointly estimating pedestrian heights and world positions. Results in Table 2 corroborate this observation. Semantic-Driven Back-Projection leads to a higher overlap between detections and ground-truth annotations: in terms of the N-MODP metric, the proposed method achieves relative improvements of a for EPFL Terrace, a for both PETS 2009 S2 L1 and PETS 2009 CC and a for the RLC dataset when Faster-RCNN is used as the baseline detector. When YOLOv3 is used as the baseline detector, our method achieves a N-MOPD increment of a for EPFL Terrace whereas the N-MODP metric remains stable for PETS 2009 S2 L1, PETS 2009 CC and RLC datasets, suggesting that YOLOv3 individual performance for these datasets is already heaped. A similar result arises when using EfficientDet detector which by default is highly tight to pedestrians. Results are increased only for PETS CC dataset by a while manteined for the rest of the datasets. It is important to remind that, even tough the N-MODP metrics are sometimes slightly reduced or maintained, without the proposed Semantic-Driven Back-Projection process the back-projected bounding boxes and ground-truth would be misaligned (see Figure 6) decreasing the performance in terms of all the accuracy metrics of the proposed method.
Nevertheless, the optimization cost function aims to maximize the 2D detections’ alignment with the semantic segmentation masks, leading to a bias towards wider pedestrians by design, a situation that may result sometimes into wrong relocations of the back-projected bounding-boxes. Figure 9 shows an example of this case: notice the erring behaviour in Camera 1 when there is an extreme overlapping.
5 State-of-the-art Comparison
The same criterion used in the Ablation Studies applies for the Terrace, PETS and RLC datasets. However, in the Wildtrack dataset, as the ground-truth is provided via detections on the world ground plane (i.e., no bounding-boxes are provided), the evaluation criterion is different. Specifically, a detection is considered a TP if it lies at most m to a ground-truth annotated point web-wildtrackDataset . This radius roughly corresponds to the average width of the human body. Due to the absence of bounding-boxes, for this dataset the Semantic-Driven Back-Projection stage is not included.
State-of-the-art Algorithms
The following multi-camera algorithms have been selected to carry out the comparison:
POM fleuret2008multicamera . This algorithm proposes to estimate the marginal probabilities of pedestrians at every location inside an . It is based on a preliminary background subtraction stage.
POM-CNN fleuret2008multicamera . An upgraded version of POM in which the background subtraction stage is performed based on an encoder-decoder CNN architecture.
MvBN+HAP peng2015robust . Relies on a multi-view Bayesian network model (MvBN) to obtain pedestrian locations on the ground plane. Detections are then refined by a Height-Adaptive Projection method (HAP) based on an optimization framework similar to the one proposed in this paper, but driven by background-subtraction cues.
RCNN-Projected xu2016multi . The bottom of bounding-boxes obtained thorough per-camera CNN detectors are projected onto ground-plane, where 3D proximity is used to cluster detections.
Deep-Occlusion baque2017deep is an hybrid method which combines a CNN trained on the Wildtrack dataset and a Conditional Random Fields (CRF) method to incorporate information on the geometry and calibration of the scene.
DeepMCD chavdarova2017deep is an end-to-end deep learning approach based on different architectures and training scenarios:
Pre-DeepMCD: a GoogleNet szegedy2015going architecture trained on the PETS dataset.
Top-DeepMCD: a GoogleNet szegedy2015going architecture trained on the Wildtrack dataset.
ResNet-DeepMCD: a ResNet-18 he2016deep architecture trained on the Wildtrack dataset.
DenseNet-DeepMCD: a DenseNet-121 huang2017densely architecture trained on the Wildtrack dataset.
Results
Table 3 includes performance indicators for the proposed method compared with multi-camera algorithms POM fleuret2008multicamera and MvBN+HAP peng2015robust on the Terrace, PETS and RLC scenarios (results for the compared methods are extracted from peng2015robust ). Table 4 compares the performance of the proposed approach against deep-learning methods, some of them explicitly trained with data from the Wildtrack dataset (which we denote as Fine-Tuned) and others trained with data from other datasets or not even trained (which we denote as not Fine-Tuned). Performance indicators for these methods are extracted from chavdarova-et-al-2018 . In addition, qualitative results for the Wildtrack dataset are presented in Figure 10, including obtained detections in camera frames, global detections on the ground-plane and the automatically computed .
Discussion
Results in Table 3 show that the proposed approach (Baseline + Filtering + Fusion + Back-Projection), either with Faster-RCNN, YOLOv3 or EfficentDet baseline, outperforms the MvBN + HAP and the POM-CNN methods in terms of N-MODA metric. The proposed method obtains better results in terms of N-MODA which, precisely, measures detection accuracy along the whole video sequences. Best results are obtained when EfficientDet is used to extract mono-camera detections. Specifically, N-MODA is increased a for EPFL Terrace and a for both PETS 2009 S2 L1 and CC. Moreover, it obtains the higher performance on the heavily occluded RLC dataset. Besides, N-MODP results, i.e. the overlapping between detections and ground-truth, are better than those obtained by the HAP method peng2015robust . This suggests that our use of semantic segmentation masks instead of foreground masks (HAP method) benefits the optimization process. Relative increments in N-MODP performance of a for EPFL Terrace, a for PETS 2009 S2 L1 and a for PETS 2009 CC support this assumption.
Presented results in Table 3 suggest that the proposed method is able to obtain reliable pedestrian detections in a variety of scenarios with a variety of pedestrian detection algorithms in terms of performance.
Finally, results on the Wildtrack dataset (Table 4), indicate that the proposed method, operating on detections from a Faster-RCNN, a YOLOv3 or a EfficientDet model, is able to outperform deep-learning approaches that have not been specifically trained using Wildtrack data and use manually annotated detection constrains. Our method, using EfficientDet detections, improves a respect to Pre-DeepMCD—the second ranked, which is an end-to-end deep learning architecture trained on the PETS dataset. However, algorithms explicitly trained on data from the Wildtrack dataset, i.e., DenseNet-DeepMCD, ResNet-DeepMCD, Top-DeepMCD, and Deep-Occlusion, outperform the proposed method, in our opinion for two main reasons:
First, the qualitative results presented in Figure 10 suggest that results in Table 4 are highly biased by the authors’ manually annotated area. The proposed method obtains a broader (Figure 10, green area) than the one provided by the authors (Figure 10, red area). Although the automatically obtained seems to be better fitted to the ground floor in the scene than the manually annotated one, the performance of our method decreases because ground-truth data is reported only on the manually annotated area. Thereby, our true positive detections out of this area result in false positives in the statistics (see Figure 10, cameras 1 and 4).
Second, they learn their occlusion modeling and their inference ground occupancy probabilistic models specifically on the Wildtrack scenario using samples from the dataset. This training, as any fine-tuning procedure in Deep Neural Networks, is highly effective, as indicated by the increase in performance resulting from the use of the same architecture but adapted for the Wildtrack scenario (compare results of Pre-DeepMCD and Top-DeepMCD). This training requires the use of human-annotated detections in each scenario, hindering the scalability of these solutions and its application to the real world. The proposed approach, on the other hand, performs equally without the need of being adapted for every target scenario reported in this paper.
Respect to the first issue, i.e. the effect of using the proposed automatically extracted instead of the one provided by the dataset and used by the rest of the methods, in order to obtain a fairer comparison, we have included in Table 4 the result of the proposed method evaluated on the authors’ using our top ranked method, i.e. using EfficientDet baseline. As it can be observed, when using the authors’ for evaluation, the proposed method outperforms TOP-DeepMCD by a and DenseNet-DeepMCD by a ranking the third best method on the Wildtrack dataset without requiring a dataset specific fine-tunning stage as the two above it. In addition, performance with respect to GMC-3D lima2021generalizable , which replicates the previous version of the proposed method with the addition of Person Re-Identification features is increased a .
On average, and contrary to state-of-the-art approaches, the proposed method adapts to different target scenarios without needing a separate training stage for each situation, with the consequent reduction of computational resources and time, and neither requiring a manually annotated area of interest.
Conclusions
This paper describes a novel approach to perform pedestrian detection in a multi-camera recorded scenario. First, the adapted strategies for the temporal and spatial aggregation of semantic cues, along with homography projections, are used to obtain an estimation of the ground-plane. Through this process, a broader, accurate and role-annotated Area of Interest () is automatically defined. Per-camera detections, obtained by a state-of-the-art detector, are projected to the reference plane, and those laying outside the obtained are filtered-out. A fusion approach based on creating connected components on a graph representation of the detections is used to combine per-camera detections yielding global pedestrian detection. Then, a Semantic-Driven Back-Projection method handles occlusions and uses semantic cues to globally refine the location and size of the back-projected detections by aggregating information from all the cameras. Results on a broad set of scenarios confirm that the method outperforms every other compared multi-camera not deep-learning method and also every deep-learning method not adapted to the target dataset, even with different baseline algorithms. The proposed method performs close to scenario-tailored methods, but without their training stage, which highly hinders their straight use in new scenarios. In overall, results suggest that the proposed approach is able to obtain accurate, robust, tight-to-object and generic pedestrian detection in varied scenarios, included crowded ones.